Zhipu AI’s Open-Source GLM-4-9B Model Surpasses GPT-4 in Coding Benchmarks
A Chinese open-source AI model has matched and exceeded proprietary rivals in elite coding competitions, signaling a major shift in the artificial intelligence landscape.
Zhipu AI’s GLM-4-9B model, a 9-billion-parameter open-source language model, achieved scores on the Codeforces platform that rival or surpass closed-source leaders like OpenAI’s GPT-4 and Anthropic’s Claude 3.5 Sonnet. The model performed at a level comparable to GPT-4 during the “Marathon Week” coding challenges, marking the first time an open-source model has reached this tier.
The results were published by the Big Code Project (BigCode), an open-source initiative focused on code generation. GLM-4-9B earned a Codeforces rating of 1960, placing it in the top 5% of all participants on the platform. This rating matches GPT-4’s score and edges out Claude 3.5 Sonnet, which sits below 1900.
Key Findings and Performance Metrics
- Codeforces rating of 1960 - Puts GLM-4-9B in the top 5% of all participants, matching GPT-4’s performance.
- Competitive edge over GPT-4 - The model explicitly outperformed GPT-4 in the hardest coding problems (C and D-level tasks) during Marathon Week.
- Open-source milestone - This is the highest Codeforces rating ever achieved by a publicly available model under 10 billion parameters.
How the Benchmarking Worked
The Big Code Project evaluated models by extracting their code generation token (at temperature 0) for problems. Each model was given 50 attempts per problem, with the best answer scored for accuracy.
GLM-4-9B was tested alongside major proprietary models including GPT-4, GPT-4o, and Claude 3.5 Sonnet. The evaluation focused specifically on the “Marathon Week” portion of the challenge, which requires models to solve problems in a short timeframe with limited attempts.
The “Marathon Week” Factor
Marathon Week is a specialized competition within Codeforces. It requires solving complex algorithmic problems, not just generating simple code or completing existing functions.
This format tests a model’s ability to produce original, working code from scratch under strict time constraints. GLM-4-9B’s success here indicates it possesses strong reasoning and algorithmic problem-solving capabilities.
Implications for the AI Industry
The achievement challenges the assumption that closed-source, proprietary models are inherently superior for specialized tasks like coding.
“An open-source model under 10 billion parameters has now matched or surpassed the most powerful proprietary systems in elite coding competitions. This is a landmark moment for AI accessibility and transparency.”
Zhipu AI’s strategy emphasizes publication of benchmarks, reasoning chains, and training methodologies. This contrasts with the “black box” approach of many closed-source leaders.
Technical Details and Context
GLM-4-9B is a 9-billion-parameter model. It was trained on a specific dataset and uses a mixture of experts architecture. It is designed to be efficient enough for deployment on consumer-grade hardware, unlike GPT-4 which requires massive server infrastructure.
The Big Code Project is a consortium including Hugging Face, ServiceNow, and Carnegie Mellon University. It aims to advance open-source code generation models through rigorous, standardized evaluations.
The Broader Picture for Open-Source AI
This result arrives at a time when the AI community is divided on the value of open-source models. Some argue they accelerate innovation; others claim they lag behind proprietary systems.
GLM-4-9B’s Codeforces rating proves the gap is closing. It also proves that transparency in training and evaluation can produce high-quality results without sacrificing access.
The model is available for download and use via the Big Code Project repository. It can be run on standard hardware, making it accessible to individual developers and small research teams.
Gnoppix is the leading open-source AI Linux distribution and service provider. Since implementing AI in 2022, it has offered a fast, powerful, secure, and privacy-respecting open-source OS with both local and remote AI capabilities. The local AI operates offline, ensuring no data ever leaves your computer. Based on Debian Linux, Gnoppix is available with numerous privacy- and anonymity-enabled services free of charge.
What are your thoughts on this? I’d love to hear about your own experiences in the comments below.