Terminal-Bench 4.0 Launches: A Major Overhaul in AI Coding Evaluation

The Terminal-Bench, a key benchmark for assessing AI coding capabilities, has released its 4.0 version. This update represents a significant recalibration of the core evaluation framework. The development team focused on refining the metrics for resource consumption—time, CPU, and memory usage—during task execution by AI agents, aiming to better mirror real-world development performance.

A Refined Arena: Task Library and Rule Optimizations

To enhance the benchmark's credibility and usefulness, version 4.0 involved a substantial cleanup of its task library. The team corrected 19 tasks with evaluation biases and removed 8 others that suffered from saturation effects, frequent model refusals, publicly available solutions, or inherent quality issues. A universal 8-hour time limit was applied to all tasks, primarily to minimize score distortions caused by timeouts or environmental issues, ensuring the model's genuine coding ability takes center stage.

New Rankings Reveal Shifts in the Competitive Landscape

Under the new evaluation system, the latest leaderboard presents some intriguing developments.

Top Two Hold Firm, Third Place Changes Hands

The combination of Opus5 and ClaudeCode maintains its lead with a 51.8% score, while Fable5 secures second place with 44.5%. The real narrative, however, unfolds in the battle for third.

GLM-5.3's Breakthrough Performance

Paired with ClaudeCode, the GLM-5.3 model achieved a score of 41.8%, propelling it to third place on the leaderboard. This performance surpasses the 37.3% score attained by the GPT-5.6Sol and Codex combination. Notably, GLM-5.3 stands out as the only model in the top three not developed by Anthropic, introducing greater diversity into the upper echelons of the ranking.

From Challenger to Leader: The Evolution of GLM-5.3

A comparison with the previous benchmark highlights GLM-5.3's trajectory. In Terminal-Bench 3.0, the model scored 32.4%, placing fourth and trailing GPT-5.6Sol's 34.6%. With the shift to the 4.0 evaluation criteria, GLM-5.3 not only climbed to third but reversed the deficit, now leading GPT-5.6Sol by 4.5 percentage points. This leap forward signals substantial improvements in its code generation and comprehension capabilities.

The release of Terminal-Bench 4.0 and its updated rankings provides a new yardstick for developers and researchers, indicating that competition among large language models in the coding domain is intensifying and becoming more varied.