Claude Opus 5 Tops AI Benchmark: Navigating the Performance-Cost Trade-off

In the latest evaluation from independent testing firm Artificial Analysis, Claude Opus 5 scored 61 points on a composite intelligence index covering nine tests, edging out Fable 5 by a single point. GPT-5.6 Sol followed with 59 points, while Kimi K3 scored 57. The competition at the top remains exceptionally close.

Substantial Cost Advantage: 26% Savings Per Task

Beyond raw performance, Opus 5 stands out for its efficiency. The model's average cost per task is $2.03, which is 26% lower than Fable 5's $2.75. For organizations scaling AI deployments, this cost differential translates to significant operational savings.

Specialized Strengths: Leading in Knowledge Work and Programming

Drilling into specific capabilities, Opus 5 secured top positions in both the GDPval-AA v2 and AA-Briefcase assessments for knowledge work. When paired with Claude Code, it also tied for first place in the programming Agent index. Its score of 89% on Terminal-Bench v2.1 roughly matched that of GPT-5.6 Sol.

Adjustable Reasoning Tiers: User-Controlled Performance Scaling

The model introduces five configurable reasoning tiers, from low to maximum. Testing reveals an approximately 8x difference in output tokens between the lowest and highest settings, with a corresponding performance gap of 407 Elo points on the GDPval-AA v2 test. This architecture allows users to actively balance performance needs against cost constraints.

Persistent Challenges: Factual Gaps and Hallucination Rates

Notable weaknesses remain. Opus 5 still trails Fable 5 in factual knowledge accuracy. More concerning is its hallucination rate, which climbed to 50% in the AA-Omniscience test at higher reasoning tiers—a 14-percentage-point increase over its predecessor, Opus 4.8. Additionally, its cost-effectiveness at lower reasoning settings slightly lags behind the GPT-5.6 series.

The evaluation underscores a key dynamic in today's AI landscape: leading models are navigating delicate trade-offs between capability, cost, and reliability. The optimal choice depends heavily on a user's specific priorities for accuracy, budget, and task type.