Claude Opus 5 Tops AI Benchmark: Navigating the Performance-Cost Trade-off
In the latest evaluation from independent testing firm Artificial Analysis, Claude Opus 5 scored 61 points on a composite intelligence index covering nine tests, edging out Fable 5 by a single point. GPT-5.6 Sol followed with 59 points, while Kimi K3 scored 57. The competition at the top remains exceptionally close.
Substantial Cost Advantage: 26% Savings Per Task
Beyond raw performance, Opus 5 stands out for its efficiency. The model's average cost per task is $2.03, which is 26% lower than Fable 5's $2.75. For organizations scaling AI deployments, this cost differential translates to significant operational savings.
Specialized Strengths: Leading in Knowledge Work and Programming
Drilling into specific capabilities, Opus 5 secured top positions in both the GDPval-AA v2 and AA-Briefcase assessments for knowledge work. When paired with Claude Code, it also tied for first place in the programming Agent index. Its score of 89% on Terminal-Bench v2.1 roughly matched that of GPT-5.6 Sol.
Adjustable Reasoning Tiers: User-Controlled Performance Scaling
The model introduces five configurable reasoning tiers, from low to maximum. Testing reveals an approximately 8x difference in output tokens between the lowest and highest settings, with a corresponding performance gap of 407 Elo points on the GDPval-AA v2 test. This architecture allows users to actively balance performance needs against cost constraints.
Persistent Challenges: Factual Gaps and Hallucination Rates
Notable weaknesses remain. Opus 5 still trails Fable 5 in factual knowledge accuracy. More concerning is its hallucination rate, which climbed to 50% in the AA-Omniscience test at higher reasoning tiers—a 14-percentage-point increase over its predecessor, Opus 4.8. Additionally, its cost-effectiveness at lower reasoning settings slightly lags behind the GPT-5.6 series.
The evaluation underscores a key dynamic in today's AI landscape: leading models are navigating delicate trade-offs between capability, cost, and reliability. The optimal choice depends heavily on a user's specific priorities for accuracy, budget, and task type.