Qwen3.8 Max Ties Top Model in Benchmark, But at What Cost?
Evaluation firm Artificial Analysis has released updated scores for Qwen3.8 Max. After retesting using the official API, the model's Intelligence Index climbed from 53 to 56, matching the latest result for Claude Opus 4.8. It still trails the leading Kimi K3 model by a single point.
Agent Capabilities Shine as Key Differentiator
The score improvement is largely driven by a leap in Agent task performance. On the GDPval-AA benchmark, Qwen3.8 Max scored 1739, surpassing both Kimi K3 (1685) and GPT-5.6 Sol (1730). Currently, only Claude Opus 5 scores higher in this category.
The model also showed widespread gains in areas like terminal operations, scientific reasoning, and coding evaluations.
The Trade-offs: Rising Costs and Reliability Concerns
The enhanced capabilities come with clear trade-offs. Although the per-token price for Qwen3.8 Max is lower, the average cost to complete a comprehensive evaluation task surged from $0.53 to $1.14, higher than Kimi K3's $0.86.
This cost increase is primarily due to the model engaging in more internal reasoning steps and generating longer outputs for complex tasks.
Perhaps more concerning is a shift in knowledge reliability. The model's hallucination rate—how often it confidently states incorrect information—jumped dramatically from 23% in the previous generation to 40%, even as its baseline knowledge accuracy remained steady.
Implications for Developers and Users
The retest paints a nuanced picture:
- Enhanced Power: The model is highly competitive for complex, multi-step Agent tasks.
- Operational Cost: Users must budget for significantly higher inference costs.
- Accuracy Risk:The doubled hallucination rate necessitates caution in fact-critical applications.
For businesses and developers considering integration, the decision now involves balancing task performance against budget constraints and tolerance for potential inaccuracies.