Radio
Now Playing
Quickyla Radio โ€” Click to play
Open โ†’
3 min left
Back to News

Alibaba's Qwen 3.8-Max underperforms against Anthropic's Claude Opus 5

Alibaba's Qwen 3.8-Max AI model was positioned as a top competitor to Claude Fable 5, but independent evaluations revealed it performed in the middle of the pack due to differing benchmark methodologโ€ฆ

Qwen 3.8-Max and Claude Opus 5 show why raw benchmark scores don't predict the bill
VentureBeat โ€” 6 August 2026
Text:
16 0 0

Alibaba launched its latest AI model, Qwen 3.8-Max, this week, positioning it as a strong competitor to Claude Fable 5. The company claimed that Qwen 3.8-Max was second only to Claude Fable 5 based on their own internal benchmarks. However, an independent evaluation suggests a different narrative, revealing that Qwen 3.8-Max placed in the middle of the pack in performance rankings.

This discrepancy arises from the differing methodologies used to measure the models' performances. Alibaba's benchmarks allowed Qwen 3.8-Max to run for significantly longer periodsโ€”up to 12 hours for some tasks. In contrast, the independent evaluation conducted by VulcanBench capped its time allowance at 60 minutes. The substantial variance in time budgetsโ€”five to sixteen times larger on Alibaba's sideโ€”has led to notably different results, underscoring the importance of context in AI performance metrics.

The implications of these findings are significant for consumers and developers assessing AI models. As the industry matures, it becomes crucial to look beyond raw benchmark scores. The key metric for evaluating these models should shift toward cost per successful task rather than just speed or raw output. This new approach could better reflect a model's overall efficiency and utility in practical applications.

As AI technology continues to evolve, understanding these nuances will be vital for businesses making investment decisions. The ongoing debate about performance measurement emphasizes the need for transparency and consistency in AI benchmarking. Moving forward, stakeholders will need to prioritize comprehensive assessments that consider both time and cost to make informed choices about which AI solutions to implement.

Read Full Story at VentureBeat โ†’
Advertisement
React:
Sources
Sponsored

More to Read

Alonso pleased with Aston Martin upgrade as Newey targets 'โ€ฆ
๐Ÿ’ป Technology
Alonso pleased with Aston Martin upgrade as Newey targets 'respectability'
Sky Sports ยท 13 days ago
Apple announces Siloโ€™s season 4 return date
๐Ÿ’ป Technology
Apple announces Siloโ€™s season 4 return date
9to5Mac ยท 10 days ago
Anthropic upgrades Claude with new Opus 5 model, details heโ€ฆ
๐Ÿ’ป Technology
Anthropic upgrades Claude with new Opus 5 model, details here
9to5Mac ยท 13 days ago
Why Tesla Stock Crashed Today
๐Ÿ“ˆ Markets & Finance
Why Tesla Stock Crashed Today
Nasdaq News ยท 13 days ago
Hereโ€™s the biggest news you missed this weekend
๐ŸŒ World News
Hereโ€™s the biggest news you missed this weekend
NBC News ยท 10 days ago
Cardinals OL Isaiah Adams practices despite his recent arreโ€ฆ
๐Ÿ’ป Technology
Cardinals OL Isaiah Adams practices despite his recent arrest
Yahoo Sports ยท 13 days ago
Full view