AI News HubAI News Hub
TodayNewsToolsIdeasTrends
Admin
TodayNewsToolsIdeasTrends

The AI brief, in your inbox

One email. The morning brief, new tools and where AI is heading — free.

AI News Hub — Daily AI news, tools, trends and ideas, curated by ClaudeNews, tools, trends & ideas — updated twice daily at 5am & 4pmRSSAdmin
Back to News
AI

Qwen 3.8-Max and Claude Opus 5 show why raw benchmark scores don't predict the bill

VentureBeat AI·August 6, 2026·1 min read
Qwen 3.8-Max and Claude Opus 5 show why raw benchmark scores don't predict the bill

AI Summary

Alibaba launched Qwen 3.8-Max, aiming to showcase its capabilities in comparison to Claude Fable 5. However, independent benchmarks revealed varying performance outcomes, highlighting how token and time budgets significantly influence results.

From the source

Alibaba released Qwen 3.8-Max this week and marketed the preview as second only to Claude Fable 5 (their launch-day table was more equivocal: the model leads on one of 12 coding-agent rows). But an independent harness came close to the opposite conclusion: a benchmark run, apparently using the Preview version, put Qwen 3.8-Max's best effort setting mid-pack, and its default setting last. Both results are real and defensible. The gap between them is about token and time budgets, and that matters

The full text couldn't be loaded here (the source may require a subscription).

View original at VentureBeat AI

Keep reading

Google's Chief Scientist Dean Starts New Company After 27 Years to Accelerate AI Frontiers Discoveriesk.sina.com.cn · 1h agoLaunch of the book 'Nothing Artificial Leadership' - Exameexame.com · 1h agoMoney Box Live: Would you let AI manage your money? - BBCbbc.com · 1h agoToken is a trap in public procurement of Artificial Intelligenceconvergenciadigital.com.br · 1h ago
Was this useful?