THE BENCH GAZETTE · 2026-07-19
Qwen 3.8 dropped today. The bench is ready.
Alibaba's Qwen team unveiled Qwen 3.8 today: a 2.4-trillion-parameter multimodal flagship — their first trillion-plus multimodal model — with a preview live now and open weights promised "soon."
The claim that matters: Alibaba says it's "one of the most powerful models available today… second only to Fable 5." The catch that matters more: no independent benchmarks exist yet. The claim rests entirely on internal evals.
That's what this bench is for.
What we know right now
- 2.4T parameters, multimodal (images, video, documents) — up from ~1T in Qwen 3.7-Max, landing in the same scale bracket as Kimi K3's 2.8T.
- Alibaba says it should beat Qwen 3.7-Max "especially in coding and complex productivity tasks" — full-stack development named specifically. Bold words. Testable words.
- Available today through Alibaba's Token Plan, Qoder, and QoderWork, at 10% of standard pricing during preview. Not on OpenRouter as of this morning; open weights not yet published.
- The subplot: Alibaba owns roughly a third of Moonshot AI — so Qwen 3.8 vs Kimi K3 is as much an in-house rivalry as a market one.
What happens next
Qwen 3.8 gets the full gauntlet: eight one-shot tests, identical conditions, run in Alibaba's own flagship coding tool per the bench's native-harness rule. Every prompt verbatim, every artifact playable, every judge vote published — and your vote counts.
Results land on the leaderboard as they happen. If the "second only to Fable 5" claim survives contact with a lava lamp, you'll see it here first.
Sources: Alibaba Qwen announcement; Bloomberg; The Decoder. Written on drop day — this post will be updated with results.
