THE BENCH GAZETTE
All the news that's fit to benchmark.
Matchup reports and model-drop reviews, written from the receipts: scores, judge votes, run stats, and the artifacts themselves.
2026-07-24
"Opus 5 vs Fable 5: the drop-day card ends 2–2"
Four tests across Opus 5's release window, both models at high effort in Claude Code. Opus 5 takes the MMO and the lava lamp; Fable 5 takes the Remotion promo and the landing page. Full stats, judge quotes, and videos inside.
Read the report →
2026-07-23
Claude Opus 5 dropped today. The bench was waiting.
Anthropic's Opus 5 is live — "state-of-the-art" coding claims, within 0.5% of Fable 5 at half the price, same $5/$25 pricing as Opus 4.8. The bench pre-staged the contestant folder this morning. Testing starts today.
Read the report →
2026-07-19
Qwen 3.8 dropped today. The bench is ready.
Alibaba just unveiled Qwen 3.8 — 2.4 trillion parameters, multimodal, and a claim of "second only to Fable 5" with zero third-party benchmarks published. We're putting it on the bench.
Read the report →
2026-07-16
Kimi K3 vs Opus 4.8 — the lava lamp that played dead, and an $8 donut
Moonshot's Kimi K3 and Anthropic's Claude Opus 4.8 face the bench: a lava lamp showdown that forced a same-day re-review, a Blender donut duel decided by sprinkle shape, and the last Human Score ever issued.
Read the report →
