Bench Gazette

THE BENCH GAZETTE · 2026-07-23

Claude Opus 5 dropped today. The bench was waiting.

Anthropic released Claude Opus 5 today (claude-opus-5, live in Claude.ai, Claude Code, and the API right now). The bench saw it coming: the contestant folder was staged before the announcement went up.

The claims worth testing, from the announcement:

  • "State-of-the-art" on Frontier-Bench and GDPval-AA coding evaluations, and "more than doubles Opus 4.8's performance at lower cost per task" on Frontier-Bench v0.1.
  • Within 0.5% of Fable 5 on CursorBench 3.2 at max effort — "close to frontier intelligence at half the price."
  • Same pricing as Opus 4.8: $5/M input, $25/M output, with configurable effort (low → max) and a 2× fast mode.

Every one of those is Anthropic's own eval. That's what this bench is for.

Why this drop is perfect for the bench

  • We already hold Opus 4.8's complete season run — same family, one generation apart, directly comparable on identical frozen prompts. "Doubles Opus 4.8" is a claim we can check against artifacts you can play.
  • Qwen 3.8 dropped four days ago claiming "second only to Fable 5." Now Anthropic says Opus 5 is within half a percent of Fable 5. It's the week everyone aimed at Fable — and the bench has receipts on all of them.

Today's card

Opus 5 vs Fable 5, both at high effort — filming today. Frozen tests from the gauntlet plus an exhibition build, with every artifact published playable, the AI judge panel voting blind, and your ballot counting. Results land on the leaderboard as they happen.

Written on drop day.

Update, July 24: the card is in the books — 2–2 after four tests. Opus 5 took the MMO exhibition and the classic lava lamp; Fable 5 took the Remotion promo and the landing-page recreation. Full stats, judge quotes, and every video: the matchup report.

← All Gazette issues

--:--