THE BENCH GAZETTE · 2026-07-24
"Opus 5 vs Fable 5: the drop-day card ends 2–2"
Anthropic shipped Opus 5 claiming it lands "within 0.5% of Fable 5 at half the price." The bench's answer: it depends which half of the job you're pricing.
The card: three brand-new prompts, run as exhibitions (new prompts debut off the 8-test grid — Ruling #7 — but everything publishes in full). Both contestants ran in Claude Code at high effort, one shot per test, fresh session per test, identical prompts fired back-to-back. No human scores — the verdicts below are the AI Arena panel's (three vision judges, no Anthropic models on the panel), and the People's Vote is open on every match.
Test 1 — Build a fantasy MMO (Three.js)
A classic early-2000s-MMO-spirit game, original assets only, single HTML file.
| Opus 5 | Fable 5 | |
|---|---|---|
| Time | 25m 36s | 8m 41s |
| Output tokens | 137.6k | 42.7k |
| API-equiv cost | ~$3.44 | ~$2.13 |
Arena: Opus 5, 3–0. The judges saw the same identically-scripted gameplay session run against both games. Gemini 3.1 Pro: "a significantly more polished and authentic MMO-style UI, including a highly detailed HUD with a minimap and action bar. The lighting and low-poly aesthetic are also more atmospheric." Grok 4.5 called out the denser terrain, rain particles, and day/night progression. Opus built EMBERHOLLOW; Fable built EMBERVALE — watch the gameplay side-by-side.
Test 2 — Produce a Remotion promo for Beckon
Scaffold a Remotion project, rebuild the heybeckon.ai UI as animated React components (no screenshot embeds), render a 30s 1080p MP4 unattended.
| Opus 5 | Fable 5 | |
|---|---|---|
| Time | 22m 41s | 10m 42s |
| Output tokens | 69.2k | 42.2k |
| API-equiv cost | ~$1.73 | ~$2.11 |
Arena: Fable 5, 3–0. Gemini: "a much higher level of polish and detail in recreating the UI components… missing elements like the navbar and buttons in the hero scene, and significantly better-designed isometric graphics and typography." Producer's note for the record: Opus lost real clock time to a genuine environment landmine — Remotion's headless-Chrome download silently corrupts on this rig, and Opus diagnosed it and vendored a working browser shell itself. Impressive agentic save; the clock ran anyway. Both rendered MP4s are on the test page.
Test 3 — Recreate a live landing page
Four reference screenshots of heybeckon.ai; recreate it as one self-contained HTML file. No external anything.
| Opus 5 | Fable 5 | |
|---|---|---|
| Time | 30m 35s | 4m 21s |
| Output tokens | 96.7k | 21.2k |
| API-equiv cost | ~$2.42 | ~$1.06 |
Arena: Fable 5, 3–0. The tale of two temperaments: Fable one-shotted its page in four minutes and stopped. Opus wrote its page in ten, then spent twenty more in a screenshot-measure-patch loop, checking itself against the references pixel by pixel. The judges preferred Fable's fuller lower sections and typography hierarchy anyway. Qwen3-VL: "more accurately replicates the reference's layout, typography hierarchy, and section transitions." Both pages are live artifacts — click through and scroll them yourself.
Test 4 — The lava lamp (frozen test 02, scored)
Added the next day by popular demand — the lava lamp is the bench's classic: season one's Sol Ultra beat Fable 5 Max 3–0 here.
| Opus 5 | Fable 5 | |
|---|---|---|
| Time | 43m 54s | 15m 36s |
| Output tokens | 185.7k | 62.1k |
| API-equiv cost | ~$4.64 | ~$3.10 |
Arena: Opus 5, 3–0. The temperaments repeated — Fable built a physics engine, unit-tested it headlessly in Node, verified 29 full rise-and-sink cycles, and shipped in a quarter hour; Opus spent three quarters of an hour on a WebGL metaball renderer with glass refraction, four colourways, drag-to-stir, and a Canvas2D fallback — and this time the judges rewarded it. Gemini: "stunning lighting, refraction, and fluid soft-body physics that perfectly captures the mesmerizing feel of a real lava lamp." Both lamps are live on the test page — leave them running a minute each.
The scoreline
The card ends 2–2. On the ELO ladder the two sit dead level at 1000 apiece (2W–2L each) — and Fable 5's max-effort season-one run still holds #1 at 1065.
The pattern across all four tests is remarkably consistent: Opus 5 works roughly 3× longer and spends 2–4× the tokens; Fable 5 ships first. Where the extra time buys visible depth — a richer game world, a lusher render — Opus wins; where the job is matching a reference faithfully, Fable's fast, complete pass wins. At $25/M out versus Fable's $50/M, Opus's per-test API-equiv bill still came out higher three times out of four, because it wrote that many more tokens. "Half the price" is a rate, not a bill.
Also published today: the season-one archive match on this same landing-page prompt — Fable 5 Max over Sol Ultra, 2–1. It's in the arena log with everything else.
Receipts, as always: every prompt is published verbatim on the test pages, every artifact is embedded live with a checksum recorded at capture, every judge ballot is public, and the People's Vote is open. Stats are recorded, never scored.
