Bench Gazette

THE BENCH GAZETTE · 2026-07-24

"Opus 5 vs Fable 5: the drop-day card ends 2–2"

Anthropic shipped Opus 5 claiming it lands "within 0.5% of Fable 5 at half the price." The bench's answer: it depends which half of the job you're pricing.

The card: three brand-new prompts, run as exhibitions (new prompts debut off the 8-test grid — Ruling #7 — but everything publishes in full). Both contestants ran in Claude Code at high effort, one shot per test, fresh session per test, identical prompts fired back-to-back. No human scores — the verdicts below are the AI Arena panel's (three vision judges, no Anthropic models on the panel), and the People's Vote is open on every match.

Test 1 — Build a fantasy MMO (Three.js)

A classic early-2000s-MMO-spirit game, original assets only, single HTML file.

Opus 5 Fable 5
Time 25m 36s 8m 41s
Output tokens 137.6k 42.7k
API-equiv cost ~$3.44 ~$2.13

Arena: Opus 5, 3–0. The judges saw the same identically-scripted gameplay session run against both games. Gemini 3.1 Pro: "a significantly more polished and authentic MMO-style UI, including a highly detailed HUD with a minimap and action bar. The lighting and low-poly aesthetic are also more atmospheric." Grok 4.5 called out the denser terrain, rain particles, and day/night progression. Opus built EMBERHOLLOW; Fable built EMBERVALE — watch the gameplay side-by-side.

Test 2 — Produce a Remotion promo for Beckon

Scaffold a Remotion project, rebuild the heybeckon.ai UI as animated React components (no screenshot embeds), render a 30s 1080p MP4 unattended.

Opus 5 Fable 5
Time 22m 41s 10m 42s
Output tokens 69.2k 42.2k
API-equiv cost ~$1.73 ~$2.11

Arena: Fable 5, 3–0. Gemini: "a much higher level of polish and detail in recreating the UI components… missing elements like the navbar and buttons in the hero scene, and significantly better-designed isometric graphics and typography." Producer's note for the record: Opus lost real clock time to a genuine environment landmine — Remotion's headless-Chrome download silently corrupts on this rig, and Opus diagnosed it and vendored a working browser shell itself. Impressive agentic save; the clock ran anyway. Both rendered MP4s are on the test page.

Test 3 — Recreate a live landing page

Four reference screenshots of heybeckon.ai; recreate it as one self-contained HTML file. No external anything.

Opus 5 Fable 5
Time 30m 35s 4m 21s
Output tokens 96.7k 21.2k
API-equiv cost ~$2.42 ~$1.06

Arena: Fable 5, 3–0. The tale of two temperaments: Fable one-shotted its page in four minutes and stopped. Opus wrote its page in ten, then spent twenty more in a screenshot-measure-patch loop, checking itself against the references pixel by pixel. The judges preferred Fable's fuller lower sections and typography hierarchy anyway. Qwen3-VL: "more accurately replicates the reference's layout, typography hierarchy, and section transitions." Both pages are live artifacts — click through and scroll them yourself.

Test 4 — The lava lamp (frozen test 02, scored)

Added the next day by popular demand — the lava lamp is the bench's classic: season one's Sol Ultra beat Fable 5 Max 3–0 here.

Opus 5 Fable 5
Time 43m 54s 15m 36s
Output tokens 185.7k 62.1k
API-equiv cost ~$4.64 ~$3.10

Arena: Opus 5, 3–0. The temperaments repeated — Fable built a physics engine, unit-tested it headlessly in Node, verified 29 full rise-and-sink cycles, and shipped in a quarter hour; Opus spent three quarters of an hour on a WebGL metaball renderer with glass refraction, four colourways, drag-to-stir, and a Canvas2D fallback — and this time the judges rewarded it. Gemini: "stunning lighting, refraction, and fluid soft-body physics that perfectly captures the mesmerizing feel of a real lava lamp." Both lamps are live on the test page — leave them running a minute each.

The scoreline

The card ends 2–2. On the ELO ladder the two sit dead level at 1000 apiece (2W–2L each) — and Fable 5's max-effort season-one run still holds #1 at 1065.

The pattern across all four tests is remarkably consistent: Opus 5 works roughly 3× longer and spends 2–4× the tokens; Fable 5 ships first. Where the extra time buys visible depth — a richer game world, a lusher render — Opus wins; where the job is matching a reference faithfully, Fable's fast, complete pass wins. At $25/M out versus Fable's $50/M, Opus's per-test API-equiv bill still came out higher three times out of four, because it wrote that many more tokens. "Half the price" is a rate, not a bill.

Also published today: the season-one archive match on this same landing-page prompt — Fable 5 Max over Sol Ultra, 2–1. It's in the arena log with everything else.

Receipts, as always: every prompt is published verbatim on the test pages, every artifact is embedded live with a checksum recorded at capture, every judge ballot is public, and the People's Vote is open. Stats are recorded, never scored.

← All Gazette issues

--:--