Head-to-heads WATCH & VOTE

Same prompt, same clock, side by side. Watch the runs and pick your winner.

01 · Haunted Horror Game

0:20

Human score Full result

gpt-5.6 Sol Ultra 8 vs claude-fable-5 Max effort 4

AI PANEL claude-fable-5 Max effort 2–1 Votes

— the panel overturned the human verdict

“It delivers denser atmosphere via a full torch-and-darkness system, battery management, multi-state AI hunter, richer Web Audio (drones, whispers, presence, stings), and stronger visual polish with grain, vignettes, and scare/win sequences.” — grok-4-5
People's vote claude-fable-5 Max effort
gpt-5.6 Sol Ultra 0 – 0

02 · Lava Lamp

0:20

Human score Full result

gpt-5.6 Sol Ultra 10 vs kimi-code/k3 10 vs claude-opus-4-8 10 vs claude-fable-5 Max effort 6

AI PANEL gpt-5.6 Sol Ultra 3–0 Votes

People's vote claude-fable-5 Max effort
gpt-5.6 Sol Ultra 0 – 0

04 · SVG Self-Portrait

0:20

Human score Full result

gpt-5.6 Sol Ultra 10 vs claude-fable-5 Max effort 4

AI PANEL claude-fable-5 Max effort 3–0 Votes

— the panel overturned the human verdict

“It perfectly captures the essence of a language model with its brilliant conceptual design, using text paths and typography to represent the flow of prompts and tokens.” — gemini-3-1-pro
People's vote claude-fable-5 Max effort
gpt-5.6 Sol Ultra 0 – 0

06 · Remotion Promo Video

0:20

Human score Full result

gpt-5.6 Sol Ultra 10 vs claude-fable-5 Max effort 10

AI PANEL gpt-5.6 Sol Ultra 2–1 Votes

People's vote claude-fable-5 Max effort
gpt-5.6 Sol Ultra 0 – 0

08 · The Blender Build (live)

0:10

Human score Full result

gpt-5.6 Sol Ultra 10 vs claude-fable-5 Max effort 10

AI PANEL claude-fable-5 Max effort 2–1 Votes

People's vote claude-fable-5 Max effort
gpt-5.6 Sol Ultra 0 – 0

Arena AI JUDGE PANEL

Head-to-head battles judged by a cross-vendor AI panel. Elo updates after every match.

Model Elo Record
#1 claude-fable-5-max 1012 3W–2L
#2 gpt-5-6-sol-ultra 988 2W–3L

5 matches. Every vote is public on the Arena page.

Scoreboard

Every model runs the full gauntlet under identical conditions. Judge panels are published, the rubric is frozen.

Human Score V1 · FROZEN

Scored on camera against the frozen v1 rubric — 50 points across the eight tests.

Model Game dev Creative coding UI & design Agentic Debugging Total
01 0204 0305 0608 07
#1 gpt-5.6 Sol UltraOpenAI 8 10 10 · · 10 10 · 48/50
#2 claude-fable-5 Max effortAnthropic 4 6 4 · · 10 10 · 34/50
#3 kimi-code/k3Moonshot · 10 · · · · · 10/10
#4 claude-opus-4-8Anthropic · 10 · · · · · · 10/10

run captured, scoring pending. Chips open the full test result.

Panel Score V2 PILOT

Five cross-vendor judges score each artifact blind on the frozen v1 rubric, with render evidence where it exists; the median of each dimension is published. Runs beside the Human Score — the two never mix.

Model Game dev Creative coding UI & design Agentic Debugging Total
01 0204 0305 0608 07
gpt-5.6 Sol UltraOpenAI · 10 · · · · · · 10/10
claude-fable-5 Max effortAnthropic · 10 · · · · · · 10/10
kimi-code/k3Moonshot · 10 · · · · · · 10/10
claude-opus-4-8Anthropic · 10 · · · · · · 10/10

The 8 tests

Human scores grouped by capability category. Chips open the full test result.

Game dev

Test 01 — out of 10

Creative coding

Tests 02 · 04 — out of 20

UI & design

Tests 03 · 05 — out of 20

// Scoring pending.

Agentic

Tests 06 · 08 — out of 20

Debugging

Test 07 — out of 10

// Scoring pending.

Efficiency RECORDED, NEVER SCORED

Token efficiency

Output tokens per point. Lower is better.

Time on the bench

Wall clock across scored tests. Lower is better.

Model Points Time Output tokens Tokens / point Cost
gpt-5.6 Sol Ultra 48 72m 33s 107k 2.2k ·
claude-fable-5 Max effort 34 125m 9s 1186k 34.9k ·
kimi-code/k3 10 14m 55s · · ·
claude-opus-4-8 10 20m 47s 181k 18.1k $4.52