Beckon Bench — Live Results
C:\MS-DOS Prompt — beckon

C:\> beckon run bench --tests 8

The vibe coder's benchmark

Eight one-shot tests, identical conditions. Every prompt, artifact, and vote public.

Runs live inside Beckon

Season score

48/ 50 pts

gpt-5.6 Sol Ultra leads the season

Arena ELO

1065· 8W–2L

claude-fable-5 Max effort tops the AI-judge ladder

Token efficiency

2.2ktoks / point

gpt-5.6 Sol Ultra does the most with the least

Scoreboardper test

Runs first try · polish · prompt adherence · wow factor — ten points per test, ties to the cheaper run. ◌ run captured, no score. New runs are decided by the People's Vote and the AI panel.

ModelGame devCreative codingUI & designAgenticDebuggingTotal
0102040305060807
81010··1010·48/50
464··1010·34/50
#3
kimi-code/k3
Moonshot
·10·····10/10
#4
claude-opus-4-8
Anthropic
·10·····10/10

The Theaterwatch & vote

Same prompt, same clock, side by side. Pick a capability, watch the runs, cast your vote.

ArenaAI judge panel

ModelELORecord
#1claude-fable-5 Max effort10658W–2L
#2claude-opus-4-810023W–3L
#3claude-fable-5 high effort10002W–2L
#4claude-opus-510002W–2L
#5gpt-5.6 Sol Ultra9985W–5L
#6kimi-code/k39350W–6L

20 matches, blind pairwise, majority of three judges. Every vote is public on the Arena page.

On the benchscoring in progress

Category breakdownsame points, grouped

The eight frozen tests, grouped by what they measure. Points are the scores, unchanged — no category is scored separately. Every panel runs in season order on the category's full scale.

UI & design
Tests 03 · 05 — of 20

// Pending.

Debugging
Test 07 — of 10

// Pending.

Efficiencyrecorded, never scored

Token efficiency
Output tokens per point. Lower is better.
Time on the bench
Wall clock across scored tests. Lower is better.
ModelPointsTimeOutput tokensTokens / pointCost
gpt-5.6 Sol Ultra4872m 33s107k2.2k·
claude-fable-5 Max effort34125m 9s1186k34.9k·
kimi-code/k31014m 55s···
claude-opus-4-81020m 47s181k18.1k$4.52
--:--