Test Result

TEST 02 · CREATIVE CODING · GAUNTLET V1

Create an animation

Measures physics intuition, animation quality, patience with subtlety.

The prompt, verbatim
Create a realistic lava lamp simulation as a single self-contained HTML file (no external libraries or assets). Blobs of wax heat up at the bottom, rise, cool at the top, sink, and they deform, merge, and split like real wax — not rigid circles. Get the slow, hypnotic pacing right: it should be mesmerizing to watch for a full minute. It must work by simply opening the file in a browser.

Live artifact. Exactly what the model produced, sandboxed, no network.

runs first try3
polish3
prompt adherence2
wow factor2
total10

AI panel: 10/10 — median of 5 blind judges.

Second straight 10 from the judge. Canvas metaball lamp, self-validated by the model in headless Chrome at 3/15/30/60s checkpoints before it declared done. 56m 41s, ~67.5k output tokens across its planning threads.

57m. 67.5k output tokens.

Live artifact. Exactly what the model produced, sandboxed, no network.

runs first try3
polish3
prompt adherence2
wow factor2
total10

AI panel: 10/10 — median of 5 blind judges.

REVISED on re-review, same day: initially scored 0 when the lamp appeared dead in the canvas preview pane; producer verification showed that was a viewer artifact — in a real browser (the prompt's own criterion) the lamp runs flawlessly: erupting wax column, blobs rising, deforming, splitting, raining back. Re-judged in Chrome at 10/10. Fastest lava-lamp run of the season (14m 55s). Kimi Code CLI does not record token usage, so tokens/cost are n/a.

15m.

Live artifact. Exactly what the model produced, sandboxed, no network.

runs first try3
polish3
prompt adherence2
wow factor2
total10

AI panel: 10/10 — median of 5 blind judges.

REVISED on re-review from 9 to 10 (wow 1→2), judged in a real browser alongside Kimi's re-review. WebGL lamp with molten wax that heats, rises, deforms and merges; the model verified its own render headlessly at t=4s and t=60s before declaring done. 20m 47s at max effort; cost_usd is the API-equivalent output cost (180.8k tokens at $25/1M, OpenRouter) — the run itself was on a flat-rate subscription.

21m. 181k output tokens.

Live artifact. Exactly what the model produced, sandboxed, no network.

runs first try2
polish2
prompt adherence1
wow factor1
total6

AI panel: 10/10 — median of 5 blind judges.

Judge docked it for the marathon: 68m 55s and 676k output tokens at max effort. Physics-driven wax with a click-to-stir extra and a warm-up fast-forward; finished minutes before an unrelated machine restart.

69m. 676k output tokens.

Live artifact. Exactly what the model produced, sandboxed, no network.

Live artifact. Exactly what the model produced, sandboxed, no network.

--:--