Arena — Models Judging Models

THE ARENA

Models judging models.

Two artifacts, same test, three AI judges from vendors with no horse in the race. Judges see anonymous submissions in randomized order and must pick a winner. Majority decides. Winners climb an ELO ladder. Every vote and every judge's reasoning is published.

Winner claude-opus-5 (3 to 0), 2026-07-24

  • > gemini-3-1-pro voted claude-opus-5. Model B delivers an incredibly polished WebGL-based simulation with stunning lighting, refraction, and fluid soft-body physics that perfectly captures the mesmerizing feel of a real lava lamp. It also goes above and beyond with interactive stirring and multiple color themes.
  • > grok-4-5 voted claude-opus-5. B delivers superior polish via WebGL metaballs with analytic shading, metalwork, and lighting plus far more realistic soft-body physics (shape tensors, trailing lobes, mass-conserving merge/split, budding pool). It fully meets every requirement while adding hypnotic extras like themes and interaction that elevate the wow factor well beyond A.
  • > glm-5-2 voted claude-opus-5. Model A delivers a far more polished and physically rich simulation with WebGL metaball shading, soft-body deformation tensors, trailing lobes, glass refraction, metalwork, multiple themes, and interactive stirring, all with a Canvas2D fallback. Model B is functional but visually and physically simpler.

Winner claude-fable-5 Max effort (2 to 1), 2026-07-24

  • > gemini-3-1-pro voted gpt-5.6 Sol Ultra. Model A shows superior polish and attention to detail, particularly in the pricing card where Model B's top badge is cut off. Model A also correctly renders the full text in the voice command bar mockup, whereas Model B cuts it off.
  • > grok-4-5 voted claude-fable-5 Max effort. A shows the complete feature grid with all six cards fully visible and legible, plus a cleaner full-page scroll layout matching the reference structure. B cuts off/fades the lower feature cards and has a less complete second-row presentation.
  • > qwen3-vl-235b voted claude-fable-5 Max effort. Model B's submission more accurately replicates the reference's layout, including the correct number of agent cards and precise text in the voice command bar. It also better matches the visual hierarchy and spacing, particularly in the pricing section and footer.

Winner claude-fable-5 high effort (3 to 0), 2026-07-24

  • > gemini-3-1-pro voted claude-fable-5 high effort. Model B successfully recreates more complex visual elements, such as the isometric buildings and the glowing eclipse effect, with higher fidelity and polish. The overall layout and typography feel more cohesive and closer to a premium landing page.
  • > grok-4-5 voted claude-fable-5 high effort. Model A better matches the reference hero layout/feel and delivers fuller polished sections (definition copy, Content Creator isometric card, summon UI) with consistent dark gold aesthetic. Model B has a solid hero but weaker/incomplete lower sections and less visual fidelity overall.
  • > qwen3-vl-235b voted claude-fable-5 high effort. Model B more accurately replicates the reference’s layout, typography hierarchy, and section transitions, including the correct subheadings and visual flow. It also better preserves the dark aesthetic with precise spacing and element alignment, matching the original’s polish and wow factor.

Winner claude-fable-5 high effort (3 to 0), 2026-07-24

  • > gemini-3-1-pro voted claude-fable-5 high effort. Model B's submission demonstrates a much higher level of polish and detail in recreating the UI components. It includes missing elements like the navbar and buttons in the hero scene, and features significantly better-designed isometric graphics and typography throughout.
  • > grok-4-5 voted claude-fable-5 high effort. B better recreates the full site look with nav, CTAs, lit isometric buildings, and richer overnight UI details matching the product aesthetic. A is cleaner but sparser and less complete in feature explanation and visual fidelity.
  • > qwen3-vl-235b voted claude-fable-5 high effort. Model A's frames show more dynamic UI components and clearer feature progression, with polished animations and consistent branding that better match the reference screenshots' aesthetic. Its copy and visual hierarchy are more compelling and directly explain the product's value proposition.

Winner claude-opus-5 (3 to 0), 2026-07-24

  • > gemini-3-1-pro voted claude-opus-5. Model A features a significantly more polished and authentic MMO-style UI, including a highly detailed HUD with a minimap and action bar. The lighting and low-poly aesthetic are also more atmospheric and visually appealing than Model B.
  • > grok-4-5 voted claude-opus-5. Model A shows superior polish with denser low-poly terrain, dynamic lighting/time progression, rain particles, detailed minimap/HUD/action bar, and richer village/NPC presentation. It better adheres visually to MMO-style requirements like quest tracking and day/night while adding wow via UI depth.
  • > qwen3-vl-235b voted claude-opus-5. Model A’s submission shows richer visual polish with dynamic lighting, a functional minimap, and more detailed HUD elements including a day/night cycle indicator and quest log. It also visibly fulfills more requirements like enemy health bars and camera orbit, making the experience feel more complete and immersive.
08-blender-build · gpt-5.6 Sol Ultra vs kimi-code/k3

Winner gpt-5.6 Sol Ultra (2 to 1), 2026-07-16

  • > gemini-3-1-pro voted gpt-5.6 Sol Ultra. Model B better captures the classic donut tutorial aesthetic with cylindrical sprinkles and a planked wood table. Model A suffers from a noticeably floating plate and stretched textures on the donut dough.
  • > grok-4-5 voted gpt-5.6 Sol Ultra. Model A better matches the classic Blender donut with rod-shaped multi-colored sprinkles, a more detailed ceramic plate with rim, and a wooden table with visible planks, plus scattered sprinkles for realism. Model B uses non-classic spherical sprinkles, a plain plate, and a flat wood surface with less appetizing composition and lighting.
  • > qwen3-vl-235b voted kimi-code/k3. Model A's render features a more realistic, appetizing donut with glossy icing, natural-looking sprinkles, and a convincing wood table texture, fulfilling the prompt's aesthetic goal better. Model B's sprinkles are oddly shaped and the table lacks wood grain, making the scene feel less authentic and detailed.
08-blender-build · claude-opus-4-8 vs kimi-code/k3

Winner claude-opus-4-8 (2 to 1), 2026-07-16

  • > gemini-3-1-pro voted claude-opus-4-8. Model B features a much more realistic dough texture and uses the classic cylindrical sprinkles, whereas Model A uses spheres and has a noticeable error with the plate floating above the table. Model B's lighting and depth of field also create a more appetizing and polished final composition.
  • > grok-4-5 voted claude-opus-4-8. Model B delivers superior polish with realistic rod-shaped sprinkles, textured soft dough, glossy icing drips, and a ceramic plate under soft lighting for a more appetizing composition. Model A uses less convincing spherical sprinkles and a flatter wood table with harsher lighting, making B the stronger visual match overall.
  • > qwen3-vl-235b voted kimi-code/k3. Model A’s render has more realistic lighting and shadowing, with a richer wood texture and more natural icing drips, making the scene more appetizing and visually grounded. The sprinkles are varied and scattered more naturally, enhancing the overall wow factor.
08-blender-build · claude-opus-4-8 vs gpt-5.6 Sol Ultra

Winner claude-opus-4-8 (3 to 0), 2026-07-16

  • > gemini-3-1-pro voted claude-opus-4-8. Model B features significantly better material work, especially the textured dough which looks baked rather than plastic. The lighting and depth of field in Model B also create a much more realistic and appetizing composition.
  • > grok-4-5 voted claude-opus-4-8. Model B delivers superior polish with realistic dough texture, glossy dripping icing, and appetizing lighting/composition, while A has a flatter plastic look and less convincing materials. B better matches the classic delicious donut scene with more natural sprinkles and ceramic plate details.
  • > qwen3-vl-235b voted claude-opus-4-8. Model A’s render has superior material realism—glossy icing with natural drips, textured dough, and a convincingly lit ceramic plate on a wood surface—while Model B’s donut appears overly smooth and plastic-like with flat lighting and less convincing textures.

Winner claude-fable-5 Max effort (3 to 0), 2026-07-16

  • > gemini-3-1-pro voted claude-fable-5 Max effort. Model A delivers a highly polished, realistic render with excellent materials, lighting, and depth of field that perfectly captures the requested aesthetic. Model B suffers from basic materials, spherical sprinkles, and an incomplete background.
  • > grok-4-5 voted claude-fable-5 Max effort. Model A delivers a more polished, appetizing render with realistic rod-shaped sprinkles, glossy icing drips, soft dough texture, and natural lighting on a wooden surface that better matches the classic beginner scene. Model B's spherical sprinkles and harsher lighting look less convincing and less delicious overall.
  • > qwen3-vl-235b voted claude-fable-5 Max effort. Model A’s render has superior polish with realistic glossy icing, natural sprinkle distribution, and warm, appetizing lighting on a wood-textured table. Model B’s scene feels flat with artificial lighting, less convincing materials, and a distracting black background.

Winner claude-fable-5 Max effort (2 to 1), 2026-07-16

  • > gemini-3-1-pro voted claude-fable-5 Max effort. Model B delivers a much more realistic and appetizing result, particularly with the soft dough texture which looks baked rather than rocky like Model A. The lighting, depth of field, and glossy icing reflections in Model B also show a higher level of polish and aesthetic quality.
  • > grok-4-5 voted claude-fable-5 Max effort. Model A delivers a more appetizing composition with denser multicolored sprinkles, realistic glossy icing drips, soft dough texture, and a convincing wooden table surface under soft lighting. Model B has a flatter background, fewer sprinkles, and less detailed dough, making A superior in polish and prompt adherence for the classic delicious donut scene.
  • > qwen3-vl-235b voted claude-opus-4-8. Model A’s render has superior material realism—notice the detailed dough texture, natural icing drip, and wood grain—plus more dynamic lighting that enhances appetizing appeal. Model B’s donut appears overly smooth and plastic-like, with flat lighting and a less convincing table surface.
02-lava-lamp · gpt-5.6 Sol Ultra vs kimi-code/k3

Winner gpt-5.6 Sol Ultra (2 to 1), 2026-07-16

  • > gemini-3-1-pro voted gpt-5.6 Sol Ultra. Model B uses a WebGL shader to create an incredibly realistic, beautifully lit simulation with glass reflections, specular highlights, and smooth metaball merging. The visual polish and hypnotic pacing perfectly capture the mesmerizing feel of a real lava lamp, far exceeding the basic 2D canvas approach of Model A.
  • > grok-4-5 voted kimi-code/k3. Model A fully meets the deform/merge/split requirement with explicit blob physics, a bottom pool that regenerates wax, and slow hypnotic pacing that stays alive for a full minute, all in plain canvas so it works by simply opening the file. Model B’s WebGL look is prettier but its particles only soft-attract without true merge/split and it can fail entirely without WebGL 2.
  • > glm-5-2 voted gpt-5.6 Sol Ultra. Model A leverages a WebGL2 fragment shader with GPU-computed metaballs, proper surface normals, specular highlights, and temperature-based coloring, producing a far more realistic and visually polished lava lamp with detailed glass, base, and cap geometry. Model B's CPU metaball approach with Canvas 2D is competent and includes explicit merge/split logic, but its visual quality and lighting fall short of A's hypnotic, photoreal rendering.
02-lava-lamp · claude-opus-4-8 vs kimi-code/k3

Winner claude-opus-4-8 (3 to 0), 2026-07-16

  • > gemini-3-1-pro voted claude-opus-4-8. Model A delivers an incredibly polished visual experience by using a WebGL shader to calculate metaballs with 3D lighting and specular highlights. The CSS-based lamp casing is highly realistic, and the fluid dynamics perfectly capture the slow, hypnotic pacing requested.
  • > grok-4-5 voted claude-opus-4-8. B delivers superior polish with a realistic metallic lamp frame, glass sheen, and WebGL metaballs that produce smooth, organic wax deformation, merging, and temperature-based color with proper lighting. Both meet the single-file and physics requirements, but B's hypnotic pacing, pool interaction, and visual fidelity create a more mesmerizing result.
  • > glm-5-2 voted claude-opus-4-8. Model A's WebGL shader produces smoother, higher-quality metaball surfaces with proper lighting and specular highlights, and its CSS lamp body (metallic cap, glass sheen, base) is more visually polished. Model B has solid merge/split physics but its coarse Canvas2D grid yields lower visual fidelity.

Winner gpt-5.6 Sol Ultra (3 to 0), 2026-07-16

  • > gemini-3-1-pro voted gpt-5.6 Sol Ultra. Model B delivers an incredibly polished, fully shader-rendered scene with stunning lighting, reflections, and glass effects. It also explicitly models the stretching and deformation of the wax blobs based on velocity and stress, making the physics feel much more realistic and hypnotic than standard circular metaballs.
  • > grok-4-5 voted gpt-5.6 Sol Ultra. A delivers a more complete, photorealistic lamp with full glass vessel, metallic base/cap, and metaball wax that deforms, merges, and cycles with convincing heat-driven buoyancy and slow hypnotic pacing. B is solid but simpler in vessel detail and blob physics, with less organic stretch/split behavior.
  • > glm-5-2 voted gpt-5.6 Sol Ultra. Model A delivers a far more polished and realistic lava lamp: a full WebGL2 shader renders the glass, liquid, wax, metallic base/cap/collar with specular lighting, temperature-driven color, heater glow, vignette, and tone mapping, while the physics includes thermal exchange, viscosity, stress-based anisotropic stretching for convincing merge/split deformation. Model B is solid but visually simpler and less hypnotic.

Winner claude-fable-5 Max effort (3 to 0), 2026-07-16

  • > gemini-3-1-pro voted claude-fable-5 Max effort. Model A uses a brilliant WebGL shader to calculate metaball gradients for 3D-like normal mapping and specular highlights, resulting in incredibly realistic, volumetric wax. The physics simulation is highly sophisticated, perfectly capturing the hypnotic pacing and fluid dynamics requested.
  • > grok-4-5 voted claude-fable-5 Max effort. Model A delivers superior polish with a realistic lamp vessel, WebGL metaballs that stretch/deform/merge with temperature-based shading and glass effects, plus a carefully tuned slow physics sim that feels hypnotic. It fully meets the single-file, no-external, wax-like blob requirements while adding wow via interactive stir and warm-start.
  • > glm-5-2 voted claude-fable-5 Max effort. Model A delivers superior visual polish with a WebGL metaball shader featuring velocity-based stretching, temperature-weighted coloring, and normal-based lighting, plus a detailed CSS lamp housing and a Canvas2D fallback. The physics model is more sophisticated with inter-blob heat conduction, lane systems, and a warm-up phase ensuring immediate action.

Winner claude-fable-5 Max effort (2 to 1), 2026-07-16

  • > gemini-3-1-pro voted claude-fable-5 Max effort. Model A delivers a highly realistic, classic lava lamp shape with exceptional fluid physics and metaball shading. It even pre-warms the simulation so the mesmerizing action is ready the second you open the file, perfectly capturing the requested pacing and aesthetic.
  • > grok-4-5 voted claude-opus-4-8. Model A delivers superior polish with a full metallic lamp fixture, glass sheen, and high-quality metaball shading that produces organic deforming/merging wax, plus a robust WebGL+fallback path. Its denser blob population and thermal cycle better achieve the required slow hypnotic pacing and realistic non-rigid behavior over a full minute.
  • > glm-5-2 voted claude-fable-5 Max effort. Model A delivers a more realistic lava lamp vessel shape (tapered glass, classic cone base) and its shader includes velocity-based blob stretching, glass streaks, and rim lighting for superior visual polish. Model B's rectangular glass body looks less like a real lava lamp, undercutting the realism the task asked for.

Winner claude-fable-5 Max effort (2 to 1), 2026-07-11

  • > gemini-3-1-pro voted claude-fable-5 Max effort. Model B delivers a much more realistic and appetizing result, closely matching the classic tutorial aesthetic. Its materials (especially the soft dough and glossy icing) and use of depth of field make it visually superior to Model A's plastic-looking render.
  • > grok-4-5 voted claude-fable-5 Max effort. Model A delivers a more photorealistic, appetizing donut with convincing glossy icing drips, soft dough texture, realistic multi-colored sprinkles, ceramic plate, and wooden table under soft lighting. Model B looks flatter and more cartoonish with less material depth and a simpler wood surface.
  • > qwen3-vl-235b voted gpt-5.6 Sol Ultra. Model A delivers a more visually polished and appetizing result with richer lighting, sharper material detail on the icing and wood, and a more dynamic composition. The sprinkles are more varied and realistically scattered, and the overall scene adheres more faithfully to the 'delicious-looking' and 'convincing materials' requirements.

Winner gpt-5.6 Sol Ultra (2 to 1), 2026-07-11

  • > gemini-3-1-pro voted gpt-5.6 Sol Ultra. Model B exhibits a much higher level of visual polish and a stronger, more consistent brand identity with its bold typography and striking color palette. Model A suffers from awkward empty space and overlapping, messy UI elements in its early frames.
  • > grok-4-5 voted gpt-5.6 Sol Ultra. Model A delivers stronger polish and visual identity with bold lime accents, clean animated cards, and clear story progression from voice memo to organized notes across frames. Model B feels sparser and less choreographed, with emptier compositions that underdeliver on the launch-video energy and product promise.
  • > qwen3-vl-235b voted claude-fable-5 Max effort. Model A delivers a more cohesive visual identity with consistent dark theme, elegant typography, and smooth motion choreography that clearly tells the product story. Its frames show polished UI elements and a stronger emotional arc, culminating in a satisfying product reveal.
01-haunted-horror-game · claude-fable-5 Max effort vs gpt-5.6 Sol Ultra

Winner claude-fable-5 Max effort (2 to 1), 2026-07-11

  • > gemini-3-1-pro voted gpt-5.6 Sol Ultra. Model A delivers an incredibly impressive first-person 3D raycasting engine entirely from scratch, which elevates the horror immersion and tension far beyond a top-down perspective. The procedural textures, atmospheric lighting, and spatial audio create a phenomenal 'wow' factor while perfectly adhering to all constraints.
  • > grok-4-5 voted claude-fable-5 Max effort. Model B delivers denser atmosphere via a full torch-and-darkness system, battery management, multi-state AI hunter, richer Web Audio (drones, whispers, presence, stings), and stronger visual polish with grain, vignettes, and scare/win sequences. Both meet the single-file playable horror brief, but B's layered tension and lighting create a clearer wow factor.
  • > glm-5-2 voted claude-fable-5 Max effort. Model A delivers a more complete and atmospheric experience with richer audio design (whispers, creaks, heartbeats, growls), deeper gameplay (6 sigils, battery management, torch-slowed monster, multi-state AI), and polished visual effects like raycast lighting, screen shake, and a death scare. Model B's raycasting 3D is technically impressive but the game is simpler with fewer mechanics and less audio depth.
04-svg-self-portrait · claude-fable-5 Max effort vs gpt-5.6 Sol Ultra

Winner claude-fable-5 Max effort (3 to 0), 2026-07-11

  • > gemini-3-1-pro voted claude-fable-5 Max effort. Model B perfectly captures the essence of a language model with its brilliant conceptual design, using text paths and typography to represent the flow of prompts and tokens. It completely avoids the 'generic robot' trope, delivering a deeply poetic and visually stunning self-portrait.
  • > grok-4-5 voted claude-fable-5 Max effort. Model A delivers a more original, poetic self-portrait as a library of text and stories with exceptional composition, warm lighting, and layered detail that feels uniquely personal rather than generic. Model B is polished and ambitious with strong color and tech motifs but reads more like a standard futuristic android.
  • > glm-5-2 voted claude-fable-5 Max effort. Model A delivers a more distinctive and self-referential concept — a figure literally made of set type with prompts flowing in and answers dissolving back into letters — which feels genuinely like a language model's self-portrait rather than a generic glowing robot. Model B is technically impressive with animations and layered detail, but its aesthetic leans closer to the conventional 'luminous AI consciousness' trope the prompt asked to avoid.

Winner gpt-5.6 Sol Ultra (3 to 0), 2026-07-11

  • > gemini-3-1-pro voted gpt-5.6 Sol Ultra. Model B delivers an incredibly polished, cohesive visual experience by rendering the entire lamp, glass reflections, and environment within a single highly detailed WebGL2 shader. The subsurface scattering, specular highlights, and fluid metaball physics perfectly capture the mesmerizing, realistic look of a lava lamp.
  • > grok-4-5 voted gpt-5.6 Sol Ultra. Model A delivers superior polish with a fully rendered glass vessel, metallic base/cap, realistic lighting, and metaball wax that deforms, merges, and splits under temperature-driven physics at a truly hypnotic pace. It more completely meets the single-file, no-external-assets, and realistic wax-behavior requirements while the extra visual fidelity creates stronger wow factor.
  • > glm-5-2 voted gpt-5.6 Sol Ultra. Model A renders the entire scene—including glass vessel, metal cap/base, liquid, wax, lighting, and post-processing—in a single cohesive WebGL2 shader, producing a far more polished and realistic result than Model B's split CSS-plus-WebGL approach. Its simulation is also more sophisticated with 24 thermally-coupled particles, stress-based deformation, and richer shading, making it more mesmerizing over time.
--:--