AI Model Reviewer

Methodology

How every result on this site was produced and labelled. ingest.py renders this file into the methodology page and the home-page summary, and publishes it in data/content.json (one entry per ## section).

Three tracks

Home harness: each vendor's own CLI (Claude Code for Anthropic, Codex CLI for OpenAI, dsh for DeepSeek, Grok Build for xAI, Gemini CLI for Google, Kimi Code CLI for Moonshot). GLM-5.3 has no coding CLI of its own, so its home track runs in Claude Code, Z.ai's recommended harness, and is labelled that way. Neutral: every model in OpenCode with the same configuration. Raw baseline: an earlier single API request with no harness and no tools. Baselines are kept for reference only; the project's rule is now harness-only. Provider = the API service that served the model for that run. Gemini 3.8 Flash stays in the roster and in the planned totals, but none of its cells is tested: we have no usable access to the model.

Max effort

Every model runs at the highest reasoning effort its API accepts (max for Claude and DeepSeek, xhigh for GPT-6 Astra and Grok 4.7, high for Gemini 3.8 Flash). A few raw baselines ran at lower effort; the effort is printed on every result.

One prompt, fully autonomous

One prompt, then no human input until the model stops. Inside a harness the model may search the web, render and inspect its own output and fix it (self-review allowed). The single exception, where the reviewer sent written feedback once, is labelled 1 fix turn.

Whose failure was it?

DNF means the model failed on its own, for example by spending its whole output budget thinking. When we end a run (our client or safety timeout, or a rule change), it is marked Stopped and not counted against the model. Harness-config failure (ours) means our own harness settings killed the run. Harness + model mismatch means the run is excluded as not a clean single-model result: either the harness and the model could not work together (for example Claude Code sending an image to GLM-5.3's text-only endpoint), or the harness ran part of the job on a different model (a mixed-model run, for example Claude Code spawning a subagent on another model); each run's note says which. Infrastructure DNF means the connection upstream of the harness dropped the run. None of the last three is counted as a model failure; the field notes describe each incident. The spend on every failed run is still counted.

Attempts lost to the harness launcher or the connection before a result existed are not results: they are listed under the results as infrastructure failures (not model results) and left out of every count, total per model and head-to-head. A planned cell that is queued but has not run yet says Pending — queued, not run yet.

Costs: reported vs estimated

Reported: the harness printed its own bill (for example Claude Code's total_cost_usd or OpenCode's session cost). When a harness reported one, it is the headline cost, labelled as that harness's figure, with the list-price estimate beside it. Estimated: token counts times the vendor's list price, with the price table and its date on every result. Unknown: the tokens were never reported, so we don't guess. When no token usage was recorded at all (for example a stream cut before its usage record), the cost reads no usage recorded, never $0. Money spent on runs that produced nothing is shown, never dropped. Totals say what they include: the result totals cover the results on this site, failed runs included; setup smoke tests and infrastructure failures are totalled separately.

15 OpenCode results ran on a box whose per-request proxy logs were lost; the served model for those runs could not be independently verified. Their per-step costs match the requested model's prices.

How results are scored

Scores have three separate sources: the AI judge, visitor votes (coming later, with a blind reveal), and the reviewer's grade. The current AI judge is the test runner, model claude-opus-5-5, described as reviewing the 1600 px render against the published, pre-registered rubric. Each scored run currently has one recorded AI-judge assessment. The reviewer has not given any numeric grades yet. Reviewer grades and visitor votes, when available, appear alongside the AI score and are not blended into it.

Only finished runs covered by a published, pre-registered rubric receive a score. Scores run from 0 to 10: up to 8 points for the prompt's requirements and up to 2 for craft, with half points for partial credit. Cost and time do not contribute to this score. Each AI score links to its rubric and remains labelled Provisional — AI judge until the reviewer checks it. Scores originate with the judge or reviewer; the site's build agent does not assign them.

We publish the per-criterion marks and calculate the total from them. When those marks cannot be read into the published breakdown, we retain the judge's stored total and label it Provisional — stored total; criterion breakdown unavailable. The displayed total has not been verified through that breakdown. A prompt's board stays unranked until every published AI score on it has a complete, validated criterion breakdown. Missing breakdowns and any disagreement between marks and totals must be resolved before ranking that board; affected runs remain visible throughout.

Eligible AI-judge boards are ranked by score within one prompt, highest first. Each row represents a particular run, identified by its model, harness, provider, effort and attempt. Equal totals share a rank, and the next rank skips the occupied positions: 1, 1, 3. Row order inside a tie has no quality meaning. These are provisional ranks of the recorded assessments. They do not establish that a model is generally better, or that a small score difference would persist in another run or under another judge. We do not publish confidence intervals or treat unequal scores as statistical ties without evidence of judging and run variability.

Scores from different prompts are never averaged or used to produce an overall model ranking. Home and neutral harness conditions stay identified, and raw baselines remain a separate reference section. Each board states how many recorded results are scored, how many are unfinished and how many remain unscored. The reviewer's written recommendations and “best” picks remain separately attributed editorial judgments.

DNF is a result. DNF, stopped, failed and no-output results remain visible on the same prompt board, with their recorded status, reason, time and cost. They receive no quality rank and no invented zero score. No output artifact describes the absence of a published deliverable; it does not by itself identify whose failure it was. A stopped run with partial output keeps that output, labelled Work in progress. Finished runs without a score say Not scored, with the reason. Prompts with no scores still have an unranked result board.

Model DNF remains distinct from our stops, harness-config failures, harness/model mismatches, infrastructure failures and unexplained errors. All records published as results stay in result totals, including failures. Separately logged infrastructure attempts that are not result records remain visible as Infrastructure failures (not model results) and outside those totals. Pending and no-access cells remain visible as Pending — queued, not run yet or Not tested (no access); neither is a scored run. Unknown cost is never displayed as zero.

Status labels

  • OKFinished on its own and produced the artifact.
  • DNFDid not finish: the model gave up or used up its own budget (e.g. hit max_tokens while still thinking). Counted against the model.
  • Stopped (our timeout)Our client or safety timeout ended the run. Not a model failure.
  • Stopped by usWe ended the run ourselves (for example after a rule change). Not a model failure.
  • Harness-config failure (ours)Our harness configuration killed the run. Not a model failure and not our safety timeout.
  • Harness + model mismatchExcluded as not a clean single-model result in this harness. Either the harness and the model could not work together (for example the harness sent an image to a text-only model endpoint), or the harness ran part of the job on a different model (a mixed-model run, for example Claude Code spawning a subagent on another model). The run's own note says which. Not counted as a result for the model and not an infrastructure failure.
  • Infrastructure DNFDid not finish because the connection upstream of the harness dropped the run (for example a silent stream cut after about 5 minutes). Not a model failure and not our harness configuration.
  • ErrorSomething else broke (a crash or an unexplained failure).

Requirement checklists

Each prompt's own numbered requirements, from the prompt registry.

PS5 Pro DualSense controller (max-fidelity SVG)

  1. Accurate DualSense silhouette and proportions: wide body, wide touchpad top centre, flared rounded grips, symmetrical, white shell with black core
  2. Face buttons in the classic diamond with correctly coloured outline PlayStation symbols on light buttons, subtle bevels
  3. D-pad as FOUR separate arrow-shaped buttons, not a single cross
  4. Two black analog sticks with concave tops, textured grip ring and shading
  5. PS logo button centred below the touchpad, mic mute button with orange indicator below it, Create (left) and Options (right) flanking the touchpad
  6. Blue light-bar glow along both touchpad edges with bloom; speaker grille holes on the black core
  7. Realistic product-render shading: gradients, specular highlights, ambient occlusion, contact shadow
  8. Dark studio background with radial vignette; front-on hero shot, centred and filling the frame

PS5 Pro DualSense controller (quick SVG)

  1. White main body with a black central section
  2. Large touchpad on top with blue light-bar edges
  3. D-pad on the left
  4. Four face buttons with triangle / circle / cross / square symbols
  5. Two black analog sticks
  6. PS button and mic mute button between the sticks; Create and Options buttons
  7. Curved grips and accurate proportions
  8. Front-on centred hero shot with gradients, highlights and soft shadows on a subtle dark studio background

PS5 Pro console and controller product shot (SVG)

  1. Slim PS5 Pro console design
  2. White side panels with a glossy black central band
  3. Three horizontal vent slits across the white panels
  4. Console standing upright
  5. White-and-black DualSense controller beside it
  6. Clean studio product shot with a soft gradient background
  7. Subtle floor reflection and soft shadows
  8. Accurate proportions, gradients and highlights for realism

Cinematic fjord boat scene (Three.js)

  1. One self-contained index.html; Three.js r160 from the pinned import map; all textures procedural
  2. Procedural fjord walls and mountains with snow caps and height/slope colouring
  3. Instanced conifer forests, distant islands, a lighthouse with a rotating beam
  4. Animated water with reflections; a wooden boat with two animated characters
  5. Weather cycle clear > fog > rain > storm > clearing with smooth transitions, lightning, wind and waves
  6. Full day/night cycle with sky, sun/moon and fog tracking time; stars, moon, lit lighthouse and lantern
  7. Keyboard and touch controls, follow camera, cinematic auto-tour, minimal UI
  8. URL parameters (tour, capture, cycle), deterministic timing, __sceneReady, smooth on phones, no console errors

Animated short film drawn in code (HTML/JS)

  1. One file, no external assets; 1920x1080 canvas letterboxed to the window
  2. Deterministic: requestAnimationFrame + performance.now only, seeded noise, FILM_DURATION, holds the last frame
  3. Original WebAudio score via buildSoundtrack / renderSoundtrack, with a tap-for-sound hint
  4. Original story with an emotional arc and one consistent main character in every scene
  5. At least four distinct scenes with real transitions
  6. Expressive character animation: walk cycle, blinking, changing emotions
  7. Virtual camera with parallax depth; lighting and time-of-day changes with grading and glows
  8. Two or more particle/weather types, storybook subtitles, cohesive art direction, closing title

Wave-1 SVG scoring rubric (AI AI Model Reviewer)

Pre-registered rubric written by the wave-1 SVG test runner (received 2026-09-23 21:25 UTC). Covers: pelican_bicycle, macbook_pro_16, lighthouse_night, isometric_town, pelican_bicycle_animated. 63 results currently carry a draft score under it. Published as received in the results folder, except that any personal name or internal file name was replaced (a name becomes "the reviewer").

Written once before any wave-1 result was scored, and kept for every wave-1 SVG card. Score 0–10, comparable only within one prompt. Scores are given only to ok runs; DNF / error / timeout cards carry no score.

Structure (matches the results format: up to 8 checklist + up to 2 craft)

Checklist, 8 points — the prompt's own requirements, split into 4 items of 2 points each (1 = partly there, 0 = missing/wrong):

promptitem 1 (2)item 2 (2)item 3 (2)item 4 (2)
pelican_bicyclereads instantly as a pelican (long beak + pouch, body, wing, legs)a real bicycle: two wheels, spokes or hubs, a diamond-ish frame, pedals/crank, handlebars, seatpelican is riding it: sits on the seat, feet on pedals, wing/beak toward the bars; plausible contact pointscomposition: whole subject in frame, ground/scene context, readable at thumbnail size
macbook_pro_16recognisable MacBook Pro 16" (proportions, notch/camera, keyboard + large trackpad, hinge)slight three-quarter view with a consistent perspective (lid and base share one vanishing logic, no warped keys)screen glow: lit display with believable content and light spillaluminium material: gradients, edge highlights, shadows that fake 3D
lighthouse_nightlighthouse on a rocky headland, correct structure (tower, lantern room, gallery)stormy sea with waves breaking; moonlight on the waterbeam cutting through rain and mist (volumetric feel, rain streaks)atmosphere/painterliness: mood, depth, colour harmony
isometric_townsprite sheet first: several distinct building designs in <defs>, reused via <use>20+ buildings placed on a true isometric grid, correct z-ordering (back-to-front, no overlaps in the wrong order)roads, trees, a river and a bridge that crosses ittheme + palette chosen and applied consistently
pelican_bicycle_animatedpelican + bicycle drawn well (same bar as the control)wheels rotate (correct direction) and the crank/pedals turnlegs pedal in sync with the crank (feet stay on the pedals, knees bend plausibly)seamless loop, SVG/SMIL/CSS only, no JavaScript (a <script> = 0 here)

Craft, 2 points — polish and finish: clean shapes, no glitches (stray paths, clipped elements, broken filters, text overflow), pleasing colour and detail, renders the same in Chromium as the model intended, file saved under the requested name. 2 = showcase quality, 1 = fine with visible flaws, 0 = rough.

Method

  • Every score is given after looking at the 1600 px PNG render (headless Chromium) and, for the animated prompt, the 6-frame contact sheet (t = 0, 0.25, 0.5, 0.75, 1.25, 2.0 s; the SVG clock is set per frame).
  • score_note lists the five sub-scores, e.g. pelican 2 · bicycle 1.5 · riding 1 · composition 2 · craft 1.5 = 8.
  • Half points allowed. Cost and time are not part of the score; they sit beside it on the card.
  • Reviewer: the test runner, acting as the AI judge; shown as provisional until the reviewer overrides in config/reviews.json.

Field notes: harness config vs model

Incidents from the test runs, written up like QA tickets. Each one shows why a run that did not finish is not automatically the model's fault, which is the point of the project: a failure is only counted against a model when the model caused it.

Claude Code's stream-idle watchdog cut silent max-effort thinking

harness configuration (ours)fixed

Symptom
Runs ended with an upstream HTTP 524 timeout and no artifact, after minutes in which the model had streamed nothing.
Cause
Claude Code kills a stream that stays silent for 180 s (its first-party default) and retries the request without streaming. Models at max effort think silently for longer than that, and the non-streaming retry then hit the upstream gateway's 120 s limit (HTTP 524).
Fix
Every Claude Code run from then on sets CLAUDE_BYTE_STREAM_IDLE_TIMEOUT_MS=1800000 (30 minutes) and CLAUDE_CODE_DISABLE_NONSTREAMING_FALLBACK=1, including this site's own build sessions.
Impact on results
Claude Fable 5.1 attempt 1 on the DualSense prompt and the first Claude Code fjord attempt with Opus 5.5 are labelled Harness-config failure (ours), not DNF; their spend is counted. This site's build session 01 also lost one turn to it (see How this site was built).

Streams silent for about 5 minutes are dropped upstream of the harness

infrastructureopen

Symptom
With the watchdog raised to 30 minutes and the fallback off, Claude Fable 5.1's stream still ended after 5 m 16 s of silent thinking: “API Error: Stream ended without receiving any events”.
Cause
A stream that emits nothing for about 5 minutes is dropped upstream of the harness, whatever the harness settings. Fable 5.1 at max effort thinks silently for longer than that.
Fix
None on our side yet. Until the upstream limit changes, Fable 5.1 at max effort cannot finish through this setup in any harness.
Impact on results
Claude Fable 5.1 attempt 2 on the DualSense prompt is labelled Infrastructure DNF (upstream silence drop): not a model failure and not a harness-config failure.

OpenCode's 5-minute stream timeouts

harness defaultsfixed

Symptom
A risk rather than an observed failure: OpenCode would abort a request that sends no chunk, or no response headers, for 5 minutes.
Cause
OpenCode defaults to a 5-minute idle-chunk timeout and a 5-minute header timeout, shorter than the silent thinking of some models at max effort.
Fix
Raised to 30 minutes for every provider from the GPT-6 Astra OpenCode run onward, and recorded in each run's meta.
Impact on results
The Claude Opus 5.5 OpenCode run started before the change and used the defaults; it was not affected and finished on its own.

The test box moved to a larger machine mid-project

infrastructure changedone

Symptom
Runs in flight when the test box was moved to a larger machine (8 GB) were cut off.
Cause
A planned infrastructure change, not a harness or model problem.
Fix
Runs cut by the move are re-run or resumed. Their cards say “restarted after machine upgrade” or “resumed after machine upgrade”, taken from the run's own record (its labels or run note).
Impact on results
Every published run that carries one of those labels is listed with this note.

Data & downloads