{
 "schema_version": 1,
 "generated_at": "2026-09-30T04:39:24Z",
 "data_hash": "93cc09f98909ae95",
 "methodology": {
  "title": "Methodology",
  "intro_html": "<p>How every result on this site was produced and labelled. <code>ingest.py</code> renders this file into the methodology page and the home-page summary, and publishes it in <code>data/content.json</code> (one entry per <code>##</code> section).</p>",
  "sections": [
   {
    "id": "three-tracks",
    "title": "Three tracks",
    "html": "<p><strong>Home harness</strong>: each vendor's own CLI (Claude Code for Anthropic, Codex CLI for OpenAI, dsh for DeepSeek, Grok Build for xAI, Gemini CLI for Google, Kimi Code CLI for Moonshot). GLM-5.3 has no coding CLI of its own, so its home track runs in Claude Code, Z.ai's recommended harness, and is labelled that way. <strong>Neutral</strong>: every model in OpenCode with the same configuration. <strong>Raw baseline</strong>: an earlier single API request with no harness and no tools. Baselines are kept for reference only; the project's rule is now harness-only. Provider = the API service that served the model for that run. Gemini 3.8 Flash stays in the roster and in the planned totals, but none of its cells is tested: we have no usable access to the model.</p>"
   },
   {
    "id": "max-effort",
    "title": "Max effort",
    "html": "<p>Every model runs at the highest reasoning effort its API accepts (max for Claude and DeepSeek, xhigh for GPT-6 Astra and Grok 4.7, high for Gemini 3.8 Flash). A few raw baselines ran at lower effort; the effort is printed on every result.</p>"
   },
   {
    "id": "one-prompt-fully-autonomous",
    "title": "One prompt, fully autonomous",
    "html": "<p>One prompt, then no human input until the model stops. Inside a harness the model may search the web, render and inspect its own output and fix it (self-review allowed). The single exception, where the reviewer sent written feedback once, is labelled <strong>1 fix turn</strong>.</p>"
   },
   {
    "id": "whose-failure-was-it",
    "title": "Whose failure was it?",
    "html": "<p><strong>DNF</strong> means the model failed on its own, for example by spending its whole output budget thinking. When <em>we</em> end a run (our client or safety timeout, or a rule change), it is marked <strong>Stopped</strong> and not counted against the model. <strong>Harness-config failure (ours)</strong> means our own harness settings killed the run. <strong>Harness + model mismatch</strong> means the run is excluded as not a clean single-model result: either the harness and the model could not work together (for example Claude Code sending an image to GLM-5.3's text-only endpoint), or the harness ran part of the job on a different model (a mixed-model run, for example Claude Code spawning a subagent on another model); each run's note says which. <strong>Infrastructure DNF</strong> means the connection upstream of the harness dropped the run. None of the last three is counted as a model failure; the field notes describe each incident. The spend on every failed run is still counted.</p>\n<p>Attempts lost to the harness launcher or the connection before a result existed are not results: they are listed under the results as <strong>infrastructure failures (not model results)</strong> and left out of every count, total per model and head-to-head. A planned cell that is queued but has not run yet says <strong>Pending — queued, not run yet</strong>.</p>"
   },
   {
    "id": "costs-reported-vs-estimated",
    "title": "Costs: reported vs estimated",
    "html": "<p><strong>Reported</strong>: the harness printed its own bill (for example Claude Code's <code>total_cost_usd</code> or OpenCode's session cost). When a harness reported one, it is the headline cost, labelled as that harness's figure, with the list-price estimate beside it. <strong>Estimated</strong>: token counts times the vendor's list price, with the price table and its date on every result. <strong>Unknown</strong>: the tokens were never reported, so we don't guess. When no token usage was recorded at all (for example a stream cut before its usage record), the cost reads <strong>no usage recorded</strong>, never $0. Money spent on runs that produced nothing is shown, never dropped. Totals say what they include: the result totals cover the results on this site, failed runs included; setup smoke tests and infrastructure failures are totalled separately.</p>\n<p>15 OpenCode results ran on a box whose per-request proxy logs were lost; the served model for those runs could not be independently verified. Their per-step costs match the requested model's prices.</p>"
   },
   {
    "id": "how-results-are-scored",
    "title": "How results are scored",
    "html": "<p>Scores have three separate sources: the <strong>AI judge</strong>, <strong>visitor votes</strong> (coming later, with a blind reveal), and the <strong>reviewer's grade</strong>. The current AI judge is the test runner, model <strong>claude-opus-5-5</strong>, described as reviewing the 1600 px render against the published, pre-registered rubric. Each scored run currently has one recorded AI-judge assessment. The reviewer has not given any numeric grades yet. Reviewer grades and visitor votes, when available, appear alongside the AI score and are not blended into it.</p>\n<p>Only finished runs covered by a published, pre-registered rubric receive a score. Scores run from 0 to 10: up to 8 points for the prompt's requirements and up to 2 for craft, with half points for partial credit. Cost and time do not contribute to this score. Each AI score links to its rubric and remains labelled <strong>Provisional — AI judge</strong> until the reviewer checks it. Scores originate with the judge or reviewer; the site's build agent does not assign them.</p>\n<p>We publish the per-criterion marks and calculate the total from them. When those marks cannot be read into the published breakdown, we retain the judge's stored total and label it <strong>Provisional — stored total; criterion breakdown unavailable</strong>. The displayed total has not been verified through that breakdown. A prompt's board stays unranked until every published AI score on it has a complete, validated criterion breakdown. Missing breakdowns and any disagreement between marks and totals must be resolved before ranking that board; affected runs remain visible throughout.</p>\n<p>Eligible <strong>AI-judge boards are ranked by score within one prompt</strong>, highest first. Each row represents a particular run, identified by its model, harness, provider, effort and attempt. Equal totals share a rank, and the next rank skips the occupied positions: 1, 1, 3. Row order inside a tie has no quality meaning. These are provisional ranks of the recorded assessments. They do not establish that a model is generally better, or that a small score difference would persist in another run or under another judge. We do not publish confidence intervals or treat unequal scores as statistical ties without evidence of judging and run variability.</p>\n<p>Scores from different prompts are never averaged or used to produce an overall model ranking. Home and neutral harness conditions stay identified, and raw baselines remain a separate reference section. Each board states how many recorded results are scored, how many are unfinished and how many remain unscored. The reviewer's written recommendations and “best” picks remain separately attributed editorial judgments.</p>\n<p><strong>DNF is a result.</strong> DNF, stopped, failed and no-output results remain visible on the same prompt board, with their recorded status, reason, time and cost. They receive no quality rank and no invented zero score. <strong>No output artifact</strong> describes the absence of a published deliverable; it does not by itself identify whose failure it was. A stopped run with partial output keeps that output, labelled <strong>Work in progress</strong>. Finished runs without a score say <strong>Not scored</strong>, with the reason. Prompts with no scores still have an unranked result board.</p>\n<p>Model DNF remains distinct from our stops, harness-config failures, harness/model mismatches, infrastructure failures and unexplained errors. All records published as results stay in result totals, including failures. Separately logged infrastructure attempts that are not result records remain visible as <strong>Infrastructure failures (not model results)</strong> and outside those totals. Pending and no-access cells remain visible as <strong>Pending — queued, not run yet</strong> or <strong>Not tested (no access)</strong>; neither is a scored run. Unknown cost is never displayed as zero.</p>"
   }
  ],
  "source": "content/methodology.md"
 },
 "statuses": [
  {
   "id": "ok",
   "label": "OK",
   "help": "Finished on its own and produced the artifact."
  },
  {
   "id": "dnf",
   "label": "DNF",
   "help": "Did not finish: the model gave up or used up its own budget (e.g. hit max_tokens while still thinking). Counted against the model."
  },
  {
   "id": "stopped_by_our_timeout",
   "label": "Stopped (our timeout)",
   "help": "Our client or safety timeout ended the run. Not a model failure."
  },
  {
   "id": "stopped_by_us",
   "label": "Stopped by us",
   "help": "We ended the run ourselves (for example after a rule change). Not a model failure."
  },
  {
   "id": "harness_config_failure",
   "label": "Harness-config failure (ours)",
   "help": "Our harness configuration killed the run. Not a model failure and not our safety timeout."
  },
  {
   "id": "harness_mismatch",
   "label": "Harness + model mismatch",
   "help": "Excluded as not a clean single-model result in this harness. Either the harness and the model could not work together (for example the harness sent an image to a text-only model endpoint), or the harness ran part of the job on a different model (a mixed-model run, for example Claude Code spawning a subagent on another model). The run's own note says which. Not counted as a result for the model and not an infrastructure failure."
  },
  {
   "id": "infra_dnf",
   "label": "Infrastructure DNF",
   "help": "Did not finish because the connection upstream of the harness dropped the run (for example a silent stream cut after about 5 minutes). Not a model failure and not our harness configuration."
  },
  {
   "id": "error",
   "label": "Error",
   "help": "Something else broke (a crash or an unexplained failure)."
  }
 ],
 "scoring": {
  "rule": "Scores come from three sources, shown side by side. The AI judge (the test runner, labelled on every score) scores finished runs against a published, pre-registered rubric; its per-criterion marks are shown and the total is computed from them, labelled a draft score until the reviewer checks it. Visitor votes are coming later. The reviewer's own grade appears when it exists. The site's build agent never scores. AI-judge boards are ranked by score within one prompt, equal totals sharing a rank; scores from different prompts are never averaged.",
  "not_scored_label": "Unscored",
  "scored_results": 63,
  "checklists": [
   {
    "prompt_id": "ps5_controller",
    "title": "PS5 Pro DualSense controller (max-fidelity SVG)",
    "checklist": [
     "Accurate DualSense silhouette and proportions: wide body, wide touchpad top centre, flared rounded grips, symmetrical, white shell with black core",
     "Face buttons in the classic diamond with correctly coloured outline PlayStation symbols on light buttons, subtle bevels",
     "D-pad as FOUR separate arrow-shaped buttons, not a single cross",
     "Two black analog sticks with concave tops, textured grip ring and shading",
     "PS logo button centred below the touchpad, mic mute button with orange indicator below it, Create (left) and Options (right) flanking the touchpad",
     "Blue light-bar glow along both touchpad edges with bloom; speaker grille holes on the black core",
     "Realistic product-render shading: gradients, specular highlights, ambient occlusion, contact shadow",
     "Dark studio background with radial vignette; front-on hero shot, centred and filling the frame"
    ]
   },
   {
    "prompt_id": "ps5_controller_quick",
    "title": "PS5 Pro DualSense controller (quick SVG)",
    "checklist": [
     "White main body with a black central section",
     "Large touchpad on top with blue light-bar edges",
     "D-pad on the left",
     "Four face buttons with triangle / circle / cross / square symbols",
     "Two black analog sticks",
     "PS button and mic mute button between the sticks; Create and Options buttons",
     "Curved grips and accurate proportions",
     "Front-on centred hero shot with gradients, highlights and soft shadows on a subtle dark studio background"
    ]
   },
   {
    "prompt_id": "ps5_console",
    "title": "PS5 Pro console and controller product shot (SVG)",
    "checklist": [
     "Slim PS5 Pro console design",
     "White side panels with a glossy black central band",
     "Three horizontal vent slits across the white panels",
     "Console standing upright",
     "White-and-black DualSense controller beside it",
     "Clean studio product shot with a soft gradient background",
     "Subtle floor reflection and soft shadows",
     "Accurate proportions, gradients and highlights for realism"
    ]
   },
   {
    "prompt_id": "fjord",
    "title": "Cinematic fjord boat scene (Three.js)",
    "checklist": [
     "One self-contained index.html; Three.js r160 from the pinned import map; all textures procedural",
     "Procedural fjord walls and mountains with snow caps and height/slope colouring",
     "Instanced conifer forests, distant islands, a lighthouse with a rotating beam",
     "Animated water with reflections; a wooden boat with two animated characters",
     "Weather cycle clear > fog > rain > storm > clearing with smooth transitions, lightning, wind and waves",
     "Full day/night cycle with sky, sun/moon and fog tracking time; stars, moon, lit lighthouse and lantern",
     "Keyboard and touch controls, follow camera, cinematic auto-tour, minimal UI",
     "URL parameters (tour, capture, cycle), deterministic timing, __sceneReady, smooth on phones, no console errors"
    ]
   },
   {
    "prompt_id": "js_short",
    "title": "Animated short film drawn in code (HTML/JS)",
    "checklist": [
     "One file, no external assets; 1920x1080 canvas letterboxed to the window",
     "Deterministic: requestAnimationFrame + performance.now only, seeded noise, FILM_DURATION, holds the last frame",
     "Original WebAudio score via buildSoundtrack / renderSoundtrack, with a tap-for-sound hint",
     "Original story with an emotional arc and one consistent main character in every scene",
     "At least four distinct scenes with real transitions",
     "Expressive character animation: walk cycle, blinking, changing emotions",
     "Virtual camera with parallax depth; lighting and time-of-day changes with grading and glows",
     "Two or more particle/weather types, storybook subtitles, cohesive art direction, closing title"
    ]
   }
  ]
 },
 "rubrics": [
  {
   "id": "rubric",
   "title": "Wave-1 SVG scoring rubric (AI AI Model Benchmark)",
   "covers": [
    "pelican_bicycle",
    "macbook_pro_16",
    "lighthouse_night",
    "isometric_town",
    "pelican_bicycle_animated"
   ],
   "received_utc": "2026-09-23 21:25 UTC",
   "author": "the wave-1 SVG test runner",
   "url": "methodology.html#rubric",
   "html": "<p>Written once before any wave-1 result was scored, and kept for every wave-1 SVG card. Score 0–10, comparable only within one prompt. Scores are given only to <code>ok</code> runs; DNF / error / timeout cards carry no score.</p>\n<h2 id=\"structure-matches-results-format-up-to-8-checklist-up-to-2-craft\">Structure (matches the results format: up to 8 checklist + up to 2 craft)</h2>\n<p><strong>Checklist, 8 points</strong> — the prompt's own requirements, split into 4 items of 2 points each (1 = partly there, 0 = missing/wrong):</p>\n<div class=\"table-wrap\"><table><thead><tr><th>prompt</th><th>item 1 (2)</th><th>item 2 (2)</th><th>item 3 (2)</th><th>item 4 (2)</th></tr></thead><tbody><tr><td>pelican_bicycle</td><td>reads instantly as a pelican (long beak + pouch, body, wing, legs)</td><td>a real bicycle: two wheels, spokes or hubs, a diamond-ish frame, pedals/crank, handlebars, seat</td><td>pelican is <em>riding</em> it: sits on the seat, feet on pedals, wing/beak toward the bars; plausible contact points</td><td>composition: whole subject in frame, ground/scene context, readable at thumbnail size</td></tr><tr><td>macbook_pro_16</td><td>recognisable MacBook Pro 16\" (proportions, notch/camera, keyboard + large trackpad, hinge)</td><td>slight three-quarter view with a consistent perspective (lid and base share one vanishing logic, no warped keys)</td><td>screen glow: lit display with believable content and light spill</td><td>aluminium material: gradients, edge highlights, shadows that fake 3D</td></tr><tr><td>lighthouse_night</td><td>lighthouse on a rocky headland, correct structure (tower, lantern room, gallery)</td><td>stormy sea with waves breaking; moonlight on the water</td><td>beam cutting through rain and mist (volumetric feel, rain streaks)</td><td>atmosphere/painterliness: mood, depth, colour harmony</td></tr><tr><td>isometric_town</td><td>sprite sheet first: several distinct building designs in <code>&lt;defs&gt;</code>, reused via <code>&lt;use&gt;</code></td><td>20+ buildings placed on a true isometric grid, correct z-ordering (back-to-front, no overlaps in the wrong order)</td><td>roads, trees, a river and a bridge that crosses it</td><td>theme + palette chosen and applied consistently</td></tr><tr><td>pelican_bicycle_animated</td><td>pelican + bicycle drawn well (same bar as the control)</td><td>wheels rotate (correct direction) and the crank/pedals turn</td><td>legs pedal in sync with the crank (feet stay on the pedals, knees bend plausibly)</td><td>seamless loop, SVG/SMIL/CSS only, no JavaScript (a <code>&lt;script&gt;</code> = 0 here)</td></tr></tbody></table></div>\n<p><strong>Craft, 2 points</strong> — polish and finish: clean shapes, no glitches (stray paths, clipped elements, broken filters, text overflow), pleasing colour and detail, renders the same in Chromium as the model intended, file saved under the requested name. 2 = showcase quality, 1 = fine with visible flaws, 0 = rough.</p>\n<h2 id=\"method\">Method</h2>\n<ul><li>Every score is given after looking at the 1600 px PNG render (headless Chromium) and, for the animated prompt, the 6-frame contact sheet (t = 0, 0.25, 0.5, 0.75, 1.25, 2.0 s; the SVG clock is set per frame).</li><li><code>score_note</code> lists the five sub-scores, e.g. <code>pelican 2 · bicycle 1.5 · riding 1 · composition 2 · craft 1.5 = 8</code>.</li><li>Half points allowed. Cost and time are <strong>not</strong> part of the score; they sit beside it on the card.</li><li>Reviewer: the test runner, acting as the AI judge; shown as provisional until the reviewer overrides in config/reviews.json.</li></ul>",
   "note": "Published as received in the results folder, except that any personal name or internal file name was replaced (a name becomes \"the reviewer\").",
   "results_scored": [
    "isometric_town__claude-code__fable-5-1",
    "isometric_town__claude-code__opus-5-5",
    "isometric_town__codex__gpt-6-astra",
    "isometric_town__dsh__deepseek-flash",
    "isometric_town__grok-build__grok-4.7",
    "isometric_town__kimi-code__kimi-k3",
    "isometric_town__opencode__deepseek-flash",
    "isometric_town__opencode__fable-5-1",
    "isometric_town__opencode__glm-5.3",
    "isometric_town__opencode__gpt-6-astra",
    "isometric_town__opencode__grok-4.7",
    "isometric_town__opencode__kimi-k3",
    "lighthouse_night__claude-code__fable-5-1",
    "lighthouse_night__claude-code__opus-5-5",
    "lighthouse_night__codex__gpt-6-astra",
    "lighthouse_night__dsh__deepseek-flash",
    "lighthouse_night__grok-build__grok-4.7",
    "lighthouse_night__kimi-code__kimi-k3",
    "lighthouse_night__opencode__deepseek-flash",
    "lighthouse_night__opencode__fable-5-1",
    "lighthouse_night__opencode__glm-5.3",
    "lighthouse_night__opencode__gpt-6-astra",
    "lighthouse_night__opencode__grok-4.7",
    "lighthouse_night__opencode__kimi-k3",
    "lighthouse_night__opencode__opus-5-5",
    "macbook_pro_16__claude-code__fable-5-1",
    "macbook_pro_16__claude-code__opus-5-5",
    "macbook_pro_16__codex__gpt-6-astra",
    "macbook_pro_16__dsh__deepseek-flash",
    "macbook_pro_16__grok-build__grok-4.7",
    "macbook_pro_16__kimi-code__kimi-k3",
    "macbook_pro_16__opencode__deepseek-flash",
    "macbook_pro_16__opencode__fable-5-1",
    "macbook_pro_16__opencode__glm-5.3",
    "macbook_pro_16__opencode__gpt-6-astra",
    "macbook_pro_16__opencode__grok-4.7",
    "macbook_pro_16__opencode__kimi-k3",
    "macbook_pro_16__opencode__opus-5-5__attempt2",
    "pelican_bicycle__claude-code__fable-5-1",
    "pelican_bicycle__claude-code__opus-5-5",
    "pelican_bicycle__codex__gpt-6-astra",
    "pelican_bicycle__dsh__deepseek-flash",
    "pelican_bicycle__grok-build__grok-4.7",
    "pelican_bicycle__kimi-code__kimi-k3",
    "pelican_bicycle__opencode__deepseek-flash",
    "pelican_bicycle__opencode__fable-5-1",
    "pelican_bicycle__opencode__glm-5.3",
    "pelican_bicycle__opencode__gpt-6-astra",
    "pelican_bicycle__opencode__grok-4.7",
    "pelican_bicycle__opencode__kimi-k3",
    "pelican_bicycle__opencode__opus-5-5",
    "pelican_bicycle_animated__claude-code__fable-5-1",
    "pelican_bicycle_animated__claude-code__opus-5-5",
    "pelican_bicycle_animated__codex__gpt-6-astra",
    "pelican_bicycle_animated__dsh__deepseek-flash",
    "pelican_bicycle_animated__kimi-code__kimi-k3",
    "pelican_bicycle_animated__opencode__deepseek-flash",
    "pelican_bicycle_animated__opencode__fable-5-1",
    "pelican_bicycle_animated__opencode__glm-5.3",
    "pelican_bicycle_animated__opencode__gpt-6-astra",
    "pelican_bicycle_animated__opencode__grok-4.7",
    "pelican_bicycle_animated__opencode__kimi-k3",
    "pelican_bicycle_animated__opencode__opus-5-5"
   ],
   "results_for_covered_prompts": 163
  }
 ],
 "field_notes": {
  "title": "Field notes: harness config vs model",
  "intro": "Incidents from the test runs, written up like QA tickets. Each one shows why a run that did not finish is not automatically the model's fault, which is the point of the project: a failure is only counted against a model when the model caused it.",
  "intro_html": "Incidents from the test runs, written up like QA tickets. Each one shows why a run that did not finish is not automatically the model's fault, which is the point of the project: a failure is only counted against a model when the model caused it.",
  "notes": [
   {
    "id": "cc-stream-idle-watchdog",
    "title": "Claude Code's stream-idle watchdog cut silent max-effort thinking",
    "category": "harness configuration (ours)",
    "status": "fixed",
    "symptom": "Runs ended with an upstream HTTP 524 timeout and no artifact, after minutes in which the model had streamed nothing.",
    "symptom_html": "Runs ended with an upstream HTTP 524 timeout and no artifact, after minutes in which the model had streamed nothing.",
    "cause": "Claude Code kills a stream that stays silent for 180 s (its first-party default) and retries the request without streaming. Models at max effort think silently for longer than that, and the non-streaming retry then hit the upstream gateway's 120 s limit (HTTP 524).",
    "cause_html": "Claude Code kills a stream that stays silent for 180 s (its first-party default) and retries the request without streaming. Models at max effort think silently for longer than that, and the non-streaming retry then hit the upstream gateway's 120 s limit (HTTP 524).",
    "fix": "Every Claude Code run from then on sets `CLAUDE_BYTE_STREAM_IDLE_TIMEOUT_MS=1800000` (30 minutes) and `CLAUDE_CODE_DISABLE_NONSTREAMING_FALLBACK=1`, including this site's own build sessions.",
    "fix_html": "Every Claude Code run from then on sets <code>CLAUDE_BYTE_STREAM_IDLE_TIMEOUT_MS=1800000</code> (30 minutes) and <code>CLAUDE_CODE_DISABLE_NONSTREAMING_FALLBACK=1</code>, including this site's own build sessions.",
    "impact": "Claude Fable 5.1 attempt 1 on the DualSense prompt and the first Claude Code fjord attempt with Opus 5.5 are labelled **Harness-config failure (ours)**, not DNF; their spend is counted. This site's build session 01 also lost one turn to it (see [How this site was built](built.html)).",
    "impact_html": "Claude Fable 5.1 attempt 1 on the DualSense prompt and the first Claude Code fjord attempt with Opus 5.5 are labelled <strong>Harness-config failure (ours)</strong>, not DNF; their spend is counted. This site's build session 01 also lost one turn to it (see <a href=\"built.html\">How this site was built</a>).",
    "results": [
     {
      "id": "ps5_controller__claude-code__fable-5-1__attempt1",
      "model": "Claude Fable 5.1",
      "harness": "Claude Code 2.1.280",
      "prompt_id": "ps5_controller",
      "status": "harness_config_failure"
     },
     {
      "id": "fjord_claude_code_opus55_attempt1",
      "model": "Claude Opus 5.5",
      "harness": "Claude Code 2.1.280",
      "prompt_id": "fjord",
      "status": "harness_config_failure"
     }
    ],
    "missing_results": []
   },
   {
    "id": "upstream-silence-drop",
    "title": "Streams silent for about 5 minutes are dropped upstream of the harness",
    "category": "infrastructure",
    "status": "open",
    "symptom": "With the watchdog raised to 30 minutes and the fallback off, Claude Fable 5.1's stream still ended after 5 m 16 s of silent thinking: “API Error: Stream ended without receiving any events”.",
    "symptom_html": "With the watchdog raised to 30 minutes and the fallback off, Claude Fable 5.1's stream still ended after 5 m 16 s of silent thinking: “API Error: Stream ended without receiving any events”.",
    "cause": "A stream that emits nothing for about 5 minutes is dropped upstream of the harness, whatever the harness settings. Fable 5.1 at max effort thinks silently for longer than that.",
    "cause_html": "A stream that emits nothing for about 5 minutes is dropped upstream of the harness, whatever the harness settings. Fable 5.1 at max effort thinks silently for longer than that.",
    "fix": "None on our side yet. Until the upstream limit changes, Fable 5.1 at max effort cannot finish through this setup in any harness.",
    "fix_html": "None on our side yet. Until the upstream limit changes, Fable 5.1 at max effort cannot finish through this setup in any harness.",
    "impact": "Claude Fable 5.1 attempt 2 on the DualSense prompt is labelled **Infrastructure DNF** (upstream silence drop): not a model failure and not a harness-config failure.",
    "impact_html": "Claude Fable 5.1 attempt 2 on the DualSense prompt is labelled <strong>Infrastructure DNF</strong> (upstream silence drop): not a model failure and not a harness-config failure.",
    "results": [
     {
      "id": "ps5_controller__claude-code__fable-5-1__attempt2",
      "model": "Claude Fable 5.1",
      "harness": "Claude Code 2.1.280",
      "prompt_id": "ps5_controller",
      "status": "infra_dnf"
     }
    ],
    "missing_results": []
   },
   {
    "id": "opencode-stream-timeouts",
    "title": "OpenCode's 5-minute stream timeouts",
    "category": "harness defaults",
    "status": "fixed",
    "symptom": "A risk rather than an observed failure: OpenCode would abort a request that sends no chunk, or no response headers, for 5 minutes.",
    "symptom_html": "A risk rather than an observed failure: OpenCode would abort a request that sends no chunk, or no response headers, for 5 minutes.",
    "cause": "OpenCode defaults to a 5-minute idle-chunk timeout and a 5-minute header timeout, shorter than the silent thinking of some models at max effort.",
    "cause_html": "OpenCode defaults to a 5-minute idle-chunk timeout and a 5-minute header timeout, shorter than the silent thinking of some models at max effort.",
    "fix": "Raised to 30 minutes for every provider from the GPT-6 Astra OpenCode run onward, and recorded in each run's meta.",
    "fix_html": "Raised to 30 minutes for every provider from the GPT-6 Astra OpenCode run onward, and recorded in each run's meta.",
    "impact": "The Claude Opus 5.5 OpenCode run started before the change and used the defaults; it was not affected and finished on its own.",
    "impact_html": "The Claude Opus 5.5 OpenCode run started before the change and used the defaults; it was not affected and finished on its own.",
    "results": [
     {
      "id": "ps5_controller__opencode__opus-5-5",
      "model": "Claude Opus 5.5",
      "harness": "OpenCode 1.18.32",
      "prompt_id": "ps5_controller",
      "status": "ok"
     }
    ],
    "missing_results": []
   },
   {
    "id": "machine-move",
    "title": "The test box moved to a larger machine mid-project",
    "category": "infrastructure change",
    "status": "done",
    "symptom": "Runs in flight when the test box was moved to a larger machine (8 GB) were cut off.",
    "symptom_html": "Runs in flight when the test box was moved to a larger machine (8 GB) were cut off.",
    "cause": "A planned infrastructure change, not a harness or model problem.",
    "cause_html": "A planned infrastructure change, not a harness or model problem.",
    "fix": "Runs cut by the move are re-run or resumed. Their cards say “restarted after machine upgrade” or “resumed after machine upgrade”, taken from the run's own record (its labels or run note).",
    "fix_html": "Runs cut by the move are re-run or resumed. Their cards say “restarted after machine upgrade” or “resumed after machine upgrade”, taken from the run's own record (its labels or run note).",
    "impact": "Every published run that carries one of those labels is listed with this note.",
    "impact_html": "Every published run that carries one of those labels is listed with this note.",
    "results": [
     {
      "id": "fjord_claudecode_opus55",
      "model": "Claude Opus 5.5",
      "harness": "Claude Code 2.1.280",
      "prompt_id": "fjord",
      "status": "ok"
     }
    ],
    "missing_results": []
   }
  ],
  "source": "content/field_notes.json"
 },
 "findings": [
  {
   "id": "best-controller",
   "title": "Best DualSense controller so far: OpenCode with Claude Opus 5.5, at a price",
   "source": "reviewer",
   "html": "<a href=\"r/ps5_controller__opencode__opus-5-5/\">OpenCode + Claude Opus 5.5 (max)</a> drew the best controller of the project so far: an accurate silhouette with an open gap between the grips, a real 4-way D-pad, correctly coloured outline symbols, knurled sticks, glowing blue light bars, a proper PS logo, a speaker grille and a small orange mute button. Its only flat spot is a plain white touchpad. It took 1h 23m and cost $18.88 (reported by OpenCode), about 1.5x the time and cost of <a href=\"r/ps5_controller__claude-code__opus-5-5/\">Claude Code with the same model</a> (56m 51s, $12.80) for a very similar look: same model, same photo-tracing approach.",
   "results": [
    "ps5_controller__claude-code__opus-5-5",
    "ps5_controller__opencode__opus-5-5"
   ]
  },
  {
   "id": "cleanest-home",
   "title": "Cleanest home-harness controller, and the best value among the strong ones: Codex with GPT-6 Astra",
   "source": "reviewer",
   "html": "<a href=\"r/ps5_controller__codex__gpt-6-astra/\">Codex CLI + GPT-6 Astra (xhigh)</a> is the cleanest result on a home harness: an open gap between the grips and no dark slab, in a flatter, more vector look. At $2.20 (estimated) and 11m 38s it is the best result per dollar among the strong results.",
   "results": [
    "ps5_controller__codex__gpt-6-astra"
   ]
  },
  {
   "id": "cc-opus",
   "title": "Claude Code with Opus 5.5: strong materials, one layout flaw",
   "source": "reviewer",
   "html": "<a href=\"r/ps5_controller__claude-code__opus-5-5/\">Claude Code + Claude Opus 5.5 (max)</a> has strong materials and detail, but a dark slab fills the gap between the grips. $12.80 (reported), 56m 51s.",
   "results": [
    "ps5_controller__claude-code__opus-5-5"
   ]
  },
  {
   "id": "dsh-value",
   "title": "dsh with DeepSeek V4.1 Flash: remarkable value, clear flaws",
   "source": "reviewer",
   "html": "<a href=\"r/ps5_controller__dsh__deepseek-flash/\">dsh + DeepSeek V4.1 Flash (max)</a> made a bold, glossy render, but a huge dark slab fills the space between the grips, the face-button diamond is lopsided, the PS logo is garbled and the touchpad is glassy. It sits behind Codex/Astra and Claude Code/Opus. The value is remarkable: $0.19 computed cost (estimated), 19m 44s and 131 tool calls, 59 of them image views.",
   "results": [
    "ps5_controller__dsh__deepseek-flash"
   ]
  },
  {
   "id": "fjord-codex",
   "title": "Fjord scene: Codex with GPT-6 Astra is night and day better than the raw one-shot",
   "source": "reviewer",
   "html": "<a href=\"r/fjord_codex_astra/\">Codex CLI + GPT-6 Astra (xhigh)</a> built green conifer slopes, snowy peaks, near-mirror reflections, a red-and-white lighthouse whose beam sweeps the water at night, a pink sunset sky and a wooden rowboat with two small rowers, in 25m 14s for about $5.57 (estimated from tokens). The <a href=\"r/fjord__raw-api__opus-5-5__low/\">raw one-shot build</a> had flat grey cliff slabs. Weak spots: the terrain reads low-poly up close, the night stretch is very dark, the characters are tiny and simple, and rain and storm barely show. It is not at the level of the viral reference post.",
   "results": [
    "fjord__raw-api__opus-5-5__low",
    "fjord_codex_astra"
   ]
  },
  {
   "id": "setup-failures",
   "title": "Three runs were lost to our setup, not to the models",
   "source": "run notes",
   "html": "Claude Code's stream-idle watchdog (our configuration) killed <a href=\"r/ps5_controller__claude-code__fable-5-1__attempt1/\">Claude Fable 5.1 attempt 1</a> and <a href=\"r/fjord_claude_code_opus55_attempt1/\">the first Claude Code fjord attempt with Opus 5.5</a>, which had spent $5.03 by then. <a href=\"r/ps5_controller__claude-code__fable-5-1__attempt2/\">Fable 5.1 attempt 2</a>, with the watchdog raised, was still cut upstream after about 5 minutes of silent thinking and is logged as an infrastructure DNF. None of these counts against the model; the spend is still counted. Details in the <a href=\"methodology.html#field-notes\">field notes</a>.",
   "results": [
    "fjord_claude_code_opus55_attempt1",
    "ps5_controller__claude-code__fable-5-1__attempt1",
    "ps5_controller__claude-code__fable-5-1__attempt2"
   ]
  },
  {
   "id": "raw-max-vs-high",
   "title": "On a raw call, max effort lost to high effort",
   "source": "data",
   "html": "Claude Opus 5.5 at max effort <a href=\"r/ps5_controller__raw-api__opus-5-5__max-128k/\">spent its entire 128,000-token output budget thinking</a> (20m 08s, $2.56) and never produced an SVG; with a 64k budget it <a href=\"r/ps5_controller__raw-api__opus-5-5__max-64k/\">did the same</a> ($1.28). Effort high <a href=\"r/ps5_controller__raw-api__opus-5-5__high/\">finished</a> in 7m 21s for $0.91. The same happened on the animated-short prompt: <a href=\"r/js_short__raw-api__opus-5-5__max/\">max effort</a> thought for 20m 42s and wrote nothing ($2.56), while <a href=\"r/js_short__raw-api__opus-5-5__low/\">effort low</a> made the whole film in 1m 38s for $0.24.",
   "results": [
    "js_short__raw-api__opus-5-5__low",
    "js_short__raw-api__opus-5-5__max",
    "ps5_controller__raw-api__opus-5-5__high",
    "ps5_controller__raw-api__opus-5-5__max-128k",
    "ps5_controller__raw-api__opus-5-5__max-64k"
   ]
  }
 ],
 "featured": {
  "why": "Picked from the reviewer's notes: the strongest harness results so far.",
  "results": [
   "ps5_controller__opencode__opus-5-5",
   "ps5_controller__codex__gpt-6-astra",
   "fjord_codex_astra",
   "ps5_controller__claude-code__opus-5-5"
  ]
 }
}
