AI Model Reviewer

Qwen 3.8 Max in OpenCode 1.18.35

skills on · OpenCode 1.18.35 + skills on (54: 33 video + 21 web/Three.js)Kit: 33 video + 21 web/Three.js · 0 of 54 installed skills used in the session

← Open in the gallerySame prompt, all models

DNF

The model did not finish, so there is no output.

DNFGame battleshipn/a effortNeutral harness

Wall time 10m 45sCost $0.23 reported

Review

No verdict yet.

The full review

DNF. Harness exited after 10.8 min without writing index.html. The model's last message stopped at the output-token limit (stop reason max_tokens / length) without a tool call, so the non-interactive run ended there.

Scores

Not scored: only runs that finished on their own are scored.

Downloads

Files actually published for this run. Original output, site-made previews and post-run conversions are labelled separately.

  • Scrubbed run transcriptgame_battleship__opencode__qwen3.8-max-skills-on__attempt1-transcript.json · 125 KB · transcript

The exact prompt

as run, 492 characters. Every result for this prompt

Build a complete, polished Battleship game playable in the browser (index.html plus any JS/CSS, no build step). Player vs computer on 10x10 grids: ship placement (drag or click, with rotate), a computer opponent with a sensible hunt/target strategy, hit/miss/sunk feedback with animation and sound, turn indicator, win/lose screen and restart. It must work with mouse and touch and look good on desktop and phone. Write everything into the working directory and finish with a short README.md.

How it was made

Wall time 10m 45s

Step by step

The prompt, every agent turn and tool call (tool name and a one-line summary; inputs and outputs are not shown), then the final message. The full machine-format transcript is published as a scrubbed download.

Show all 4 steps

Loading the timeline…

Download the full transcript OpenCode session export, 125 KB. Scrubbed: anything that looks like a key, token, e-mail address or phone number is replaced, and account details and the test machine's file paths are removed.

Cost breakdown

Headline cost $0.23, as reported by OpenCode (its own bill for the run).

Tokens by typeTokens$ per 1MCost
Input (uncached)15,010n/an/a
Cache write0n/an/a
Cache read14,890n/an/a
Output32,047n/an/a
Total computed from tokens
unknown
Total reported by the harness
$0.2260

OpenCode 1.18.35 reported its own total; token counts x list prices shown beside it. No list price for qwen3.8-max in config/prices.json; no estimate computed. The harness reports reasoning tokens separately from output; here output includes them (30 visible + 32,017 reasoning), and both bill at the output rate.

Run notes

one prompt, agentic harness (self-review allowed)skills=on

  • frozen prompt (written): core requirements sha256 43877af111c4, full harness prompt sha256 43877af111c4; exact prompt text used by the existing harness rows of this test on the runs branch (copied from game_battleship__claude-code__haiku-5-5__attempt1/meta.json, sha256 verified); Beta half of the MR matr
  • no TUI: the harness ran in its non-interactive mode, so there are no TUI screenshots; the timelapse is built from version renders only
  • OpenCode config (providers only, permission allow-all except question; no agents/instructions) + both default skill kits in xdg/config/opencode/skill (skills=on)
  • infra retries: 0
  • in the published transcript the sandbox HOME directory is shown as <HOME>
  • skills=on: 54 skills installed in the harness skill directory before launch (video-33 remotion-dev/skills + heygen-com/hyperframes@0b5db9da + f; web-threejs-21 [redacted]/ai-AI Model Reviewer skills/ (alpha/mr-homebar)@7521074d98c9).
  • Mode: headless run mode (opencode run --format json), inside bubblewrap sandbox
  • Isolation audit: none
  • Skills on: the run had 54 skills installed (33 video + 21 web/Three.js), as its meta records.
  • Qwen 3.8 Max served via Fireworks (OpenCode provider fireworks-ai).
  • Same frozen prompt as the other rows on this test; configurations differ and all rows are unscored, so this is not a head-to-head result.