AI Model Reviewer

Kimi K3 in Kimi Code CLI 2.1.0

skills off

← Open in the gallerySame prompt, all models

Zoom, inline player and more ways to view
Kimi K3's SVG for MacBook Pro 16" realistic 3D render (Playcode benchmark)

OKMacBook Pro 16" realistic 3D render (Playcode benchmark)max effortHome harness

Wall time 27m 44sCost $1.68 estimated8.0/10 · Draft score — stored total; criterion breakdown unavailable

Review

Solid MacBook Pro in three-quarter view on a dark backdrop: notch, a vivid wallpaper with a window and dock, blue light spill, speaker grilles, a big trackpad, side ports and a front notch.

The full review

Solid MacBook Pro in three-quarter view on a dark backdrop: notch, a vivid wallpaper with a window and dock, blue light spill, speaker grilles, a big trackpad, side ports and a front notch. The keyboard is a flat unlabelled grid, the trackpad sits off-centre toward the right, and the aluminium shading is flat, with little edge highlight.

Verdict by the AI judge (provisional).

OK. Finished on its own after 27.7 min and produced macbook.svg.

Bugs

  • keyboard is a flat grid
  • trackpad off-centre
  • flat aluminium

Scores

AI judge

8.0/10

Draft score — stored total; criterion breakdown unavailable

Visitor votes and reviewer grades are not available yet.

Draft score — stored total; criterion breakdown unavailable. the per-criterion marks could not be read, so the stored score is shown; the displayed total has not been verified through a criterion breakdown.

AI judge
AI judge: the test runner (model claude-opus-5-5), scoring by visual review of the 1600 px render against the pre-registered rubric. A draft score until the reviewer checks it. Rubric: Wave-1 SVG scoring rubric (AI Model Reviewer).
Visitors
Public voting with a blind reveal is coming: you score the output first, then see the other scores.
Reviewer
The reviewer's own grade appears here once it is added.

Scores compare results for the same prompt only. AI-judge boards rank runs within one prompt, equal totals sharing a rank; scores from different prompts are never averaged.

Downloads

Files actually published for this run. Original output, site-made previews and post-run conversions are labelled separately.

The exact prompt

agentic harness (save the file; self-render allowed), 373 characters. Every result for this prompt

Generate an SVG that looks like a realistic 3D render of a MacBook Pro 16" (opened, slight three-quarter view, screen glow, aluminum body). Use gradients and perspective to fake the 3D. Save it as macbook.svg in the current directory. You may render and inspect your own output (headless chromium is available at /usr/bin/chromium) and iterate on your own before finishing.

How it was made

Wall time 27m 44sFirst artifact at 15m 43s5 saved versions

Versions saved during the run

5 intermediate outputs, in the order the model wrote them.

Version 5 of 5
  1. Version 1 at 15m 40s thumbnailv1 15:40
  2. Version 2 at 19m 01s thumbnailv2 19:01
  3. Version 3 at 22m 33s thumbnailv3 22:33
  4. Version 4 at 25m 54s thumbnailv4 25:54
  5. Version 5 at 26m 46s thumbnailv5 26:46

Step by step

The prompt, every agent turn and tool call (tool name and a one-line summary; inputs and outputs are not shown), then the final message. The full transcript is 2.6 MB; it is published as a scrubbed, compressed download rather than shown on the page.

Show all 2 steps

Loading the timeline…

Download the full transcript transcript, 151 KB, gzip. Scrubbed: anything that looks like a key, token, e-mail address or phone number is replaced, and account details and the test machine's file paths are removed.

Cost breakdown

Headline cost $1.68, estimated from the token counts at list price.

Tokens by typeTokens$ per 1MCost
Input (uncached)82,919$3.00$0.2488
Cache write0$0.00$0
Cache read2,338,634$0.30$0.7016
Output48,318$15.00$0.7248
of which reasoningnot recordedbilled as output
Total computed from tokens
$1.6751
Total reported by the harness
none

Tokens summed from every step.end usage in Kimi Code's wire.jsonl (main agent + any subagent wires): inputOther = uncached input, inputCacheRead, inputCacheCreation, output (reasoning is included in output, not split). Kimi Code prints no dollar figure: cost = tokens x Moonshot list price.

Prices: models.dev catalog as loaded by OpenCode 1.18.32 (read 2026-09-23) (Moonshot kimi-k3 list price), 2026-09-23.

Run notes

one prompt, agentic harness (self-review allowed)wave 1

  • Playcode MacBook SVG Benchmark prompt verbatim, except its last sentence ("Respond with ONLY the SVG markup, no explanation.") is swapped for the save-to-file harness suffix, as the prompt research recommends for harness runs.
  • Mode: interactive TUI in tmux (160x48), inside bubblewrap sandbox
  • Isolation: bwrap: private /tmp, /run, PID/IPC/UTS ns; HOME, TMPDIR and the harness home (sessions, config copy) inside this run dir; only /usr, /etc, the harness binary dir(s) (ro) and this run dir bound; the project files, results/, other runs, /tmp/hx_* not mounted