AI Model Reviewer

Claude Opus 5.5 in Claude Code 2.1.280

skills off

← Open in the gallerySame prompt, all models

Zoom, inline player and more ways to view
Open the original page

The model's own page, run in a sandboxed frame only when you ask (a page that needs its own origin opens in a new tab on a separate demo site). Phones get a full-screen player capped at 1.5x pixel density (2x on larger screens) so it stays smooth; the original page runs uncapped. It starts in its auto-tour (T switches to play mode on a keyboard).

OKSlime softbodyhigh effortNeutral harness

Wall time 55m 03sCost $4.70 estimated

Review

No verdict yet.

The full review

OK. Finished on its own after 55.1 min.

Attempt 2 of 2

Every try of this model in this harness on this prompt, oldest first. This run is the cell's result.

  1. Attempt 1 OK1m 18s$0.232026-09-30
  2. Attempt 2 OK this run55m 03s$4.702026-09-30

Scores

AI judge

Not scored

No rubric for this prompt

Visitor votes and reviewer grades are not available yet.

AI judge
AI judge: the test runner (model claude-opus-5-5), scoring by visual review of the 1600 px render against the pre-registered rubric. No AI-judge score exists for this result: there is no rubric for this prompt, so no score is shown and none is made up.
Visitors
Public voting with a blind reveal is coming: you score the output first, then see the other scores.
Reviewer
The reviewer's own grade appears here once it is added.

Scores compare results for the same prompt only. AI-judge boards rank runs within one prompt, equal totals sharing a rank; scores from different prompts are never averaged.

Downloads

Files actually published for this run. Original output, site-made previews and post-run conversions are labelled separately.

The exact prompt

as run, 624 characters. Every result for this prompt

Build a single-page Three.js scene (index.html, no build step, three.js from a pinned CDN or vendored) showing a glossy translucent slime/jelly blob resting on a floor. It must behave as a soft body: clicking or dragging pokes and squishes it, it wobbles and settles with visible jiggle and damping, and it casts a soft contact shadow on the floor. Include subsurface-like translucency, specular highlights, and a small UI to reset and to drop the blob from height. Target 60 fps on a mid-range phone; mobile touch must work. Write everything into the working directory and finish with a short README.md describing controls.

How it was made

Wall time 55m 03sFirst artifact at 13m 38s

Step by step

The prompt, every agent turn and tool call (tool name and a one-line summary; inputs and outputs are not shown), then the final message. The full transcript is 4.4 MB; it is published as a scrubbed, compressed download rather than shown on the page.

Show all 76 steps

Loading the timeline…

Download the full transcript Claude Code session JSONL, 157 KB, gzip. Scrubbed: anything that looks like a key, token, e-mail address or phone number is replaced, and account details and the test machine's file paths are removed.

Cost breakdown

Headline cost $4.70, estimated from the token counts at list price.

Tokens by typeTokens$ per 1MCost
Input (uncached)122$4.00$0.0005
Cache write180,714$5.00$0.9036
Cache read7,788,602$0.20$1.5577
Output111,764$20.00$2.2353
of which reasoning58,089billed as output
Total computed from tokens
$4.6971
Total reported by the harness
none

Computed from harness-reported token counts x list prices.

Prices: Project price table (vendor list prices), 2026-09-22.

Run notes

one prompt, agentic harness (self-review allowed)labelled re-run (of threejs_slime_softbody__claude-code__opus-5-5-high)

  • frozen prompt (written), used verbatim
  • plain wave-1-equivalent cc_config copy (settings.json + .claude.json; no ECC files)
  • Mode: interactive TUI in tmux (160x48), inside bubblewrap sandbox
  • Context setting: 1M ([1m] model name)
  • Isolation audit: none