AI Model Reviewer

Claude Opus 5.5 in a raw API call

raw API baseline

← Open in the gallerySame prompt, all models

DNF

The model did not finish, so there is no output.

DNFAnimated short film drawn in code (HTML/JS)max effortRaw baseline

Wall time 20m 42sCost $2.56 estimated

Review

DNF. At max effort it spent its entire 128,000-token output ceiling thinking for 20.7 minutes and never wrote a line of HTML. $2.56 spent, nothing produced. Effort low made the whole film in 98 s for $0.24.

Verdict by site build agent (draft written from the run records and the render; not reviewed yet).

The full review

DNF. stop_reason max_tokens after 1,242 s: all 128,000 output tokens were thinking, 0 characters of HTML. The model's own output ceiling, not our timeout.

Attempt 1 of 3

Every try of this model in this harness on this prompt, oldest first. A later attempt is the cell's result: attempt 3.

  1. Attempt 1 DNF this run20m 42s$2.562026-09-23
  2. Attempt 2 DNF22m 33s$2.562026-09-25
  3. Attempt 3 DNF the cell's result21m 53s$2.562026-09-25

Scores

Not scored: only runs that finished on their own are scored.

Downloads

Files actually published for this run. Original output, site-made previews and post-run conversions are labelled separately.

The exact prompt

raw API call, 3,292 characters. Every result for this prompt

Write a single self-contained HTML file that plays an original animated short film, about 50 seconds long, drawn entirely in code.

Hard technical rules:
- One file: HTML + inline <script> (+ optional inline <style>). No external images, fonts, libraries, audio files or network requests of any kind. Every frame is drawn procedurally with Canvas 2D (or WebGL if you prefer).
- The canvas renders at a fixed internal resolution of 1920x1080 and is scaled with CSS to fit the window (letterboxed, black background).
- The animation starts automatically on page load and is driven ONLY by requestAnimationFrame and the timestamp from performance.now() (measure elapsed time from the first frame). Everything on screen must be a pure function of elapsed time t, so the film is deterministic and can be captured frame by frame. No Math.random() at draw time — use a seeded PRNG / hash noise so frames are identical run to run. No setTimeout/setInterval for animation.
- Define window.FILM_DURATION (seconds). After the film ends, hold the final frame.
- Soundtrack: write an original WebAudio score (melody, chords, a little percussion or ambience, and sound cues that match story moments) synced to the same timeline. Put ALL audio scheduling in a function window.buildSoundtrack(ctx, startTime) that schedules every note/effect on any BaseAudioContext (so it also works on an OfflineAudioContext). Also expose window.renderSoundtrack = () => an OfflineAudioContext(2, 44100*FILM_DURATION, 44100) render of it returning a Promise<AudioBuffer>. For live viewing, start the soundtrack on the first click/tap/keypress (browsers block autoplay) aligned to the current film time; show a small unobtrusive "tap for sound" hint until then. Use only oscillators, noise buffers you generate, filters, gain envelopes — no samples.
- No console errors.

The film (this is a showcase of what you can animate — make it genuinely beautiful and moving):
- An ORIGINAL story with a small emotional arc (setup, a problem or longing, a turn, a warm resolution). A single consistent main character — any character you like — who appears in every scene and is recognisably the same design throughout.
- At least 4 distinct scenes, with real transitions between them (e.g. wipes, iris, match cuts, cross-dissolves, camera moves through space).
- Expressive character animation: a proper walk cycle (legs, arms, body bob), blinking, and facial emotion changes (e.g. curious, sad, surprised, joyful) that follow the story.
- A virtual camera that pans, zooms and tracks, with parallax layers for depth.
- Lighting and time-of-day changes across the film (e.g. dawn → day → dusk → night, or a storm clearing), with colour grading, glows, and shadows.
- Particles and weather (rain, snow, fireflies, leaves, dust, stars — your choice, at least two kinds).
- On-screen story text / subtitles in a tasteful storybook style that tell the story in a few short lines.
- A cohesive art direction: a warm, hand-made "paper cut-out / storybook" look is welcome (textured paper, soft edges, slight wobble), but choose what serves your story.
- End with a title of your film and a gentle closing frame.

Output ONLY the complete HTML file, starting with <!DOCTYPE html> and ending with </html>. No explanation, no markdown fences.

How it was made

Wall time 20m 42s

Step by step

The prompt, every agent turn and tool call (tool name and a one-line summary; inputs and outputs are not shown), then the final message. The full transcript is published as text (under 2 MB, secrets redacted).

Show all 5 steps

Loading the timeline…

Download the full transcript text record, 4 KB. Scrubbed: anything that looks like a key, token, e-mail address or phone number is replaced, and account details and the test machine's file paths are removed.

Cost breakdown

Headline cost $2.56, estimated from the token counts at list price.

Tokens by typeTokens$ per 1MCost
Input (uncached)1,174$4.00$0.0047
Cache write0$5.00$0
Cache read0$0.20$0
Output128,000$20.00$2.5600
of which reasoning128,000billed as output
Total computed from tokens
$2.5647
Total reported by the harness
none

Raw API call: token counts from the response usage block x list price. Thinking/reasoning tokens are billed as output. No harness, so there is no reported total to compare.

Prices: Project price table (vendor list prices), 2026-09-22.

Run notes

raw one-shot (one call, no tools, no self-review)max efforthit max_tokens