AI Model Reviewer

GPT-6 Sol in a raw API call

raw API baseline

← Open in the gallerySame prompt, all models

Zoom, inline player and more ways to view
Poster frame of GPT-6 Sol's page
Open the original page

The model's own page, run in a sandboxed frame only when you ask (a page that needs its own origin opens in a new tab on a separate demo site). Phones get a full-screen player capped at 1.5x pixel density (2x on larger screens) so it stays smooth; the original page runs uncapped. Tap inside for sound.

OKAnimated short film drawn in code (HTML/JS)max effortRaw baseline

Wall time 15m 34sCost $0.59 estimated

Snapshot 25 of 25 saved during GPT-6 Sol's run
  1. Snapshot 1 of 25 saved during GPT-6 Sol's run
  2. Snapshot 4 of 25 saved during GPT-6 Sol's run
  3. Snapshot 8 of 25 saved during GPT-6 Sol's run
  4. Snapshot 11 of 25 saved during GPT-6 Sol's run
  5. Snapshot 15 of 25 saved during GPT-6 Sol's run
  6. Snapshot 18 of 25 saved during GPT-6 Sol's run
  7. Snapshot 22 of 25 saved during GPT-6 Sol's run
  8. Snapshot 25 of 25 saved during GPT-6 Sol's run
Still frames from a low-frame-rate recording made during the run (under 12 fps, so not shown as video): 25 distinct frames.

Review

No verdict yet.

The full review

OK. Finished on its own after 15.6 min.

Scores

AI judge

Not scored

No rubric for this prompt

Visitor votes and reviewer grades are not available yet.

AI judge
AI judge: the test runner (model claude-opus-5-5), scoring by visual review of the 1600 px render against the pre-registered rubric. No AI-judge score exists for this result: there is no rubric for this prompt, so no score is shown and none is made up.
Visitors
Public voting with a blind reveal is coming: you score the output first, then see the other scores.
Reviewer
The reviewer's own grade appears here once it is added.

Scores compare results for the same prompt only. AI-judge boards rank runs within one prompt, equal totals sharing a rank; scores from different prompts are never averaged.

Downloads

Files actually published for this run. Original output, site-made previews and post-run conversions are labelled separately.

The exact prompt

raw API call, 3,292 characters. Every result for this prompt

Write a single self-contained HTML file that plays an original animated short film, about 50 seconds long, drawn entirely in code.

Hard technical rules:
- One file: HTML + inline <script> (+ optional inline <style>). No external images, fonts, libraries, audio files or network requests of any kind. Every frame is drawn procedurally with Canvas 2D (or WebGL if you prefer).
- The canvas renders at a fixed internal resolution of 1920x1080 and is scaled with CSS to fit the window (letterboxed, black background).
- The animation starts automatically on page load and is driven ONLY by requestAnimationFrame and the timestamp from performance.now() (measure elapsed time from the first frame). Everything on screen must be a pure function of elapsed time t, so the film is deterministic and can be captured frame by frame. No Math.random() at draw time — use a seeded PRNG / hash noise so frames are identical run to run. No setTimeout/setInterval for animation.
- Define window.FILM_DURATION (seconds). After the film ends, hold the final frame.
- Soundtrack: write an original WebAudio score (melody, chords, a little percussion or ambience, and sound cues that match story moments) synced to the same timeline. Put ALL audio scheduling in a function window.buildSoundtrack(ctx, startTime) that schedules every note/effect on any BaseAudioContext (so it also works on an OfflineAudioContext). Also expose window.renderSoundtrack = () => an OfflineAudioContext(2, 44100*FILM_DURATION, 44100) render of it returning a Promise<AudioBuffer>. For live viewing, start the soundtrack on the first click/tap/keypress (browsers block autoplay) aligned to the current film time; show a small unobtrusive "tap for sound" hint until then. Use only oscillators, noise buffers you generate, filters, gain envelopes — no samples.
- No console errors.

The film (this is a showcase of what you can animate — make it genuinely beautiful and moving):
- An ORIGINAL story with a small emotional arc (setup, a problem or longing, a turn, a warm resolution). A single consistent main character — any character you like — who appears in every scene and is recognisably the same design throughout.
- At least 4 distinct scenes, with real transitions between them (e.g. wipes, iris, match cuts, cross-dissolves, camera moves through space).
- Expressive character animation: a proper walk cycle (legs, arms, body bob), blinking, and facial emotion changes (e.g. curious, sad, surprised, joyful) that follow the story.
- A virtual camera that pans, zooms and tracks, with parallax layers for depth.
- Lighting and time-of-day changes across the film (e.g. dawn → day → dusk → night, or a storm clearing), with colour grading, glows, and shadows.
- Particles and weather (rain, snow, fireflies, leaves, dust, stars — your choice, at least two kinds).
- On-screen story text / subtitles in a tasteful storybook style that tell the story in a few short lines.
- A cohesive art direction: a warm, hand-made "paper cut-out / storybook" look is welcome (textured paper, soft edges, slight wobble), but choose what serves your story.
- End with a title of your film and a gentle closing frame.

Output ONLY the complete HTML file, starting with <!DOCTYPE html> and ending with </html>. No explanation, no markdown fences.

How it was made

Wall time 15m 34sFirst artifact at 15m 23s1 saved version

Versions saved during the run

1 intermediate output, in the order the model wrote them.

Version 1 of 1
  1. Version 1 at 15m 25s thumbnailv1 15:25

Step by step

The prompt, every agent turn and tool call (tool name and a one-line summary; inputs and outputs are not shown), then the final message. The full machine-format transcript is published as a scrubbed download.

Show all 3 steps

Loading the timeline…

Download the full transcript raw API response, 17 KB, gzip. Scrubbed: anything that looks like a key, token, e-mail address or phone number is replaced, and account details and the test machine's file paths are removed.

Cost breakdown

Headline cost $0.59, estimated from the token counts at list price.

Tokens by typeTokens$ per 1MCost
Input (uncached)741$2.00$0.0015
Cache write0n/a$0
Cache read0$0.20$0
Output58,854$10.00$0.5885
of which reasoning36,881billed as output
Total computed from tokens
$0.5900
Total reported by the harness
none

Computed from harness-reported token counts x list prices. No request above 272,000 input tokens (1 checked): standard rates.

Prices: Project price table, 2026-09-25.

Run notes

raw one-shot (one call, no tools, no self-review)

  • frozen prompt (verbatim): results.json: the text all 2 published raw-API runs of js_short used
  • raw API baseline: one call, no tools, no self-review
  • raw response JSON (model, stop_reason, usage, reasoning summary, content)
  • Mode: single API call (raw baseline, no harness)
  • Isolation audit: none