GLM-5.3 in Claude Code 2.1.280
skills off
This is the unfinished output at the moment the run ended (Harness + model mismatch). See the review below.
Wall time 13m 43sCost $0.47 estimated
Review
No verdict yet.
The full review
Harness + model mismatch. Reclassified 2026-09-24 (published as ok): the run ended on 'API Error: 400 This model does not support image inputs' when GLM-5.3 read its own render (/tmp/chrome/preview1.png) with Claude Code's Read tool; the artifact was already written but the self-review was cut short (Claude Code x text-only model). Re-run on the text-only relay: brown_pelican_detailed__claude-code__glm-5.3__attempt2.
Attempt 1 of 2
Every try of this model in this harness on this prompt, oldest first. A later attempt is the cell's result: attempt 2.
- Attempt 1 Harness + model mismatch this run13m 43s$0.472026-09-24
- Attempt 2 OK the cell's result17m 09s$0.772026-09-25
Scores
Not scored: only runs that finished on their own are scored.
Downloads
Files actually published for this run. Original output, site-made previews and post-run conversions are labelled separately.
The exact prompt
agentic harness (save the file; self-render allowed), 561 characters. Every result for this prompt
Generate an SVG of a California brown pelican riding a bicycle. The bicycle must have spokes and a correctly shaped bicycle frame. The pelican must have its characteristic large pouch, and there should be a clear indication of feathers. The pelican must be clearly pedaling the bicycle. The image should show the full breeding plumage of the California brown pelican. Save it as brown_pelican.svg in the current directory. You may render and inspect your own output (headless chromium is available at /usr/bin/chromium) and iterate on your own before finishing.
How it was made
Step by step
The prompt, every agent turn and tool call (tool name and a one-line summary; inputs and outputs are not shown), then the final message. The full machine-format transcript is published as a scrubbed download.
Show all 15 steps
Loading the timeline…
Download the full transcript Claude Code session JSONL, 366 KB. Scrubbed: anything that looks like a key, token, e-mail address or phone number is replaced, and account details and the test machine's file paths are removed.
Cost breakdown
Headline cost $0.47, estimated from the token counts at list price.
| Tokens by type | Tokens | $ per 1M | Cost |
|---|---|---|---|
| Input (uncached) | 78,160 | $1.40 | $0.1094 |
| Cache write | 0 | $0.00 | $0 |
| Cache read | 392,431 | $0.26 | $0.1020 |
| Output | 58,579 | $4.40 | $0.2577 |
| of which reasoning | not recorded | billed as output |
- Total computed from tokens
- $0.4692
- Total reported by the harness
- none
Tokens summed from the Claude Code session JSONL (per message id, incl. subagent files); if Claude Code logged zero usage for this unrecognised model, the /cost screen's per-model usage is used instead. Claude Code tags glm-5p3 as unrecognized_model and its own $ figure has costBasis unknown, so it is NOT used: cost = tokens x Fireworks GLM-5.3 price.
Prices: Fireworks GLM-5.3 serverless price (STATE.md 22:42Z; same as the OpenCode/models.dev catalog entry fireworks-ai/glm-5p3), 2026-09-23.
Run notes
- A well-known benchmark author's detailed pelican prompt, verbatim from the image alt text of the post (his own comment in brackets after it left out).
- Mode: interactive TUI in tmux (160x48), inside bubblewrap sandbox
- Isolation: bwrap: private /tmp, HOME and TMPDIR inside run dir, own CLAUDE_CONFIG_DIR copy; the project files, other runs, /tmp/hx_* not mounted

