Pi 0.87.1
1 model ran in it: Claude Opus 5.5.
8 runs on 8 prompts · 88% finished
Results
The best run of each of its 8 tests first, then every other run. Filter them in the explorer
7 finished runs
Numbers: finish rate, cost, wall time, draft scores
- Runs
- 8
- on 8 prompts
- Finished
- 88%
- 7 of 8
- Median cost, finished run
- $4.89
- n=7
- Median wall time, finished run
- 29m 04s
- n=7
- Recorded spend
- $59.11
- $7.01 on unfinished runs
From the numbers
Set against the other harnesses with at least 3 runs (top or bottom third), and its draft-score ranks within single prompts. Nothing here is averaged across prompts.
Strengths
- 7 of 8 runs finished: 4th of 14 harnesses with at least 3 (finish rate)
Weaknesses
- 29m 04s (n=7): 9th of 12 harnesses with at least 3 (median wall time of a finished run)
- $4.89 (n=7): 10th of 12 harnesses with at least 3 (median cost of a finished run)
- Other unfinished runs, not counted against the model: 1 TIMEOUT
Best and hardest run
Best run
Picked as its fastest finished run (no scored result yet).
-
Claude Opus 5.5, Pi 0.87.1, remotion_sales_chart
Pi 0.87.1, max effort, via Anthropic
skills off
22m 36s
$4.12 reported by Pi 0.87.1 est. $4.12
OK
Unscored: no rubric for this prompt
Hardest run
Picked as its costliest run that did not finish.
-
Did not finish · re-run queued
Claude Opus 5.5, Pi 0.87.1, Animated short (HTML/JS)
Pi 0.87.1, max effort, via Anthropic
skills off
Stopped by us
41m 34s $7.01 reported by Pi 0.87.1 est. $7.01
Draft scores by prompt
No scored result yet (only prompts with a published rubric are scored).
By model
| Claude Opus 5.5 | 8 | 7 | 88% | $4.89 n=7 | 29m 04s n=7 |
|---|
All 8 tests
- Artemis II (1 run)
- Sales chart (1 run)
- California brown pelican riding a bicycle (detailed spec) (1 run)
- Dolphin vlog bicycle (1 run)
- Pelican car france (1 run)
- Pelican riding a bicycle (control) (1 run)
- Animated pelican riding a bicycle (SMIL/CSS) (1 run)
- Animated short (HTML/JS) (1 run)