AI Model Reviewer

Pi 0.87.1 + skills kit (ponytail, ECC, vetted skills)

1 model ran in it: Claude Opus 5.5.

16 runs on 13 prompts · 56% finished

Results

The best run of each of its 13 tests first, then every other run. Filter them in the explorer

9 finished runs

Numbers: finish rate, cost, wall time, draft scores
Runs
16
on 13 prompts
Finished
56%
9 of 16
Median cost, finished run
$2.75
n=9
Median wall time, finished run
16m 06s
n=9
Recorded spend
$83.90
$36.77 on unfinished runs

From the numbers

Set against the other harnesses with at least 3 runs (top or bottom third), and its draft-score ranks within single prompts. Nothing here is averaged across prompts.

Strengths

  • 16m 06s (n=9): 4th of 12 harnesses with at least 3 (median wall time of a finished run)

Weaknesses

  • 9 of 16 runs finished: 11th of 14 harnesses with at least 3 (finish rate)
  • 2 runs ended as the model's own failure (DNF, class MODEL)
  • Other unfinished runs, not counted against the model: 5 TIMEOUT

Best and hardest run

Best run

Picked as its fastest finished run (no scored result yet).

Hardest run

Picked as its costliest run that did not finish.

Draft scores by prompt

No scored result yet (only prompts with a published rubric are scored).

By model

Claude Opus 5.516956%$2.75 n=916m 06s n=9
All 13 tests