AI Model Reviewer

Oh My Pi (omp) 18.3.1

1 model ran in it: Claude Opus 5.5.

7 runs on 6 prompts · 86% finished

Results

The best run of each of its 6 tests first, then every other run. Filter them in the explorer

6 finished runs

Numbers: finish rate, cost, wall time, draft scores
Runs
7
on 6 prompts
Finished
86%
6 of 7
Median cost, finished run
$7.36
n=6
Median wall time, finished run
33m 25s
n=6
Recorded spend
$80.32
$0 on unfinished runs; 1 without a recorded cost

From the numbers

Set against the other harnesses with at least 3 runs (top or bottom third), and its draft-score ranks within single prompts. Nothing here is averaged across prompts.

Strengths

  • Nothing stands out with enough runs to say so.

Weaknesses

  • 33m 25s (n=6): 10th of 12 harnesses with at least 3 (median wall time of a finished run)
  • $7.36 (n=6): 12th of 12 harnesses with at least 3 (median cost of a finished run)
  • Other unfinished runs, not counted against the model: 1 INFRA

Best and hardest run

Best run

Picked as its fastest finished run (no scored result yet).

Hardest run

Picked as its slowest finished run.

Draft scores by prompt

No scored result yet (only prompts with a published rubric are scored).

By model

Claude Opus 5.57686%$7.36 n=633m 25s n=6
All 6 tests