AI Model Reviewer

Hermes Agent 0.21.5 + Oh-My-Hermes 2.0.5

1 model ran in it: Claude Opus 5.5.

8 runs on 6 prompts · 0% finished

Results

The best run of each of its 6 tests first, then every other run. Filter them in the explorer

0 finished runs

    Numbers: finish rate, cost, wall time, draft scores
    Runs
    8
    on 6 prompts
    Finished
    0%
    0 of 8
    Median cost, finished run
    not recorded
    n=0
    Median wall time, finished run
    not recorded
    n=0
    Recorded spend
    $24.64
    $24.64 on unfinished runs; 7 without a recorded cost

    From the numbers

    Set against the other harnesses with at least 3 runs (top or bottom third), and its draft-score ranks within single prompts. Nothing here is averaged across prompts.

    Strengths

    • Nothing stands out with enough runs to say so.

    Weaknesses

    • 0 of 8 runs finished: 14th of 14 harnesses with at least 3 (finish rate)
    • 1 run ended as the model's own failure (DNF, class MODEL)
    • Other unfinished runs, not counted against the model: 1 INFRA, 6 TIMEOUT

    Best and hardest run

    Hardest run

    Picked as its costliest run that did not finish.

    Draft scores by prompt

    No scored result yet (only prompts with a published rubric are scored).

    By model

    Claude Opus 5.5800%not recorded: no finished run with a recorded costnot recorded: no finished run with a recorded time
    All 6 tests