Models & harnesses
Browse by model or by harness: each page opens on its best run per test, then every other run.
Models
- Claude Opus 5.5 Anthropic · 354 runs
- Claude Sonnet 5.5 Anthropic · 84 runs
- Claude Fable 5.1 Anthropic · 181 runs
- Claude Haiku 5.5 Anthropic · 7 runs
- GPT-6 Astra OpenAI · 264 runs
- DeepSeek V4.1 Flash DeepSeek · 248 runs
- DeepSeek V4 Pro DeepSeek · 171 runs
- Grok 4.7 xAI · 10 runs
- Gemini 3.8 Flash Google · 16 runs
- Kimi K3 Moonshot · 106 runs
- GLM-5.3 Z.ai · 110 runs
- GPT-6 Sol OpenAI · 238 runs
- Space Bunny Alpha Undisclosed (stealth) · 24 runs
- Muse Spark 1.3 Meta · 24 runs
- LongCat 2.5 Preview Meituan · 52 runs
- MiMo V2.6 Flash Xiaomi · 16 runs
- Claude Sonnet 5 Anthropic · 99 runs
- GPT-6 Luna OpenAI · 134 runs
- Gemini 3.6 Flash Google · 11 runs
- Gemini 3.7 Flash Google · 2 runs
- MiniMax M3 MiniMax (via Fireworks) · 64 runs
- Qwen3.8 27B Alibaba Qwen (via Cerebras) · 340 runs
- Qwen3.8 Max Alibaba Qwen (via Fireworks) · 71 runs
- gemini-3.5-flash Google · 2 runs
- gpt-6.1-sol OpenAI · 31 runs
- gpt-oss-120b OpenAI open-weight (via Cerebras) · 354 runs
Harnesses
- Raw API call 2,097 runs
- OpenCode 307 runs
- Claude Code 260 runs
- Codex CLI 186 runs
- Cline CLI 67 runs
- dsh 22 runs
- Pi 0.87.1 + skills kit (ponytail, ECC, vetted skills) 16 runs
- Kimi Code CLI 15 runs
- Hermes Agent 0.21.5 10 runs
- Pi 0.87.1 8 runs
- Hermes Agent 0.21.5 + Oh-My-Hermes 2.0.5 8 runs
- Oh My Pi (omp) 18.3.1 7 runs
- Kilo Code CLI 7.7.9 6 runs
- Grok Build 4 runs
All models, in numbers
Show the table (26 rows)
| Claude Opus 5.5Anthropic | 354 | 80% 283 of 354 | $2.37 n=275 | 14m 39s n=283 | $1,853.09 |
|---|---|---|---|---|---|
| Claude Sonnet 5.5Anthropic | 84 | 64% 54 of 84 | $6.65 n=10 | 26m 03s n=54 | $177.92 |
| Claude Fable 5.1Anthropic | 181 | 89% 161 of 181 | $0.97 n=161 | 4m 02s n=161 | $1,001.68 |
| Claude Haiku 5.5Anthropic | 7 | 100% 7 of 7 | not recorded | 7m 56s n=7 | $0 |
| GPT-6 AstraOpenAI | 264 | 98% 260 of 264 | $0.87 n=258 | 5m 14s n=260 | $743.43 |
| DeepSeek V4.1 FlashDeepSeek | 248 | 88% 219 of 248 | $0.12 n=61 | 1m 16s n=219 | $8.86 |
| DeepSeek V4 ProDeepSeek | 171 | 98% 167 of 171 | $0.052 n=1 | 2m 23s n=167 | $0.052 |
| Grok 4.7xAI | 10 | 100% 10 of 10 | $6.29 n=10 | 36m 49s n=10 | $65.07 |
| Gemini 3.8 FlashGoogle | 16 | 50% 8 of 16 | $0 n=3 | 9m 44s n=8 | $0.073 |
| Kimi K3Moonshot | 106 | 90% 95 of 106 | $1.28 n=30 | 11m 50s n=95 | $67.64 |
| GLM-5.3Z.ai | 110 | 41% 45 of 110 | $1.08 n=31 | 17m 09s n=45 | $73.00 |
| GPT-6 SolOpenAI | 238 | 100% 238 of 238 | $1.31 n=104 | 2m 25s n=238 | $153.40 |
| Space Bunny AlphaUndisclosed (stealth) | 24 | 96% 23 of 24 | $0 n=22 | 5m 43s n=23 | $0 |
| Muse Spark 1.3Meta | 24 | 96% 23 of 24 | $0 n=18 | 5m 20s n=23 | $0 |
| LongCat 2.5 PreviewMeituan | 52 | 81% 42 of 52 | $0 n=42 | 9m 18s n=42 | $0 |
| MiMo V2.6 FlashXiaomi | 16 | 12% 2 of 16 | $0 n=2 | 49m 26s n=2 | $0 |
| Claude Sonnet 5Anthropic | 99 | 98% 97 of 99 | not recorded | 2m 17s n=97 | $0 |
| GPT-6 LunaOpenAI | 134 | 100% 134 of 134 | not recorded | 1m 10s n=134 | $0 |
| Gemini 3.6 FlashGoogle | 11 | 27% 3 of 11 | $0 n=3 | 3m 12s n=3 | $0 |
| Gemini 3.7 FlashGoogle | 2 | 0% 0 of 2 | not recorded | not recorded | $0 |
| MiniMax M3MiniMax (via Fireworks) | 64 | 75% 48 of 64 | not recorded | 2m 39s n=48 | $0 |
| Qwen3.8 27BAlibaba Qwen (via Cerebras) | 340 | 44% 150 of 340 | not recorded | 26 s n=150 | $0 |
| Qwen3.8 MaxAlibaba Qwen (via Fireworks) | 71 | 83% 59 of 71 | not recorded | 3m 56s n=59 | $0 |
| gemini-3.5-flashGoogle | 2 | 50% 1 of 2 | not recorded | 2m 15s n=1 | $0 |
| gpt-6.1-solOpenAI | 31 | 94% 29 of 31 | not recorded | 19m 46s n=29 | $0 |
| gpt-oss-120bOpenAI open-weight (via Cerebras) | 354 | 92% 324 of 354 | not recorded | 8.5 s n=324 | $0 |
Harnesses
Show the table (14 rows)
| Raw API callone call, no harness | 2097 | 83% 1744 of 2097 | $0.59 n=371 | 1m 34s n=1744 | $274.74 |
|---|---|---|---|---|---|
| OpenCodeneutral harness | 307 | 80% 247 of 307 | $1.06 n=246 | 19m 49s n=247 | $1,281.79 |
| Claude Codethe vendor's own harness | 260 | 80% 208 of 260 | $4.57 n=161 | 26m 43s n=208 | $1,570.71 |
| Codex CLIthe vendor's own harness | 186 | 99% 184 of 186 | $1.80 n=166 | 14m 29s n=184 | $601.04 |
| Cline CLIneutral harness | 67 | 52% 35 of 67 | $0 n=23 | 5m 34s n=35 | $0 |
| dshthe vendor's own harness | 22 | 82% 18 of 22 | $0.14 n=18 | 16m 40s n=18 | $3.06 |
| Pi 0.87.1 + skills kit (ponytail, ECC, vetted skills)neutral harness | 16 | 56% 9 of 16 | $2.75 n=9 | 16m 06s n=9 | $83.90 |
| Kimi Code CLIthe vendor's own harness | 15 | 93% 14 of 15 | $1.30 n=14 | 22m 50s n=14 | $35.36 |
| Hermes Agent 0.21.5neutral harness | 10 | 10% 1 of 10 | $23.83 n=1 | 2h 30m n=1 | $50.71 |
| Pi 0.87.1neutral harness | 8 | 88% 7 of 8 | $4.89 n=7 | 29m 04s n=7 | $59.11 |
| Hermes Agent 0.21.5 + Oh-My-Hermes 2.0.5neutral harness | 8 | 0% 0 of 8 | not recorded | not recorded | $24.64 |
| Oh My Pi (omp) 18.3.1neutral harness | 7 | 86% 6 of 7 | $7.36 n=6 | 33m 25s n=6 | $80.32 |
| Kilo Code CLI 7.7.9neutral harness | 6 | 83% 5 of 6 | $6.83 n=5 | 33m 50s n=5 | $59.17 |
| Grok Buildthe vendor's own harness | 4 | 100% 4 of 4 | $4.50 n=4 | 38m 25s n=4 | $19.64 |
Every attempt counts, failures included; runs lost to infrastructure are not results. Cost against quality · Leaderboards