AI Model Reviewer

Harnesses

Which agentic harness each model family should run in, and which harness is fair for everyone. The research behind the home and neutral tracks, researched 2026-09-23.

Scope: which harness to use per model vendor, and which harnesses can run models from several vendors, for the models reachable from our test machine (Anthropic, OpenAI, DeepSeek, xAI, Google, Moonshot, Cerebras and Fireworks). Method: web search plus pages fetched and read in this session; versions pulled live from the npm registry and GitHub API on 2026-09-23. No harness was installed or test-run for this note. Anything marked (unverified) comes from a snippet or a single secondary source.

1. The headline evidence (who wins per model family)

Terminal-Bench 4.0, official board (tbench.ai, read from the page's own data 2026-09-23). It only lists first-party harnesses plus mini-SWE-agent:

HarnessModel (effort)Resolution
CodexGPT-6 Astra (max)58.2%
CodexGPT-6 Astra (xhigh / high)57.9%
Claude CodeFable 5.1 (xhigh / max)57.9%
Claude CodeOpus 5 (xhigh)53.9%
Claude CodeGLM-5.3 (max)41.8%
Grok BuildGrok 4.7 (xhigh)37.6%
CodexGPT-5.6 Sol (max)37.3%
mini-SWE-agentGemini 3.8 Flash (high)19.1%
Claude CodeSonnet 5 (max)12.4%

(Opus 5.5 isn't on the official TB4 board yet.)

Artificial Analysis Coding Agent Index v1.5 (artificialanalysis.ai/agents/coding-agents; DeepSWE v1.1 + Terminal-Bench 4.0 + SWE-Atlas-QnA, 3 attempts each; read from the page's data):

Harness – modelIndexTB4
Claude Code – Opus 5.5 (max)66.063.1
Claude Code – Fable 5.1 (max)62.257.6
Codex – GPT-6 Astra (max)61.755.6
Devin Fusion CLI – Fable 5.1 xhigh + SWE-261.756.1
Claude Code – Opus 5 (max)59.754.5
Devin Fusion CLI – GPT-6 Astra xhigh + SWE-258.950.0
Codex – GPT-6 Sol (max)56.643.4
Grok Build – Grok 4.7 (xhigh)56.333.3
OpenCode – GLM-5.3 (max)53.639.9
Kimi Code CLI – Kimi K351.921.2
Claude Code – Qwen3.8 Max43.316.7
Codex – DeepSeek V4 Pro 0813 (max)43.010.1
Codex – GPT-6 Luna (max)41.115.2
Codex – DeepSeek V4 Flash 0731 (max)38.710.6

AA's own "Harness Comparison" tab says "Coming soon", so there's no clean same-model, many-harness table yet.

What the numbers say:

  • The first-party harness wins where we can compare. GPT-6 Astra: Codex 61.7 vs Devin Fusion 58.9. Fable 5.1: Claude Code 62.2 vs Devin Fusion 61.7.
  • Codex is a bad home for DeepSeek. DeepSeek V4 Pro/Flash in Codex score 10–11% on TB4 (AA). DeepSeek reports 31.2% on TB4 for V4.1 Flash in its own harness (vendor-reported, not on the official board; via mightybot.ai).
  • GLM-5.3 does about as well in Claude Code (41.8%, official TB4) as in OpenCode (39.9%, AA's TB4).
  • This hasn't always held. On the older Terminal-Bench 2.0, third-party harnesses sometimes beat the vendor's: Factory Droid 64.9% vs Codex CLI 62.9% on GPT-5.2 (reddit r/codex), and Letta Code 59.1% vs Claude Code 41.6% on Opus 4.5 (winder.ai). The arXiv position paper "Stop Comparing LLM Agents Without Disclosing the Harness" (2605.23950, May 2026) finds that on SWE-bench Pro, changing the harness around one model moved its score twice as much as changing the model in a fixed harness (Opus 4.5: 45.9% SEAL vs 55.4% Claude Code).
  • DoltHub's practitioner review (Aug 5, 2026) sums up the trend: "The models were trained to use their own harnesses… so they just worked better with their native harnesses."

2. Recommendation per vendor family (our models)

Vendor (our models)Best harnessWhyRunner-up
Anthropic (opus-5-5, fable-5-1, opus-5, sonnet-5, haiku-4-5)Claude Code 2.1.280 on npm (GitHub tag 2.1.281 today)Top of AA's index (Opus 5.5 66.0) and joint top of TB4 (Fable 5.1 57.9%)OpenCode (neutral)
OpenAI (gpt-6-astra/sol/luna, gpt-5.5, 5.4-mini/nano)Codex CLI 0.156.1#1 on TB4 (Astra max 58.2%); beats Devin Fusion on the same model in AA. I found no 2026 result where another harness beats Codex on GPT-6Factory Droid (droid exec, closed, BYOK over the Responses API; beat Codex on GPT-5.2 in TB2.0). OpenCode for neutral runs
DeepSeek (deepseek-v4-pro, deepseek-flash)DeepSeek Harness (dsh): npm 0.1.5-rc.3, GitHub tag dsh-v0.1.7-rc.1 (Sep 23)Vendor reports 31.2% TB4 in its own harness, vs about 10% for DeepSeek in Codex (AA). Developer preview with "breaking changes" warnings, web-UI-firstClaude Code via DeepSeek's official Anthropic endpoint; OpenCode (built-in DeepSeek provider)
xAI (grok-4.7, grok-4.6, grok-build-0.1)Grok Build (open-sourced under Apache-2.0 in July, v1.0 on Aug 7)xAI's own harness; the only Grok entries on TB4/AA are in it (Grok 4.7 37.6% / index 56.3); headless grok -p --output-format streaming-json; reads XAI_API_KEY and custom base_urlOpenCode; Codex also works (xAI's recommended API is Responses)
Google (gemini flash / flash-lite / pro)Gemini CLI (v0.60.0, Sep 15) with an API keyGoogle shut Gemini CLI off for free/Pro/Ultra accounts on Jun 18, 2026 and replaced it with the closed Antigravity CLI (agy). Gemini CLI still works with paid Gemini API keys, and GOOGLE_GEMINI_BASE_URL lets it point at a custom endpoint. agy is OAuth-first (API-key auth is an open feature request, issue #78), so it's awkward on a headless boxOpenCode / Qwen Code (the open Gemini-CLI fork). No strong harness result exists for Gemini: the only TB4 entry is in the minimal mini-SWE-agent (19.1%)
Moonshot (kimi-k3, kimi-k2.7-code)Kimi Code CLI 2.1.0 (released today, MIT)Moonshot's own harness, on AA's index (K3 51.9). Non-interactive -p mode (docs)Claude Code via Moonshot's official api.moonshot.ai/anthropic endpoint (Moonshot documents it)
Fireworks / Cerebras open models (glm-5.3, minimax-m3, kimi-k3-fast, qwen3p8-max, gpt-oss-120b, qwen-3.8-27b)OpenCodeBuilt-in Fireworks and Cerebras providers; AA runs GLM-5.3 in OpenCode (53.6). GLM-5.3 also runs well in Claude Code (41.8% TB4) but needs an Anthropic endpoint for it. Z.AI's own ZCode is a closed desktop app, not a CLIQwen Code for the Qwen models; Claude Code for GLM where there's an Anthropic endpoint

3. Best multi-vendor harness for fair same-harness comparisons: OpenCode

  • Open source and active: MIT, about 210k GitHub stars, v1.18.32 released Sep 21, 2026. The repo moved from sst/opencode to anomalyco/opencode.
  • Headless: opencode run "<prompt>" -m provider/model --format json streams raw JSON events. opencode serve plus run --attach avoids MCP cold starts. opencode session stats gives token and cost stats.
  • Speaks all three of our API formats natively:
    • Anthropic Messages (@ai-sdk/anthropic)
    • OpenAI Responses (@ai-sdk/openai; the docs say to use it for /v1/responses)
    • Google native (@ai-sdk/google)
    • OpenAI-compatible chat (@ai-sdk/openai-compatible) for everything else.
  • Custom endpoints: a baseURL per provider, so each provider can point at a custom endpoint.
  • Built-in providers for every vendor on our list: Anthropic, OpenAI, Google, DeepSeek, xAI, Moonshot, Cerebras, Fireworks, Groq, Z.AI and MiniMax, plus 75+ via Models.dev.
  • It's what benchmarkers use: Artificial Analysis already uses it as a neutral harness. It's the most-used provider-agnostic harness in 2026 roundups (pinggy.io, winder.ai, nimbalyst, mightybot).
  • Caveats:
    • Anthropic forced it to drop Claude Pro/Max subscription login in January 2026. API keys still work, which is our setup.
    • Before v1.2.23 it sent session titles to a hosted small model. That's fixed.
    • A neutral harness still penalises models tuned for their own harness. So run two tracks and label them: "home harness" (each model in its vendor's CLI) and "neutral harness" (every model in OpenCode, same config). The gap between the two is itself a QA finding.

Close alternatives for the neutral track:

  • Pi (earendil-works/pi, MIT, about 109k stars, v0.87.1 Sep 22): a tiny, token-efficient loop with little harness bias. Unified OpenAI/Anthropic/Google provider layer; print/JSON mode.
  • Goose (aaif-goose/goose, Apache-2.0, now under the Linux Foundation, v1.52.0 today): 15+ providers, goose run headless.
  • Crush (Charm, FSL-1.1-MIT, v0.96.1): OpenAI, openai-compat, Anthropic and Gemini providers. crush run headless (always auto-approve). Pre-1.0.
  • Qwen Code (Apache-2.0, npm 0.24.4): the Gemini-CLI fork; configures OpenAI, Anthropic and Gemini APIs.
  • OpenHands (MIT, v1.23.0 today): openhands --headless -t "…" --json, always auto-approve; LiteLLM underneath, so any provider.

4. Harness-by-harness facts

HarnessOpen sourceHeadlessAPI formats / providersNotes / first-party pairing
Claude Code (Anthropic)No (source-available repo, custom licence)claude -pAnthropic Messages only; ANTHROPIC_BASE_URL for gatewaysAnthropic's home harness. Non-Claude models work only where the vendor exposes an Anthropic-format endpoint: DeepSeek api.deepseek.com/anthropic (official), Moonshot api.moonshot.ai/anthropic (official), Z.AI/MiniMax/Qwen (community guide cc-compatible-models). OpenAI, Gemini and Cerebras need a translation layer (LiteLLM, claude-code-router). Anthropic's gateway docs: it "doesn't support routing Claude Code to non-Claude models through any gateway".
Codex CLI (OpenAI)Yes, Apache-2.0, about 126k starscodex execResponses API only: the config reference says wire_api "responses is the only supported value". Chat-completions providers are outOpenAI's home harness. Also runs xAI (Responses is xAI's recommended API) and DeepSeek (official Responses endpoint, stateless: no previous_response_id/store/background), but DeepSeek scores poorly in it (AA). Anthropic, Gemini, Cerebras and Moonshot need a gateway that translates to Responses (e.g. LiteLLM).
OpenCode (Anomaly, ex-SST)Yes, MITopencode run --format jsonAnthropic, OpenAI Responses, Google, OpenAI-compat; 75+ providersBest neutral harness (section 3).
Gemini CLI (Google)Yes, Apache-2.0gemini -pGemini API only; GOOGLE_GEMINI_BASE_URLRetired for consumer accounts Jun 18, 2026; paid API keys still work.
Antigravity CLI agy (Google)No (closed Go)agy -p (print mode, stream-JSON)Google models; OAuth-firstGemini CLI's successor; API-key auth not yet supported (issue #78).
Qwen Code (Alibaba)Yes, Apache-2.0-p (inherited from Gemini CLI)OpenAI, Anthropic, Gemini, Qwen APIs; any base URLHome harness for Qwen models.
Kimi Code CLI (Moonshot)Yes, MITkimi -p (non-interactive, auto permissions)Kimi first; "other compatible providers" configurableMoonshot's home harness; v2.1.0 today; desktop client Sep 17.
Grok Build (xAI)Yes, Apache-2.0 (open-sourced Jul 2026)grok -p, --output-format streaming-json; ACPxAI API key; any custom model via ~/.grok/config.toml base_url + env_keyxAI's home harness; ports tool code from Codex and OpenCode (its NOTICE). Cursor CLI is a separate product from the same camp.
DeepSeek Harness dshYes, MIT, about 234k starsWeb UI first (dsh web --no-open); headless also mentioned (pinggy). Exact CLI flag (unverified)Multi-provider; community reports it works with "almost all major providers" (community report); can drive Claude Code/Codex as sub-agents (winder.ai)DeepSeek's home harness; developer preview, breaking changes promised.
Crush (Charm)Source-available FSL-1.1 → MIT after 2 yrscrush runOpenAI, openai-compat, Anthropic, Gemini, Bedrock, OpenRouter…Best-looking TUI (good for screen recordings); switches models mid-session.
Goose (Block → Linux Foundation)Yes, Apache-2.0goose run -t15+ providers, plus ACP to subscriptionsGeneral agent; MCP-heavy.
AiderYes, Apache-2.0aider --messageAny (LiteLLM)No commits since May 22, 2026; last tagged release v0.86.0 (Aug 2025). Stalled; avoid.
Cline CLIYes, Apache-2.0 (npm cline 3.0.64)Has a CLI; headless flags (unverified)Model-agnosticActive.
Roo CodeArchived May 15, 2026——Dead; its users moved to Kilo Code.
Kilo Code CLIYes, MIT (npm @kilocode/cli 7.7.9)Yes (built on OpenCode)Same breadth as OpenCodeRoo/Cline fork; CLI is OpenCode underneath, so it adds nothing for us.
PiYes, MITprint / JSON / RPC modesOpenAI, Anthropic, Google, …Lean, low-bias harness.
OpenHandsYes, MITopenhands --headless -t … --jsonAny via LiteLLMSandboxed, CI-oriented.
Factory DroidNodroid execBYOK: anthropic (Messages), openai (Responses), generic-chat-completion-api (chat), each with a custom baseUrlStrong on OpenAI models historically (TB2.0 GPT-5.2: 64.9% vs Codex 62.9%). npm @factory/cli 0.225.2. Account needed.
Cursor CLI (agent)NoPrint modeCursor-hosted models; the CLI doesn't accept your own API key or endpoint (Cursor forum)Not usable with custom endpoints.
Amp (Sourcegraph)NoExecute modeHosted only; the mode picks the model (e.g. Astra, Fable 5.1, GLM-5.3-Flash)You can't pin one model, so it's unsuitable for per-model QA.
Warp Agent CLIClient AGPL-3.0YesWarp-hosted modelsNot a fit for BYOK testing.
Devin Fusion CLI (Cognition)No—Pairs a frontier model with SWE-2On AA's index, but it's a model blend, not a single-model harness.

5. What works with each vendor's API (per harness × vendor)

✅ = native format, ⚠️ = possible via the vendor's alternate endpoint or unverified, ❌ = needs a translation gateway.

AnthropicOpenAIDeepSeekxAIGeminiMoonshotCerebrasFireworks
Claude Code✅❌✅ (/anthropic)❌❌✅ (/anthropic)❌⚠️
Codex CLI❌✅✅ (Responses, but scores poorly)✅ (Responses)❌❌❌⚠️
OpenCode✅✅✅✅✅✅✅✅
Qwen Code✅✅✅✅✅✅✅✅
Crush / Goose / Pi / OpenHands✅✅✅✅✅✅✅✅
Grok Build⚠️⚠️⚠️✅⚠️⚠️⚠️⚠️
Kimi Code⚠️⚠️⚠️⚠️⚠️✅⚠️⚠️
Gemini CLI❌❌❌❌✅❌❌❌
Factory Droid (BYOK)✅✅✅✅⚠️✅✅✅

Claude Code quirks with third-party Anthropic endpoints:

  • Map every model tier explicitly: ANTHROPIC_MODEL, ANTHROPIC_DEFAULT_OPUS_MODEL / SONNET / HAIKU, CLAUDE_CODE_SUBAGENT_MODEL.
  • DeepSeek silently maps any unknown model name to deepseek-flash, and claude-opus* to deepseek-v4-pro. A wrong name therefore still "works", on the wrong model. Check the model field in every response.
  • DeepSeek's endpoint ignores anthropic-beta, anthropic-version, mcp_servers and container.
  • Anthropic's server-side tools and effort semantics won't carry over one-to-one.

Sources (fetched and read this session unless marked)

Rendered from the project's research note (harness_research.md). Lightly edited for the web: one personal attribution was removed, one personal-blog domain is shown as "community report", and notes on how the test machine reaches each vendor were left out.