Harnesses
Which agentic harness each model family should run in, and which harness is fair for everyone. The research behind the home and neutral tracks, researched 2026-09-23.
Scope: which harness to use per model vendor, and which harnesses can run models from several vendors, for the models reachable from our test machine (Anthropic, OpenAI, DeepSeek, xAI, Google, Moonshot, Cerebras and Fireworks). Method: web search plus pages fetched and read in this session; versions pulled live from the npm registry and GitHub API on 2026-09-23. No harness was installed or test-run for this note. Anything marked (unverified) comes from a snippet or a single secondary source.
1. The headline evidence (who wins per model family)
Terminal-Bench 4.0, official board (tbench.ai, read from the page's own data 2026-09-23). It only lists first-party harnesses plus mini-SWE-agent:
| Harness | Model (effort) | Resolution |
|---|---|---|
| Codex | GPT-6 Astra (max) | 58.2% |
| Codex | GPT-6 Astra (xhigh / high) | 57.9% |
| Claude Code | Fable 5.1 (xhigh / max) | 57.9% |
| Claude Code | Opus 5 (xhigh) | 53.9% |
| Claude Code | GLM-5.3 (max) | 41.8% |
| Grok Build | Grok 4.7 (xhigh) | 37.6% |
| Codex | GPT-5.6 Sol (max) | 37.3% |
| mini-SWE-agent | Gemini 3.8 Flash (high) | 19.1% |
| Claude Code | Sonnet 5 (max) | 12.4% |
(Opus 5.5 isn't on the official TB4 board yet.)
Artificial Analysis Coding Agent Index v1.5 (artificialanalysis.ai/agents/coding-agents; DeepSWE v1.1 + Terminal-Bench 4.0 + SWE-Atlas-QnA, 3 attempts each; read from the page's data):
| Harness – model | Index | TB4 |
|---|---|---|
| Claude Code – Opus 5.5 (max) | 66.0 | 63.1 |
| Claude Code – Fable 5.1 (max) | 62.2 | 57.6 |
| Codex – GPT-6 Astra (max) | 61.7 | 55.6 |
| Devin Fusion CLI – Fable 5.1 xhigh + SWE-2 | 61.7 | 56.1 |
| Claude Code – Opus 5 (max) | 59.7 | 54.5 |
| Devin Fusion CLI – GPT-6 Astra xhigh + SWE-2 | 58.9 | 50.0 |
| Codex – GPT-6 Sol (max) | 56.6 | 43.4 |
| Grok Build – Grok 4.7 (xhigh) | 56.3 | 33.3 |
| OpenCode – GLM-5.3 (max) | 53.6 | 39.9 |
| Kimi Code CLI – Kimi K3 | 51.9 | 21.2 |
| Claude Code – Qwen3.8 Max | 43.3 | 16.7 |
| Codex – DeepSeek V4 Pro 0813 (max) | 43.0 | 10.1 |
| Codex – GPT-6 Luna (max) | 41.1 | 15.2 |
| Codex – DeepSeek V4 Flash 0731 (max) | 38.7 | 10.6 |
AA's own "Harness Comparison" tab says "Coming soon", so there's no clean same-model, many-harness table yet.
What the numbers say:
- The first-party harness wins where we can compare. GPT-6 Astra: Codex 61.7 vs Devin Fusion 58.9. Fable 5.1: Claude Code 62.2 vs Devin Fusion 61.7.
- Codex is a bad home for DeepSeek. DeepSeek V4 Pro/Flash in Codex score 10–11% on TB4 (AA). DeepSeek reports 31.2% on TB4 for V4.1 Flash in its own harness (vendor-reported, not on the official board; via mightybot.ai).
- GLM-5.3 does about as well in Claude Code (41.8%, official TB4) as in OpenCode (39.9%, AA's TB4).
- This hasn't always held. On the older Terminal-Bench 2.0, third-party harnesses sometimes beat the vendor's: Factory Droid 64.9% vs Codex CLI 62.9% on GPT-5.2 (reddit r/codex), and Letta Code 59.1% vs Claude Code 41.6% on Opus 4.5 (winder.ai). The arXiv position paper "Stop Comparing LLM Agents Without Disclosing the Harness" (2605.23950, May 2026) finds that on SWE-bench Pro, changing the harness around one model moved its score twice as much as changing the model in a fixed harness (Opus 4.5: 45.9% SEAL vs 55.4% Claude Code).
- DoltHub's practitioner review (Aug 5, 2026) sums up the trend: "The models were trained to use their own harnesses… so they just worked better with their native harnesses."
2. Recommendation per vendor family (our models)
| Vendor (our models) | Best harness | Why | Runner-up |
|---|---|---|---|
| Anthropic (opus-5-5, fable-5-1, opus-5, sonnet-5, haiku-4-5) | Claude Code 2.1.280 on npm (GitHub tag 2.1.281 today) | Top of AA's index (Opus 5.5 66.0) and joint top of TB4 (Fable 5.1 57.9%) | OpenCode (neutral) |
| OpenAI (gpt-6-astra/sol/luna, gpt-5.5, 5.4-mini/nano) | Codex CLI 0.156.1 | #1 on TB4 (Astra max 58.2%); beats Devin Fusion on the same model in AA. I found no 2026 result where another harness beats Codex on GPT-6 | Factory Droid (droid exec, closed, BYOK over the Responses API; beat Codex on GPT-5.2 in TB2.0). OpenCode for neutral runs |
| DeepSeek (deepseek-v4-pro, deepseek-flash) | DeepSeek Harness (dsh): npm 0.1.5-rc.3, GitHub tag dsh-v0.1.7-rc.1 (Sep 23) | Vendor reports 31.2% TB4 in its own harness, vs about 10% for DeepSeek in Codex (AA). Developer preview with "breaking changes" warnings, web-UI-first | Claude Code via DeepSeek's official Anthropic endpoint; OpenCode (built-in DeepSeek provider) |
| xAI (grok-4.7, grok-4.6, grok-build-0.1) | Grok Build (open-sourced under Apache-2.0 in July, v1.0 on Aug 7) | xAI's own harness; the only Grok entries on TB4/AA are in it (Grok 4.7 37.6% / index 56.3); headless grok -p --output-format streaming-json; reads XAI_API_KEY and custom base_url | OpenCode; Codex also works (xAI's recommended API is Responses) |
| Google (gemini flash / flash-lite / pro) | Gemini CLI (v0.60.0, Sep 15) with an API key | Google shut Gemini CLI off for free/Pro/Ultra accounts on Jun 18, 2026 and replaced it with the closed Antigravity CLI (agy). Gemini CLI still works with paid Gemini API keys, and GOOGLE_GEMINI_BASE_URL lets it point at a custom endpoint. agy is OAuth-first (API-key auth is an open feature request, issue #78), so it's awkward on a headless box | OpenCode / Qwen Code (the open Gemini-CLI fork). No strong harness result exists for Gemini: the only TB4 entry is in the minimal mini-SWE-agent (19.1%) |
| Moonshot (kimi-k3, kimi-k2.7-code) | Kimi Code CLI 2.1.0 (released today, MIT) | Moonshot's own harness, on AA's index (K3 51.9). Non-interactive -p mode (docs) | Claude Code via Moonshot's official api.moonshot.ai/anthropic endpoint (Moonshot documents it) |
| Fireworks / Cerebras open models (glm-5.3, minimax-m3, kimi-k3-fast, qwen3p8-max, gpt-oss-120b, qwen-3.8-27b) | OpenCode | Built-in Fireworks and Cerebras providers; AA runs GLM-5.3 in OpenCode (53.6). GLM-5.3 also runs well in Claude Code (41.8% TB4) but needs an Anthropic endpoint for it. Z.AI's own ZCode is a closed desktop app, not a CLI | Qwen Code for the Qwen models; Claude Code for GLM where there's an Anthropic endpoint |
3. Best multi-vendor harness for fair same-harness comparisons: OpenCode
- Open source and active: MIT, about 210k GitHub stars, v1.18.32 released Sep 21, 2026. The repo moved from sst/opencode to anomalyco/opencode.
- Headless:
opencode run "<prompt>" -m provider/model --format jsonstreams raw JSON events.opencode serveplusrun --attachavoids MCP cold starts.opencode session statsgives token and cost stats. - Speaks all three of our API formats natively:
- Anthropic Messages (
@ai-sdk/anthropic) - OpenAI Responses (
@ai-sdk/openai; the docs say to use it for/v1/responses) - Google native (
@ai-sdk/google) - OpenAI-compatible chat (
@ai-sdk/openai-compatible) for everything else.
- Anthropic Messages (
- Custom endpoints: a
baseURLper provider, so each provider can point at a custom endpoint. - Built-in providers for every vendor on our list: Anthropic, OpenAI, Google, DeepSeek, xAI, Moonshot, Cerebras, Fireworks, Groq, Z.AI and MiniMax, plus 75+ via Models.dev.
- It's what benchmarkers use: Artificial Analysis already uses it as a neutral harness. It's the most-used provider-agnostic harness in 2026 roundups (pinggy.io, winder.ai, nimbalyst, mightybot).
- Caveats:
- Anthropic forced it to drop Claude Pro/Max subscription login in January 2026. API keys still work, which is our setup.
- Before v1.2.23 it sent session titles to a hosted small model. That's fixed.
- A neutral harness still penalises models tuned for their own harness. So run two tracks and label them: "home harness" (each model in its vendor's CLI) and "neutral harness" (every model in OpenCode, same config). The gap between the two is itself a QA finding.
Close alternatives for the neutral track:
- Pi (earendil-works/pi, MIT, about 109k stars, v0.87.1 Sep 22): a tiny, token-efficient loop with little harness bias. Unified OpenAI/Anthropic/Google provider layer; print/JSON mode.
- Goose (aaif-goose/goose, Apache-2.0, now under the Linux Foundation, v1.52.0 today): 15+ providers,
goose runheadless. - Crush (Charm, FSL-1.1-MIT, v0.96.1): OpenAI, openai-compat, Anthropic and Gemini providers.
crush runheadless (always auto-approve). Pre-1.0. - Qwen Code (Apache-2.0, npm 0.24.4): the Gemini-CLI fork; configures OpenAI, Anthropic and Gemini APIs.
- OpenHands (MIT, v1.23.0 today):
openhands --headless -t "…" --json, always auto-approve; LiteLLM underneath, so any provider.
4. Harness-by-harness facts
| Harness | Open source | Headless | API formats / providers | Notes / first-party pairing |
|---|---|---|---|---|
| Claude Code (Anthropic) | No (source-available repo, custom licence) | claude -p | Anthropic Messages only; ANTHROPIC_BASE_URL for gateways | Anthropic's home harness. Non-Claude models work only where the vendor exposes an Anthropic-format endpoint: DeepSeek api.deepseek.com/anthropic (official), Moonshot api.moonshot.ai/anthropic (official), Z.AI/MiniMax/Qwen (community guide cc-compatible-models). OpenAI, Gemini and Cerebras need a translation layer (LiteLLM, claude-code-router). Anthropic's gateway docs: it "doesn't support routing Claude Code to non-Claude models through any gateway". |
| Codex CLI (OpenAI) | Yes, Apache-2.0, about 126k stars | codex exec | Responses API only: the config reference says wire_api "responses is the only supported value". Chat-completions providers are out | OpenAI's home harness. Also runs xAI (Responses is xAI's recommended API) and DeepSeek (official Responses endpoint, stateless: no previous_response_id/store/background), but DeepSeek scores poorly in it (AA). Anthropic, Gemini, Cerebras and Moonshot need a gateway that translates to Responses (e.g. LiteLLM). |
| OpenCode (Anomaly, ex-SST) | Yes, MIT | opencode run --format json | Anthropic, OpenAI Responses, Google, OpenAI-compat; 75+ providers | Best neutral harness (section 3). |
| Gemini CLI (Google) | Yes, Apache-2.0 | gemini -p | Gemini API only; GOOGLE_GEMINI_BASE_URL | Retired for consumer accounts Jun 18, 2026; paid API keys still work. |
Antigravity CLI agy (Google) | No (closed Go) | agy -p (print mode, stream-JSON) | Google models; OAuth-first | Gemini CLI's successor; API-key auth not yet supported (issue #78). |
| Qwen Code (Alibaba) | Yes, Apache-2.0 | -p (inherited from Gemini CLI) | OpenAI, Anthropic, Gemini, Qwen APIs; any base URL | Home harness for Qwen models. |
| Kimi Code CLI (Moonshot) | Yes, MIT | kimi -p (non-interactive, auto permissions) | Kimi first; "other compatible providers" configurable | Moonshot's home harness; v2.1.0 today; desktop client Sep 17. |
| Grok Build (xAI) | Yes, Apache-2.0 (open-sourced Jul 2026) | grok -p, --output-format streaming-json; ACP | xAI API key; any custom model via ~/.grok/config.toml base_url + env_key | xAI's home harness; ports tool code from Codex and OpenCode (its NOTICE). Cursor CLI is a separate product from the same camp. |
DeepSeek Harness dsh | Yes, MIT, about 234k stars | Web UI first (dsh web --no-open); headless also mentioned (pinggy). Exact CLI flag (unverified) | Multi-provider; community reports it works with "almost all major providers" (community report); can drive Claude Code/Codex as sub-agents (winder.ai) | DeepSeek's home harness; developer preview, breaking changes promised. |
| Crush (Charm) | Source-available FSL-1.1 → MIT after 2 yrs | crush run | OpenAI, openai-compat, Anthropic, Gemini, Bedrock, OpenRouter… | Best-looking TUI (good for screen recordings); switches models mid-session. |
| Goose (Block → Linux Foundation) | Yes, Apache-2.0 | goose run -t | 15+ providers, plus ACP to subscriptions | General agent; MCP-heavy. |
| Aider | Yes, Apache-2.0 | aider --message | Any (LiteLLM) | No commits since May 22, 2026; last tagged release v0.86.0 (Aug 2025). Stalled; avoid. |
| Cline CLI | Yes, Apache-2.0 (npm cline 3.0.64) | Has a CLI; headless flags (unverified) | Model-agnostic | Active. |
| Roo Code | Archived May 15, 2026 | — | — | Dead; its users moved to Kilo Code. |
| Kilo Code CLI | Yes, MIT (npm @kilocode/cli 7.7.9) | Yes (built on OpenCode) | Same breadth as OpenCode | Roo/Cline fork; CLI is OpenCode underneath, so it adds nothing for us. |
| Pi | Yes, MIT | print / JSON / RPC modes | OpenAI, Anthropic, Google, … | Lean, low-bias harness. |
| OpenHands | Yes, MIT | openhands --headless -t … --json | Any via LiteLLM | Sandboxed, CI-oriented. |
| Factory Droid | No | droid exec | BYOK: anthropic (Messages), openai (Responses), generic-chat-completion-api (chat), each with a custom baseUrl | Strong on OpenAI models historically (TB2.0 GPT-5.2: 64.9% vs Codex 62.9%). npm @factory/cli 0.225.2. Account needed. |
Cursor CLI (agent) | No | Print mode | Cursor-hosted models; the CLI doesn't accept your own API key or endpoint (Cursor forum) | Not usable with custom endpoints. |
| Amp (Sourcegraph) | No | Execute mode | Hosted only; the mode picks the model (e.g. Astra, Fable 5.1, GLM-5.3-Flash) | You can't pin one model, so it's unsuitable for per-model QA. |
| Warp Agent CLI | Client AGPL-3.0 | Yes | Warp-hosted models | Not a fit for BYOK testing. |
| Devin Fusion CLI (Cognition) | No | — | Pairs a frontier model with SWE-2 | On AA's index, but it's a model blend, not a single-model harness. |
5. What works with each vendor's API (per harness × vendor)
✅ = native format, ⚠️ = possible via the vendor's alternate endpoint or unverified, ❌ = needs a translation gateway.
| Anthropic | OpenAI | DeepSeek | xAI | Gemini | Moonshot | Cerebras | Fireworks | |
|---|---|---|---|---|---|---|---|---|
| Claude Code | ✅ | ❌ | ✅ (/anthropic) | ❌ | ❌ | ✅ (/anthropic) | ❌ | ⚠️ |
| Codex CLI | ❌ | ✅ | ✅ (Responses, but scores poorly) | ✅ (Responses) | ❌ | ❌ | ❌ | ⚠️ |
| OpenCode | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| Qwen Code | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| Crush / Goose / Pi / OpenHands | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ | ✅ |
| Grok Build | ⚠️ | ⚠️ | ⚠️ | ✅ | ⚠️ | ⚠️ | ⚠️ | ⚠️ |
| Kimi Code | ⚠️ | ⚠️ | ⚠️ | ⚠️ | ⚠️ | ✅ | ⚠️ | ⚠️ |
| Gemini CLI | ❌ | ❌ | ❌ | ❌ | ✅ | ❌ | ❌ | ❌ |
| Factory Droid (BYOK) | ✅ | ✅ | ✅ | ✅ | ⚠️ | ✅ | ✅ | ✅ |
Claude Code quirks with third-party Anthropic endpoints:
- Map every model tier explicitly:
ANTHROPIC_MODEL,ANTHROPIC_DEFAULT_OPUS_MODEL/SONNET/HAIKU,CLAUDE_CODE_SUBAGENT_MODEL. - DeepSeek silently maps any unknown model name to deepseek-flash, and
claude-opus*to deepseek-v4-pro. A wrong name therefore still "works", on the wrong model. Check the model field in every response. - DeepSeek's endpoint ignores
anthropic-beta,anthropic-version,mcp_serversandcontainer. - Anthropic's server-side tools and effort semantics won't carry over one-to-one.
Sources (fetched and read this session unless marked)
- Terminal-Bench 4.0 official board (data embedded in page): https://www.tbench.ai/leaderboard/terminal-bench/4.0
- Artificial Analysis Coding Agent Index (harness+model rows): https://artificialanalysis.ai/agents/coding-agents
- CodingFleet TB4 summary (Sep 8/11, 2026): https://codingfleet.com/blog/terminal-bench-4-leaderboard-2026/
- Codex config reference, wire_api = responses only: https://developers.openai.com/codex/config-reference
- OpenCode providers / CLI docs: https://opencode.ai/docs/providers/ , https://opencode.ai/docs/cli/
- Claude Code LLM gateway docs: https://code.claude.com/docs/en/llm-gateway
- DeepSeek Anthropic API / Responses API: https://api-docs.deepseek.com/guides/anthropic_api , https://api-docs.deepseek.com/guides/responses_api
- Moonshot "Use Kimi in Claude Code": https://platform.kimi.ai/docs/guide/claude-code-kimi
- Grok Build docs: https://docs.x.ai/build/overview ; repo https://github.com/xai-org/grok-build
- xAI Responses vs Chat Completions: https://docs.x.ai/developers/model-capabilities/text/comparison
- Factory BYOK: https://docs.factory.ai/model-independence/byok
- Antigravity CLI headless + API-key issue: https://antigravity.google/docs/cli/headless/ , https://github.com/google-antigravity/antigravity-cli/issues/78
- Gemini CLI sunset + config (snippets): https://geminicli.com/docs/cli/headless/ , https://geminicli.com/docs/reference/configuration/
- OpenHands headless: https://docs.openhands.dev/openhands/usage/cli/headless
- Amp modes: https://ampcode.com/modes
- Cursor CLI status: https://forum.cursor.com/t/will-cursor-discontinue-agent-cli-after-grok-build-acquisition/171257 ; no BYOK in CLI (snippet): https://forum.cursor.com/t/how-to-bring-my-own-key-with-cursor-cli/127862
- Open-source harness landscape, Aug 28, 2026: https://pinggy.io/blog/best_open_source_cli_coding_agents/
- Harness comparison (Letta vs Claude Code, etc.): https://winder.ai/ai-agent-harness-comparison/
- DoltHub "Best Coding Agent 2026" (Aug 5, 2026): https://www.dolthub.com/blog/2026-08-05-best-coding-agent-2026/
- MightyBot ranking (Sep 18, 2026; DeepSeek own-harness TB4 claim): https://mightybot.ai/blog/coding-ai-agents-for-accelerating-engineering-workflows/
- arXiv 2605.23950 "Stop Comparing LLM Agents Without Disclosing the Harness": https://arxiv.org/html/2605.23950v1
- Droid vs Codex on GPT-5.2, TB2.0 (snippet only): https://www.reddit.com/r/codex/comments/1q52mwr/gpt52_hits_629_codex_cli_and_649_droid_on/
- Versions: npm registry and GitHub API, read 2026-09-23 (codex 0.156.1, claude-code 2.1.280 npm / v2.1.281 GH, opencode-ai 1.18.32, @deepseek-ai/dsh 0.1.5-rc.3 npm / dsh-v0.1.7-rc.1 GH, @qwen-code/qwen-code 0.24.4, @moonshot-ai/kimi-code 2.1.0, @charmland/crush 0.96.1, @kilocode/cli 7.7.9, cline 3.0.64, @google/gemini-cli 0.60.0, @factory/cli 0.225.2, goose v1.52.0, OpenHands v1.23.0, pi v0.87.1)
Rendered from the project's research note (harness_research.md). Lightly edited for the web: one personal attribution was removed, one personal-blog domain is shown as "community report", and notes on how the test machine reaches each vendor were left out.