🤖 Coding agent harness barely affects quality, but changes price up to 5×
Arena.ai published the HarnessTax study on September 16, 2026: 7 models on three harnesses (Claude Code, Codex CLI, Pi) were run on SWE-bench Lite and Terminal-Bench 2.0. Success variance — ±2–5%, prices — up to 5×: Claude Fable 5 gives 97.8% in Claude Code vs. 96.7% in Pi, but $1.33 vs. $0.67 per task.
🌍 In 9 out of 12 comparisons, a "foreign" harness outperformed the native one: Sonnet 4.6 — 68.9% in Codex vs. 66.7% in Claude Code. The default harness is a hidden confounder in leaderboards: comparing models without accounting for the harness is incorrect.
👤 The same model in a different CLI — almost the same success, savings up to 5×. The open Pi with four tools is on the Pareto frontier, its starting context is 10+ times smaller. Data and traces are open at harnesstax.github.io — check on your own tasks.
Source 1: https://arena.ai/blog/coding-agents-harness-tax Source 2: https://harnesstax.github.io/
