Submitted:
09 May 2026
Posted:
12 May 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Why Current Evaluation Conflates Model and Harness
2.1. How the Harness Enters Every Benchmark Score
2.2. The Structural Reason: Agents Are Closed-Loop Systems
2.3. Why Standard Responses Fall Short
3. The Binding Constraint Thesis
| The Binding Constraint Thesis |
|
For LLM agents operating on long-horizon tasks with comparable frontier models, let denote the benchmark score of model under harness . Define: |
3.1. Variance Decomposition
3.2. Reliability Quantities of the Harness Controller
3.3. Scope, Falsifiability, and Implications
4. Evidence
4.1. Evidence from Public Leaderboards
| Benchmark | Model (Fixed) | Harness Change | Layer | |
|---|---|---|---|---|
| SWE-bench Pro | Claude Opus 4.5 | SEAL → Claude Code | E, T, C | +9.5pp |
| (45.9% → 55.4%) | ||||
| SWE-bench Verified | Grok 4 | SWE-agent → xAI scaffold | T, C, O | +14–16pp |
| (58.6% → 72–75%) | ||||
| TerminalBench 2.0 | Fixed model | Prompt + middleware + verif. | C, S, V | +13.7pp |
| (52.8% → 66.5%) | ||||
| TerminalBench-2 | Fixed model | Automated harness opt. | All | 76.4% |
| TerminalBench 2 | GPT-5.4 (high) | AHE harness evolution | C, T, S, O | +7.3pp |
| (69.7% → 77.0%) | ||||
| Coding benchmarks | Fixed model | Context format + tool defs | T, C | ∼10× |
| Agent tasks (Vercel) | Fixed model | 15 tools → 2 tools | T | 80%→100% |
| SWE-bench Pro | Various | +WarpGrep search subagent | T | +2.1–2.2pp |
4.2. Controlled Factorial
- (Minimal): no context compression, verbose tool schema, no retry logic, no verification hooks, and no anomaly recovery. High drift rate, high control lag, and no stability guarantees. This is the baseline open-loop harness.
- (Improved): compressed context with task-relevance retrieval, minimal tool schema, and structured retry on tool failure with exponential backoff. This configuration reduces drift and provides moderate control lag with basic closed-loop feedback.
- (Full): plus per-step self-checking, KL-style drift checks every five steps, anomaly-detection middleware for repeated action loops and context contradictions, full output validation, and checkpoint rollback retaining the last 10 states. This configuration has low drift, low control lag, and explicit stability enforcement.
| (Minimal) | (Improved) | (Full) | HV | |
| GLM-5.1 | 52.5 | 56.5 | 65.5 | 29.56 |
| GPT-5.4 | 55.0 | 58.5 | 63.5 | 12.17 |
| Kimi K2.6 | 52.0 | 59.0 | 60.5 | 13.72 |
| MV | 1.72 | 1.17 | 4.22 |
5. A Harness-Aware Evaluation Framework
6. Alternative Views and Counterarguments
7. Conclusion and Discussion
Appendix A. Example ETCSOVG Disclosure Card
| Layer | Minimum disclosure |
|---|---|
| Execution | Runtime substrate, sandboxing, filesystem access, network access, task timeout, step timeout, maximum steps, and evaluation entry point. |
| Tool | Tool list, schema style, tool-selection policy, error format, retryable error classes, and whether tools are model-visible or hidden middleware. |
| Context | Context cap, ordering policy, summarization/compression policy, retrieval method, persistent memory, and cache policy. |
| Scheduling | Agent loop, stopping rule, retry policy, escalation behavior, delegation policy, and rollback behavior. |
| Observability | Logged artifacts, trajectory format, checkpoints, traces, validation logs, and whether failures can be audited after the run. |
| Verification | Output parser, schema validation, self-checking, test execution, anomaly detection, and final patch validation. |
| Governance | Permission model, allowlist/denylist rules, side-effect boundaries, secret handling, and human approval points. |
| Layer | Disclosure field | setting |
|---|---|---|
| Execution | Runtime and limits | Docker SWE-bench runner; 50 steps; 120s step timeout |
| Tool | Interface and errors | Task-relevant tools; minimal schema; structured errors |
| Context | Memory policy | 32k token cap; summarize old steps; BM25 top-5 retrieval |
| Scheduling | Retry and recovery | 3-attempt backoff; checkpoint rollback depth 3 |
| Observability | Trace surface | Full trace logs; 10 retained checkpoints |
| Verification | Validation hooks | Self-checking; full output validation; anomaly detection |
| Governance | Permission model | Allowlist-oriented tool governance |
Appendix B. Harness Configuration Specifications
| Field | Value |
|---|---|
| Source dataset | princeton-nlp/SWE-bench_Verified |
| Split | test |
| Full split size | 500 instances |
| Subset size | 100 instances |
| Sampling rule | stratified by difficulty |
| Random seed | 42 |
| min fix | 39 |
| 15 min–1 hour | 52 |
| 1–4 hours | 8 |
| hours | 1 |
| Model | Provider | Generation settings | Timeout |
|---|---|---|---|
| GPT-5.4 | OpenAI | max output 4096 | 180s |
| Kimi K2.6 | Moonshot | max output 4096 | 180s |
| GLM-5.1 | ZAI | max output 4096 | 180s |
| (Minimal) | (Improved) | (Full) | HV | |
| First final run | ||||
| GLM-5.1 | 52.0 | 57.0 | 66.0 | 33.56 |
| GPT-5.4 | 55.0 | 59.0 | 64.0 | 13.56 |
| Kimi K2.6 | 52.0 | 59.0 | 61.0 | 14.89 |
| MV | 2.00 | 0.89 | 4.22 | |
| Second final run | ||||
| GLM-5.1 | 53.0 | 56.0 | 65.0 | 26.00 |
| GPT-5.4 | 55.0 | 58.0 | 63.0 | 10.89 |
| Kimi K2.6 | 52.0 | 59.0 | 60.0 | 12.67 |
| MV | 1.56 | 1.56 | 4.22 | |
| Field | Minimal | Improved | Full |
|---|---|---|---|
| Source tasks | SWE-bench Verified subset100, seed 42, fixed task order | Same as | Same as |
| Runtime | Docker SWE-bench execution and evaluation pipeline | Same as | Same as |
| Step budget | 50 agent steps; 120s per-step timeout | Same as | Same as |
| Design goal | Open-loop baseline with minimal intervention | Tool-robust closed-loop harness | plus self-checking and recovery controls |
| Context strategy | Append all prior steps chronologically | Compressed retrieval context | Same as |
| Context cap | 200k token estimate | 32k token estimate | 32k token estimate |
| History compression | None | Summarize older steps outside the recent window; compression window 8 | Same as |
| Retrieval | None | BM25 top-5 over older trajectory steps | Same as |
| Context ordering | Chronological | Relevance then recency | Relevance then recency |
| Drift monitoring | Disabled | Disabled | KL-style drift check every 5 steps; threshold 2.0 |
| Tool schema | Verbose tool descriptions | Minimal task-focused tool schema | Minimal task-focused tool schema |
| Tool exposure | All available tools | Task-relevant tool subset | Task-relevant tool subset |
| Tool error format | Raw tool errors | Structured error feedback | Structured error feedback |
| Runtime retry | No retry on tool failure | 3-attempt exponential backoff | retry policy plus malformed-output retry |
| Retryable errors | None | timeout, rate limit, transient network | timeout, rate limit, transient network, malformed output |
| Failure escalation | No explicit recovery policy | Continue after recoverable tool failures and feed the error back as an observation | Checkpoint-style recovery on detected anomalies |
| Parser behavior | Basic parser with one tolerated malformed-output retry | Parser normalization for natural model outputs | Parser normalization plus validation layer |
| Output validation | Disabled | Schema-only validation | Full output validation |
| Empty patch handling | Allowed | Rejected when validation is active | Rejected when validation is active |
| Self-verification | Disabled | Disabled | Enabled; lightweight per-step self-check |
| Anomaly detection | Disabled | Disabled | Enabled for repeated action loops and context contradictions |
| Checkpointing | Disabled | Disabled | Enabled; retain last 10 state checkpoints |
| Rollback policy | None | None | Roll back up to 3 checkpoints on handled anomalies |
| Observability | Minimal logs | Full trajectory-level logs | Full trace logs with verification metadata |
| Governance | Permissive permission model | Permissive permission model | Allowlist-oriented tool governance |
| Model | Harness | Cost (USD) | Tokens | Sequential time |
|---|---|---|---|---|
| GPT-5.4 | ||||
| Minimal | $34.19 | 18.9M | 3.6h | |
| Improved | $22.19 | 9.7M | 2.0h | |
| Full | $23.05 | 10.0M | 2.4h | |
| Subtotal | $79.43 | 38.5M | 8.0h | |
| Kimi K2.6 | ||||
| Minimal | $24.29 | 60.4M | 19.6h | |
| Improved | $30.45 | 43.7M | 11.0h | |
| Full | $33.83 | 47.7M | 20.7h | |
| Subtotal | $88.57 | 151.8M | 51.3h | |
| GLM-5.1 | ||||
| Minimal | $12.54 | 29.7M | 9.2h | |
| Improved | $13.86 | 16.5M | 7.1h | |
| Full | $21.39 | 24.5M | 8.8h | |
| Subtotal | $47.79 | 70.7M | 25.1h | |
| Total | $215.79 | 261.1M | 84.4h | |
Appendix C. Perturbation Stress-Test Protocol Details
Appendix D. Industry Evidence Summary
Appendix E. Mechanism Analysis from Trajectory Logs
References
- Jimenez, C.E.; Yang, J.; Wettig, A.; Yao, S.; Pei, K.; Press, O.; Narasimhan, K. Swe-bench: Can language models resolve real-world github issues? arXiv 2023, arXiv:2310.06770. [Google Scholar]
- Merrill, M.A.; Shaw, A.G.; Carlini, N.; Li, B.; Raj, H.; Bercovich, I.; Shi, L.; Shin, J.Y.; Walshe, T.; Buchanan, E.K.; et al. Terminal-bench: Benchmarking agents on hard, realistic tasks in command line interfaces. arXiv 2026, arXiv:2601.11868. [Google Scholar] [CrossRef]
- Liu, X.; Yu, H.; Zhang, H.; Xu, Y.; Lei, X.; Lai, H.; Gu, Y.; Ding, H.; Men, K.; Yang, K.; et al. Agentbench: Evaluating llms as agents. arXiv 2023, arXiv:2308.03688. [Google Scholar] [CrossRef]
- Mialon, G.; Fourrier, C.; Wolf, T.; LeCun, Y.; Scialom, T. Gaia: a benchmark for general ai assistants. In Proceedings of the The Twelfth International Conference on Learning Representations, 2023. [Google Scholar]
- Lin, J.; Liu, S.; Pan, C.; Lin, L.; Dou, S.; Huang, X.; Yan, H.; Han, Z.; Gui, T. Agentic Harness Engineering: Observability-Driven Automatic Evolution of Coding-Agent Harnesses. arXiv 2026. [arXiv:cs.CL/2604.25850]. [Google Scholar]
- Brand, F.; JSD. Why Benchmarking Is Hard. Epoch AI, Gradient Updates, 2025.
- Yao, S.; Zhao, J.; Yu, D.; Du, N.; Shafran, I.; Narasimhan, K.; Cao, Y. React: Synergizing reasoning and acting in language models. arXiv 2022, arXiv:2210.03629. [Google Scholar]
- Yang, J.; Jimenez, C.E.; Wettig, A.; Lieret, K.; Yao, S.; Narasimhan, K.R.; Press, O. SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. In Proceedings of the The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. [Google Scholar]
- Wang, X.; Li, B.; Song, Y.; Xu, F.F.; Tang, X.; Zhuge, M.; Pan, J.; Song, Y.; Li, B.; Singh, J.; et al. Openhands: An open platform for ai software developers as generalist agents. arXiv 2024, arXiv:2407.16741. [Google Scholar]
- Lee, Y.; Nair, R.; Zhang, Q.; Lee, K.; Khattab, O.; Finn, C. Meta-Harness: End-to-End Optimization of Model Harnesses. arXiv 2026, arXiv:2603.28052. [Google Scholar]
- Lou, X.; Lázaro-Gredilla, M.; Dedieu, A.; Wendelken, C.; Lehrach, W.; Murphy, K.P. AutoHarness: improving LLM agents by automatically synthesizing a code harness. arXiv 2026, arXiv:2603.03329. [Google Scholar] [CrossRef]
- Zhang, Q.; Hu, C.; Upasani, S.; Ma, B.; Hong, F.; Kamanuru, V.; Rainton, J.; Wu, C.; Ji, M.; Li, H.; et al. Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models. In Proceedings of the The Fourteenth International Conference on Learning Representations, 2026. [Google Scholar]
- Morph. SWE-Bench Pro Leaderboard (2026): Why 46% Beats 81%, 2026.
- Kapoor, S.; Stroebl, B.; Kirgis, P.; Nadgir, N.; Siegel, Z.S.; Wei, B.; Xue, T.; Chen, Z.; Chen, F.; Utpala, S.; et al. Holistic agent leaderboard: The missing infrastructure for ai agent evaluation. arXiv 2025, arXiv:2510.11977. [Google Scholar] [CrossRef]
- Li, X.; Jiao, W.; Jin, J.; Dong, G.; Jin, J.; Wang, Y.; Wang, H.; Zhu, Y.; Wen, J.R.; Lu, Y.; et al. Deepagent: A general reasoning agent with scalable toolsets. In Proceedings of the Proceedings of the ACM Web Conference 2026, 2026, pp. 2219–2230.
- Pan, J.; Wang, X.; Neubig, G.; Jaitly, N.; Ji, H.; Suhr, A.; Zhang, Y. Training software engineering agents and verifiers with swe-gym. arXiv 2024, arXiv:2412.21139. [Google Scholar] [CrossRef]
- Cai, Y.; Chen, L.; Chen, Q.; Ding, Y.; Fan, L.; Fu, W.; Gao, Y.; Guo, H.; Guo, P.; Han, Z.; et al. Nex-n1: Agentic models trained via a unified ecosystem for large-scale environment construction. arXiv 2025, arXiv:2512.04987. [Google Scholar]
- Xia, C.S.; Deng, Y.; Dunn, S.; Zhang, L. Demystifying llm-based software engineering agents. Proc. ACM Softw. Eng. 2025, 2, 801–824. [Google Scholar] [CrossRef]
- Sutton, R.S.; Barto, A.G.; et al. Reinforcement learning: An introduction; MIT press: Cambridge, 1998; Vol. 1. [Google Scholar]
- Deng, X.; Da, J.; Pan, E.; He, Y.Y.; Ide, C.; Garg, K.; Lauffer, N.; Park, A.; Pasari, N.; Rane, C.; et al. Swe-bench pro: Can ai agents solve long-horizon software engineering tasks? arXiv 2025, arXiv:2509.16941. [Google Scholar]
- Shihipar, T. Seeing like an Agent: How We Design Tools in Claude Code, 2026.
- Martin, L.; Cemaj, G.; Cohen, M. Scaling Managed Agents: Decoupling the Brain from the Hands. 2026. [Google Scholar]
- Miserendino, S.; Wang, M.; Patwardhan, T.; Heidecke, J. SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering? In Proceedings of the International Conference on Machine Learning. PMLR, 2025; pp. 44412–44450. [Google Scholar]
- Chan, J.S.; Chowdhury, N.; Jaffe, O.; Aung, J.; Sherburn, D.; Mays, E.; Starace, G.; Liu, K.; Maksin, L.; Patwardhan, T.; et al. MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering. In Proceedings of the The Thirteenth International Conference on Learning Representations.
- Jain, N.; Han, K.; Gu, A.; Li, W.D.; Yan, F.; Zhang, T.; Wang, S.; Solar-Lezama, A.; Sen, K.; Stoica, I. LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code. In Proceedings of the The Thirteenth International Conference on Learning Representations.
- Zhuo, T.Y.; Chien, V.M.; Chim, J.; Hu, H.; Yu, W.; Widyasari, R.; Yusuf, I.N.B.; Zhan, H.; He, J.; Paul, I.; et al. BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions. In Proceedings of the The Thirteenth International Conference on Learning Representations, 2025. [Google Scholar]
- Yang, J.; Jimenez, C.E.; Zhang, A.L.; Lieret, K.; Yang, J.; Wu, X.; Press, O.; Muennighoff, N.; Synnaeve, G.; Narasimhan, K.R.; et al. SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains? In Proceedings of the The Thirteenth International Conference on Learning Representations, 2025. [Google Scholar]
- Vidgen, B.; Mann, A.; Fennelly, A.; Stanly, J.W.; Rothman, L.; Burstein, M.; Benchek, J.; Ostrofsky, D.; Ravichandran, A.; Sur, D.; et al. APEX-agents. arXiv 2026, arXiv:2601.14242. [Google Scholar] [CrossRef]
- Hu, S.; Lu, C.; Clune, J. Automated design of agentic systems. arXiv 2024, arXiv:2408.08435. [Google Scholar] [CrossRef]
- Rajasekaran, P. Harness Design for Long-Running Application Development, 2026.
- Zunic, G. The Bitter Lesson of Agent Harnesses, 2026.
- Ge, C.; Kryvosheieva, D.; Fried, D.; Girit, U.; Hariharan, K. Agent psychometrics: Task-level performance prediction in agentic coding benchmarks. arXiv 2026, arXiv:2604.00594. [Google Scholar] [CrossRef]
- Vals, A.I. SWE-Bench. 2026. [Google Scholar]
- Schluntz, E. Raising the Bar on SWE-bench Verified with Claude 3.5 Sonnet. 2025. [Google Scholar]
- Trivedy, V. Improving Deep Agents with Harness Engineering; 2026. [Google Scholar]
- Qu, A. We Removed 80% of Our Agent’s Tools. 2025. [Google Scholar]
- Bölük, C. I Improved 15 LLMs at Coding in One Afternoon. Only the Harness Changed, 2026. [Google Scholar]
- OpenAI. Introducing GPT-5.4. 2026. [Google Scholar]
- Moonshot, A.I. Kimi-K2.6, 2026.
- ai, Z. GLM-5.1. 2026. [Google Scholar]
- OpenAI. Codex CLI. 2025.
- Anthropic. Claude-Code, 2025.
- Anomaly. Opencode: The Open Source Coding Agent., 2025.
- Harbor. Terminus-2 2026.
- Research, N. Hermes Agent — The Agent That Grows With You. 2026. [Google Scholar]
- DeepSeek-AI. DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence, 2026.
- Team, Q. Qwen3.6-Plus: Towards Real World Agents. 2026. [Google Scholar]
| 1 | LLM Stats coding leaderboard, https://llm-stats.com/, accessed April 23, 2026. |
| 2 |
https://llm-stats.com/leaderboards/best-ai-for-coding, accessed April 23, 2026. |
| 3 | An open-source personal-AI variant is OpenClaw, described at https://openclaw.ai/. |

Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).