Harness Engineering is the discipline of designing the scaffolding around an AI model – the prompts, tools, sandboxes, hooks, and feedback loops – so the model can complete real, multi-step work. The formulation of Harness Engineering in engineering terms would be as follows: Agent = Model + Harness. Upon being provided with the model, the harness will be the sole thing left for the engineer to develop. This guide gives an understanding of the meaning of the term, the issue of the originator of the term, the components of the harness, how five well-known programming agents perform harnessing, and ends with a checklist.
What Will I Learn?
What Is a Harness, Really?
A harness is everything that surrounds an AI model which is not the model itself: system prompts, tool definitions, the file system, sandboxed execution environment, hooks, sub-agents orchestration, observability logging. A bare model – the weights without the tools and execution loop – can neither perform an action nor read a file or execute a command. It is made an agent only when a harness provides it with a state, tools, and loop to run in.
A harness drives more behavior than the model does in many cases. For example, on Terminal-Bench 2.0, an open-source benchmarking suite for multi-step agent tasks using bash extensively, the exact same underlying model – Claude Opus 4.6 – achieved 79.8% performance in the ForgeCode harness versus 75.3% in the Capy harness, which is a 4.5 percentage point difference between the two harnesses with identical weights underneath.
There are three important implications that result from this definition:
- Two engineers utilizing the same model can arrive at different conclusions since their harnesses are different and not because their prompts are different.
- Only components of a harness which are able to make up for what the model itself is unable to do alone deserve to be included in the system: reliable session storage, security, durable state, etc.
- The improvement of a harness is a reproducible engineering task, unlike prompt tuning or model selection.
Artificial Intelligence (AI) Course
Average time: 4 month(s)
Skills you’ll build: Python for AI, Machine Learning, Neural Networks, NLP Basics, AI Tools (ChatGPT, Copilot)
Where the Term Actually Comes From
“Harness Engineering” is a term with a controversial history as it was coined by Mitchell Hashimoto, the co-founder of HashiCorp, for the first time on February 5, 2026, in his personal blog, and it was then formally defined by Vivek Trivedy of LangChain after twelve days.
The full timeline, as documented independently by multiple sources including a dedicated Wikipedia entry on “Agent harness”:
| Date (2026) | Author / Organization | Publication | Contribution |
|---|---|---|---|
| February 5 | Mitchell Hashimoto, HashiCorp co-founder | “My AI Adoption Journey” (personal blog) | First documented use of the phrase, describing the practice of engineering a permanent fix into an agent’s environment each time it makes a mistake |
| February 11 | Ryan Lopopolo, OpenAI | “Harness Engineering: Leveraging Codex in an Agent-First World” | OpenAI’s institutional definition of the discipline |
| February 17 | Vivek Trivedy, LangChain | “Improving Deep Agents with Harness Engineering” | First LangChain use of the term |
| March 10 | Vivek Trivedy, LangChain | “The Anatomy of an Agent Harness” | Formalized the Agent = Model + Harness equation and an 11-component breakdown; this post is the one most secondary sources credit as the term’s origin |
| April 2 | Birgitta Böckeler, Thoughtworks | “Harness engineering for coding agent users” (martinfowler.com) | Extended the concept with a feedforward/feedback framework, building on an earlier internal memo |
| July 4 | Lilian Weng, Thinking Machines Lab | “Harness Engineering for Self-Improvement” | Applied the concept to recursive self-improvement research |
Most of the individual articles, such as the highly read article by Addy Osmani, mention Trivedy alone as he is the one who brought the idea to light for everyone in the technical community and provided the formula that people have come to use. The accurate framing is that Hashimoto named the practice first, Trivedy gave it the definition that stuck.
The Anatomy of an Agent Harness
A harness has five recurring components: durable state, safe execution, tool access, memory and context management, and orchestration across multiple agents. Each feature is meant to address a particular weakness of the model.
Filesystem and State
A model can only reason about what fits in its context window. Without a filesystem, there is no way to persist work between turns. A harness provides the agent with a place to work — a place where it can read code and data, write out its intermediary outputs, and, if combined with Git, save changes, revert mistakes, and perform branching. Almost all other elements of the harness rely on the filesystem for some reason.
Sandboxes and Execution
Executing the code generated by an agent directly on a production server or a developer’s computer poses a threat. This is because of dangerous commands, unexpected network interactions, and resource consumption. The sandbox is used to isolate the execution, which allows the harness to whitelist certain commands, restrict network usage, and clean up after the task is done. OpenAI Codex enforces this at the kernel level, using Linux security mechanisms such as Landlock and seccomp; Claude code enforces it through configurable hooks – a more deterministic approach versus a more programmable one, respectively (Firecrawl comparison, June 206).
Tools and the Model Context Protocol
The agent must have well-defined actions that it can perform, such as reading a file, executing a shell command, and accessing the database. The Model Context Protocol (MCP) is an open protocol developed by Anthropic to integrate an agent with various tools and data sources using a universal interface instead of developing custom integrations for each tool. The internal coding assistant at Spotify, named Honk, utilizes MCP as the protocol between the agent’s reasoning and its verification system.
Memory and Context Management
The knowledge of a model is determined by the training phase and whatever information can fit within the context window size. Memory is maintained by harnesses using memory files, where AGENTS.md is one of the most frequently used conventions, which are reloaded every time before the session starts. While the context window reaches its limit, harnesses implement three strategies to avoid performance loss: compaction (context summarization), offloading (storing tool outputs on disk rather than context), and progressive disclosure (tool loading on demand).
Sub-Agents and Orchestration
Difficult tasks may involve the creation of several specialized agents rather than a single all-purpose agent. A harness may create several sub-agents with individual context windows, delegate individual tasks to them, and gather their results for coordination. This approach involves the separation of two functions that are unreliable when executed by a single agent – creating a solution and validating it.
Harness Engineering vs. Context Engineering vs. Prompt Engineering
Prompt Engineering focuses on optimizing one interaction between the user and the model; Context Engineering is about optimizing whatever goes into the context window; Harness Engineering deals with the complete environment within which the agent works.
| Discipline | Scope | Example activity |
|---|---|---|
| Prompt engineering | One interaction | Rewording an instruction to get a more accurate single response |
| Context engineering | One context window | Deciding which files, examples, and instructions to include before a single model call |
| Harness engineering | The full agent lifecycle, across sessions | Building hooks, memory files, sandboxes, and verification loops that persist across every run |
Harness engineering is the broader discipline; context engineering is one of the problems a harness has to solve, alongside execution safety, tool access, and multi-session memory.
The Feedforward and Feedback Model
A harness constrains an agent’s behavior in two ways: feedforward guides, which mold the agent’s output before acting, and feedback sensors, which measure the outcome after acting.
Coding convention documents, AGENTS.md, and skill definitions all are examples of feedforward guides; they’re things the agent reads before creating output. Static analysis, test suites, linters, and code reviews all are examples of feedback sensors; they’re things that evaluate the agent’s output after the fact and report back into the loop. If a harness gives feedforward guides but not feedback sensors, it can’t learn from its mistakes. If it gives feedback sensors but not feedforward guides, it spends time correcting mistakes it could have prevented.
Computational vs. Inferential Controls
Harness control mechanisms can be grouped in two groups. Computational control mechanisms make use of deterministic tooling, while inferential control mechanisms make use of a model to evaluate the outputs.
| Control type | How it works | Speed and cost | Examples |
|---|---|---|---|
| Computational | Deterministic tooling checks output against fixed rules | Fast, cheap, runs on every change | Linters, type checkers, eslint, architecture-boundary tests |
| Inferential | A model evaluates output using semantic judgment | Slower, more expensive, non-deterministic | AI code review, an “LLM as judge” step, semantic diff analysis |
Computational controls make it likely that the correct answer will be found because of fixed instrumentation; inferential controls involve judgments that cannot be made computationally, like judging if the change is right according to the ticket it is supposed to satisfy. Computational controls are usually used by production harnesses on every change, while inferential controls are used only for important checks.
Harness Engineering by the Numbers
Three figures quantify why harness design has become a distinct engineering discipline in 2026 rather than a footnote to model selection.
- 4.5 percentage points: the Terminal-Bench 2.0 score gap between the same model (Claude Opus 4.6) run in two different harnesses — 79.8% in ForgeCode versus 75.3% in Capy.
- 10x: the increase in new security findings per month in AI-generated code between December 2024 and June 2025, based on Apiiro’s analysis of tens of thousands of repositories — over 10,000 new findings per month by June 2025, up from roughly 1,000 six months earlier.
- 1,500+: the number of AI-generated pull requests Spotify’s internal coding agent, Honk, had merged into production as of its first public engineering post, running on the Claude Agent SDK with a generate-verify-retry loop that decouples code generation from CI verification.
Reading together, these figures support the same conclusion from three directions: the harness measurably changes benchmark performance, unchecked AI-generated code introduces measurable security risk, and a harness built around explicit verification can scale AI-generated code into production at volume.
How Harnesses Compare Across Coding Agents
In mid-2026, the five coding agents that are most frequently used have different approaches to the way that harness primitives work: the two that have the most sophisticated built-in scaffolding are Claude Code and OpenAI Codex, but Aider and Cline provide more control to the user.
| Coding agent | Interface | Execution model | Extensibility primitives | Verification behavior |
|---|---|---|---|---|
| Claude Code | Terminal-first CLI | Sandboxing enforced through configurable hooks | Skills, Hooks, Subagents, Plugins — described by multiple 2026 comparisons as the deepest extensibility stack among coding agents | Hooks can run linters and tests before a commit and block the commit on failure |
| OpenAI Codex | Cloud CLI, web app, IDE extension, GitHub integration | Kernel-level sandboxing in a cloud VM | Manifest-driven plugins | Automatic PR review runs through GitHub integration before human review |
| Cursor | IDE (VS Code fork) | Background cloud agents | Composer modes | Local and cloud agent modes, less standardized verification pipeline |
| Aider | CLI, git-native | Runs locally under user control; less autonomous by default | Bring-your-own-model | Every change becomes a discrete, reviewable git commit |
| Cline | VS Code extension (also CLI and SDK) | Runs against the user’s own API key and infrastructure | Plugin format described as close to and cross-compatible with Claude Code’s Skills | Verification depends on the user’s own CI setup |
Long-Horizon Execution and the Ratchet Method
Long-running agent tasks fail in three predictable ways: early stopping before the task is complete, poor decomposition of a large task into smaller steps, and incoherence once work spans multiple context windows. A harness address each failure mode with a specific technique.
Planning involves dividing the task into steps in order. The steps are usually written in the plan file that the agent uses in its process of achieving the goal. Verification of self involves verification of each step using the test suite or explicit criteria before proceeding to the next step. It is better to have a separate evaluation from the generator of the answer since when one agent evaluates its output there is a tendency towards false assurance.
Ratchet technique considers each mistake made by any agent as an input in the harness. As soon as an agent makes a certain mistake the correction for it is placed in the harness in the layer:
- If the agent violates a project convention, add the rule to AGENTS.md.
- If the agent runs a destructive command, add a hook that blocks that command pattern.
- If the agent produces code that fails a specific test category, add that test category to the verification loop that runs before every commit.
This means that each fix can be attributed to a particular failure that has been observed. This is different from using the ratchet approach because, when creating a big set of rules upfront, many rules that do not have a basis of a particular failure become mere noise for the model.
Artificial Intelligence (AI) Course
Average time: 4 month(s)
Skills you’ll build: Python for AI, Machine Learning, Neural Networks, NLP Basics, AI Tools (ChatGPT, Copilot)
Harness-as-a-Service: Build or Buy?
Anthropic, OpenAI, and other providers now package prebuilt harness infrastructure into software development kits to minimize engineering effort necessary for building an agent loop from scratch. The Claude Agent SDK, the Codex SDK, and the Open AI Agents SDK all provide the agent loop, tool calling infrastructure, context management, and sandboxing primitives as the starting point rather than a from-scratch build.
This changes the default starting point for creating a new agent project. In absence of any SDKs, developing an agent involved implementation of execution loop, tool calling logic, and conversation context management separately. Having the prebuilt harness SDK, these primitives are included into the default setup; further engineering effort involves the domain-specific details: which tools to expose, what prompt to use, and what hooks are needed for the particular codebase.
The prebuilt harness SDK does not make the complete harness on its own. Spotify’s Honk system uses the Claude Agent SDK but relies heavily on the additional abstraction on top of it: a verification service that provides an abstraction layer for the CI system, review-and-merge workflow, and specific codebase tools. The SDK provides the loop; the organization-specific harness is to be developed on top of it.
Self-Improving Harnesses
A harness can evolve through three pathways: the mining of logs from the agent itself in order to suggest changes to its own configuration, an evolutionary search among the candidates for the harness itself, and optimization jointly of the harness and the model weights that underlie it.
This is an area of ongoing research work at mid-2026 rather than something that’s widely implemented in production. A survey by Lilian Weng from July 2026 on the topic covers around 35 papers on the topic and mentions some specific implementations, one of them known as Self-Harness, where an agent evolves its own failures to prove suggested changes to its harness configuration. Users of the technique to build their own production harness will find this as something to watch rather than to try.
The Role of the Human
A harness must seek human authorization at certain specified points in time instead of being fully autonomous or reviewing each individual action manually. The common points are seeking authorization before undertaking a destructive task (pushing to the common branch forcefully, deleting a database table), before merging to a production branch, and reviewing the harness configuration itself periodically as failure patterns evolve.
These points will vary depending on the risk associated with the actions themselves and not because there is some set limit on the amount of autonomy that an agent should be granted in general. A low-risk, easily reversible task such as editing a test file does not require human authorization whereas a high-risk and irreversible action does.
Common Mistakes When Building a Harness
The following mistakes recur across public write-ups on harness design and account for most reported harness failures.
- Writing a large, speculative AGENTS.md before any failure has occurred. Rules not traceable to an observed mistake add noise without adding reliability, since every additional rule competes for the model’s attention on each turn.
- Skipping hooks and relying on prompt instructions alone. The instruction in a prompt is something that the model can choose not to do; the hook is something that the model has to do.
- Using the same agent to generate and evaluate its own work. Self-evaluation is unreliable because a model grading its own output tends to rate it more favorably than an independent evaluator would.
- Providing too many overlapping tools. Both name and description of every tool take space in the system prompt with each request; a smaller set of clearly differentiated tools outperforms a large set of redundant ones.
- Treating the harness as a one-time setup rather than a living system. The harness should be updated as new patterns of failure emerge and the underlying model evolves.
A Starter Harness Checklist
The following sequence gives a working starting point for a coding-agent harness.
- Create an AGENTS.md file containing only rules traceable to an actual observed failure, kept under roughly 60 lines.
- Add a pre-commit hook that will run the linter and type checker for the project and prevent the commit on failure.
- Add a destructive-command block list that includes commands such as rm -rf, force pushes to the main branch, and drop statements for databases.
- Configure a sandbox that allows you to execute code that is separate from the credentials of production and the rest of the file system not within the project directory.
- Define an approval checkpoint before any action that is high-risk or difficult to reverse, such as a production deployment.
- Set up a feedback sensor — a test suite, static analysis tool, or review step — that runs automatically after every change and reports failures back to the agent.
- Review the harness monthly against recent failures and remove rules that no longer apply as the model’s capabilities change.
FAQ
Q1. What is the difference between harness engineering and context engineering?
Ans. While context engineering deals with optimizing the information to fit into a single context window of a model, harness engineering deals with optimizing everything about the environment of an agent throughout multiple sessions, which includes execution safety, tool availability, and persistence of memory. Context engineering is one of the problems that have to be solved by a harness, rather than a distinct discipline.
Q2. Who coined the term harness engineering?
Ans. The term was first used by Mitchell Hashimoto of HashiCorp in his personal blog post on February 5, 2026; the formal definition, along with the formula, was published by Vivek Trivedy of LangChain twelve days later, on February 17, 2026.
Q3. Do I need to build a custom harness, or is a default agent’s built-in harness enough?
Ans. A default harness, such as the one built into Claude Code or OpenAI Codex, is sufficient for general-purpose tasks; a custom harness becomes necessary once a team needs organization – specific conventions, compliance requirements, or verification steps the default configuration does not enforce.
Q4. Does harness engineering apply to agents outside of coding?
Ans. Yes. The five main elements that make up harness engineering (durable state, safe execution, tools, memory, and feedback loop) apply to any AI agent (research, customer support, data processing, etc.) as coding agents are just the most public example of 2026.
Q5. How do I measure whether a harness is working?
Ans. Track task completion rate on a fixed benchmark or representative task set before and after a harness change, along with the rate of repeated failures the harness was designed to prevent, a working harness should reduce repeat failures over time without reducing completion rate.
Q6. What tools are required to start building a harness?
Ans. A version control system, a way to define hooks or pre-commit checks, a sandboxed execution environment, and a memory file convention such as AGENTS.md cover the minimum requirements; specific tooling depends on the coding agent or SDK in use.
Where Harness Engineering Is Going
The major coding agents in 2026 – Claude Code, Cursor, OpenAI Codex, Aider, and Cline — differ in underlying model but increasingly converge on the same harness primitives: hooks, sandboxes, memory files, and sub-agent orchestration. Prebuilt SDKs are lowering the cost of building a harness from scratch, which shifts the remaining engineering effort toward the domain-specific configuration a generic SDK cannot supply: which conventions to enforce, which actions require approval, and which failures to design against next.