Loop Engineering: The Practitioner’s Guide to Agent Loops

|
11 min read
|
36 views

What Is Loop Engineering, Really?

Loop engineering is the design of an automated process that will prompt an AI agent, evaluate its answer, and decide what to tell it next without a user typing out each instruction manually. A loop engineer will create the goal, the evaluation function, and the trigger point once, and the loop will keep running itself thereafter.

Loop engineering is distinct from AI prompts in the sense that once the loop starts working, it does not stop until the output matches the specified goal or the specified maximum number of loop cycles is achieved.

Where the Term Came From

The term “loop engineering” spread during a short span in June 2026. 

DatePerson or TeamContribution
Earlier in 2026Boris Cherny, AnthropicDescribed his role as writing the loops that prompt Claude Code, rather than prompting it directly.
June 7, 2026Addy OsmaniPublished an essay defining loop engineering as replacing the person who prompts an agent with a designed system.
June 8, 2026Peter SteinbergerStated publicly that engineers should design loops that prompt coding agents, instead of writing prompts by hand.
June 12, 2026Shawn Wang (“Swyx”)Named the broader practice “loopcraft” in an essay on stacking multiple loops.
June 16, 2026Sydney Runkle, LangChainPublished a four-level loop framework mapped to specific tooling.

Loop Engineering vs. Prompt Engineering

The engineering of the prompt fine-tunes an individual instruction for an individual AI output. The engineering of the loop develops the process of writing and rewriting that instruction repeatedly.

AspectPrompt EngineeringLoop Engineering
Unit of workOne instructionA repeating cycle of instructions
Who evaluates the outputA person, after each responseA grader or rule built into the system
Best suited forOne-off tasks, single model callsMulti-step tasks: code generation, monitoring, research
Human effort per cycleHigh — a new prompt each timeLow after setup — the system re-prompts itself

Loop Engineering vs. Context Engineering vs. Harness Engineering

Four interrelated techniques provide descriptions for the various levels of an AI agent system: prompt engineering, context engineering, harness engineering, and loop engineering.

  • Prompt engineering composes the instruction text itself.
  • Context engineering chooses which pieces of data, tools, and files will be made available to the model prior to responding.
  • Harness engineering constructs the execution environment in which the agent operates, including file access and feedback paths.
  • Loop engineering dictates how frequently the harness will be executed and when each iteration will be initiated.

There is no agreement among sources regarding the relationship between these levels. While one source puts loop engineering at a higher level than harness engineering, another considers it to be a separate technique rather than a higher level. Both agree, however, that a loop requires the harness, context pipeline, and prompt to exist first.

The Anatomy of a Loop: Every Stage, Explained

A loop runs through four repeating stages: a goal, an action, an observation, and an adjustment.

The Core Loop — Goal, Action, Observation, Adjustment

Loop Engineering
  1. Goal: Defined goal with a measurable stopping criteria e.g., “All the tests in the authentication module succeed.”
  2. Action: The agent writes code, modifies a file, or executes a command, taking it closer to achieving the goal.
  3. Observation: The system evaluates the outcome of the action; this could be through automated testing or continuous integration build.
  4. Adjustment: The system modifies the next command and begins the process again until the stopping criteria for the goal is met or the cycle threshold is exceeded.

Verification Loops

The verification loop introduces an additional step of grading to make sure that the output of the agent meets the criteria according to the rubric.

There are two types of grading:

  • Deterministic grading performs a set number of tests, like verifying whether a link works or if a build works.
  • Model-based grading involves using another language model to grade the output on the criteria written down. This type of grading is referred to as LLM-as-judge.

If the output fails the check, the system returns it to the agent with the specific failure identified, rather than a generic instruction to try again.

Event-Driven & Scheduled Loops

In an event-driven loop, the execution of an agent is initiated based on an event instead of the execution of a task by a user.

Examples of events are:

  • Scheduled time, which can be set using a cron job or an equivalent scheduler.
  • Webhook, which might be a pull request or a support ticket.
  • Message in a channel that is being watched, such as a Slack channel.

Hill-Climbing / Self-Improving Loops

The hill climbing loop is designed to evaluate past runs in order to optimize future instruction for the system instead of performing just one task.

During each agent run, a trace is generated which is a list of actions performed, tools invoked, and the grade received at the end of the run. The analysis phase examines the trace and finds a repeated problem.

The Building Blocks Every Loop Needs

A loop depends on five components: worktrees, skills, connectors, subagents, and a persistent state file.

Loop Engineering

Worktrees

Worktree refers to an isolated working copy of the source code repository, which allows various agents to modify files independently of one another. Git provides this functionality out of the box in that all worktrees share the same version history but have different checkouts. Various agents can operate in parallel this way. Each worktree is examined by a person/evaluator prior to merge with the master branch.

Skills (SKILL.md)

Skills are documents that contain instructions on how to carry out tasks that the agent needs to perform regularly in relation to the project. The standard format for skills involves a folder with an SKILL.md file and any number of other scripts and documents. An absence of the skill document means that the agent does not have information regarding the project’s specific conventions.

Plugins, Connectors & MCP

Connector is a connection from the agent to some other system like issue tracker, database, or messaging application. Most connectors support the Model Context Protocol (MCP), which is an open protocol used for connecting language models with other systems. Plugin packages connectors along with some skills and thus a team needs to configure one complete plugin rather than individual connectors.

Subagents — Maker vs. Checker

Subagents are distinct instances of agents each designated for performing an individual task of either coding or proofreading or verification of claims. The separation of the role of the coder and the role of the verifier leads to better results because it is difficult for the model to find its mistakes. One model will perform the coding while another will perform verification against the initial instruction.

State & Memory — the “Spine”

A loop requires a persistence of its memory in the form of a spine that keeps track of what has been tried by the system and what is still left to do. The most common examples are markdown files stored inside the repository or task board like Linear. An absence of persistent memory makes an agent forget the result of the previous iteration after closing the session.

What This Actually Costs — Tokens, Time, and Trade-offs

The loop approach increases the number of model calls needed to accomplish a task compared to a single prompt.

There are three ways token usage will increase:

  • Verification passes. Grading the completion pass is a new model call along with the original action.
  • Subagent splits. The maker-checker approach will make at least two model calls instead of one.
  • Retry cycles. A failure to complete the check means retrying the action pass, meaning another token usage for that action.
Loop Engineering

The trade-off is time spent versus detection of errors: fewer model calls but a greater risk of a missed error, or more model calls but catching more mistakes in advance.

Should You Actually Build a Loop? A Decision Framework

Not all tasks which use the help of AI require loops. There are four criteria which determine if a loop will add value to a task.

  1. Does the task have repetitions? A single-shot task does not warrant the additional effort to create a loop.
  2. Is it possible to automatically check the output? A loop needs an automatic check/rubric/regulation to check itself. Otherwise, there is no way to self-verify a task.
  3. What are the consequences of a mistake? A task which is prone to doing irreparable damage like sending money to someone requires human oversight in a loop, not automation.
  4. Is the token budget sufficient? A loop is more costly than a single prompt per task due to the verification procedure and presence of several agents.

A task that repeats, can be checked automatically, carries a reversible cost of error, and fits the available token budget is a strong candidate for a loop.

Build Your First Loop — A Tool-Agnostic Walkthrough

To set up the first loop, one must go through five actions, following their order.

  • Define the goal. Provide a clear, quantifiable criterion for halting the loop, e.g., “all unit tests succeed and the linter shows no errors.”
  • Create an isolated worktree. Create a dedicated version of the branch that will be unaffected by the loop until further examination.
  • Describe the skill file. Document all project conventions to ensure that the agent does not need to learn them each time it runs.
  • Include a validation action. Either configure a testing framework or another instance of the model that will validate the results relative to the goal.
  • Invoke the loop. Schedule a cron job, a webhook, or an in-session invocation of the loop.

Below is the table showing the correlation of functions and tools. 

FunctionClaude CodeOpenAI Codex
Repeat a prompt on a cadence/loop commandAutomations tab, scheduled run
Run until a condition is met/goal command/goal command
Isolate parallel workgit worktree, –worktree flagBuilt-in worktree per thread
Store project instructionsSKILL.md fileSKILL.md file, invoked with $name
Delegate to a subagentSubagent defined in .claude/agents/Subagent defined as TOML in .codex/agents/

A team should complete these five steps on a single repository before adding a second worktree or a second subagent.

agentic-ai
Professional Certificate

Artificial Intelligence (AI) Course

A foundational AI course covering machine learning, neural networks and applied AI tools for career-switchers and working professionals.

4.8 (86,542 ratings)  •  199,046 already enrolled  •  Beginner level

Class Starts on 11 Oct, 2026 — SAT & SUN (Weekend Batch)

Average time: 4 month(s)

Skills you’ll build: Python for AI, Machine Learning, Neural Networks, NLP Basics, AI Tools (ChatGPT, Copilot)

When Loops Go Wrong — Failure Modes and Real Fixes

Four risks increase as a loop runs with less human review: comprehension debt, intent debt, cognitive surrender, and unverified code.

Comprehension Debt

The comprehension debt is defined as the difference between the quantity of code in the system and the quantity understood by the people in the team. Comprehension debt builds up because of the fact that the agent produces code faster than the person can read and verify it, and the automated test may succeed without verification by the person of the correctness of the logic implemented.

Intent Debt

Intent debt refers to the absence of the rationale for a decision, if this rationale has not been documented. The agent without documented rationale for adherence to a particular convention in carrying out a project will end up optimizing for an outcome that is technically right but strategically wrong. Intent can be recorded either in the skill or project instruction files.

Cognitive Surrender

Cognitive surrender, therefore, refers to the full handing over of the decision-making power from an individual to the computerized system as opposed to the use of the system under continued supervision by the individual. An individual demonstrates cognitive surrender where he/she accepts the outcome of the loop without any further verification.

Unverified Code & Accountability

A checker sub-agent is still an agent rather than a person; hence, there will be no elimination of accountability from the human side. Even when a particular sub-agent produces code, a development team will be accountable for that code.

A Loop Failure, Walked Through

The most common form of failure behavior is a retry loop where a failure action is repeated without any restriction on the number of cycles.

For example, an agent tasked with fixing a test failure makes changes to the same file at every cycle without ensuring whether the last change made had succeeded in fixing the problem. In the absence of both the restriction on the number of cycles and the escalation rule, the loop will continue repeating the changes until it runs out of tokens or is stopped manually. The two measures that ensure that this does not happen include:

Human Oversight — Where People Still Belong in the Loop

Human oversight is necessary for four points in the loop: the action itself, the verification step, the output which gets sent to the end user, and any modification of the loop itself.

  • The human will review and sign off on the sensitive actions before they are carried out by the agent, e.g., transactions or production deployments.
  • The human is the grader for the workflow which does not have sufficient verification capability to make such judgments.
  • The human signs off on the output itself before it is sent to the end user in the customer-facing workflows.
  • The human reviews any modifications to the prompts or tools in the loop itself.

Beyond Code — Where Else Loop Engineering Applies

Loop Engineering is applicable to processes other than programming that exhibit similar characteristics of repetitive behavior, testability of results, and clear termination criteria.

Examples include:

  • A research bot that finds sources, tests whether each source satisfies a citation policy, and terminates after finding a minimum number of sources that pass the test.
  • A ticket classification process that categorizes incoming tickets, verifies the category using a threshold value for certainty, and refers difficult cases to a human operator.
  • A monitoring process for data pipelines that examines a data source for schema updates periodically and signals if a mismatch is found before reaching a report stage.

Frequently Asked Questions

Q1. Is loop engineering safe for production use? 

Ans. Loop engineering is considered safe for production use when there is a verification phase, cycle limit, and a point of human decision in the loop for irreversible actions.

Q2. How much does a loop cost to run compared with a single prompt? 

Ans. Loop engineering consumes more tokens than a single prompt since each verification attempt and subagent requires separate model calls. The exact multiplier varies by the amount of verification and retry attempts.

Q3. Do I need loop engineering, or is a single agent enough? 

Ans. Single-agent execution is sufficient for a one-time task with no repetition and no automated validation. Otherwise, loop engineering becomes necessary when the task becomes repetitive and produces output that can be graded.

Q4. What tools support loop engineering out of the box? 

Ans. Both Claude Code and OpenAI Codex have scheduled runs, goal-driven stop conditions, worktree isolation, and subagents at their disposal in 2026.

Q5. Is loop engineering the same as a cron job with extra steps? 

Ans. A cron job triggers execution of a task based on a schedule. Loop engineering builds on top of that by introducing goal, verification, and adjustment steps not found in a regular cron job.

Q6. Who coined the term loop engineering? 

Ans. The term itself was not coined by any single individual. Multiple people, including Anthropic researcher Boris Cherny and software developer Addy Osmani, independently described the concept of the same name within days from each other in June 2026. 

Key Takeaway

The loop substitutes the individual’s constant reminder with an automated process that evaluates itself based on its outcome and chooses its actions accordingly. However, such an automation process cannot get rid of the necessity of human assessment completely. The review happens at certain points during the process: the objective setting, the validation criteria, and every non-reversible action. A team that does not assess itself at these points will accrue intent debt and comprehension debt faster than one which retains individuals at each checkpoint.

Shalki Aggarwal is a Software Engineer II at Microsoft and an AI & Data Science expert specializing in Generative AI, Agentic AI, Python, LangChain, LangGraph, CrewAI, Deep Agents, and Loop Engineering. She is also a corporate trainer for leading organizations including L&T, Bharat Petroleum, Luminous, Denso, and Toshiba Midea, helping teams apply AI and emerging technologies to real-world business challenges.