Loop Engineering
Loop Engineering is a new concept that spread rapidly in the AI programming community in June 2026. It was systematically organized by Google engineer Addy Osmani, and Anthropic Claude Code lead Boris Cherny and developer Peter Steinberger both publicly proposed the same viewpoint.
This article will guide you from zero to understand what Loop Engineering is, its essential differences from Prompt Engineering, the six key elements that constitute a complete Agent Loop, and how to design your first runnable Loop.
The Evolution of AI Engineering
Over the past two years, the focus of AI development has been constantly shifting: from researching how to write good prompts, gradually evolving to organizing context and orchestrating tools and workflows (Harness), and then to building loop systems (Loop) that can run autonomously and continuously deliver results.
Prompt:怎么问 AI ↓ Context:给 AI 什么信息 ↓ Harness:如何组织 AI 的能力 ↓ Loop:如何让 AI 持续创造结果

| Engineering Stage | Core Idea | Focus | Input Content | AI Capability | Human Role | Typical Scenario |
|---|---|---|---|---|---|---|
| Prompt Engineering | Obtain better output by designing prompts | How to ask questions | Prompt / Instruction | Single-turn generation | Questioner | Chat, writing, code generation |
| Context Engineering | Organize and provide complete background information | What information to give AI | Knowledge base, history records, constraints, context | Context understanding | Information organizer | RAG, AI search, code assistant |
| Harness Engineering (Orchestration Engineering) | Connect models, tools, and data to form workflows | How to invoke capabilities | Context + API + toolchain | Execute tasks | System designer | Agent, automated processes, multi-tool collaboration |
| Loop Engineering | Build goal-driven autonomous closed-loop systems | How to continuously complete goals | Goals, state, memory, verification mechanisms | Plan → Execute → Verify → Fix → Run continuously | Rule maker | Claude Code, AI programming, automated operations, AI employees |
The Origins of Loop Engineering
Loop Engineering emerged at a clear point in time: June 2026.
The Tipping Point: Two Sentences, Millions of Shares
In early June 2026, the head of Anthropic Claude CodeBoris Chernysaid in a public speech:
"I don't prompt Claude anymore. I have loops running. They're the ones prompting Claude and figuring out what to do. My job is to write loops."
(I no longer prompt Claude directly. I have a set of Loops running that prompt Claude and decide what to do next. My job is writing Loops.)
A few days later, on June 7, 2026, developerPeter Steinberger—creator of the open-source AI Agent project OpenClaw (the fastest-starred new repository in GitHub history)—posted a twelve-character tweet:
"You shouldn't be prompting coding agents anymore. You should be designing loops that prompt your agents."
(You should no longer manually prompt AI coding assistants. You should design Loops that let Agents prompt themselves.)
- Prompt Engineeringis "teaching AI how to do things"
- Loop Engineeringis "designing a system where AI continuously does things itself"

These two sentences sparked huge discussion in the AI developer community because they precisely described a trend that many had felt but not yet named:Whether prompts are well-written is no longer the bottleneck; the bottleneck lies in the entire running system you design for the Agent.。
Subsequently, Addy Osmani published a long article on Substack, formally naming and systematizing this practice asLoop Engineering, making it an independent engineering discipline.
Background: Three Generations of AI Programming Tools
Loop Engineering did not come out of nowhere; it is the natural result of the evolution of AI programming tool capabilities.
| Stage | Representative tool | Working method | Bottleneck |
|---|---|---|---|
| First generation: Auto-completion | Early versions of GitHub Copilot | Completes the current line or function; humans lead all decisions | Can only assist, cannot act autonomously |
| Second generation: Conversational | ChatGPT、Claude.ai | Ask a question, get an answer; humans manually advance each step | Humans become the bottleneck; speed is limited by typing speed |
| Third generation: Agent autonomous loop | Claude Code、OpenAI Codex Agent | The Agent autonomously plans, executes, verifies, and iterates in a loop until completion | How to design a system that makes Loops run reliably |
The emergence of third-generation tools means that an engineer's core competency has shifted from writing prompts to designing Loops.
What is Loop Engineering
Loop Engineering is the engineering practice of designing, operating, and continuously improving feedback loops that enable AI coding Agents to autonomously plan, execute code modifications, observe results, and complete tasks across multiple iterations.
In one sentence:
Loop Engineering turns you from "the person who prompts the Agent" into "the engineer who designs the system that prompts the Agent."
What is a Loop
A Loop is a recursive goal system: you define a purpose, and the Agent iterates until the work is truly done.
Every Agent already has a built-in "inner loop" when executing tasks: Perceive → Reason → Act → Observe, then loop again.
Loop Engineering works at thenext level above:
| layer | Who drives | What is done |
|---|---|---|
| Inner loop (built into the Agent) | The Agent itself | Read file → modify code → run tests → read errors → modify again |
| Outer loop (designed by you) | The system you design | Discover tasks per plan → dispatch Agents → verify results → record status → start next round |
You no longer sit beside the Agent, typing a command for every step. You are designing an external system that drives the inner loop for you, while you do the things that require more judgment.
The Difference Between Loop Engineering and Prompt Engineering
| Dimension | Prompt Engineering | Loop Engineering |
|---|---|---|
| Optimization target | A single instruction you write by hand | The entire system that automatically decides "what to prompt, when to prompt, and whether the result is acceptable" |
| Unit of work | One conversation turn you manually input | A complete workflow that spans multiple turns and runs automatically |
| Success measurement | Quality of the first reply | Quality of the final output |
| Failure mode | The model gives a poor answer | The system's Loop is poorly designed: the loop stops too early, ignores error signals, or cannot verify completion |
| Perspective on the Agent | A tool you hold | A long-running process you schedule |

Prompt Engineering is not dead. A Loop is composed of multiple Prompts, and poorly written Prompts placed into a Loop only produce bad work at a faster rate. Loop Engineering is a layer above Prompt Engineering, not a replacement for it.
The full three-layer technical stack:
| Layer | What is optimized | Unit of work |
|---|---|---|
| Prompt Engineering | How to phrase a single instruction | A single conversation you type manually |
| Context Engineering | What content goes into the context window: documents, history, tool definitions | The environmental conditions surrounding a single response |
| Loop Engineering | A self-running system that decides what to prompt, when to prompt, and whether results are acceptable | Automated workflows spanning multiple inner loops |
The Core Loop
The foundational structure of an Agent Loop consists of five stages connected end-to-end, iterating continuously.
| Stage | English name | What it does | Typical signal sources |
|---|---|---|---|
| Intent | Intent | Define the target outcome: what success looks like, what the constraints are | Developer or external system (Issue, CI report) |
| Context | Context | Gather relevant code, documentation, error logs, conventions | Codebase, test output, conversation history |
| Action | Action | Edit files, run commands, call tools, draft plans | Agent executes autonomously |
| Observation | Observation | Obtain test results, compilation errors, runtime output, code diffs | Test framework, type checker, CI, human review |
| Adjustment | Adjustment | Update the plan based on observations, repeat the loop until the task is complete or blocked | Next inner loop round |
The power of a Loop lies not in any single step, but in theclosed loop. A test failure is not just an error message; it is new context. A type error is not just a blocker; it is a signal about a wrong assumption. A Code Review comment is not just feedback; it is a new observation driving the next action.
Six Core Components
A Loop that can truly run independently requires five core components, plus a memory system that runs throughout. None can be missing.
Component 1: Automations
Automations are the heartbeat of a Loop. Without automatic triggering, a Loop is just "an operation you did once," not a true cycle.
Triggers define:When? Do what?
In Claude Code, use/loopcommand to create a scheduled Loop:
Example
# Write findings to TODO.md, and draft fixes for issues marked as quick-win
/loop "Read yesterday's CI failures and open issues, write findings \
to TODO.md, and draft fixes for anything labeled quick-win" \
--schedule "0 9 * * 1-5"
# /goal: run until a verifiable condition is met
# The following command will keep running until all tests for the auth module pass and lint is clean
/goal "All tests in test/auth pass and lint is clean"
In OpenAI Codex, Automations has a dedicated UI panel where you can set the project, prompt, cadence, and whether to run in a local checkout or a background Worktree. Runs that find content go into a Triage inbox; runs that find nothing are archived automatically.
Token cost warning:Scheduled Loops with verification sub-agents consume tokens on every trigger, and consumption varies significantly with task complexity. It's recommended to start with a slower cadence (e.g., once a day), observe costs for a few days, then increase frequency.
Component 2: Parallel Isolation (Worktrees)
When you run multiple Agents at the same time, file conflicts are the most likely source of problems. Two Agents writing to the same file simultaneously is like two engineers committing to the same line of code—the result is catastrophic.
Git Worktreeis the solution: give each Agent its own working directory, operate on separate branches, share the same Git history, but keep file changes completely isolated.
Example
git worktree add ../agent-fix-auth feature/fix-auth-tests
git worktree add ../agent-upgrade-deps feature/upgrade-axios
# In Claude Code, configure worktree isolation for sub-agents
# Add to the frontmatter of .claude/agents/reviewer.md:
# isolation: worktree → each Agent gets an independent checkout, automatically cleaned up when done
Worktrees eliminate mechanical file conflicts, but they cannot eliminatethe review bottleneck. The speed at which you process and approve code changes is the real ceiling on how many Agents you can run in parallel, not how many Worktrees the tool can open.
Component 3: Skills Files
Skills (skill files) address a cost that is wasted every time:With every new conversation, the Agent has to infer your project conventions from scratch。
A Skill is a package containingSKILL.mdThe folder of files, which documents project conventions, build steps, and knowledge like "we don't do it this way because of that incident". The Agent loads the Skill at the start of each session instead of re-guessing every time.
Example
---
name: project-conventions
description: Project coding standards and build steps. Any task involving code modifications should load this skill.
---
## Tech Stack
- Backend: Node.js20 + TypeScript 5.4 + Fastify
- Database: PostgreSQL16, ORM uses Drizzle
- Testing: Vitest, test files placed under __tests__/ at the same level as src/
## Build Commands
- Install dependencies: pnpm install
- Run tests: pnpm test
- Type check: pnpm typecheck
- Build artifacts: pnpm build
## Core Conventions
- All database queries must go through wrapper functions in src/db/queries/; writing raw SQL directly in the business layer is prohibited
- Use the AppError class uniformly for errors (src/errors/AppError.ts); throwing bare strings is prohibited
- API route file naming:{resource}.routes.ts, placed in src/routes/
- New APIs must also update docs/api.md
## Prohibitions
- Do not directly modify historical migration files under the migrations/ directory
- Do not commit real secrets in the .env file; use .env.example as a placeholder
Component 4: Connectors / MCP
A Loop that can only see the local filesystem can do very limited things. Connectors (based on MCP—Model Context Protocol) enable the Agent to read issue trackers, query databases, call APIs, and send messages in Slack.
This is the core difference between "the Agent saying 'here's the fix'" and "the Loop automatically opening a PR, linking the ticket, and notifying the channel after CI passes."
Both Claude Code and OpenAI Codex natively support the MCP protocol, and connectors written for one tool can generally be used in both.
Example: Configuring MCP Connectors in Claude Code
// Configure MCP connectors so the Agent can operate GitHub and send Slack notifications
{
"connectors": [
{
"name": "github",
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-github"],
"env": {
"GITHUB_TOKEN": "${GITHUB_TOKEN}"
}
},
{
"name": "slack",
"command": "npx",
"args": ["-y", "@modelcontextprotocol/server-slack"],
"env": {
"SLACK_BOT_TOKEN": "${SLACK_BOT_TOKEN}"
}
}
]
}
Connecting connectors to the Agent means it can take real actions in real environments. Each connector must be configured with least privilege, and high-risk operations (pushing code, merging PRs, sending external notifications) must require human approval; they cannot be executed fully automatically.
Component 5: Sub-Agents
One of the most important architectural decisions in the Loop:Separate the Agent that writes code from the Agent that checks code。
The model that wrote the code is too lenient when grading its own work. An independent checker with different instructions—sometimes using a different model—can catch problems that the first Agent rationalized away.
This "Maker-Checker" pattern is also applied to the Loop's stopping condition: Claude Code's/goalcommand uses a separate model after each iteration to judge whether it is "done", rather than letting the model that did the work make that judgment.
Example: Defining a Reviewer Sub-Agent
---
name: spec-reviewer
description: Adversarially review completed code changes to verify they meet specifications and test requirements.
model: opus # Use a stronger model as the verifier
isolation: worktree # Check out independently to avoid polluting the maker's workspace
---
You are an adversarial code reviewer. Your job is not to approve, but to question.
After receiving a diff, you need to:
1. Run the full test suite (pnpm test), record failures
2. Run type checking (pnpm typecheck)
3. Check the diff against the project specifications in SKILL.md
4. Verify that user-visible behavior changes meet the original requirements
Only output APPROVED when all tests pass, type checking is clean, and there are no specification violations.
Otherwise, output specific rejection reasons; do not give vague feedback.
Component 6: Persistent Memory
Memory is the core pillar of any Loop that can run persistently across conversations.
The model completely forgets between conversations. Without external memory, every time the Loop triggers, the Agent starts from zero, not knowing what was done yesterday, which fixes have been merged, and which tasks are still open.
The solution is extremely simple:Write state to files, and keep the files in the repository.. The repository remembers even if the model does not.
Example: Loop State File
# Loop Task Status
Last updated:2026-06-1409:03 UTC (updated by automatic Loop)
## In Progress
- [ ]Flaky test in test/auth/login.spec.ts (CI Run#4821, failed 3 times)
- Hypothesis: session state leakage between concurrent tests
- Tried: isolating test database connection → ineffective
- Next step: check cleanup logic in beforeEach
## Pending
- [ ]Upgrade axios to1.7.x (security vulnerability CVE-2026-xxxxx)
- [ ]API documentation update (PR#308 is behind the code after merge)
## Completed
- [x]Fixed the billing module [issue] caused by company names containing single quotes500error (PR#312, merged)
- [x]Upgrade Node.js version to20.x(PR #307, merged)
Five Common Loop Patterns
Different types of tasks require different feedback signals and stopping conditions, therefore several classic Loop patterns have emerged.
| Pattern | Core observation signal | Stopping condition | Typical scenario |
|---|---|---|---|
| Test-driven Loop | Test pass / fail | All target tests pass | Bug fixes, regression tests, data transformation logic |
| Compiler-driven Loop | Type errors, compilation error list | Zero type-check errors | TypeScript migration, dependency upgrades, refactoring |
| Review-driven Loop | Human review comments | All comments processed or documented as ignored | Mechanical follow-up on PR reviews |
| Runtime debugging Loop | Logs, stack traces, HTTP responses | Reproduce issue → form hypothesis → verify fix | Production bugs, performance issues, API exceptions |
| Product iteration Loop | Screenshots, browser checks, accessibility reports | Aligned with design mockups, responsive correct, no spec violations | Landing pages, UI adjustments, marketing components |
Building Your First Loop
Don't try to build a fully automated, multi-agent Loop that auto-merges PRs from the start. Start with the smallest usable Loop, understand its behavior, then gradually expand.
Step 1: Start with a Narrow Task
The narrower the task, the clearer the Agent is about which files matter and which validation signals are relevant.
| Bad task definition | Good task definition |
|---|---|
| "Optimize dashboard performance" | "Reduce the dashboard's initial load time by 30% by deferring non-critical chart loading while keeping existing filters working" |
| "Fix checkout issues" | "Fix the failing tax calculation test in test/checkout/tax.spec.ts" |
| "Improve the settings page" | "Fix the layout issue where the account deletion button is truncated on mobile (375px width)" |
Step 2: Clearly Specify the Verification Method
The Loop needs to know what "done" means. If you know the verification command, write it directly into the instructions.
Example: Loop Instructions with Verification Conditions
/goal "Fix the auth bug"
# Good: has specific verification commands and success criteria
/goal "Fix the session leak causing flaky tests in test/auth/login.spec.ts.
Success condition: run 'pnpm test test/auth/login.spec.ts' 5 times
consecutively with zero failures. Do not touch other test files."
Step 3: Set Up Safety Mechanisms
Before the Loop can automatically open PRs, have it only write files and perform no external actions; you review the diff and decide whether to commit.
Example: Daily CI Failure Triage Loop (Minimal Safe Version)
# - Read-only operations (read CI logs, read Issues)
# - Only write TODO.md (touch no other files)
# - No PRs, no code merges
# - You check TODO.md every morning and manually decide the next step
/loop "Read yesterday's CI failure logs and GitHub Issues labeled 'bug'.
Categorize findings by likely cause.
Write a summary with prioritized action items to TODO.md.
Do NOT edit any source files. Do NOT open any PRs." \
--schedule "0 8 * * 1-5"
This minimal version of the Loop is already valuable: it organizes yesterday's issue list for you every morning, and when you get to the office you can open TODO.md and know what to do today instead of spending half an hour digging through CI logs. This is the best starting point for understanding Loop behavior, with almost zero risk.
Step 4: Gradually Increase Autonomy
| Stage | What the Loop can do | What the human does |
|---|---|---|
| Stage 1 (read-only) | Detect issues, triage tasks, write status files | Review TODO.md, manually decide the order of handling |
| Stage 2 (draft) | Draft fixes, run tests, write to a branch | Review the diff, manually execute git push |
| Stage 3 (Semi-automatic) | Open a draft PR, run CI, notify Slack | Review the PR, manually click Merge |
| Stage 4 (Fully automatic) | Producer + Reviewer dual agents, auto-merge after CI passes | Human intervention when anomalies occur, periodically audit merge history |
Three Major Risks of Loop Engineering
Loop changes the way you work, but it doesn't eliminate the engineer's responsibility. There are three problems that become more severe as Loop improves, not easier.
Risk 1: Verification Remains Your Responsibility
A Loop that runs unattended is also a Loop that makes mistakes unattended. Setting up a reviewer sub-agent is a good way to reduce risk, but "passed verification" is a claim, not proof.
No matter how reliable the Loop is, human review of merged code cannot disappear.
Risk 2: Comprehension Debt Accumulates Faster
The faster Loop produces code, the lower the proportion of code you actually understand. This isn't a problem unique to AI programming, but Loop has accelerated it.
The only antidote is:Read the code Loop producesDon't stop understanding what it's doing just because Loop runs smoothly.
Risk 3: Cognitive Surrender
When Loop runs automatically, accepting whatever result it returns is the most comfortable choice. This is the most hidden danger of Loop Engineering.
Two people can build exactly the same Loop yet get completely opposite results: one uses it to advance faster on the basis of deep understanding, the other uses it to avoid the work of understanding itself. Loop doesn't know the difference. You do.
Four types of Loop failure modes and countermeasures:
| Failure mode | Symptoms | Root cause | Solution |
|---|---|---|---|
| Thrashing | Agent repeatedly modifies code, but doesn't converge | Unclear goals, noisy verification signals, or each modification's scope is too large | Narrow the goal scope, reduce the size of each diff, use more reliable verification commands |
| Overfitting tests | All tests pass, but the functionality is actually wrong | Test coverage is too narrow, real user scenarios are not verified | Combine automated tests with manual acceptance checks, add end-to-end tests |
| Context drift | Agent keeps working based on stale assumptions, ignoring newly emerging changes | Context is not refreshed after key observations | After important observations, re-collect context; don't treat the initial plan as sacred |
| Unsafe autonomy | Agent performs destructive operations without authorization | Permission scope too broad with no clear stopping conditions | Principle of least privilege; high-risk operations require human approval; set clear stopping rules |
Best Practices Summary
The following are core principles for building reliable Loops; each corresponds to a common class of failure.
| Principles | Specific practices |
|---|---|
| Start with narrow tasks | Define only one clearly scoped goal at a time; broad goals make it impossible for the Agent to determine which files and which validation signals are relevant |
| Tell the Agent how to validate | Write the validation command directly in the instructions (pnpm test test/auth), acceptance scenarios, or API endpoints, making "completion" measurable |
| Prefer small, reversible changes | Require the Agent to make the smallest coherent modification, run validation, then expand; large-scale speculative rewrites make it hard to tell which assumption went wrong |
| Respect existing code patterns | Have the Agent first inspect adjacent implementations, reuse existing components, follow existing naming conventions, and avoid introducing unnecessary new abstractions |
| Keep humans in the judgment seat | The Agent handles evidence collection and mechanical fixes; product judgment, architectural decisions, and final review remain with humans |
| Codify reusable Loops | When a Loop runs well, solidify it into a Skill file or standardized trigger to reduce future repetition costs |