Loop Engineering

Loop Engineering is a new concept that spread rapidly in the AI programming community in June 2026. It was systematically organized by Google engineer Addy Osmani, and Anthropic Claude Code lead Boris Cherny and developer Peter Steinberger both publicly proposed the same viewpoint.

This article will guide you from zero to understand what Loop Engineering is, its essential differences from Prompt Engineering, the six key elements that constitute a complete Agent Loop, and how to design your first runnable Loop.

The Evolution of AI Engineering

Over the past two years, the focus of AI development has been constantly shifting: from researching how to write good prompts, gradually evolving to organizing context and orchestrating tools and workflows (Harness), and then to building loop systems (Loop) that can run autonomously and continuously deliver results.

Prompt:怎么问 AI
↓
Context:给 AI 什么信息
↓
Harness:如何组织 AI 的能力
↓
Loop:如何让 AI 持续创造结果

Engineering Stage Core Idea Focus Input Content AI Capability Human Role Typical Scenario
Prompt Engineering Obtain better output by designing prompts How to ask questions Prompt / Instruction Single-turn generation Questioner Chat, writing, code generation
Context Engineering Organize and provide complete background information What information to give AI Knowledge base, history records, constraints, context Context understanding Information organizer RAG, AI search, code assistant
Harness Engineering (Orchestration Engineering) Connect models, tools, and data to form workflows How to invoke capabilities Context + API + toolchain Execute tasks System designer Agent, automated processes, multi-tool collaboration
Loop Engineering Build goal-driven autonomous closed-loop systems How to continuously complete goals Goals, state, memory, verification mechanisms Plan → Execute → Verify → Fix → Run continuously Rule maker Claude Code, AI programming, automated operations, AI employees

The Origins of Loop Engineering

Loop Engineering emerged at a clear point in time: June 2026.

The Tipping Point: Two Sentences, Millions of Shares

In early June 2026, the head of Anthropic Claude CodeBoris Chernysaid in a public speech:

"I don't prompt Claude anymore. I have loops running. They're the ones prompting Claude and figuring out what to do. My job is to write loops."
(I no longer prompt Claude directly. I have a set of Loops running that prompt Claude and decide what to do next. My job is writing Loops.)

A few days later, on June 7, 2026, developerPeter Steinberger—creator of the open-source AI Agent project OpenClaw (the fastest-starred new repository in GitHub history)—posted a twelve-character tweet:

"You shouldn't be prompting coding agents anymore. You should be designing loops that prompt your agents."
(You should no longer manually prompt AI coding assistants. You should design Loops that let Agents prompt themselves.)

  • Prompt Engineeringis "teaching AI how to do things"
  • Loop Engineeringis "designing a system where AI continuously does things itself"

These two sentences sparked huge discussion in the AI developer community because they precisely described a trend that many had felt but not yet named:Whether prompts are well-written is no longer the bottleneck; the bottleneck lies in the entire running system you design for the Agent.。

Subsequently, Addy Osmani published a long article on Substack, formally naming and systematizing this practice asLoop Engineering, making it an independent engineering discipline.

Background: Three Generations of AI Programming Tools

Loop Engineering did not come out of nowhere; it is the natural result of the evolution of AI programming tool capabilities.

StageRepresentative toolWorking methodBottleneck
First generation: Auto-completionEarly versions of GitHub CopilotCompletes the current line or function; humans lead all decisionsCan only assist, cannot act autonomously
Second generation: ConversationalChatGPT、Claude.aiAsk a question, get an answer; humans manually advance each stepHumans become the bottleneck; speed is limited by typing speed
Third generation: Agent autonomous loopClaude Code、OpenAI Codex AgentThe Agent autonomously plans, executes, verifies, and iterates in a loop until completionHow to design a system that makes Loops run reliably

The emergence of third-generation tools means that an engineer's core competency has shifted from writing prompts to designing Loops.


What is Loop Engineering

Loop Engineering is the engineering practice of designing, operating, and continuously improving feedback loops that enable AI coding Agents to autonomously plan, execute code modifications, observe results, and complete tasks across multiple iterations.

In one sentence:

Loop Engineering turns you from "the person who prompts the Agent" into "the engineer who designs the system that prompts the Agent."

What is a Loop

A Loop is a recursive goal system: you define a purpose, and the Agent iterates until the work is truly done.

Every Agent already has a built-in "inner loop" when executing tasks: Perceive → Reason → Act → Observe, then loop again.

Loop Engineering works at thenext level above:

layerWho drivesWhat is done
Inner loop (built into the Agent)The Agent itselfRead file → modify code → run tests → read errors → modify again
Outer loop (designed by you)The system you designDiscover tasks per plan → dispatch Agents → verify results → record status → start next round

You no longer sit beside the Agent, typing a command for every step. You are designing an external system that drives the inner loop for you, while you do the things that require more judgment.

The Difference Between Loop Engineering and Prompt Engineering

DimensionPrompt EngineeringLoop Engineering
Optimization targetA single instruction you write by handThe entire system that automatically decides "what to prompt, when to prompt, and whether the result is acceptable"
Unit of workOne conversation turn you manually inputA complete workflow that spans multiple turns and runs automatically
Success measurementQuality of the first replyQuality of the final output
Failure modeThe model gives a poor answerThe system's Loop is poorly designed: the loop stops too early, ignores error signals, or cannot verify completion
Perspective on the AgentA tool you holdA long-running process you schedule

Prompt Engineering is not dead. A Loop is composed of multiple Prompts, and poorly written Prompts placed into a Loop only produce bad work at a faster rate. Loop Engineering is a layer above Prompt Engineering, not a replacement for it.

The full three-layer technical stack:

LayerWhat is optimizedUnit of work
Prompt EngineeringHow to phrase a single instructionA single conversation you type manually
Context EngineeringWhat content goes into the context window: documents, history, tool definitionsThe environmental conditions surrounding a single response
Loop EngineeringA self-running system that decides what to prompt, when to prompt, and whether results are acceptableAutomated workflows spanning multiple inner loops

The Core Loop

The foundational structure of an Agent Loop consists of five stages connected end-to-end, iterating continuously.

StageEnglish nameWhat it doesTypical signal sources
IntentIntentDefine the target outcome: what success looks like, what the constraints areDeveloper or external system (Issue, CI report)
ContextContextGather relevant code, documentation, error logs, conventionsCodebase, test output, conversation history
ActionActionEdit files, run commands, call tools, draft plansAgent executes autonomously
ObservationObservationObtain test results, compilation errors, runtime output, code diffsTest framework, type checker, CI, human review
AdjustmentAdjustmentUpdate the plan based on observations, repeat the loop until the task is complete or blockedNext inner loop round

The power of a Loop lies not in any single step, but in theclosed loop. A test failure is not just an error message; it is new context. A type error is not just a blocker; it is a signal about a wrong assumption. A Code Review comment is not just feedback; it is a new observation driving the next action.


Six Core Components

A Loop that can truly run independently requires five core components, plus a memory system that runs throughout. None can be missing.

Component 1: Automations

Automations are the heartbeat of a Loop. Without automatic triggering, a Loop is just "an operation you did once," not a true cycle.

Triggers define:When? Do what?

In Claude Code, use/loopcommand to create a scheduled Loop:

Example

# Run every weekday at 9 AM: read yesterday's CI failures and Issues,
# Write findings to TODO.md, and draft fixes for issues marked as quick-win
/loop "Read yesterday's CI failures and open issues, write findings \
      to TODO.md, and draft fixes for anything labeled quick-win"
\
       --schedule "0 9 * * 1-5"

# /goal: run until a verifiable condition is met
# The following command will keep running until all tests for the auth module pass and lint is clean
/goal "All tests in test/auth pass and lint is clean"

In OpenAI Codex, Automations has a dedicated UI panel where you can set the project, prompt, cadence, and whether to run in a local checkout or a background Worktree. Runs that find content go into a Triage inbox; runs that find nothing are archived automatically.

Token cost warning:Scheduled Loops with verification sub-agents consume tokens on every trigger, and consumption varies significantly with task complexity. It's recommended to start with a slower cadence (e.g., once a day), observe costs for a few days, then increase frequency.

Component 2: Parallel Isolation (Worktrees)

When you run multiple Agents at the same time, file conflicts are the most likely source of problems. Two Agents writing to the same file simultaneously is like two engineers committing to the same line of code—the result is catastrophic.

Git Worktreeis the solution: give each Agent its own working directory, operate on separate branches, share the same Git history, but keep file changes completely isolated.

Example

# Manually create a worktree (both Claude Code and Codex support automatic management)
git worktree add ../agent-fix-auth feature/fix-auth-tests
git worktree add ../agent-upgrade-deps feature/upgrade-axios

# In Claude Code, configure worktree isolation for sub-agents
# Add to the frontmatter of .claude/agents/reviewer.md:
# isolation: worktree → each Agent gets an independent checkout, automatically cleaned up when done

Worktrees eliminate mechanical file conflicts, but they cannot eliminatethe review bottleneck. The speed at which you process and approve code changes is the real ceiling on how many Agents you can run in parallel, not how many Worktrees the tool can open.

Component 3: Skills Files

Skills (skill files) address a cost that is wasted every time:With every new conversation, the Agent has to infer your project conventions from scratch。

A Skill is a package containingSKILL.mdThe folder of files, which documents project conventions, build steps, and knowledge like "we don't do it this way because of that incident". The Agent loads the Skill at the start of each session instead of re-guessing every time.

Example

<!-- File path: .claude/skills/project-conventions/SKILL.md -->
---
name
: project-conventions
description
: Project coding standards and build steps. Any task involving code modifications should load this skill.
---

## Tech Stack
- Backend: Node.js20 + TypeScript 5.4 + Fastify
- Database: PostgreSQL16, ORM uses Drizzle
- Testing: Vitest, test files placed under __tests__/ at the same level as src/

## Build Commands
- Install dependencies: pnpm install
- Run tests: pnpm test
- Type check: pnpm typecheck
- Build artifacts: pnpm build

## Core Conventions
- All database queries must go through wrapper functions in src/db/queries/; writing raw SQL directly in the business layer is prohibited
- Use the AppError class uniformly for errors (src/errors/AppError.ts); throwing bare strings is prohibited
- API route file naming:{resource}.routes.ts, placed in src/routes/
- New APIs must also update docs/api.md

## Prohibitions
- Do not directly modify historical migration files under the migrations/ directory
- Do not commit real secrets in the .env file; use .env.example as a placeholder

Component 4: Connectors / MCP

A Loop that can only see the local filesystem can do very limited things. Connectors (based on MCP—Model Context Protocol) enable the Agent to read issue trackers, query databases, call APIs, and send messages in Slack.

This is the core difference between "the Agent saying 'here's the fix'" and "the Loop automatically opening a PR, linking the ticket, and notifying the channel after CI passes."

Both Claude Code and OpenAI Codex natively support the MCP protocol, and connectors written for one tool can generally be used in both.

Example: Configuring MCP Connectors in Claude Code

// File path: .claude/mcp.json
// Configure MCP connectors so the Agent can operate GitHub and send Slack notifications
{
  "connectors": [
    {
      "name": "github",
      "command": "npx",
      "args": ["-y", "@modelcontextprotocol/server-github"],
      "env": {
        "GITHUB_TOKEN": "${GITHUB_TOKEN}"
      }
    },
    {
      "name": "slack",
      "command": "npx",
      "args": ["-y", "@modelcontextprotocol/server-slack"],
      "env": {
        "SLACK_BOT_TOKEN": "${SLACK_BOT_TOKEN}"
      }
    }
  ]
}

Connecting connectors to the Agent means it can take real actions in real environments. Each connector must be configured with least privilege, and high-risk operations (pushing code, merging PRs, sending external notifications) must require human approval; they cannot be executed fully automatically.

Component 5: Sub-Agents

One of the most important architectural decisions in the Loop:Separate the Agent that writes code from the Agent that checks code。

The model that wrote the code is too lenient when grading its own work. An independent checker with different instructions—sometimes using a different model—can catch problems that the first Agent rationalized away.

This "Maker-Checker" pattern is also applied to the Loop's stopping condition: Claude Code's/goalcommand uses a separate model after each iteration to judge whether it is "done", rather than letting the model that did the work make that judgment.

Example: Defining a Reviewer Sub-Agent

<!-- File path: .claude/agents/spec-reviewer.md -->
---
name
: spec-reviewer
description
: Adversarially review completed code changes to verify they meet specifications and test requirements.
model
: opus            # Use a stronger model as the verifier
isolation
: worktree    # Check out independently to avoid polluting the maker's workspace
---

You are an adversarial code reviewer. Your job is not to approve, but to question.

After receiving a diff, you need to:
1. Run the full test suite (pnpm test), record failures
2. Run type checking (pnpm typecheck)
3. Check the diff against the project specifications in SKILL.md
4. Verify that user-visible behavior changes meet the original requirements

Only output APPROVED when all tests pass, type checking is clean, and there are no specification violations.
Otherwise, output specific rejection reasons; do not give vague feedback.

Component 6: Persistent Memory

Memory is the core pillar of any Loop that can run persistently across conversations.

The model completely forgets between conversations. Without external memory, every time the Loop triggers, the Agent starts from zero, not knowing what was done yesterday, which fixes have been merged, and which tasks are still open.

The solution is extremely simple:Write state to files, and keep the files in the repository.. The repository remembers even if the model does not.

Example: Loop State File

<!-- File path: TODO.md (Loop's status file, committed with the code) -->
# Loop Task Status

Last updated:2026-06-1409:03 UTC (updated by automatic Loop)

## In Progress
- [ ]Flaky test in test/auth/login.spec.ts (CI Run#4821, failed 3 times)
- Hypothesis: session state leakage between concurrent tests
- Tried: isolating test database connection → ineffective
- Next step: check cleanup logic in beforeEach

## Pending
- [ ]Upgrade axios to1.7.x (security vulnerability CVE-2026-xxxxx)
- [ ]API documentation update (PR#308 is behind the code after merge)

## Completed
- [x]Fixed the billing module [issue] caused by company names containing single quotes500error (PR#312, merged)
- [x]Upgrade Node.js version to20.x(PR #307, merged)

Five Common Loop Patterns

Different types of tasks require different feedback signals and stopping conditions, therefore several classic Loop patterns have emerged.

PatternCore observation signalStopping conditionTypical scenario
Test-driven LoopTest pass / failAll target tests passBug fixes, regression tests, data transformation logic
Compiler-driven LoopType errors, compilation error listZero type-check errorsTypeScript migration, dependency upgrades, refactoring
Review-driven LoopHuman review commentsAll comments processed or documented as ignoredMechanical follow-up on PR reviews
Runtime debugging LoopLogs, stack traces, HTTP responsesReproduce issue → form hypothesis → verify fixProduction bugs, performance issues, API exceptions
Product iteration LoopScreenshots, browser checks, accessibility reportsAligned with design mockups, responsive correct, no spec violationsLanding pages, UI adjustments, marketing components

Building Your First Loop

Don't try to build a fully automated, multi-agent Loop that auto-merges PRs from the start. Start with the smallest usable Loop, understand its behavior, then gradually expand.

Step 1: Start with a Narrow Task

The narrower the task, the clearer the Agent is about which files matter and which validation signals are relevant.

Bad task definitionGood task definition
"Optimize dashboard performance""Reduce the dashboard's initial load time by 30% by deferring non-critical chart loading while keeping existing filters working"
"Fix checkout issues""Fix the failing tax calculation test in test/checkout/tax.spec.ts"
"Improve the settings page""Fix the layout issue where the account deletion button is truncated on mobile (375px width)"

Step 2: Clearly Specify the Verification Method

The Loop needs to know what "done" means. If you know the verification command, write it directly into the instructions.

Example: Loop Instructions with Verification Conditions

# Bad: no verification condition, the Agent doesn't know when to stop
/goal "Fix the auth bug"

# Good: has specific verification commands and success criteria
/goal "Fix the session leak causing flaky tests in test/auth/login.spec.ts.
       Success condition: run 'pnpm test test/auth/login.spec.ts' 5 times
       consecutively with zero failures. Do not touch other test files."

Step 3: Set Up Safety Mechanisms

Before the Loop can automatically open PRs, have it only write files and perform no external actions; you review the diff and decide whether to commit.

Example: Daily CI Failure Triage Loop (Minimal Safe Version)

# This is the recommended first Loop:
# - Read-only operations (read CI logs, read Issues)
# - Only write TODO.md (touch no other files)
# - No PRs, no code merges
# - You check TODO.md every morning and manually decide the next step

/loop "Read yesterday's CI failure logs and GitHub Issues labeled 'bug'.
       Categorize findings by likely cause.
       Write a summary with prioritized action items to TODO.md.
       Do NOT edit any source files. Do NOT open any PRs."
\
       --schedule "0 8 * * 1-5"

This minimal version of the Loop is already valuable: it organizes yesterday's issue list for you every morning, and when you get to the office you can open TODO.md and know what to do today instead of spending half an hour digging through CI logs. This is the best starting point for understanding Loop behavior, with almost zero risk.

Step 4: Gradually Increase Autonomy

StageWhat the Loop can doWhat the human does
Stage 1 (read-only)Detect issues, triage tasks, write status filesReview TODO.md, manually decide the order of handling
Stage 2 (draft)Draft fixes, run tests, write to a branchReview the diff, manually execute git push
Stage 3 (Semi-automatic)Open a draft PR, run CI, notify SlackReview the PR, manually click Merge
Stage 4 (Fully automatic)Producer + Reviewer dual agents, auto-merge after CI passesHuman intervention when anomalies occur, periodically audit merge history

Three Major Risks of Loop Engineering

Loop changes the way you work, but it doesn't eliminate the engineer's responsibility. There are three problems that become more severe as Loop improves, not easier.

Risk 1: Verification Remains Your Responsibility

A Loop that runs unattended is also a Loop that makes mistakes unattended. Setting up a reviewer sub-agent is a good way to reduce risk, but "passed verification" is a claim, not proof.

No matter how reliable the Loop is, human review of merged code cannot disappear.

Risk 2: Comprehension Debt Accumulates Faster

The faster Loop produces code, the lower the proportion of code you actually understand. This isn't a problem unique to AI programming, but Loop has accelerated it.

The only antidote is:Read the code Loop producesDon't stop understanding what it's doing just because Loop runs smoothly.

Risk 3: Cognitive Surrender

When Loop runs automatically, accepting whatever result it returns is the most comfortable choice. This is the most hidden danger of Loop Engineering.

Two people can build exactly the same Loop yet get completely opposite results: one uses it to advance faster on the basis of deep understanding, the other uses it to avoid the work of understanding itself. Loop doesn't know the difference. You do.

Four types of Loop failure modes and countermeasures:

Failure modeSymptomsRoot causeSolution
ThrashingAgent repeatedly modifies code, but doesn't convergeUnclear goals, noisy verification signals, or each modification's scope is too largeNarrow the goal scope, reduce the size of each diff, use more reliable verification commands
Overfitting testsAll tests pass, but the functionality is actually wrongTest coverage is too narrow, real user scenarios are not verifiedCombine automated tests with manual acceptance checks, add end-to-end tests
Context driftAgent keeps working based on stale assumptions, ignoring newly emerging changesContext is not refreshed after key observationsAfter important observations, re-collect context; don't treat the initial plan as sacred
Unsafe autonomyAgent performs destructive operations without authorizationPermission scope too broad with no clear stopping conditionsPrinciple of least privilege; high-risk operations require human approval; set clear stopping rules

Best Practices Summary

The following are core principles for building reliable Loops; each corresponds to a common class of failure.

PrinciplesSpecific practices
Start with narrow tasksDefine only one clearly scoped goal at a time; broad goals make it impossible for the Agent to determine which files and which validation signals are relevant
Tell the Agent how to validateWrite the validation command directly in the instructions (pnpm test test/auth), acceptance scenarios, or API endpoints, making "completion" measurable
Prefer small, reversible changesRequire the Agent to make the smallest coherent modification, run validation, then expand; large-scale speculative rewrites make it hard to tell which assumption went wrong
Respect existing code patternsHave the Agent first inspect adjacent implementations, reuse existing components, follow existing naming conventions, and avoid introducing unnecessary new abstractions
Keep humans in the judgment seatThe Agent handles evidence collection and mechanical fixes; product judgment, architectural decisions, and final review remain with humans
Codify reusable LoopsWhen a Loop runs well, solidify it into a Skill file or standardized trigger to reduce future repetition costs
Other extensions