Agent Architecture
An Agent (intelligent agent) refers to an AI system that can autonomously perceive its environment, reason, and take actions to achieve goals.
From the simplest loop to multi-agent collaboration, this article gives you a thorough understanding of the principles, diagrams, and applicable scenarios of six mainstream architectures, helping you make appropriate technical choices in real projects.

What is Agent Architecture
In AI application development,Agent (intelligent agent)Refers to an AI system that can perceive its environment, make autonomous decisions, and take actions.
Unlike traditional "question-answer" style large model calls, an Agent can continuously perform multi-step operations, call tools, and even coordinate other Agents to complete complex tasks.
Agent ArchitectureRefers to the organization of components in an Agent system, which determines the capability boundaries, reliability, flexibility, and applicable scenarios of the Agent.
The way an Agent works is essentially aLoop— perceive the current state, reason about the next step, execute an action, perceive again... until the task is complete.
The difference between architectures lies in how to organize and extend this basic loop.
This article assumes you are already familiar with the basic concepts of large language models (LLMs) and the basic idea of tool calling (Function Calling). If you are not familiar yet, you can learn these two fundamentals first.
Architecture 1: Single Agent Loop
Best for beginners Simple to implement
The most basic and intuitive architecture: one Agent independently completes all tasks from start to finish.
The single Agent loop directly embodies theReAct pattern(Reasoning + Acting): every step is "think first, then act." The LLM serves as the brain, and tool calling is its hands.
How it works
Perceive:Read the current state — file contents, environment variables, output from previous steps — and integrate it into the current context.
Reason:The LLM decides the next action based on the context — which tool to call, what parameters to pass — or determines whether the task is complete.
Act:Execute tool calls, such as reading/writing files or searching the web. The tool's execution results are appended to the context, and it proceeds to the next loop iteration.
Every tool call result is written back to the context window. Therefore, as the task progresses, the context keeps growing until it reaches the LLM's context window limit — this is the primary bottleneck of the single Agent loop.
Example
class SimpleAgent:
"""Basic structure of the single Agent loop"""
def __init__(self, model, tools, max_turns=10):
self.model = model # Large language model
self.tools = tools # Available tool list
self.max_turns = max_turns # Maximum loop iterations, preventing infinite loop
def run(self, task: str) -> str:
"""Main loop for executing the task"""
context = f"User task: {task}"
for turn in range(self.max_turns):
# Step 1: Think — let the model decide the next step
response = self.model.think(context)
# If the model believes the task is complete, return the final answer
if response.is_final():
return response.content
# Step 2: Act — call the tool chosen by the model
tool_name = response.tool_name
tool_args = response.tool_args
tool_result = self.tools[tool_name](**tool_args)
# Step 3: Feed the tool result back to the model and proceed to the next round
context += f"\nTool {tool_name} returned: {tool_result}"
return "Maximum rounds reached, task not completed"
# Usage example
agent = SimpleAgent(model=llm, tools={
"read_file": read_file,
"search_code": search_code,
"run_test": run_test
})
result = agent.run("Fix the type error in user.py in the example project")
Advantages
- Simplest to implement, easy to debug
- Suitable for scenarios with clear task boundaries
- Supported by almost all Agent frameworks
Disadvantages
- Context window easily fills up
- Prone to "going off course" in complex tasks
- Cannot process multiple subtasks in parallel
Best use cases:Fix a bug, write a function, answer a specific question. The task is clear, with moderate complexity, and does not require parallelism or multi-role collaboration.
If your task is expected to require more than 15 rounds of tool calls, a single Agent loop may not be the best choice. Consider using multi-Agent collaboration or a plan-and-execute architecture to break down the complexity.
Architecture 2: Plan & Execute
Intuitive Reviewable
Separates "figuring out what to do" and "actually doing it" into two independent phases, improving task predictability and auditability.
The Plan + Execute architecture splits the Agent's work into two distinct phases: firstPlan, thenExecute。
In the planning phase, the model does not perform any actions; it only generates a detailed list of execution steps. In the execution phase, the system completes each step in order. This separation allows users to review the plan before execution, similar to Claude Code's Plan Mode.
Two variants
| Variant | Behavior | Typical scenario |
|---|---|---|
| Static planning | The plan is generated all at once and executed linearly in order without mid-course adjustments. | Tasks with fixed processes and well-defined steps, such as data migration scripts. |
| Dynamic planning | Re-evaluate after each step and adjust the subsequent plan based on the results. | Tasks with uncertain outcomes, such as debugging and exploratory data analysis. |
Dynamic planning is more robust, but it is more complex to implement, and re-planning at each step consumes extra tokens.
Claude Code's Plan Mode embodies this architecture—after clicking "Plan", the AI first outputs a detailed plan for your review, and only begins execution after confirmation. This greatly enhances the user's sense of control.
Example
class PlanExecuteAgent:
"""Agent that plans first, then executes"""
def plan(self, task: str) -> list:
"""Phase 1: Generate an execution plan"""
plan = self.model.generate(f"""
Please break down the following task into a list of executable steps:
Task: {task}
Return a JSON-formatted list of steps, each containing:
- step_id: step number
- description: step description
- tool: name of the tool to call
""")
return plan
def execute(self, plan: list, dynamic: bool = False) -> str:
"""Phase 2: Execute the plan step by step"""
results = []
remaining_plan = plan.copy()
while remaining_plan:
step = remaining_plan.pop(0)
output = self.tools[step["tool"]](step["description"])
results.append({"step": step["step_id"], "output": output})
if dynamic and remaining_plan:
# Dynamic planning: re-evaluate the subsequent plan based on current results
remaining_plan = self.replan(remaining_plan, results)
return self.summarize(results)
# Usage example
agent = PlanExecuteAgent()
plan = agent.plan("Add user authentication functionality to the example project")
# Humans can review the plan first and execute after confirming it is reasonable
result = agent.execute(plan, dynamic=True)
Advantages
- The plan can be manually reviewed before execution.
- Clear separation of reasoning and execution responsibilities.
- Friendly for long tasks.
Disadvantages
- The initial plan may not be accurate enough.
- The two-phase approach adds latency.
- The static version struggles to handle unexpected situations.
The cost of Plan & Execute is increased reasoning rounds, which is wasteful for simple tasks. If a task can be completed in 3 steps, using a single-agent loop directly is more efficient.
Architecture 3: Multi-Agent Collaboration
Production recommendation Complex tasks
An Orchestrator is responsible for task decomposition and scheduling; multiple Subagents each perform their own duties, completing subtasks in parallel or sequentially, and results are gathered back to the Orchestrator for synthesis.
When a single Agent faces an insufficient context window or overly complex tasks, the multi-Agent architecture offers an elegant solution:Have multiple specialized sub-Agents work in parallel, while an Orchestrator coordinates the overall situation.
Independent context is a core advantage
Each sub-Agent has anindependent context window. The code review Agent's deep reading of auth.py does not affect the performance analysis Agent's judgment; the large amount of intermediate output from the security detection Agent does not crowd out other Agents' space.
A Subagent is transient and isolated—it is destroyed after completing a task. Agent Teams, on the other hand, are multiple independent Agent instances collaborating over a long period and messaging each other, more like a real team.
Example
class Orchestrator:
"""Orchestrator: responsible for task decomposition, distribution, and result aggregation"""
def __init__(self):
self.subagents = {
"code_review": Subagent(
name="Code review",
tools=["read_file", "static_analysis"],
system_prompt="You are a code review expert..."
),
"security": Subagent(
name="Security detection",
tools=["scan_vulnerability", "check_deps"],
system_prompt="You are a security detection expert..."
),
"performance": Subagent(
name="Performance analysis",
tools=["profile_code", "analyze_complexity"],
system_prompt="You are a performance analysis expert..."
)
}
def handle_task(self, task: str) -> dict:
# Step 1: Analyze the task and decide which Subagents are needed
needed = self.plan(task)
# Step 2: Distribute in parallel (each Subagent works simultaneously with independent context)
results = {}
for agent_name in needed:
sub_task = self.decompose(task, agent_name)
results[agent_name] = self.subagents[agent_name].run(sub_task)
# Step 3: Aggregate the results of each Subagent and produce a comprehensive output
return self.synthesize(task, results)
# Usage Example: One Run, Three-Dimensional Parallel Analysis
orch = Orchestrator()
report = orch.handle_task(Review PR #42 of the example project)
Advantages
- Naturally supports parallelism, fast
- Sub-agents are independent, contexts do not interfere with each other
- Can specialize the role of each sub-agent
Disadvantages
- Complex coordination logic, difficult to debug
- Higher token cost for parallel multi-agent execution
- The Orchestrator itself may become a bottleneck
The main cost of multi-agent collaboration is orchestration overhead. If subtasks are very simple (each requiring only 1-2 steps), orchestration overhead may exceed the cost of the actual work, in which case a single agent is more appropriate.
Architecture 4: Reflection and Self-Correction
High-quality output Easy to integrate
Add quality assessment to the agent's output stage; if unsatisfactory, regenerate or correct, forming an internal iteration loop.
The reflection architecture adds a "quality inspection step" for the agent: after each output generation, aCriticevaluates the quality, and if it does not meet the standard, requests corrections until the output satisfies the criteria.
This is like a developer running tests on their own code after writing it—self-checking before delivery.
Two implementation approaches
| Approach | Mechanism | Advantages | Disadvantages |
|---|---|---|---|
| Self-reflection | The same model executes first, then evaluates its own output | Simple implementation, no additional model cost | The model may be "blind" to its own errors |
| Critic model | Use an independent critic model to evaluate the executor model's output | More objective, can detect blind spots of the executor model | Adds model call cost and latency |
"Write unit tests → Run tests → Observe failures → Fix code → Run again" — this is a classic application of the reflection architecture. The test results themselves serve as the Critic's feedback signal.
Example
class ReflectiveAgent:
"""Agent with self-reflection capability"""
def __init__(self, model, tools, max_reflections=3):
self.model = model
self.tools = tools
self.max_reflections = max_reflections # Maximum number of corrections to prevent infinite loops
def run(self, task: str) -> str:
# Step 1: Execute normally, produce initial output
output = self.model.generate(task)
for i in range(self.max_reflections):
# Step 2: Reflection — evaluate output quality
critique = self.model.generate(f"""
Please strictly evaluate the following output:
Original task: {task}
Current output: {output}
Check: factual errors? logic gaps? missing information? formatting issues?
If the output is flawless, reply "PASS".
""")
if "PASS" in critique:
break # Output passed review
# Step 3: Correction — improve based on critique
output = self.model.generate(f"""
Original task: {task}
Previous output: {output}
Issue feedback: {critique}
Please correct the output based on the feedback.
""")
return output
# Usage example
agent = ReflectiveAgent(model=llm, tools={})
code = agent.run("Write a Python function that implements AES encryption for the EXAMPLE string")
# After generating the code, the Agent self-checks the encryption implementation and key handling,
# automatically fixes vulnerabilities after detection, ensuring secure and reliable output
Advantages
- Significantly improves output quality
- Allows setting clear quality standards
- Suitable for tasks with objective evaluation criteria
Disadvantages
- Multiple iterations increase latency and cost
- Need to set a maximum iteration count to prevent infinite loops
- Limited effectiveness when evaluation criteria are difficult to formalize
Each iteration of the reflection loop is an additional LLM call, which significantly increases latency. In addition, an upper limit on the number of reflections needs to be set, otherwise the model may fall into an "never satisfied" infinite loop.
Architecture 5: RAG + Agent (Retrieval-Augmented Agent)
Knowledge-intensive tasks Large knowledge bases
Add vector retrieval capability to the Agent's toolset, allowing the Agent to dynamically query external knowledge bases during reasoning, overcoming the limitations of the context window.
RAG (Retrieval-Augmented Generation) is originally a technique that allows LLMs to query external knowledge bases. When combined with an Agent, it becomes more powerful: the Agent canproactively decide when to retrieve and what to retrieve, rather than passively retrieving once each time.
Key differences from ordinary RAG
Ordinary RAG is "passive and one-shot": when a user asks a question, it retrieves once and stuffs the results into the Prompt.
RAG + Agent is differentThe Agent autonomously determines which stage of reasoning needs additional knowledge and what to retrieve, and can query the knowledge base multiple times until it obtains enough information to complete the task.
When the codebase Q&A Agent is asked "Why is the singleton pattern used here?", it proactively retrieves project documentation, design decision records, and relevant code files, rather than guessing only from the LLM's training memory.
Example
class RAGAgent:
"""Agent with dynamic retrieval capability"""
def __init__(self, model, vector_db, max_retrievals=5):
self.model = model
self.vector_db = vector_db # Vector Database
self.max_retrievals = max_retrievals
def should_retrieve(self, context: str, question: str) -> bool:
"""Agent decides whether it needs to retrieve more information"""
decision = self.model.generate(f"""
Current known information: {context}
Current question: {question}
Is the existing information sufficient to answer the question? Answer YES or NO.
""")
return "NO" in decision
def run(self, task: str) -> str:
context = ""
retrieval_count = 0
while retrieval_count < self.max_retrievals:
# Agent autonomously determines whether retrieval is needed
if not self.should_retrieve(context, task):
break
# Agent autonomously decides what to retrieve
search_query = self.model.generate(f"""
Task: {task}
Existing information: {context}
To complete the task, what information should be retrieved next?
""")
# Perform retrieval, append results to context
docs = self.vector_db.search(search_query)
context += "\n".join(docs)
retrieval_count += 1
# Synthesize all information to generate the final answer
return self.model.generate(f"Task: {task}\nReference materials: {context}")
# Usage Example
agent = RAGAgent(model=llm, vector_db=example_docs_db)
answer = agent.run("How do you configure a database connection pool in the EXAMPLE framework?")
# Agent first retrieves "connection pool configuration" and finds mentions of "maximum connections"
# If it doesn't understand, it retrieves "maximum connections best practices" again
# Finally, it synthesizes multi-round retrieval results to provide a complete answer
Advantages
- Break through context window limitations
- Output is well-documented, reducing hallucinations
- Knowledge base can be updated independently
Disadvantages
- Retrieval quality affects overall results
- Maintenance cost of vector databases
- Retrieval latency increases response time
Architecture 6: Workflow Orchestration (Workflow / DAG)
Preferred for production High reliability
Solidify Agent behavior into a directed acyclic graph (DAG), where each node is an LLM call or tool call, edges represent data dependencies, and execution is driven by the framework.
This is an Agent architecture closest to traditional software engineering. The biggest difference from the previous architectures is:The Agent's autonomous decision-making space is limited to within a single node, and the flow between nodes is predefined and unchangeable.
Pure Agent vs DAG Workflow
| Feature | Pure Agent | DAG Workflow |
|---|---|---|
| Flow control | Model autonomously decides the next step | Determined by predefined DAG graph |
| Predictability | Low, execution path may differ each time | High, execution path is fixed |
| Debuggability | Difficult, relies on log tracing | Easy, each node's input and output are clear |
| Fault tolerance | Relies on model self-recovery | Framework provides retry and resume from checkpoint |
| Flexibility | High, can handle unexpected situations | Low, can only follow predefined paths |
The "Acyclic" property of DAG means the workflow isDeterministic— no infinite loops, execution path can be fully predicted, and failed nodes can be retried individually.
DAG is the extreme of "low autonomy, high predictability." You need to pre-design the entire process. This is not a disadvantage, but a deliberate design trade-off — in production environments, determinism is sometimes more important than flexibility.
Example
from langgraph import StateGraph
# Define workflow state — data object passed between nodes
class PipelineState:
raw_data: str = "" # Raw input data
cleaned_data: str = "" # Cleaned data
analysis_result: dict = {} # Analysis result
final_report: str = "" # Final report
# Define DAG nodes — each node is an independent processing unit
def extract_data(state: PipelineState) -> PipelineState:
"""Node 1: Extract raw data from example database"""
state.raw_data = query_database("SELECT * FROM logs")
return state
def clean_data(state: PipelineState) -> PipelineState:
"""Node 2: Clean data (deduplicate, standardize format)"""
state.cleaned_data = preprocess(state.raw_data)
return state
def analyze_data(state: PipelineState) -> PipelineState:
"""Node 3: Statistical analysis"""
state.analysis_result = statistical_analysis(state.cleaned_data)
return state
def generate_report(state: PipelineState) -> PipelineState:
"""Node 4: Generate report using LLM"""
state.final_report = llm.generate(
f"Generate report based on the following analysis results: {state.analysis_result}"
)
return state
# Build DAG: define nodes and edges (data flow)
workflow = StateGraph(PipelineState)
workflow.add_node("extract", extract_data)
workflow.add_node("clean", clean_data)
workflow.add_node("analyze", analyze_data)
workflow.add_node("report", generate_report)
# Define edges: extract → clean → analyze → report
workflow.add_edge("extract", "clean")
workflow.add_edge("clean", "analyze")
workflow.add_edge("analyze", "report")
workflow.set_entry_point("extract")
workflow.set_finish_point("report")
# Compile and run
app = workflow.compile()
result = app.invoke(PipelineState())
print(result.final_report)
Advantages
- Predictable, auditable, retryable
- Supports parallel acceleration
- High engineering maturity, ops-friendly
Disadvantages
- Flow requires pre-design, low flexibility
- Difficult to handle unexpected situations
- Requires learning orchestration frameworks
Horizontal Comparison and How to Choose
Compare six architectures across multiple dimensions to help you quickly identify the right option.
| Architecture | Autonomy | Predictability | Parallel capability | Suitable task complexity | Typical implementation |
|---|---|---|---|---|---|
| Single Agent loop | High | Low | None | Medium-low | Claude Code default mode |
| Planning + Execution | Medium | Medium | Partial | Medium-high | Claude Code Plan Mode |
| Multi-Agent collaboration | High | Low | strong | High | AutoGen, CrewAI |
| Reflection / self-correction | Medium | Medium | None | Medium | Reflexion, Self-Refine |
| RAG + Agent | High | Medium | None | Medium-high | LangChain RAG Agent |
| Workflow orchestration | Low | High | strong | High (fixed flow) | LangGraph, Prefect |
Common combination patterns
Real production systems often combine multiple architectures. Here are several mature combination patterns:
Workflow orchestration + Multi-Agent: Use DAG to define the main flow, with each node being an independent Agent. For example, in a CI/CD pipeline, the code review node is a Code Review Agent, and the security scan node is a Security Agent.
Multi-Agent + RAG: Multiple Subagents share the same vector knowledge base, each retrieving on demand based on its subtask. For example, in a customer service system, both the order query Agent and the refund processing Agent query the same knowledge base but retrieve different content.
Planning execution + Reflection: Plan first, then execute, but add a reflection step after each step to ensure quality. Suitable for tasks with extremely high quality requirements.
If you are unsure where to start, start with a single-Agent loop. It is the easiest to implement and debug. When you find the context window is no longer sufficient, consider multi-Agent; when you need quality assurance, add a reflection layer; when the process stabilizes, refactor it into a DAG to improve reliability.
Common misconceptions
Misconception 1: The more complex the architecture, the better
Multi-Agent collaboration looks powerful, but if your task can be completed in 5 steps with a single-Agent loop, introducing orchestration overhead actually reduces efficiency. Principle: use the simplest architecture that meets the requirements.
Misconception 2: Reflection will definitely improve quality
The effectiveness of self-reflection depends on the model's self-evaluation ability. If quality requirements are extremely strict, consider using a Critic model or introducing external validation (such as automated code testing).
Misconception 3: DAG workflows do not need Agents
DAG defines the process skeleton, but each node can still be an Agent call internally. Workflow orchestration and Agent capabilities are not mutually exclusive but complementary — DAG provides reliability, and Agents provide flexibility.
Misconception 4: A larger context window means RAG is not needed
Even if the model supports a 1M token context window, stuffing all documents into it is still not optimal. The value of RAG is not just "being able to fit it all," but alsoprecise retrieval— reducing noise, lowering inference costs, and improving answer accuracy.
Summary
Six Agent architectures cover the full spectrum from highly autonomous to highly controllable.
When choosing an architecture, there are only two core considerations:How much flexibility you need to handle unexpected situations, andHow much certainty you need to ensure reliable results。
Start simple and add complexity only when truly needed — this is the first principle of Agent architecture selection.
Other extensions