LangChain Middleware

LangChain Middleware is LangChain's most powerful feature. It allows you to insert custom logic at various stages of Agent execution to implement retries, degradation, caching, content filtering, logging, and other functions — without modifying the Agent's own code.


What is Middleware

Middleware consists ofhooks (Hook)in the Agent execution flow. Each hook lets you execute custom code at specific points in time:

Example

# Intuitive understanding of Middleware:
# Assume the Agent execution flow is like this:

# 1. User input → 2. Model thinking → 3. May call tools → 4. Model thinks again → 5. Output result

# Middleware lets you insert custom logic between these 5 stages:
# 1. User input
# ↓ [before_agent hook: logging, permission checks]
# 2. Model thinking
# ↓ [before_model hook: message preprocessing]
# ↓ [wrap_model_call hook: retry, degradation, caching]
# ↓ [after_model hook: content moderation]
# 3. Tool execution
# ↓ [wrap_tool_call hook: tool call retry]
# 4. Back to model thinking (loop until complete)
# ↓ [after_agent hook: result formatting, statistical analysis]
# 5. Output result

Six Hook Points

LangChain's Middleware provides 6 hooks, divided into two categories by execution timing:

HookExecution FrequencyExecution PositionMain Purpose
before_agentOnceBefore Agent startsInitialization, permission checks, input preprocessing
before_modelEvery loop iterationBefore model callMessage preprocessing, dynamic context injection
wrap_model_callEvery loop iterationWraps model callRetry, degradation, caching, request rewriting
after_modelEvery loop iterationAfter model callContent moderation, response filtering, logging
wrap_tool_callEvery tool callWraps tool executionTool retry, result caching, parameter rewriting
after_agentOnceAfter Agent endsOutput formatting, statistics, resource cleanup

Two Ways to Use

Middleware can be used in two ways: class inheritance or decorators.

Method 1: Decorator (Recommended)

Example

from langchain.agents.middleware import before_model, after_model

# Decorator approach: simple and intuitive
@before_model
def log_before(state, runtime):
    """Log before each model call"""
    msg_count = len(state.get("messages", []))
    print(f"[before_model] Current message count: {msg_count}")
    return None


@after_model
def log_after(state, runtime):
    """Log after each model call"""
    last_msg = state["messages"][-1] if state.get("messages") else None
    if last_msg and hasattr(last_msg, 'tool_calls') and last_msg.tool_calls:
        print(f"[after_model] Model requested a tool call")
    return None

Method 2: Class Inheritance (Suitable for Complex Logic)

Example

from langchain.agents.middleware import AgentMiddleware


class LoggingMiddleware(AgentMiddleware):
    """Custom logging middleware"""

    @property
    def name(self) -> str:
        # Custom middleware name (defaults to class name)
        return "logging"

    def before_agent(self, state, runtime):
        """Logic before Agent starts"""
        print("[Logging] Agent started executing")
        return None

    def before_model(self, state, runtime):
        """Logic before model call"""
        msg_count = len(state.get("messages", []))
        print(f"[Logging] Preparing to call model, currently {msg_count} messages")
        return None

    def after_model(self, state, runtime):
        """Logic after model call"""
        print("[Logging] Model call completed")
        return None

    def after_agent(self, state, runtime):
        """Logic after Agent ends"""
        print("[Logging] Agent execution finished")
        return None

Complete Lifecycle Example

Example

from dotenv import load_dotenv
load_dotenv()

from langchain.agents import create_agent
from langchain.agents.middleware import (
    before_agent, after_agent,
    before_model, after_model,
)
from langchain.chat_models import init_chat_model
from langchain.messages import HumanMessage
from langchain.tools import tool


@before_agent
def start_log(state, runtime):
    """Before Agent starts"""
    print(">>> [before_agent] Agent started <<<")
    runtime.stream_writer({"type": "lifecycle", "phase": "start"})
    return None


@before_model
def pre_model(state, runtime):
    """Before each model call"""
    msg_count = len(state.get("messages", []))
    print(f" -> [before_model] Message {msg_count}")
    return None


@after_model
def post_model(state, runtime):
    """After each model call"""
    last = state["messages"][-1] if state.get("messages") else None
    if hasattr(last, 'tool_calls') and last.tool_calls:
        tools = [tc['name'] for tc in last.tool_calls]
        print(f" <- [after_model] Requested tool: {tools}")
    else:
        content = str(last.content)[:50] if last and hasattr(last, 'content') else ""
        print(f" <- [after_model] Direct reply: {content}...")
    return None


@after_agent
def end_log(state, runtime):
    """After Agent ends"""
    total_msgs = len(state.get("messages", []))
    print(f"<<< [after_agent] Agent ended, {total_msgs} messages in total <<<")
    return None


@tool
def get_weather(city: str) -> str:
    """Query weather"""
    return f"{city}: Sunny, 25°C"


model = init_chat_model("deepseek:deepseek-v4-flash", temperature=0)
agent = create_agent(
    model=model,
    tools=[get_weather],
    middleware=[start_log, pre_model, post_model, end_log],
    system_prompt="You are an assistant.",
)

print("\n========== First question (requires tool) ==========")
result = agent.invoke({
    "messages": [HumanMessage(content="What's the weather in Hangzhou?")]
})

print(f"\nFinal reply: {result['messages'][-1].content}")

print("\n========== Second question (no tool needed) ==========")
result = agent.invoke({
    "messages": [HumanMessage(content="Hello")]
})
print(f"\nFinal reply: {result['messages'][-1].content}")

Output:

========== 第一个问题(需要工具) ==========
>>> [before_agent] Agent 开始 <<<
  -> [before_model] 第 2 条消息
  <- [after_model] 请求工具: ['get_weather']
  -> [before_model] 第 3 条消息
  <- [after_model] 直接回复: 杭州今天晴,气温25°C。...
<<< [after_agent] Agent 结束,共 4 条消息 <<<

Final reply: 杭州今天晴,气温25°C。

========== 第二个问题(无需工具) ==========
>>> [before_agent] Agent 开始 <<<
  -> [before_model] 第 1 条消息
  <- [after_model] 直接回复: 你好!有什么可以帮你的?...
<<< [after_agent] Agent 结束,共 2 条消息 <<<

Final reply: Hello! How can I help you?

From the output, you can see:

  • before_agent and after_agent: executed only once per question
  • before_model and after_model: executed on every model call (the first question called the model twice, so each ran twice)

Middleware Return Value

Middleware's return value determines whether to modify the Agent state or control the flow:

Return ValueEffectExample
NoneDoes not modify any state; continues the normal flowLogging only
dictUpdates Agent state (merged into the current state)Return {"custom_field": "value"}
Dict containing jump_toJump to the specified nodeReturn {"jump_to": "end"}

The returned dict is merged through the Agent state's reducer. For the messages field, the add_messages reducer is used, so returned messages are appended rather than overwritten.

Other Extensions