LangChain Streaming Output

Streaming output makes the AI's replies display character by character like typing, greatly improving the user experience. LangChain's Agent has built-in comprehensive streaming output support.


Why Streaming Output is Needed

If invoke() is used, users have to wait for the Agent to complete all steps (multiple model calls + tool execution) before seeing the result. For complex tasks, this can take more than ten seconds or even longer.

Streaming output solves this problem: each generated Token is returned immediately, so users can see progress in real time.

MethodUser ExperienceApplicable Scenarios
invoke()Wait → See the complete result at onceScripts, API, batch processing
stream()See every Token in real timeChat interfaces, real-time display

stream_mode="messages" — Token-by-Token Streaming

This is the finest-grained streaming mode, where each chunk corresponds to one Token:

Example

from dotenv import load_dotenv
load_dotenv()

from langchain.agents import create_agent
from langchain.chat_models import init_chat_model
from langchain.messages import HumanMessage

model = init_chat_model("deepseek:deepseek-v4-flash")
agent = create_agent(
    model=model,
    system_prompt="You are the assistant of EXAMPLE Tutorial.",
)

# stream_mode="messages" returns tokens one by one
print("Real-time streaming output:")
for msg_chunk, metadata in agent.stream(
    {"messages": [HumanMessage(content="Introduce Python in one sentence")]},
    stream_mode="messages",
):
    # msg_chunk is an AIMessageChunk
    # each chunk contains only a small piece of content
    if msg_chunk.content:
        print(msg_chunk.content, end="", flush=True)
print()  # newline at the end

Run result:

Live streaming output:
Example is a free online learning platform for programming beginners, offering rich technical tutorials and hands-on examples.

Understanding metadata

metadata contains source information about this chunk:

Example

print("View metadata information:"\n")
for msg_chunk, metadata in agent.stream(
    {"messages": [HumanMessage(content="Hello, introduce yourself")]},
    stream_mode="messages",
):
    if msg_chunk.content and len(msg_chunk.content) > 5:
        print(f"Content: {msg_chunk.content}")
        print(f"Source node: {metadata.get('langgraph_node')}")
        print(f"Message type: {type(msg_chunk).__name__}")
        break  # view only the first meaningful chunk

Run result:

Inspect metadata:

Content: Hello! I am an AI assistant
Source node: model
Message type: AIMessageChunk

stream_mode="updates" — Viewing the Agent Execution Process Step by Step

This mode is very useful when building interfaces that need to display the "thinking process":

Example

from langchain.tools import tool


@tool
def search_course(keyword: str) -> str:
    """Search for courses on EXAMPLE"""
    courses = {
        "python": "Python3 Basic Tutorial (30 chapters, 20 hours)",
        "html": "HTML Basic Tutorial (25 chapters, 15 hours)",
    }
    return courses.get(keyword.lower(), "No related courses found")


agent = create_agent(
    model=init_chat_model("deepseek:deepseek-v4-flash", temperature=0),
    tools=[search_course],
    system_prompt="You are a course consultant at EXAMPLE.",
)

# use updates mode to view each step
print("=== Agent execution process ==="\n")
for chunk in agent.stream(
    {"messages": [HumanMessage(content=Help me check Python courses)]},
    stream_mode="updates",
):
    for node_name, update in chunk.items():
        print(f"[{node_name}]", end=" ")
        if "messages" in update:
            for msg in update["messages"]:
                if msg.type == "ai":
                    if hasattr(msg, 'tool_calls') and msg.tool_calls:
                        calls = [tc['name'] for tc in msg.tool_calls]
                        print(fRequest calls: {calls})
                    elif msg.content:
                        print(fResponse: {msg.content[:80]})
                elif msg.type == "tool":
                    print(fTool returned [{msg.name}]: {msg.content})

Run result:

=== Agent 执行过程 ===

[model] 请求调用: ['search_course']
[tools] 工具返回 [search_course]: Python3 基础教程(30章,20小时)
[model] Reply: Example has the Python3 basics tutorial, 30 chapters, about 20 hours of study, perfect for Python beginners.

stream_mode="custom" — Sending Custom Events

Through the Middleware's runtime.stream_writer(), you can send custom events to the stream:

Example

from langchain.agents.middleware import before_model, after_model


@before_model
def notify_before(state, runtime):
    Send custom events before model invocation
    runtime.stream_writer({
        "type": "status",
        "message": Thinking...,
    })
    return None


@after_model
def notify_after(state, runtime):
    Send custom events after model invocation
    last_msg = state["messages"][-1] if state.get("messages") else None
    has_tools = hasattr(last_msg, 'tool_calls') and last_msg.tool_calls

    if has_tools:
        tool_names = [tc['name'] for tc in last_msg.tool_calls]
        runtime.stream_writer({
            "type": "status",
            "message": fCalling tools: {', '.join(tool_names)}...,
        })
    else:
        runtime.stream_writer({
            "type": "status",
            "message": Answer completed,
        })
    return None


agent = create_agent(
    model=init_chat_model("deepseek:deepseek-v4-flash", temperature=0),
    tools=[search_course],
    middleware=[notify_before, notify_after],
    system_prompt="You are a course consultant at EXAMPLE.",
)

# Use stream_mode=["updates", "custom"] to receive both events at the same time
print(=== Mixed streaming output ===\n")
for mode, chunk in agent.stream(
    {"messages": [HumanMessage(content=Check Python courses)]},
    stream_mode=["updates", "custom"],
):
    if mode == "custom":
        print(f[Custom event] Status: {chunk['message']})
    elif mode == "updates":
        for node_name, update in chunk.items():
            if "messages" in update:
                for msg in update["messages"]:
                    if msg.type == "ai" and msg.content:
                        print(f[Response] {msg.content})

Run result:

=== 混合流式输出 ===

[custom event] Status: thinking...
[custom event] Status: calling tool: search_course...
[reply] Example has the Python3 basics tutorial, 30 chapters, about 20 hours of study.
[custom event] Status: thinking...
[custom event] Status: answer completed

stream_mode modes can be combined, such as stream_mode=["updates", "custom", "messages"]. However, too many modes will increase the number of events in the stream, so it is recommended to select them as needed.


Asynchronous Streaming Output

In web services, using asynchronous streaming can avoid blocking the event loop:

Example

import asyncio


async def stream_agent():
    Asynchronously run Agent in streaming mode
    agent = create_agent(
        model=init_chat_model("deepseek:deepseek-v4-flash"),
        system_prompt="You are the assistant of EXAMPLE Tutorial.",
    )

    full_response = ""
    async for msg_chunk, metadata in agent.astream(
        {"messages": [HumanMessage(content=Introduce Python in one sentence)]},
        stream_mode="messages",
    ):
        if msg_chunk.content:
            full_response += msg_chunk.content
            print(msg_chunk.content, end="", flush=True)

    print(f"\n\nFull response length: {len(full_response)} characters)


# Run the async function
asyncio.run(stream_agent())

FastAPI Integration Example

Example

from fastapi import FastAPI
from fastapi.responses import StreamingResponse
from langchain.agents import create_agent
from langchain.chat_models import init_chat_model
from langchain.messages import HumanMessage

app = FastAPI()
agent = create_agent(
    model=init_chat_model("deepseek:deepseek-v4-flash"),
    system_prompt="You are the assistant of EXAMPLE Tutorial.",
)


@app.get("/chat")
async def chat(message: str):
    Chat interface, returns SSE streaming response
    async def generate():
        async for msg_chunk, metadata in agent.astream(
            {"messages": [HumanMessage(content=message)]},
            stream_mode="messages",
        ):
            if msg_chunk.content:
                # SSE format: data: xxx
                yield f"data: {msg_chunk.content}\n\n"
        yield "data: [DONE]\n\n"

    return StreamingResponse(
        generate(),
        media_type="text/event-stream",
    )

# Start: uvicorn main:app --reload

In a production environment, it is recommended to create the Agent instance as a global singleton to avoid recreating it on every request. The cost of creating an Agent is small (mainly compiling the graph), but reusing the instance is more efficient.


stream_mode Quick Reference

ModeGranularityIteration ObjectTypical Use
messagesToken level(AIMessageChunk, metadata)Typing effect, real-time chat
updatesNode level{node_name: state_update}Display thinking process
valuesNode level (full)Complete stateState snapshot, debugging
customCustomArbitrary dictProgress notification, status push
debugDetailedDebug informationTroubleshooting during development
Other extensions