LangChain Streaming Output
Streaming output makes the AI's replies display character by character like typing, greatly improving the user experience. LangChain's Agent has built-in comprehensive streaming output support.
Why Streaming Output is Needed
If invoke() is used, users have to wait for the Agent to complete all steps (multiple model calls + tool execution) before seeing the result. For complex tasks, this can take more than ten seconds or even longer.
Streaming output solves this problem: each generated Token is returned immediately, so users can see progress in real time.
| Method | User Experience | Applicable Scenarios |
|---|---|---|
| invoke() | Wait → See the complete result at once | Scripts, API, batch processing |
| stream() | See every Token in real time | Chat interfaces, real-time display |
stream_mode="messages" — Token-by-Token Streaming
This is the finest-grained streaming mode, where each chunk corresponds to one Token:
Example
load_dotenv()
from langchain.agents import create_agent
from langchain.chat_models import init_chat_model
from langchain.messages import HumanMessage
model = init_chat_model("deepseek:deepseek-v4-flash")
agent = create_agent(
model=model,
system_prompt="You are the assistant of EXAMPLE Tutorial.",
)
# stream_mode="messages" returns tokens one by one
print("Real-time streaming output:")
for msg_chunk, metadata in agent.stream(
{"messages": [HumanMessage(content="Introduce Python in one sentence")]},
stream_mode="messages",
):
# msg_chunk is an AIMessageChunk
# each chunk contains only a small piece of content
if msg_chunk.content:
print(msg_chunk.content, end="", flush=True)
print() # newline at the end
Run result:
Live streaming output: Example is a free online learning platform for programming beginners, offering rich technical tutorials and hands-on examples.
Understanding metadata
metadata contains source information about this chunk:
Example
for msg_chunk, metadata in agent.stream(
{"messages": [HumanMessage(content="Hello, introduce yourself")]},
stream_mode="messages",
):
if msg_chunk.content and len(msg_chunk.content) > 5:
print(f"Content: {msg_chunk.content}")
print(f"Source node: {metadata.get('langgraph_node')}")
print(f"Message type: {type(msg_chunk).__name__}")
break # view only the first meaningful chunk
Run result:
Inspect metadata: Content: Hello! I am an AI assistant Source node: model Message type: AIMessageChunk
stream_mode="updates" — Viewing the Agent Execution Process Step by Step
This mode is very useful when building interfaces that need to display the "thinking process":
Example
@tool
def search_course(keyword: str) -> str:
"""Search for courses on EXAMPLE"""
courses = {
"python": "Python3 Basic Tutorial (30 chapters, 20 hours)",
"html": "HTML Basic Tutorial (25 chapters, 15 hours)",
}
return courses.get(keyword.lower(), "No related courses found")
agent = create_agent(
model=init_chat_model("deepseek:deepseek-v4-flash", temperature=0),
tools=[search_course],
system_prompt="You are a course consultant at EXAMPLE.",
)
# use updates mode to view each step
print("=== Agent execution process ==="\n")
for chunk in agent.stream(
{"messages": [HumanMessage(content=Help me check Python courses)]},
stream_mode="updates",
):
for node_name, update in chunk.items():
print(f"[{node_name}]", end=" ")
if "messages" in update:
for msg in update["messages"]:
if msg.type == "ai":
if hasattr(msg, 'tool_calls') and msg.tool_calls:
calls = [tc['name'] for tc in msg.tool_calls]
print(fRequest calls: {calls})
elif msg.content:
print(fResponse: {msg.content[:80]})
elif msg.type == "tool":
print(fTool returned [{msg.name}]: {msg.content})
Run result:
=== Agent 执行过程 === [model] 请求调用: ['search_course'] [tools] 工具返回 [search_course]: Python3 基础教程(30章,20小时) [model] Reply: Example has the Python3 basics tutorial, 30 chapters, about 20 hours of study, perfect for Python beginners.
stream_mode="custom" — Sending Custom Events
Through the Middleware's runtime.stream_writer(), you can send custom events to the stream:
Example
@before_model
def notify_before(state, runtime):
Send custom events before model invocation
runtime.stream_writer({
"type": "status",
"message": Thinking...,
})
return None
@after_model
def notify_after(state, runtime):
Send custom events after model invocation
last_msg = state["messages"][-1] if state.get("messages") else None
has_tools = hasattr(last_msg, 'tool_calls') and last_msg.tool_calls
if has_tools:
tool_names = [tc['name'] for tc in last_msg.tool_calls]
runtime.stream_writer({
"type": "status",
"message": fCalling tools: {', '.join(tool_names)}...,
})
else:
runtime.stream_writer({
"type": "status",
"message": Answer completed,
})
return None
agent = create_agent(
model=init_chat_model("deepseek:deepseek-v4-flash", temperature=0),
tools=[search_course],
middleware=[notify_before, notify_after],
system_prompt="You are a course consultant at EXAMPLE.",
)
# Use stream_mode=["updates", "custom"] to receive both events at the same time
print(=== Mixed streaming output ===\n")
for mode, chunk in agent.stream(
{"messages": [HumanMessage(content=Check Python courses)]},
stream_mode=["updates", "custom"],
):
if mode == "custom":
print(f[Custom event] Status: {chunk['message']})
elif mode == "updates":
for node_name, update in chunk.items():
if "messages" in update:
for msg in update["messages"]:
if msg.type == "ai" and msg.content:
print(f[Response] {msg.content})
Run result:
=== 混合流式输出 === [custom event] Status: thinking... [custom event] Status: calling tool: search_course... [reply] Example has the Python3 basics tutorial, 30 chapters, about 20 hours of study. [custom event] Status: thinking... [custom event] Status: answer completed
stream_mode modes can be combined, such as stream_mode=["updates", "custom", "messages"]. However, too many modes will increase the number of events in the stream, so it is recommended to select them as needed.
Asynchronous Streaming Output
In web services, using asynchronous streaming can avoid blocking the event loop:
Example
async def stream_agent():
Asynchronously run Agent in streaming mode
agent = create_agent(
model=init_chat_model("deepseek:deepseek-v4-flash"),
system_prompt="You are the assistant of EXAMPLE Tutorial.",
)
full_response = ""
async for msg_chunk, metadata in agent.astream(
{"messages": [HumanMessage(content=Introduce Python in one sentence)]},
stream_mode="messages",
):
if msg_chunk.content:
full_response += msg_chunk.content
print(msg_chunk.content, end="", flush=True)
print(f"\n\nFull response length: {len(full_response)} characters)
# Run the async function
asyncio.run(stream_agent())
FastAPI Integration Example
Example
from fastapi.responses import StreamingResponse
from langchain.agents import create_agent
from langchain.chat_models import init_chat_model
from langchain.messages import HumanMessage
app = FastAPI()
agent = create_agent(
model=init_chat_model("deepseek:deepseek-v4-flash"),
system_prompt="You are the assistant of EXAMPLE Tutorial.",
)
@app.get("/chat")
async def chat(message: str):
Chat interface, returns SSE streaming response
async def generate():
async for msg_chunk, metadata in agent.astream(
{"messages": [HumanMessage(content=message)]},
stream_mode="messages",
):
if msg_chunk.content:
# SSE format: data: xxx
yield f"data: {msg_chunk.content}\n\n"
yield "data: [DONE]\n\n"
return StreamingResponse(
generate(),
media_type="text/event-stream",
)
# Start: uvicorn main:app --reload
In a production environment, it is recommended to create the Agent instance as a global singleton to avoid recreating it on every request. The cost of creating an Agent is small (mainly compiling the graph), but reusing the instance is more efficient.
stream_mode Quick Reference
| Mode | Granularity | Iteration Object | Typical Use |
|---|---|---|---|
| messages | Token level | (AIMessageChunk, metadata) | Typing effect, real-time chat |
| updates | Node level | {node_name: state_update} | Display thinking process |
| values | Node level (full) | Complete state | State snapshot, debugging |
| custom | Custom | Arbitrary dict | Progress notification, status push |
| debug | Detailed | Debug information | Troubleshooting during development |