LangChain Building RAG Agent
In the previous two articles, we prepared the vector store and retriever.
This article integrates them into an Agent to build a complete RAG Agent — an intelligent assistant that can answer questions based on a private knowledge base.
Creating a Retriever Tool
Wrap the retriever into a tool, and the Agent can automatically search the knowledge base when needed.
If you don't have an OpenAI key, you can use Alibaba Bailian's Embedding service; see configuration.https://www.example.com/langchain/langchain-alibailian.html:
# OpenAI 的文本嵌入模型
# 将文本转换为向量(一组浮点数)
# 使用阿里云百炼(DashScope)的通义千问 Embedding 服务
embeddings = OpenAIEmbeddings(
model="text-embedding-v4",
api_key=os.getenv("DASHSCOPE_API_KEY"),
base_url="https://dashscope.aliyuncs.com/compatible-mode/v1",
check_embedding_ctx_length=False,
chunk_size=10,
)
Example (OpenAI)
from dotenv import load_dotenv
load_dotenv()
from langchain.tools import tool
from langchain.agents import create_agent
from langchain.chat_models import init_chat_model
from langchain.messages import HumanMessage
from langchain_openai import OpenAIEmbeddings
from langchain_chroma import Chroma
from langchain_text_splitters import RecursiveCharacterTextSplitter
# ----- Step 1: Prepare the knowledge base -----
# Simulate the knowledge documents of EXAMPLE
knowledge_docs = [
"EXAMPLE was founded in 2013 and is a completely free programming learning platform.",
"The platform has launched 300+ tutorials, covering front-end, back-end, databases, mobile development, and other fields.",
"The Python3 Basic Tutorial is the most popular course on the platform, with over 5 million cumulative learners.",
"The Python3 Basic Tutorial has 30 chapters, including environment setup, basic syntax, functions, classes, exception handling, and more.",
"The HTML Basic Tutorial has 25 chapters, covering everything from basic HTML structure to forms and multimedia elements.",
"EXAMPLE supports running code online; learners can write and run code without installing any software.",
"The platform is mobile-friendly, allowing users to learn programming anytime, anywhere on their phones.",
"EXAMPLE's membership service provides value-added services such as video courses, hands-on projects, and one-on-one Q&A.",
]
# Split
text_splitter = RecursiveCharacterTextSplitter(
chunk_size=200, chunk_overlap=30
)
chunks = text_splitter.create_documents(knowledge_docs)
# Vectorized storage -- this can be changed to Alibaba Bailian if you don't have an OpenAI key
embeddings = OpenAIEmbeddings(model="text-embedding-3-small")
vector_store = Chroma.from_documents(
documents=chunks,
embedding=embeddings,
)
retriever = vector_store.as_retriever(search_kwargs={"k": 3})
# ----- Step 2: Create a retrieval tool -----
@tool
def search_knowledge_base(query: str) -> str:
"""Search for relevant information in the EXAMPLE knowledge base.
When users ask about specific information about EXAMPLE (such as number of courses, platform history, features, etc.),
you must use this tool to query the knowledge base for accurate information.
Args:
query: search keyword or question
"""
docs = retriever.invoke(query)
if not docs:
return "No relevant information found in the knowledge base."
results = []
for i, doc in enumerate(docs, 1):
results.append(f"[{i}] {doc.page_content}")
return "\n\n".join(results)
# ----- Step 3: Create the RAG Agent -----
model = init_chat_model("deepseek:deepseek-v4-flash", temperature=0)
agent = create_agent(
model=model,
tools=[search_knowledge_base],
system_prompt="""You are the intelligent customer service assistant of EXAMPLE.
## Rules
1. When users ask about specific information about EXAMPLE, you must use the search_knowledge_base tool to query.
2. Answer based on the retrieved information; do not make up content that is not in the knowledge base.
3. If the knowledge base has no relevant information, honestly tell the user.
4. Be friendly, concise, and accurate in your answers.""",
)
# ----- Step 4: Test -----
questions = [
"When was EXAMPLE founded?",
"How many chapters are in the Python3 Basic Tutorial?",
"How many tutorials does EXAMPLE have in total?",
]
for q in questions:
result = agent.invoke({"messages": [HumanMessage(content=q)]})
print(f"Q: {q}")
print(f"A: {result['messages'][-1].content}")
print("-" * 60)
Run results:
Q: Example是什么时候创立的? A: Example创立于 2013 年,是一个完全免费的编程学习平台。 ------------------------------------------------------------ Q: Python3 基础教程有多少章? A: Python3 基础教程共 30 章,包含环境搭建、基本语法、函数、类、异常处理等内容。 ------------------------------------------------------------ Q: Example一共有多少套教程? A: Example已上线 300+ 套教程,涵盖前端、后端、数据库、移动开发等多个领域。 ------------------------------------------------------------
Execution Flow of RAG Agent
For the third question above, "How many tutorials does EXAMPLE have in total?", the Agent's execution flow is:
- The user asks a question
- The model determines it needs to query the knowledge base → calls search_knowledge_base("EXAMPLE number of tutorials")
- The retriever searches the vector database for the semantically most similar document chunks.
- Returns the retrieval results to the model.
- The model generates an accurate answer based on the retrieval results.
Adding Citation Sources
Professional RAG systems usually include citation sources so users know where the information comes from:
Example
@tool
def search_with_sources(query: str) -> str:
"""Search the EXAMPLE knowledge base and return results with source annotations.
Args:
query: search keyword
"""
docs = retriever.invoke(query)
if not docs:
return "No relevant information found."
results = []
for i, doc in enumerate(docs, 1):
source = doc.metadata.get("source", "EXAMPLE Knowledge Base")
results.append(f"[Source {i}: {source}]"\n{doc.page_content}")
return "\n\n".join(results)
# To retain source information in the document, you can add metadata when creating it.
doc_with_meta = Document(
page_content="The Python3 Basic Tutorial has 30 chapters...",
metadata={"source": "Python3 Basic Tutorial - Course Introduction", "url": "https://www.example.com/python3/"}
)
Persistence of Vector Store
In real projects, you don't rebuild the vector index every time. Chroma supports persisting to local storage:
Example
embeddings = OpenAIEmbeddings(model="text-embedding-3-small")
vector_store = Chroma.from_documents(
documents=chunks,
embedding=embeddings,
persist_directory="./example_vector_db", # Persistence directory
)
# Load directly on subsequent runs
loaded_store = Chroma(
persist_directory="./example_vector_db",
embedding_function=embeddings,
)
retriever = loaded_store.as_retriever()
# No need to recompute vectors!
Other ExtensionsPersisting the vector store can greatly improve startup speed. With large document volumes (tens of thousands of documents), recomputing the Embeddings for all vectors can take dozens of minutes. Once persisted, you only need to load it.