LangChain Document Loading and Splitting
In previous articles we manually entered text, but in real projects, documents may come from PDFs, web pages, Markdown files, etc.
This section introduces how to use Document Loader to load various types of documents, and how to use Text Splitter to split documents into small chunks suitable for retrieval.
Document Loader—Loading Documents
LangChain provides dozens of document loaders, covering common file formats:
| Loader | Source | Installation Package |
|---|---|---|
| TextLoader | .txt files | langchain (built-in) |
| PyPDFLoader | PDF files | langchain-community + pypdf |
| WebBaseLoader | Webpage URL | langchain-community + beautifulsoup4 |
| CSVLoader | CSV files | langchain-community |
| UnstructuredMarkdownLoader | Markdown files | langchain-community + unstructured |
Example
from langchain_community.document_loaders import TextLoader
loader = TextLoader("knowledge.txt", encoding="utf-8")
docs = loader.load()
print(f"Loaded {len(docs)} documents")
print(f"Content preview: {docs)
# Load a webpage
# pip install langchain-community beautifulsoup4
from langchain_community.document_loaders import WebBaseLoader
loader = WebBaseLoader("https://www.example.com/python/python-tutorial.html")
docs = loader.load()
print(f"\nWebpage content: {docs)
Text Splitter—Document Splitting
Documents are usually too long and need to be split into small chunks for effective retrieval. The splitting strategy directly affects RAG performance:
Example
# Create a splitter
text_splitter = RecursiveCharacterTextSplitter(
chunk_size=500, # Maximum of 500 characters per chunk
chunk_overlap=50, # 50-character overlap between chunks
separators=["\n\n", "\n", "。", "!", "?", ";", ",", " ", ""],
# Split by paragraph first, then sentence, and finally by character
)
# Example document
long_text = """Example Tutorial (EXAMPLE) is a free programming learning platform.
The platform provides rich programming language tutorials, including but not limited to:
- Python tutorial: from basic syntax to data analysis
- Java tutorial: object-oriented programming to Spring framework
- Frontend tutorial: HTML, CSS, JavaScript, and their frameworks
All tutorials come with detailed code examples and an online running environment.
Learners can quickly master programming skills by learning while practicing."""
# Split the document
chunks = text_splitter.split_text(long_text)
print(f"Original length: {len(long_text)} characters")
print(f"After splitting: {len(chunks)} chunks\n")
for i, chunk in enumerate(chunks):
print(f"--- Chunk {i+1} ({len(chunk)} characters) ---")
print(chunk)
print()
Running result:
Original length: 153 字 Chunks after split: 3 块 --- 块 1 (54 字) --- Example is a free programming learning platform. 平台提供了丰富的编程语言教程,包括但不限于: --- 块 2 (49 字) --- - Python 教程:从基础语法到数据分析 - Java 教程:面向对象编程到 Spring 框架 --- 块 3 (50 字) --- - 前端教程:HTML、CSS、JavaScript 及其框架 所有教程都配有详细的代码示例和在线运行环境。
chunk_overlap is very important. If there is no overlap between chunks, a complete sentence may be split in half, causing key information to be missed during retrieval. An overlap of 50-100 characters is a common setting.
Chunking Parameter Settings Guide
| Scenario | chunk_size | chunk_overlap | Reason |
|---|---|---|---|
| FAQ Q&A | 200~500 | 20~50 | Q&A pairs are short, small chunks are sufficient |
| Technical documentation | 500~1000 | 50~100 | Technical content requires more context |
| Long articles/papers | 1000~2000 | 100~200 | Need to preserve paragraph completeness |
| Code repository | 500~1500 | 0~50 | Functions/classes as natural boundaries |
Complete Flow: Load → Split → Vectorize
Example
from langchain_openai import OpenAIEmbeddings
from langchain_chroma import Chroma
# from langchain_community.document_loaders import TextLoader
# Step 1: Load
# loader = TextLoader("example_knowledge.txt", encoding="utf-8")
# docs = loader.load()
# For demonstration, directly use example text
docs = [
"Example Tutorial (EXAMPLE) is a free programming learning website.",
"The website provides tutorials for multiple programming languages such as Python, Java, and HTML.",
"The Python3 basic tutorial has 30 chapters, suitable for beginners with zero foundation.",
"The HTML basic tutorial has 25 chapters, including forms, multimedia, and more.",
"All basic tutorials on Example are free.",
]
# Step 2: Split
text_splitter = RecursiveCharacterTextSplitter(
chunk_size=100,
chunk_overlap=20,
)
chunks = text_splitter.create_documents(docs)
# Step 3: Vectorize and store
embeddings = OpenAIEmbeddings(model="text-embedding-3-small")
vector_store = Chroma.from_documents(
documents=chunks,
embedding=embeddings,
persist_directory="./example_db",
)
print(f"Index built: {len(chunks)} document chunks")
# Step 4: Retrieve
results = vector_store.similarity_search("How many chapters does the Python tutorial have?", k=2)
for doc in results:
print(f"Search result: {doc.page_content}")
Running result:
Index built:5 个文档块 Search result: Python3 基础教程共 30 章,适合零基础入门学习。 Search result: Example is a free programming learning website.Other extensions