LangChain Document Loading and Splitting

In previous articles we manually entered text, but in real projects, documents may come from PDFs, web pages, Markdown files, etc.

This section introduces how to use Document Loader to load various types of documents, and how to use Text Splitter to split documents into small chunks suitable for retrieval.


Document Loader—Loading Documents

LangChain provides dozens of document loaders, covering common file formats:

LoaderSourceInstallation Package
TextLoader.txt fileslangchain (built-in)
PyPDFLoaderPDF fileslangchain-community + pypdf
WebBaseLoaderWebpage URLlangchain-community + beautifulsoup4
CSVLoaderCSV fileslangchain-community
UnstructuredMarkdownLoaderMarkdown fileslangchain-community + unstructured

Example

# Load a text file (built-in, no additional installation required)
from langchain_community.document_loaders import TextLoader

loader = TextLoader("knowledge.txt", encoding="utf-8")
docs = loader.load()

print(f"Loaded {len(docs)} documents")
print(f"Content preview: {docs)

# Load a webpage
# pip install langchain-community beautifulsoup4
from langchain_community.document_loaders import WebBaseLoader

loader = WebBaseLoader("https://www.example.com/python/python-tutorial.html")
docs = loader.load()
print(f"\nWebpage content: {docs)

Text Splitter—Document Splitting

Documents are usually too long and need to be split into small chunks for effective retrieval. The splitting strategy directly affects RAG performance:

Example

from langchain_text_splitters import RecursiveCharacterTextSplitter

# Create a splitter
text_splitter = RecursiveCharacterTextSplitter(
    chunk_size=500,         # Maximum of 500 characters per chunk
    chunk_overlap=50,       # 50-character overlap between chunks
    separators=["\n\n", "\n", "。", "!", "?", ";", ",", " ", ""],
    # Split by paragraph first, then sentence, and finally by character
)

# Example document
long_text = """Example Tutorial (EXAMPLE) is a free programming learning platform.

The platform provides rich programming language tutorials, including but not limited to:
- Python tutorial: from basic syntax to data analysis
- Java tutorial: object-oriented programming to Spring framework
- Frontend tutorial: HTML, CSS, JavaScript, and their frameworks

All tutorials come with detailed code examples and an online running environment.
Learners can quickly master programming skills by learning while practicing."""


# Split the document
chunks = text_splitter.split_text(long_text)

print(f"Original length: {len(long_text)} characters")
print(f"After splitting: {len(chunks)} chunks\n")

for i, chunk in enumerate(chunks):
    print(f"--- Chunk {i+1} ({len(chunk)} characters) ---")
    print(chunk)
    print()

Running result:

Original length: 153 字
Chunks after split: 3 块

--- 块 1 (54 字) ---
Example is a free programming learning platform.
平台提供了丰富的编程语言教程,包括但不限于:

--- 块 2 (49 字) ---
- Python 教程:从基础语法到数据分析
- Java 教程:面向对象编程到 Spring 框架

--- 块 3 (50 字) ---
- 前端教程:HTML、CSS、JavaScript 及其框架
所有教程都配有详细的代码示例和在线运行环境。

chunk_overlap is very important. If there is no overlap between chunks, a complete sentence may be split in half, causing key information to be missed during retrieval. An overlap of 50-100 characters is a common setting.


Chunking Parameter Settings Guide

Scenariochunk_sizechunk_overlapReason
FAQ Q&A200~50020~50Q&A pairs are short, small chunks are sufficient
Technical documentation500~100050~100Technical content requires more context
Long articles/papers1000~2000100~200Need to preserve paragraph completeness
Code repository500~15000~50Functions/classes as natural boundaries

Complete Flow: Load → Split → Vectorize

Example

from langchain_text_splitters import RecursiveCharacterTextSplitter
from langchain_openai import OpenAIEmbeddings
from langchain_chroma import Chroma
# from langchain_community.document_loaders import TextLoader


# Step 1: Load
# loader = TextLoader("example_knowledge.txt", encoding="utf-8")
# docs = loader.load()

# For demonstration, directly use example text
docs = [
    "Example Tutorial (EXAMPLE) is a free programming learning website.",
    "The website provides tutorials for multiple programming languages such as Python, Java, and HTML.",
    "The Python3 basic tutorial has 30 chapters, suitable for beginners with zero foundation.",
    "The HTML basic tutorial has 25 chapters, including forms, multimedia, and more.",
    "All basic tutorials on Example are free.",
]

# Step 2: Split
text_splitter = RecursiveCharacterTextSplitter(
    chunk_size=100,
    chunk_overlap=20,
)
chunks = text_splitter.create_documents(docs)

# Step 3: Vectorize and store
embeddings = OpenAIEmbeddings(model="text-embedding-3-small")
vector_store = Chroma.from_documents(
    documents=chunks,
    embedding=embeddings,
    persist_directory="./example_db",
)

print(f"Index built: {len(chunks)} document chunks")

# Step 4: Retrieve
results = vector_store.similarity_search("How many chapters does the Python tutorial have?", k=2)
for doc in results:
    print(f"Search result: {doc.page_content}")

Running result:

Index built:5 个文档块
Search result: Python3 基础教程共 30 章,适合零基础入门学习。
Search result: Example is a free programming learning website.
Other extensions