RAG and Knowledge Retrieval

RAG (Retrieval-Augmented Generation) is currently one of the most mainstream LLM deployment architectures.

The core idea of RAG is:When answering questions, let the LLM first retrieve relevant content from an external knowledge base, then generate a response based on the retrieval results, rather than relying only on knowledge memorized during model training.

This solves two core pain points of LLMs: knowledge cutoff (the model does not know events that occurred after training) and hallucination (the model fabricates answers when uncertain).


Basic Principles of RAG

A complete RAG system consists of two pipelines:Offline indexing pipeline(preprocesses documents and stores them in a vector database) andonline query pipeline(receives user questions, retrieves, and generates).

In the offline stage, raw documents are split into small chunks, converted into vectors via an Embedding model, and stored in a vector database.

In the online stage, the user question is also converted into a vector, the most similar document chunks are found from the database, concatenated into context, and passed to the LLM to generate an answer.

The figure below shows the complete request flow of RAG:

Basic RAG Flowchart RAG complete request flow: user questions are converted into query vectors via the Embedding model, similar document chunks are retrieved in the vector database, the question and retrieval results are concatenated into a Prompt, and input to the LLM to generate the final answer. User question Embedding Convert to query vector Vector database Similarity retrieval Top-K Prompt concatenation Question + document chunks LLM generation Answer based on retrieval results Offline indexing (one-time preprocessing) Original documents Chunking / cleaning Embedding Convert to document vectors Write Vector database Online query flow (every request goes through)

Data Preprocessing and Document Chunking

Prerequisite Challenge: Complex Document Parsing

Before chunking, RAG often facesformat parsingchallenges. Especially for tables, images, and multi-column layouts in PDFs, Word documents, or scanned files, ordinary text extraction can easily cause semantic confusion.

The current mainstream industry approach is to introducedocument parsing engines(such as LlamaParse, Unstructured) or multimodal large models, converting complex images and text into structured Markdown, laying the foundation for subsequent high-quality chunking.

Document Chunking Strategies

Document chunking is the foundation of RAG effectiveness; chunk granularity directly affects retrieval quality. Chunks that are too large introduce noise, while chunks that are too small lose context. Common strategies are as follows:

Chunking strategy Applicable scenarios Advantages Disadvantages
Fixed-size chunking General text Simple to implement, fast May cut off semantically complete sentences
Recursive character chunking Structured text (Markdown, code) Prefers splitting along semantic boundaries such as paragraphs and sentences Slightly complex to implement; requires setting a reasonable list of separators
Semantic chunking Long documents, books Uses Embedding to compute similarity between adjacent sentences and automatically finds semantic turning points for chunking High computational cost, slow preprocessing
Parent-child document retrieval
(Small-to-Big)
Comprehensive coverage scenarios Uses "small chunks" for high-precision vector retrieval, and upon a hit returns the corresponding "large chunk" (parent document) to the LLM, balancing retrieval precision and context completeness. Database design and maintenance costs double

In practice, it is common to add during chunkingoverlap, i.e., adjacent chunks share several characters to prevent important information from being truncated at boundaries. Typical configuration: chunk size 512 tokens, overlap 50~100 tokens.

Example: Recursive Chunking with LangChain

from langchain.text_splitter import RecursiveCharacterTextSplitter

splitter = RecursiveCharacterTextSplitter(
    chunk_size=512,        # Maximum number of tokens per chunk
    chunk_overlap=50,      # Overlap token count between adjacent chunks to prevent information loss at boundaries
    separators=["\n\n", "\n", "。", ".", " ", ""]  # Prioritize splitting by paragraphs and sentences
)

chunks = splitter.split_text(document_text)
print(f"Split into {len(chunks)} document chunks")

Vector Retrieval

Embedding Models

Embedding models convert text into dense vectors (usually 768- or 1536-dimensional float arrays). Semantically similar texts are closer in vector space, which is the mathematical basis of similarity retrieval.

Comparison of common embedding models:

Model Dimensions Supported languages Features
text-embedding-3-small(OpenAI) 1536 Multilingual High cost-effectiveness, suitable for large-scale indexing
text-embedding-3-large(OpenAI) 3072 Multilingual Highest accuracy, higher cost
BAAI/bge-m3 1024 Chinese and English Open source, excellent Chinese performance, supports multiple languages
sentence-transformers/all-MiniLM-L6-v2 384 English Small size, fast speed, suitable for extremely lightweight local deployment

Similarity Computation and ANN Algorithms

The core of retrieval is measuring distance. The most commonly used isCosine Similarity, which computes the cosine of the angle between two vectors, with a range of [-1, 1]; the closer to 1, the more similar. Additionally, there are Dot Product and Euclidean Distance (L2 Distance).

To achieve millisecond-level retrieval among millions of vectors, databases typically useApproximate Nearest Neighbor (ANN) algorithms(such asHNSW, IVF). HNSW is currently the most mainstream algorithm. It builds a multi-layer skip graph network, sacrificing very little accuracy in exchange for an order-of-magnitude improvement in search speed.


Advanced RAG (Advanced Architecture)

The basic architecture (Naive RAG) often faces issues such as inaccurate retrieval and "context flooding" caused by redundant information. Advanced RAG addresses these throughpre-retrieval optimization → retrieval fusion → post-retrieval optimizationusing a three-stage architecture to solve them.

1. Pre-retrieval: Query Optimization

Users' original questions are often not precise:

  • Query Rewriting: Use LLM to rewrite colloquial queries into standardized retrieval terms.
  • HyDE(Hypothetical Document Embedding): Let the LLM first "blind-guess" a hypothetical answer. Since the generated answer usually contains more industry terminology than the original question, using the vector of this hypothetical answer for retrieval often recalls higher-quality documents.

2. Hybrid Search

willVector Retrieval(Understands semantics, high fault tolerance) andKeyword RetrievalThe results are fused by weight. This is especially important when encountering proper nouns, product models, and code snippets, because traditional vector retrieval is prone to "failing" on specific proper nouns.

3. Post-retrieval Optimization: Reranking

This is acoarse ranking → fine rankingtwo-stage design. Although vector retrieval is fast, its scoring is not precise enough. Reranking introducesCross-Encoder model(such as `bge-reranker`), which inputs "question" and "document" pairs into the model for joint inference and scoring. It is computationally heavy and only responsible for selecting the Top-20 down to Top-5.

Example: Reranking Pipeline Pseudocode

from sentence_transformers import CrossEncoder

reranker = CrossEncoder("BAAI/bge-reranker-v2-m3")

# 1. Coarse ranking: vector retrieval rapidly recalls Top-50
candidates = vector_store.similarity_search(query, k=50)

# 2. Fine ranking: construct [question, document] pairs for precise scoring
pairs = [[query, doc.page_content] for doc in candidates]
scores = reranker.predict(pairs)

# 3. Filter the final Top-5 to pass into the LLM
ranked_docs = sorted(zip(scores, candidates), reverse=True)
final_docs = [doc for _, doc in ranked_docs[:5]]

4. Self-RAG and CRAG (Corrective RAG)

A self-reflection mechanism is added. For example, CRAG (Corrective RAG), after obtaining retrieval results, first has the LLM act as a "judge" to score them. If the local knowledge base has no matching document or the quality is extremely low, the system automatically triggers Web Search (such as Google API) as a supplement, greatly reducing hallucinations.


GraphRAG: Knowledge Graph + Retrieval Integration

Traditional RAG treats the knowledge base as independent text fragments and cannot answer questions such as "find all companies founded by the current CEO and with a market value exceeding 100 billion" that requirecross-document, multi-hop reasoningcomplex problems.GraphRAGIntroducing a Knowledge Graph to explicitly model entities and relationships.

GraphRAG architecture diagram GraphRAG fuses vector retrieval with the knowledge graph: the user question triggers both vector retrieval and graph retrieval, and the results of the two paths are fused and fed into the LLM to generate an answer. User question Vector retrieval Relevant document chunks Graph retrieval Entity-relationship subgraph Context fusion Document chunks + graph paths Entity A → Relationship → Entity B Entity B → Relationship → Entity C

GraphRAG Core Steps

  1. Knowledge Construction: In the offline phase, use LLM to extract triples (subject, relationship, object) from documents and write them into graph databases such as Neo4j.
  2. Dual-Path Retrieval: For entities in the query, not only is traditional vector retrieval performed, but graph traversal is also triggered in the knowledge graph to extract multi-hop relationship chains.
  3. Graph-Text Fusion Generation: Assemble the "chunks" retrieved by vector retrieval and the "path structures" retrieved by graph retrieval into the Prompt, giving the LLM both a global view and specific details.

GraphRAG Content Reference:https://www.example.com/ai-agent/graphrag-usage.html


Technology and Database Selection Recommendations

Database/Tool Selection Type Recommended Implementation Scenarios
Pinecone / Zilliz Cloud Fully Managed Cloud Service Out-of-the-box, no need to maintain infrastructure. Combined with Cohere Rerank + GPT-4o, it is the fastest commercial solution.
Qdrant Open Source + Managed Written in Rust, excellent memory management, extremely high performance. Suitable for enterprise-level private deployment.
Weaviate / Elasticsearch Open Source + Managed Comes with an extremely mature BM25 + vector hybrid search (Hybrid Search), and is the top choice for scenarios with many proper nouns.
Milvus Open Source Distributed Suitable for ultra-large-scale enterprise retrieval platforms at the billion to ten-billion level.
Chroma / FAISS Local Library/Embedded Extremely lightweight, no need to deploy a standalone service. Very suitable for local development and personal knowledge base project validation.

RAG Evaluation Metrics (RAGAS Framework)

RAG system evaluation cannot rely on intuition alone; it mainly usesRAGASa framework for automated quantitative testing from the two dimensions of "retrieval" and "generation":

  • Context Recall: What proportion of the information in the standard answer can be retrieved.
  • Context Precision: What proportion of the retrieved documents are truly relevant.
  • Faithfulness (Faithfulness/Hallucination Metric): Whether the generated answers are all supported by the retrieved documents.
  • Answer Relevance: Whether the generated answer truly addresses the user's question, avoiding irrelevant responses.
Other Extensions