Vector Database
A vector database is a database system specifically designed for storing, indexing, and retrieving high-dimensional vector data.
You can think of it as: storing things with similar meanings together, and quickly finding the ones most similar to a given item.
Unlike traditional databases, which query through exact matching (WHERE name = 'Alice'), vector databases query through similarity (finding the 10 images most similar to this image).
An intuitive analogy
Imagine a library scenario:
| Database type | Retrieval method | Analogy |
|---|---|---|
| Traditional database | Exact search by book number and title | Find a book with a specified number |
| Vector database | Search by content relevance | Find "all science fiction novels similar in style to The Three-Body Problem" |
This semantic similarity is precisely the core problem that vector databases solve.
Why do we need a vector database?
Before diving into technical details, let's first understand what problems vector databases solve.
Limitations of traditional databases
Traditional relational databases (MySQL, PostgreSQL) are very good at handling structured data, but they fall short when faced with the following requirements:
- Image search (find visually similar images)
- Semantic search (when a user searches for "Apple phone", it can find related content about "iPhone")
- Recommendation systems (find "songs similar in style to the ones you like")
- Anomaly detection (find "logs that differ the most from normal behavior")
The common characteristic of these problems is the need to understand the "meaning" of content, rather than performing literal matching.
Problems with traditional approaches
用 LIKE '%苹果%' 搜索 → 找不到 "iPhone"、"Apple" 用全文索引搜索 → 找不到语义相关但用词不同的内容
Comparison diagram
The diagram below intuitively shows the fundamental differences in query methods between traditional databases and vector databases.
Core concepts: vectors and embeddings
Understanding vectors and embeddings is the first step to mastering vector databases.
What is a vector?
In mathematics, a vector is an ordered set of numbers.
[0.12, -0.54, 0.87, 0.03, ..., 0.61] ← 这就是一个向量
In machine learning, this set of numbers represents the semantic features of an object, typically with dimensions ranging from 128 to 4096.
What is an embedding?
Embedding is the process and result of converting real-world objects (text, images, audio, etc.) into vectors.
This conversion is performed by an embedding model, whose core idea is: objects with similar semantics have vectors that are closer in space.
Similar semantics means similar vectors
Use a simplified 2D example to understand (in reality it is hundreds to thousands of dimensions):
Key understanding: Two vectors that are close in vector space also have semantically similar original content. This is the foundation of all capabilities of vector databases.
Similarity calculation methods
The core of finding the "most similar vector" is to calculate the distance or similarity between two vectors. Here are the three most commonly used methods.
Cosine Similarity
Cosine similarity measures the angle between two vectors, ignoring their magnitude. It is the most commonly used method, especially suitable for text scenarios.
Formula:
\[ \text{CosineSimilarity}(A, B) = \frac{A \cdot B} {\|A\|\|B\|} = \frac{\sum_{i=1}^{n} A_i B_i} {\sqrt{\sum_{i=1}^{n} A_i^2}\sqrt{\sum_{i=1}^{n} B_i^2}} \]
- Result range: -1 to 1, larger value means more similar
- Applicable scenarios: text semantic search, document similarity
Euclidean Distance
Euclidean distance measures the straight-line distance between two points; the smaller the distance, the more similar they are.
Formula:
\[ d(A,B) = \sqrt{\sum_{i=1}^{n}(A_i-B_i)^2} \]
- Result range: 0 to ∞, smaller value means more similar
- Applicable scenarios: image retrieval, location-related applications
Dot Product
The dot product is the sum of vector component products, combining both direction and magnitude information.
Formula:
\[ A \cdot B = \sum_{i=1}^{n} A_i B_i \]
- Applicable scenarios: recommender systems (equivalent to cosine similarity when vectors are normalized)
Comparison of the three methods
Python code example
The following example demonstrates Python implementations of three similarity calculation methods:
Example
# Cosine similarity: measures directional similarity
def cosine_similarity(a, b):
return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))
# Euclidean distance: measures absolute position differences
def euclidean_distance(a, b):
return np.linalg.norm(a - b)
# Dot product: combines direction and magnitude
def dot_product(a, b):
return np.dot(a, b)
# Example vectors
v1 = np.array([0.12, -0.54, 0.87, 0.03])
v2 = np.array([0.10, -0.50, 0.90, 0.05])
v3 = np.array([-0.80, 0.20, -0.30, 0.70])
print(f"v1 vs v2 cosine similarity: {cosine_similarity(v1, v2):.4f}") # Approximately 0.9997 (very similar)
print(fv1 vs v3 cosine similarity: {cosine_similarity(v1, v3):.4f}) # Approximately -0.55 (not similar)
v1 vs v2 余弦相似度: 0.9997 v1 vs v3 余弦相似度: -0.5512
Vector indexing algorithms
When the data volume is large (millions, billions), computing similarity for every piece of data (brute-force search) is too slow. Vector databases use specialized indexing algorithms to accelerate queries.
Brute-force search (Flat / Brute-force)
Brute-force search traverses all vectors and computes similarity one by one.
| Dimension | Description |
|---|---|
| Principle | Traverse all vectors, compute similarity one by one |
| Advantages | 100% accurate results |
| Disadvantages | Extremely slow with large data volumes, O(n) complexity |
| Applicable | Data volume less than 100k, extremely high precision requirements |
IVF (Inverted File Index)
IVF execution steps:
- Training phase: Use K-Means to cluster all vectors into N clusters, record the center of each cluster
- Query phase: First find the centers of the nearest clusters, then perform exact search only within these clusters
HNSW (Hierarchical Navigable Small World graph)
HNSW is currently the most mainstream vector indexing algorithm, balancing speed and accuracy.
HNSW core idea:
- Build a multi-layer graph structure, sparse at the top, dense at the bottom
- At query time, start from the top-layer entry point and play a "hopscotch game": greedily jump to closer nodes at each layer, then descend to the next layer
- Greatly reduces the number of nodes that need to be compared, time complexity approximately O(log n)
Other common indexes
| Index type | Features | Applicable scenarios |
|---|---|---|
| Flat (brute force) | Accurate but slow | Small datasets, accuracy first |
| IVF_Flat | Exact search after clustering, fast | Medium to large scale, sufficient memory |
| IVF_PQ | Quantization compression, saves memory | Ultra-large scale, memory constrained |
| HNSW | Fast, high accuracy, high memory usage | Most commonly used, recommended first choice |
| ScaNN | Made by Google, optimized throughput | High-concurrency production environments |
Comparison of mainstream vector databases
The following is a horizontal comparison of the most mainstream vector databases currently available, helping you make choices in different scenarios.
Beginner advice: Start with Chroma or pgvector. The former is suitable for AI application prototypes, the latter for projects with existing PostgreSQL.
Quick start: Python examples
Below we use Chroma (easiest to get started) to demonstrate the complete CRUD workflow.
Installation
Example
Complete example: building a document semantic search system
The following code demonstrates end-to-end how to use Chroma to build a semantic-based document search system.
Example
from chromadb.utils import embedding_functions
# ─── 1. Initialize client ───────────────────────────────────────────
# Persist locally (recommended)
client = chromadb.PersistentClient(path="./my_vector_db")
# Use OpenAI embedding model (could also switch to a local model)
openai_ef = embedding_functions.OpenAIEmbeddingFunction(
api_key="your-openai-api-key", # Required: replace with your API Key
model_name="text-embedding-3-small" # 1536 dimensions, cost-effective
)
# ─── 2. Create collection (like a "table" in relational databases)────────────────────────
collection = client.get_or_create_collection(
name="my_documents", # Collection name
embedding_function=openai_ef, # Bind embedding function
metadata={"hnsw:space": "cosine"} # Use cosine similarity
)
# ─── 3. Insert Documents ──────────────────────────────────────────────
documents = [
Python is an object-oriented interpreted programming language, widely used in data science and AI development,
Machine learning is a subfield of artificial intelligence that enables computers to learn patterns from data,
Deep learning uses multi-layer neural networks and performs excellently in image recognition and NLP tasks,
Vector databases are specifically designed to store high-dimensional vectors and support semantic similarity search,
PostgreSQL is a powerful open-source relational database,
Redis is an in-memory high-performance key-value database, often used for caching,
Docker containerization technology allows applications to run consistently in any environment,
Git is a distributed version control system and a fundamental tool for modern software development,
]
ids = [f"doc_{i}" for i in range(len(documents))]
# Batch insert (Chroma automatically calls the embedding model to convert to vectors before storing)
collection.add(
documents=documents,
ids=ids,
metadatas=[{"source": "tutorial", "index": i} for i in range(len(documents))]
)
print(fInserted {len(documents)} documents)
已插入 8 条文档
Example
query = How to do artificial intelligence with Python
results = collection.query(
query_texts=[query],
n_results=3, # Return the top 3 most similar
include=["documents", "distances", "metadatas"]
)
print(f"\n"Query: {query}")
print("-" * 50)
for i, (doc, dist) in enumerate(zip(
results["documents"][0],
results["distances"][0]
)):
similarity = 1 - dist # Convert cosine distance to similarity
print(fRank {i+1} (similarity {similarity:.4f}):)
print(f" {doc}")
print()
查询:如何用 Python 做人工智能 -------------------------------------------------- 第 1 名(相似度 0.9231):Python 是一种面向对象的解释型编程语言... 第 2 名(相似度 0.8874):机器学习是人工智能的子领域... 第 3 名(相似度 0.8612):深度学习使用多层神经网络...
Example
results_filtered = collection.query(
query_texts=[Database technology],
n_results=2,
where={"source": "tutorial"}, # Search only among documents where source=tutorial
include=["documents", "distances"]
)
# ─── 6. Update Documents ──────────────────────────────────────────────
collection.update(
ids=["doc_0"],
documents=[Python is currently the most popular programming language, widely used in AI, data analysis, and web development],
metadatas=[{"source": "tutorial", "index": 0, "updated": True}]
)
# ─── 7. Delete Documents ──────────────────────────────────────────────
collection.delete(ids=["doc_7"]) # Delete Git-related documents
# ─── 8. View Collection Statistics ──────────────────────────────────────────
print(fCurrent collection document count: {collection.count()})
Without using third-party embedding APIs (fully local)
If you don't want to use the OpenAI API, you can run completely offline with local embedding models.
Example
from sentence_transformers import SentenceTransformer
# Use local embedding models (no API key required, fully offline)
model = SentenceTransformer("paraphrase-multilingual-MiniLM-L12-v2") # Supports Chinese
client = chromadb.Client()
collection = client.create_collection("local_demo")
texts = ["The weather is nice today", "It's sunny and perfect for going out", "The stock market surged", "It might rain tomorrow"]
# Manually generate vectors before insertion
embeddings = model.encode(texts).tolist()
collection.add(
embeddings=embeddings,
documents=texts,
ids=[f"id_{i}" for i in range(len(texts))]
)
# Query
query_embedding = model.encode(["What's the weather like today?"]).tolist()
results = collection.query(query_embeddings=query_embedding, n_results=2)
print(results["documents"])
# Output: [['The weather is nice today', 'It's sunny and perfect for going out']]
pgvector example (for PostgreSQL users)
If your project already uses PostgreSQL, pgvector is the lightest way to integrate.
Example
CREATE EXTENSION IF NOT EXISTS vector;
-- Create table: store article titles and their vectors (1536 dimensions)
CREATE TABLE articles (
id SERIAL PRIMARY KEY,
title TEXT NOT NULL,
content TEXT,
embedding vector(1536) -- Vector column, 1536 dimensions
);
-- Create HNSW index to speed up queries
CREATE INDEX ON articles
USING hnsw (embedding vector_cosine_ops)
WITH (m = 16, ef_construction = 64);
-- Insert data (vectors are generated in the application layer and passed in)
INSERT INTO articles (title, embedding)
VALUES ('Python Beginner's Guide', '[0.12, -0.54, 0.87, ...]'::vector);
-- Semantic search: find the 5 most similar articles
SELECT id, title,
1 - (embedding <=> '[0.10, -0.50, 0.90, ...]'::vector) AS similarity
FROM articles
ORDER BY embedding <=> '[0.10, -0.50, 0.90, ...]'::vector
LIMIT 5;
Note: <=> is a vector distance operator provided by pgvector, used to compute cosine distance. Subtract the cosine distance from 1 to get cosine similarity.
Typical application scenarios
Vector databases have a wide range of applications in the AI era; here's an overview of the six most typical scenarios.
RAG (Retrieval-Augmented Generation) architecture
RAG is currently one of the primary application scenarios for vector databases. Below is its core workflow:
Selection recommendations and best practices
Selection decision tree
Based on your specific situation, choose an appropriate vector database according to the following decision tree:
你的情况是什么?
│
├─── 已有 PostgreSQL,且数据量 < 500 万
│ └──> 用 pgvector,无缝集成,零额外运维
│
├─── 做 AI/LLM 应用原型,快速验证
│ └──> 用 Chroma,几行代码跑起来
│
├─── 需要生产级部署,性能优先,数据量 500 万 ~ 1 亿
│ └──> 用 Qdrant,Rust 实现,性能强
│
├─── 超大规模(> 1 亿),有 K8s 运维能力
│ └──> 用 Milvus,分布式,功能最全
│
└─── 团队没有运维能力,愿意付费
└──> 用 Pinecone 云服务,开箱即用
Choosing an embedding model
Choosing a suitable embedding model is the key first step in vector database applications.
| Requirements | Recommended Model |
|---|---|
| Chinese and English text (High Quality) | OpenAI text-embedding-3-small |
| Chinese text (Local Offline) | BAAI/bge-large-zh-v1.5 |
| Multilingual universal. | paraphrase-multilingual-MiniLM-L12-v2 |
| Image-text multimodal. | OpenAI CLIPseries |
Performance optimization tips
1. Batch insertion: Insert multiple entries at once to avoid frequent single-entry writes.
Example
collection.add(documents=docs_list, ids=ids_list)
# Not recommended: looping one by one (re-indexes every time, extremely inefficient)
# for doc, id in zip(docs_list, ids_list):
# collection.add(documents=[doc], ids=[id])
2. Vector normalization: Before using cosine similarity, normalizing vectors in advance can speed up computation.
Example
def normalize(v):
"""Perform L2 normalization on vectors so that the norm is 1"""
return v / np.linalg.norm(v)
3. Set n_results reasonably: Don’t blindly set a very large top_k; 3~10 results are usually enough for RAG scenarios.
4. Make good use of metadata filtering: Combine with where conditions during search to narrow the search scope.
Example
collection.query(
query_texts=["Python exception handling"],
where={"category": "tech_doc"},
n_results=5
)
5. Rebuild indexes regularly: After data volume grows, rebuild the HNSW index in a timely manner to maintain query performance.
Common pitfalls
The following are the most common problems beginners encounter when using vector databases, along with solutions.
| Problem | Description | Solution |
|---|---|---|
| Embedding model must be consistent | The same model must be used for insertion and query | Pin the model version in the configuration file |
| Dimension mismatch | Switched models but didn’t rebuild the collection | Delete and rebuild the collection when changing models |
| Text too long | Most models have token limits (512~8192) | Chunk overly long text first |
| Inaccurate similarity | Text is not chunked, so semantics are diluted | Split documents by paragraph or fixed length |
| Slow cold start | When data volume is large, the initial index loading takes time | Pre-warm in advance, or use persistent indexes |
Summary
Let’s review the core knowledge points of this tutorial:
Vector databases are a crucial part of AI-era infrastructure.
It solves the semantic similarity search problem that traditional databases cannot handle, and is a core component for building RAG systems, recommendation systems, and multimodal search.
Learning path recommendations
- Step 1: Understand the concepts of vectors and embeddings, and run the Chroma example from this article.
- Step 2: Try building a simple document Q&A system with LangChain + Chroma.
- Step 3: Learn indexing algorithms such as HNSW, and understand the trade-off between accuracy and speed.
- Step 4: Based on actual project requirements, select an appropriate vector database and deploy it in a production environment.