Cosine Similarity for Text Recommendation

Represent sentences as bag-of-words vectors, compute cosine similarity, and find the most similar sentence.

After completing this case, you will understand:The mathematical foundation of search and recommendation systems — the dot product measures 'same direction', and the norm is used for normalization.


Everyday Life Introduction

Taobao search for 'Bluetooth headphones'

When you type 'Bluetooth headphones', Taobao instantly finds the most relevant items from millions of products. It doesn't rank by whether titles contain exact matching words — instead, it converts each product title into a string of numbers (a vector), and then compares whether the 'direction' of your search term matches that of each title.

The closer the directions, the more similar the meanings. Whether the title is long or short doesn't affect the judgment — this iscosine similarity's intuition.


Intuitive Understanding

Imagine arrows shot from two bows. The more consistent the directions of the two arrows, the closer the cosine of the angle is to 1; when the two arrows are perpendicular, the cosine is 0; when the two arrows are in opposite directions, the cosine is -1.

~ 1
Same direction
The two sentences have the same meaning
~ 0
Perpendicular to each other
The two sentences are completely unrelated
~ -1
Opposite directions
Rare in the bag-of-words model

The key question is: how do you turn a piece of text into an arrow? This requires introducing the bag-of-words model.


Mathematical Definition

Cosine Similarity Formula

\[ \cos(\theta) = \frac{\mathbf{a} \cdot \mathbf{b}}{\|\mathbf{a}\| \cdot \|\mathbf{b}\|} = \frac{\sum_{i=1}^{n} a_i \cdot b_i}{\sqrt{\sum a_i^2} \cdot \sqrt{\sum b_i^2}} \]

The numerator is the dot product (measuring how much the two vectors point in the 'same direction'), and the denominator is the product of the two norms (eliminating the effect of vector length).

Bag of Words Model

Method to turn a sentence into a vector:

  1. Determine a vocabulary (all words that may appear)
  2. For each sentence, if a word appears, mark it 1; if it doesn't, mark it 0.
  3. Obtain a vector with length equal to the vocabulary size.

Here we use the simplest 'whether it appears' (one-hot bag of words). In practice, TF-IDF or word embeddings are commonly used instead. But the core idea is the same: turn text into vectors that can participate in mathematical operations.


Hands-on Python Practice

Use 5 Chinese sentences to construct a bag-of-words vector, compute the cosine similarity with the query sentence, and find the most similar sentence.

Example

import numpy as np

# 1. Prepare corpus - 5 Chinese sentences, covering two topics: AI learning and daily life
sentences = [
    "I like to use python to learn machine learning",
    "python is a good tool for learning artificial intelligence",
    "The weather is nice today, suitable for going out for a walk",
    "Deep learning requires a lot of data and computing power",
    "Walking is a good way to relax",
]

# 2. Build vocabulary (simplified teaching version, manually define possible words)
vocab_words = ["I", "like", "use", "python", "learn", "machine learning", "is",
               "artificial intelligence", "good", "tool", "today", "weather", "very", "suitable",
               "go out", "walk", "deep learning", "need", "a lot of", "of", "data",
               "computing power", "a kind of", "relax", "way"]

def to_vector(sentence, vocab):
    """Convert sentences to bag-of-words vectors — mark appearing words as 1"""
    vec = np.zeros(len(vocab))
    for i, w in enumerate(vocab):
        if w in sentence:
            vec[i] = 1  # If a word appears, mark it as 1
    return vec

# 5 sentences → 5 vectors
vectors = np.array([to_vector(s, vocab_words) for s in sentences])
print("EXAMPLE Vector dimension of each sentence:", vectors.shape[1],
      "(vocabulary size)")

# 3. Write cosine similarity function manually (without calling sklearn)
def cosine_similarity(a, b):
    dot = np.dot(a, b)             # dot product
    norm_a = np.linalg.norm(a)     # magnitude of a
    norm_b = np.linalg.norm(b)     # magnitude of b
    if norm_a == 0 or norm_b == 0:
        return 0.0                 # zero vector has no direction
    return dot / (norm_a * norm_b)

# 4. Given a query sentence, find the most similar sentence
query = "I am learning python and artificial intelligence"
query_vec = to_vector(query, vocab_words)

sims = [cosine_similarity(query_vec, v) for v in vectors]

print(f"\nEXAMPLE Query sentence: {query!r}\n")
for s, sim in sorted(zip(sentences, sims),
                     key=lambda x: -x[1]):
    print(f" Similarity {sim:.3f} <- {s}")

# 5. Verification: similarity of unrelated sentences is 0
print("\nEXAMPLE Verification: two completely unrelated sentences")
v1 = to_vector("python machine learning", vocab_words)
v2 = to_vector("walk relax", vocab_words)
print(f" 'python machine learning' vs 'walk relax': "
      f"cos = {cosine_similarity(v1, v2):.3f}")
EXAMPLE 每句话的向量维度: 25 (词表大小)

EXAMPLE 查询句子: '我在学习python和人工智能'

  相似度 0.577  <-  python是学习人工智能的好工具
  相似度 0.500  <-  我喜欢用python学习机器学习
  相似度 0.000  <-  今天天气很好适合出去散步
  相似度 0.000  <-  深度学习需要大量的数据和算力
  相似度 0.000  <-  散步是一种很好的放松方式

EXAMPLE 验证:两个完全不相关的句子
  'python机器学习' vs '散步放松': cos = 0.000

The result is intuitive: sentences containing the same keywords have high similarity, and sentences with completely unrelated content have similarity of 0.


Application Scenarios in AI

AI scenariosHow to use cosine similarity
Search enginesConvert both search terms and documents into vectors, sort by cosine similarity, and return results.
Recommender systemsFind products most similar to the user's historical behavior vector.
Semantic searchUse word embeddings (Word2Vec/BERT) for vectorization to capture synonyms—"Bluetooth headphones" and "wireless earbuds" will also match.
Face recognitionEncode face images into vectors (Face Embedding), and compare the cosine similarity of two faces to determine whether they are the same person.
RAG retrievalIn large model applications, use cosine similarity to retrieve the most relevant document fragments from a knowledge base.

In the "attention mechanism" used in large models such as GPT, the dot product of Q and K (QK^T) is essentially also computing similarity—cosine similarity omits the magnitude normalization of the denominator.

Other extensions