Skip to content
beginner

Embedding Similarity Math: Dot Products, Cosine, and Retrieval Ranking

A retrieval system returns a top result with a similarity score of 0.83. Most people read that number as "83% relevant" or "83% likely to be correct." It…

Published 2026-10-03Updated 2026-10-049 min read
Two backpackers waiting at a railway station, ready for their travel adventure.
Two backpackers waiting at a railway station, ready for their travel adventure. Photo by Ketut Subiyanto on Pexels.

A retrieval system returns a top result with a similarity score of 0.83. Most people read that number as "83% relevant" or "83% likely to be correct." It is neither. It is a geometric measurement between two lists of numbers, and once you can compute it by hand, you stop trusting the number and start trusting the ranking.

This article assumes you already know that embeddings turn meaning into vectors. Here we work on the arithmetic of comparing them: how the score is built, how a similarity measure turns a pile of candidate vectors into an ordered list, and why a high score is evidence of similarity, not proof of truth.

What the Similarity Score Actually Measures

An embedding places a piece of text as a point in a high-dimensional space. Comparing two embeddings means comparing two arrows: how aligned are they, and how far apart do their tips sit?

A similarity score answers one narrow question: how aligned are these two vectors? It does not answer "is this document true," "is this the answer," or "should the model trust this." Those are separate judgments that happen after retrieval, not inside the score.

To keep every calculation visible, we will use a tiny 3-dimensional vector set as a stand-in for real embeddings, which often have hundreds or thousands of dimensions. The arithmetic is identical; only the number of terms changes.

Three measures show up constantly in retrieval work, and they are related but not interchangeable:

  • Dot product — the sum of component-wise products.
  • Cosine similarity — the dot product with vector length divided out.
  • Euclidean distance — the straight-line distance between vector tips.

We will derive the first two and use the third as a reference point.

Notation and Assumptions Before We Calculate

A vector is an ordered list of numbers. We write a vector aa with components a1,a2,…,ana_1, a_2, \dots, a_n, where nn is the dimension — the number of numbers in the list.

The dot product of two vectors multiplies matching components and sums the results:

a⋅b=a1b1+a2b2+⋯+anbna \cdot b = a_1 b_1 + a_2 b_2 + \dots + a_n b_n

The magnitude (length) of a vector is the square root of the vector dotted with itself:

∥a∥=a⋅a=a12+a22+⋯+an2\|a\| = \sqrt{a \cdot a} = \sqrt{a_1^2 + a_2^2 + \dots + a_n^2}

Three assumptions matter for everything that follows:

  1. Vectors are real-valued and non-zero.
  2. The query and the documents were embedded by the same model.
  3. Both sides share the same dimension nn.

Warning: Comparing vectors from different models, or vectors with different dimensions, is meaningless. The numbers do not share a coordinate system, so any score you compute is noise dressed as arithmetic.

Deriving the Dot Product Step by Step

A query vector and a document vector point in similar directions; a second document vector points the same way but is twice as long. A compact comparison shows dot products of 16 and 32, while both cosine similarities are about 0.933.
Scaling a vector changes its dot product, but not its cosine similarity or direction.

Let's compute a dot product on real numbers. Suppose our query vector is:

q=[1,2,3]q = [1, 2, 3]

And one candidate document vector is:

d=[2,1,4]d = [2, 1, 4]

Multiply matching components and add:

q⋅d=(1)(2)+(2)(1)+(3)(4)=2+2+12=16q \cdot d = (1)(2) + (2)(1) + (3)(4) = 2 + 2 + 12 = 16

That is the whole calculation. A large positive dot product means the vectors point broadly in the same direction and/or are long. Both effects are baked into the same number.

Here is the trap. Take a document vector that is simply longer, but pointed in the same direction:

d′=[4,2,8]d' = [4, 2, 8]

q⋅d′=(1)(4)+(2)(2)+(3)(8)=4+4+24=32q \cdot d' = (1)(4) + (2)(2) + (3)(8) = 4 + 4 + 24 = 32

The dot product jumped from 16 to 32, but d′d' is just dd scaled by 2. The score doubled because the vector got longer, not because it got more relevant. This is why raw dot product is sensitive to magnitude, and why some systems use it deliberately: when length carries a real signal — popularity, frequency, recency — you may want that signal in the score. When length is noise, you do not.

Knowledge check

Check your understanding

Answer this question before you continue.

For q = [1, 2, 3] and d′ = [4, 2, 8], what is q · d′?
Single Choice

Focus: Calculate a dot product and recognize its sensitivity to vector magnitude.

Deriving Cosine Similarity from the Dot Product

The dot product has a geometric identity behind it:

a⋅b=∥a∥ ∥b∥cos⁡θa \cdot b = \|a\| \, \|b\| \cos\theta

where θ\theta is the angle between the two vectors. Rearranging isolates the cosine:

cos⁡θ=a⋅b∥a∥ ∥b∥\cos\theta = \frac{a \cdot b}{\|a\| \, \|b\|}

That expression is cosine similarity. It is the dot product with both lengths divided out, so only direction survives.

Let's compute it for q=[1,2,3]q = [1, 2, 3] and d=[2,1,4]d = [2, 1, 4]. We already have q⋅d=16q \cdot d = 16. Now the magnitudes:

∥q∥=12+22+32=14≈3.742\|q\| = \sqrt{1^2 + 2^2 + 3^2} = \sqrt{14} \approx 3.742

∥d∥=22+12+42=21≈4.583\|d\| = \sqrt{2^2 + 1^2 + 4^2} = \sqrt{21} \approx 4.583

Divide:

cos⁡θ=163.742×4.583≈1617.15≈0.933\cos\theta = \frac{16}{3.742 \times 4.583} \approx \frac{16}{17.15} \approx 0.933

The result is bounded. Cosine similarity runs from -1 (opposite directions) through 0 (orthogonal, no shared direction) to 1 (identical direction). Because length is divided out, a long document and a short document can score identically if they point the same way — which is usually what you want for text embeddings, where document length is not a signal of relevance.

Knowledge check

Check your understanding

Answer this question before you continue.

The article's d = [2, 1, 4] is changed to d′ = [4, 2, 8], which points in the same direction. What happens to cosine similarity with q = [1, 2, 3]?
Comparison Reasoning

Focus: Explain why scaling a vector without changing its direction leaves cosine similarity unchanged.

Worked Example: Ranking Three Documents Against One Query

Now let's turn the formula into a nearest neighbor retrieval ranking. Keep the query:

q=[1,2,3]q = [1, 2, 3]

And add three candidate documents:

d1=[2,1,4],d2=[3,3,3],d3=[−1,0,1]d_1 = [2, 1, 4], \quad d_2 = [3, 3, 3], \quad d_3 = [-1, 0, 1]

Cosine similarity for d1d_1 we already computed: ≈0.933\approx 0.933.

Cosine similarity for d2d_2:

q⋅d2=(1)(3)+(2)(3)+(3)(3)=3+6+9=18q \cdot d_2 = (1)(3) + (2)(3) + (3)(3) = 3 + 6 + 9 = 18

∥d2∥=9+9+9=27≈5.196\|d_2\| = \sqrt{9 + 9 + 9} = \sqrt{27} \approx 5.196

cos⁡θ=183.742×5.196≈1819.44≈0.926\cos\theta = \frac{18}{3.742 \times 5.196} \approx \frac{18}{19.44} \approx 0.926

Cosine similarity for d3d_3:

q⋅d3=(1)(−1)+(2)(0)+(3)(1)=−1+0+3=2q \cdot d_3 = (1)(-1) + (2)(0) + (3)(1) = -1 + 0 + 3 = 2

∥d3∥=1+0+1=2≈1.414\|d_3\| = \sqrt{1 + 0 + 1} = \sqrt{2} \approx 1.414

cos⁡θ=23.742×1.414≈25.29≈0.378\cos\theta = \frac{2}{3.742 \times 1.414} \approx \frac{2}{5.29} \approx 0.378

Sort descending: d1d_1 (0.933), d2d_2 (0.926), d3d_3 (0.378). The retrieval rule is plain: the highest score wins the top slot. Take the top kk and hand them to the next stage.

Now re-run the same comparison with raw dot product:

DocumentDot productCosine similarity
d1d_1160.933
d2d_2180.926
d3d_320.378

The ranking flips at the top. Dot product puts d2d_2 first because it is longer; cosine puts d1d_1 first because it points more directly at the query. Neither is "wrong." The measure you choose is a design decision about what you want the score to reward.

Knowledge check

Check your understanding

Answer this question before you continue.

Using the article's scores for d₁, d₂, and d₃, which document is first under each ranking rule?
Output Prediction

Focus: Derive the top-ranked document under dot-product and cosine ranking from the worked example.

Why High Scores Are Not Proof of Relevance

Here is the misconception that causes the most damage: treating the score as an absolute measure of relevance or truth.

Many embedding models compress scores into a narrow high band. In practice, unrelated sentence pairs from some models can still land above 0.7, because the model was trained to push relevant pairs higher than irrelevant ones — not to spread scores across the full -1 to 1 range. A 0.75 can be a weak match in one model and a strong one in another. The absolute number is not portable across models.

Note: What matters is the relative order of scores, not their absolute values. An embedding model is useful as long as relevant pairs score higher than irrelevant pairs, even when the gap is small.

That gives you a practical rule: use scores to form a model-specific candidate ordering, then evaluate whether that ordering serves your task. Thresholds must be calibrated per model and per dataset, usually by inspecting the score distribution on your own data rather than borrowing a cutoff from a blog post.

The score also cannot answer three separate questions:

  • Is this relevant? The score is evidence, not a verdict.
  • Is this correct? Similarity is not truth. A confidently wrong passage can sit close to the query.
  • Is this sufficient to answer the user? Retrieval finds candidates; it does not decide whether they cover the question.

Knowledge check

Check your understanding

Answer this question before you continue.

A retrieved passage has cosine similarity 0.75. Which conclusion is supported by the article?
Misconception Check

Focus: Interpret a similarity score as model-dependent ranking evidence rather than a probability or proof of relevance.

Choosing a Measure and Knowing When It Breaks

The decision rule is short:

  • Use cosine similarity when direction should dominate and length is noise — the usual case for text embeddings.
  • Use dot product when length carries real signal, such as popularity or frequency, and you want that signal in the score.

One simplification is worth memorizing: if you normalize every vector to unit length, the dot product and cosine similarity produce the same ranking. Many vector indexes exploit this by storing normalized vectors and using a fast inner-product search.

Failure modes to watch:

  • Mixing embedding models. Query and documents must come from the same model.
  • Comparing across dimensions. Different nn means no shared space.
  • Unnormalized vectors in a dot-product index. Length silently dominates the ranking.
  • Treating a score as a probability. It is not one, and it does not sum to anything meaningful across candidates.

Common mistake: Setting a hard similarity threshold and assuming everything above it is relevant. Calibrate the threshold on your own data, or skip the threshold and rely on top-kk ranking plus a re-ranking step.

That last point is the practical bridge. Retrieval is a coarse filter. The usual way to recover precision is to retrieve more candidates than you need, then re-rank them with a stronger, slower model. The math in this article gets you a good shortlist; it does not get you a final answer.

What to Do Next

Compute the score. Rank the candidates. Then judge the ranking, not the number.

Here is a concrete exercise that will teach you more than another article. Start with the toy vectors from this article: q=[1,2,3]q = [1, 2, 3], d1=[2,1,4]d_1 = [2, 1, 4], d2=[3,3,3]d_2 = [3, 3, 3], d3=[−1,0,1]d_3 = [-1, 0, 1]. Compute both dot product and cosine similarity for each pair by hand. Confirm the ranking flip you saw in the table. Then change one component of d3d_3 and watch how the ranking shifts.

Once the arithmetic feels mechanical, embed three short sentences with a single model — one that clearly matches a query, one loosely related, one unrelated — and compare the model's cosine scores to your hand calculations. If the unrelated sentence scores surprisingly high, you have just seen the narrow-band problem in your own data, and you will never read a similarity score the same way again.

From here, the next design problem is ranking quality: how many candidates to retrieve, how to re-rank them, and how to measure whether the top slot is actually the one you want.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A retrieval index normalizes the query and every document vector to unit length. What does the article predict about rankings by dot product and cosine similarity?
Question 1 of 2Scenario Interpretation

Focus: Predict the relationship between dot-product and cosine rankings when all vectors are normalized to unit length.

A system ranks a passage first for a query. What is the appropriate next interpretation according to the article?
Question 2 of 2Comparison Reasoning

Focus: Distinguish geometric retrieval ranking from deciding whether a candidate answers correctly and sufficiently.

References

  1. Measuring similarity from embeddings  |  Machine Learning  |  Google for Developersdevelopers.google.com
  2. intfloat/multilingual-e5-large · cosine similarity very high for any pair of sentenceshuggingface.co
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.