Embedding Similarity Math: Dot Products, Cosine, and Retrieval Ranking
A retrieval system returns a top result with a similarity score of 0.83. Most people read that number as "83% relevant" or "83% likely to be correct." It…

Key topics
A retrieval system returns a top result with a similarity score of 0.83. Most people read that number as "83% relevant" or "83% likely to be correct." It is neither. It is a geometric measurement between two lists of numbers, and once you can compute it by hand, you stop trusting the number and start trusting the ranking.
This article assumes you already know that embeddings turn meaning into vectors. Here we work on the arithmetic of comparing them: how the score is built, how a similarity measure turns a pile of candidate vectors into an ordered list, and why a high score is evidence of similarity, not proof of truth.
What the Similarity Score Actually Measures
An embedding places a piece of text as a point in a high-dimensional space. Comparing two embeddings means comparing two arrows: how aligned are they, and how far apart do their tips sit?
A similarity score answers one narrow question: how aligned are these two vectors? It does not answer "is this document true," "is this the answer," or "should the model trust this." Those are separate judgments that happen after retrieval, not inside the score.
To keep every calculation visible, we will use a tiny 3-dimensional vector set as a stand-in for real embeddings, which often have hundreds or thousands of dimensions. The arithmetic is identical; only the number of terms changes.
Three measures show up constantly in retrieval work, and they are related but not interchangeable:
- Dot product — the sum of component-wise products.
- Cosine similarity — the dot product with vector length divided out.
- Euclidean distance — the straight-line distance between vector tips.
We will derive the first two and use the third as a reference point.
Notation and Assumptions Before We Calculate
A vector is an ordered list of numbers. We write a vector with components , where is the dimension — the number of numbers in the list.
The dot product of two vectors multiplies matching components and sums the results:
The magnitude (length) of a vector is the square root of the vector dotted with itself:
Three assumptions matter for everything that follows:
- Vectors are real-valued and non-zero.
- The query and the documents were embedded by the same model.
- Both sides share the same dimension .
Warning: Comparing vectors from different models, or vectors with different dimensions, is meaningless. The numbers do not share a coordinate system, so any score you compute is noise dressed as arithmetic.
Deriving the Dot Product Step by Step
Let's compute a dot product on real numbers. Suppose our query vector is:
And one candidate document vector is:
Multiply matching components and add:
That is the whole calculation. A large positive dot product means the vectors point broadly in the same direction and/or are long. Both effects are baked into the same number.
Here is the trap. Take a document vector that is simply longer, but pointed in the same direction:
The dot product jumped from 16 to 32, but is just scaled by 2. The score doubled because the vector got longer, not because it got more relevant. This is why raw dot product is sensitive to magnitude, and why some systems use it deliberately: when length carries a real signal — popularity, frequency, recency — you may want that signal in the score. When length is noise, you do not.
Knowledge check
Check your understanding
Answer this question before you continue.
Deriving Cosine Similarity from the Dot Product
The dot product has a geometric identity behind it:
where is the angle between the two vectors. Rearranging isolates the cosine:
That expression is cosine similarity. It is the dot product with both lengths divided out, so only direction survives.
Let's compute it for and . We already have . Now the magnitudes:
Divide:
The result is bounded. Cosine similarity runs from -1 (opposite directions) through 0 (orthogonal, no shared direction) to 1 (identical direction). Because length is divided out, a long document and a short document can score identically if they point the same way — which is usually what you want for text embeddings, where document length is not a signal of relevance.
Knowledge check
Check your understanding
Answer this question before you continue.
Worked Example: Ranking Three Documents Against One Query
Now let's turn the formula into a nearest neighbor retrieval ranking. Keep the query:
And add three candidate documents:
Cosine similarity for we already computed: .
Cosine similarity for :
Cosine similarity for :
Sort descending: (0.933), (0.926), (0.378). The retrieval rule is plain: the highest score wins the top slot. Take the top and hand them to the next stage.
Now re-run the same comparison with raw dot product:
| Document | Dot product | Cosine similarity |
|---|---|---|
| 16 | 0.933 | |
| 18 | 0.926 | |
| 2 | 0.378 |
The ranking flips at the top. Dot product puts first because it is longer; cosine puts first because it points more directly at the query. Neither is "wrong." The measure you choose is a design decision about what you want the score to reward.
Knowledge check
Check your understanding
Answer this question before you continue.
Why High Scores Are Not Proof of Relevance
Here is the misconception that causes the most damage: treating the score as an absolute measure of relevance or truth.
Many embedding models compress scores into a narrow high band. In practice, unrelated sentence pairs from some models can still land above 0.7, because the model was trained to push relevant pairs higher than irrelevant ones — not to spread scores across the full -1 to 1 range. A 0.75 can be a weak match in one model and a strong one in another. The absolute number is not portable across models.
Note: What matters is the relative order of scores, not their absolute values. An embedding model is useful as long as relevant pairs score higher than irrelevant pairs, even when the gap is small.
That gives you a practical rule: use scores to form a model-specific candidate ordering, then evaluate whether that ordering serves your task. Thresholds must be calibrated per model and per dataset, usually by inspecting the score distribution on your own data rather than borrowing a cutoff from a blog post.
The score also cannot answer three separate questions:
- Is this relevant? The score is evidence, not a verdict.
- Is this correct? Similarity is not truth. A confidently wrong passage can sit close to the query.
- Is this sufficient to answer the user? Retrieval finds candidates; it does not decide whether they cover the question.
Knowledge check
Check your understanding
Answer this question before you continue.
Choosing a Measure and Knowing When It Breaks
The decision rule is short:
- Use cosine similarity when direction should dominate and length is noise — the usual case for text embeddings.
- Use dot product when length carries real signal, such as popularity or frequency, and you want that signal in the score.
One simplification is worth memorizing: if you normalize every vector to unit length, the dot product and cosine similarity produce the same ranking. Many vector indexes exploit this by storing normalized vectors and using a fast inner-product search.
Failure modes to watch:
- Mixing embedding models. Query and documents must come from the same model.
- Comparing across dimensions. Different means no shared space.
- Unnormalized vectors in a dot-product index. Length silently dominates the ranking.
- Treating a score as a probability. It is not one, and it does not sum to anything meaningful across candidates.
Common mistake: Setting a hard similarity threshold and assuming everything above it is relevant. Calibrate the threshold on your own data, or skip the threshold and rely on top- ranking plus a re-ranking step.
That last point is the practical bridge. Retrieval is a coarse filter. The usual way to recover precision is to retrieve more candidates than you need, then re-rank them with a stronger, slower model. The math in this article gets you a good shortlist; it does not get you a final answer.
What to Do Next
Compute the score. Rank the candidates. Then judge the ranking, not the number.
Here is a concrete exercise that will teach you more than another article. Start with the toy vectors from this article: , , , . Compute both dot product and cosine similarity for each pair by hand. Confirm the ranking flip you saw in the table. Then change one component of and watch how the ranking shifts.
Once the arithmetic feels mechanical, embed three short sentences with a single model — one that clearly matches a query, one loosely related, one unrelated — and compare the model's cosine scores to your hand calculations. If the unrelated sentence scores surprisingly high, you have just seen the narrow-band problem in your own data, and you will never read a similarity score the same way again.
From here, the next design problem is ranking quality: how many candidates to retrieve, how to re-rank them, and how to measure whether the top slot is actually the one you want.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Want a more structured LLMOps path?
Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.
Large Language Models Starter Pack
A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.
- 227-page Illustrated PDF edition
- 12 guided LLM engineering chapters
- Visual concept diagrams
- Self-assessment quizzes
- Bonus deep-dive sections
- Prompt design, structured output, context windows & RAG pipelines
- Agents, tool calling, prompt injection, evaluation & application lifecycles
Coming soon


