Skip to content
intermediate

Attention and Token Representations: A Worked Mathematical Walkthrough

Most people can recite the attention formula. Far fewer can compute it. That gap is worth closing, because the formula is not a wall of symbols — it is a…

Published 2026-10-03Updated 2026-10-0410 min read
Appetizing grilled squid served with garnishes on a stylish plate with a spoon.
Appetizing grilled squid served with garnishes on a stylish plate with a spoon. Photo by ธันยกร ไกรสร on Pexels.

Most people can recite the attention formula. Far fewer can compute it. That gap is worth closing, because the formula is not a wall of symbols — it is a short sequence of tensor operations you can trace with a pencil.

You already know the shape of the story: text becomes tokens, tokens become embeddings, and transformers stack attention layers to turn those embeddings into predictions. What we are going to do here is shrink the problem until every number fits on one line, then walk the whole mechanism end to end. Three tokens. Two dimensions. One head. No bias terms, no dropout, one layer. By the end you should be able to say what shape every intermediate tensor has, why the scaling factor exists, and why the mask has to be applied before softmax rather than after.

Why the Formula Feels Opaque

The usual failure mode is treating Q, K, and V as three mysterious learned objects that appear from nowhere. They are not. They are three different projections of the same token vectors. Same input, three different weight matrices, three different jobs.

The real question attention answers is simple: for this token, which earlier tokens carry information I should pull in, and how much of each? Everything else — the dot products, the scaling, the softmax, the weighted sum — is machinery for answering that question with numbers.

A three-token example is not a toy dodge. It is the fastest way to see the mechanism, because every value stays checkable by hand. If you can trace three tokens, you can trace three thousand; the arithmetic just gets longer.

Notation and Shapes Before Any Numbers

Before we compute anything, let us fix the notation. Skipping this step is why the formula feels like a wall.

Let XX be the input matrix with shape [ntokens×dmodel][n_{\text{tokens}} \times d_{\text{model}}] — one row per token. In our example, ntokens=3n_{\text{tokens}} = 3 and dmodel=2d_{\text{model}} = 2.

We have three learned projection matrices:

  • WQW_Q with shape [dmodel×dk][d_{\text{model}} \times d_k]
  • WKW_K with shape [dmodel×dk][d_{\text{model}} \times d_k]
  • WVW_V with shape [dmodel×dk][d_{\text{model}} \times d_k]

And we compute:

Q=XWQ,K=XWK,V=XWVQ = XW_Q, \quad K = XW_K, \quad V = XW_V

Each has shape [ntokens×dk][n_{\text{tokens}} \times d_k]. In our example, dk=2d_k = 2.

Why three separate projections? If we used the raw token vectors for both queries and keys, the score for a token against itself would tend to be large relative to its scores against other tokens, because a vector's dot product with itself is its squared length. That does not make self-attention inevitable, but it biases the matching rule toward self-similarity and leaves the model with less freedom to route information elsewhere. The projections give the model learned spaces where a query can encode what a token is looking for and a key can encode what a token advertises, so the two roles can be tuned independently.

The score matrix is:

S=QKTS = QK^T

This has shape [ntokens×ntokens][n_{\text{tokens}} \times n_{\text{tokens}}]. Each row is one query scored against every key. Entry SijS_{ij} measures how much token ii cares about token jj.

The final output has shape [ntokens×dk][n_{\text{tokens}} \times d_k]. Each row is a context-enriched representation of one token.

Note: We are working with a single attention head, no bias terms, no dropout, and one layer. Real models add all of those. The mechanism below is the same; the bookkeeping just grows.

Knowledge check

Check your understanding

Answer this question before you continue.

If the input has shape 3 × 2 and each projection matrix has shape 2 × 2, what are the shapes of Q and QKᵀ, respectively?
Single Choice

Focus: Determine the shapes of projected token representations and the attention score matrix from the input and projection dimensions.

A Three-Token Worked Example: Scores and Scaling

Let us set up three tokens with 2-dimensional embeddings. Keep the numbers small so every step is checkable.

X=[100111]X = \begin{bmatrix} 1 & 0 \\ 0 & 1 \\ 1 & 1 \end{bmatrix}

Token A is (1,0)(1, 0), token B is (0,1)(0, 1), token C is (1,1)(1, 1).

Now define small projection matrices:

WQ=[1001],WK=[1001],WV=[1001]W_Q = \begin{bmatrix} 1 & 0 \\ 0 & 1 \end{bmatrix}, \quad W_K = \begin{bmatrix} 1 & 0 \\ 0 & 1 \end{bmatrix}, \quad W_V = \begin{bmatrix} 1 & 0 \\ 0 & 1 \end{bmatrix}

Using identity matrices keeps the arithmetic transparent. In a real model these would be learned values.

Since all three projections are the identity, Q=K=V=XQ = K = V = X:

TokenQuery q\mathbf{q}Key k\mathbf{k}Value v\mathbf{v}
A(1,0)(1, 0)(1,0)(1, 0)(1,0)(1, 0)
B(0,1)(0, 1)(0,1)(0, 1)(0,1)(0, 1)
C(1,1)(1, 1)(1,1)(1, 1)(1,1)(1, 1)

Now compute raw dot-product scores. For token A's query (1,0)(1, 0) against each key:

  • Against A's key (1,0)(1, 0): 1⋅1+0⋅0=11 \cdot 1 + 0 \cdot 0 = 1
  • Against B's key (0,1)(0, 1): 1⋅0+0⋅1=01 \cdot 0 + 0 \cdot 1 = 0
  • Against C's key (1,1)(1, 1): 1⋅1+0⋅1=11 \cdot 1 + 0 \cdot 1 = 1

Doing this for all queries gives the raw score matrix:

Sraw=[101011112]S_{\text{raw}} = \begin{bmatrix} 1 & 0 & 1 \\ 0 & 1 & 1 \\ 1 & 1 & 2 \end{bmatrix}

Now the scaling. We divide by dk=2≈1.414\sqrt{d_k} = \sqrt{2} \approx 1.414:

Sscaled=[0.7100.7100.710.710.710.711.41]S_{\text{scaled}} = \begin{bmatrix} 0.71 & 0 & 0.71 \\ 0 & 0.71 & 0.71 \\ 0.71 & 0.71 & 1.41 \end{bmatrix}

Why divide by dk\sqrt{d_k}? As dkd_k grows, dot products grow in magnitude — each additional dimension adds another term to the sum. Large scores push softmax into saturation, where one weight approaches 1.0 and the rest approach 0.0. When that happens, gradients shrink toward zero and learning stalls. The scale factor keeps the distribution in a range where softmax can still produce a useful spread of weights.

Read the scaled score matrix as a table of "how much does token ii care about token jj." Token C, with the largest self-score, currently cares most about itself.

Knowledge check

Check your understanding

Answer this question before you continue.

Why does scaled dot-product attention divide scores by √dₖ before applying softmax?
Misconception Check

Focus: Explain how scaling dot-product scores affects softmax behavior as the key/query dimension grows.

Softmax and the Weighted Sum of Values

Softmax is applied row-wise, not to the whole matrix. Each row becomes a probability distribution that sums to 1.

For token A's row [0.71,0,0.71][0.71, 0, 0.71]:

e0.71≈2.03,e0=1.00,e0.71≈2.03e^{0.71} \approx 2.03, \quad e^{0} = 1.00, \quad e^{0.71} \approx 2.03

Sum: 2.03+1.00+2.03=5.062.03 + 1.00 + 2.03 = 5.06

Weights: [2.03/5.06,  1.00/5.06,  2.03/5.06]≈[0.40,  0.20,  0.40][2.03/5.06, \; 1.00/5.06, \; 2.03/5.06] \approx [0.40, \; 0.20, \; 0.40]

For token C's row [0.71,0.71,1.41][0.71, 0.71, 1.41]:

e0.71≈2.03,e0.71≈2.03,e1.41≈4.10e^{0.71} \approx 2.03, \quad e^{0.71} \approx 2.03, \quad e^{1.41} \approx 4.10

Sum: 2.03+2.03+4.10=8.162.03 + 2.03 + 4.10 = 8.16

Weights: [2.03/8.16,  2.03/8.16,  4.10/8.16]≈[0.25,  0.25,  0.50][2.03/8.16, \; 2.03/8.16, \; 4.10/8.16] \approx [0.25, \; 0.25, \; 0.50]

The full attention weight matrix:

A=[0.400.200.400.200.400.400.250.250.50]A = \begin{bmatrix} 0.40 & 0.20 & 0.40 \\ 0.20 & 0.40 & 0.40 \\ 0.25 & 0.25 & 0.50 \end{bmatrix}

Each row sums to 1. Now multiply by VV to get the output. For token A:

outputA=0.40⋅(1,0)+0.20⋅(0,1)+0.40⋅(1,1)=(0.80,  0.60)\text{output}_A = 0.40 \cdot (1,0) + 0.20 \cdot (0,1) + 0.40 \cdot (1,1) = (0.80, \; 0.60)

For token C:

outputC=0.25⋅(1,0)+0.25⋅(0,1)+0.50⋅(1,1)=(0.75,  0.75)\text{output}_C = 0.25 \cdot (1,0) + 0.25 \cdot (0,1) + 0.50 \cdot (1,1) = (0.75, \; 0.75)

The output row is a weighted average of all value vectors. Token A's original representation was (1,0)(1, 0); its output is (0.80,0.60)(0.80, 0.60) — enriched with information routed from tokens B and C. Token C's original was (1,1)(1, 1); its output is (0.75,0.75)(0.75, 0.75), pulled slightly toward the other tokens by the weight it placed on them.

Tip: Sanity-check your work by looking at a row that puts most weight on one key. That output should be close to that key's value vector. If it is not, you have an arithmetic slip.

Knowledge check

Check your understanding

Answer this question before you continue.

Using token A’s attention weights [0.40, 0.20, 0.40] and values (1, 0), (0, 1), and (1, 1), what is its output?
Output Prediction

Focus: Compute one token's context representation as the weighted sum of the value vectors.

Causal Masking: Why the Future Must Stay Hidden

Two side-by-side calculations for token A. Masking future positions before softmax turns scores into weights [1.00, 0, 0], which sum to 1. Zeroing future weights after softmax leaves [0.40, 0, 0], which sum to 0.40.
Applying the causal mask before softmax keeps each attention row normalized; masking afterward does not.

When a model predicts token tt, the representation for position ii must not depend on positions after ii. Otherwise the model could peek at the answer during training.

The mask is an upper-triangular pattern of −∞-\infty added to the score matrix before softmax:

M=[0−∞−∞00−∞000]M = \begin{bmatrix} 0 & -\infty & -\infty \\ 0 & 0 & -\infty \\ 0 & 0 & 0 \end{bmatrix}

Adding this to the scaled scores and then applying softmax gives:

Amasked=[1.00000.500.5000.250.250.50]A_{\text{masked}} = \begin{bmatrix} 1.00 & 0 & 0 \\ 0.50 & 0.50 & 0 \\ 0.25 & 0.25 & 0.50 \end{bmatrix}

Token A can only attend to itself. Token B attends to A and B. Token C attends to all three — and its row is unchanged from the unmasked matrix, because C is the final position and has no future keys to exclude.

Why before softmax and not after? Because e−∞=0e^{-\infty} = 0. Softmax of −∞-\infty yields exactly zero weight with no special-case code. If you masked after softmax, you would zero out entries that had already been normalized — and the remaining weights would no longer sum to 1.

The diagonal is never masked. A token always attends to itself.

Common mistake: Masking after softmax. The resulting row no longer sums to 1, and your weighted average is no longer an average.

Encoder-style models like BERT use bidirectional attention — no causal mask — because they see the full input at once. Padding masks (which hide filler tokens) combine with causal masks via logical OR: any position masked by either constraint is masked in the combined mask.

Knowledge check

Check your understanding

Answer this question before you continue.

With tokens ordered A, B, C, which positions may token B attend to under the article’s causal mask?
Scenario Interpretation

Focus: Identify which token positions remain available to a query under a causal attention mask.

What This Predicts About Real Models

The score matrix is [n×n][n \times n], so attention cost grows quadratically with sequence length. Double the tokens, quadruple the score matrix. This is the mechanical reason long context is expensive — not a policy choice, but arithmetic.

Multi-head attention splits dmodeld_{\text{model}} across heads. Each head runs this same computation on a slice, then the outputs are concatenated and mixed with another learned matrix WOW_O. Different heads can learn to attend to different patterns.

Key/value caching works because of causal masking. Earlier tokens' keys and values do not change when new tokens are appended — they cannot see the future. So during generation, only the newest query needs computing; the keys and values from prior positions are reused.

Where the analogy stops: real models add residual connections, layer normalization, feed-forward blocks, and many stacked layers. One attention head is a component, not the whole model.

Common Mistakes When Tracing Attention by Hand

  • Transposing the wrong matrix. QKTQK^T is [n×dk][n \times d_k] times [dk×n][d_k \times n], producing [n×n][n \times n]. Not QQ times KK.
  • Applying softmax to the entire matrix. Softmax is row-wise. Each row is an independent distribution.
  • Forgetting the scale factor. Then wondering why one weight is nearly 1.0 and the rest are nearly 0.
  • Assuming self-attention is symmetric. It is not. QQ and KK use different projections, so token ii's attention to token jj need not equal token jj's attention to token ii.
  • Masking after softmax. Produces zeros that no longer sum to 1.
  • Treating the output as a new token embedding. It is a context-mixed representation that still flows through the rest of the block — layer norm, feed-forward, residual add.

Your Next Step: Change One Number

Reading a derivation is not the same as owning it. Here is the drill that converts one into the other.

Take the worked example and change token B's embedding from (0,1)(0, 1) to (1,1)(1, 1). Before you recompute anything, predict which attention weights will shift. Then do the arithmetic and check.

Next, increase dkd_k to 4 or 8 and watch how much the raw scores grow. The scaling factor's purpose becomes visible when you see what happens without it.

Finally, remove the causal mask and recompute token A's output. Watch it change. That change is the mask doing real work.

From here, the natural next concepts are multi-head attention and the feed-forward block — or, if you want the practical side, how context length and cost follow directly from that [n×n][n \times n] score matrix. Either path builds on what you just traced by hand.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

Why can a causal language model reuse earlier tokens’ keys and values when generating a new token?
Question 1 of 2Comparison Reasoning

Focus: Relate causal masking to reusing earlier keys and values during autoregressive generation.

Which sequence correctly describes the attention computation taught in the article?
Question 2 of 2Comparison Reasoning

Focus: Trace the order of operations that converts token representations into context-mixed outputs.

References

  1. Scaled Dot-Product Attention: The Core Transformer Mechanism - Interactivembrenndoerfer.com
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.

A white humanoid toy robot standing on a reflective black surface in a studio setting with a blue and pink gradient background.
beginner
8 min read

A Brief History of LLMs

Large language models look like overnight magic, but they are the visible tip of decades of compounding engineering. If you want to use, trust, and build…

Read tutorial