Skip to content
intermediate

Temperature and Token Sampling: The Probability Math

Temperature does not add randomness to a model. It reshapes a probability distribution the model already produced — and once you see the arithmetic, you…

Published 2026-10-03Updated 2026-10-0411 min read
Close-up of a large pot filled with black dye used in traditional incense stick production indoors.
Close-up of a large pot filled with black dye used in traditional incense stick production indoors. Photo by HONG SON on Pexels.

Temperature does not add randomness to a model. It reshapes a probability distribution the model already produced — and once you see the arithmetic, you stop treating the dial as a creativity knob.

You already know the intuition: low temperature means predictable output, high temperature means varied output. That sentence is true and almost useless. It does not tell you why a logit gap of 2 becomes a probability ratio of 54.6 at one setting and 2.7 at another, or why lowering temperature can make a wrong answer more consistent rather than more correct.

To reason about that, we need the actual mechanism. The good news: it is one equation, and a four-token example carries the whole story.

From Logits to Probabilities: The Baseline Softmax

Before temperature enters, the model produces a vector of logits — raw, unnormalized scores, one per token in the vocabulary. Call them z1,z2,…,zVz_1, z_2, \dots, z_V, where VV is the vocabulary size. Logits are not probabilities. They can be negative, they do not sum to anything meaningful, and their absolute values carry no interpretation on their own.

To turn logits into a probability distribution, we apply the softmax function:

pi=exp⁡(zi)∑j=1Vexp⁡(zj)p_i = \frac{\exp(z_i)}{\sum_{j=1}^{V} \exp(z_j)}

Two properties make the exponential the right choice here. First, exp⁡(z)\exp(z) is always positive, so every token gets a non-negative probability. Second, the exponential is strictly increasing, so if za>zbz_a > z_b, then pa>pbp_a > p_b. Softmax preserves the ranking of the logits — it only changes their scale.

Let's pin this down with a running example. Suppose the model has narrowed the next token to four candidates, with these logits:

TokenLogit
Paris5.0
London4.0
Berlin3.0
banana1.0

At the baseline (no temperature adjustment), we compute exp⁡\exp of each logit: e5.0≈148.4e^{5.0} \approx 148.4, e4.0≈54.6e^{4.0} \approx 54.6, e3.0≈20.1e^{3.0} \approx 20.1, e1.0≈2.7e^{1.0} \approx 2.7. The denominator is their sum: 148.4+54.6+20.1+2.7=225.8148.4 + 54.6 + 20.1 + 2.7 = 225.8.

Dividing through gives the baseline distribution: Paris ≈ 0.657, London ≈ 0.242, Berlin ≈ 0.089, banana ≈ 0.012. That is the distribution the model actually produced. Everything from here is about how temperature transforms it.

Knowledge check

Check your understanding

Answer this question before you continue.

If one token has a larger logit than another, what does applying baseline softmax guarantee?
Single Choice

Focus: Infer what baseline softmax preserves about the ordering of token logits.

The Temperature Formula, Term by Term

The temperature-scaled softmax is:

pi(T)=exp⁡(zi/T)∑j=1Vexp⁡(zj/T)p_i(T) = \frac{\exp(z_i / T)}{\sum_{j=1}^{V} \exp(z_j / T)}

Read the numerator carefully. Temperature divides the logits, not the probabilities. This is the single most common point of confusion. If you think temperature multiplies the output probabilities, you will predict the wrong behavior at every setting.

Because the division happens before the exponential, temperature controls how much the logit gaps matter. To see this, take any two tokens aa and bb and form their probability ratio:

pa(T)pb(T)=exp⁡(za/T)exp⁡(zb/T)=exp⁡(za−zbT)\frac{p_a(T)}{p_b(T)} = \frac{\exp(z_a / T)}{\exp(z_b / T)} = \exp\left(\frac{z_a - z_b}{T}\right)

The denominator cancels. What remains is a clean statement: temperature only affects the relative gap between logits. It never changes which token ranks highest. If Paris has the largest logit at T=1T = 1, it has the largest logit at every positive temperature. Temperature changes the spread, not the order.

A few boundary cases follow directly from that ratio:

  • T=1T = 1 is the identity case. The formula reduces to the baseline softmax.
  • T→0T \to 0 makes the ratio explode for the top token. The distribution collapses toward the argmax — greedy selection.
  • T→∞T \to \infty makes every ratio approach 1. The distribution flattens toward uniform.

Note: T=0T = 0 is a limit, not a value. You cannot divide by zero, so implementations handle it as a special case — usually by switching to greedy decoding or clamping to a small positive number. When you set temperature to 0 in an API, you are asking for argmax, not for the formula to be evaluated at zero.

One more thing worth flagging early: providers do not apply temperature at the same point in the pipeline. Some scale logits before truncation (top-k or top-p), some after. The same numeric temperature is not portable across APIs, and I have watched that assumption cause real debugging pain.

Knowledge check

Check your understanding

Answer this question before you continue.

Paris and Berlin have logits 5 and 3. At temperature 0.5, approximately how many times as likely is Paris as Berlin?
Output Prediction

Focus: Use the temperature ratio formula to calculate relative likelihood from a logit gap.

Worked Example: One Distribution, Three Temperatures

Three side-by-side probability columns compare temperatures 0.2, 1.0, and 1.5 for Paris, London, Berlin, and banana. Paris has the longest bar at every temperature; as temperature rises, its bar shortens while the other tokens’ bars grow.
The same logits produce a sharply concentrated distribution at low temperature and a flatter one at higher temperature, without changing token order.

Let's run the four-token example through three temperatures. I will show T=0.2T = 0.2 in full, then tabulate the rest.

At T=0.2T = 0.2, we divide each logit by 0.2, which is the same as multiplying by 5. The scaled logits become 25.0, 20.0, 15.0, and 5.0. Now exponentiate: e25≈7.2×1010e^{25} \approx 7.2 \times 10^{10}, e20≈4.85×108e^{20} \approx 4.85 \times 10^{8}, e15≈3.27×106e^{15} \approx 3.27 \times 10^{6}, e5≈148.4e^{5} \approx 148.4.

The denominator is dominated by the first term: roughly 7.25×10107.25 \times 10^{10}. Dividing through gives Paris ≈ 0.993, London ≈ 0.0067, Berlin ≈ 0.000045, banana ≈ 0.000000002.

Here is the full comparison:

TokenLogitP(T=0.2)P(T=1.0)P(T=1.5)
Paris5.00.9930.6570.553
London4.00.00670.2420.284
Berlin3.00.0000450.0890.146
banana1.0~0.0000000020.0120.038

Look at what happened to banana. At T=0.2T = 0.2 it is effectively unreachable. At T=1.5T = 1.5 it has a 3.8% chance — small, but real. Same logits. Different world.

The ratio formula explains the shift. Take Paris and Berlin, a logit gap of 2. At T=0.5T = 0.5, their probability ratio is e4≈54.6e^{4} \approx 54.6: Paris is roughly 55 times more likely. At T=2.0T = 2.0, the same gap gives e1≈2.7e^{1} \approx 2.7: Paris is now less than three times more likely. Temperature did not touch the logits. It changed how much the gap mattered.

This is the formal version of what people loosely call "creativity." Lower temperature concentrates probability mass on the top tokens; higher temperature spreads it across the tail. The technical name for that concentration is entropy, and temperature is a direct lever on it — not a linear one, but a monotonic one.

Knowledge check

Check your understanding

Answer this question before you continue.

For the article's four-token logits, how does Paris's probability compare at temperature 0.2 and 1.5?
Comparison Reasoning

Focus: Compare how the article's example distribution changes as temperature increases.

What Sampling Actually Does With Those Probabilities

Here is where a lot of confusion lives: the model outputs a distribution, and the sampler draws one token from it. These are two separate steps, and temperature only touches the first.

Greedy decoding ignores the distribution entirely. It takes the argmax. If your pipeline uses greedy, temperature is irrelevant — you can set it to anything and the output will not change.

Stochastic sampling draws from the distribution. This is where temperature matters, and it is also where variation comes from. The same prompt at the same temperature can produce different tokens on different runs. That is not model inconsistency; it is the sampler doing its job.

Truncation methods sit alongside temperature and reshape the candidate set:

  • Top-k keeps only the kk highest-probability tokens and renormalizes.
  • Top-p (nucleus sampling) keeps the smallest set of tokens whose cumulative probability exceeds pp.

The interaction between temperature and truncation depends on order. If temperature is applied first, it changes which tokens survive the truncation cut. If truncation is applied first, temperature only redistributes mass within the surviving set. Providers differ here, and the difference is observable.

Common mistake: Assuming a fixed seed guarantees reproducibility. On some platforms it does, but not when the model is genuinely uncertain between near-equal candidates. I have seen runs with a fixed seed and near-zero temperature still diverge on questions the model was "wrestling with." Determinism is a property of the whole pipeline, not a parameter you can set and forget.

Knowledge check

Check your understanding

Answer this question before you continue.

A generation pipeline always chooses the argmax token using greedy decoding. What should you expect if you change only its temperature setting?
Scenario Interpretation

Focus: Distinguish greedy decoding from stochastic sampling when reasoning about temperature.

Sampling Variation Is Not Correctness

This is the section I would underline twice if I were handing you the article on paper.

Temperature changes which plausible token gets picked. It does not change which token is correct. Those are different questions, and conflating them is the most expensive mistake in this territory.

Consider a factual question where the model is confidently wrong. Its distribution puts high probability on the wrong answer. Lowering temperature makes that wrong answer more consistent — you get the same error every time, which feels like reliability but is just reproducibility. Raising temperature gives you a chance of sampling a different token, but only if the correct answer was already somewhere in the distribution with non-trivial mass. If it was not, no temperature setting will surface it.

The mirror failure is creative work. At low temperature, the model keeps picking the same high-probability token, and the output goes flat and repetitive. Raising temperature adds variety — and also adds incoherence risk, because you are now sampling from the tail where the model's judgment is weakest.

The decision rule I use:

Temperature controls the shape of the output distribution, not the content of the answer. If the model has the right answer somewhere in its distribution, temperature helps you sample it. If it does not, temperature is the wrong dial.

Correctness comes from the prompt, the context, the model's knowledge, and — when you need external facts — retrieval. Temperature is downstream of all of those.

When to Reach for Temperature — and When Not To

Rough guidance, with the caveat that provider defaults differ and the numbers are not portable:

GoalTemperature rangeWhy
Structured extraction, classification, deterministic pipelines0 – 0.3You want the highest-probability token almost every time
Code generation where one correct form exists0 – 0.3Same reasoning; variation is usually error
General chat, drafting, summarization0.7 – 1.0Some variation is acceptable and often useful
Brainstorming, creative writing, generating diverse candidates1.0 – 1.5You want the tail to have real probability mass

The more important guidance is when not to reach for temperature at all. If the problem is a vague prompt, missing context, or a model that simply does not know the answer, changing temperature is a distraction. You will spend an afternoon tuning a dial that cannot fix the actual bottleneck. I have done this. It feels productive. It is not.

Also remember that temperature interacts with top-p and top-k. Tuning one in isolation can mislead you, because the other two may be doing most of the work.

What the Math Cannot Tell You

The formula describes how temperature transforms a given distribution. It says nothing about whether that distribution is well-calibrated, whether the model's logits reflect reality, or whether the top token is actually the right answer. Those are separate questions, and the math is silent on all of them.

Provider implementations vary in ways that matter. Some clamp temperature to a maximum. Some apply frequency or presence penalties before truncation, some after. Some treat T=0T = 0 as greedy, others clamp to a tiny positive value. The same number is not the same behavior across APIs, and treating it as portable is a reliable way to ship a bug.

Sampling settings are also one lever among many. Prompt quality, context, model choice, and retrieval quality usually dominate the outcome. If you are debugging a bad response, I would check those first and reach for temperature last.

One open question worth watching: reasoning-style models may not expose or honor the temperature parameter the same way standard models do. The behavior is evolving, and the safe assumption is that the parameter's effect is model-specific until you have tested it.

The One Question That Replaces the Dial

Before you touch temperature, ask a single question: is the problem distribution shape or distribution content?

If the model has the right answer somewhere in its distribution and you just need to sample it more reliably, temperature is your lever. Lower it for consistency, raise it for variety, and watch the ratio formula predict what happens.

If the model does not have the right answer in its distribution, temperature is the wrong dial. Fix the prompt, add context, retrieve better evidence, or choose a different model. The math will not save you from a distribution that never contained the answer.

The best way to internalize this is to run it. Take the four logits from this article — [5.0, 4.0, 3.0, 1.0] — and compute the softmax at three temperatures in a notebook. Change one value. Watch the probabilities shift. The formula stops being abstract the moment you have executed it and seen the numbers move.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A model's distribution strongly favors a wrong factual answer, while the correct answer has negligible probability. What is the most accurate expectation when temperature is lowered?
Question 1 of 2Misconception Check

Focus: Explain why lowering temperature cannot make an absent or low-probability correct answer reliably appear.

A model repeatedly misses a fact because the prompt lacks necessary context. According to the article's decision rule, what is the best next step?
Question 2 of 2Scenario Interpretation

Focus: Choose a response to a generation problem based on whether it concerns distribution shape or missing content.

References

  1. Decoding Strategies in Large Language Modelshuggingface.co
  2. Enhancing LLM's Exploration via Attention Temperature ...aclanthology.org
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.