
Estimate the Memory a Local LLM Needs
Two people run the same 7B model. One says it needs 4 GB. The other says it needs 9 GB. Neither is lying.
Read tutorialHow a trained model is used at runtime to process context and generate outputs, including decoding and inference resource behavior.
Tagged articles
18 articles in this tag.

Two people run the same 7B model. One says it needs 4 GB. The other says it needs 9 GB. Neither is lying.
Read tutorial
A language model never scores a sentence. It scores one token at a time — and the sentence score is just those pieces multiplied together.
Read tutorial
Large language models are not digital minds. They are probability engines that turn a conversation into a series of next-token guesses. The guesswork is…
Read tutorial
A builder doubles the candidate count, watches the benchmark number climb, ships it, and then discovers that p95 latency and per-request cost moved in ways…
Read tutorial
Frameworks change. The model underneath them doesn't. Learn the mechanism first, and every new tool becomes just another wrapper around something you…
Read tutorial
A single fast response tells you almost nothing about whether your LLM system can handle real load. The request that feels snappy in isolation is often the…
Read tutorial
Throughput climbs, the dashboard looks healthy, and p95 quietly doubles. Here is the arithmetic that explains why.
Read tutorial
You asked for a summary and got a rambling essay. You asked for a story and got three bullet points that stopped mid-thought. You asked a factual question…
Read tutorial
The weights fit. The model loads. Then the second concurrent request arrives, and the server refuses to admit it.
Read tutorial
You turned the temperature down to 0.1 because you wanted the AI to stop making mistakes. It still gave you a wrong answer—just a more confident, more…
Read tutorial
You correct a chatbot's mistake and expect it to remember next time. It won't. You ask the same question twice and get two different answers, which makes…
Read tutorial
When you watch a large language model type out a response, it feels like the system thought through the whole answer before writing a single word. It…
Read tutorial
The output looks wrong, and your hand is already on the temperature slider. Before you drag it, notice what you are actually doing: guessing. The same…
Read tutorial
The model finishes downloading. You type a prompt. The cursor blinks, the fan spins up, and nothing happens for a while. That pause is the whole question:…
Read tutorial
Temperature does not add randomness to a model. It reshapes a probability distribution the model already produced — and once you see the arithmetic, you…
Read tutorial
You know an LLM predicts the next word. The real question is how it decides which word deserves to come next—and the answer lives in a design called the…
Read tutorial
A large language model is a pattern-prediction engine for language: it learns how words and ideas tend to follow one another, then uses that fluency to…
Read tutorial
You ask a chatbot for a quick fact. The answer comes back in clean, confident sentences—specific dates, plausible names, a citation that looks real. It…
Read tutorial