Skip to content
beginner

Estimate the Memory a Local LLM Needs

Two people run the same 7B model. One says it needs 4 GB. The other says it needs 9 GB. Neither is lying.

Published 2026-10-03Updated 2026-10-048 min read
Detailed close-up of blue soap foam showcasing abstract geometric patterns and texture.
Detailed close-up of blue soap foam showcasing abstract geometric patterns and texture. Photo by Antonio Friedemann on Pexels.

Two people run the same 7B model. One says it needs 4 GB. The other says it needs 9 GB. Neither is lying.

That gap is the whole problem. "Model size" feels like one number, so we treat memory as one number too. It isn't. Memory is a sum of parts, and once you can name the parts, you can estimate them in your head before you download a single gigabyte.

Here's the deal: you'll get a formula for the predictable part, a separate allowance for the messy part, and a clear statement of what the estimate cannot tell you.

Why One Memory Number Is Always Wrong

Ask "how much memory does this model need?" and you've asked a question with no single answer. The same model produces different figures because three things change underneath the label:

  • Precision — how many bits each stored parameter occupies.
  • Context length — how much text the model must hold in working memory at once.
  • Runtime — the software loading the model, plus whatever else your machine is doing.

Change any one and the number moves. That's why the 4 GB person and the 9 GB person can both be right about the same weights.

So split the question in two:

  1. Weights — the learned values the model must load. Fixed for a given model and precision, and predictable.
  2. Runtime overhead — the working room the system needs on top. Variable, and workload-dependent.

If you've read our piece on what parameters actually are, you already know parameter count is an incomplete signal for capability. Here it becomes a genuinely useful sizing input — once you pair it with precision. Parameters tell you how much there is to store. Precision tells you how tightly it's packed.

One scope note before the math: this is a capacity screen for one user running one model. It is not a serving model for many concurrent requests. Keep that boundary in mind, because it's exactly where the simple version stops working.

The Two Numbers You Actually Need

Before any formula, let's name the inputs. There are only two.

P — parameter count. The number of learned values in the model. A "7B" model has P ≈ 7,000,000,000. That's it. No mystery.

b — effective weight bits. How many bits each stored parameter occupies at load time. Common values:

  • 32 for full precision
  • 16 for half precision
  • 8 or 4 for common quantized formats

The word effective is doing real work there. Quantization doesn't always deliver exactly b bits per parameter once you count metadata and any layers the runtime keeps at higher precision. Treat b as the label on the box, not a guarantee about every byte inside it.

Where do you find these? The model card or config file gives you P and the storage type. The quantization label in the filename — something like Q4 — gives you b.

One assumption to state up front: every parameter is stored at the same precision. That's a simplification. We'll come back to where it breaks.

Knowledge check

Check your understanding

Answer this question before you continue.

A model is labeled 7B and uses a Q4 quantization format. Which interpretation supplies the two inputs for a rough weight estimate?
Single Choice

Focus: Identify parameter count and effective weight bits as the two inputs to the idealized weight estimate.

Deriving the Weight Estimate

Let's build the formula one step at a time, because each step is just unit conversion.

Step 1: one parameter. A parameter stored at b bits occupies b/8 bytes. The division by 8 is nothing more than bits-to-bytes. One byte holds eight bits.

Step 2: all parameters. Multiply by P:

Mweights≈P×b8 bytesM_{weights} \approx \frac{P \times b}{8} \text{ bytes}

Step 3: convert to gigabytes. Here's a trap. Decimal gigabytes (10⁹ bytes) and binary gibibytes (2³⁰ bytes) differ by about 7%. Same number, different label, different answer. Pick one and stay consistent. I'll use decimal throughout.

Two anchors worth memorizing:

  • Roughly 2 GB per billion parameters at 16-bit.
  • Roughly 0.5 GB per billion at 4-bit.

Those two numbers will carry you through most quick comparisons.

Knowledge check

Check your understanding

Answer this question before you continue.

Using the article's idealized formula, approximately how many decimal gigabytes of weights does a 2-billion-parameter model at 8 bits need, before overhead?
Single Choice

Focus: Apply the bits-to-bytes conversion and parameter count to estimate idealized weight storage.

Worked Example: A 1B Model at 4 Bits

Let's run it end to end.

P = 1,000,000,000 and b = 4:

Mweights≈1,000,000,000×48=500,000,000 bytes≈0.5 GBM_{weights} \approx \frac{1{,}000{,}000{,}000 \times 4}{8} = 500{,}000{,}000 \text{ bytes} \approx 0.5 \text{ GB}

Half a gigabyte of idealized weights. Now the important part — what that 0.5 GB excludes:

  • Quantization metadata such as per-group scales and zero-points.
  • Embedding and output layers, which are often kept at higher precision.
  • Any layers the runtime refuses to quantize.

This is why a real 4-bit download usually lands a bit above the idealized figure. The gap is metadata, not a mistake in your arithmetic. If your math says 0.5 and the file says 0.6, your math is fine.

The practical read: a 1B model at 4 bits is small enough that overhead, not weights, decides whether it fits comfortably. That flips your attention to the right place.

Knowledge check

Check your understanding

Answer this question before you continue.

Your calculation gives about 0.5 GB for a 1B model at 4 bits, but its downloaded file is larger. Which explanation best matches the article?
Scenario Interpretation

Focus: Interpret the 1B-at-4-bit result as an idealized weight estimate that excludes real-world storage additions.

Changing Precision Changes Everything

Same model, different b. The estimate scales linearly, so you can predict the result before computing it.

PrecisionBytes per parameter1B model7B model
32-bit4~4 GB~28 GB
16-bit2~2 GB~14 GB
8-bit1~1 GB~7 GB
4-bit0.5~0.5 GB~3.5 GB

Read the table as idealized weights only. Every cell needs overhead added before it means anything about fit.

Lower precision is not free. Aggressive quantization can change outputs, and it can even slow inference, because the runtime has to quantize and dequantize during decoding. Smaller is not automatically faster.

My decision rule: pick the highest precision that fits your memory budget with room left for overhead — not the smallest number that technically loads.

Adding Runtime Overhead Without Fooling Yourself

Parameter count and weight precision determine the weights block; runtime overhead is added alongside it, with context length and concurrent requests shown as factors that can enlarge the overhead block.
Separate the predictable weight estimate from workload-dependent overhead before comparing the total with available memory.

Now the second bucket. This is where most beginner sizing errors live, because the weight estimate feels like an answer and it isn't.

The largest runtime cost is usually the KV cache — the stored attention keys and values that let the model reuse previous tokens instead of recomputing them. It grows with context length and with how many requests are in flight at once.

Other overhead: temporary activations, framework and driver allocations, and a fragmentation buffer the allocator needs.

A common shortcut is a percentage multiplier on the weight estimate — often somewhere in the 10–30% range. Useful as a rough screen. Misleading as a precise model. The KV cache depends on sequence length, batch size, layer count, and hidden size, so long-context workloads can need far more headroom than a flat percentage suggests.

Let's apply that shortcut to the worked example so the two buckets actually combine. Start with the 0.5 GB of idealized weights for the 1B model at 4 bits. Choose a 20% overhead allowance — a stated assumption, not a universal requirement:

Mtotal≈0.5 GB×1.20=0.6 GBM_{total} \approx 0.5 \text{ GB} \times 1.20 = 0.6 \text{ GB}

So the capacity screen says: plan for roughly 0.6 GB before you count the operating system, the runtime process, and anything else running on the machine. That is the number to hold against your available memory — not the 0.5 GB weight figure alone.

And keep the caveat attached. If you push context length up or run a second request, the real overhead can blow past that 20% allowance. The multiplier is a starting guess, not a ceiling.

Common mistake: Treating the weight estimate as the total requirement. Weights are the fixed block. Overhead is the block that grows while you're not looking.

Knowledge check

Check your understanding

Answer this question before you continue.

Starting with 0.5 GB of idealized weights, what total does the article's example estimate when it assumes 20% overhead?
Output Prediction

Focus: Add a stated runtime-overhead allowance to the idealized weight estimate without treating it as universal.

What the Estimate Cannot Tell You

Set the boundaries honestly, or you'll trust the number for jobs it can't do.

Fitting is not running well. Once weights are loaded, memory bandwidth often gates tokens per second more than raw capacity does. A model that fits can still crawl.

CPU offloading can fake a fit. Spilling layers to system RAM makes a model "fit" at a real cost in throughput. It loads. It may not be usable.

Hardware and software support vary. Not every GPU accelerates every precision format natively, so a format that fits may not be the fast path.

Treat the estimate as a capacity screen. It rules models out cheaply and tells you what to measure next. It does not tell you what performance you'll get.

Use it when shortlisting models before downloading tens of gigabytes. Don't use it for capacity planning across multiple users or long-context workloads — that's a different problem with different math.

Your Next Step: Screen, Then Measure

Run the formula on the two or three models you're actually considering, at the precision you'd realistically download. Write the weight estimate next to your available memory, add a stated overhead allowance on top, and leave explicit headroom instead of assuming zero.

The estimate tells you what to try. An actual run tells you what happens. Measure elapsed time and observed memory use on your own machine — and remember that measurements from one machine are observations, not benchmarks.

Weights are predictable. Overhead is not. Keep those two buckets separate and you'll stop being surprised by numbers that disagree. Then go run one and see what your hardware says.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

For the same 1B model, which comparison is consistent with the article's idealized estimates?
Question 1 of 2Comparison Reasoning

Focus: Compare idealized weight estimates across precisions and use the estimate with an overhead margin when screening capacity.

A model's estimated memory is below the machine's available capacity. What conclusion is justified by the article?
Question 2 of 2Misconception Check

Focus: Distinguish a rough capacity screen from guarantees about fit, speed, and performance.

References

  1. Optimizing LLMs for Speed and Memory · Hugging Facehuggingface.co
  2. Calculating GPU memory for serving LLMs | LLM Inference Handbookhandbook.modular.com
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.