Skip to content
beginner

Run a Small Local LLM and Measure Its Resource Use

The model finishes downloading. You type a prompt. The cursor blinks, the fan spins up, and nothing happens for a while. That pause is the whole question:…

Published 2026-10-03Updated 2026-10-049 min read
An exciting underwater scene of people diving with a shark in Jupiter, Florida.
An exciting underwater scene of people diving with a shark in Jupiter, Florida. Photo by Al F on Pexels.

The model finishes downloading. You type a prompt. The cursor blinks, the fan spins up, and nothing happens for a while. That pause is the whole question: can this machine actually run a language model, or not?

Most beginners treat "local LLM" as a yes/no capability. It is not. It is a budget question with three numbers: memory (does it fit), latency (is it fast enough to tolerate), and quality (is the answer good enough for the job). You can measure all three in about ten minutes, and the numbers you record will tell you more than any spec sheet.

This tutorial walks you through one reproducible local inference using Ollama and a small model, then shows you how to record what actually happened. If you have already read about open-weight models and how model files are downloaded rather than trained, you have the background you need. If not, that is the one prerequisite worth covering first.

What You Are Actually Measuring

Running a model locally means inference — generating responses from downloaded weights. You are not training anything. You are loading a file into memory and asking it to produce tokens.

Three quantities decide whether the setup is usable:

QuantityWhat it answersHow you observe it
Memory footprintDoes the model fit without swapping?Peak RAM/VRAM during generation
LatencyIs it fast enough for your task?Wall-clock time from prompt to final token
Output qualityIs the answer good enough?Your own read of the response

A measurement is a result under specific conditions: your CPU or GPU, your RAM, your model file, your prompt length. It is not a benchmark. It does not transfer to other machines, and it should not be quoted as a general performance claim.

Before you measure anything, write down a task and a budget. Without them, the numbers mean nothing. For example: "Summarize a 200-word note in under 30 seconds on a laptop with 8 GB RAM." That sentence gives you a pass/fail criterion. Everything below is in service of testing it.

Knowledge check

Check your understanding

Answer this question before you continue.

After timing one prompt on your laptop, what can you responsibly conclude from the result?
Misconception Check

Focus: Distinguish a local inference observation from a general performance benchmark.

Check Your Machine Before You Download Anything

You want to know how much memory you can spare before you commit to a model size.

Windows: Open Task Manager (Ctrl+Shift+Esc) and check the Performance tab for available RAM. If you have a discrete GPU, the GPU tab shows dedicated VRAM.

macOS: Open Activity Monitor and check the Memory tab. Apple Silicon shares memory between CPU and GPU, so the total unified memory is your ceiling.

Linux: Run free -h for RAM. For GPU memory, nvidia-smi works on NVIDIA cards.

Now the key concept: quantization. A model stored in 4-bit form takes roughly a quarter of the memory of its full-precision version. That is why small quantized models fit on ordinary laptops. A 0.5B parameter model at 4-bit is tiny — well under a gigabyte — while the same model at full precision would be several times larger.

Warning: Leave headroom for your operating system and the app itself. A model that exactly fills your RAM will run, badly. If the machine starts swapping to disk, latency collapses and the fan becomes the loudest part of the experiment.

"Small" is relative. A 0.5B model is small enough to be a first experiment, not a general-purpose assistant. Treat it as a measurement subject, not a replacement for the tools you already use.

Knowledge check

Check your understanding

Answer this question before you continue.

A model seems close to filling the computer's available memory. Which adjustment best follows the article's guidance?
Scenario Interpretation

Focus: Choose a model-memory strategy that accounts for both quantization and system headroom.

Install Ollama and Pull One Small Model

Ollama is a command-line runtime that downloads and runs models locally. Install it for your platform from the official site. The install and the first model download need internet; inference afterwards can run offline.

Once installed, verify the runtime is present:

ollama --version

Then pull a small model. Qwen2.5 0.5B is a reasonable first subject:

ollama pull qwen2.5:0.5b

The download is small — a few hundred megabytes for a model this size. Confirm the runtime sees it:

ollama list

You should see the model name and its size in the output. If you do not, the pull did not complete.

Assumptions for this tutorial: you are comfortable in a terminal, you have roughly 2 GB of free disk for the runtime plus a tiny model, and you are running on a machine with at least 4 GB of usable RAM.

Why the command line first? A GUI hides the numbers you came for. The CLI gives you clean, copyable output and timing, which is exactly what you need to measure.

Run One Reproducible Inference

Use one fixed prompt with a fixed output length so runs are comparable. Changing the prompt between runs invalidates the comparison.

ollama run qwen2.5:0.5b "Summarize this in exactly three sentences: A local language model runs on your own hardware. It loads weights from disk into memory. Generation happens one token at a time."

A successful response looks like a coherent three-sentence summary followed by a clean return to your shell prompt. No error output, no hang.

Here is what happens under the hood in one paragraph: your prompt is tokenized — split into pieces the model understands. The model then generates one token at a time, each token requiring a full forward pass through the network. Generation stops when the model emits an end token or hits your length limit. That sequential process is why output length dominates latency: more tokens means more forward passes, almost linearly.

Success criteria to check off:

  • A coherent answer that addresses the prompt
  • No error output in the terminal
  • A clean exit back to your shell

If all three hold, you have a working local inference. Now measure it.

Knowledge check

Check your understanding

Answer this question before you continue.

You want to compare two runs of the same local model. Which setup makes the comparison most meaningful?
Single Choice

Focus: Set up comparable inference runs by holding the prompt and output length fixed.

Record Elapsed Time and Memory

Two identical timed runs lead to a comparison: the first includes model loading and generation, while the second is a rough estimate of generation time. Peak memory is observed during generation, and both measures are recorded.
Repeating the same prompt helps estimate loading overhead, while peak memory shows whether the run fits your machine.

The simplest honest measurement is total elapsed time for the whole command. On macOS and Linux:

time ollama run qwen2.5:0.5b "Summarize this in exactly three sentences: A local language model runs on your own hardware. It loads weights from disk into memory. Generation happens one token at a time."

On Windows PowerShell, wrap the same call in a timer:

Measure-Command { ollama run qwen2.5:0.5b "Summarize this in exactly three sentences: A local language model runs on your own hardware. It loads weights from disk into memory. Generation happens one token at a time." }

Both commands report one number: the wall-clock time from launch to exit. That number mixes two different costs together:

  • Load time — reading the model file from disk into memory. Paid once per session.
  • Generation time — producing tokens. Paid on every run.

You cannot cleanly separate them from a single timed command. But you can estimate the split with a repeat run. Run the exact same command twice in a row. The second run usually skips most of the disk read because the model is already cached in memory. The difference between run one and run two is a rough estimate of load time; the second run's total is a rough estimate of generation time for that prompt.

Note: This is an estimate, not a precise decomposition. Caching behavior, disk speed, and background processes all blur the line. Record it as an estimate and say so when you report it.

For memory, open your system monitor and watch peak usage during generation, not idle usage. On Windows, Task Manager. On macOS, Activity Monitor. On Linux, htop or nvidia-smi if you have a GPU.

Record everything in a small table:

FieldYour value
Model nameqwen2.5:0.5b
Quantization(check the model card)
Prompt(your fixed prompt)
Output length(token count or sentence count)
Total time, run 1
Total time, run 2
Estimated load time(run 1 minus run 2)
Estimated generation time(run 2 total)
Peak memory

Now compare your generation estimate against the budget you wrote down. If the answer arrives after you have lost interest, the setup fails the task — regardless of how good the answer is.

Knowledge check

Check your understanding

Answer this question before you continue.

You time the exact same command twice in a row. How should you use the two totals?
Scenario Interpretation

Focus: Interpret repeat-run timings as rough estimates rather than a precise separation of loading and generation costs.

When It Goes Wrong: Common Failures and What They Mean

A failed run is evidence about the system, not a dead end. Here is a debugging map:

Out-of-memory or heavy swapping. The model is too large for available memory. Drop to a smaller model or a more aggressive quantization.

Very slow first token. Usually model loading, not generation. Check whether the model is being read from disk on every run.

Garbage or repetitive output. Often a prompt or template mismatch, or a model too small for the task. Not necessarily a broken install.

Command not found or connection refused. The runtime is not installed, or its background service is not running. Restart the Ollama service.

Common mistake: Changing three things at once when a run fails. Change one variable at a time and re-measure, so you know which change moved the number.

One Experiment: Change a Variable and Re-Measure

A single run is a data point. Two runs with one variable changed is an experiment.

Pick one variable: a larger model in the same family, a longer output limit, or a longer prompt. Predict the direction of the change before you run it. More parameters should mean more memory and slower generation. A longer output should mean longer total time.

Then re-run with the same fixed prompt and compare against your table.

What the comparison teaches: the tradeoff curve between quality, memory, and speed is the real decision surface, and it is specific to your machine. Keep the scope honest — two data points on one laptop describe your laptop, not the model.

Decide Whether This Setup Fits Your Task

Apply the criteria you wrote at the start:

  • Does it fit in memory without swapping?
  • Does it answer within the task's time budget?
  • Is the quality good enough for the specific job?

When local fits: offline work, privacy-sensitive text, high-volume experimentation where per-call cost matters, and learning how inference behaves.

When local does not fit: tasks needing frontier-level reasoning, long-context document work, or multimodal input beyond what a tiny model supports.

State the uncertainty plainly. Your numbers are observations from one machine under one configuration. A different laptop, model file, or prompt length will produce different results.

The durable takeaway is the measurement habit — task, budget, run, record, compare — not the specific model you happened to pull.

Your next step: run the same measurement with a second model size in the same family. Two points give you a line. That line is your tradeoff curve, built from your hardware, not someone else's benchmark.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A run stays within the memory budget without swapping and finishes on time, but its answer is not good enough for the stated task. What is the best conclusion?
Question 1 of 2Comparison Reasoning

Focus: Judge a local setup against task-specific memory, latency, and output-quality criteria.

A local run produces repetitive output but no terminal error. Which next step best matches the article's troubleshooting approach?
Question 2 of 2Debugging

Focus: Respond to poor local-model output by making a focused change and treating measurements as configuration-specific.

References

  1. Use AI Models Locally · Hugging Facehuggingface.co
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.