Skip to content
intermediate

How Multimodal Inputs Use Model Capacity Beyond Text Tokens

You type a short question, attach one screenshot, and hit send. The words are maybe forty. The request behaves like it swallowed a page.

Published 2026-10-03Updated 2026-10-048 min read
Detailed view of rough and dry brown soil, showcasing its natural texture.
Detailed view of rough and dry brown soil, showcasing its natural texture. Photo by Roy Photos on Pexels.

You type a short question, attach one screenshot, and hit send. The words are maybe forty. The request behaves like it swallowed a page.

Nothing is broken. You are watching a budget you cannot see on screen.

Why the Visible Token Count Stops Being the Whole Budget

Most of us carry the same mental model: characters and words map to tokens, tokens map to context usage, done. That model is genuinely useful. For plain text, with one model and one tokenizer, it predicts behavior well enough to plan a prompt.

Then an image enters the request, and the visible text stops predicting the capacity consumed.

The reason is structural. A multimodal model does not read your image the way you do. It converts that image into a set of internal representations — patches, embeddings, vision tokens — and those representations occupy the same finite space as your words. Context is a finite resource with diminishing returns, so anything that consumes it matters, even when it never appears as text you can count.

I want to be honest about scope before we go further. This article builds a model of the accounting. It is not a lookup table of real vendor numbers, and you should not treat any figure below as a prediction of what your provider will charge or accept.

Knowledge check

Check your understanding

Answer this question before you continue.

A short text prompt includes an image. Why might the request use much more capacity than its visible text-token count suggests?
Misconception Check

Focus: Explain why visible text-token counts do not fully describe capacity use for multimodal requests.

Define a Capacity Unit Before Doing Any Arithmetic

To make the mechanism visible, we need a unit. Real systems expose different accounting surfaces, so I will define a deliberately hypothetical one.

Call it a capacity unit (CU): the model's internal accounting slot for one piece of input, text or otherwise. Think of it as a seat in a theater. Every seat is identical to the model, but different kinds of input claim different numbers of seats.

Now the assumptions, stated plainly so the arithmetic is checkable:

  • One fixed total capacity per request.
  • Consumption is additive — text plus image plus history, no overlap.
  • No compression, pruning, or eviction of representations.
  • One model, one encoder configuration, held constant.

Define the conversion factors as symbols:

  • rtr_t — capacity units consumed per text token.
  • rir_i — capacity units consumed per image representation.
  • CC — total capacity available for the request.

The relationship is simple: total consumption equals text tokens times rtr_t, plus images times rir_i. Remaining capacity is CC minus that total.

Note: If the model compresses, prunes, or caches representations, this additive model becomes an approximation. It still explains the direction of the effect, but the exact number drifts.

Knowledge check

Check your understanding

Answer this question before you continue.

Under the article's additive assumptions, a request has 1,000 CU total, 50 text tokens at 2 CU each, and one image representation at 300 CU. How much capacity remains?
Single Choice

Focus: Apply the article's additive capacity-unit model to calculate remaining request capacity.

Worked Example: Text-Only Request vs. One Image

Two equal-width capacity bars compare requests. The text-only bar has 120 CU of text and 880 CU remaining. The image request has the same 120 CU of text, a 400 CU image segment, and 480 CU remaining.
With total capacity held fixed, the illustrative image uses 400 CU that would otherwise remain available.

Here are illustrative assumptions. They are chosen to make the arithmetic clean, not to describe any real product.

SymbolMeaningIllustrative value
CCTotal capacity per request1,000 CU
rtr_tCapacity per text token1 CU
rir_iCapacity per image representation400 CU

Step 1 — Text-only request.

Suppose your prompt is 120 text tokens.

consumption=120×1=120 CU\text{consumption} = 120 \times 1 = 120 \text{ CU} remaining=1000−120=880 CU\text{remaining} = 1000 - 120 = 880 \text{ CU}

You have 880 CU left for history, retrieved documents, and the model's response.

Step 2 — Add one image.

Same 120 text tokens, plus one image representation.

consumption=(120×1)+(1×400)=520 CU\text{consumption} = (120 \times 1) + (1 \times 400) = 520 \text{ CU} remaining=1000−520=480 CU\text{remaining} = 1000 - 520 = 480 \text{ CU}

Step 3 — Name the delta.

RequestText CUImage CUTotal CURemaining CU
Text only1200120880
With one image120400520480

The image displaced 400 CU — more than three times the entire text portion of the request. Your forty-word question was never the expensive part.

Step 4 — Translate remaining capacity into a consequence.

With 480 CU left, how much additional text fits before the request is truncated or rejected?

additional text tokens=4801=480 tokens\text{additional text tokens} = \frac{480}{1} = 480 \text{ tokens}

That sounds generous until you remember what competes for it: conversation history, retrieved documents, system instructions, and the model's own output. Add a second image and you spend another 400 CU, leaving 80 CU — barely enough for a short instruction, and nothing else.

Tip: Change one number and re-run it. Raise rir_i to 600 and the single-image request leaves only 280 CU. The mechanism does not change; the pressure does.

Knowledge check

Check your understanding

Answer this question before you continue.

In the worked example, the one-image request leaves 480 CU and text costs 1 CU per token. If all remaining capacity were devoted to additional text, how many additional text tokens would fit?
Output Prediction

Focus: Calculate the additional text-token capacity left in the article's one-image example.

What the Numbers Actually Predict About Behavior

The arithmetic earns its place only if it predicts something you can observe.

Capacity consumed by non-text input is capacity unavailable to instructions, retrieved documents, and conversation history. When that squeeze happens, you tend to see a recognizable cluster of signals:

  • Context gets dropped or truncated.
  • Instruction-following degrades — the model ignores a constraint you clearly stated.
  • Prefill slows down because there is more to process before the first output token.
  • Cost rises, because you are paying for representations you never typed.
  • Answers quietly ignore earlier parts of the request.

The dangerous failure mode is the silent one. The request still succeeds. It returns a fluent answer built on less evidence than you assumed, and nothing in the response tells you that something got squeezed out.

Common mistake: Treating the image as a free attachment. When an image is central to the task, cut text elsewhere — trim history, drop marginal retrieved chunks, shorten the system prompt. Pay for the image by spending less somewhere else.

Picture a stacked bar for each case. The text-only bar is almost all headroom. The image bar is a thin text sliver, a fat image block, and a much shorter remaining segment. Same total width. Very different room to work.

Knowledge check

Check your understanding

Answer this question before you continue.

A multimodal request succeeds and returns a fluent answer, but the answer ignores earlier context. Which interpretation best matches the article's warning?
Scenario Interpretation

Focus: Interpret fluent but incomplete responses as a possible sign of capacity pressure.

Why This Arithmetic Does Not Transfer Between Models

Here is where beginners get burned: they take a conversion rate from one explanation and apply it everywhere.

The conversion rate is a property of the model's encoder, its resolution handling, and its tokenizer. It is not a constant of nature. Change the model and you change rir_i — sometimes dramatically.

Three reasons the number moves:

Resolution and partitioning change the count. Higher-fidelity input generally produces more internal units. The growth is not linear in a way you can guess from pixel dimensions alone, because models partition images differently.

Accounting surfaces differ. The same image can look cheap in one interface and expensive in another, because each provider exposes a different view of consumption.

Compression, pruning, and caching shift the effective cost. The nominal count and the actual work done are not always the same thing.

So what can you safely generalize?

  • Non-text input consumes capacity.
  • It is usually larger than the visible text suggests.
  • The only reliable number comes from the model or provider you are actually using.

Everything else is a thinking tool, not a fact.

When to Reason This Way — and When Not To

Use the capacity model when you are budgeting a prompt with images, audio, or video; when you are debugging truncated or ignored context; or when you are comparing cost across providers and need a consistent frame.

Do not use it when the request is text-only and your tokenizer already gives you a real count. Do not use it when you need a billing-accurate figure — that requires the provider's own accounting, not a hypothetical unit.

Two mistakes I see repeatedly:

Common mistake: Treating the illustrative conversion rate as a real constant and building a cost estimate on top of it. The estimate will be wrong in a direction you cannot predict.

Common mistake: Assuming a smaller image is proportionally cheaper without checking how the model handles resolution. Downscaling does not always reduce the count the way you expect.

The practical habit that survives all of this: measure once with your actual model, record the observed consumption, and use that as your working estimate. One real measurement beats ten borrowed assumptions.

The Next Move

Treat every non-text input as a capacity purchase, not a free attachment. That single reframe will change how you compose prompts with images, audio, or video.

Here is your concrete next step. Take a prompt you already use and trust. Run it text-only and note the behavior. Now add one image and run it again. Watch for the squeeze: does the model follow your instructions as tightly? Does it reference earlier context as reliably? Does the answer feel thinner?

If something got squeezed out, you just found the budget — not in a dashboard, but in the behavior. That observation is worth more than any conversion rate I could hand you, because it came from the model you actually run.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

You want to understand whether adding an image affects a prompt you use with a particular model. Which first test follows the article's practical recommendation?
Question 1 of 2Scenario Interpretation

Focus: Use a controlled text-only versus image comparison to observe capacity effects on a particular model.

You need a billing-accurate estimate for an image request. How should you use the article's hypothetical capacity-unit arithmetic?
Question 2 of 2Misconception Check

Focus: Distinguish a hypothetical capacity model from a provider-accurate budget or billing estimate.

References

  1. Daily Papers - Hugging Facehuggingface.co
  2. Effective context engineering for AI agentswww.anthropic.com
  3. Cost tracking - Docs by LangChaindocs.langchain.com
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.