Model LLM Agent Permission Risk: Compare Scope, Impact, and Reversibility
Two permission grants land on the review table. One lets the agent read every ticket and send email. The other lets it read assigned tickets and draft…

Key topics
Two permission grants land on the review table. One lets the agent read every ticket and send email. The other lets it read assigned tickets and draft replies. Someone says the first one "feels riskier." Everyone nods. Nobody can say by how much, or which part of it is the problem.
That is the moment this article is built for. You already know the vocabulary from bounded permissions and least privilege: approval points, reversibility, blast radius. What you do not have yet is a way to put two candidate grants side by side and argue about them with numbers instead of adjectives. So we are going to build a small exposure model. Its only job is to make the comparison explicit, inspectable, and arguable.
One warning before we start, and I mean it: the number this model produces is a thinking tool, not a security guarantee. It ranks grants against each other. It does not predict incidents, and it never authorizes anything.
Why "Broad Feels Riskier" Is Not a Comparison
The default review conversation compares adjectives. Broad versus narrow. Scary versus safe. The problem is that two grants can differ along three independent axes at the same time: how many actions they reach, how bad the worst outcome of each action is, and how hard that outcome is to undo. A single gut ranking collapses all three into one feeling, and feelings do not survive contact with a skeptical reviewer.
Worse, the axes trade off against each other. A grant that reaches many actions, each trivial and reversible, can be safer than a grant that reaches one action that quietly deletes a production table. "Broad" and "narrow" tell you about scope. They tell you nothing about impact or reversibility, and those are where the real damage lives.
So the payoff of building a model is concrete: one comparable exposure figure per grant, plus a decision rule for narrowing, approving, or denying. Let's define the pieces.
Notation: Scope, Impact, Reversibility, Probability
Every symbol below maps to a mechanism you can point at in a real system. Define them once, then we can compute.
Let be the action set: every distinct action the agent could take if it had unlimited authority. A grant is the subset of it actually authorizes. Scope, in plain terms, is the size and content of .
For each action , define three quantities:
- Impact : a bounded severity score for the worst realistic outcome of . Use a stated scale, say 1–5, with named anchors. A 1 might be "no persistent change," a 3 "recoverable data change," a 5 "irreversible external or financial consequence." The anchors matter more than the numbers.
- Reversibility : a recovery factor in . Here means fully recoverable at negligible cost, and means effectively permanent. Note the direction: high is good.
- Probability : the estimated chance the agent actually attempts during the task window. This is an estimate, not a measurement. You are guessing, and you should write down why.
Now combine them into exposure:
Read it in plain causal language: for each action the grant allows, multiply how likely the agent is to try it, by how bad it would be, by how much of that badness you cannot take back. Sum across the grant. The factor is the part of the impact that sticks.
That is the whole model. It is deliberately hypothetical. It ranks grants; it does not forecast reality.
Knowledge check
Check your understanding
Answer this question before you continue.
What the Equation Actually Measures
Before we run numbers, we need to be precise about what is — because the name "exposure" invites a mistake.
The formula measures irreversible-impact exposure only. It is not total action risk. When , the term collapses to zero, and that zero means "no modeled irreversible impact," not "no consequence." A draft reply can still leak information, consume budget, or create a state someone has to clean up. A read can still expose data to the wrong person. None of that appears in , because none of it is irreversible in the sense the equation scores.
So treat the score as one column in a wider review, not the whole ledger. The model is good at one thing: ranking grants by the damage you cannot take back. It is silent on everything you can.
Common mistake: Reading as "this grant is safe." A zero here means the grant's actions are all reversible at negligible cost. It says nothing about scope, data exposure, cost, or who gets affected.
Knowledge check
Check your understanding
Answer this question before you continue.
Assumptions the Model Depends On
A formula is only as honest as its assumptions, so let's name them before we trust a single output.
Actions are scored independently. The sum treats each action as its own event. Real agents chain actions: read a credential, then use it, then exfiltrate. Chained and correlated actions are under-counted, sometimes badly.
is fixed for the task window. We treat probability as a constant. But a manipulated agent — one fed untrusted content through a ticket, a document, a tool output — can raise the probability of high-impact actions. The model's does not know that.
Impact and reversibility are ordinal judgments dressed as numbers. A 4 is not twice as bad as a 2. Small differences in are noise. Only large gaps and dominant terms carry signal.
The model ignores who else is affected. A write to your own scratch account and a customer-visible write can share a score but not a consequence. Blast radius is not in the formula.
When these assumptions fail, the ranking can invert — the "safer" grant scores higher. That is exactly why the number informs a decision instead of making it.
Worked Example: One Agent, Two Grants
Let's run it end to end. Scenario: a support agent that reads tickets and drafts replies. Two proposed grants.
- Broad grant: read all tickets, write to the CRM, send email, run a shell command.
- Restricted grant: read assigned tickets, draft only.
I'll assign , , and with a one-line justification each, because the justification is the real work.
| Action | Justification | |||
|---|---|---|---|---|
| Read all tickets | 0.9 | 2 | 1.0 | Read-only, no persistent change |
| Write to CRM | 0.5 | 4 | 0.6 | Recoverable but customer-visible |
| Send email | 0.4 | 4 | 0.1 | Cannot be unsent |
| Run shell command | 0.2 | 5 | 0.2 | Broad, hard to bound or undo |
| Read assigned tickets | 0.9 | 2 | 1.0 | Read-only, scoped |
| Draft only | 0.8 | 1 | 1.0 | No external side effect |
Now compute the broad grant term by term. Remember is the sticky fraction.
- Read all tickets:
- Write to CRM:
- Send email:
- Run shell command:
The restricted grant:
- Read assigned tickets:
- Draft only:
The gap is the entire 3.04, and one action dominates it: send email, at 1.44, nearly half the total. Notice what that tells you. Trimming the CRM write and the shell command together removes 1.60. Removing email alone removes 1.44. The dominant term is where the leverage is, and it is not always the scariest-sounding action — the shell command scored lower because its probability was low.
But here is the honest part. The restricted grant's zero does not mean it is risk-free. It means every action in that grant is reversible at negligible cost, so the irreversible-impact column is empty. The draft can still leak a customer's data into the wrong hands. The read can still expose a ticket the agent should not have seen. Those are real exposures; they just live outside this equation.
And the restricted grant is not automatically correct. If the task genuinely requires sending email, zero exposure is a fiction. The right move is to keep the action and add an approval boundary — a human checkpoint before send — rather than pretend the risk vanished. The model told you which action to gate. It did not tell you to delete it.
Knowledge check
Check your understanding
Answer this question before you continue.
Reading the Number: Narrow, Approve, or Deny
A total is a summary. The dominant term is the decision. Here is how I turn the calculation into a choice.
Start with the action contributing most of . That is the one worth narrowing or gating first. Reversibility is the lever with the largest effect on the score, because multiplies the whole term: a high-impact action that is fully reversible contributes zero to , while a modest action you cannot undo can dominate. But "contributes zero to " is not the same as "safe." A fully reversible action still has scope, exposure, and cost consequences the equation does not see.
So the decision rule has two stages. First, use to prioritize: the dominant term is where review effort pays off. Second, before you narrow, gate, or deny, ask the questions the score cannot answer:
- Does the task require this capability? If yes, an approval boundary is usually better than removal.
- What residual consequences remain even when the action is reversible? Data exposure, cost, and customer-visible state changes survive a rollback.
- Who else is affected? A write to your own scratch space and a write to a customer record can share a score but not a consequence.
Use this model when you are comparing two or three candidate grants for the same task, or justifying a scope reduction to a reviewer. Do not use it as a compliance artifact, as a substitute for runtime enforcement, or as evidence that a grant is safe. It is a pre-decision tool. It says nothing about what the agent does once it starts running.
Knowledge check
Check your understanding
Answer this question before you continue.
Common Mistakes When Scoring Permissions
The model is easy to run and easy to run wrong. These are the failures I see most.
Scoring the tool name instead of the action. A general-purpose shell tool hides dozens of actions behind one label. You cannot score "shell" as a single row; decompose it into the operations that actually matter, or the score is meaningless.
Treating as a property of the model. Probability belongs to the task and the input channel, not the model. Untrusted content raises the probability of high-impact actions. If your agent reads tickets, documents, or tool output, those are inputs to , and should move.
Averaging away the tail. A low-probability, irreversible action can dominate exposure while looking harmless in a mean. The sum catches it; an average hides it.
Reusing one score sheet across tasks. Scope should expire when the task ends. A sheet built for a read-heavy research task is not valid for a write-heavy ops task.
Confusing low exposure with authorization. A small number is not a permission. The calculation never grants access; a human or a policy does.
Where This Model Stops
Let's close the loop honestly. This model ranks proposed grants by irreversible-impact exposure. It does not guarantee safety, detect prompt injection, or replace runtime enforcement and audit trails. It is a pre-decision tool, not a monitoring tool — it says nothing about what the agent does after it starts running. And because the scores are judgments with stated assumptions, the useful output is the argument behind each number, not the number itself.
So here is your next step. Take one real agent you are working on. List every action it can reach. Score , , and for each, with a written justification per number. Sum it. Then look at which single action dominates — and decide whether to narrow it, gate it behind approval, or deny it outright. Write down what the score leaves out: scope, exposure, cost, affected parties. That habit, repeated, is worth more than any total the formula hands you.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Want a more structured LLMOps path?
Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.
Large Language Models Starter Pack
A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.
- 227-page Illustrated PDF edition
- 12 guided LLM engineering chapters
- Visual concept diagrams
- Self-assessment quizzes
- Bonus deep-dive sections
- Prompt design, structured output, context windows & RAG pipelines
- Agents, tool calling, prompt injection, evaluation & application lifecycles
Coming soon


