Practice Identifying Prompt Injection and Choosing the Right Boundary
You can recite that untrusted content can carry instructions. Then someone shows you a proposed action and asks, "Should this run?" and the recitation…

Key topics
You can recite that untrusted content can carry instructions. Then someone shows you a proposed action and asks, "Should this run?" and the recitation doesn't help. This lab moves that decision out of the model's head and into code you can read, run, and break.
A quick bridge from what you already know: prompt injection is an instruction-conflict problem, and the durable control lives at the boundary where actions are taken, not in the prompt text. If that framing is new, read the prerequisite article first. Here, we assume it and go straight to building.
What This Lab Actually Tests
The harness you're about to write is a policy checker for proposed actions. It is not a prompt sanitizer, and it is not a model-level guardrail. It does not read the text and decide whether the text looks suspicious. It looks at three things:
- The trust label of the content that produced the action — trusted instruction or untrusted content.
- The action itself — what the system is about to do.
- The authorization for that action — is this action within the scope the user actually granted?
From those inputs, it returns one of three decisions: allow, block, or escalate (send to a human for review).
Two things are explicitly out of scope. First, we are not writing attack payloads. The scenarios describe situations; they don't reproduce exploits. Second, we are not treating delimiters, "ignore previous instructions" phrasing, or any other text-level trick as a defense. Those are not reliable boundaries, and a lab that pretends otherwise teaches the wrong lesson.
The mechanism is the point: a policy on capability survives prompt wording changes that a text filter would not. If the model's instructions are rewritten tomorrow, your boundary still holds, because it never depended on the wording.
Set Up the Scenario Data
Start with the smallest thing that can run. A scenario is a dict with an id, a source trust label, an action, and whether that action is authorized for this task. We keep the action types fixed so the policy has something concrete to reason about.
# scenarios.py — standard library only, no API keys, no model calls
SCENARIOS = [
{"id": "s1", "source": "trusted", "action": "read_lookup", "authorized": True},
{"id": "s2", "source": "untrusted", "action": "read_lookup", "authorized": True},
{"id": "s3", "source": "trusted", "action": "send_message", "authorized": True},
{"id": "s4", "source": "untrusted", "action": "send_message", "authorized": True},
{"id": "s5", "source": "untrusted", "action": "write_data", "authorized": True},
{"id": "s6", "source": "untrusted", "action": "external_call", "authorized": True},
{"id": "s7", "source": "trusted", "action": "external_call", "authorized": True},
{"id": "s8", "source": "trusted", "action": "write_data", "authorized": False},
]
Four action types cover the interesting space:
| Action type | What it does | Reversible? | Blast radius |
|---|---|---|---|
read_lookup | Reads a permitted, non-sensitive field | Yes | None |
send_message | Sends something outward | No | External party sees it |
write_data | Modifies stored state | Sometimes | Corrupts downstream state |
external_call | Reaches a network endpoint | No | Leaves the system |
The split that matters is read-only versus state-changing or outbound. Reversibility and blast radius drive every rule that follows. A read you can repeat costs nothing if it's wrong. A message you can't unsend, or a write that corrupts state, is a different category of decision.
Two clarifications before you run anything. The read_lookup action is deliberately narrow — a permitted, non-sensitive field. Reading a sensitive record is not harmless just because it doesn't mutate state; it can leak data, and that case belongs in the escalate or block column, not the allow column. And authorized is separate from source. A trusted instruction can still ask for something outside the scope the user granted, and the policy has to catch that.
The expected output shape is a decision per scenario: allow, block, or escalate. Know what success looks like before you run anything.
Note: The scenario descriptions are deliberately descriptive, not instructional. The goal is to classify the situation, not to reproduce an attack. If you find yourself writing payload text, you've left the lab.
Knowledge check
Check your understanding
Answer this question before you continue.
Write the Boundary Policy
The policy is a small set of explicit rules. No model judgment call, no scoring, no fuzzy matching. Here's the core function:
# policy.py
ALLOWED = "allow"
BLOCKED = "block"
ESCALATE = "escalate"
def decide(source, action, authorized):
# Unauthorized actions are blocked regardless of source trust.
if not authorized:
return BLOCKED
# Permitted, non-sensitive read-only actions: allow regardless of source.
if action == "read_lookup":
return ALLOWED
# Untrusted content that changes state or leaves the system: block.
if source == "untrusted" and action in ("write_data", "external_call"):
return BLOCKED
# Untrusted content that sends a message: escalate, don't block.
if source == "untrusted" and action == "send_message":
return ESCALATE
# Trusted source, authorized, state-changing or outbound: allow.
if source == "trusted" and action in ("send_message", "write_data", "external_call"):
return ALLOWED
# Default deny-by-omission: no matching rule means no silent pass.
return BLOCKED
Read the rules in order and notice the reasoning behind each one.
Authorization is checked first, and it is independent of trust. A trusted instruction that asks for an action outside the granted scope is still blocked. Trust tells you where the instruction came from; authorization tells you whether the action was ever permitted. Conflating the two is the mistake this rule exists to prevent.
Permitted, non-sensitive reads are allowed regardless of source. A lookup on a permitted field has no side effects. If untrusted content triggers one, the worst case is a wasted query. Blocking it would cost usefulness for no safety gain. Sensitive reads are a different case and should route through review.
Untrusted content that writes or calls out is blocked. These actions change state or leave the system. The trust label is weak, and the action is not reversible. Block.
Untrusted content that sends a message is escalated, not blocked. This is the interesting rule. A message might be legitimate — maybe the untrusted content is a customer email and the reply is exactly what the user wanted. Blocking it makes the system useless. Escalating keeps it useful while putting a human at the decision point. The rule encodes a judgment: when the action is plausibly legitimate but the trust label is weak, route it to review rather than killing it.
Trusted, authorized sources get to act. If the instruction came from a trusted source and the action is within scope, the action is allowed. That's the whole point of a trust label — it buys capability, but only inside the boundary the user already granted.
The default is deny. If an action type has no matching rule, it returns BLOCKED. This is the most important line in the file. A new action type added later will not silently pass; it will fail closed until someone writes a rule for it.
The mechanism here is a boundary on capability, not on text. The policy never looks at what the content says. It looks at what the system is about to do, who asked for it, and whether that action was ever authorized. That's why it survives prompt rewording.
Knowledge check
Check your understanding
Answer this question before you continue.
Run It and Read the Outcomes
Wire the two files together and print the decisions:
# run.py
from scenarios import SCENARIOS
from policy import decide
for s in SCENARIOS:
result = decide(s["source"], s["action"], s["authorized"])
print(f'{s["id"]} {s["source"]:9} {s["action"]:13} auth={str(s["authorized"]):5} -> {result}')
Expected output:
s1 trusted read_lookup auth=True -> allow
s2 untrusted read_lookup auth=True -> allow
s3 trusted send_message auth=True -> allow
s4 untrusted send_message auth=True -> escalate
s5 untrusted write_data auth=True -> block
s6 untrusted external_call auth=True -> block
s7 trusted external_call auth=True -> allow
s8 trusted write_data auth=False -> block
Walk through four of these.
s1 — allow. Trusted source, authorized, read-only action. The first two rules pass through. No side effects, no risk.
s5 — block. Untrusted source, write action. The third rule fires. The content came from somewhere you don't control, and the action changes stored state. Block.
s4 — escalate. Untrusted source, send action. The fourth rule fires. The message might be legitimate, but the trust label is weak, so a human decides.
s8 — block. Trusted source, but the action is not authorized. The first rule fires before trust is even considered. This is the scenario that separates provenance from permission: the instruction is trusted, and the action is still out of scope.
Now look at the pair that proves the boundary is doing its job: s3 and s4. Same action type — send_message — but different decisions. s3 is allowed because the source is trusted and the action is authorized; s4 is escalated because the source is untrusted. The decision came from the action, its source, and its authorization, not from scanning the text for suspicious phrases. That's the whole design.
And here's the cost. Look at s7: trusted source, external call, allowed. Now imagine the same call arriving through untrusted content — s6 — and it's blocked outright. If that external call were a legitimate part of a workflow triggered by a document the user uploaded, the policy just blocked something a human would have approved. That's the price of the control. Blunt rules are safe and dumb. The next section is about tuning that bluntness.
Knowledge check
Check your understanding
Answer this question before you continue.
Modify the Policy and Predict the Change
Reading output is passive. Predicting it is where the learning happens. Pick one modification, write down your predicted decision for all eight scenarios, then re-run and compare.
Try this one: require escalation for all outbound actions, regardless of trust label. Change the send_message and external_call rules so that any outbound action escalates, trusted or not.
Before you touch the code, write your predictions. For each scenario, what does decide return now?
Here's what actually changes:
- s3 flips from
allowtoescalate. A trusted message now needs review. - s7 flips from
allowtoescalate. A trusted external call now needs review. - s4 stays
escalate. It was already escalated. - s6 flips from
blocktoescalate. Untrusted outbound now goes to review instead of being killed. - s1, s2, s5, s8 are unchanged.
If your predictions matched, your mental model of the rule ordering is solid. If any mismatched, that's not a failure — it's evidence about where your model of the policy diverges from the policy itself. That gap is exactly what you want to find in a lab instead of in production.
Notice what a single rule change did: it shifted four outcomes at once. That's the real lesson about boundary design. Rules aren't independent. Tightening one edge moves several decisions, and the tradeoff is always the same — tighter policy means less risk and less usefulness. The right setting depends on the action's reversibility and impact. A message you can't unsend deserves more caution than a lookup you can repeat.
Knowledge check
Check your understanding
Answer this question before you continue.
Failure Modes and What This Harness Does Not Guarantee
This is where most labs get dishonest, so let's be precise about the limits.
The trust label is an input you supply. The policy trusts the label. If upstream labeling is wrong — if untrusted content gets tagged trusted — the policy makes confident wrong decisions. The checker is only as good as the labeling that feeds it, and labeling is a separate, harder problem.
Authorization is also an input you supply. The authorized flag is only as good as the scope definition behind it. If the scope is too broad, the policy will happily allow actions the user never intended. If it's too narrow, legitimate work gets blocked and people route around the boundary.
The policy only governs actions that pass through it. Any code path that takes an action without calling decide is completely unprotected. A single bypass — a direct tool call, a cached shortcut, a debug route — undoes the whole boundary. The checker is a chokepoint, and chokepoints only work if everything goes through them.
A passing test set proves the rules behave as written, not that the system is safe. Eight scenarios passing tells you the code does what you told it to do. It says nothing about the scenario you didn't write. Your test set is a sample, not a proof.
Text-level defenses are not part of this harness on purpose. Delimiters and instruction wording can be spoofed, and the model still decides how much weight to give each source. They are not boundaries. That's why the policy never looks at the text.
Real systems need layers. This checker is one layer. Least privilege, human review at high-impact decisions, monitoring, and evaluation are others. A boundary that blocks the wrong things is a boundary people route around, and a routed-around boundary protects nothing.
Common mistake: Treating a green test run as a safety guarantee. The harness proves your rules are consistent. It does not prove your system is safe against scenarios you haven't imagined.
Where to Take This Next
The lab is useful as-is, but it becomes a real artifact with one addition: logging. Add a reason to every decision so you can audit which rule fired.
# policy.py — same rules, now returning a reason
def decide(source, action, authorized):
if not authorized:
return BLOCKED, "unauthorized"
if action == "read_lookup":
return ALLOWED, "permitted_read"
if source == "untrusted" and action in ("write_data", "external_call"):
return BLOCKED, "untrusted_state_change"
if source == "untrusted" and action == "send_message":
return ESCALATE, "untrusted_outbound"
if source == "trusted":
return ALLOWED, "trusted_authorized"
return BLOCKED, "default_deny"
The return shape changed from a single string to a tuple, so the caller has to unpack it. Update run.py to match:
# run.py — updated for the (decision, reason) return shape
from scenarios import SCENARIOS
from policy import decide
for s in SCENARIOS:
decision, reason = decide(s["source"], s["action"], s["authorized"])
print(f'{s["id"]} -> {decision:8} ({reason})')
Expected output:
s1 -> allow (permitted_read)
s2 -> allow (permitted_read)
s3 -> allow (trusted_authorized)
s4 -> escalate (untrusted_outbound)
s5 -> block (untrusted_state_change)
s6 -> block (untrusted_state_change)
s7 -> allow (trusted_authorized)
s8 -> block (unauthorized)
Now every decision carries its reason. When something gets blocked in production, you can read the log and see exactly which rule fired — the policy becomes auditable instead of mysterious.
Two more moves worth making. Add a scenario that is legitimate but arrives through untrusted content, and decide whether escalation or a narrower action is the better answer. Then reuse the same policy shape when you review an agent's tool permissions: list the actions, label the sources, mark which actions are authorized, and decide allow, block, or escalate for each. The shape transfers.
The decision rule to carry away: when you can't separate instructions from data in the text, move the decision to the boundary where actions happen. Make that boundary explicit, inspectable, and honest about its limits. Then add the logging, re-run the harness, and leave with a working artifact instead of a summary.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Want a more structured LLMOps path?
Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.
Large Language Models Starter Pack
A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.
- 227-page Illustrated PDF edition
- 12 guided LLM engineering chapters
- Visual concept diagrams
- Self-assessment quizzes
- Bonus deep-dive sections
- Prompt design, structured output, context windows & RAG pipelines
- Agents, tool calling, prompt injection, evaluation & application lifecycles
Coming soon


