Skip to content
intermediate

Practice Choosing an Agent Memory Strategy for Each Task

The agent remembered everything and still gave the wrong answer. That is the failure this drill is built to prevent.

Published 2026-10-03Updated 2026-10-0410 min read
Close-up view of rippled sand texture at Leba Beach, Poland, showcasing natural patterns.
Close-up view of rippled sand texture at Leba Beach, Poland, showcasing natural patterns. Photo by Jan Kopřiva on Pexels.

The agent remembered everything and still gave the wrong answer. That is the failure this drill is built to prevent.

Memory is not a feature you add once. It is a per-task allocation decision, and most bad agent memory comes from choosing a mechanism by habit instead of by constraint. This exercise gives you a runnable fixture of synthetic tasks, each carrying its own persistence, freshness, retrieval, and privacy constraints. Your job is to pick the smallest mechanism that satisfies the task, cite the evidence, and flag anything that needs verification before it can be trusted.

We are drilling the bottleneck skill here: choosing well, and knowing when the choice is not enough.

Why Memory Choices Fail Before the Code Does

The default move is to reach for retrieval or bolt on a memory store as a blanket upgrade. That over-builds simple tasks and under-serves the ones with real freshness or privacy constraints. A support agent that needs to remember a preference for the next ten minutes does not need a vector database. A multi-day project tracker does not survive on a summary.

Four mechanisms cover almost everything you will face:

  • Context-only — the information lives inside the current session and dies with it.
  • Summary — long history compressed to a gist, with detail loss accepted.
  • Retrieval — facts pulled on demand from a corpus too large to hold in context.
  • Maintained state — values that survive the session and change over time.

The constraint types discriminate between them: persistence, freshness, retrieval need, and privacy scope. If you already know what these mechanisms are, good — this article assumes that and moves straight to the choosing. The second half of the skill is the part people skip: a chosen mechanism is not a trusted one. Stale or conflicting state needs a verification or escalation disposition before the agent acts on it.

What You Will Build and What You Need

The fixture is a single self-contained Python file. No external services, no API keys, no network calls — synthetic tasks only.

Inputs: a list of task records, each carrying constraint fields.

Outputs: a completed record per task with a chosen mechanism, cited task evidence, and a verification disposition, plus a checker report listing missing criteria.

Environment: standard library only, Python 3.x, run from the command line.

Success criteria: every case has a justified mechanism choice and a required verification disposition, and the checker reports no missing criteria.

One boundary up front: the checker is a consistency aid. It confirms that a choice exists, that evidence is cited, and that a disposition is present. It cannot judge whether your choice is correct, and it cannot guarantee the memory stays fresh or correct in production. That part is on you.

Read the Fixture: Tasks, Constraints, and the Record Shape

Here is the complete task set. Each field tests a specific decision:

TASKS = [
    {
        "id": "T1",
        "description": "Answer follow-up questions within one chat session.",
        "needs_persistence": False,
        "freshness_required": False,
        "retrieval_required": False,
        "privacy_scope": "session",
        "conflicting_state": False,
    },
    {
        "id": "T2",
        "description": "Summarize a 40-turn debugging conversation for a handoff note.",
        "needs_persistence": False,
        "freshness_required": False,
        "retrieval_required": False,
        "privacy_scope": "session",
        "conflicting_state": False,
    },
    {
        "id": "T3",
        "description": "Answer questions across a 500-page internal policy corpus.",
        "needs_persistence": True,
        "freshness_required": True,
        "retrieval_required": True,
        "privacy_scope": "org",
        "conflicting_state": False,
    },
    {
        "id": "T4",
        "description": "Track a project's current deployment target across multiple days.",
        "needs_persistence": True,
        "freshness_required": True,
        "retrieval_required": False,
        "privacy_scope": "team",
        "conflicting_state": True,
    },
    {
        "id": "T5",
        "description": "Remember a user's preferred notification channel across sessions.",
        "needs_persistence": True,
        "freshness_required": False,
        "retrieval_required": False,
        "privacy_scope": "user",
        "conflicting_state": False,
    },
]

And the record you fill in:

ANSWER = {
    "id": "T1",
    "mechanism": None,          # "context" | "summary" | "retrieval" | "state"
    "evidence": "",             # which constraint fields justify the choice
    "verification": None,       # "none" | "verify_before_use" | "escalate"
}

Evidence citation is required because this exercise is about justifying a choice from constraints, not guessing an intended answer. Some tasks are deliberately ambiguous or carry conflicting state. Those are the ones that need a verification or escalation disposition rather than a confident mechanism choice.

A rough decision flow from constraints to candidates:

needs_persistence? --no--> retrieval_required? --no--> context-only
                   |                          --yes--> retrieval
                   --yes--> values change? --yes--> maintained state
                                          --no--> summary or state

Knowledge check

Check your understanding

Answer this question before you continue.

For T3, which fixture field most directly signals that retrieval is required?
Scenario Interpretation

Focus: Identify which task constraint explicitly requires retrieval from a large corpus.

Run It: Baseline Output and What the Checker Reports

Here is the checker. It validates structure, not judgment:

import json

REQUIRED = {"id", "mechanism", "evidence", "verification"}
VALID_MECHANISMS = {"context", "summary", "retrieval", "state"}
VALID_VERIFICATIONS = {"none", "verify_before_use", "escalate"}

def check(answers):
    missing = []
    for a in answers:
        gaps = []
        if a.get("mechanism") not in VALID_MECHANISMS:
            gaps.append("mechanism")
        if not a.get("evidence"):
            gaps.append("evidence")
        if a.get("verification") not in VALID_VERIFICATIONS:
            gaps.append("verification")
        if gaps:
            missing.append((a["id"], gaps))
    return missing

if __name__ == "__main__":
    answers = [
        {"id": t["id"], "mechanism": None, "evidence": "", "verification": None}
        for t in TASKS
    ]
    report = check(answers)
    for tid, gaps in report:
        print(f"{tid}: missing {', '.join(gaps)}")
    print(f"{len(report)} cases incomplete.")

Run it with placeholder records first. You should see:

$ python memory_drill.py
T1: missing mechanism, evidence, verification
T2: missing mechanism, evidence, verification
T3: missing mechanism, evidence, verification
T4: missing mechanism, evidence, verification
T5: missing mechanism, evidence, verification
5 cases incomplete.

The checker validates structure, not judgment. If it reports nothing missing but your choices are wrong, the checker is working exactly as designed — your reasoning is the thing under test.

Common mistake: Treating a clean checker run as proof of a correct memory strategy. A passing run means you filled the fields. It says nothing about whether the memory stays fresh.

Knowledge check

Check your understanding

Answer this question before you continue.

The checker reports no missing criteria for a completed record. What does that establish?
Misconception Check

Focus: Distinguish structural validation by the checker from judgment about a memory choice.

Fill the Records: Choosing the Smallest Mechanism

Now the core loop. Match the constraint signal to the mechanism:

MechanismPersistenceFreshness handlingRetrieval needTypical failure mode
Context-onlyNoneAlways currentNoneLost when session ends
SummaryWithin sessionCurrent at write timeNoneSilently drops the detail that mattered
RetrievalExternal corpusDepends on index freshnessRequiredStale index returns old facts
Maintained stateAcross sessionsMust be updated explicitlyOptionalConflicting writes, no authority

Two worked cases show the trap.

Context-only is correct here (T1): "Answer follow-up questions within one chat session." No persistence, small working set. Reaching for retrieval is over-building — you would add an index, a query path, and a freshness problem to solve a problem that does not exist.

Maintained state is required here (T4): "Track a project's current deployment target across multiple days." A summary would compress the history and silently lose the update when the target changes. The value must survive the session and be revised, so it belongs in maintained state.

Privacy and scope are constraints on the mechanism, not separate mechanisms. If a task says the state is scoped to one user, that shapes where and how you store it — it does not add a fifth option.

Knowledge check

Check your understanding

Answer this question before you continue.

A 40-turn debugging conversation must be compressed into a handoff note, with no need to preserve every detail or retrieve from a corpus. Which mechanism best fits this task?
Comparison Reasoning

Focus: Choose a memory mechanism that compresses session history for a handoff when detail loss is acceptable.

Flag What Must Be Verified: Stale and Conflicting State

A chosen memory value is checked for conflicting records first. An unresolved conflict leads to escalation; otherwise, a freshness concern leads to verification before use, and no concern leads to no additional check.
Choose the memory mechanism first, then check for conflict or staleness before trusting its contents.

Choosing a mechanism is half the job. The other half is deciding when the mechanism is not enough.

Stale state: the stored value was true earlier and may no longer be. The record needs a freshness check before the value is trusted.

Conflicting state: two records disagree and neither is obviously authoritative. The record needs a resolution rule or an escalation.

Distinguish the two dispositions carefully. Verify before use is a cheap check the agent can run itself. Escalate is a boundary the agent should not cross alone — when the conflict cannot be resolved without human judgment.

Common mistake: Writing a mechanism choice and leaving the verification field empty because the choice felt confident. Confidence in the mechanism says nothing about the freshness of the data inside it.

And the honest limit: a verification disposition reduces the chance of acting on bad state. It does not make the state correct.

Knowledge check

Check your understanding

Answer this question before you continue.

A project's deployment target is maintained across days, but two records disagree and neither is clearly authoritative. Which disposition fits the article's guidance?
Scenario Interpretation

Focus: Select an appropriate disposition when persistent state conflicts and there is no clear authority.

Compare Against the Explained Solutions

Here are the explained solutions for all five cases. Compare your records against these:

T1 — Context-only, verification: none. No persistence, no retrieval need, session-scoped. The information lives and dies with the conversation. Nothing to verify because nothing outlives the turn.

T2 — Summary, verification: none. The 40-turn history needs compression for a handoff note. Detail loss is acceptable here because the goal is a gist, not a replay. No persistence beyond the session, no retrieval corpus, no conflicting state.

T3 — Retrieval, verification: verify_before_use. The 500-page corpus is too large for context, and retrieval is explicitly required. Freshness is required, which means the index must be checked before the answer is trusted. The org scope constrains where the index lives, not which mechanism you pick.

T4 — Maintained state, verification: escalate. The deployment target persists across days and changes. Conflicting state is flagged, meaning two records may disagree with no clear authority. That is an escalation boundary — the agent should not silently pick one. This is the case where a confident mechanism choice is not enough.

T5 — Maintained state, verification: none. The notification preference persists across sessions but does not change frequently. No freshness requirement, no conflicting state, no retrieval need. Maintained state is the smallest fit. If the preference later conflicts with a newer statement, that would trigger verification — but the task as written does not carry that constraint.

When your choice differs, the useful question is not whether you matched the key — it is which constraint you weighted differently. Some cases have more than one defensible mechanism. The tipping factor is usually the constraint you ranked highest: if persistence dominates, maintained state wins; if detail loss is acceptable and the session is short, summary is fine. Other cases have a clearly smallest mechanism, and the tempting larger one is over-building. Learn to spot both.

Change One Constraint and Re-Run

This is the real test. If your choice does not move when the constraint moves, you memorized cases instead of learning the rule.

  • Take T1 and set needs_persistence: True. The correct mechanism should move to maintained state.
  • Take T4 and set conflicting_state: False. The verification disposition should drop from escalate to verify_before_use or none, depending on whether freshness still matters.
  • Take T3 and set freshness_required: False. The verification disposition should disappear, but the mechanism stays retrieval because the corpus size and retrieval need did not change.

Run the checker again after each mutation. Watch the records change with the constraints.

The rule to carry into real agent work: name the constraint first, pick the smallest mechanism that satisfies it, then decide what must be verified before the state is trusted. When you wire memory into an actual agent loop, apply that same constraint-first reasoning to every category of information the agent touches — current task, user preferences, and history each land in a different place.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

T1 originally needs only within-session follow-up. If its only changed constraint is `needs_persistence: True`, what mechanism should the learner choose under the article's stated mutation?
Question 1 of 2Output Prediction

Focus: Predict how a mechanism choice changes when a session-only task gains a persistence requirement.

T3 still concerns questions across a 500-page corpus and still has `retrieval_required: True`, but `freshness_required` changes to false. What should happen?
Question 2 of 2Comparison Reasoning

Focus: Separate the mechanism choice driven by retrieval need from the verification disposition driven by freshness.

References

  1. Effective context engineering for AI agentswww.anthropic.com
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.