Test Conversation-History Strategies on a Fixed Task
A chat assistant repeats a detail from ten turns ago, then forgets the order number you gave it two turns ago. Nothing about the model changed between…

Key topics
A chat assistant repeats a detail from ten turns ago, then forgets the order number you gave it two turns ago. Nothing about the model changed between those two moments. The history policy did.
Beginners usually assume conversation history is either "kept" or "lost." That model breaks the moment you watch a real chat. History handling is a selection policy: every strategy keeps some facts and drops others, and the only honest question is which facts, and why. This exercise makes that question measurable. You will run two fixed strategies against one fixed task, score the results with one rubric, and read the tradeoffs instead of hunting for a winner.
What This Experiment Actually Tests
Before any code, fix the scope. If you change the task between runs, you learn nothing about the strategies.
The unit of measurement here is a task-relevant fact, not a message. A message is a blob of text; a fact is a specific piece of information the task genuinely needs to succeed. For a support-style task, that might be:
- An order ID
- A constraint the user stated ("must ship before Friday")
- A decision that was made ("refund approved")
- A correction the user issued ("actually, it's the blue one")
- A preference ("email me, don't call")
Define five or six of these up front. They become your scorecard.
Two strategies are held fixed:
- Sliding window — keep the last N turns verbatim, drop everything older.
- Summarization — compress older turns into a short summary, inject it as context ahead of the recent turns.
The task stays identical across every run. The only variable is the history strategy.
Your success criterion is simple: for each run, count how many of the predefined facts are still recoverable from the assembled context. Not from the model's answer — from the context you handed it.
If you have read how apps manage long chats and what a context window is, this is the next step. A context window limits what one request can see. This exercise measures what a chosen policy puts inside that limit.
Knowledge check
Check your understanding
Answer this question before you continue.
Set Up the Scenario and the Fact Checklist
You need a conversation you can reuse. Here is a compact support-style scenario, 14 turns, with a task switch in the middle — because task switches are where history handling gets interesting.
Turn 1 user: Hi, I need help with an order.
Turn 2 assistant: Sure — what's the order ID?
Turn 3 user: It's ORD-4471. It was supposed to arrive Tuesday.
Turn 4 assistant: Got it. Let me check ORD-4471.
Turn 5 user: Also, I need it before Friday. I'm traveling.
Turn 6 assistant: Understood — Friday deadline noted.
Turn 7 user: Actually, can we also talk about my subscription?
Turn 8 assistant: Of course. What's going on with it?
Turn 9 user: I want to downgrade to the basic plan.
Turn 10 assistant: Noted. Anything else on the order?
Turn 11 user: Yes — the shipping address is wrong.
Turn 12 assistant: What's the correct address?
Turn 13 user: Actually, the order ID is ORD-4472, not ORD-4471.
Turn 14 user: And the correct address is 88 Pine Street.
Now build the fact checklist. This is the artifact you will reuse for every run.
| Fact | Where it appears | Why the task needs it |
|---|---|---|
| Order ID is ORD-4472 | Turn 13 (corrects turn 3) | Wrong ID means wrong order |
| Deadline is before Friday | Turn 5 | Determines shipping method |
| User is traveling | Turn 5 | Context for the deadline |
| Subscription downgrade requested | Turn 9 | Separate task, still active |
| Correct address is 88 Pine Street | Turn 14 | Required to fix shipping |
| Original ID was ORD-4471 | Turn 3 (superseded) | Stale — should not be treated as current |
Notice the correction pair: turn 3 states ORD-4471, turn 13 states ORD-4472. A policy that keeps the old value and drops the update has failed, even though a fact survived. Hold onto that distinction — it is where most beginners get fooled.
One assumption to state plainly: this is a text-only, single-user scenario with no retrieval layer. Results describe history policy alone, not the full behavior of a production app.
Build the Sliding-Window Baseline
Get something runnable first. The smallest useful version is a list of message dictionaries, a max_turns constant, and a function that returns the last N turns plus the current user message.
history = [
{"role": "user", "content": "Hi, I need help with an order."},
{"role": "assistant", "content": "Sure — what's the order ID?"},
{"role": "user", "content": "It's ORD-4471. It was supposed to arrive Tuesday."},
{"role": "assistant", "content": "Got it. Let me check ORD-4471."},
{"role": "user", "content": "Also, I need it before Friday. I'm traveling."},
{"role": "assistant", "content": "Understood — Friday deadline noted."},
{"role": "user", "content": "Actually, can we also talk about my subscription?"},
{"role": "assistant", "content": "Of course. What's going on with it?"},
{"role": "user", "content": "I want to downgrade to the basic plan."},
{"role": "assistant", "content": "Noted. Anything else on the order?"},
{"role": "user", "content": "Yes — the shipping address is wrong."},
{"role": "assistant", "content": "What's the correct address?"},
{"role": "user", "content": "Actually, the order ID is ORD-4472, not ORD-4471."},
{"role": "user", "content": "And the correct address is 88 Pine Street."},
]
def build_context(history, new_message, max_turns):
recent = history[-max_turns:]
return recent + [{"role": "user", "content": new_message}]
for n in (2, 6):
print(f"--- max_turns = {n} ---")
for msg in build_context(history, "Please fix the shipping.", n):
print(f"{msg['role']}: {msg['content']}")
Dependencies are minimal. Any chat model works, but you do not need one to observe the selection behavior — a local script that prints the assembled context is enough. No API key required for the baseline.
Run it with max_turns = 2. The context contains only the last two turns plus your new message. Turn 13 (the correction) survives. Turn 5 (the Friday deadline) is gone. Turn 3 (the stale ID) is also gone — which happens to be lucky, not correct.
Run it with max_turns = 6. Now turn 5 comes back, but so does the stale ORD-4471 from turn 3. The window does not know which one is current. It only knows which one is recent.
That is the mechanism: the window is a recency filter, and recency is blind to importance. An early constraint and an early greeting are dropped with equal indifference. Fill in your fact checklist for each run and note what fell off the edge.
Knowledge check
Check your understanding
Answer this question before you continue.
Add Summarization and Compare Retention
Now the second strategy. Compress turns older than the window into a short summary, then inject that summary ahead of the verbatim recent turns.
def build_summarized_context(history, new_message, max_turns, summary):
recent = history[-max_turns:]
context = []
if summary:
context.append({"role": "system", "content": f"Earlier conversation summary:\n{summary}"})
return context + recent + [{"role": "user", "content": new_message}]
The summarization instruction matters more than the model you use. Ask for key facts, decisions, names, and numbers — and frame the output explicitly as a summary, not a new instruction. A summary that reads like a command can shift the model's behavior in ways you did not intend.
A reasonable summary of turns 1–12 might read:
User reported order ORD-4471 arriving late, needs delivery before Friday due to travel. User also requested a subscription downgrade to the basic plan. User says the shipping address is wrong.
Run the identical fact checklist against this summarized context. Record the score side by side with the window run.
The typical pattern: summarization recovers distant facts the window lost — the Friday deadline, the subscription request — but it can blur or drop the exact detail the task depends on. Notice the summary above says ORD-4471. It was written before the correction in turn 13, so it carries the stale value forward. A window with max_turns = 6 would have done the same thing. Different mechanism, same failure.
Your output is a two-column score table, not a single winner.
Knowledge check
Check your understanding
Answer this question before you continue.
Score Results With One Rubric
Fix the scoring rules before you look at outcomes. Otherwise you will unconsciously grade the strategy you already prefer.
Use three levels per fact:
| Level | Meaning |
|---|---|
| Retained exactly | The fact appears with its precise value — the exact ID, number, or constraint |
| Retained approximately | Paraphrased but still usable — "needs it soon" for "before Friday" |
| Lost | Not recoverable from the assembled context |
Two rules keep this fair:
Score the assembled context, not the model's final answer. If you score the answer, you are measuring generation variance alongside policy. Isolate the policy.
Handle the correction pair explicitly. A policy that keeps ORD-4471 and drops ORD-4472 is a retention failure, even though a fact survived. Mark it as lost for the "current order ID" row.
Keep the same rubric across both strategies and across repeated runs. Changing the rubric mid-experiment invalidates the comparison. Report the score as a small table, and remember: a single run is a signal, not a benchmark.
Knowledge check
Check your understanding
Answer this question before you continue.
Read the Tradeoffs, Not a Winner
Turn the scores into a decision rule instead of a ranking.
Sliding window is cheap, predictable, and easy to reason about. It is strong on recency and weak when the task depends on an early constraint or a decision made long ago. You can predict exactly what it will keep: the last N turns, nothing more.
Summarization reaches further back, but its losses are hard to predict. Exact identifiers and numbers are the first casualties, because a summarizer naturally compresses "ORD-4472" into "the order" unless you explicitly tell it not to.
Watch for three failure modes:
- A summary that quietly drops a correction, leaving the stale value in place.
- A window that keeps a stale fact because it happens to be recent.
- A summary that reads like an instruction and shifts model behavior.
The decision rule: choose by what the task cannot afford to lose. Recency-dominant tasks — interactive coding, step-by-step troubleshooting — favor a window. Long-horizon tasks with early constraints favor summarization. Tasks that need exact values favor keeping those values verbatim outside the summary.
State the boundary honestly: this small experiment shows retention behavior under fixed conditions. It does not establish a universally best policy.
Change One Variable and Re-run
Pick exactly one change and predict the outcome before you run it. The gap between your prediction and the new score table is the real lesson.
Three options:
- Shrink the window from 6 to 3 and watch which facts survive.
- Tighten the summary length and see which details the summarizer sacrifices first.
- Move the correction later in the conversation and observe whether the window catches it.
Optional extension: pin one critical fact — the current order ID — verbatim outside both strategies, and observe how much the score recovers. This is the seed of a hybrid approach, and it is worth seeing with your own eyes before you trust it.
Keep the rubric identical so the new run is comparable to the first two.
Where This Leaves You
The score table is the artifact worth keeping. It tells you, for your task, which facts your policy protects and which it quietly sacrifices.
My rule: match the history strategy to what the task cannot afford to lose, and keep exact identifiers verbatim outside any summary. A summary is a compression, and compression always chooses what to discard — you may as well choose deliberately.
The natural next step is combining a recency window with a summary, then adding retrieval for the facts that must survive any window. That is where conversation history management stops being a single choice and becomes a small system.
Knowledge check
Final check
Finish the article by checking the ideas you just learned.
References
Want a more structured LLMOps path?
Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.
Large Language Models Starter Pack
A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.
- 227-page Illustrated PDF edition
- 12 guided LLM engineering chapters
- Visual concept diagrams
- Self-assessment quizzes
- Bonus deep-dive sections
- Prompt design, structured output, context windows & RAG pipelines
- Agents, tool calling, prompt injection, evaluation & application lifecycles
Coming soon


