Skip to content
intermediate

Practice Propagating RAG Document Corrections and Deletions

A query returns a sentence the source no longer contains. The team blames the model. The model is innocent — the index is still holding a chunk that was…

Published 2026-10-03Updated 2026-10-049 min read
Artistic sand texture with natural shadows creating a serene abstract pattern, perfect for background use.
Artistic sand texture with natural shadows creating a serene abstract pattern, perfect for background use. Photo by Martin Péchy on Pexels.

A query returns a sentence the source no longer contains. The team blames the model. The model is innocent — the index is still holding a chunk that was edited or deleted days ago, and nothing told it otherwise.

This is a propagation problem, not a generation problem. Your vector index is a derived cache of the source of truth. Like every cache, it goes incoherent the moment the source can change without notifying it. Prompt fixes cannot repair an index that still holds a superseded chunk.

So we are going to build a tiny harness where every layer is inspectable — source records, chunks, vectors, ranking, cache, invalidation log — and then mutate the source underneath it. When something stale survives, it will have nowhere to hide.

Why Deleted Documents Keep Winning Retrieval

A source delete is not an index delete. When you remove a row from your database, you have removed exactly one artifact. The chunks derived from it, the embeddings computed from those chunks, the metadata filters that reference it, and any cached answers keyed to it are separate derived artifacts, each needing its own removal step.

The vector index behaves like a cache of embeddings computed at indexing time. It has no built-in knowledge that the source row changed or vanished. The ranking function happily scores a chunk that no longer exists in the world.

The layers we will instrument:

  • Source records — the authoritative text with an explicit version.
  • Chunks — deterministic splits with stable IDs.
  • Vectors — embeddings with source, chunk, and version attached.
  • Ranking function — cosine similarity with a fixed threshold.
  • Cache — a simulated answer cache keyed by document ID.
  • Invalidation log — an append-only record of every removal event.

If you have already traced the RAG pipeline stages, this exercise makes the propagation between those stages observable. Ingestion, chunking, embedding, retrieval, and caching each fail in their own way. Here we watch a single mutation travel — or fail to travel — across all of them.

What You Need Before You Start

Python 3 standard library only. No model, no API credential, no vector database, no vendor cache. Every behavior in the trace is code you can read.

You should already know what a chunk is, what an embedding represents, and how cosine similarity ranks candidates. We will not re-teach those.

Two deliberate stand-ins keep the exercise deterministic:

  1. A fixed feature-vector mapping over a small controlled vocabulary replaces a real embedding model. Each word maps to a hand-written vector, so the same text always produces the same vector.
  2. A simulated cache keyed by document ID replaces a real answer cache.

The stand-in preserves versioned vectors, ID bookkeeping, ranking, and threshold behavior. It does not preserve semantic generalization — two paraphrases of the same idea will not land near each other. That is fine. We are testing propagation, not meaning.

Two rules matter for correctness:

  • Fixed relevance threshold. A candidate counts as relevant only if its cosine similarity meets or exceeds 0.80.
  • Zero-vector rule. A vector with zero magnitude has undefined cosine similarity. Exclude it from scoring entirely rather than treating it as 0.0.

Build the Versioned Store and Ranking Function

The identity model is layered: source_id, document_id, version_id, chunk_id, vector_id. Never use a filename as identity — filenames get reused, and reused identity is how stale data sneaks back in.

import math

VOCAB = {
    "refund":  (1.0, 0.0, 0.0),
    "policy":  (0.8, 0.2, 0.0),
    "days":    (0.0, 1.0, 0.0),
    "window":  (0.0, 0.9, 0.1),
    "shipping":(0.0, 0.0, 1.0),
    "returns": (0.1, 0.9, 0.0),
}

def embed(text):
    vec = [0.0, 0.0, 0.0]
    for word in text.lower().split():
        if word in VOCAB:
            for i, v in enumerate(VOCAB[word]):
                vec[i] += v
    return tuple(vec)

def cosine(a, b):
    mag_a = math.sqrt(sum(x * x for x in a))
    mag_b = math.sqrt(sum(x * x for x in b))
    if mag_a == 0.0 or mag_b == 0.0:
        return None  # zero-vector rule: undefined, exclude
    dot = sum(x * y for x, y in zip(a, b))
    return dot / (mag_a * mag_b)

THRESHOLD = 0.80

def rank(query, vectors):
    q = embed(query)
    scored = []
    for row in vectors:
        score = cosine(q, row["vector"])
        if score is None:
            continue
        scored.append({**row, "score": score, "relevant": score >= THRESHOLD})
    return sorted(scored, key=lambda r: r["score"], reverse=True)

Each vector row carries source_id, chunk_id, version, and the text it was computed from. Ranking returns scored candidates with IDs and versions attached — not bare text. That is what makes later diagnosis possible.

Knowledge check

Check your understanding

Answer this question before you continue.

A candidate's stored vector has zero magnitude. What should the ranking function do with it?
Single Choice

Focus: Apply the harness's zero-vector rule when deciding which candidates may be scored.

Ingest Versioned Sources and Run the Baseline Query

Ingest two small source records with explicit versions and deterministic chunks.

sources = {
    "doc-refund": {"version": 1, "text": "refund policy window days"},
    "doc-ship":   {"version": 1, "text": "shipping returns policy"},
}

vectors = []
cache = {}
invalidation_log = []

def ingest(source_id, version, text):
    chunk_id = f"{source_id}:v{version}:c0"
    vectors.append({
        "vector_id": f"vec-{chunk_id}",
        "source_id": source_id,
        "chunk_id": chunk_id,
        "version": version,
        "text": text,
        "vector": embed(text),
    })

for sid, rec in sources.items():
    ingest(sid, rec["version"], rec["text"])

query = "refund policy days"
print("BASELINE:", [(r["source_id"], r["version"], round(r["score"], 3), r["relevant"])
                    for r in rank(query, vectors)])

Expected output:

BASELINE: [('doc-refund', 1, 0.998, True), ('doc-ship', 1, 0.577, False)]

doc-refund ranks above threshold; doc-ship does not. Print a full state dump now — source versions, chunk IDs, vector IDs and versions, cache contents (empty), invalidation log (empty). Without a recorded before-state, a later failure is indistinguishable from a query that never worked.

Knowledge check

Check your understanding

Answer this question before you continue.

Given the tutorial's two version-1 sources and query `refund policy days`, which result matches the baseline output?
Output Prediction

Focus: Predict the baseline ranking and relevance labels from the tutorial's deterministic query and source records.

Apply a Correction and Verify the Old Vector Is Gone

A correction means replacing the affected chunk and its vector, not appending a second vector for the same logical content.

def correct(source_id, new_version, new_text):
    global vectors
    vectors = [v for v in vectors if v["source_id"] != source_id]
    ingest(source_id, new_version, new_text)
    sources[source_id] = {"version": new_version, "text": new_text}

correct("doc-refund", 2, "refund policy window days returns")
print("AFTER CORRECTION:", [(r["source_id"], r["version"], round(r["score"], 3), r["relevant"])
                            for r in rank(query, vectors)])

Expected output:

AFTER CORRECTION: [('doc-refund', 2, 0.996, True), ('doc-ship', 1, 0.577, False)]

The corrected source is retrievable. The old version is absent from the candidate set entirely — not merely outranked.

The subtle failure to watch for: both versions present, both above threshold, and the stale one winning on a slightly higher score. If you see two doc-refund rows in the output, your correction appended instead of replaced. Inspect the diff between before and after state dumps — vector IDs, versions, scores — rather than trusting the final answer text.

Knowledge check

Check your understanding

Answer this question before you continue.

After a correction, the same source appears twice in ranked results, once for each version. Which change addresses the stale-version failure?
Debugging

Focus: Diagnose and correct an implementation that leaves a superseded vector eligible after a source correction.

Apply a Deletion and Invalidate the Cache

A deleted source branches to two cleanup paths: remove its vectors from the index and invalidate its cached answer. The vector index feeds ranking, while the cache can return an answer directly; both paths are marked clear after deletion.
A clean vector index is not enough: invalidate the cache too, or it can still return a deleted answer.

Deletion is a cascade. Remove the source's chunks and vectors, invalidate the cache entry keyed by document ID, then append an invalidation log record.

def delete(source_id):
    global vectors
    removed = [v for v in vectors if v["source_id"] == source_id]
    vectors = [v for v in vectors if v["source_id"] != source_id]
    if source_id in cache:
        del cache[source_id]
    invalidation_log.append({"event": "delete", "source_id": source_id,
                             "vectors_removed": len(removed)})

# Simulate a cached answer for the source we are about to delete
cache["doc-refund"] = "Refunds are available within 30 days."

delete("doc-refund")
print("AFTER DELETE:", [(r["source_id"], r["version"], round(r["score"], 3), r["relevant"])
                        for r in rank(query, vectors)])
print("CACHE HIT:", cache.get("doc-refund"))
print("LOG:", invalidation_log)

Expected output:

AFTER DELETE: [('doc-ship', 1, 0.577, False)]
CACHE HIT: None
LOG: [{'event': 'delete', 'source_id': 'doc-refund', 'vectors_removed': 1}]

The deleted source cannot appear in results at any rank. The cache no longer returns its old answer. The invalidation log shows the event.

Why the cache is a separate layer: it can serve a deleted answer even when the vector store is perfectly clean. A cache hit after deletion is a stale-evidence path that bypasses ranking entirely. That is why deletion must propagate to every derived representation, not just the one you remember.

Knowledge check

Check your understanding

Answer this question before you continue.

After deleting `doc-refund` in the exercise, which state is expected?
Scenario Interpretation

Focus: Identify the required vector and cache effects of deleting a source in the tutorial's harness.

Debugging Cues: Locating the Stale Layer

Read the failure by symptom, not by guessing:

SymptomLikely layer
Stale text appears in ranked resultsVectors or chunk rows
Stale answer returned, ranking is cleanCache
Correct result missing after correctionCorrection step or threshold
Two versions of the same source both rankCorrection appended instead of replaced
A candidate scores 0.0 and still appearsZero-vector rule not enforced

Check IDs and versions first, scores second, text last. Text is the least informative signal — it tells you what leaked, never where.

Common mistakes I see learners make:

  • Deleting the source row only, leaving orphan chunks searchable.
  • Appending a new vector on correction instead of replacing the old one.
  • Keying the cache by query string instead of document ID, so a reworded query misses the invalidation.
  • Letting a zero vector score as a valid candidate.

One boundary worth naming: this harness is synchronous and single-process. It cannot reproduce in-flight races where a request reads a chunk that is deleted mid-query. Those races are real, but they belong to a concurrency exercise, not this one.

Extend the Drill: Tombstones and Re-Ingestion

Add a tombstone record for the deleted source_id, then re-run ingestion of the same source. Watch whether deleted content returns as a fresh chunk.

tombstones = {"doc-refund"}

def ingest_guarded(source_id, version, text):
    if source_id in tombstones:
        print(f"BLOCKED: {source_id} is tombstoned")
        return
    ingest(source_id, version, text)

ingest_guarded("doc-refund", 3, "refund policy window days")

Then add an is_current / is_deleted filter to rank and confirm it excludes superseded and deleted rows by default. Compare the two approaches: hard removal versus filtered soft state. Hard removal is simpler and cheaper; filtered soft state preserves history and makes audit trails trivial. Pick based on whether you need to answer "what did the system know last Tuesday?"

Success looks like this: the same query returns the same clean result set after re-ingestion as it did immediately after deletion.

The Rule This Exercise Earns

Treat the source of truth as authoritative. Every index, vector, and cache is a derived projection that must be explicitly updated or invalidated. Deletion propagation is not a cleanup chore — it is a correctness property of the pipeline.

Your next move: add a reconciliation check that asks two questions on a schedule. Does every active chunk point to an active source? Is any deleted or superseded chunk still searchable? Those two queries catch the majority of RAG index drift before a user ever sees a stale citation.

From there, the natural next step in this path is building retrieval evaluation that includes deletion and stale-version cases as first-class test scenarios — not as an afterthought once something breaks in production.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A deleted source is absent from the ranked results, but a user still receives its old answer. Which layer should you inspect first?
Question 1 of 2Debugging

Focus: Use observed retrieval and cache behavior to locate a stale answer in the cache rather than the vector-ranking layer.

Which pair of scheduled checks best applies the article's reconciliation guidance?
Question 2 of 2Comparison Reasoning

Focus: Select reconciliation checks that detect active chunks pointing to inactive sources and deleted or superseded chunks remaining searchable.

References

  1. RAG Security - OWASP Cheat Sheet Seriescheatsheetseries.owasp.org
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.