Skip to content
intermediate

Test LLM Routing Fallbacks Under Controlled Failures

Your fallback chain looks correct in review. Then the primary stalls at 900ms, the backup is rate-limited, and nobody can say whether the request…

Published 2026-10-03Updated 2026-10-0410 min read
Detailed texture of pebbles and sand, perfect for backgrounds and design projects.
Detailed texture of pebbles and sand, perfect for backgrounds and design projects. Photo by Atlantic Ambience on Pexels.

Your fallback chain looks correct in review. Then the primary stalls at 900ms, the backup is rate-limited, and nobody can say whether the request completed, how much it cost, or how many tokens you paid for and threw away.

That is the gap this tutorial closes. A fallback chain is a list of targets. A fallback policy is a set of decisions: when to retry the same target, when to reroute, how many attempts are allowed, and what happens on exhaustion. You cannot review a policy into correctness. You have to inject failures on purpose and measure what comes out.

We will build a small deterministic simulator — standard library only, no API keys, no network — and run two policies against the same event schedule. Then we will compare them by completion, duplicate work, latency, and cost, and pick one for a stated objective.

Why a Fallback Chain Is Not a Fallback Policy

A chain answers "where could this request go?" A policy answers "what do we do when it breaks?" Those are different artifacts, and only one of them is testable.

The distinction that matters most is retry versus fallback. A retry re-sends to the same target and fits transient blips — a momentary connection reset, a single dropped packet. A fallback reroutes to a backup after the retry budget fails to clear the condition. Blur the two and you get the classic production bug: a rate-limited primary absorbing ten retries while the user waits.

Worth simulating, because each behaves differently:

Failure classObservable signalFirst action
Connectivity errorError raised at the clientReroute permitted
Provider 5xxHTTP 500–504 from the providerReroute permitted
Request timeoutElapsed time crosses your thresholdReroute permitted
Auth failureProvider rejects credentialsReroute permitted
Rate limitingHTTP 429Retry with backoff first; reroute only if the limit persists

Four outcomes decide whether a policy is good: completion, duplicate work, latency, and cost. A chain tuned to maximize completion can quietly maximize the other three. That is the whole reason to simulate.

This assumes you already have a routing path — classification decides where a request goes. Here we only test what happens when that path breaks.

Knowledge check

Check your understanding

Answer this question before you continue.

A routed request receives HTTP 429 from its primary. Which initial action matches the policy distinctions in the article?
Scenario Interpretation

Focus: Distinguish when a rate-limited request should retry the same target versus reroute to a backup.

Define the Service Objective Before You Simulate

Write the objective as a sentence with numbers, or the comparison later produces a table instead of a decision:

Complete at least 98% of requests, keep p95 end-to-end latency under 4 seconds, keep duplicate work under 5% of requests, and hold cost per completed request under $0.004.

Three details trip people up.

Per-attempt versus per-request budgets. A timeout applied to every attempt can fail the whole request even when a fast fallback exists. If your budget is 1000ms per attempt and your fastest backup has a typical time-to-first-token above that, every attempt fails and the request dies with a 502. Set the budget above the typical latency of your fastest fallback.

What counts as complete. A returned response, or a response that also passes your output contract? A fast backup that violates your schema is not a completion. Decide before you measure.

Capability compatibility. Rerouting is only free if the backup honors the same context window, tool support, and response format. If it does not, your fallback is a second failure mode wearing a success costume.

One limit to state plainly: a simulator measures policy behavior under modeled failures. It does not measure real provider quality drift.

Build a Deterministic Failure Simulator

Determinism is the point. Same seed, same table, twice. If a failing run is not reproducible, you cannot diff it against the fix.

import random
from dataclasses import dataclass

@dataclass
class Attempt:
    target: str
    outcome: str        # "success" | "timeout" | "error" | "rate_limit"
    latency_ms: int
    tokens: int         # billable tokens, even if discarded

@dataclass
class Provider:
    name: str
    base_latency_ms: int
    cost_per_1k: float

    def call(self, event, attempt_no, rng):
        # event drives the failure; rng only adds latency noise
        if event == "timeout":
            return Attempt(self.name, "timeout", 10_000, 0)
        if event == "error":
            return Attempt(self.name, "error", 50, 0)
        if event == "rate_limit":
            return Attempt(self.name, "rate_limit", 30, 0)
        latency = self.base_latency_ms + rng.randint(-40, 40)
        tokens = 800
        return Attempt(self.name, "success", latency, tokens)

The critical design choice: failures come from an explicit event schedule, not from randomness. Randomness only perturbs latency. That way a run that fails is a run you can replay.

# One event per request index. None means healthy.
SCHEDULE = ["timeout", "error", None, "rate_limit", None, None, "error", None]

def run(policy, schedule, seed=7):
    rng = random.Random(seed)
    rows = []
    for i, event in enumerate(schedule):
        result = policy(i, event, rng)
        rows.append(result)
    return rows

Track per-request state: attempts used, targets tried, elapsed latency, tokens billed, and whether work was duplicated. Duplicate work means an attempt produced billable output that the policy then discarded — you paid for tokens you never used.

Your success criterion for this step: the same seed produces the same table twice. Print it, run it again, compare.

Knowledge check

Check your understanding

Answer this question before you continue.

A team wants each failure run to be reproducible, while still modeling latency variation. Which simulator design follows the article?
Debugging

Focus: Design simulator inputs so controlled failures can be replayed consistently while latency remains variable.

Encode Two Policies and Run the Same Events

Side-by-side lanes start with the same primary timeout. Policy A retries the primary three times before using the backup; Policy B reroutes to the backup immediately, with one backup retry.
Holding the failure event constant makes the retry tax and reroute timing easy to compare.

Now the comparison. Both policies share the same caps — max attempts, max fallbacks per request, hard per-request deadline — so the only variable is strategy.

Policy A (retry-heavy): bounded retries with exponential backoff and jitter on the primary; reroute only after the retry budget is exhausted.

Policy B (reroute-heavy): immediate reroute to the backup on the first qualifying failure, with a single retry on the backup.

def policy_a(i, event, rng):
    primary = Provider("primary", 600, 0.002)
    backup = Provider("backup", 900, 0.004)
    attempts, tokens, elapsed, dup = 0, 0, 0, False
    for n in range(3):                       # bounded retry budget
        attempts += 1
        a = primary.call(event, n, rng)
        elapsed += a.latency_ms
        tokens += a.tokens
        if a.outcome == "success":
            return dict(req=i, policy="A", outcome="success", attempts=attempts,
                        latency=elapsed, cost=tokens / 1000 * primary.cost_per_1k,
                        duplicate=False)
        if a.outcome == "timeout":
            dup = True                       # provider likely billed, we discarded
        backoff = (2 ** n) * 100 + rng.randint(0, 50)   # jitter
        elapsed += backoff
    a = backup.call(event, 0, rng)           # reroute after budget exhausted
    attempts += 1
    elapsed += a.latency_ms
    tokens += a.tokens
    return dict(req=i, policy="A",
                outcome="success" if a.outcome == "success" else "failed",
                attempts=attempts, latency=elapsed,
                cost=tokens / 1000 * backup.cost_per_1k, duplicate=dup)

Policy B is the same shape with the reroute moved to the front. Run both against SCHEDULE and read the tables side by side.

Where does each win? Policy A wins on cost when failures are rare, because it stays on the cheaper primary. Policy B wins on latency when the primary is genuinely sick, because it stops paying the retry tax. Where they tie is the interesting part — usually on healthy requests, which is most of them.

Jitter is not decoration. Synchronized retries against a strict-limit backup turn one outage into a retry storm, and the storm outlives the original failure.

Measure Duplicate Work, Not Just Completion

Completion-rate dashboards hide the bill. Duplicate work appears when a slow attempt is abandoned after the provider already did the work.

Count two things explicitly:

  • Attempts that produced billable output but were discarded.
  • Requests where two targets both generated a full response.

The failure mode to watch for: a tight timeout raises completion rate while raising duplicate spend, because the primary keeps getting cut off mid-generation. You reroute faster, you complete more requests, and you pay twice for a growing share of them.

There is a correctness edge here too. Rerouting a request that already triggered a side effect — a write, a tool call, an email — is not a cost problem. It is a duplicate-action problem. Idempotency keys or a no-reroute rule for side-effecting requests belong in the policy, not in a comment.

Decision rule: if duplicate work exceeds the ceiling in your objective, widen the timeout or reduce attempts before you add another fallback target. More targets multiply the duplicate surface.

Knowledge check

Check your understanding

Answer this question before you continue.

A primary attempt is abandoned after generating billable tokens, and the request then succeeds on a backup. What should the simulator count as duplicate work?
Misconception Check

Focus: Identify duplicate work using billed output that a policy discards rather than merely counting retries.

Inject Recovery and Watch the Policy Adapt

Most fallback tests stop at the outage. The transition back to health is where policies quietly burn money.

Add a recovery event: the primary becomes healthy again at a known request index. Then compare a policy with no health tracking against one with a circuit breaker that opens on repeated failures and closes after a probe succeeds.

Two failure modes, opposite directions:

  • Breaker open too long: traffic keeps hitting the more expensive backup after the primary recovered. You pay a premium for a problem that no longer exists.
  • Breaker closes too eagerly: the primary gets hammered again, fails again, and the cycle repeats.

The observable signals are fallback share over time and time-to-first-successful-primary-request after recovery. Plot both. A breaker that never closes shows up as a flat fallback share after the recovery index; one that closes too fast shows up as a sawtooth.

Knowledge check

Check your understanding

Answer this question before you continue.

After the primary recovers, a plot shows fallback share staying flat at a high level. Which failure mode does the article associate with this pattern?
Output Prediction

Focus: Interpret fallback-share trends to diagnose a circuit breaker that remains open after primary recovery.

Choose a Policy for Your Objective

Map each policy to the objective it satisfies:

ObjectivePolicyWhy
Latency-sensitive interactive pathEarly rerouteStops paying retry tax on a sick primary
Cost-sensitive batch pathBounded retryStays on the cheaper target through transient blips
Side-effecting requestNo reroute, or idempotentDuplicate actions are a correctness failure

Name the conditions that flip the recommendation. A fallback with materially different output quality, a strict-limit backup, or a request with side effects can each invalidate the table above.

And know when not to use aggressive fallback at all: requests that must not be duplicated, and paths where the backup cannot honor your output contract. For those, failing fast and surfacing the error beats a confident wrong answer.

Keep the simulator as a regression harness. Re-run it when timeouts, provider mix, or pricing change, and diff the tables. Record the decision and its criteria next to the config, so the next engineer sees why the policy is what it is.

Extend the Experiment

One modification deepens this without expanding scope: sweep the timeout value across a small range and plot completion against duplicate work. There is a knee in that curve — the point where more completion starts costing more duplicates than it is worth. Finding it by measurement beats arguing about it in review.

Two more if you want them. Add a third target with different latency and cost characteristics and re-run the same schedule. Or add a quality gate so a fast but contract-violating response counts as a failure and triggers the next target.

Then close the loop: compare your simulated failure classes against real traces from your routing layer. Where the model of failure is too clean — no partial responses, no slow successes, no correlated outages — your simulator is flattering you. Fix the model before you trust the policy.

Knowledge check

Final check

Finish the article by checking the ideas you just learned.

A batch service prioritizes low cost, failures are usually transient, and the primary is cheaper than the backup. Which policy is the best fit among the article's choices?
Question 1 of 2Comparison Reasoning

Focus: Select bounded retry or early reroute according to the service's latency or cost priority and the primary's health.

A request can trigger an external side effect, and the backup cannot honor the required output contract. The request has no idempotency protection. What is the safest policy direction described in the article?
Question 2 of 2Scenario Interpretation

Focus: Recognize when aggressive fallback is inappropriate because rerouting risks duplicate side effects or violates the output contract.

References

  1. LLM Failover and A/B Testing Tutorial - Inworld AIinworld.ai
  2. Provider fallbacks: Ensuring LLM availabilitywww.statsig.com
  3. What Is LLM Fallback? Meaning and How It Workswww.truefoundry.com
Practical resource

Want a more structured LLMOps path?

Use the LLMOps Practical Starter Bundle to connect RAG, evaluation, observability, and production patterns.

View the bundle
Coming soon

Large Language Models Starter Pack

A 12-chapter guide connecting LLM fundamentals with prompting, RAG, agents, tool calling, evaluation, security, and application engineering.

$9
PDF BundleLarge Language ModelsRAG and AgentsAI Engineering
  • 227-page Illustrated PDF edition
  • 12 guided LLM engineering chapters
  • Visual concept diagrams
  • Self-assessment quizzes
  • Bonus deep-dive sections
  • Prompt design, structured output, context windows & RAG pipelines
  • Agents, tool calling, prompt injection, evaluation & application lifecycles

Coming soon

Keep learning

Related tutorials

Continue with nearby topics and beginner-friendly explanations.