Model Routing & Fallback
This is where running both vendors pays off. Route each task to the model that fits — the strong reasoner for review, the cheap one for generation — and fall back to the other provider when one has an outage or rate-limits you. One small router, used by every AI job in your pipeline, turns two vendors from a complication into a resilience and cost advantage.
Learning objectives
- Route a task to the right model by difficulty and cost.
- Fall back to a second provider when the first fails.
- Wrap both vendors behind one uniform
complete()interface. - Reason about A/B comparison and when multi-vendor is worth the complexity.
ANTHROPIC_API_KEY and OPENAI_API_KEY.1 · Why route at all essential
By now you've seen the split in action: gl2 pointed review at a strong reasoning model; gl3 pointed changelog generation at a cheap one. Routing is just making that choice systematic — a single place that decides, per task, which model to call. And once you're deliberately choosing a model, calling a second vendor as a backup is a tiny extra step that buys real resilience.
Two distinct motivations, often conflated. Routing is about fit: send each task to the cheapest model that does it well (hard reasoning → flagship; easy generation → mini). Fallback is about resilience: if your primary provider is down, rate-limiting, or erroring, retry the same task on a different provider so the pipeline doesn't fail. Routing optimizes cost and quality; fallback optimizes uptime. A mature pipeline wants both.
The common mistake is scattering model choice across every job — claude-opus hardcoded here, gpt-5.5 there — so changing strategy means editing ten files, and there's no fallback anywhere. The fix is one small router module every job imports: the policy lives in one place, and fallback is built in once.
- The blue task box enters the router, which decides by difficulty which model to call.
- The solid edge to the green primary model is the normal path (routing = fit).
- The dashed edge to the amber fallback vendor is taken only on error/outage (fallback = resilience).
In short: routing and fallback are two jobs of one small module — pick the right model, and have a backup when it's unavailable.
2 · One uniform interface over both vendors intermediate
The enabling trick is a single complete() function that hides which vendor is being called. Each job asks for text; the router decides the vendor and handles the SDK differences. Here are the two vendor adapters it wraps — identical signature, different SDK.
The vendor adapter — in Claude or OpenAI
Each adapter has the same signature — (system, user) -> text — so the router can call either interchangeably. This uniform shape is what makes routing and fallback possible. Toggle the tab to see each side.
adapters.pyfrom anthropic import Anthropic
_anthropic = Anthropic()
def call_claude(system, user, model="claude-opus-4-8", effort=None):
kw = {"output_config": {"effort": effort}} if effort else {}
resp = _anthropic.messages.create(
model=model, max_tokens=1500, system=system,
messages=[{"role":"user","content": user}], **kw)
return next(b.text for b in resp.content if b.type=="text")
adapters.pyfrom openai import OpenAI
_openai = OpenAI()
def call_openai(system, user, model="gpt-5.5", effort=None):
kw = {"reasoning": {"effort": effort}} if effort else {}
resp = _openai.responses.create(
model=model, max_output_tokens=1500,
instructions=system, input=user, **kw)
return resp.output_text
- Both adapters take the same arguments — a
systemprompt, theusertext, an optional model and effort — and return a plain string. - Inside, each speaks its own SDK: Claude's
messages.create+ content blocks +output_config.effort; OpenAI'sresponses.create+output_text+reasoning.effort. - Because the signature is identical, the router (next) can treat them as interchangeable — the whole point of the uniform interface.
Try this: this is the same dual-vendor pattern as the rest of the course, now formalized into two functions with one signature — the seam that routing and fallback are built on.
3 · The router: pick by task, fall back on failure advanced
Now the module every job imports. complete() routes by a task label to a (vendor, model, effort) choice, and wraps the call so that if the primary vendor raises, it retries on the other.
router.pyimport anthropic, openai
from adapters import call_claude, call_openai
# routing policy: task -> primary, then fallback
ROUTES = {
"review": [(call_claude, "claude-opus-4-8", "high"), # reasoning-heavy
(call_openai, "gpt-5.5", "high")], # fallback
"generate": [(call_openai, "gpt-5.5", "low"), # cheap/easy
(call_claude, "claude-haiku-4-5", None)], # fallback
}
TRANSIENT = (anthropic.APIError, openai.APIError)
def complete(task, system, user):
for fn, model, effort in ROUTES[task]:
try:
return fn(system, user, model=model, effort=effort)
except TRANSIENT as e:
print(f"{fn.__name__} failed ({e}); trying fallback")
raise RuntimeError(f"all providers failed for task={task}")
ROUTESis the whole policy in one place: each task maps to an ordered list of(adapter, model, effort)— primary first, fallback second.complete(task, …)tries each in order; on a transient API error from one vendor it logs and moves to the next — that's fallback.- Review routes to the strong model at high effort (fit); generation routes to the cheap one at low effort (fit); each has the other vendor as backup (resilience). A job just calls
complete("review", sys, diff)and never knows which vendor answered.
Try this: add a third task ("summarize") to ROUTES without touching any job — that's the payoff of centralizing the policy. Changing strategy is a one-file edit.
4 · A/B comparison advanced
Running both vendors also lets you compare them on your real workload. An A/B job sends the same task to both and logs the two outputs (and their cost/latency) so you can judge which model actually does your work better — not which benchmarks higher in general. This is how you make model choice evidence-based (the eval discipline from Ch 5), rather than guessing. In CI you'd run A/B on a sample of tasks, not every one, and feed the comparison into which model becomes the ROUTES primary.
ROUTES table should be set by measurement: periodically A/B both vendors on a sample of your real tasks, compare quality/cost/latency, and promote the winner to primary. That closes the loop — routing isn't a one-time guess, it's a policy you tune as the models (and their prices) change.5 · Tech-lead — is multi-vendor worth it? tech-lead
Two vendors is real complexity: two sets of keys, two SDKs, two bills, two sets of quirks. Be honest about when it's worth it. For a small team or a low-stakes pipeline, one well-chosen vendor behind the complete() interface (so you could add a second later) is often the right call — don't pay the multi-vendor tax for resilience you don't need yet. The uniform interface is the cheap insurance: build it from day one so adding a fallback vendor is a ROUTES edit, not a rewrite. Reach for genuine multi-vendor when uptime is critical (a provider outage would hurt), when cost at scale justifies routing cheap tasks to the cheapest provider, or when you're contractually hedging against one vendor. The portability theme from the whole OpenAI track lands here: keep model choice in config and behind an interface, and multi-vendor becomes a dial you turn when the need arrives.
🪜 Practice ladder beginner → industry
- Beginner: write the two adapters (Lab G4.1) with one shared signature.
- Easy: build the
ROUTEStable and acomplete()that picks primary by task. - Core: add fallback — make
complete()try the second provider on an API error. - Stretch: add a new task to
ROUTESand use it from a job without editing the job. - Hard: write an A/B job that runs both vendors on a sample and logs quality/cost/latency.
- Industry: write the decision memo: does your pipeline need multi-vendor now, or one vendor behind the interface?
✓ Checkpoint — you can move on when you can…
- Wrap both vendors behind one uniform
complete()interface. - Route a task to a model by difficulty and cost.
- Fall back to a second provider on failure.
- Decide when multi-vendor is actually worth the complexity.
Knowledge check check yourself
What's the difference between routing and fallback, and why must fallback cross vendors?
Show answer
Why wrap both vendors behind one complete() interface, and how should the ROUTES policy be set?
Show answer
(system, user) → text signature for both vendors) means jobs don't know or care which vendor answered, so routing and fallback live in one module and changing strategy is a one-file edit instead of touching every job. The ROUTES policy should be set by evidence: periodically A/B both vendors on a sample of your real tasks, compare quality/cost/latency, and promote the winner to primary — routing is a tuned policy, not a one-time guess.