AI EngineeringZero to ProductionHome·About·What’s new·Contact
GitLab CI/CD with AI · Part 4

Model Routing & Fallback

This is where running both vendors pays off. Route each task to the model that fits — the strong reasoner for review, the cheap one for generation — and fall back to the other provider when one has an outage or rate-limits you. One small router, used by every AI job in your pipeline, turns two vendors from a complication into a resilience and cost advantage.

⏱️ ~1.5 hours🔀 Routing & fallback🎯 Advanced→Tech-lead

Learning objectives

  • Route a task to the right model by difficulty and cost.
  • Fall back to a second provider when the first fails.
  • Wrap both vendors behind one uniform complete() interface.
  • Reason about A/B comparison and when multi-vendor is worth the complexity.
The router runs wherever your CI jobs run and needs both vendors' API keys as CI secrets (gl5). The code is complete; adapt it into your pipeline and provide ANTHROPIC_API_KEY and OPENAI_API_KEY.

1 · Why route at all essential

By now you've seen the split in action: gl2 pointed review at a strong reasoning model; gl3 pointed changelog generation at a cheap one. Routing is just making that choice systematic — a single place that decides, per task, which model to call. And once you're deliberately choosing a model, calling a second vendor as a backup is a tiny extra step that buys real resilience.

Two distinct motivations, often conflated. Routing is about fit: send each task to the cheapest model that does it well (hard reasoning → flagship; easy generation → mini). Fallback is about resilience: if your primary provider is down, rate-limiting, or erroring, retry the same task on a different provider so the pipeline doesn't fail. Routing optimizes cost and quality; fallback optimizes uptime. A mature pipeline wants both.

The common mistake is scattering model choice across every job — claude-opus hardcoded here, gpt-5.5 there — so changing strategy means editing ten files, and there's no fallback anywhere. The fix is one small router module every job imports: the policy lives in one place, and fallback is built in once.

one router: pick by task (routing) + retry on the other vendor (fallback) taskreview? generate? routerpick by difficulty primary modele.g. Claude review fallback vendoron error / outage routing = fit (cost/quality) · fallback = resilience (uptime) Route by fit, fall back for uptime. One router picks the model per task (routing) and retries on the other vendor if the primary fails (fallback) — so every AI job is both cost-tuned and outage-resilient.
🗺️ How to read this diagram
  • The blue task box enters the router, which decides by difficulty which model to call.
  • The solid edge to the green primary model is the normal path (routing = fit).
  • The dashed edge to the amber fallback vendor is taken only on error/outage (fallback = resilience).

In short: routing and fallback are two jobs of one small module — pick the right model, and have a backup when it's unavailable.

2 · One uniform interface over both vendors intermediate

The enabling trick is a single complete() function that hides which vendor is being called. Each job asks for text; the router decides the vendor and handles the SDK differences. Here are the two vendor adapters it wraps — identical signature, different SDK.

Lab G4.1

The vendor adapter — in Claude or OpenAI

Each adapter has the same signature — (system, user) -> text — so the router can call either interchangeably. This uniform shape is what makes routing and fallback possible. Toggle the tab to see each side.

adapters.pyfrom anthropic import Anthropic
_anthropic = Anthropic()

def call_claude(system, user, model="claude-opus-4-8", effort=None):
    kw = {"output_config": {"effort": effort}} if effort else {}
    resp = _anthropic.messages.create(
        model=model, max_tokens=1500, system=system,
        messages=[{"role":"user","content": user}], **kw)
    return next(b.text for b in resp.content if b.type=="text")
▶ How this works
  1. Both adapters take the same arguments — a system prompt, the user text, an optional model and effort — and return a plain string.
  2. Inside, each speaks its own SDK: Claude's messages.create + content blocks + output_config.effort; OpenAI's responses.create + output_text + reasoning.effort.
  3. Because the signature is identical, the router (next) can treat them as interchangeable — the whole point of the uniform interface.

Try this: this is the same dual-vendor pattern as the rest of the course, now formalized into two functions with one signature — the seam that routing and fallback are built on.

3 · The router: pick by task, fall back on failure advanced

Now the module every job imports. complete() routes by a task label to a (vendor, model, effort) choice, and wraps the call so that if the primary vendor raises, it retries on the other.

Lab G4.2
router.pyimport anthropic, openai
from adapters import call_claude, call_openai

# routing policy: task -> primary, then fallback
ROUTES = {
    "review":   [(call_claude, "claude-opus-4-8", "high"),   # reasoning-heavy
                 (call_openai, "gpt-5.5", "high")],         # fallback
    "generate": [(call_openai, "gpt-5.5", "low"),          # cheap/easy
                 (call_claude, "claude-haiku-4-5", None)],   # fallback
}
TRANSIENT = (anthropic.APIError, openai.APIError)

def complete(task, system, user):
    for fn, model, effort in ROUTES[task]:
        try:
            return fn(system, user, model=model, effort=effort)
        except TRANSIENT as e:
            print(f"{fn.__name__} failed ({e}); trying fallback")
    raise RuntimeError(f"all providers failed for task={task}")
▶ How this works
  1. ROUTES is the whole policy in one place: each task maps to an ordered list of (adapter, model, effort) — primary first, fallback second.
  2. complete(task, …) tries each in order; on a transient API error from one vendor it logs and moves to the next — that's fallback.
  3. Review routes to the strong model at high effort (fit); generation routes to the cheap one at low effort (fit); each has the other vendor as backup (resilience). A job just calls complete("review", sys, diff) and never knows which vendor answered.

Try this: add a third task ("summarize") to ROUTES without touching any job — that's the payoff of centralizing the policy. Changing strategy is a one-file edit.

Fallback needs genuinely independent providersFalling back from one model to another at the same provider doesn't help when the provider itself is down or rate-limiting you. Cross-vendor fallback (Claude ↔ OpenAI) is what buys real uptime — which is exactly why this track is dual-vendor. Two providers is a resilience asset, not just a cost lever.

4 · A/B comparison advanced

Running both vendors also lets you compare them on your real workload. An A/B job sends the same task to both and logs the two outputs (and their cost/latency) so you can judge which model actually does your work better — not which benchmarks higher in general. This is how you make model choice evidence-based (the eval discipline from Ch 5), rather than guessing. In CI you'd run A/B on a sample of tasks, not every one, and feed the comparison into which model becomes the ROUTES primary.

Route by evidence, not vibesThe ROUTES table should be set by measurement: periodically A/B both vendors on a sample of your real tasks, compare quality/cost/latency, and promote the winner to primary. That closes the loop — routing isn't a one-time guess, it's a policy you tune as the models (and their prices) change.

5 · Tech-lead — is multi-vendor worth it? tech-lead

Two vendors is real complexity: two sets of keys, two SDKs, two bills, two sets of quirks. Be honest about when it's worth it. For a small team or a low-stakes pipeline, one well-chosen vendor behind the complete() interface (so you could add a second later) is often the right call — don't pay the multi-vendor tax for resilience you don't need yet. The uniform interface is the cheap insurance: build it from day one so adding a fallback vendor is a ROUTES edit, not a rewrite. Reach for genuine multi-vendor when uptime is critical (a provider outage would hurt), when cost at scale justifies routing cheap tasks to the cheapest provider, or when you're contractually hedging against one vendor. The portability theme from the whole OpenAI track lands here: keep model choice in config and behind an interface, and multi-vendor becomes a dial you turn when the need arrives.

🪜 Practice ladder beginner → industry

  1. Beginner: write the two adapters (Lab G4.1) with one shared signature.
  2. Easy: build the ROUTES table and a complete() that picks primary by task.
  3. Core: add fallback — make complete() try the second provider on an API error.
  4. Stretch: add a new task to ROUTES and use it from a job without editing the job.
  5. Hard: write an A/B job that runs both vendors on a sample and logs quality/cost/latency.
  6. Industry: write the decision memo: does your pipeline need multi-vendor now, or one vendor behind the interface?

✓ Checkpoint — you can move on when you can…

  • Wrap both vendors behind one uniform complete() interface.
  • Route a task to a model by difficulty and cost.
  • Fall back to a second provider on failure.
  • Decide when multi-vendor is actually worth the complexity.

Knowledge check check yourself

✓ Knowledge check

What's the difference between routing and fallback, and why must fallback cross vendors?

Show answer
Routing is about fit — sending each task to the cheapest model that does it well (hard reasoning → strong model, easy generation → cheap model). Fallback is about resilience — retrying a failed task on a different provider so the pipeline doesn't break. Fallback must cross vendors (Claude ↔ OpenAI) because falling back to another model at the same provider doesn't help when that provider itself is down or rate-limiting you.
✓ Knowledge check

Why wrap both vendors behind one complete() interface, and how should the ROUTES policy be set?

Show answer
A uniform interface (same (system, user) → text signature for both vendors) means jobs don't know or care which vendor answered, so routing and fallback live in one module and changing strategy is a one-file edit instead of touching every job. The ROUTES policy should be set by evidence: periodically A/B both vendors on a sample of your real tasks, compare quality/cost/latency, and promote the winner to primary — routing is a tuned policy, not a one-time guess.
© 2026 studybydoing.in · AI Engineering: Zero to Production · All rights reserved. · About · Privacy Policy · Terms · Contact
Educational content, provided as-is and without warranty. Code samples are examples — review, test, and adapt them before using in production. See the Terms of Use & Disclaimer. Use at your own risk.
© studybydoing.in