Azure OpenAI & Running Anywhere
The same code that calls OpenAI directly can run against Azure OpenAI — the enterprise deployment where the models live in your cloud tenancy, under your compliance and networking rules. This chapter covers the one client swap that gets you there, why enterprises need it, and the broader "same SDK, different endpoint" pattern.
Learning objectives
- Explain why an enterprise runs OpenAI models via Azure rather than the public API.
- Swap
OpenAI()forAzureOpenAI(...)with no other code change. - Understand deployments, API versions, and Entra ID auth.
- Reason about the general "same SDK, different endpoint" portability pattern.
pip install openai.1 · Why run OpenAI through Azure essential
For a lot of companies, "just call api.openai.com" is a non-starter — and not for technical reasons. A regulated bank, a hospital, a government contractor often cannot send data to a third-party API over the public internet: they have data-residency rules (data must stay in a region), network rules (no egress to the open web), procurement rules (vendors must be on an approved cloud), and compliance certifications they must inherit. Azure OpenAI exists to satisfy exactly those constraints — the same OpenAI models, but deployed inside your Azure tenancy, in your chosen region, behind your private networking, under Microsoft's enterprise agreements.
The beautiful part for you as the engineer: almost none of your code changes. The OpenAI Python SDK ships an AzureOpenAI client that speaks the same methods — you build the client differently (endpoint, deployment, API version instead of a bare key), and from there client.responses.create(...) (or chat.completions) works the same. The application logic is portable; only the door changes.
The common mistake is assuming Azure OpenAI is a different product you must re-learn. It isn't — it's the same models and SDK with an enterprise front door. Learn the client-construction difference and the deployment-vs-model naming, and everything else you know carries over.
OpenAI() hits the public API; AzureOpenAI(...) hits models deployed in your Azure tenancy under your compliance and networking — and the responses.create calls after that are identical.
- The blue box is the direct path — the public API with a key, what you've used so far.
- The green box is Azure — the same models inside your cloud tenancy, constructed with an endpoint, a deployment name, and an API version.
- Everything downstream of the client is identical; the choice is purely about where the request goes and under whose rules.
In short: enterprises pick the green door for compliance; your app code doesn't notice the difference.
2 · The client swap essential
Here's the whole change — construct AzureOpenAI instead of OpenAI. Note that on Azure you call a deployment name (what you named your model deployment in the Azure portal), not the bare model id.
direct_vs_azure.py# --- direct (public API) ---
from openai import OpenAI
client = OpenAI() # OPENAI_API_KEY from env
# --- Azure OpenAI (same methods after this) ---
from openai import AzureOpenAI
client = AzureOpenAI(
azure_endpoint="https://my-resource.openai.azure.com",
api_version="2024-10-01-preview", # pin the API version
# api_key from AZURE_OPENAI_API_KEY, or use Entra ID (below)
)
# identical from here — on Azure, model = your DEPLOYMENT name
resp = client.responses.create(model="my-gpt-deployment", input="Hello!")
print(resp.output_text)
AzureOpenAI(...)replacesOpenAI(). It needsazure_endpoint(your resource URL) andapi_version(Azure pins behavior to a dated version); the key comes fromAZURE_OPENAI_API_KEYor Entra ID.- The
modelargument on Azure is your deployment name — the label you gave the deployed model in the portal — not the public model id. - Everything else —
responses.create,output_text, usage, tools — is identical, which is the whole point.
Try this: factor client construction into one make_client() function behind an env flag, so the same app runs against the public API in dev and Azure in production with zero logic changes.
model="gpt-5.5"; on Azure model="<your-deployment-name>". The deployment is your named instance of a model in your resource — it may or may not match the underlying model id.3 · API versions & Entra ID auth intermediate
Two Azure-specific details matter in production. API version: Azure pins API behavior to a dated api_version string, so upgrades are deliberate — you bump the version when you're ready, rather than being moved automatically. Auth: beyond a static key, Azure supports Entra ID (formerly Azure AD) — token-based auth via azure_ad_token or an azure_ad_token_provider callback — which is how enterprises avoid long-lived keys and use managed identities instead.
entra_auth.pyfrom openai import AzureOpenAI
from azure.identity import DefaultAzureCredential, get_bearer_token_provider
# managed identity / Entra ID instead of a static key
token_provider = get_bearer_token_provider(
DefaultAzureCredential(),
"https://cognitiveservices.azure.com/.default")
client = AzureOpenAI(
azure_endpoint="https://my-resource.openai.azure.com",
api_version="2024-10-01-preview",
azure_ad_token_provider=token_provider, # no static key in code
)
4 · The broader "run anywhere" pattern advanced
Azure is the concrete case of a general pattern: the OpenAI SDK is a client, and clients can point at different endpoints. The SDK lets you set a custom base_url, which is why the same code can talk to any OpenAI-compatible endpoint — a local inference server (vLLM, Ollama expose an OpenAI-compatible API), a gateway/proxy that adds logging and routing, or a third-party host. The application logic — your prompts, your tools, your parsing — rides on top unchanged.
This is the portability dividend: by building against the OpenAI SDK's interface, you keep the freedom to move where the model runs — public API, Azure, a private cluster, a self-hosted open model behind a compatible shim — without rewriting the app. Combined with keeping the model id in config (ox1), you've decoupled your code from both which model and where it runs.
base_url at them and your code runs against a local open model. The interface is the portability layer.5 · Tech-lead — a portability strategy tech-lead
The lead-level takeaway: treat "which provider/endpoint" as a configuration boundary, not something baked through your codebase. One make_client() factory, driven by environment, returns the right client (public OpenAI, AzureOpenAI, or a custom base_url); the model/deployment id is a config value; everything above that line is provider-agnostic application code. That discipline is what lets a company start on the public API, move to Azure for compliance, and keep a self-hosted fallback — all without a rewrite, and all testable by swapping one env var.
🪜 Practice ladder beginner → industry
- Beginner: read Lab OP6.1 and name every argument that differs between
OpenAI()andAzureOpenAI(). - Easy: write a
make_client()that returns public or Azure based on an env flag. - Core: explain why
modelis a deployment name on Azure and where that name comes from. - Stretch: swap static-key auth for the Entra ID token-provider pattern.
- Hard: point the SDK's
base_urlat a local OpenAI-compatible server and run the same code. - Industry: design the config boundary for an app that must run public in dev, Azure in prod, and self-hosted as fallback.
✓ Checkpoint — you can move on when you can…
- Explain why enterprises run OpenAI models via Azure.
- Swap
OpenAI()forAzureOpenAI(...)and name the extra arguments. - Use an API version and Entra ID auth.
- Describe the general same-SDK-different-endpoint portability pattern.
Knowledge check check yourself
Why do enterprises use Azure OpenAI, and what actually changes in your code to target it?
Show answer
OpenAI() for AzureOpenAI(azure_endpoint=…, api_version=…) and pass a deployment name as model; everything after that — responses.create, output_text, usage, tools — is identical.What is the general "run anywhere" pattern, and how does it relate to keeping the model id in config?
Show answer
base_url), so the same application code can target the public API, Azure, a gateway, or a local OpenAI-compatible server (vLLM/Ollama) without rewriting logic. Combined with keeping the model/deployment id in one config value, you decouple your code from both which model runs and where it runs — a factory like make_client() driven by environment makes the provider a configuration boundary.