Building an AI Landing Zone on Azure — Part 3: One front door for many models

How a single, VNet-injected API Management instance becomes the only way to reach Azure OpenAI, Mistral OCR and Azure Maps, with policy fragments as composable building blocks and an identity per hop.

This is part 3 of a five-part series on building an AI Landing Zone on Azure. Part 1: Why every enterprise needs an AI gateway · Part 2: The platform underneath · Part 3: One front door for many models (this post) · Part 4: Metering every call · Part 5: Turning tokens into euros

One front door for many models

With the foundation from part 2 in place, this post is about the gateway itself. The job of the gateway is deceptively simple: be the only way to reach any AI model, and make every call look the same to the platform regardless of which model it goes to. Azure API Management does that job well, but only if you make a few decisions early and stick to them.

Which APIM SKU

The gateway is APIM Premium v2, and it is VNet-injected, not VNet-integrated. The distinction is a bit confusing, so let me explain the difference.

With VNet integration (available on Standard v2 and Premium v2), APIM stays publicly reachable on its own public IP; only its outbound traffic to backends goes through your VNet. It is the right choice when you want a public API gateway that can reach private backends.

With VNet injection (Premium v2 only in the v2 tiers), APIM is deployed into a subnet of your VNet and gets a private IP. There is no public IP for the gateway. Inbound and outbound both flow through the VNet, and the NSG on the APIM subnet actually controls who can reach it.

For a regulated environment the choice was obvious, which is VNet injection. The gateway is reachable only from inside the spoke and peered networks, which in practice means the workloads on the AKS cluster and developer tooling on the private network, the developer portal and management endpoints are private, and the security team can point at an NSG rule as the control. The cost is that the APIM subnet needs service endpoints and NSG rules for everything APIM's control plane talks to (Entra ID, Event Hub, Key Vault, SQL, Storage, Azure Monitor), and that took some iteration to get right. Microsoft's documentation on the differences is quite good, have a look at https://learn.microsoft.com/en-us/azure/api-management/virtual-network-concepts.

A big downside of this solution is obviously cost. Premium v2 is the most exepnsive SKU in the APIM arsenal, but in the case of a regulated financial, there is not much choice. I scaled down to 1 instance in the non-prod environment to keep costs at a "minimum".

One identity for the gateway

APIM runs with a single user-assigned managed identity. I prefer user-assigned over system-assigned, because the identity's lifecycle is independent of the resource. I can create it, assign roles, and then create the APIM instance, rather than the other way around, which makes Terraform dependency ordering much cleaner and means rebuilding APIM does not mess up role assignments.

That identity has exactly two capabilities: Cognitive Services User on the AI Services accounts (so it can call models and Content Safety), and Event Hubs Data Sender on the usage Event Hub (so it can emit usage events, the subject of part 4). It cannot read from Event Hub, cannot touch Cosmos, cannot read Key Vault secrets it does not need. If the gateway were somehow made to misbehave, that list is the full extent of the damage.

The identity's client ID is stored as an APIM named value, along with the audience for AI Services. Policy fragments reference those named values, which means the same policy XML works unchanged across environments.

Policy fragments as the unit of reuse

The single best decision in the gateway design was to build almost nothing directly into an API's policy and instead compose each API from policy fragments. A fragment is a reusable chunk of policy XML that you include by ID. The pattern is borrowed from Microsoft's AI Hub Gateway solution accelerator, which I used as the reference architecture.

The library ended up looking like this:

Fragment What it does
aad-auth Optionally validates an Entra ID bearer token from the caller, toggled by a named value
strip-caller-credentials Deletes the caller's api-key and subscription key headers so they are never forwarded upstream
backend-mi-auth Acquires a token with the gateway's managed identity and sets it as the upstream Authorization header
content-safety-aoai-inbound Sends the user prompt to Content Safety (prompt shield + harm analysis) in audit or enforce mode
content-safety-mistral-outbound Same idea for OCR output, scanning the extracted text on the way back
throttling-events Emits a custom metric whenever a backend answers 429 (rate limiting), dimensioned by backend and subscriptiion
rate-limit-logging Logs rate-limit outcomes for the operations team
usage-ingestion-aoai / usage-ingestion-mistral Emit the usage record to Event Hub.

An API's policy is then mostly a list of includes:

<inbound>
    <base />
    <include-fragment fragment-id="aad-auth" />
    <include-fragment fragment-id="strip-caller-credentials" />
    <include-fragment fragment-id="content-safety-aoai-inbound" />
    <include-fragment fragment-id="backend-mi-auth" />
    <azure-openai-token-limit ... />
    <azure-openai-emit-token-metric ... />
    <set-backend-service backend-id="azure-openai-backend" />
</inbound>

Two things make this work well in practice. First, the fragments are managed by Terraform from XML files in the repository, so a change to how caller credentials are stripped is one file, one merge request, and it takes effect on every API at once. Second, behaviour that differs per environment or per rollout stage, like whether Entra authentication is required or whether Content Safety blocks or only audits, is driven by named values, not by editing policy. Switching Content Safety from audit to enforce in production is a Terraform variable change.

The credential handover

The most important sequence in the gateway is what happens to authentication between the caller and the model, and it is the reason no application in the eco-system holds a model key.

The caller authenticates to APIM. Depending on the product this is an APIM subscription key, an Entra ID token, or both. APIM validates that, and then the strip-caller-credentials fragment deletes every credential header the caller sent. The backend-mi-auth fragment then asks Entra ID for a token for the AI Services audience using APIM's own managed identity and puts it in the Authorization header. The model backend sees a request from APIM's identity and nothing else.

The consequence is that the model deployments have local key authentication effectively unused (result being they cannot be used directly), the consumer's identity is known to APIM but not to the model, and swapping the backend from one Azure OpenAI account to another, or from one model version to another, is a change to an APIM backend definition that no consumer sees. I used that twice during the build when a preview model version turned out to be unstable and had to be rolled back.

Many backends, one shape

Behind the gateway there are currently six backends. Azure OpenAI in the gpt-5 family serves three distinct use cases, each on its own AI Services account: a document-extraction workload and a summarisation workload that use Chat Completions, and a developer coding-assistant workload that uses the newer Responses API through the unified /openai/v1 endpoint. Mistral Document AI, served through an AI Foundry serverless endpoint, handles OCR, and there are two versions of it live so that consumers can be migrated one at a time. Content Safety is itself a backend, called from policy.

Each backend is exposed as its own API with its own path (/openai, /chatgpt, /mistral, and so on) and its own policy, but because the policies are built from the same fragments, every one of them has the same authentication handover, the same safety scanning where it applies, and the same usage emission.

Where the backends differ, the policies differ in the obvious places. The OpenAI APIs use the built-in azure-openai-token-limit policy for per-subscription tokens-per-minute limits and azure-openai-emit-token-metric to push token counts to Application Insights. The Mistral API is billed per page, not per token, so it uses generic rate and quota limits per subscription instead, and its usage record carries pages processed and document size rather than token counts.

Streaming or not, you must choose

There is one design decision in the gateway that I want to point out here because it comes back in a big way in part 4. Chat completions can be streamed, with the response arriving as a stream of server-sent events, or returned as a single JSON body. Whether APIM buffers the response is set per route in the backend section of the policy, and it has consequences for both the caller experience and for what the gateway can observe.

If APIM buffers the response, it can read the whole body on the way out, including the usage block with exact token counts, but the caller only receives the answer once it is complete. If APIM does not buffer, the caller gets token-by-token output, but the outbound policy cannot read the body and the exact token counts are gone.

I made it a per-route choice. The coding-assistant route, where cost accuracy matters more than a few seconds of latency, buffers and gets exact numbers. Interactive chat routes stream live and fall back to an estimate. How that fallback works is explained in part 4.

Products as the unit of chargeback

APIM products are what turn "a gateway" into "a platform with tenants". Each consuming team or use case gets a product, such as Code Assist, Summariser, or a POC product for experiments, and each product bundles the APIs that team is allowed to call. Subscriptions live under products, and every usage record the gateway emits carries both the subscription and the product name.

That is the whole basis for chargeback in part 5. Finance does not need to know which model a team used; they need to know that the Summariser product cost this much in September. The product is the cost centre. Approval-required products give the platform team a gate for new consumers, and the developer portal (private, of course) gives teams self-service for keys and documentation.

Observability at the gateway

Every API's diagnostics go to Application Insights, and the policies emit custom metrics: token consumption per subscription, product and deployment; OCR requests per subscription; throttling events per backend. Those are what the operations team alerts on. When a backend starts returning 429 because an Azure OpenAI deployment ran out of capacity, the throttling metric spikes with the backend name as a dimension, and the platform team knows which quota to raise before anyone opens a ticket.

Next: the pipeline that watches every call

The gateway now does what it needs to do on the request path. Part 4 is about what happens after the response: the usage-emission fragment, the Event Hub, the Function App and Cosmos DB, and the bugs I tackled, each of which silently produced zero records while everything looked healthy.

As always, if you are working on something similar, or want to see how this gateway pattern could land in your own Azure estate, connect with me on LinkedIn or reach out via ascode.nl.

Written by

Erik Christiaans

Independent cloud and AI platform architect in Velsen, NL. Twenty years of Azure, AWS, Entra ID and Kubernetes for insurers, banks and other regulated enterprises - and the control mappings that make those platforms defensible.

Book a 30-minute call