How inference request routing works
The setup combines the Gateway API Inference Extension’s InferencePool contract with Envoy AI Gateway. Four components have separate responsibilities. The picker that satisfies the contract is a pluggable implementation: the generated bundle installs the upstream Gateway API Inference Extension endpoint picker, while the author’s cluster runs the llm-d router build of the same protocol, which scores with a predicted-latency model and an exact prefix index instead of approximate scoring. Both speak ext-proc on 9002 to the same Envoy data plane; the routing guide records what only the router build does.
Why add these components
Section titled “Why add these components”You can expose one vLLM server with a Kubernetes Service. That is the simpler choice when callers already know the backend and there is only one replica. Envoy AI Gateway and EPP become useful when the cluster must route by model, apply one set of traffic policy, or choose among several serving replicas without putting that logic in every client.
Envoy AI Gateway provides the client-facing routing layer. It understands model-oriented routes, so a request can name a served model while the Gateway maps it to an InferencePool. It also gives operators one place for route timeouts, request-size limits and, when configured separately, edge controls such as TLS and authentication. Clients keep using an OpenAI-compatible endpoint instead of discovering pod or Service addresses.
EPP provides the endpoint-selection layer. Kubernetes load balancing normally treats ready endpoints as interchangeable. LLM replicas are not interchangeable under load: one may have a shorter queue, enough KV-cache capacity, or a useful prefix already cached. EPP receives metadata about the eligible endpoints, scores those signals, and returns a rank-zero endpoint to Envoy. Envoy still carries the request body directly to vLLM; EPP is not a proxy in the data path.
How much of this is useful depends on how many serving groups exist. A recipe applied once has one TP=2 group and therefore one eligible rank-zero endpoint, so EPP cannot improve placement; it still establishes the routing contract that supports additional replicas later. The author’s cluster runs the Qwen recipe with replicas: 2, which puts two independent rank-zero endpoints in one pool and makes the selection real. If the deployment will remain a single group and needs no model-aware gateway policy, use the rank-zero Service and omit this routing stack.
Open the architecture figure by itself.
Request path
Section titled “Request path”- A client sends an OpenAI-compatible request with
x-ai-eg-model: deepseek-v4-flash-visionto Gateway port 10080. The header value is the engine’s--served-model-name, which is what the generator writes into both the engine flag and the route match. - The AI route matches that exact model header and refers to the
InferencePool. A deployment that also serves other models carries a second, unmatched rule as the default route, and a native Anthropic/v1/messagesrequest needs its ownHTTPRoute, because it is not an OpenAI request and never matches the header rule. - Envoy sends request metadata to EPP over gRPC port 9002. EPP evaluates only endpoints admitted by the pool.
- EPP returns its endpoint choice.
FailClosemakes selection fail if EPP cannot decide; it does not silently bypass policy. - Envoy sends the original inference payload directly to the chosen rank-zero API on port 8888. The payload is not proxied through EPP.
- Rank zero coordinates tensor-parallel execution with rank one. Only rank zero returns the HTTP response.
The pool selector admits LWS worker index 0 only. Rank one is never an HTTP backend even though it is essential to inference.
Collective path
Section titled “Collective path”Rank zero and rank one exchange distributed-engine rendezvous and GPU collective traffic. Kubernetes control, DNS, Gateway/EPP and the model API use the Cilium primary network. The configured high-volume collective path uses the Multus/macvlan attachment and host RDMA devices.
These networks have different owners and policy boundaries. Cilium policy does not automatically protect the secondary macvlan interface, and Multus does not replace Cilium as the pod’s primary network.
Why the debug Service exists
Section titled “Why the debug Service exists”The rank-zero Service gives operators a stable way to port-forward directly to /v1/models or /v1/chat/completions. It proves the serving group independently of the routing stack. It also bypasses EPP, model-header matching and Gateway policy, so it cannot prove scheduled routing.
What changes with replicas
Section titled “What changes with replicas”With one TP group there is one eligible API endpoint, so EPP still enforces the protocol and failure behavior while queue/KV/prefix-cache scoring has no alternative replica to choose. Add an independent TP group and those signals start to differentiate endpoints. The 2/2/3 weights the generator ships describe a configuration validated locally, not a universal optimum, and the llm-d router in the author’s cluster does not use numeric weights at all: it filters on an exact prefix-cache affinity threshold, scores a predicted time-to-first-token, and picks randomly among the survivors.
Security boundary
Section titled “Security boundary”EPP gRPC is plaintext inside the cluster. Namespace/network isolation must restrict it to the Gateway data plane. Port-forward is used for initial acceptance so the tutorial does not imply an authenticated public service. Production exposure requires a separately accepted TLS, authentication and authorization boundary.