llm-d request path Envoy AI Gateway asks the endpoint picker for a rank-zero destination, then sends the inference payload directly to that endpoint while rank one participates only through tensor-parallel collectives. ARCHITECTURE / GATEWAY API INFERENCE EXTENSION Inference request and endpoint-selection paths LWS manages the model group; EPP selects an eligible rank-zero endpoint. KUBERNETES CLUSTER HTTP :10080 PICK gRPC SELECT ENDPOINT INFERENCE PAYLOAD METRICS :8888 TP=2 CLIENT OpenAI client x-ai-eg-model header GATEWAY Envoy AI Gateway Gateway + AIGatewayRoute EPP Endpoint picker scores queue / KV / prefix failureMode: FailClose payload never enters EPP INFERENCEPOOL Eligible API endpoints selector: worker-index=0 WATCH POD vLLM rank 0 API :8888 / selected endpoint LeaderWorkerSet member POD vLLM rank 1 headless never pooled Debug Service bypasses scheduled routing LEGEND REQUEST PAYLOAD CONTROL / METRICS GPU COLLECTIVE

Endpoint selection is out of band

The Gateway asks EPP over gRPC. EPP scores endpoints represented by the InferencePool and returns a choice; it does not proxy the image or text payload.

Only rank zero is pooled

The pool selector admits worker-index=0, so one TP=2 group contributes one eligible endpoint and scoring has nothing to choose between. Add an independent group and the same contract places work across two rank-zero APIs; rank one never leaves the group, it only contributes compute through collectives.