Skip to content

Install serving and routing controllers

Experienced operators can enter here if their cluster already satisfies the host, GPU, primary-network and RDMA contracts. Fresh k3s readers should complete the linear tutorial first.

Run from a workstation with cluster-admin credentials. Stop on the first failed condition:

Terminal window
kubectl get nodes
cilium status --wait
kubectl -n kube-system rollout status deployment/coredns --timeout=5m
kubectl get runtimeclass nvidia
kubectl get nodes -o custom-columns=NAME:.metadata.name,GPU:.status.allocatable.nvidia\.com/gpu,RDMA:.status.allocatable.rdma\.com/roce
kubectl get network-attachment-definition -A

Require two distinct eligible GPU nodes, one allocatable GPU and one selected RDMA unit per rank, a tested secondary network and persistent storage with at least 200 GiB free per node. Existing clusters may use parameterized resource names/counts only when their device plugins define those allocator units.

cert-manager must precede controllers whose webhooks or certificates depend on it.

Terminal window
helm repo add jetstack https://charts.jetstack.io
helm repo update
helm upgrade --install cert-manager jetstack/cert-manager \
--namespace cert-manager --create-namespace \
--version v1.21.2 \
--set crds.enabled=true
kubectl -n cert-manager wait --for=condition=Available deployment --all --timeout=5m

LWS orchestrates a leader and workers as a group with startup ordering. It is not a scheduler and does not select inference endpoints.

Terminal window
helm upgrade --install lws oci://registry.k8s.io/lws/charts/lws \
--namespace lws-system --create-namespace \
--version 0.10.0
kubectl -n lws-system wait --for=condition=Available deployment --all --timeout=5m
kubectl get crd leaderworkersets.leaderworkerset.x-k8s.io

Stop if the CRD is absent or the webhook/controller is unavailable. Source: LWS v0.10.0 release.

The Gateway API Inference Extension defines InferencePool and the Endpoint Picker Protocol. The recipe later creates one endpoint picker per model. The picker image is a separate choice from these CRDs: the generated bundle installs the upstream Gateway API Inference Extension picker, and the author’s cluster substitutes the llm-d router build, which additionally needs llm-d.ai InferenceObjective and InferenceModelRewrite CRDs. Nothing in this guide or the generated bundle requires those two, so they are not installed here.

Terminal window
kubectl apply --server-side -f \
https://github.com/kubernetes-sigs/gateway-api/releases/download/v1.5.0/standard-install.yaml
kubectl apply --server-side -f \
https://github.com/kubernetes-sigs/gateway-api/releases/download/v1.5.0/experimental-install.yaml
kubectl apply --server-side -f \
https://raw.githubusercontent.com/kubernetes-sigs/gateway-api-inference-extension/v1.5.0/config/crd/bases/inference.networking.k8s.io_inferencepools.yaml
kubectl get crd inferencepools.inference.networking.k8s.io

The final URL is the v1.5.0-tagged CRD source, not main. Review the v1.5.0 installation guides before adding other experimental APIs.

Terminal window
cat > eg-values.yaml <<'EOF'
config:
envoyGateway:
gateway:
controllerName: gateway.envoyproxy.io/gatewayclass-controller
extensionApis:
enableEnvoyPatchPolicy: true
enableBackend: true
extensionManager:
hooks:
xdsTranslator:
translation:
listener: {includeAll: true}
route: {includeAll: true}
cluster: {includeAll: true}
secret: {includeAll: true}
post: [Translation, Cluster, Route]
service:
fqdn:
hostname: ai-gateway-controller.envoy-ai-gateway-system.svc.cluster.local
port: 1063
backendResources:
- group: inference.networking.k8s.io
kind: InferencePool
version: v1
EOF
helm upgrade --install eg oci://docker.io/envoyproxy/gateway-helm \
--namespace envoy-gateway-system --create-namespace \
--version 1.8.3 -f eg-values.yaml
kubectl -n envoy-gateway-system wait --for=condition=Available deployment --all --timeout=10m

Those values are what make an InferencePool backend reachable at all. backendResources lets a route name that kind as a backend; extensionManager hands xDS translation to the Envoy AI Gateway controller on port 1063, which is where the external processor gets attached to the generated route. A plain helm install of the chart without them produces a Gateway that accepts and programs but cannot resolve an InferencePool reference. Install Envoy AI Gateway below before applying any routing bundle, or that hostname does not resolve when the first Gateway is translated.

Create the fresh-cluster GatewayClass with Envoy Gateway’s official controller name:

Terminal window
cat <<'EOF' | kubectl apply -f -
apiVersion: gateway.networking.k8s.io/v1
kind: GatewayClass
metadata:
name: vllm-envoy
spec:
controllerName: gateway.envoyproxy.io/gatewayclass-controller
EOF
kubectl wait gatewayclass/vllm-envoy --for=condition=Accepted --timeout=5m

Existing-cluster readers may instead supply an already-Accepted class to the routing builder. Do not assume the private deployment’s class name. Sources: Envoy Gateway install and 1.8.3 release.

Install CRDs before the controller, following the release’s Helm installation:

Terminal window
helm upgrade --install envoy-ai-gateway-crds \
oci://docker.io/envoyproxy/ai-gateway-crds-helm \
--namespace envoy-ai-gateway-system --create-namespace \
--version 1.1.0
helm upgrade --install envoy-ai-gateway \
oci://docker.io/envoyproxy/ai-gateway-helm \
--namespace envoy-ai-gateway-system \
--version 1.1.0
kubectl -n envoy-ai-gateway-system wait --for=condition=Available deployment --all --timeout=10m
kubectl get crd aigatewayroutes.aigateway.envoyproxy.io clienttrafficpolicies.gateway.envoyproxy.io

Stop if either chart/reference is unavailable in the v1.1.0 release or the installed CRD group differs. Do not substitute an unpinned latest chart. Source: Envoy AI Gateway v1.1.0.

Terminal window
kubectl api-resources | grep -E 'LeaderWorkerSet|InferencePool|AIGatewayRoute|ClientTrafficPolicy|BackendTrafficPolicy'
kubectl get gatewayclass vllm-envoy -o jsonpath='{range .status.conditions[*]}{.type}={.status}{" "}{end}{"\n"}'

Only proceed to TP=2 deployment when every resource type is discoverable and the GatewayClass is Accepted=True.