Install serving and routing controllers
Experienced operators can enter here if their cluster already satisfies the host, GPU, primary-network and RDMA contracts. Fresh k3s readers should complete the linear tutorial first.
Readiness checklist
Section titled “Readiness checklist”Run from a workstation with cluster-admin credentials. Stop on the first failed condition:
kubectl get nodescilium status --waitkubectl -n kube-system rollout status deployment/coredns --timeout=5mkubectl get runtimeclass nvidiakubectl get nodes -o custom-columns=NAME:.metadata.name,GPU:.status.allocatable.nvidia\.com/gpu,RDMA:.status.allocatable.rdma\.com/rocekubectl get network-attachment-definition -ARequire two distinct eligible GPU nodes, one allocatable GPU and one selected RDMA unit per rank, a tested secondary network and persistent storage with at least 200 GiB free per node. Existing clusters may use parameterized resource names/counts only when their device plugins define those allocator units.
Install cert-manager first
Section titled “Install cert-manager first”cert-manager must precede controllers whose webhooks or certificates depend on it.
helm repo add jetstack https://charts.jetstack.iohelm repo updatehelm upgrade --install cert-manager jetstack/cert-manager \ --namespace cert-manager --create-namespace \ --version v1.21.2 \ --set crds.enabled=truekubectl -n cert-manager wait --for=condition=Available deployment --all --timeout=5mInstall LeaderWorkerSet v0.10.0
Section titled “Install LeaderWorkerSet v0.10.0”LWS orchestrates a leader and workers as a group with startup ordering. It is not a scheduler and does not select inference endpoints.
helm upgrade --install lws oci://registry.k8s.io/lws/charts/lws \ --namespace lws-system --create-namespace \ --version 0.10.0kubectl -n lws-system wait --for=condition=Available deployment --all --timeout=5mkubectl get crd leaderworkersets.leaderworkerset.x-k8s.ioStop if the CRD is absent or the webhook/controller is unavailable. Source: LWS v0.10.0 release.
Install Inference Extension v1.5.0
Section titled “Install Inference Extension v1.5.0”The Gateway API Inference Extension defines InferencePool and the Endpoint Picker Protocol. The recipe later creates one endpoint picker per model. The picker image is a separate choice from these CRDs: the generated bundle installs the upstream Gateway API Inference Extension picker, and the author’s cluster substitutes the llm-d router build, which additionally needs llm-d.ai InferenceObjective and InferenceModelRewrite CRDs. Nothing in this guide or the generated bundle requires those two, so they are not installed here.
kubectl apply --server-side -f \ https://github.com/kubernetes-sigs/gateway-api/releases/download/v1.5.0/standard-install.yamlkubectl apply --server-side -f \ https://github.com/kubernetes-sigs/gateway-api/releases/download/v1.5.0/experimental-install.yamlkubectl apply --server-side -f \ https://raw.githubusercontent.com/kubernetes-sigs/gateway-api-inference-extension/v1.5.0/config/crd/bases/inference.networking.k8s.io_inferencepools.yamlkubectl get crd inferencepools.inference.networking.k8s.ioThe final URL is the v1.5.0-tagged CRD source, not main. Review the v1.5.0 installation guides before adding other experimental APIs.
Install Envoy Gateway 1.8.3
Section titled “Install Envoy Gateway 1.8.3”cat > eg-values.yaml <<'EOF'config: envoyGateway: gateway: controllerName: gateway.envoyproxy.io/gatewayclass-controller extensionApis: enableEnvoyPatchPolicy: true enableBackend: true extensionManager: hooks: xdsTranslator: translation: listener: {includeAll: true} route: {includeAll: true} cluster: {includeAll: true} secret: {includeAll: true} post: [Translation, Cluster, Route] service: fqdn: hostname: ai-gateway-controller.envoy-ai-gateway-system.svc.cluster.local port: 1063 backendResources: - group: inference.networking.k8s.io kind: InferencePool version: v1EOFhelm upgrade --install eg oci://docker.io/envoyproxy/gateway-helm \ --namespace envoy-gateway-system --create-namespace \ --version 1.8.3 -f eg-values.yamlkubectl -n envoy-gateway-system wait --for=condition=Available deployment --all --timeout=10mThose values are what make an InferencePool backend reachable at all. backendResources lets a route name that kind as a backend; extensionManager hands xDS translation to the Envoy AI Gateway controller on port 1063, which is where the external processor gets attached to the generated route. A plain helm install of the chart without them produces a Gateway that accepts and programs but cannot resolve an InferencePool reference. Install Envoy AI Gateway below before applying any routing bundle, or that hostname does not resolve when the first Gateway is translated.
Create the fresh-cluster GatewayClass with Envoy Gateway’s official controller name:
cat <<'EOF' | kubectl apply -f -apiVersion: gateway.networking.k8s.io/v1kind: GatewayClassmetadata: name: vllm-envoyspec: controllerName: gateway.envoyproxy.io/gatewayclass-controllerEOFkubectl wait gatewayclass/vllm-envoy --for=condition=Accepted --timeout=5mExisting-cluster readers may instead supply an already-Accepted class to the routing builder. Do not assume the private deployment’s class name. Sources: Envoy Gateway install and 1.8.3 release.
Install Envoy AI Gateway v1.1.0
Section titled “Install Envoy AI Gateway v1.1.0”Install CRDs before the controller, following the release’s Helm installation:
helm upgrade --install envoy-ai-gateway-crds \ oci://docker.io/envoyproxy/ai-gateway-crds-helm \ --namespace envoy-ai-gateway-system --create-namespace \ --version 1.1.0helm upgrade --install envoy-ai-gateway \ oci://docker.io/envoyproxy/ai-gateway-helm \ --namespace envoy-ai-gateway-system \ --version 1.1.0kubectl -n envoy-ai-gateway-system wait --for=condition=Available deployment --all --timeout=10mkubectl get crd aigatewayroutes.aigateway.envoyproxy.io clienttrafficpolicies.gateway.envoyproxy.ioStop if either chart/reference is unavailable in the v1.1.0 release or the installed CRD group differs. Do not substitute an unpinned latest chart. Source: Envoy AI Gateway v1.1.0.
Final controller gate
Section titled “Final controller gate”kubectl api-resources | grep -E 'LeaderWorkerSet|InferencePool|AIGatewayRoute|ClientTrafficPolicy|BackendTrafficPolicy'kubectl get gatewayclass vllm-envoy -o jsonpath='{range .status.conditions[*]}{.type}={.status}{" "}{end}{"\n"}'Only proceed to TP=2 deployment when every resource type is discoverable and the GatewayClass is Accepted=True.