Skip to content
- Node: a physical or virtual machine registered with Kubernetes. This tutorial has one control-plane node and two GPU agent nodes.
- Pod: Kubernetes’ smallest deployable unit; one or more containers sharing network and lifecycle.
- Controller: software that watches Kubernetes objects and reconciles actual state toward their declared state.
- CRD: CustomResourceDefinition, which adds a resource type such as
LeaderWorkerSet or InferencePool to the Kubernetes API.
- CNI: Container Network Interface; the plugin contract used to connect pods. Cilium is primary, while Multus attaches an additional network.
- Runtime: the container execution integration beneath Kubernetes. k3s uses containerd and discovers
nvidia-container-runtime for GPU injection.
- RuntimeClass: Kubernetes object selecting a configured runtime handler. Recipe pods explicitly use
runtimeClassName: nvidia.
- GPU device plugin: component that advertises GPU extended resources to Kubernetes and assigns devices to containers.
- Extended resource: integer scheduling/assignment contract such as
nvidia.com/gpu or rdma.com/roce. Its count is allocator-defined and need not equal physical-device count.
- RDMA: Remote Direct Memory Access, used here for the secondary high-throughput collective path between GPU ranks.
- RoCE: RDMA over Converged Ethernet; the selected fabric carried by a macvlan secondary interface.
- Multus: CNI multiplexer that requests one or more extra pod-network attachments after the primary CNI.
- Whereabouts: cluster-wide IP address manager for the secondary Multus network.
- LWS: LeaderWorkerSet, a Kubernetes controller/API that manages a leader and workers as a group. It is not an inference scheduler.
- Tensor parallelism (TP): one model computation split across ranks. TP=2 here means two ranks across two nodes.
- Rank zero: leader process exposing the HTTP model API and coordinating distributed execution.
- Rank one: headless worker process participating in collectives; never an InferencePool HTTP endpoint.
- InferencePool: Gateway API Inference Extension resource declaring eligible model-serving endpoints.
- EPP: Endpoint Picker Protocol/Picker; evaluates eligible endpoints and returns a selection to Gateway. The request payload does not traverse EPP. The upstream Gateway API Inference Extension image and the llm-d router build both implement it, and they score with different signals.
- ext_proc: Envoy’s external processing gRPC filter. Envoy AI Gateway attaches it to the route generated for an
InferencePool backend, and it is the channel on which Envoy asks the picker and receives an endpoint.
- KV events: engine-published records of which cache blocks a rank owns, sent over ZeroMQ from the serving pod. A router that indexes real block ownership instead of guessing from token prefixes needs them, plus a topic naming the pod address the pool advertises.
- EndpointPickerConfig: the picker’s own configuration document: which data producers run, which filters and scorers apply and in what order. Its api group differs between the upstream extension (
inference.networking.x-k8s.io) and the llm-d router (llm-d.ai).
- InferenceObjective: an llm-d router resource naming a request priority band; requests carry the band in a header, and flow control admits each band up to its own cap.
- BackendTrafficPolicy / ClientTrafficPolicy: Envoy Gateway policies that raise the upstream request timeout ceiling and the per-connection buffer limit respectively; the first bounds long prefills, the second bounds request-body size in buffered external processing.
- GatewayClass: cluster-scoped declaration of which Gateway controller owns a class of Gateways.
- AIGatewayRoute: Envoy AI Gateway route that matches AI model metadata and sends traffic toward the InferencePool.
- GSP: GPU System Processor firmware. Its version must match the NVIDIA kernel modules/userspace contract.
- DKMS: Dynamic Kernel Module Support; builds and installs the pinned open driver against the exact kernel ABI.