2026 AI Computing Infrastructure: Kubernetes GPU Operator, MIG Partitioning, Topology-Aware Scheduling, and Auto-scaling Inference Engines
Why does traditional CPU/memory scalar scheduling fail in the large language model era? A rigorous architectural deconstruction of AI infrastructure on Kubernetes in 2026: automated zero-touch lifecycles via the NVIDIA GPU Operator; hardware-isolated Multi-Instance GPU (MIG) slicing versus time-slicing; overcoming cross-NUMA interconnect bottlenecks with Kubelet Topology Manager and Dynamic Resource Allocation (DRA); and production auto-scaling for vLLM and SGLang inference clusters leveraging KEDA and real-time KV cache saturation metrics.
From Docker to Kubernetes: Cloud-Native Container & Cluster Orchestration Handbook
Master containerization and cluster orchestration from first principles: Linux namespaces, cgroups v2, SwarmKit, K8s control plane; real-world software architecture on K8s (gRPC/zero-downtime, Local NVMe/RocksDB/fencing); hacking K8s internals (Scheduling Framework plugins, NRI runtimes, Aggregated APIServer); and 2026 AI GPU scheduling.
Browse all 13 chapters in this series ▾
- 01 Docker Internals from First Principles: Linux Namespaces, cgroups v2, and OverlayFS Union Mounts Deep Dive
- 02 Container Networking Deep Dive: veth-pair, Linux Bridge, iptables NAT, and Cross-Host Topology
- 03 Lightweight Cluster Orchestration: Docker SwarmKit Architecture, Raft Consensus, and Ingress Routing Mesh
- 04 Kubernetes Control Plane Deep Dive: Declarative APIs, etcd Consensus, Scheduler, and Controller Reconciliation Loops
- 05 Kubernetes Networking Panorama: CNI Specification, Calico BGP, Cilium eBPF, and Gateway API Architecture
- 06 Kubernetes Storage Architecture: CSI Specification, Dynamic PV/PVC Provisioning, and StatefulSet Guarantees
- 07 Application Networking on Kubernetes: The gRPC Load Balancing Trap, Service Mesh, and Zero-Downtime Draining Sequences
- 08 Production Databases and Storage on Kubernetes: Local NVMe Passthrough, RocksDB/WAL Tuning, and Fencing Split-Brain Protection
- 09 Hacking the Kubernetes Scheduler: Custom Plugins via Scheduling Framework, Volcano DRF Math, and Descheduler Dynamic Rebalancing
- 10 Hacking Kubernetes Nodes & Runtimes: Breaking PLEG Bottlenecks, NRI Plugins, Kata/gVisor Sandboxes, and cgroups v2 Tuning
- 11 Hacking the Kubernetes Control Plane: WatchCache Internals, Aggregated APIServers, APF Shuffle Sharding, and 10k-Node etcd Sharding
- 12 Kubernetes Extensibility: The Operator Pattern, Custom Resource Definitions (CRD), and KubeBuilder in Production
- 13 2026 AI Computing Infrastructure: Kubernetes GPU Operator, MIG Partitioning, Topology-Aware Scheduling, and Auto-scaling Inference Engines Reading
Introduction: The Paradigm Shift from Scalar CPUs to AI GPU Topologies
Throughout the first decade of cloud-native evolution, the core resource abstractions of Kubernetes scheduling were fundamentally scalar and homogeneous:
- A standard Pod requests
cpu: "2"andmemory: "4Gi"; - The scheduler executes straightforward integer subtraction: . Whether those two CPU cores reside on socket 0 or socket 1, and which memory bus feeds that RAM, has negligible performance impact on conventional web microservices.
However, in the 2026 era of Large Language Models (LLMs), multimodal foundation models, and embodied intelligence, these simplistic scalar scheduling assumptions collapse:
- Capital Expenditure & Extreme Topology Sensitivity: An 8x NVIDIA H100/H200/B200 computing node represents significant hardware investment;
- The "Topology Cliff" of NUMA and NVLink Interconnects: When executing Tensor Parallelism (TP=4) across multiple GPUs, if the scheduler assigns 4 GPUs split across disparate NVSwitch interconnect groups or across CPU NUMA boundaries, cross-device communication latency increases by to , degrading token generation rates (Tokens/s);
- Sub-Device Granularity: Dedicated full-GPU allocation under-utilizes hardware for auxiliary workloads, requiring hardware-enforced Multi-Instance GPU (MIG) physical slicing with strict multi-tenant isolation.
flowchart TD
subgraph GPUHost["Physical Topology of Modern 8-GPU AI Node"]
direction TB
subgraph NUMA0["NUMA Socket 0 (CPU 0 + Memory 0)"]
GPU0["GPU 0 (B200)"] <== "NVLink (900 GB/s)" ==> GPU1["GPU 1 (B200)"]
GPU2["GPU 2 (B200)"] <== "NVLink (900 GB/s)" ==> GPU3["GPU 3 (B200)"]
GPU0 <==> GPU2
GPU1 <==> GPU3
NIC0["InfiniBand / RoCE NIC 0
(PCIe Switch Direct to GPU 0-3)"]
end
subgraph NUMA1["NUMA Socket 1 (CPU 1 + Memory 1)"]
GPU4["GPU 4 (B200)"] <== "NVLink (900 GB/s)" ==> GPU5["GPU 5 (B200)"]
GPU6["GPU 6 (B200)"] <== "NVLink (900 GB/s)" ==> GPU7["GPU 7 (B200)"]
GPU4 <==> GPU6
GPU5 <==> GPU7
NIC1["InfiniBand / RoCE NIC 1
(PCIe Switch Direct to GPU 4-7)"]
end
NUMA0 <-.->|"Cross-NUMA UPI/QPI Interconnect (Bottleneck: 32-64 GB/s)"| NUMA1
end
As the capstone chapter of From Docker to Kubernetes: Cloud-Native Container & Cluster Orchestration Handbook, this guide explores AI computing infrastructure: the NVIDIA GPU Operator automated lifecycle, MIG hardware slicing, topology-aware scheduling, and production autoscaling architectures for vLLM and SGLang inference clusters.
1. NVIDIA GPU Operator: Zero-Touch Management at Scale
Historically, enabling GPU workloads on Kubernetes was an operational bottleneck. Administrators had to manually SSH into bare-metal servers, compile kernel-matching NVIDIA drivers, configure CUDA runtimes, set up nvidia-container-toolkit, and deploy the Device Plugin DaemonSet.
In hyperscale AI clusters subject to continuous node cycling and autoscaling, manual provisioning causes driver version drift and node initialization failures.
The NVIDIA GPU Operator applies the Operator Pattern to hardware management. It automates node enablement via a declarative ClusterPolicy, delivering zero-touch node configuration:
flowchart LR
K8sNode["Clean Linux Bare Metal
(Only containerd required)"] --> GPUOperator["NVIDIA GPU Operator
(ClusterPolicy Controller)"]
subgraph AutoComponents["Automated Containerized DaemonSets"]
direction TB
C1["1. Driver Container
Dynamically compiles & loads host nvidia.ko modules"]
C2["2. Container Toolkit
Configures containerd CDI to inject /dev/nvidia* devices"]
C3["3. Device Plugin
Advertises nvidia.com/gpu scalar resources to Kubelet"]
C4["4. DCGM Exporter
Streams GPU memory, power, & compute metrics to Prometheus"]
C5["5. GPU Feature Discovery (GFD)
Labels nodes with hardware specs (e.g., Blackwell architecture)"]
C1 --> C2 --> C3 --> C4 --> C5
end
GPUOperator --> AutoComponents
AutoComponents --> ReadyNode["Fully Provisioned AI Worker Node
Ready for LLM Workloads!"]
Component Architecture:
- Driver Container: Compiles and injects
nvidia.koandnvidia-uvm.kokernel modules into the host kernel at boot, keeping the underlying Linux host filesystem clean; - NVIDIA Container Toolkit: Modifies the container runtime (
/etc/containerd/config.toml) with Container Device Interface (CDI) hooks, ensuring containers securely bind/dev/nvidia*devices and CUDA libraries at launch; - NVIDIA Device Plugin: Communicates with Kubelet over gRPC to register assignable GPU inventory;
- DCGM Exporter: Collects high-frequency GPU telemetry—such as VRAM consumption, Tensor Core utilization, NVLink throughput, and thermal metrics—via the Data Center GPU Manager (DCGM);
- GPU Feature Discovery (GFD): Applies descriptive Kubernetes labels to nodes (e.g.,
nvidia.com/gpu.product=NVIDIA-H100-80GB-HBM3), enabling targeted workload scheduling.
2. GPU Slicing Architecture: Time-Slicing, MIG, and DRA
Not all AI tasks require an entire 80GB or 140GB HBM GPU. To maximize hardware utilization across enterprise clusters, three distinct allocation mechanisms are used:
| GPU Slicing Mechanism | Isolation Architecture | Memory & Compute Independence | Recommended Use Case |
|---|---|---|---|
| Time-Slicing | Software-level temporal sharing. Exposes multiple virtual devices to Kubelet. | No Hardware Isolation. Shared memory space; an out-of-memory error in one process terminates all co-located containers. | Dev/test staging, lightweight CI/CD, tiny feature-extraction models. |
| MIG (Multi-Instance GPU) | Hardware-Level Partitioning. Divides physical silicon into dedicated Compute Instances (CI) and GPU Instances (GI). | Strict Physical Isolation. Dedicated Streaming Multiprocessors, independent memory controllers, and isolated DMA engines. | Enterprise multi-tenancy, inference serving for small models (Embedding, Reranking). |
| DRA (Dynamic Resource Allocation) | Kubernetes 1.30+ claim-based allocation model replacing scalar integer counters. | Topology-Aware Claims. Requests specific device pairings (e.g., "Pair 2 GPUs connected via NVLink"). | Large-scale LLM training, disaggregated heterogeneous inference clusters. |
NVIDIA MIG Physical Partitioning (Example: H100 80GB):
┌─────────────────────────────────────────────────────────────┐
│ Physical NVIDIA H100 80GB HBM3 GPU │
├──────────────────┬──────────────────┬───────────────────────┤
│ Instance 1 │ Instance 2 │ Instance 3 │
│ Profile: 3g.40gb │ Profile: 2g.20gb │ Profile: 1g.10gb (x2) │
│ • 42 SM Units │ • 28 SM Units │ • 14 SM Units │
│ • 40GB VRAM │ • 20GB VRAM │ • 10GB VRAM │
│ • 1.6 TB/s Band │ • 800 GB/s Band │ • 400 GB/s Band │
│ (Deploy 14B LLM) │ (Deploy 7B LLM) │ (Deploy Embedding) │
└──────────────────┴──────────────────┴───────────────────────┘ Configuring mig.strategy=mixed within the GPU Operator enables declarative partitioning of physical cards into distinct MIG profiles (such as nvidia.com/mig-3g.40gb: 2), which the Kubernetes scheduler treats as non-preemptible hardware assets.
3. Topology-Aware Scheduling: NUMA and NVLink Co-location
Distributed LLM inference using Tensor Parallelism (TP) relies on ultra-low latency inter-card communication. Suboptimal placement decisions introduce significant performance degradation:
3.1 The Cross-NUMA Interconnect Bottleneck
On a typical dual-socket platform, CPU 0 controls Socket 0 local memory and GPUs 0–3, while CPU 1 controls Socket 1 and GPUs 4–7.
- If a Pod is assigned
CPU Core 2(Socket 0) alongsideGPU 5(Socket 1 PCIe root complex); - Every host-to-device tensor memory transfer must cross CPU interconnect buses (such as Intel UPI or AMD Infinity Fabric), reducing bandwidth from 64 GB/s down to sub-15 GB/s!
3.2 Kubelet Topology Manager Alignment
To resolve this, Kubernetes provides the Topology Manager on worker nodes. It coordinates the CPU Manager, Memory Manager, and Device Plugin to enforce co-location:
# Kubelet node configuration: /var/lib/kubelet/config.yaml
topologyManagerPolicy: single-numa-node
topologyManagerScope: container
cpuManagerPolicy: static Under single-numa-node:
- When Kubelet receives a container assignment, the Topology Manager evaluates NUMA boundaries;
- The container launches only if all requested exclusive CPU cores, HugePages memory allocations, and GPU devices reside within the identical physical NUMA node;
- If alignment cannot be satisfied, Kubelet rejects the placement (
TopologyAffinityError), forcingkube-schedulerto find an aligned node and preventing cross-NUMA interconnect stalls.
4. Production AI Serving: Autoscaling vLLM & SGLang on Kubernetes
Operating production inference engines like vLLM and SGLang on Kubernetes requires accounting for two architectural constraints:
- Aggressive VRAM Pre-allocation: Upon initialization, vLLM allocates up to 90% of available VRAM to establish its PagedAttention KV Cache pool. Consequently, conventional Kubernetes CPU/Memory HPA policies fail, as GPU memory utilization remains constant regardless of active traffic;
- Prefill/Decode (P/D) Disaggregation: Prefill operations are compute-bound, whereas Decode operations are memory-bandwidth-bound.
The modern production architecture leverages KEDA (Kubernetes Event-driven Autoscaling) triggered by vLLM metrics scraped via Prometheus:
flowchart TD
subgraph TrafficEntry["Inference Ingress Layer"]
ClientReq["Concurrent Client Prompts"] --> Gateway["Envoy / Cilium Gateway API"]
end
Gateway --> Router["Prefill-Decode Disaggregation Router"]
subgraph K8sCluster["Kubernetes Production AI Cluster"]
direction TB
Router -->|"Compute-Intensive Prefill"| PrefillPool["Prefill Pods (High-Throughput Batching)"]
Router -->|"Latency-Sensitive Decode"| DecodePool["Decode Pods (KV Cache Sensitive)"]
PrefillPool -.->|"Metrics (:8000/metrics)"| Prometheus["Prometheus Server"]
DecodePool -.->|"vllm:num_requests_waiting
vllm:gpu_cache_usage_factor"| Prometheus
Prometheus --> KEDA["KEDA Autoscaling Engine
Evaluates Queue Depth & KV Saturation"]
KEDA -->|"HPA Dynamic Scaling"| DecodePool
end
4.1 Production KEDA ScaledObject Configuration
The following manifest demonstrates autoscaling for a production vLLM deployment:
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: vllm-inference-autoscaler
namespace: llm-serving
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: vllm-qwen-72b-instruct
minReplicaCount: 2
maxReplicaCount: 16
cooldownPeriod: 300 # 300s scale-down delay prevents thrashing from large model reload cycles
pollingInterval: 5 # Polls telemetry every 5s for rapid burst response
advanced:
horizontalPodAutoscalerConfig:
behavior:
scaleUp:
stabilizationWindowSeconds: 0
policies:
- type: Percent
value: 100 # Permits instantaneous 100% replica expansion during surges
periodSeconds: 15
scaleDown:
stabilizationWindowSeconds: 300
triggers:
# Trigger 1: Queued inference requests awaiting compute capacity
- type: prometheus
metadata:
serverAddress: http://prometheus-k8s.monitoring.svc.cluster.local:9090
metricName: vllm_num_requests_waiting
query: sum(vllm:num_requests_waiting{model="qwen-72b"})
threshold: '5' # Scales up when queue backlog exceeds 5 requests
# Trigger 2: KV cache block allocation saturation
- type: prometheus
metadata:
serverAddress: http://prometheus-k8s.monitoring.svc.cluster.local:9090
metricName: vllm_gpu_cache_usage_factor
query: avg(vllm:gpu_cache_usage_factor{model="qwen-72b"})
threshold: '0.80' # Triggers expansion when KV cache saturation crosses 80% 5. Comprehensive Series Review: Full Curriculum Retrospective
This concludes all 13 chapters of From Docker to Kubernetes: Cloud-Native Container & Cluster Orchestration Handbook!
Reviewing the architectural progression across the curriculum:
flowchart TD
Ch1["Ch 1: Docker Internals
Linux namespaces, cgroups v2, OverlayFS CoW, and the runc execution chain"]
Ch2["Ch 2: Container Networking
veth-pairs, Linux Bridge docker0, iptables NAT, and multi-host topologies"]
Ch3["Ch 3: Docker SwarmKit
SwarmKit topology, embedded Raft consensus, VXLAN, and IPVS Routing Mesh"]
Ch4["Ch 4: Kubernetes Control Plane
Declarative control theory, etcd MVCC, two-phase scheduling, and Informer reconcilers"]
Ch5["Ch 5: Kubernetes Networking
The four network axioms, CNI plugins, Calico BGP routing, Cilium eBPF, and Gateway API"]
Ch6["Ch 6: Kubernetes Storage
The CSI four-stage lifecycle, dynamic PV/PVC provisioning, and StatefulSet topology guarantees"]
Ch7["Ch 7: Application Networking
gRPC load balancing trap, Service Mesh, and zero-downtime draining sequences"]
Ch8["Ch 8: Production Storage & Databases
Local NVMe passthrough, RocksDB/WAL tuning, and Node Fencing split-brain defense"]
Ch9["Ch 9: Scheduler Customization
Scheduling Framework plugins, Gang Scheduling, and Descheduler rebalancing"]
Ch10["Ch 10: Node & Runtime Customization
NRI resource plugins, Kata/gVisor sandboxed runtimes, and sysctl isolation"]
Ch11["Ch 11: Control Plane Customization
Aggregated APIServers, APF traffic control with Shuffle Sharding, and etcd sharding"]
Ch12["Ch 12: The Operator Pattern
CRD domain models, KubeBuilder scaffolding, Split Clients, Reconcilers, and Finalizers"]
Ch13["Ch 13: 2026 AI Infrastructure
NVIDIA GPU Operator, MIG hardware partitioning, Topology Manager, and vLLM/SGLang autoscaling"]
Ch1 --> Ch2 --> Ch3 --> Ch4 --> Ch5 --> Ch6 --> Ch7 --> Ch8 --> Ch9 --> Ch10 --> Ch11 --> Ch12 --> Ch13
- Linux Kernel Foundations to Single-Host Runtimes (Ch 1–2): We deconstructed the abstraction of containers into standard Linux processes bounded by kernel namespaces and cgroups, mapped through virtual bridges and iptables routing rules;
- Lightweight Clustering to Declarative Control Planes (Ch 3–6): We examined Docker Swarm's embedded Raft consensus, then progressed to Kubernetes declarative control loops, CNI flat networking, and CSI stateful volume architectures;
- Application Engineering & Hacking Kubernetes (Ch 7–11): We resolved application-layer gRPC stream pinning, rollout 502 race conditions, Local NVMe direct passthrough, and split-brain fencing; then customized the Kubernetes core via Scheduling Framework plugins, NRI hardware interception, and Aggregated APIServers;
- Automation to the 2026 AI Frontier (Ch 12–13): We codified operational heuristics into autonomous production Operators, synthesizing modern cloud-native orchestration with hyperscale accelerator topologies from physical NUMA alignment to distributed LLM serving.
Frequently Asked Questions (FAQ)
Q1: Why does vLLM report GPU VRAM allocation above 90% immediately after booting, and will this trigger Kubernetes OOMKills?
No. This is an intentional design pattern of PagedAttention KV Cache pre-allocation:
- During startup, vLLM loads model weights (e.g., Qwen-72B requires ~144GB across multiple cards);
- Following model initialization, vLLM pre-allocates 90% of remaining physical VRAM (
--gpu-memory-utilization=0.9) to manage dynamic paging tables; - Kubernetes OOM Behavior: The Kubernetes OOMKiller monitors Host Physical RAM, not device VRAM. As long as host memory stays within
resources.limits.memory, the container will not be terminated; - Autoscaling Guidance: Do not base HPA rules on raw VRAM consumption reported by DCGM. Instead, autoscale against vLLM's internal metric
vllm:gpu_cache_usage_factor.
Q2: On an 8-GPU node, why does distributing a Tensor Parallel (TP=4) model across arbitrary GPUs introduce severe latency, and how does the Topology Manager resolve it?
Root Cause: Asymmetric Interconnect Topologies. In HGX architectures, GPUs communicate in distinct sub-clusters via high-bandwidth NVSwitch fabrics (up to 900 GB/s bidirectional bandwidth per card). If the scheduler places a TP=4 workload across non-co-located GPUs (e.g., GPUs 0, 1, 4, and 5):
- GPUs 0 and 1 communicate over high-speed NVLink;
- GPUs 1 and 4 are forced over lower-bandwidth PCIe or cross-socket CPU interconnects (dropping bandwidth to sub-30 GB/s);
- Because Tensor Parallelism requires frequent synchronization via
All-Reduceoperations, overall throughput drops significantly; Resolution: EnabletopologyManagerPolicy: single-numa-nodeor leverage Kubernetes 1.30+ Dynamic Resource Allocation (DRA) to schedule GPU workloads within verified interconnect boundaries.
Q3: What is the primary operational advantage of Kubernetes Dynamic Resource Allocation (DRA) over traditional Device Plugins for AI workloads?
Limitations of Traditional Device Plugins: Device Plugins advertise resources as scalar integers (e.g., nvidia.com/gpu: 4). The scheduler cannot evaluate whether the assigned GPUs share an NVLink fabric, connect to the same NUMA socket, or align with local InfiniBand NICs.
Capabilities of Dynamic Resource Allocation (DRA):
- Claim-Based Parameterization: Workloads submit a
ResourceClaimspecifying structured hardware constraints; - Topology Discovery: DRA enables vendor drivers to expose hardware topology (NVLink connectivity matrices, PCIe switch mappings) directly to the scheduling framework;
- Coordinated Multi-Device Scheduling: The scheduler atomically pairs GPU allocations with corresponding local network controllers on the same PCIe root complex, preventing interconnect bottlenecks.