Kubenatives
Production Kubernetes for ML/AI workloads: GPU infrastructure, control plane internals, and model serving patterns for engineers running inference at scale.
- Indexed issues, last 90 days
- 5
- Latest publication
- Jul 31, 2026
- Audience
- Checking…
- Earliest in this view
- Jul 3, 2026
Latest issues
Triton Inference Server on Kubernetes: Multi-Model Serving (opens the original)
Read excerpt
vLLM has one job. Serve transformer language models with PagedAttention and continuous batching. It does that job better than anything else.Triton has a different job. Serve any model, in any framework, on any hardware, from a single endpoint. PyTorch, TensorFlow, ONNX, TensorRT, Python code, OpenVINO, even custom backends. All from the same server.That flexibility is the point. And the reason most teams end up running Triton alongside vLLM rather than instead of it.This article covers what Trit
RDMA and InfiniBand: Why GPU Networking Is Different on Kubernetes (opens the original)
Read excerpt
Single-node GPU workloads are simple. The GPU talks to CPU memory over PCIe, to other GPUs in the same box over NVLink. No network involved. Kubernetes networking never enters the picture.Multi-node GPU workloads are different. A Llama 70B model split across 2 nodes with 4 GPUs each requires all-reduce operations at every training step or tensor-parallel inference call. That is hundreds of gigabytes per second of cross-node traffic.The default Kubernetes network stack uses kernel TCP over a CNI
A/B Testing LLM Models in Production with Kubernetes (opens the original)
Read excerpt
Your team just finished evaluating a new model. Llama 3.3 70B outperforms your current Llama 3.1 8B on every internal benchmark. The engineering review is done. The cost analysis is done. Time to ship.The wrong way: update the Deployment image, wait for the rollout, monitor Grafana for 10 minutes, call it a success.The right way: split 5% of production traffic to the new model, compare quality metrics side by side for 48 hours, ramp traffic if quality holds, roll back instantly if it drops.This
API Server Latency: What Is Normal and What Is a Red Flag (opens the original)
Read excerpt
Every Kubernetes operation goes through the API server. kubectl commands, controller reconciliation loops, kubelet heartbeats, webhook calls, custom operators. Thousands of requests per minute on a busy cluster.When the API server slows down, everything slows down. Deployments stall. Pods stay in Pending longer. Controllers fall behind their desired state. The cluster feels sluggish even though CPU and memory on the nodes look fine.The hard part is knowing when slow is slow. A 200ms p99 on a lis
GPU Monitoring with DCGM Exporter: The Metrics That Matter (opens the original)
Read excerpt
nvidia-smi is the first tool engineers reach for when checking GPU status. It shows utilization, temperature, memory, and power draw. It works for a single node.It does not work for a 20 node GPU cluster. You cannot SSH into 20 nodes every time someone reports slow inference. You need GPU metrics in Prometheus, dashboards in Grafana, and alerts that page you before users notice.DCGM Exporter is the NVIDIA Data Center GPU Manager running as a DaemonSet on every GPU node. It collects GPU health an
Publishing over time
Last 90 days. Choose a month to open its work.
Recurring subjects
Named in the text we hold. One piece can cover several.
Audience
No verified audience measurement yet.
About this data
Counts cover the work we have indexed. Tone needs enough text and a confident classification. Excerpts and episode notes are not full articles or transcripts.
Identity or attribution wrong? Suggest a correction.