Skip to content
HeyJared

Kubenatives

Production Kubernetes for ML/AI workloads: GPU infrastructure, control plane internals, and model serving patterns for engineers running inference at scale.

Newsletter · By Sharon Sahadevan · English · Paid tier available · Official site

Indexed issues, last 90 days
5
Latest publication
Jul 31, 2026
Audience
Checking…
Earliest in this view
Jul 3, 2026
The latest indexed work is over 30 days old. There may be a gap in what we hold.

Latest issues

  1. Issue · Jul 31, 2026

    Triton Inference Server on Kubernetes: Multi-Model Serving (opens the original)

    Excerpt

    Read excerpt

    vLLM has one job. Serve transformer language models with PagedAttention and continuous batching. It does that job better than anything else.Triton has a different job. Serve any model, in any framework, on any hardware, from a single endpoint. PyTorch, TensorFlow, ONNX, TensorRT, Python code, OpenVINO, even custom backends. All from the same server.That flexibility is the point. And the reason most teams end up running Triton alongside vLLM rather than instead of it.This article covers what Trit

  2. Issue · Jul 27, 2026

    RDMA and InfiniBand: Why GPU Networking Is Different on Kubernetes (opens the original)

    Excerpt

    Read excerpt

    Single-node GPU workloads are simple. The GPU talks to CPU memory over PCIe, to other GPUs in the same box over NVLink. No network involved. Kubernetes networking never enters the picture.Multi-node GPU workloads are different. A Llama 70B model split across 2 nodes with 4 GPUs each requires all-reduce operations at every training step or tensor-parallel inference call. That is hundreds of gigabytes per second of cross-node traffic.The default Kubernetes network stack uses kernel TCP over a CNI

  3. Issue · Jul 17, 2026

    A/B Testing LLM Models in Production with Kubernetes (opens the original)

    Excerpt

    Read excerpt

    Your team just finished evaluating a new model. Llama 3.3 70B outperforms your current Llama 3.1 8B on every internal benchmark. The engineering review is done. The cost analysis is done. Time to ship.The wrong way: update the Deployment image, wait for the rollout, monitor Grafana for 10 minutes, call it a success.The right way: split 5% of production traffic to the new model, compare quality metrics side by side for 48 hours, ramp traffic if quality holds, roll back instantly if it drops.This

  4. Issue · Jul 10, 2026

    API Server Latency: What Is Normal and What Is a Red Flag (opens the original)

    Excerpt

    Read excerpt

    Every Kubernetes operation goes through the API server. kubectl commands, controller reconciliation loops, kubelet heartbeats, webhook calls, custom operators. Thousands of requests per minute on a busy cluster.When the API server slows down, everything slows down. Deployments stall. Pods stay in Pending longer. Controllers fall behind their desired state. The cluster feels sluggish even though CPU and memory on the nodes look fine.The hard part is knowing when slow is slow. A 200ms p99 on a lis

  5. Issue · Jul 3, 2026

    GPU Monitoring with DCGM Exporter: The Metrics That Matter (opens the original)

    Excerpt

    Read excerpt

    nvidia-smi is the first tool engineers reach for when checking GPU status. It shows utilization, temperature, memory, and power draw. It works for a single node.It does not work for a 20 node GPU cluster. You cannot SSH into 20 nodes every time someone reports slow inference. You need GPU metrics in Prometheus, dashboards in Grafana, and alerts that page you before users notice.DCGM Exporter is the NVIDIA Data Center GPU Manager running as a DaemonSet on every GPU node. It collects GPU health an

Publishing over time

Last 90 days. Choose a month to open its work.

Recurring subjects

Named in the text we hold. One piece can cover several.

Audience

No verified audience measurement yet.

About this data

Counts cover the work we have indexed. Tone needs enough text and a confident classification. Excerpts and episode notes are not full articles or transcripts.

Identity or attribution wrong? Suggest a correction.

See coverage about Kubenatives