Distributed Training
Kubeflow’s Graduation Is a Vote for Kubernetes as the AI Control Plane
Kubeflow’s CNCF graduation signals growing confidence in Kubernetes as a common control plane for production AI workloads, from training and pipelines to governance and inference ...
Alan Shimel | | agentic AI, AI infrastructure, AI lifecycle, AI platform, AI Workloads, cloud native AI, cncf, Distributed Training, enterprise AI, GPU scheduling, KServe, Kubeflow, Kubeflow graduation, Kubeflow Pipelines, Kubeflow Trainer, kubernetes, Kubernetes AI, MLOps, OpenTelemetry, platform engineering
Kubernetes v1.36 Promotes Stability, Compatibility & Reproducibility
Kubernetes v1.36 (Spring 2026) introduces 70 enhancements, including major security hardening for the Kubelet API and the debut of Workload-Aware Scheduling (WAS) for AI/ML. This release focuses on fine-grained resource health, stable ...
Adrian Bridgwater | | AI/ML Infrastructure, CI/CD, cloud native security, cloud-native applications, Cluster Hardening, container security, containers, CSI Token Redaction, developers, Distributed Training, DRA, Dynamic Resource Allocation, External Token Signing, Gang Scheduling, K8s v1.36, Kubelet API Authorization, kubernetes, Kubernetes Enhancements 2026., Kubernetes v1.36, microservices, Node Logs, open source, PodGroup API, Resource Health Status, storage, Volume Group Snapshots, WAS, workload-aware scheduling

