cloud native AI
Autoscaling AI Workloads on Kubernetes With KEDA and What it Means for Agentic Systems
KEDA can scale Kubernetes AI workloads on real demand signals such as queue depth, helping model-serving and agent workloads respond faster while reducing idle compute costs ...
Kishor Patil | | agentic AI, AI agents, AI infrastructure, AI model serving, cloud native AI, devops, Event-Driven Autoscaling, horizontal pod autoscaling, HPA, inference scaling, KEDA, Kubernetes AI workloads, Kubernetes autoscaling, Kubernetes Event-Driven Autoscaling, platform engineering, Pub/Sub, queue depth, RabbitMQ, Redis, scale-to-zero, SQS
Kubeflow’s Graduation Is a Vote for Kubernetes as the AI Control Plane
Kubeflow’s CNCF graduation signals growing confidence in Kubernetes as a common control plane for production AI workloads, from training and pipelines to governance and inference ...
Alan Shimel | | agentic AI, AI infrastructure, AI lifecycle, AI platform, AI Workloads, cloud native AI, cncf, Distributed Training, enterprise AI, GPU scheduling, KServe, Kubeflow, Kubeflow graduation, Kubeflow Pipelines, Kubeflow Trainer, kubernetes, Kubernetes AI, MLOps, OpenTelemetry, platform engineering
Kubernetes Wasn’t Built for GPUs. Make It Behave
Kubernetes counts whole GPUs and treats pods as disposable. An LLM pod is neither. Share the silicon with MIG/MPS/time-slicing and stop paying for idle ...
Sneha Gullapalli | | A100, AI infrastructure, AI Workloads, cloud native AI, Dynamic Resource Allocation, GPU autoscaling, GPU cost reduction, GPU optimization, GPU partitioning, GPU sharing, GPU time-slicing, GPU utilization, H100, Karpenter, KServe, Kubernetes DRA, Kubernetes GPU scheduling, LLM Inference, model caching, multi-instance GPU, NVIDIA GPU Operator, NVIDIA MIG, NVIDIA MPS, scale-to-zero, VRAM
Stop Treating GPUs Like Web Pods
Kubernetes schedules accelerators as opaque integers, and your bill pays for it. Share the silicon, scale on the right signal and keep weights out of the image ...
Veera Ravindra Divi | | AI infrastructure, AI serving, autoscaling, cloud costs, cloud native AI, DCGM exporter, DRA, Dynamic Resource Allocation, GPU costs, GPU scheduling, GPU sharing, GPU utilization, GPUs, inference workloads, KEDA, kubernetes, Kubernetes GPU scheduling, LLM Inference, MIG, model weights, MPS, NVIDIA GPUs, NVIDIA MIG, Prometheus, scale-to-zero, time-slicing
NVIDIA Is Putting Real Skin in the Open AI Game
Open AI requires community-governed infrastructure and companies willing to contribute code, engineering and costly GPU cycles. NVIDIA is doing exactly that ...
Beyond the Model: Why AI Agent Orchestration Requires Cloud-Native Engineering
AI agent orchestration is a distributed systems challenge. Cloud-native engineering provides the resilience, observability, security and scalability needed for production AI ...
Nithiya Dharshini | | agent communication, agentic AI infrastructure, AI agent orchestration, AI governance, AI infrastructure, AI observability, AI scalability, AI security, AI workflow monitoring, AI workload management, automated scaling, cloud native AI, cloud-native engineering, containerized AI services, distributed AI systems, distributed tracing, enterprise AI agents, event-driven architecture, GitOps, Kubernetes for AI, multi-agent systems, platform engineering, production AI systems, resilient AI systems
AI-driven Kubernetes in Action: Exploring AI-Assisted Kubernetes Operations
Discover how AI is transforming Kubernetes from reactive troubleshooting to proactive, intelligent automation. Learn about essential AIOps tools, resource optimization strategies, and the challenges of managing AI-enabled container orchestration at scale ...
Measuring AI-Driven Automation: The Metrics That Prove Whether Your Platform is Actually Getting Smarter
AI is reshaping cloud-native operations, but old automation metrics mislead. Learn the modern KPIs—MTTR, action quality, autonomy and cognitive load reduction ...
Ankush Dhar | | agentic AI systems, AI action quality, AI automation metrics, AI cloud operations, AI governance metrics, AI incident response, AI-driven reliability engineering, AIOps performance metrics, autonomous remediation, cloud cost optimization AI, cloud native AI, cloud platform automation, cloud-native automation KPIs, cognitive load reduction SRE, explainable AI operations, false action rate, MTTR reduction, operational intelligence, predictive incident prevention, SRE automation
Google Extends Kubernetes Service to Safely Run Agentic AI Workloads
At KubeCon + CloudNativeCon North America 2025, Google unveiled major GKE upgrades — including an AI sandbox, inference gateway, pod snapshots, and 130,000-node clusters — to optimize and secure agentic AI workloads ...
CNCF Adds Program to Standardize AI Workloads on Kubernetes Clusters
CNCF introduces the Certified Kubernetes AI Conformance Program to standardize AI and ML workload deployment, ensuring interoperability and sovereign cloud compliance ...
Mike Vizard | | AI deployment, AI infrastructure, AI on Kubernetes, AI portability, AI scalability, AI Workloads, Certified Kubernetes AI Conformance Program, cloud native AI, cloud-native ecosystem, CloudNativeCon, cncf, data science, hybrid cloud, IT operations, KubeCon 2025, kubernetes, Kubernetes certification, Kubernetes conformance, Kubernetes interoperability, Kubernetes standards, ML deployment, ML frameworks, ML workloads, sovereign cloud

