Cloud-Native Development
CNCF Graduates Kubeflow for Production AI on Kubernetes
The Cloud Native Computing Foundation has graduated Kubeflow, giving the open source AI and machine learning platform CNCF’s highest maturity designation as enterprises move more AI workloads into production. Kubeflow runs on ...
How We Cut Kubernetes Deployment Validation From 45 Minutes to 2 minutes
There is a moment every release engineer knows well. The CI/CD pipeline turns green. The deployment job reports success. Everyone exhales for a second and thinks, “Okay, the release is done.” But ...
Docker Desktop Gets a Hypervisor of its Own
Docker is bringing full backend parity across all of its Docker Desktop editions, ensuring that macOS, Windows, and (eventually) Linux users get the same performance and polish. The company unveiled a new ...
Kubernetes Wasn’t Built for GPUs. Make It Behave
Kubernetes counts whole GPUs and treats pods as disposable. An LLM pod is neither. Share the silicon with MIG/MPS/time-slicing and stop paying for idle ...
Sneha Gullapalli | | A100, AI infrastructure, AI Workloads, cloud native AI, Dynamic Resource Allocation, GPU autoscaling, GPU cost reduction, GPU optimization, GPU partitioning, GPU sharing, GPU time-slicing, GPU utilization, H100, Karpenter, KServe, Kubernetes DRA, Kubernetes GPU scheduling, LLM Inference, model caching, multi-instance GPU, NVIDIA GPU Operator, NVIDIA MIG, NVIDIA MPS, scale-to-zero, VRAM
How Open-Source Automation Tools Handle the Testing Problem That Cloud-Native Independent Deployment Creates
Independent deployment creates a coverage currency problem manual maintenance cannot scale to address. Learn how open-source automation tools handle it structurally. ...
Sancharini Panda | | API mocking, behavioral drift, CI/CD testing, Cloud-Native Testing, contract testing, coverage currency, eBPF, go-vcr, independent deployment, integration test fixtures, integration testing, Keploy, Kubernetes testing, Microcks, microservices testing, open-source automation tools, Pact, platform engineering, record and replay testing, service dependencies, test automation, Testcontainers, VCR, VCR.py, WireMock
Stop Treating GPUs Like Web Pods
Kubernetes schedules accelerators as opaque integers, and your bill pays for it. Share the silicon, scale on the right signal and keep weights out of the image ...
Veera Ravindra Divi | | AI infrastructure, AI serving, autoscaling, cloud costs, cloud native AI, DCGM exporter, DRA, Dynamic Resource Allocation, GPU costs, GPU scheduling, GPU sharing, GPU utilization, GPUs, inference workloads, KEDA, kubernetes, Kubernetes GPU scheduling, LLM Inference, MIG, model weights, MPS, NVIDIA GPUs, NVIDIA MIG, Prometheus, scale-to-zero, time-slicing
A Green Kubernetes Deployment Does Not Mean a Healthy Application
The deployment finishes, kubectl rollout status reports success, and every pod shows Running and Ready. For most teams, that is the moment the release is considered done. Then a customer transaction fails ...
The Foundation Was Already Poured
Techstrong's Experts Exchange this October, Cloud Native Now: The AI Stack, and this November's KubeCon in Salt Lake City are both making the same case for cloud native and AI. The argument ...
Alan Shimel | | agent governance, agent identity, agentic AI, AI agents, AI governance, AI infrastructure, AI security, AI stack, AI strategy, AI Workloads, cloud native, cloud native developers, cncf, enterprise AI, GPU scheduling, KubeCon, kubernetes, Kubernetes AI, MLOps, model serving, observability, platform engineering, sigstore, SLSA, software supply chain security
Rust Rewrite Readies Kata Containers for Agent Sandboxing
The OpenInfra Foundation has released version 4 of Kata Containers, which arrives with a new Rust-based default runtime that brings new memory safety and performance gains. With the Rust rewrite, the Foundation ...
Beyond the Model: Why AI Agent Orchestration Requires Cloud-Native Engineering
AI agent orchestration is a distributed systems challenge. Cloud-native engineering provides the resilience, observability, security and scalability needed for production AI ...
Nithiya Dharshini | | agent communication, agentic AI infrastructure, AI agent orchestration, AI governance, AI infrastructure, AI observability, AI scalability, AI security, AI workflow monitoring, AI workload management, automated scaling, cloud native AI, cloud-native engineering, containerized AI services, distributed AI systems, distributed tracing, enterprise AI agents, event-driven architecture, GitOps, Kubernetes for AI, multi-agent systems, platform engineering, production AI systems, resilient AI systems

