platform engineering
Containers Became the Unit of Speed. AI Agents Are Making VMs the Unit of Trust
The container is not disappearing. But as autonomous agents generate code, install packages and invoke tools, cloud-native infrastructure is placing a VM-grade security boundary around it ...
Alan Shimel | | Agent Sandboxing, agent security, agentic AI, AI agents, AI infrastructure, autonomous agents, cloud native security, container security, containers, Docker Sandboxes, Firecracker, gVisor, Kata Containers, kubernetes, MCP security, MicroVMs, platform engineering, RuntimeClass, secure execution, virtualization, workload isolation, zero-trust
Autoscaling AI Workloads on Kubernetes With KEDA and What it Means for Agentic Systems
KEDA can scale Kubernetes AI workloads on real demand signals such as queue depth, helping model-serving and agent workloads respond faster while reducing idle compute costs ...
Kishor Patil | | agentic AI, AI agents, AI infrastructure, AI model serving, cloud native AI, devops, Event-Driven Autoscaling, horizontal pod autoscaling, HPA, inference scaling, KEDA, Kubernetes AI workloads, Kubernetes autoscaling, Kubernetes Event-Driven Autoscaling, platform engineering, Pub/Sub, queue depth, RabbitMQ, Redis, scale-to-zero, SQS
Cloud-Native Complexity Is a Cost: When More Platform Layers Stop Adding Value
Cloud-native environments rarely become complex overnight. In most teams, complexity builds gradually. A platform may begin with containers and a basic deployment process, then grow to include orchestration, CI/CD, observability, security controls, ...
Kubeflow’s Graduation Is a Vote for Kubernetes as the AI Control Plane
Kubeflow’s CNCF graduation signals growing confidence in Kubernetes as a common control plane for production AI workloads, from training and pipelines to governance and inference ...
Alan Shimel | | agentic AI, AI infrastructure, AI lifecycle, AI platform, AI Workloads, cloud native AI, cncf, Distributed Training, enterprise AI, GPU scheduling, KServe, Kubeflow, Kubeflow graduation, Kubeflow Pipelines, Kubeflow Trainer, kubernetes, Kubernetes AI, MLOps, OpenTelemetry, platform engineering
How Open-Source Automation Tools Handle the Testing Problem That Cloud-Native Independent Deployment Creates
Independent deployment creates a coverage currency problem manual maintenance cannot scale to address. Learn how open-source automation tools handle it structurally. ...
Sancharini Panda | | API mocking, behavioral drift, CI/CD testing, Cloud-Native Testing, contract testing, coverage currency, eBPF, go-vcr, independent deployment, integration test fixtures, integration testing, Keploy, Kubernetes testing, Microcks, microservices testing, open-source automation tools, Pact, platform engineering, record and replay testing, service dependencies, test automation, Testcontainers, VCR, VCR.py, WireMock
The Telemetry Debt Crisis: Why Cloud-Native Teams are Optimizing the Wrong Metric
Telemetry debt is overwhelming engineering teams with noisy alerts, unused dashboards and rising observability costs. Here’s how to identify, reduce and prevent it ...
David Iyanu Jonathan | | adaptive sampling, AI observability, alert fatigue, dashboard sprawl, eBPF observability, FinOps, incident response, log management, metric cardinality, MTTR, observability as code, observability costs, observability maturity, observability strategy, OpenTelemetry, platform engineering, telemetry governance, telemetry ROI, trace data
A Green Kubernetes Deployment Does Not Mean a Healthy Application
The deployment finishes, kubectl rollout status reports success, and every pod shows Running and Ready. For most teams, that is the moment the release is considered done. Then a customer transaction fails ...
The Foundation Was Already Poured
Techstrong's Experts Exchange this October, Cloud Native Now: The AI Stack, and this November's KubeCon in Salt Lake City are both making the same case for cloud native and AI. The argument ...
Alan Shimel | | agent governance, agent identity, agentic AI, AI agents, AI governance, AI infrastructure, AI security, AI stack, AI strategy, AI Workloads, cloud native, cloud native developers, cncf, enterprise AI, GPU scheduling, KubeCon, kubernetes, Kubernetes AI, MLOps, model serving, observability, platform engineering, sigstore, SLSA, software supply chain security
Beyond the Model: Why AI Agent Orchestration Requires Cloud-Native Engineering
AI agent orchestration is a distributed systems challenge. Cloud-native engineering provides the resilience, observability, security and scalability needed for production AI ...
Nithiya Dharshini | | agent communication, agentic AI infrastructure, AI agent orchestration, AI governance, AI infrastructure, AI observability, AI scalability, AI security, AI workflow monitoring, AI workload management, automated scaling, cloud native AI, cloud-native engineering, containerized AI services, distributed AI systems, distributed tracing, enterprise AI agents, event-driven architecture, GitOps, Kubernetes for AI, multi-agent systems, platform engineering, production AI systems, resilient AI systems
Cloud-Native’s Interest Payment Just Came Due
We spent a decade telling each other that cloud-native was how you move fast. Break the monolith into services. Put everything in containers. Declare your infrastructure. Add a service mesh, a GitOps ...

