observability
Cloud Observability Is More Than a Cloud-Native Story
Cloud-native systems have come to define much of the public conversation about observability. Discussions often begin with Kubernetes, microservices, OpenTelemetry, and distributed tracing. But enterprise teams are responsible for a much wider ...
Your Service Is Healthy, but Its Data Isn’t: Rethinking Cloud-Native Health Checks
Some of the most difficult production issues do not start with a failed service. The application is running, the database is reachable, requests are completed successfully and the monitoring dashboard is green ...
Write Access Is the Easy Part: The Verification Gap in Agentic Kubernetes Remediation
Giving an AI agent the power to change a cluster is now straightforward. Confirming the change landed, did not produce unintended duplicate effects and achieved the outcome the operator actually wanted is ...
DataAgent Emerges From Stealth To Bring Autonomous Remediation to Kubernetes
DataAgent emerged from stealth today with $10 million in pre-seed funding and an agentic AI platform designed to fix production problems inside Kubernetes environments without waiting for a site reliability engineer to ...
The Foundation Was Already Poured
Techstrong's Experts Exchange this October, Cloud Native Now: The AI Stack, and this November's KubeCon in Salt Lake City are both making the same case for cloud native and AI. The argument ...
Alan Shimel | | agent governance, agent identity, agentic AI, AI agents, AI governance, AI infrastructure, AI security, AI stack, AI strategy, AI Workloads, cloud native, cloud native developers, cncf, enterprise AI, GPU scheduling, KubeCon, kubernetes, Kubernetes AI, MLOps, model serving, observability, platform engineering, sigstore, SLSA, software supply chain security
When Your Cluster Won’t Sit Still: The Hidden Cost of Kubernetes Autonomy During Incidents
I’ve spent the better part of the last few years on the receiving end of Kubernetes pages, both as an operator and as someone building tooling for platform teams. The pattern I’ve ...
Stop Treating Your Models Like Microservices
A few years ago, it felt like Kubernetes had become the universal answer to infrastructure problems. Teams wanted resiliency? Kubernetes. Faster deployments? Kubernetes. Scalability? Kubernetes again. Eventually, the industry stopped treating cloud-native ...
Why Observability is Critical for Modern Cloud‑Native Systems
In the future, observability will be a key factor for any organization looking to succeed with the concept of cloud native architectures ...
Istio Weaves ‘Future-Ready’ Service Mesh for AI
At KubeCon + CNC 2026, Istio unveils Ambient Multicluster and the Gateway API Inference Extension to simplify AI infrastructure. Learn how sidecar-less mesh and agentgateway secure agentic workloads and boost deployment velocity ...
Adrian Bridgwater | | agentgateway, AI infrastructure, AI Workloads, Ambient Multi-cluster, cloud native, cncf, data plane, Gateway API Inference Extension, generative AI, Istio, KubeCon 2026, kubernetes, microservices, Node Proxy, observability, platform engineering, service mesh, Sidecar-less Mesh, traffic management, Waypoint Proxy
Why Your Kubernetes Network is Still a Black Box — And How to Fix It
Kubernetes networking failures are hard to diagnose. Learn how eBPF and Microsoft Retina provide real-time network observability across your cluster ...

