Friday, August 14, 2026
Cloud Native Now

Cloud Native Now


MENUMENU
  • Home
  • Webinars
    • Upcoming
    • Calendar View
    • On-Demand
  • Podcasts
    • Cloud Native Now Podcast
    • Techstrong.tv Podcast
    • Techstrong.tv - Twitch
  • About
  • Sponsor
MENUMENU
  • News
    • Latest News
    • News Releases
  • Cloud-Native Development
  • Cloud-Native Platforms
  • Cloud-Native Networking
  • Cloud-Native Security
Cloud-Native Networking Container/Kubernetes Management Contributed Content Service Mesh Social - Facebook Social - LinkedIn Social - X 

The Hidden Cost of “Just Works” Load Balancing in a Service Mesh

August 14, 2026August 14, 2026 Sai Aneesh Mullapudi aws, Istio, kubernetes, load balancing, service mesh, site reliability engineering
by Sai Aneesh Mullapudi

If you’re running a multi-AZ Kubernetes cluster with a service mesh on top, there’s a good chance you’re paying a tax you never signed up for, and it won’t show up as a line item on any dashboard you’re already watching.

The Default Nobody Configures

Kubernetes Services and Istio’s Envoy sidecars both default to distributing traffic randomly across all healthy endpoints, with zero awareness of which availability zone a pod happens to live in. That’s a perfectly reasonable default for a single-AZ deployment. But once you spread pods across three AZs for resilience (which is more or less table stakes at this point), that same default quietly becomes a liability. A pod in us-east-1a calling a downstream service now has roughly a two-in-three chance of landing on a pod in a different AZ. Multiply that across a request path touching three or four internal services plus a database replica, and a single user request can rack up multiple cross-AZ hops before a response ever goes out.

Techstrong Gang Youtube

AWS makes this worse by default, too. Cross-zone load balancing on an NLB sitting in front of an Istio ingress gateway will happily route an incoming connection to a target in any AZ, even when a perfectly healthy target is sitting right next to the load balancer node that received the request in the first place.

What It Actually Costs

Two things, concretely.

Latency. In a mid-sized production cluster running roughly 3,500 RPS across three AZs, tracing data showed same-AZ requests landing at 15-18ms p50, while cross-AZ requests to the same service ran 25-30ms p50, 40% to 65% slower. With about two-thirds of requests crossing AZs, the weighted p50 sat around 24ms instead of the ~17ms it could have been. That gap compounds fast on any request path with multiple internal hops.

Money. AWS bills $0.01/GB for traffic crossing AZs within a region. That sounds trivial on a per-request basis, but service-to-service traffic (the kind that doesn’t show up in your external bandwidth numbers) tends to run 5x to 10x the volume of external traffic once you count internal API calls, database replica reads and Kafka consumer traffic. In the environment analyzed here, a conservative accounting of ingress, service-to-service, database and Kafka cross-AZ traffic landed around $600/month before any fix, and that’s before factoring in automatic Aurora cross-AZ replication and log shipping, which push the real number higher in practice. The exact dollar figure depends heavily on your own traffic mix, so pull your own numbers from AWS Cost Explorer’s DataTransfer-Regional-Bytes usage type before treating any of this as gospel, including mine.

The Fix: Locality-Aware, Not Locality-Only

Istio supports locality-aware load balancing through the DestinationRule traffic policy. The naive fix is to route 100% of traffic to the local AZ and 0% everywhere else. Don’t do that. It removes your ability to absorb a bad deploy or a hot pod in the local AZ, and it makes a full AZ failure much harder to recover from cleanly.

The better pattern is a weighted distribution: 80% same-AZ, split 10/10 across the other two, combined with outlier detection so unhealthy endpoints get ejected automatically.

apiVersion: networking.istio.io/v1beta1
kind: DestinationRule
metadata:
  name: checkout-service-locality-lb
spec:
  host: checkout-service.prod.svc.cluster.local
  trafficPolicy:
    loadBalancer:
      localityLbSetting:
        enabled: true
        distribute:
          - from: us-east-1a/*
            to:
              "us-east-1a/*": 80
              "us-east-1b/*": 10
              "us-east-1c/*": 10
    outlierDetection:
      consecutiveErrors: 5
      interval: 30s
      baseEjectionTime: 30s
      maxEjectionPercent: 50

Three things have to be true for this to actually work, and all three are easy to skip:

  1. Even pod distribution across AZs. If one AZ holds 50% of a service’s pods, an 80/10/10 policy sends that AZ far more than half of all traffic. Use topologySpreadConstraints with maxSkew: 1 to keep pod counts balanced. Otherwise the locality policy just relocates the imbalance instead of fixing it.
  2. Ingress gateway pods spread across all AZs, too. Locality-aware routing at the service level doesn’t help if every request enters through ingress pods concentrated in one AZ to begin with.
  3. NLB cross-zone load balancing disabled, via the service.beta.kubernetes.io/aws-load-balancer-cross-zone-load-balancing-enabled: "false" annotation. Skip this and the load balancer itself reintroduces the exact randomness you just eliminated at the mesh layer.

What You Get, and What You Give Up

Latency improves, but the more interesting change is in failure blast radius. Under the old random-distribution model, losing an AZ means the remaining two suddenly absorb 50% more load than they were running a moment earlier, a fast path to cascading saturation. Under an 80/10/10 locality policy, losing an AZ only redistributes the 80% slice that AZ was already handling for itself; the other AZs see a much smaller relative increase, and outlier detection reroutes around the failure automatically.

The tradeoff is operational complexity. Traffic distribution stops being an intuitive 33/33/33, and anyone debugging “why is this AZ getting more traffic than that one” now needs to know the locality policy exists. Document it before you ship it, not after someone gets paged asking why a dashboard looks lopsided.

Before You Roll This Out

Validate in a lower environment first, with realistic load. A k6 or Locust run against a staging cluster is usually enough to surface load imbalance caused by uneven pod counts. Canary on your lowest-traffic services in production before touching anything customer-critical, and watch error rate and p95 latency, not just the cross-AZ percentage, before calling it done.

Locality-aware load balancing isn’t exotic. It’s an Istio feature that’s existed for years, and it gets skipped mostly because the mesh works fine without it, right up until someone reconciles the AWS data transfer line item or gets paged during an AZ event. Both are worth checking before you assume your multi-AZ cluster is actually behaving like one.

  • Click to share on X (Opens in new window) X
  • Click to share on Facebook (Opens in new window) Facebook
  • Click to share on LinkedIn (Opens in new window) LinkedIn
  • Click to share on Reddit (Opens in new window) Reddit

Related

  • ← How We Cut Kubernetes Deployment Validation From 45 Minutes to 2 minutes

Techstrong TV

Click full-screen to enable volume control
Watch latest episodes and shows

Tech Field Day Events

UPCOMING WEBINARS

  • CloudNativeNow.com
  • Error
  • SecurityBoulevard.com
Migrating Apache Solr Workloads to Amazon OpenSearch Service
29 September 2026
Migrating Apache Solr Workloads to Amazon OpenSearch Service
Modernize for the AI Era
16 September 2026
Modernize for the AI Era
The Strategic Imperative: Embracing Agentic B2B Selling
8 September 2026
The Strategic Imperative: Embracing Agentic B2B Selling

RSS Error: Retrieved unsupported status code "403"

Closing the Loop on AI Coders: 3,200 Vulns Fixed with Zero Human Triage
1 October 2026
Closing the Loop on AI Coders: 3,200 Vulns Fixed with Zero Human Triage
Agentic Runtime Security on AWS: Identity, Least Privilege, and Audit for AI Agents with IBM Verify Identity Access and HashiCorp Vault
29 September 2026
Agentic Runtime Security on AWS: Identity, Least Privilege, and Audit for AI Agents with IBM Verify Identity Access and HashiCorp Vault
Migrate and Modernize your applications with at Zscaler Zero Trust Security Platform on AWS
17 September 2026
Migrate and Modernize your applications with at Zscaler Zero Trust Security Platform on AWS

Podcast


Listen to all of our podcasts

Press Releases

ThreatHunter.ai Halts Hundreds of Attacks in the past 48 hours: Combating Ransomware and Nation-State Cyber Threats Head-On

ThreatHunter.ai Halts Hundreds of Attacks in the past 48 hours: Combating Ransomware and Nation-State Cyber Threats Head-On

Deloitte Partners with Memcyco to Combat ATO and Other Online Attacks with Real-Time Digital Impersonation Protection Solutions

Deloitte Partners with Memcyco to Combat ATO and Other Online Attacks with Real-Time Digital Impersonation Protection Solutions

SUBSCRIBE TO CNN NEWSLETTER

MOST READ

BellSoft Rings Change Bringing Zero-CVE Images to Buildpacks Users

July 21, 2026

Rust Rewrite Readies Kata Containers for Agent Sandboxing

July 24, 2026

NVIDIA Is Putting Real Skin in the Open AI Game

July 28, 2026

Kubernetes Key Management Streamlined by HashiCorp Vault Plug-In

August 6, 2026

The Foundation Was Already Poured

July 27, 2026

RECENT POSTS

The Hidden Cost of “Just Works” Load Balancing in a Service Mesh
Cloud-Native Networking Container/Kubernetes Management Contributed Content Service Mesh Social - Facebook Social - LinkedIn Social - X 

The Hidden Cost of “Just Works” Load Balancing in a Service Mesh

August 14, 2026 Sai Aneesh Mullapudi 0
How We Cut Kubernetes Deployment Validation From 45 Minutes to 2 minutes
Cloud-Native Development Contributed Content Kubernetes Social - Facebook Social - LinkedIn Social - X 

How We Cut Kubernetes Deployment Validation From 45 Minutes to 2 minutes

August 13, 2026 Sai Joshitha Kathari 0
Docker Desktop Gets a Hypervisor of its Own
Cloud-Native Development Containers Docker Features Social - Facebook Social - LinkedIn Social - X Virtualization 

Docker Desktop Gets a Hypervisor of its Own

August 13, 2026 Joab Jackson 0
Kubernetes Wasn’t Built for GPUs. Make It Behave
Contributed Content Kubernetes Social - Facebook Social - LinkedIn Social - X Topics 

Kubernetes Wasn’t Built for GPUs. Make It Behave

August 12, 2026 Sneha Gullapalli 0
How Open-Source Automation Tools Handle the Testing Problem That Cloud-Native Independent Deployment Creates  
Cloud-Native Development Contributed Content Social - Facebook Social - LinkedIn Social - X Topics 

How Open-Source Automation Tools Handle the Testing Problem That Cloud-Native Independent Deployment Creates  

August 12, 2026 Sancharini Panda 0
  • About
  • Media Kit
  • Sponsor Info
  • Write for Cloud Native Now
  • Copyright
  • TOS
  • Privacy Policy
Powered by Techstrong Group
Copyright © 2026 Techstrong Group, Inc. All rights reserved.
×

Modern Software Development and Delivery 

1Q1
2Q2
3Q3
4Q4
5Q5
6Q6
How would you best describe your organization's current software delivery environment(s) on mainframe computers? (Select all that apply)(Required)
Which outcomes are most important to your organization's software delivery strategy today? (Select up to three)(Required)
What, if anything, is limiting your organization's mainframe software delivery progress? (Select up to three)(Required)
In which areas of your mainframe software delivery environment are you currently using AI? (Select all that apply)(Required)
As software delivery responsibilities expand beyond traditional build, test, and deploy, which of the following areas is the most challenging for your organization today with respect to mainframe software delivery? (Select one)(Required)
Which of the following do you expect is most likely to accelerate your organization's mainframe software delivery progress over the next 12-18 months? (Select one)(Required)

×