Monday, September 14, 2026
Cloud Native Now
MENU
MENU
Home
Webinars
Upcoming
On-Demand
Calendar View
Podcasts
Cloud Native Now Podcast
Techstrong.tv Podcast
Techstrong.tv - Twitch
About
Sponsor
MENU
MENU
News
Latest News
News Releases
Cloud-Native Development
Cloud-Native Platforms
Cloud-Native Networking
Cloud-Native Security
A100
Kubernetes Wasn’t Built for GPUs. Make It Behave
Kubernetes counts whole GPUs and treats pods as disposable. An LLM pod is neither. Share the silicon with MIG/MPS/time-slicing and stop paying for idle ...
Sneha Gullapalli
|
August 12, 2026
|
A100
,
AI infrastructure
,
AI Workloads
,
cloud native AI
,
Dynamic Resource Allocation
,
GPU autoscaling
,
GPU cost reduction
,
GPU optimization
,
GPU partitioning
,
GPU sharing
,
GPU time-slicing
,
GPU utilization
,
H100
,
Karpenter
,
KServe
,
Kubernetes DRA
,
Kubernetes GPU scheduling
,
LLM Inference
,
model caching
,
multi-instance GPU
,
NVIDIA GPU Operator
,
NVIDIA MIG
,
NVIDIA MPS
,
scale-to-zero
,
VRAM
×