Monday, September 14, 2026
Cloud Native Now
MENU
MENU
Home
Webinars
Upcoming
On-Demand
Calendar View
Podcasts
Cloud Native Now Podcast
Techstrong.tv Podcast
Techstrong.tv - Twitch
About
Sponsor
MENU
MENU
News
Latest News
News Releases
Cloud-Native Development
Cloud-Native Platforms
Cloud-Native Networking
Cloud-Native Security
GPU autoscaling
Kubernetes Wasn’t Built for GPUs. Make It Behave
Kubernetes counts whole GPUs and treats pods as disposable. An LLM pod is neither. Share the silicon with MIG/MPS/time-slicing and stop paying for idle ...
Sneha Gullapalli
|
August 12, 2026
|
A100
,
AI infrastructure
,
AI Workloads
,
cloud native AI
,
Dynamic Resource Allocation
,
GPU autoscaling
,
GPU cost reduction
,
GPU optimization
,
GPU partitioning
,
GPU sharing
,
GPU time-slicing
,
GPU utilization
,
H100
,
Karpenter
,
KServe
,
Kubernetes DRA
,
Kubernetes GPU scheduling
,
LLM Inference
,
model caching
,
multi-instance GPU
,
NVIDIA GPU Operator
,
NVIDIA MIG
,
NVIDIA MPS
,
scale-to-zero
,
VRAM
×