Saturday, September 26, 2026
Cloud Native Now
MENU
MENU
Home
Webinars
Upcoming
On-Demand
Calendar View
Podcasts
Cloud Native Now Podcast
Techstrong.tv Podcast
Techstrong.tv - Twitch
About
Sponsor
MENU
MENU
News
Latest News
News Releases
Cloud-Native Development
Cloud-Native Platforms
Cloud-Native Networking
Cloud-Native Security
time-slicing
Stop Treating GPUs Like Web Pods
Kubernetes schedules accelerators as opaque integers, and your bill pays for it. Share the silicon, scale on the right signal and keep weights out of the image ...
Veera Ravindra Divi
|
August 12, 2026
|
AI infrastructure
,
AI serving
,
autoscaling
,
cloud costs
,
cloud native AI
,
DCGM exporter
,
DRA
,
Dynamic Resource Allocation
,
GPU costs
,
GPU scheduling
,
GPU sharing
,
GPU utilization
,
GPUs
,
inference workloads
,
KEDA
,
kubernetes
,
Kubernetes GPU scheduling
,
LLM Inference
,
MIG
,
model weights
,
MPS
,
NVIDIA GPUs
,
NVIDIA MIG
,
Prometheus
,
scale-to-zero
,
time-slicing
×