CloudBolt Adds Ability to Optimize GPU Consumption by Kubernetes Workload
TL;DR — Key Takeaways
– CloudBolt has introduced a StormForge capability that provides workload-level visibility into GPU utilization and memory consumption across Kubernetes clusters.
– The platform maps GPU processes to Kubernetes pods, enabling IT teams to track resource consumption and costs by cluster, namespace and workload.
CloudBolt Software this week added an ability to optimize consumption of graphics processing units (GPUs) running on Kubernetes clusters at the individual workload level.
Company COO Yasmin Rajabi said the challenge IT teams are running into is that the Data Center GPU Manager (DCGM) tool provided by NVIDIA today only provides insights into utilization and memory per physical device rather than by the workload. DCGM also can only be invoked one node at a time and there is no history generated that can be accessed later, she noted.
The StormForge IT automation platform from CloudBolt instead tracks GPU state per process and maps those processes back to pods to reconcile how resources are actually allocated. As an extension of that platform, it becomes possible to track GPU use and memory per workload, even on time-sliced GPUs where the standard exporter can’t, and map it back to the pods actually running. Armed with insights into GPU costs broken down by cluster, namespace, and workload, the StormForge platform then recommends an optimization at the node level, she added.
In the absence of that capability, each IT team would otherwise have to manually map those workloads in a way that still doesn’t provide the level of visibility needed, noted Rajabi.
Mitch Ashley, vice president and practice lead for the Futurum Group, said IT teams can’t manage GPU spend that they can’t attribute to a workload. Closing that gap is the prerequisite for chargeback, capacity planning, and deciding which AI projects earn more hardware, he added.
It’s not clear just how many AI workloads are running on Kubernetes clusters in production environments, but as they increase, so too will the need to optimize consumption of highly limited resources. In fact, before too long, the dominant class of workloads running on Kubernetes clusters might in some way or another be AI-related.
In general, IT teams are under more pressure than ever to optimize consumption of IT infrastructure. As providers of AI services increase demand for processors and memory rises, the cost of a limited supply of physical IT infrastructure is increasing. At the same time, IT teams are starting to deploy more AI applications on Kubernetes clusters. Unfortunately, because of overprovisioning of clusters, the utilization rates of GPUs in servers are often in the single digits, so IT teams now need to find ways to share limited IT resources across multiple Kubernetes clusters.
At the moment, most IT teams recognize there is a current level of chaos that they are now moving to rein in. The challenge is that, given the limited IT infrastructure resources, many IT teams will need to start prioritizing their allocation by workload. Not all those workloads, however, will require access to the latest GPU, so IT teams will need to be able to route requests across clusters that are invoking a wide range of GPUs, other classes of AI accelerators and even traditional CPUs.
Hopefully, there will come a day soon when AI agents automatically right size Kubernetes clusters. In the meantime, however, both IT administrators and those AI agents will need access to reliable telemetry data to automate the management of Kubernetes clusters that, while powerful, remain the most challenging platform to manage in the modern enterprise.
Frequently Asked Questions
What is CloudBolt's new GPU optimization capability?
CloudBolt has added a capability to its StormForge platform that tracks GPU utilization and memory consumption at the individual workload level across Kubernetes clusters, helping IT teams identify inefficiencies and optimize resource allocation.
How does StormForge improve GPU visibility?
StormForge tracks GPU activity by process and maps those processes to Kubernetes pods. This enables organizations to understand how GPU resources are consumed by individual workloads, including those running on time-sliced GPUs.
How does CloudBolt help reduce GPU costs?
The platform provides insights into GPU costs by cluster, namespace and workload, then recommends better-fit GPU node types to help reduce overprovisioning and improve infrastructure efficiency.



