When the Scheduler is Full but the Autoscaler Is Not
TL;DR — Key Takeaways
- A successful CPU-triggered scale-out proved the Azure VMSS path worked, but downsizing workers exposed a different limit: Docker Swarm could run out of schedulable capacity before host metrics looked stressed.
- Resource reservations describe capacity the scheduler must commit, not actual runtime usage, so Azure Monitor and Swarm can disagree while both remain technically correct.
- Autoscaling design should test scheduler pressure directly and combine infrastructure metrics with signals showing when tasks cannot be placed.
Our autoscaling design looked healthy in the first test. Docker Swarm workers ran in an Azure Virtual Machine Scale Set, Azure Monitor watched infrastructure metrics and a CPU spike caused a new worker to be created and join the cluster. From the outside, the system behaved exactly as expected.
Then we reduced the worker size from 16 GiB to 8 GiB while keeping the existing Docker service reservations. Around 6.5 GiB of memory was already reserved across the workload we were testing, and several services could no longer be placed normally. The surprising part was that the infrastructure autoscaler did not necessarily see the same shortage that the Swarm scheduler saw.
That test changed how I think about autoscaling: Host utilization and schedulable capacity are related, but they are not the same signal.
The First Test Proved the Scale-Out Path
Our production-oriented design used CPU, memory and queue signals with cooldown periods to avoid flapping. The CPU rule was designed around sustained high utilization, and the scale-out path added one worker at a time. During validation, we shortened the observation window so we could prove the end-to-end behavior without waiting through the full production threshold.
The important result was not the exact test duration. It was that the complete path worked: The metric crossed the threshold, Azure created a new VMSS instance, the startup process joined it to the Swarm and the cluster gained another worker.
A representative CPU rule looked like this:
az monitor autoscale rule create \
–resource-group <resource-group> \
–autoscale-name swarm-autoscale \
–condition “Percentage CPU > 80 avg 10m” \
–scale out 1 \
–cooldown 5
Downsizing Exposed a Different Limit
The next test changed the worker shape. The larger worker class provided 16 GiB of memory; the smaller option provided 8 GiB. We kept the service reservations in place because that was the point of the test:
Could the existing scheduling policy still work on the smaller node size?
It could not. Multiple services were unable to run normally even though the VM-level memory signal did not necessarily resemble an exhausted host.
One of the reservation patterns in the stack was straightforward:
deploy:
resources:
reservations:
cpus: “0.25”
memory: 512M
placement:
constraints:
– “node.role==worker”
Reserved Capacity Isn’t Runtime Usage
That YAML is a scheduling instruction. A 512 MiB reservation tells Swarm that the task should only be placed where that amount of memory is available to reserve. It does not mean the container is continuously consuming 512 MiB at runtime.
This is where the two control loops diverged. Swarm was reasoning about committed scheduling capacity. Azure Monitor was reasoning about the metrics configured on the VM or scale set, such as CPU or host memory. If a task cannot be scheduled because its reservation cannot be satisfied, it may never create the runtime pressure that a host-utilization rule is waiting to observe.
In other words, the cluster can be full from the scheduler’s perspective while the VM still looks healthy enough to the autoscaler. Both systems can be correct and still disagree about whether more capacity is needed.
The Diagnostic Signal Has to Come From the Scheduler Too
When this happens, looking only at host CPU or memory is not enough. The fastest way to confirm the mismatch is to inspect the desired versus running service state and the reason a task is pending.
A minimal diagnostic sequence is:
docker service ls
docker service ps <service> –no-trunc
docker node inspect –pretty <node>
The key question is whether the service is waiting because the node is actually out of runtime memory or because the scheduler cannot satisfy a declared reservation or placement constraint. Those are different failure modes, and they require different scaling signals.
We also ran a controlled test with the reservations removed on the smaller workers. More services could then be scheduled, which confirmed that reservation pressure was part of the placement problem. However, removing reservations was not the conclusion I wanted to carry into production. Reservations exist to prevent the scheduler from promising capacity that the workload is expected to need.
What I Changed in the Way I Test Autoscaling
I no longer consider a successful CPU scale-out test enough evidence that a container platform is properly autoscaled. I now test four things independently:
- Can every required task be scheduled at the minimum worker count?
- What happens to schedulable capacity when the worker size changes?
- Which metric actually triggers infrastructure scale-out?
- Does that metric react before the scheduler starts leaving tasks pending?
The design lesson is simple: Autoscaling policy and scheduler reservation policy have to be designed together. If the infrastructure layer cannot see the pressure that the orchestrator is experiencing, adding more autoscaling rules around the same host metrics does not close the gap.
The most useful autoscaling signal is not always the one that says a VM is busy. Sometimes it is the one that says the scheduler has work it cannot place.
Frequently Asked Questions
Why did Azure Monitor not necessarily trigger when Swarm could not place services?
Because Azure Monitor was watching host-level utilization while Swarm was enforcing declared CPU and memory reservations. A task can remain pending because reservations cannot be satisfied even when runtime memory use still looks reasonable.
What does a Docker memory reservation mean?
It tells Swarm how much capacity must be available before placing the task. It does not mean the container continuously consumes that amount of memory.
What should platform teams test when validating autoscaling?
They should verify minimum-worker schedulability, test different node sizes, identify the actual scale-out signal and confirm that signal reacts before the scheduler begins leaving tasks pending.


