Autoscaling Simulator — Target Utilization, Startup Delay & Stabilization

Interactive autoscaling simulator — apply a demand surge, set target utilization, replica startup delay and maximum replicas, and watch the controller, request backlog and downscale stabilization window respond.

← Cloud Computing Labs
About this tool — how it works & FAQOpen ▾Close ▴

About the Autoscaling Simulator

This simulator connects a demand trace to a pool of replicas through a metrics sampler and a replica controller. Each ready replica can serve a fixed number of requests per second; whatever demand exceeds capacity collects as backlog. You adjust the controller and the environment and watch capacity chase demand.

What the simulator shows

• A 3D scene containing the demand trace, the metrics sampler, the replica controller, the starting and ready replica pool, and an unserved request reservoir. • Controls for base demand (2 to 24 requests per second), a switch to triple demand from 6 to 18 seconds, target utilization (40 to 90 percent), new replica startup delay (1 to 6 seconds), maximum replicas (2 to 8) and an 8-second downscale stabilization checkbox. • Readouts for ready replicas, desired replicas, request backlog, measured utilization, service throughput and a request-balance residual that should stay at zero. • Experiments for a demand surge, a long startup delay and a capacity ceiling.

How the controller decides

Every ready replica serves up to 4 requests per second, and the backlog changes by arrivals minus requests served. The controller computes desired replicas as the ceiling of ready replicas times measured utilization divided by target utilization, sampled on a two-second control interval. A scale-up decision does not create capacity instantly: new replicas must finish their startup delay first, so the backlog can keep growing after the decision. When the stabilization option is on, downscaling uses the maximum recommendation from the previous 8 seconds, which prevents the controller from removing replicas after a brief dip.

What the model is and is not

This is a deterministic fluid service model inspired by utilization-based autoscaling. The two-second control interval and eight-second stabilization are teaching settings, not the defaults of any product, and the model has no CPU request calibration, metric errors, readiness exclusions or scale-rate policies. The capacity ceiling experiment makes one point plainly: if maximum replicas cannot cover demand, no controller setting removes the backlog.

Frequently asked questions

Does a scale-up decision instantly add capacity?

No. New replicas must complete their startup delay before they serve requests, so desired replicas can exceed ready replicas for several seconds. Lengthen the startup delay in the lab to see the backlog grow while the pool warms.

How are desired replicas calculated here?

Desired equals the ceiling of ready replicas multiplied by measured utilization, divided by the target utilization. If measured utilization is above target the result exceeds the ready count, and the controller asks for more.

Why stabilize downscaling?

To reduce rapid removal of replicas after a brief dip in demand. With the 8-second stabilization window enabled, the controller keeps the highest recommendation it made in the previous 8 seconds, so capacity is retained through short lulls.

Why does the backlog keep growing at the maximum replica count?

Because demand exceeds the service capacity that the maximum replicas can provide. Each ready replica serves at most 4 requests per second, so once the ceiling is reached the unserved difference accumulates.

Related tools & guides