This simulator connects a demand trace to a pool of replicas through a metrics sampler and a replica controller. Each ready replica can serve a fixed number of requests per second; whatever demand exceeds capacity collects as backlog. You adjust the controller and the environment and watch capacity chase demand.
• A 3D scene containing the demand trace, the metrics sampler, the replica controller, the starting and ready replica pool, and an unserved request reservoir. • Controls for base demand (2 to 24 requests per second), a switch to triple demand from 6 to 18 seconds, target utilization (40 to 90 percent), new replica startup delay (1 to 6 seconds), maximum replicas (2 to 8) and an 8-second downscale stabilization checkbox. • Readouts for ready replicas, desired replicas, request backlog, measured utilization, service throughput and a request-balance residual that should stay at zero. • Experiments for a demand surge, a long startup delay and a capacity ceiling.
Every ready replica serves up to 4 requests per second, and the backlog changes by arrivals minus requests served. The controller computes desired replicas as the ceiling of ready replicas times measured utilization divided by target utilization, sampled on a two-second control interval. A scale-up decision does not create capacity instantly: new replicas must finish their startup delay first, so the backlog can keep growing after the decision. When the stabilization option is on, downscaling uses the maximum recommendation from the previous 8 seconds, which prevents the controller from removing replicas after a brief dip.
This is a deterministic fluid service model inspired by utilization-based autoscaling. The two-second control interval and eight-second stabilization are teaching settings, not the defaults of any product, and the model has no CPU request calibration, metric errors, readiness exclusions or scale-rate policies. The capacity ceiling experiment makes one point plainly: if maximum replicas cannot cover demand, no controller setting removes the backlog.
No. New replicas must complete their startup delay before they serve requests, so desired replicas can exceed ready replicas for several seconds. Lengthen the startup delay in the lab to see the backlog grow while the pool warms.
Desired equals the ceiling of ready replicas multiplied by measured utilization, divided by the target utilization. If measured utilization is above target the result exceeds the ready count, and the controller asks for more.
To reduce rapid removal of replicas after a brief dip in demand. With the 8-second stabilization window enabled, the controller keeps the highest recommendation it made in the previous 8 seconds, so capacity is retained through short lulls.
Because demand exceeds the service capacity that the maximum replicas can provide. Each ready replica serves at most 4 requests per second, so once the ceiling is reached the unserved difference accumulates.