Load Balancing Simulator — Round Robin vs Least Outstanding & Health Checks

Interactive load balancer simulator — route request tokens to three backends with different service times, compare round robin with least-outstanding routing, and see what a delayed health check does when a server fails.

← Cloud Computing Labs
About this tool — how it works & FAQOpen ▾Close ▴

About the Load Balancing Simulator

This simulator follows individual request tokens from a client through a load balancer to three servers that work at different speeds. You choose the arrival rate and the routing policy, optionally make Backend 2 fail, and see how queues, latency and failures develop.

What the simulator shows

• A 3D scene with a client request source, a load balancer and health monitor, and Backend 1 (0.3 s), Backend 2 (1.2 s) and Backend 3 (0.6 s) service times. • Controls for arrival rate (1 to 8 requests per second), routing policy (round robin or least outstanding), whether Backend 2 fails after 5 seconds, and the failure detection delay (0 to 4 seconds). • Readouts for requests received, completed requests, rejected or failed requests, outstanding requests, mean completed latency and the number of advertised healthy backends. • Three experiments: round-robin overload, least-outstanding routing and delayed failure detection.

Routing policy and failure detection

Round robin rotates through healthy targets without looking at their queues, so the slow 1.2 s backend builds a longer queue than the fast ones when load rises. Least-outstanding routing sends each request to the shortest queue and adapts to the different service rates, which lowers mean latency under the same arrival rate. The accounting identity is received equals completed plus failed plus outstanding. When Backend 2 fails, its outstanding requests fail at once, but until the health check detects the failure the balancer still advertises it and new requests routed there also fail — the longer the detection delay, the more failures.

What the model is and is not

Each backend is a single FIFO server with eight outstanding slots, with no retries, no transport latency and no connection draining. Service times and capacities are teaching parameters. The lab isolates routing policy and health-check delay so the effect of each is visible, rather than modeling a specific cloud load balancer product.

Frequently asked questions

Does round robin account for how long each request takes?

No. Round robin rotates across healthy targets without observing their current queues, so a slow backend can accumulate a backlog while faster ones sit idle. Switch to least-outstanding in the lab to compare.

Why would a failed server still receive traffic?

Health checks take time to detect a failure. During the detection window the load balancer still advertises the server as healthy, so new requests continue to be routed to it and fail. Setting the failure detection delay to 0 removes that window in the model.

What does least-outstanding routing do?

It chooses the backend with the fewest requests currently outstanding, so faster backends naturally receive more of the traffic. In the lab it adapts to the 0.3 s, 1.2 s and 0.6 s service times and generally lowers mean completed latency.

What happens when a backend's queue is full?

Each backend holds at most eight outstanding requests. Requests that cannot be queued are counted as rejected or failed, which is what you see when a high arrival rate overloads the slower servers.

Related tools & guides