Kubernetes Orchestration Simulator — Pod Scheduling, Requests & Node Failure

Interactive Kubernetes scheduling simulator — set desired replicas and per-pod CPU and memory requests, watch pending pods placed onto three worker nodes, then fail a worker and see eviction and rescheduling.

← Cloud Computing Labs
About this tool — how it works & FAQOpen ▾Close ▴

About the Kubernetes Orchestration Simulator

This simulator shows how an orchestrator turns a desired replica count into placed, running pods. Pods wait in a pending queue, a simple first-fit scheduler places one every 0.5 seconds using their resource requests, and each placed pod takes two seconds to become ready. Take a worker away and watch the cluster respond.

What the simulator shows

• A 3D cluster with an API and reconciliation controller, a pending pod queue and Worker 1, Worker 2 and Worker 3. • Controls for desired replicas (1 to 10), CPU request per pod (0.5 to 3 cores), memory request per pod (1 to 4 GiB) and a Worker 2 unavailable checkbox, with Restart trial. • Readouts for ready pods, pending pods, starting pods, reserved cluster CPU, reserved cluster memory and evicted pods. • Three experiments: fill the cluster so all ten pods fit, create a resource shortage where only three fit and seven stay pending, and lose a worker.

Requests, capacity and readiness

Each worker has 4 CPU cores and 8 GiB of memory, and a pod fits only if the summed requests on the node stay within both limits. Scheduling uses the requested amounts, not observed low utilization, so a generous request reserves capacity even if the pod is idle. Placement and readiness are separate events: one pending pod is placed every 0.5 seconds, then it starts for two seconds before counting as ready, and ready plus starting plus pending always equals the desired replica count. When Worker 2 becomes unavailable, its pods are evicted and go back into the pending queue for the remaining workers.

What the model is and is not

The scheduler is a simplified, deterministic first-fit policy with immediate failure detection and eviction. It does not implement real Kubernetes scoring, taints, affinity, image pulls, probes or production eviction delays, and the timings are teaching parameters. Use it to build the core intuition — requests reserve capacity, pending means no node fits, and the controller keeps reconciling — rather than to predict the behavior of a particular cluster.

Frequently asked questions

Does the scheduler place pods based on requests or actual usage?

On requests. Requested resources reserve scheduling capacity on a node, which is why the lab lets you raise CPU and memory requests until pods can no longer fit and remain pending, even though nothing is actually running.

Is a placed pod immediately ready?

No. The model separates placement from a two-second startup. A pod shows as starting after it is placed and becomes ready only when the startup completes, so ready, starting and pending counts move at different times.

What happens when Worker 2 becomes unavailable?

Its pods are evicted, the evicted-pods readout increases, and the replacements return to the pending queue to be scheduled onto the healthy workers within their capacities. If the remaining capacity is too small, some pods stay pending.

Why do some pods stay pending?

Pending means no worker has enough unreserved CPU or memory for that pod's requests. With large requests, only three pods fit across the three workers in the lab and seven remain pending until you lower the requests or reduce replicas.

Related tools & guides