A traffic peak is 30 minutes away, and hello-web needs one more replica. The new Pod is created, but it stays Pending. You check the monitoring, and the Nodes' CPU usage is clearly low. What would you change first?
To answer that, we first need to know what the Scheduler actually looks at when it chooses a Node for a Pod. This article breaks down how the Scheduler decides, then comes back to this question.
1. After a Pod is created, how does the Scheduler pick it up?
When a Deployment needs a new replica, a new Pod object first appears in the api-server (to see how it gets created, read How does a Deployment become Pods?).
The Scheduler keeps watching the api-server for Pods that do not have a Node yet, and finds a place for each of them in turn. How the Scheduler watches Pods When a Pod's turn comes, the Scheduler has only one question to answer: which Node should this Pod go to?
2. First find where it fits, then choose the best fit
The Scheduler chooses a Node for a Pod in two steps. First it asks "which Nodes cannot take this Pod at all?" and removes the Nodes that fail the required conditions. This step is called Filtering. If several Nodes are left, it asks "which one fits best?" and gives each candidate a score. This step is called Scoring. How the Scheduler selects a Node
This order leads to a direct conclusion: if a Pod cannot be scheduled, the problem is always in Filtering. Scoring only happens when there are candidates. No score can bring back a Node that has already been removed. So for the stuck replica at the start, the question is "which required conditions blocked it?"
Required conditions: three common gates
Filtering checks many kinds of conditions. In practice, these three block Pods most often:
- Resource room: the Pod declares that it needs 1 CPU, but the Node only has 0.5 left, so it does not fit. How "what is left" is calculated is the focus of the next section, and it is also the answer to "the CPU is clearly idle" from the start.
- Taints and tolerations: a Node can set its own gate. For example, a Node with GPUs gets the taint
dedicated=gpu:NoSchedule, which means "keep this Node for work that needs GPUs." Pods without a matching toleration are removed. A toleration only removes this block so the Pod can become a candidate. The Scheduler does not have to choose that Node. Taints and tolerations - Required node affinity: a Pod can also choose Nodes from its side. If a training job requires
accelerator=gpuinrequiredDuringSchedulingIgnoredDuringExecution, only Nodes with that label can become candidates. Node affinity
The last two gates work in opposite directions. A taint lets a Node keep out Pods that should not come. Affinity lets a Pod choose the Nodes it wants. So when you want a group of Nodes to run only a certain kind of work, the official docs suggest using both. With only a taint, the work with the toleration may still be scheduled on other Nodes. With only affinity, other Pods can still take up those Nodes. Examples: dedicated Nodes and special hardware
If a Node fails any one gate, it is removed. So whether a Pod can be scheduled does not depend on "how much CPU the whole cluster has left." It depends on whether one Node passes every gate at the same time.
Preferences: how Scoring chooses
If more than one Node passes Filtering, the Scheduler gives each one a score. The score comes from several scoring items, such as preferred node affinity (a Pod's "better if it matches" preference) and other factors like how resources are spread. They are combined by weight, and the Node with the highest total wins. Filter and Score
Keep required and preferred apart. Required decides "can this Node be a candidate?" Preferred only adds points between candidates. Preferring a Node does not guarantee it will be chosen, because the total also depends on other items.
Of the three gates, taints and affinity can be judged by reading the settings. The confusing one is resources: if the monitoring shows the Node is idle, why does the Pod "not fit"?
3. The CPU is idle, so why does the Pod not fit?
To decide whether a Pod fits, the Scheduler does not look at the usage in your monitoring. It looks at the requests. Each Node has an amount of CPU that Pods can use, called Allocatable. Each Pod already on the Node has declared a request. Subtract the requests from Allocatable, and you get what is left.
| One Node's capacity | CPU |
|---|---|
| Allocatable | 4 |
| Total requests of Pods already on it | 3.5 |
| Left | 0.5 |
| New Pod's request | 1 |
The new Pod needs 1, and this Node has only 0.5 left, so it is removed. Even if monitoring shows this Node actually using only 0.4 CPU, the Scheduler does not change "3.5 already counted" to 0.4. CPU capacity check
If monitoring already shows live usage, why not schedule based on how busy each Node is right now? There are three main reasons:
- Usage changes. A Node that is idle right now may just be waiting for its workloads to reach their peak. Five minutes later, it may not be idle.
- A request is a promise. When there is not enough CPU to go around, containers get CPU in proportion to their requests. Scheduling makes sure these promises do not add up to more than the Node can give; otherwise the promises would mean nothing. How requests are applied at run time
- Scheduling is a one-time decision. Once a Pod is placed, it is not moved when usage changes. So the decision has to hold up for a while.
That is why Kubernetes uses requests as the basis for capacity checks by default, and a moment of idle CPU does not become more room for scheduling. How Pods with resource requests are scheduled Load-aware scheduling does exist, for example Trimaran in scheduler-plugins, but it is an extra plugin you install, not the default behavior.
So "the CPU is idle" and "the requests are full" can both be true at once. That is half of the answer to the question at the start.
A request itself may be too high or too low, and should be reviewed with workload data. But "how this capacity check works" and "whether this request is set well" are two separate questions. This accounting also does not reserve dedicated CPU cores for a container. How CPU is shared at run time is covered in CPU is only 30% used, so why is it throttled?. This example only counts a single regular container; other calculation details are in the further reading.
After Filtering, and after choosing the Node with the highest score, the Scheduler has to hand the decision over.
4. The assignment succeeds, or this attempt fails
Submitting the Node assignment to the api-server is called binding. The default handling sends a Binding request that includes the Pod's identity and the target Node. When it succeeds, the Pod in the api-server records the assignment in .spec.nodeName. DefaultBinder
The kubelet on the target Node watches the api-server for Pods assigned to it, then handles startup. The Scheduler does not need to call the kubelet and say "please start this Pod." Where the kubelet gets its Pods
The diagram shows each component's responsibility, not the full call sequence. After a Node is chosen, other checks may still run; the full flow is in the further reading.
But a place cannot always be found. If no Node is left after Filtering, the Scheduler has nothing to submit. On the usual path where no suitable Node exists, it records a FailedScheduling Event and updates the PodScheduled condition. That condition is the part of the Pod that records whether scheduling is done and why, so we can see what limit this attempt hit. Scheduling failure handling
But the Pod does not lose its chance forever. A new Node, freed resources, or other changed conditions can put it back into the scheduling queue to wait for the next attempt. Retries also use backoff: after a failure, wait for a while, so the same useless work is not repeated over and over. Which changes can trigger a retry depends on the queue and its rules. Do not assume every update causes an immediate retry. Requeueing and QueueingHint
The lab article shows this: after a Node frees up capacity, a Pod that could not be scheduled gets placed on its own, with the same UID and without being recreated.
That completes the Scheduler's decision and handoff: filter, score, count capacity, then pass the assignment to the api-server and the kubelet. Now we can go back to the replica that could not be scheduled.
5. Back to the start: what would you change first?
Let's put these mechanisms into the on-call scenario from the start. This is a made-up exercise based on interview notes, not a record of a real incident.
The hello-web Pod requests 1 CPU, must run on a Node labeled workload=web, and has no tolerations. The cluster has four workers, and this is their current state:
| Node | Label | Taint | CPU left |
|---|---|---|---|
| node-a | workload=web | None | 0.5 |
| node-b | workload=web | dedicated=batch:NoSchedule | 2 |
| node-c | workload=web | None | 0.5 |
| node-d | workload=batch | None | 2 |
Check them against the three gates from section 2. node-a and node-c are web Nodes, but they do not have room for 1 CPU. node-b has room, but it has a taint the new replica does not tolerate. node-d also has room, but its label does not match the required affinity. Each of the four Nodes is blocked by one gate, so the new replica has no candidates at all. The idle CPU in the monitoring is exactly the case from section 3: the requests on node-a and node-c are full, but usage has not gone up yet.
kubectl describe pod shows an Event like this. It is the real output from reproducing the same state in kind (steps in the lab article). scheduling-lab-worker to -worker4 map to node-a through node-d, plus one control-plane Node:
$ kubectl get pod hello-web-new -o wide
NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES
hello-web-new 0/1 Pending 0 1s <none> <none> <none> <none>
$ kubectl describe pod hello-web-new
...
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Warning FailedScheduling 1s default-scheduler 0/5 nodes are available:
1 node(s) didn't match Pod's node affinity/selector,
1 node(s) had untolerated taint {dedicated: batch},
1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: },
2 Insufficient cpu.
no new claims to deallocate,
preemption: 0/5 nodes are available:
2 No preemption victims found for incoming pod,
3 Preemption is not helpful for scheduling.
The message is one long line. It is split at commas and periods here to make it easier to read; the text is unchanged.
The opening 0/5 nodes are available means none of the five Nodes can be used. The reasons follow, grouped and counted: 2 Insufficient cpu is node-a and node-c, the taints are node-b and the control-plane Node, and the affinity mismatch is node-d. Each group points back to a setting to check.
The later parts, no new claims to deallocate and preemption:, are recovery attempts the Scheduler makes after filtering fails: freeing Dynamic Resource Allocation resources, and evicting lower-priority Pods. Neither helps in this scenario. The lab article explains each part.
You have confirmed that the scheduling state and Events point to these limits. The existing replicas are healthy, and error rates and latency have not gotten worse. But a traffic peak is expected in 30 minutes.
A few more conditions for the decision:
- The matching web node pool can scale out. Quota and budget are confirmed, and in the past a new Node took about 10 minutes. That is not a guarantee for this time.
- The batch Nodes still have dedicated needs that must be protected. The web workload's required placement has not been shown to be safe to relax.
- You only have data showing low CPU right now. You do not have enough startup, peak, or load test data to prove the request is too high.
Which option would you pick first? Which blocking condition does it change? What would you need to see to call this a success?
| Option | Action |
|---|---|
| A | Lower the new Pod's CPU request from 1 to 0.5 |
| B | Add capacity with matching web Nodes |
| C | Add a toleration for node-b's taint to the workload |
| D | Change the required node affinity to preferred |
| E | Delete the Pending Pod and let the ReplicaSet Controller recreate it |
Expand: my choice and an evaluation of each option
Under these conditions, I would pick B first and keep checking whether scaling can keep up with the timeline. It directly adds capacity that meets the placement requirements. A, C, and D all need evidence we do not have yet. This is an engineering judgment based on the stated conditions, not an order of steps set by Kubernetes.
| Option | My evaluation | What to watch when verifying or rolling back |
|---|---|---|
| A: lower the request | It lets node-a and node-c pass this CPU check, but low current usage does not prove 0.5 is enough. A lower request also changes how the container competes for CPU at run time, so a successful schedule is not enough. | Test on a small scale only with data to support it, and watch peak latency and resource contention. Rolling back means restoring the workload settings, and the larger request then needs enough capacity. |
| B: add matching capacity | First choice here. The costs are money and waiting time. Adding Nodes that do not match the labels, taints, or other conditions still does not solve the problem. | Check that the Node is Ready, the effective capacity left, and that the Pod is Ready, then look at service metrics. Temporary capacity can be removed later, but first make sure the workloads can move safely. Do not just delete Nodes as a rollback. |
| C: add a toleration | It only removes the block from that taint. First confirm that sharing is allowed and will not interfere with the dedicated workloads. This scenario has no basis for that. | Watch both web and batch metrics. Removing the toleration from the template does not move Pods already running there; you need a separate safe replacement. |
| D: relax the affinity | It may let node-d become a candidate, but it changes the original placement requirement. If it was really only a preference, it can be redesigned; this scenario cannot assume that. | Check that the new location meets the workload's needs. Restoring the required rule in the template does not immediately move existing Pods back. |
| E: recreate the Pod | With the same settings and capacity, it removes no blocking condition. Deleting is also not needed to get a retry. | It replaces the Pod's identity and breaks existing tracking, with no clear benefit. Keep the evidence and deal with the limits instead. |
For the resource meaning of A, see Resource Management. For what C and D allow, and what happens after scheduling, see Taints and Tolerations and Node Affinity.
Real changes should be made to the workload that manages the Pod, such as the Pod template in a Deployment. These options do not mean every field can be patched directly on an existing Pod. If you use GitOps, keep the declared source in sync with the change, and consider the rollout it triggers. Updating a Deployment's Pod template
Success means the service gets the capacity it needs. A value in nodeName only shows the Pod passed assignment. You still need to check that the Pod is Ready, latency and error rates, and whether other workloads were affected. In the lab article, A, C, and D all got the Pod a Node, and each landed exactly where the table above warns about.
Now change one condition: what if the quota is full and there is no time to scale out?
I still would not switch to A automatically. First, find evidence for the alternatives: is there lower-priority work that can be paused and is allowed to give up capacity? Is the affinity just a leftover setting? Is there a reliable measurement for the request? Also prepare a plan to shed traffic or reduce features if capacity is short before the peak. These are service-level decisions. They cannot be ranked only by "which change makes the Pod stop being Pending fastest."
From Node assignment to actual startup
The Scheduler's job ends at assignment: filter, score, then hand the decision to the api-server. When the replica at the start could not be scheduled, the Scheduler was actually working exactly as our conditions told it to. What needs to change is the conditions or the capacity, not the Scheduler.
This article covers the Scheduler up to completing the assignment. How the Pod starts on the Node has not been covered yet. The next question is: after the kubelet sees this Pod, how does it turn the description in the api-server into container processes on the Node?
That is the next handoff to follow.
Further reading: useful limits to know when troubleshooting
Pending does not always mean no Node has been chosen
A Pod's phase is a summary. Pending also covers part of container preparation, such as downloading an image. The STATUS column in kubectl get pods may show a more specific hint, such as SchedulingGated, ContainerCreating, or ImagePullBackOff, which is not always the phase. Pod lifecycle
To find where a Pod is stuck, check .spec.nodeName first. If it is empty, look at scheduling. If it has a value, follow the startup messages on that Node. When there is no Node yet, the PodScheduled condition separates the reasons:
PodScheduled | What it means | What to check next |
|---|---|---|
| The condition stays missing | No scheduler has handled the Pod | Whether a scheduler is running for .spec.schedulerName |
False, reason SchedulingGated | Scheduling gates hold the Pod before scheduling | Who is responsible for removing .spec.schedulingGates |
False, reason Unschedulable | The Scheduler tried and found no suitable Node | The condition message and FailedScheduling Events |
True | A Node has been assigned | Container states, Events, and that Node's status |
A new Pod also has no such condition for a short time, before the Scheduler first handles it. That is why the first row says "stays missing." The first two cases never reach Node selection, so they have no FailedScheduling either: no failure Event does not mean everything is fine. You can see all four cases in the lab article.
To check, use your Pod's name:
kubectl get pod <pod> -o wide
kubectl get pod <pod> \
-o jsonpath='{.status.phase}{"\n"}{.spec.nodeName}{"\n"}{range .status.conditions[?(@.type=="PodScheduled")]}{.status}{" "}{.reason}{" "}{.message}{"\n"}{end}'
kubectl describe pod <pod>
You can bypass the Scheduler by setting nodeName yourself, so a value there does not prove the Pod went through the normal checks. Setting nodeName
Also compare Events with their time. An earlier failure may be out of date. Once a Filter rejects a Node, later checks may not run for it, so the message may not list every problem that Node has. Filter
What else happens after a Node is chosen?
This article narrows the main path to filtering, scoring, and assignment. In the real flow, after a Node is chosen there may also be steps that reserve resources (Reserve), decide whether to wait or let the Pod through (Permit), and prepare before assignment (PreBind). These steps explain why "a Node has been chosen" still does not guarantee binding completes.
There is also a timing gap between choosing a Node and completing the assignment. When the Scheduler decides, it reads a copy of cluster data it keeps locally: which Nodes exist, their conditions, and which Pods are already assigned. When it chooses a Node, it first records the Pod's request against that Node in this local data. This is called assume. Binding then happens asynchronously. While it waits, the Scheduler can move on to the next Pod, and that Pod's capacity check already counts this request. When the Scheduler later sees the api-server complete the assignment, it does not count the request twice. If a later step fails, the record is undone. AssumePod and ForgetPod · ScheduleOne
This local data is updated by watching the api-server, so it can lag behind. Assume only covers decisions the Scheduler itself just made. That is why there is one more checkpoint: before starting a Pod, the kubelet checks resources again against the Node state it sees. If the Pod does not fit, the kubelet rejects it, and the Pod fails with a reason such as OutOfcpu. kubelet admission check Assume does not mean the api-server has accepted the assignment, and it certainly does not mean the containers have started.
In the v1.34.0 code, scoring is also skipped when only one suitable Node is left. To follow the full order, see Scheduling Framework and schedule_one.go.
When checking capacity, look at the settings stored in the api-server
A request missing from your YAML does not mean the Scheduler sees zero. If a container sets only a limit, and no other admission step sets a default request, Kubernetes uses the limit as the request. Admission is the step where the api-server applies defaults, validates, or changes an object as it accepts it. So when troubleshooting, look at the Pod in the api-server. Requests and limits
With init containers, Pod overhead, and similar settings, the single-container calculation here does not apply directly. When the numbers really do not add up, go back to the official resource management docs and check each item.
Storage can also affect "which Node"
A Pending PVC is not always a mount problem after assignment. Storage topology can take part in scheduling. A StorageClass with WaitForFirstConsumer delays volume binding or provisioning, so the choice can follow the Pod's scheduling conditions. Volume binding mode
The implementation details in this article were checked against Kubernetes v1.34.0. The Node numbers are a thought experiment, while the Event in section 5 is real output reproduced in a local kind cluster. The evaluation of the production question is an engineering trade-off based on the stated conditions. The official sources support the underlying mechanisms; they do not recommend one single order for handling incidents.