A new replica stays Pending, but the Nodes' CPU usage is low. Lower the request, add a toleration, relax the affinity, delete and recreate, or scale out: what actually happens with each of these fixes? Instead of guessing, try it on your own machine.
This article uses kind to create a throwaway Kubernetes cluster on your machine, builds a state where a new replica cannot be scheduled, then applies five common fixes one by one to see which Node the Pod ends up on.
The scenario comes from section 5 of Where does a new Pod run?. That article explains how the Scheduler decides, and how to weigh the options in production. This one is only the lab. It is easier to follow if you have read that article, but you can follow the steps without it.
The scenario: four Nodes and a replica that does not fit
The cluster has four workers with the conditions below. The new replica requests 1 CPU, must run on a Node labeled workload=web, and has no tolerations:
| Node | Label | Taint | CPU left | Can the new replica go here? |
|---|---|---|---|---|
| node-a | workload=web | None | 0.5 | No, not enough room |
| node-b | workload=web | dedicated=batch:NoSchedule | Plenty | No, no matching toleration |
| node-c | workload=web | None | 0.5 | No, not enough room |
| node-d | workload=batch | None | Plenty | No, does not match the required node affinity |
"CPU left" is based on requests: the CPU a Node can give to Pods (Allocatable) minus the requests already assigned. When the Scheduler decides whether a Pod fits, it uses this number, not the usage in your monitoring. This lab deliberately fills up the requests on node-a and node-c while keeping their actual usage near zero.
The five fixes we will try:
| Option | What changes |
|---|---|
| A | Lower the new replica's CPU request from 1 to 0.5 |
| B | Add capacity with matching web Nodes |
| C | Add a toleration for node-b's taint |
| D | Change the required node affinity to preferred |
| E | Delete the Pending Pod and create it again |
Before you start
- You need Docker and kind. This article uses kind v0.30.0, whose default Kubernetes version is v1.34.0, and kubectl v1.34.0.
- All Pods only run
pausecontainers and use almost no resources, so you do not need a powerful machine. - By default, kind switches the current context in
~/.kube/configto the new cluster. This article uses a separate kubeconfig file instead, and checks the context before each step, so commands cannot accidentally reach a work cluster or any other cluster. - Run every command in the same empty directory. At the end, the whole cluster and the generated files are deleted together.
1. Create the cluster
One control-plane Node plus four workers. The worker labels are set directly in the kind config:
cat > kind-config.yaml <<'EOF'
kind: Cluster
apiVersion: kind.x-k8s.io/v1alpha4
nodes:
- role: control-plane
- role: worker # node-a
labels:
workload: web
- role: worker # node-b
labels:
workload: web
- role: worker # node-c
labels:
workload: web
- role: worker # node-d
labels:
workload: batch
EOF
kind create cluster --name scheduling-lab --config kind-config.yaml --kubeconfig ./scheduling-lab.kubeconfig
export KUBECONFIG="$PWD/scheduling-lab.kubeconfig"
kubectl config current-context # should print kind-scheduling-lab
kubectl wait --for=condition=Ready node --all --timeout=180s
kubectl create namespace scheduling-lab
kubectl config set-context --current --namespace=scheduling-lab
kubectl get nodes -L workload
kind names the workers scheduling-lab-worker through scheduling-lab-worker4, which map to node-a through node-d. kubectl config set-context --current --namespace=scheduling-lab only changes the separate kubeconfig you just created, so later commands do not need -n every time.
Wait until every Node is Ready. A Node that is not Ready has the node.kubernetes.io/not-ready taint, which would show up in the scheduling messages later and make the results hard to read.
2. Set up the state: leave exactly 0.5 CPU on two web Nodes
Add the taint to node-b. On node-a and node-c, create a filler Pod that stands for "requests already counted from other workloads."
kubectl taint node scheduling-lab-worker2 dedicated=batch:NoSchedule
milli() { case "$1" in *m) echo "${1%m}" ;; *) echo $(( $1 * 1000 )) ;; esac; }
for node in scheduling-lab-worker scheduling-lab-worker3; do
alloc=$(milli "$(kubectl get node "$node" -o jsonpath='{.status.allocatable.cpu}')")
used=0
for r in $(kubectl get pods -A --field-selector spec.nodeName="$node" \
-o jsonpath='{range .items[*].spec.containers[*]}{.resources.requests.cpu}{" "}{end}'); do
used=$(( used + $(milli "$r") ))
done
request="$(( alloc - used - 500 ))m"
echo "$node allocatable=${alloc}m requested=${used}m filler=$request"
kubectl apply -f - <<EOF
apiVersion: v1
kind: Pod
metadata:
name: filler-${node##*-}
spec:
nodeSelector:
kubernetes.io/hostname: $node
containers:
- name: app
image: registry.k8s.io/pause:3.10
resources:
requests:
cpu: "$request"
EOF
done
kubectl wait --for=condition=Ready pod --all --timeout=120s
kubectl describe node scheduling-lab-worker3 | grep -A5 'Allocated resources'
The filler's request has to be exact, for two reasons:
- Every kind Node reports the CPU count of the whole machine as Allocatable, so the numbers depend on your machine and must be calculated on the spot.
- Nodes already carry requests from system Pods, such as kindnet's 100m. If you only subtracted 0.5, less than 0.5 would actually be left. Option A tests "a 0.5 request on a Node with 0.5 left," so the existing requests must be subtracted too.
The machine used for this article has 20 CPUs. The output looks like this:
scheduling-lab-worker allocatable=20000m requested=100m filler=19400m
scheduling-lab-worker3 allocatable=20000m requested=100m filler=19400m
...
Allocated resources:
(Total limits may be over 100 percent, i.e., overcommitted.)
Resource Requests Limits
-------- -------- ------
cpu 19500m (97%) 100m (0%)
memory 50Mi (0%) 50Mi (0%)
19500m (97%) means this Node's requests already add up to 19.5 CPU, leaving only 0.5. But the filler only runs pause, so its actual CPU usage is close to zero. If you opened your monitoring now, you would see a Node that looks idle. The Scheduler sees a Node that is almost full.
3. The new replica does not fit: reading FailedScheduling
Create the new replica. The manifest is saved to a file, because the options below are based on it:
cat > hello-web-new.yaml <<'EOF'
apiVersion: v1
kind: Pod
metadata:
name: hello-web-new
spec:
affinity:
nodeAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: workload
operator: In
values: ["web"]
containers:
- name: app
image: registry.k8s.io/pause:3.10
resources:
requests:
cpu: "1"
EOF
kubectl apply -f hello-web-new.yaml
kubectl wait --for=condition=PodScheduled=false pod/hello-web-new --timeout=60s
kubectl get pod hello-web-new -o wide
kubectl get pod hello-web-new \
-o jsonpath='{.metadata.uid}{"\n"}{.status.phase}{"\n"}{.spec.nodeName}{"\n"}{range .status.conditions[?(@.type=="PodScheduled")]}{.status}{" "}{.reason}{"\n"}{end}'
kubectl describe pod hello-web-new | sed -n '/^Events:/,$p'
Output:
$ kubectl get pod hello-web-new -o wide
NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES
hello-web-new 0/1 Pending 0 1s <none> <none> <none> <none>
$ kubectl describe pod hello-web-new
...
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Warning FailedScheduling 1s default-scheduler 0/5 nodes are available:
1 node(s) didn't match Pod's node affinity/selector,
1 node(s) had untolerated taint {dedicated: batch},
1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: },
2 Insufficient cpu.
no new claims to deallocate,
preemption: 0/5 nodes are available:
2 No preemption victims found for incoming pod,
3 Preemption is not helpful for scheduling.
The message is one long line. It is split at commas and periods here to make it easier to read; the text is unchanged. It can be read in three parts.
Part one: the result of this filtering pass. 0/5 nodes are available means none of the five Nodes can be used. The reasons follow, grouped and counted:
| Message | Node | What to compare |
|---|---|---|
1 node(s) didn't match Pod's node affinity/selector | node-d | The Pod's required conditions and the Node's actual labels |
1 node(s) had untolerated taint {dedicated: batch} | node-b | The Node's taint and the Pod's tolerations |
1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: } | control-plane | kind's control-plane Node; normal workloads never run there |
2 Insufficient cpu | node-a, node-c | The Pod's requests, the Node's Allocatable, and the allocated requests |
Each Node shows only one reason here. Once a Filter finds the first condition a Node fails, the Scheduler may stop checking that Node's other conditions. So in a real cluster, the message may not list every problem each Node has. This lab gives each Node only one limit on purpose, so the message lines up with the table above. Filter
Part two: no new claims to deallocate. After filtering fails, the Dynamic Resource Allocation (DRA) plugin checks whether freeing an allocated ResourceClaim would make room. This Pod uses no ResourceClaim, so there is nothing to do. DRA PostFilter
Part three: preemption:. This is a preemption dry run: could evicting lower-priority Pods make room for the new replica?
- A taint or affinity mismatch is a failure that no eviction can fix. These three Nodes (node-b, node-d, and control-plane) are marked
Preemption is not helpful for schedulingright away. - The CPU shortage on node-a and node-c could in theory be fixed by eviction, so they go into the dry run. But none of the Pods on them (the filler and system Pods) has a lower priority than the new replica, so there is nothing to evict:
No preemption victims found for incoming pod.
Preemption candidates · Victim selection · Pod Priority and Preemption
4. Try the five options one by one
In a real environment, these changes should be made to the Deployment's Pod template. To isolate each change, the lab creates standalone Pods instead. Each option uses kubectl patch --local to make a copy of hello-web-new.yaml with one change. We delete each Pod after looking at it, so the capacity it uses does not affect the next option. --local only changes the manifest on your machine; it does not touch any object in the cluster.
A: lower the request to 0.5
# A: lower the request from 1 to 0.5
kubectl patch --local -f hello-web-new.yaml --type=merge -o yaml \
-p '{"metadata":{"name":"option-a"},"spec":{"containers":[{"name":"app","image":"registry.k8s.io/pause:3.10","resources":{"requests":{"cpu":"500m"}}}]}}' \
| kubectl apply -f -
kubectl wait --for=condition=PodScheduled pod/option-a --timeout=60s
kubectl get pod option-a -o wide
kubectl delete pod option-a --wait
$ kubectl get pod option-a -o wide
NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES
option-a 0/1 ContainerCreating 0 0s <none> scheduling-lab-worker3 <none> <none>
The Pod landed on node-c. 0.5 CPU is exactly what is left, and the capacity check asks whether the request is larger than what is left, so an equal amount still fits. CPU capacity check node-a also has 0.5 left. Landing on node-c this time is the result of scoring; another run might pick node-a.
What this shows: scheduling only looks at the declared request, so a lower request fits. What it does not show: whether 0.5 CPU is enough for this Pod at peak. The request also affects how the container competes for CPU at run time, and this lab has no load, so it cannot show that side.
C: add node-b's toleration
# C: add a toleration for node-b's taint
kubectl patch --local -f hello-web-new.yaml --type=merge -o yaml \
-p '{"metadata":{"name":"option-c"},"spec":{"tolerations":[{"key":"dedicated","operator":"Equal","value":"batch","effect":"NoSchedule"}]}}' \
| kubectl apply -f -
kubectl wait --for=condition=PodScheduled pod/option-c --timeout=60s
kubectl get pod option-c -o wide
kubectl delete pod option-c --wait
$ kubectl get pod option-c -o wide
NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES
option-c 0/1 ContainerCreating 0 0s <none> scheduling-lab-worker2 <none> <none>
The Pod landed on node-b. node-a and node-c still do not have enough room, and node-d still does not match the affinity, so node-b is the only suitable Node.
What this shows: a toleration only removes the taint's block and lets node-b become a candidate. What it does not show: whether the batch work on node-b allows sharing, or whether it will be disturbed. A taint usually means the Node has a dedicated purpose, and that is an organizational decision the lab cannot see.
D: change required to preferred
# D: change required to preferred
kubectl patch --local -f hello-web-new.yaml --type=merge -o yaml \
-p '{"metadata":{"name":"option-d"},"spec":{"affinity":{"nodeAffinity":{"requiredDuringSchedulingIgnoredDuringExecution":null,"preferredDuringSchedulingIgnoredDuringExecution":[{"weight":100,"preference":{"matchExpressions":[{"key":"workload","operator":"In","values":["web"]}]}}]}}}}' \
| kubectl apply -f -
kubectl wait --for=condition=PodScheduled pod/option-d --timeout=60s
kubectl get pod option-d -o wide
kubectl delete pod option-d --wait
$ kubectl get pod option-d -o wide
NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES
option-d 0/1 ContainerCreating 0 0s <none> scheduling-lab-worker4 <none> <none>
The Pod ran on node-d, the Node labeled workload=batch. After the change to preferred, workload=web only adds points; it is no longer a requirement. With no room on the web Nodes, the Scheduler chose node-d, the only Node with room.
What this shows: required decides "can this Node be a candidate?" and preferred only affects "which candidate wins?" What it does not show: whether this service can run on batch Nodes. The original required rule may have had a reason, such as hardware, network location, or isolation.
E: delete and recreate
kubectl get pod hello-web-new -o jsonpath='{.metadata.uid}{"\n"}'
kubectl delete pod hello-web-new --wait
kubectl apply -f hello-web-new.yaml
kubectl wait --for=condition=PodScheduled=false pod/hello-web-new --timeout=60s
kubectl get pod hello-web-new -o jsonpath='{.metadata.uid}{"\n"}{range .status.conditions[?(@.type=="PodScheduled")]}{.status}{" "}{.reason}{"\n"}{end}'
354cd96f-8db4-4142-be94-f1963812a205
pod "hello-web-new" deleted from scheduling-lab namespace
pod/hello-web-new created
pod/hello-web-new condition met
007312d9-1e0c-49ca-a679-221a3543395b
False Unschedulable
The UID changed, but the state did not: Unschedulable. None of the limits changed. Recreating only replaced the Pod's identity, and lost the Events on the original Pod along the way.
B: add matching capacity
kind cannot easily add a Node on the fly, so we simulate "one more matching web Node's worth of room" by deleting the filler on node-c. Note that this step does not touch the Pending Pod at all:
kubectl get pod hello-web-new -o jsonpath='{.metadata.uid}{"\n"}'
kubectl delete pod filler-worker3 --wait
kubectl wait --for=condition=PodScheduled pod/hello-web-new --timeout=120s
kubectl get pod hello-web-new -o jsonpath='{.metadata.uid}{"\n"}{.spec.nodeName}{"\n"}'
kubectl describe pod hello-web-new | sed -n '/^Events:/,$p'
007312d9-1e0c-49ca-a679-221a3543395b
pod "filler-worker3" deleted from scheduling-lab namespace
pod/hello-web-new condition met
007312d9-1e0c-49ca-a679-221a3543395b
scheduling-lab-worker3
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Warning FailedScheduling 1s default-scheduler 0/5 nodes are available: ...
Normal Scheduled 1s default-scheduler Successfully assigned scheduling-lab/hello-web-new to scheduling-lab-worker3
Normal Pulled 0s kubelet Container image "registry.k8s.io/pause:3.10" already present on machine
Normal Created 0s kubelet Created container: app
Normal Started 0s kubelet Started container app
The same UID was placed on node-c on its own. The Events show the earlier FailedScheduling, then Scheduled and the container starting. Once capacity appeared, the Scheduler let the Pod that did not fit try again, and no one had to recreate it. Requeueing and QueueingHint
What this shows: adding matching capacity is the only fix that does not change the workload's requirements and puts the Pod where it was meant to go. What it does not show: how long scaling takes in a real environment. Deleting a Pod here takes one second. Adding a Node to a cloud node pool usually takes minutes, and that is the race against time in the production decision.
Results of the five options
| Option | Where the Pod landed | Scheduled? | Does the placement meet the original requirements? |
|---|---|---|---|
| A | node-c (0.5 left) | Yes | The location does, but the request was lowered |
| B | node-c (after capacity was added) | Yes | Yes |
| C | node-b (dedicated to batch) | Yes | No, it entered a dedicated Node |
| D | node-d (batch label) | Yes | No, it left the web Nodes |
| E | None | No | The state did not change |
A, C, and D all got the Pod a Node, and they all look "fixed." A value in nodeName only shows that scheduling passed. It does not show the decision was right. Which fix makes sense depends on the workload's needs, the Nodes' purpose, and the time pressure. That discussion is in section 5 of Where does a new Pod run?
5. Three other kinds of Pending
Not every Pending Pod gets a FailedScheduling. In the same cluster, create three more Pods: one with a scheduling gate, one that names a scheduler that does not exist, and one that uses an image that does not exist:
kubectl apply -f - <<'EOF'
apiVersion: v1
kind: Pod
metadata:
name: gated
spec:
schedulingGates:
- name: example.com/wait-for-approval
containers:
- name: app
image: registry.k8s.io/pause:3.10
---
apiVersion: v1
kind: Pod
metadata:
name: no-scheduler
spec:
schedulerName: my-scheduler # No such scheduler in this cluster
containers:
- name: app
image: registry.k8s.io/pause:3.10
EOF
kubectl run bad-image --image=registry.k8s.io/pause:does-not-exist
kubectl wait --for=condition=PodScheduled=false pod/gated --timeout=60s
kubectl wait --for=jsonpath='{.status.containerStatuses[0].state.waiting.reason}'=ImagePullBackOff pod/bad-image --timeout=120s
kubectl get pod gated no-scheduler bad-image -o wide
kubectl get pod gated -o jsonpath='{.status.conditions}{"\n"}'
kubectl get pod no-scheduler -o jsonpath='conditions={.status.conditions}{"\n"}'
kubectl get events --field-selector involvedObject.name=no-scheduler
kubectl get pod bad-image \
-o jsonpath='{.status.phase}{"\n"}{.spec.nodeName}{"\n"}{range .status.conditions[?(@.type=="PodScheduled")]}{.status}{"\n"}{end}'
kubectl describe pod bad-image | sed -n '/^Events:/,$p'
$ kubectl get pod gated no-scheduler bad-image -o wide
NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES
gated 0/1 SchedulingGated 0 17s <none> <none> <none> <none>
no-scheduler 0/1 Pending 0 17s <none> <none> <none> <none>
bad-image 0/1 ImagePullBackOff 0 17s 10.244.2.3 scheduling-lab-worker4 <none> <none>
$ kubectl get pod gated -o jsonpath='{.status.conditions}'
[{"lastProbeTime":null,"lastTransitionTime":"2026-09-23T18:57:45Z","message":"Scheduling is blocked due to non-empty scheduling gates","reason":"SchedulingGated","status":"False","type":"PodScheduled"}]
$ kubectl get pod no-scheduler -o jsonpath='conditions={.status.conditions}'
conditions=
$ kubectl get events --field-selector involvedObject.name=no-scheduler
No resources found in scheduling-lab namespace.
$ kubectl get pod bad-image -o jsonpath='...'
Pending
scheduling-lab-worker4
True
$ kubectl describe pod bad-image
...
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Normal Scheduled 17s default-scheduler Successfully assigned scheduling-lab/bad-image to scheduling-lab-worker4
Normal BackOff 15s kubelet Back-off pulling image "registry.k8s.io/pause:does-not-exist"
Warning Failed 15s kubelet Error: ImagePullBackOff
Normal Pulling 1s (x2 over 16s) kubelet Pulling image "registry.k8s.io/pause:does-not-exist"
Warning Failed 0s (x2 over 16s) kubelet Failed to pull image "registry.k8s.io/pause:does-not-exist": rpc error: code = NotFound desc = failed to pull and unpack image "registry.k8s.io/pause:does-not-exist": failed to resolve reference "registry.k8s.io/pause:does-not-exist": registry.k8s.io/pause:does-not-exist: not found
Warning Failed 0s (x2 over 16s) kubelet Error: ErrImagePull
None of the three STATUS values look "normal," but the reasons are completely different:
gated: STATUS showsSchedulingGateddirectly, and thePodScheduledreason is alsoSchedulingGated. Until the gate is removed, the Scheduler does not try to choose a Node, so there is noFailedScheduling. Pod Scheduling Readinessno-scheduler:schedulerNamepoints to a scheduler that does not exist, so no component is responsible for it. It has noPodScheduledcondition and no Events at all, yet its STATUS shows Pending just like a Pod that cannot be scheduled. From STATUS alone, it is easy to confuse it with the new replica in section 3.bad-image: the phase is still Pending, butnodeNamehas a value andPodScheduledisTrue. Scheduling finished long ago. What is stuck is the kubelet pulling the image, and the Scheduler is no longer involved.
After you remove the gate, gated enters scheduling. Once a Pod is created, scheduling gates can only be removed, not added:
kubectl patch pod gated --type=json -p='[{"op":"remove","path":"/spec/schedulingGates"}]'
kubectl wait --for=condition=PodScheduled pod/gated --timeout=60s
kubectl get pod gated -o wide
So when you see Pending, first check .spec.nodeName and PodScheduled. Whether there is a Node, whether the condition exists, and what the reason is will tell you whether the Pod has not been scheduled yet, is being held back, cannot be scheduled, or was scheduled long ago but is stuck starting.
What this lab can and cannot show
The lab clearly shows the Scheduler's decisions: it counts capacity by requests, not usage; each fix changes the candidate Nodes in a specific way; and once capacity appears, it retries on its own.
It cannot show the hard parts of the production decision:
- Run-time effects: with no load, you cannot see the CPU contention after lowering a request, or how batch and web work disturb each other on a shared Node.
- Organizational and architectural limits: the reasons behind a taint or a required affinity are not stored in cluster objects.
- Time: the real wait for scaling is the biggest risk of choosing B.
- kind specifics: all Nodes share the host's CPUs, so Allocatable equals the host CPU count, and the control-plane Node shows up in the messages. Whether option A lands on node-a or node-c depends on scoring and is not fixed. "Rescheduled within one second" is only what happened this time, not a timing the Scheduler guarantees.
Clean up
Delete the whole kind cluster and the files created here:
kind delete cluster --name scheduling-lab --kubeconfig ./scheduling-lab.kubeconfig
rm -f ./scheduling-lab.kubeconfig ./kind-config.yaml ./hello-web-new.yaml
unset KUBECONFIG