新副本一直 Pending,Node 的 CPU 使用率卻很低。降 request、加 toleration、放寬 affinity、刪掉重建、擴容,這幾種改法各自會發生什麼事? 與其用想的,不如在自己電腦上做一次。
這篇用 kind 在本機建立一個一次性的 Kubernetes 叢集,做出一個「新副本排不上」的狀態,然後逐一套用五種常見的處理方式,看 Pod 最後落到哪台 Node。
情境來自〈Pod 建立之後,要跑在哪裡?〉的第 5 節,那篇負責解釋 Scheduler 怎麼做決定,以及在 production 裡該怎麼取捨。這篇只做實驗,讀過那篇會更容易理解,但沒讀也能照著做。
情境:四台 Node,一個放不下的副本
叢集裡有四台 worker,各自的條件如下。新副本要求 1 CPU,而且必須跑在 workload=web 的節點上,沒有任何 toleration:
| Node | Label | Taint | 剩餘 CPU 額度 | 新副本能不能放? |
|---|---|---|---|---|
| node-a | workload=web | 無 | 0.5 | 不行,額度不足 |
| node-b | workload=web | dedicated=batch:NoSchedule | 充足 | 不行,沒有相符 toleration |
| node-c | workload=web | 無 | 0.5 | 不行,額度不足 |
| node-d | workload=batch | 無 | 充足 | 不行,不符合必要 node affinity |
「剩餘 CPU 額度」是依 requests 計算的:Node 可供 Pod 使用的 CPU(Allocatable),扣掉已經分配出去的 requests。Scheduler 判斷放不放得下時看的是這筆帳,不是監控上的使用率。這個實驗會刻意讓 node-a、node-c 的 requests 幾乎計滿,但實際使用量接近零。
接著要試的五種處理:
| 選項 | 改了什麼 |
|---|---|
| A | 把新副本的 CPU request 從 1 降到 0.5 |
| B | 增加符合條件的 web 節點容量 |
| C | 加上 node-b 對應的 toleration |
| D | 把 required node affinity 改為 preferred |
| E | 刪掉 Pending 的 Pod,重新建立 |
準備
- 需要 Docker 與 kind。本文使用 kind v0.30.0,它預設的 Kubernetes 版本是 v1.34.0;kubectl 使用 v1.34.0。
- 所有 Pod 都只跑
pause容器,幾乎不消耗資源,不需要高規格的電腦。 - kind 建立叢集時,預設會把
~/.kube/config的 current-context 切到新叢集。本文改用獨立的 kubeconfig 檔,並在每一步前確認 context,避免指令誤用到公司或其他叢集。 - 所有指令都在同一個空目錄執行,結束時整個叢集和產生的檔案一起刪除。
1. 建立叢集
一個 control-plane,加上四個 worker。worker 的 label 直接寫在 kind 設定裡:
cat > kind-config.yaml <<'EOF'
kind: Cluster
apiVersion: kind.x-k8s.io/v1alpha4
nodes:
- role: control-plane
- role: worker # node-a
labels:
workload: web
- role: worker # node-b
labels:
workload: web
- role: worker # node-c
labels:
workload: web
- role: worker # node-d
labels:
workload: batch
EOF
kind create cluster --name scheduling-lab --config kind-config.yaml --kubeconfig ./scheduling-lab.kubeconfig
export KUBECONFIG="$PWD/scheduling-lab.kubeconfig"
kubectl config current-context # 應該顯示 kind-scheduling-lab
kubectl wait --for=condition=Ready node --all --timeout=180s
kubectl create namespace scheduling-lab
kubectl config set-context --current --namespace=scheduling-lab
kubectl get nodes -L workload
kind 的 worker 名稱是 scheduling-lab-worker 到 scheduling-lab-worker4,依序對應 node-a 到 node-d。kubectl config set-context --current --namespace=scheduling-lab 只修改剛產生的獨立 kubeconfig,之後的指令不必每次加 -n。
一定要等所有節點都 Ready。還沒 Ready 的節點帶有 node.kubernetes.io/not-ready taint,會混進後面的排程訊息,讓結果不好判讀。
2. 布置狀態:讓兩台 web 節點剛好剩 0.5 CPU
node-b 加上 taint。node-a 與 node-c 各放一個 filler Pod,代表「其他工作負載已經計入的 requests」。
kubectl taint node scheduling-lab-worker2 dedicated=batch:NoSchedule
milli() { case "$1" in *m) echo "${1%m}" ;; *) echo $(( $1 * 1000 )) ;; esac; }
for node in scheduling-lab-worker scheduling-lab-worker3; do
alloc=$(milli "$(kubectl get node "$node" -o jsonpath='{.status.allocatable.cpu}')")
used=0
for r in $(kubectl get pods -A --field-selector spec.nodeName="$node" \
-o jsonpath='{range .items[*].spec.containers[*]}{.resources.requests.cpu}{" "}{end}'); do
used=$(( used + $(milli "$r") ))
done
request="$(( alloc - used - 500 ))m"
echo "$node allocatable=${alloc}m requested=${used}m filler=$request"
kubectl apply -f - <<EOF
apiVersion: v1
kind: Pod
metadata:
name: filler-${node##*-}
spec:
nodeSelector:
kubernetes.io/hostname: $node
containers:
- name: app
image: registry.k8s.io/pause:3.10
resources:
requests:
cpu: "$request"
EOF
done
kubectl wait --for=condition=Ready pod --all --timeout=120s
kubectl describe node scheduling-lab-worker3 | grep -A5 'Allocated resources'
filler 的 request 要算準,理由有兩個:
- kind 的每個節點都回報整台電腦的 CPU 數作為 Allocatable,所以數字因機器而異,必須現場計算。
- 節點上本來就有系統 Pod 的 requests,例如 kindnet 的 100m。只扣 0.5 的話,實際剩下的會比 0.5 少;後面選項 A 要測的正是「0.5 的 request 放進剩 0.5 的節點」,所以要把既有的 requests 一起扣掉。
本文實驗的電腦有 20 個 CPU,輸出如下:
scheduling-lab-worker allocatable=20000m requested=100m filler=19400m
scheduling-lab-worker3 allocatable=20000m requested=100m filler=19400m
...
Allocated resources:
(Total limits may be over 100 percent, i.e., overcommitted.)
Resource Requests Limits
-------- -------- ------
cpu 19500m (97%) 100m (0%)
memory 50Mi (0%) 50Mi (0%)
19500m (97%) 表示這台節點的 requests 已經計到 19.5 CPU,只剩 0.5。但 filler 只跑 pause,實際 CPU 使用量接近零。如果這時打開監控,會看到一台「很閒」的節點;Scheduler 看到的,卻是一台幾乎放滿的節點。
3. 新副本排不上:讀懂 FailedScheduling
建立新副本。manifest 存成檔案,後面的選項都以它為基礎修改:
cat > hello-web-new.yaml <<'EOF'
apiVersion: v1
kind: Pod
metadata:
name: hello-web-new
spec:
affinity:
nodeAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: workload
operator: In
values: ["web"]
containers:
- name: app
image: registry.k8s.io/pause:3.10
resources:
requests:
cpu: "1"
EOF
kubectl apply -f hello-web-new.yaml
kubectl wait --for=condition=PodScheduled=false pod/hello-web-new --timeout=60s
kubectl get pod hello-web-new -o wide
kubectl get pod hello-web-new \
-o jsonpath='{.metadata.uid}{"\n"}{.status.phase}{"\n"}{.spec.nodeName}{"\n"}{range .status.conditions[?(@.type=="PodScheduled")]}{.status}{" "}{.reason}{"\n"}{end}'
kubectl describe pod hello-web-new | sed -n '/^Events:/,$p'
輸出:
$ kubectl get pod hello-web-new -o wide
NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES
hello-web-new 0/1 Pending 0 1s <none> <none> <none> <none>
$ kubectl describe pod hello-web-new
...
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Warning FailedScheduling 1s default-scheduler 0/5 nodes are available:
1 node(s) didn't match Pod's node affinity/selector,
1 node(s) had untolerated taint {dedicated: batch},
1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: },
2 Insufficient cpu.
no new claims to deallocate,
preemption: 0/5 nodes are available:
2 No preemption victims found for incoming pod,
3 Preemption is not helpful for scheduling.
Message 原本是一整行,這裡依逗號與句點換行方便閱讀,文字沒有改動。這一行可以拆成三段來讀。
第一段:這次篩選的結果。 0/5 nodes are available 表示五台 Node 一台都不可行,後面依原因分組計數:
| 訊息 | 對應的 Node | 回頭比對什麼 |
|---|---|---|
1 node(s) didn't match Pod's node affinity/selector | node-d | Pod 的必要條件與 Node 的實際 labels |
1 node(s) had untolerated taint {dedicated: batch} | node-b | Node 的 taint 與 Pod 的 tolerations |
1 node(s) had untolerated taint {node-role.kubernetes.io/control-plane: } | control-plane | kind 的控制平面節點,一般工作負載本來就不會排上去 |
2 Insufficient cpu | node-a、node-c | Pod 的 requests、Node 的 Allocatable 與已分配的 requests |
每台 Node 在這裡只列出一個原因。Scheduler 的 Filter 遇到第一個不符合的條件,就可能不再檢查同一台 Node 的其他條件,所以實際叢集裡的訊息不保證列出每台 Node 的所有問題。這個實驗刻意讓每台 Node 只有一個限制,訊息才會剛好對應上表。Filter 階段
第二段:no new claims to deallocate。 篩選失敗後,Dynamic Resource Allocation(DRA)的 plugin 會檢查能不能釋放已分配的 ResourceClaim 來騰出位置。這個 Pod 沒有使用 ResourceClaim,所以沒有可做的事。DRA PostFilter 實作
第三段:preemption:。 這是搶占的試算:能不能驅逐優先權較低的 Pod,替新副本騰出位置?
- taint 與 affinity 不符屬於「驅逐誰都解決不了」的失敗,這三台(node-b、node-d、control-plane)直接標為
Preemption is not helpful for scheduling。 - node-a、node-c 的 CPU 不足,理論上可以靠驅逐解決,所以會進入試算。但上面的 Pod(filler 與系統 Pod)優先權都不比新副本低,找不到可以驅逐的對象,於是是
No preemption victims found for incoming pod。
搶占的候選節點 · 受害者篩選 · Pod Priority and Preemption
4. 逐一試五個選項
實際環境裡,這些改動應該作用在 Deployment 的 Pod template 上。實驗為了隔離變因,直接建立獨立的 Pod:每個選項都用 kubectl patch --local 從 hello-web-new.yaml 產生一個只改一處的副本,看完就刪掉,避免它佔用的容量影響下一個選項。--local 只在本機修改 manifest,不會碰到叢集裡的物件。
A:把 request 降到 0.5
# A:request 從 1 降到 0.5
kubectl patch --local -f hello-web-new.yaml --type=merge -o yaml \
-p '{"metadata":{"name":"option-a"},"spec":{"containers":[{"name":"app","image":"registry.k8s.io/pause:3.10","resources":{"requests":{"cpu":"500m"}}}]}}' \
| kubectl apply -f -
kubectl wait --for=condition=PodScheduled pod/option-a --timeout=60s
kubectl get pod option-a -o wide
kubectl delete pod option-a --wait
$ kubectl get pod option-a -o wide
NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES
option-a 0/1 ContainerCreating 0 0s <none> scheduling-lab-worker3 <none> <none>
Pod 排上了 node-c。0.5 CPU 剛好等於剩餘額度,容量檢查比較的是「是否超過」,相等仍然放得下。CPU 容量檢查實作 node-a 同樣剩 0.5,這次落在 node-c 是評分的結果,換一次也可能落在 node-a。
這個結果說明了:排程只看宣告的 request,request 降下來就放得進去。它沒有說明:這個 Pod 在尖峰時只拿 0.5 CPU 夠不夠用。request 也影響執行時爭用 CPU 的待遇,這個實驗沒有任何負載,看不到那一面。
C:加上 node-b 的 toleration
# C:加上 node-b 對應的 toleration
kubectl patch --local -f hello-web-new.yaml --type=merge -o yaml \
-p '{"metadata":{"name":"option-c"},"spec":{"tolerations":[{"key":"dedicated","operator":"Equal","value":"batch","effect":"NoSchedule"}]}}' \
| kubectl apply -f -
kubectl wait --for=condition=PodScheduled pod/option-c --timeout=60s
kubectl get pod option-c -o wide
kubectl delete pod option-c --wait
$ kubectl get pod option-c -o wide
NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES
option-c 0/1 ContainerCreating 0 0s <none> scheduling-lab-worker2 <none> <none>
Pod 排上了 node-b。node-a、node-c 仍然額度不足,node-d 仍然不符 affinity,node-b 成為唯一可行的節點。
這個結果說明了:toleration 只是解除 taint 的阻擋,讓 node-b 進入候選。它沒有說明:node-b 上的 batch 工作是否允許共用、會不會被干擾。taint 通常代表節點有專門用途,這是實驗看不到的組織決策。
D:把 required 改成 preferred
# D:required 改成 preferred
kubectl patch --local -f hello-web-new.yaml --type=merge -o yaml \
-p '{"metadata":{"name":"option-d"},"spec":{"affinity":{"nodeAffinity":{"requiredDuringSchedulingIgnoredDuringExecution":null,"preferredDuringSchedulingIgnoredDuringExecution":[{"weight":100,"preference":{"matchExpressions":[{"key":"workload","operator":"In","values":["web"]}]}}]}}}}' \
| kubectl apply -f -
kubectl wait --for=condition=PodScheduled pod/option-d --timeout=60s
kubectl get pod option-d -o wide
kubectl delete pod option-d --wait
$ kubectl get pod option-d -o wide
NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES
option-d 0/1 ContainerCreating 0 0s <none> scheduling-lab-worker4 <none> <none>
Pod 跑到了 node-d,也就是 label 為 workload=batch 的節點。改成 preferred 之後,workload=web 只是加分條件,不再是門檻;web 節點都放不下時,Scheduler 就選了唯一放得下的 node-d。
這個結果說明了:required 決定「能不能進候選」,preferred 只影響「候選之間怎麼選」。它沒有說明:這個服務能不能跑在 batch 節點上。原本寫 required 可能有理由,例如硬體、網路位置或隔離需求。
E:刪掉重建
kubectl get pod hello-web-new -o jsonpath='{.metadata.uid}{"\n"}'
kubectl delete pod hello-web-new --wait
kubectl apply -f hello-web-new.yaml
kubectl wait --for=condition=PodScheduled=false pod/hello-web-new --timeout=60s
kubectl get pod hello-web-new -o jsonpath='{.metadata.uid}{"\n"}{range .status.conditions[?(@.type=="PodScheduled")]}{.status}{" "}{.reason}{"\n"}{end}'
354cd96f-8db4-4142-be94-f1963812a205
pod "hello-web-new" deleted from scheduling-lab namespace
pod/hello-web-new created
pod/hello-web-new condition met
007312d9-1e0c-49ca-a679-221a3543395b
False Unschedulable
UID 換了,狀態卻和原本一樣:Unschedulable。限制完全沒有改變,重建只是換掉 Pod 的身分,還丟掉了原本 Pod 上的 Events 紀錄。
B:補上符合條件的容量
kind 不方便即時加 Node,這裡用刪除 node-c 上的 filler 模擬「多出一台相符的 web 容量」。注意這一步完全不動 Pending 的 Pod:
kubectl get pod hello-web-new -o jsonpath='{.metadata.uid}{"\n"}'
kubectl delete pod filler-worker3 --wait
kubectl wait --for=condition=PodScheduled pod/hello-web-new --timeout=120s
kubectl get pod hello-web-new -o jsonpath='{.metadata.uid}{"\n"}{.spec.nodeName}{"\n"}'
kubectl describe pod hello-web-new | sed -n '/^Events:/,$p'
007312d9-1e0c-49ca-a679-221a3543395b
pod "filler-worker3" deleted from scheduling-lab namespace
pod/hello-web-new condition met
007312d9-1e0c-49ca-a679-221a3543395b
scheduling-lab-worker3
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Warning FailedScheduling 1s default-scheduler 0/5 nodes are available: ...
Normal Scheduled 1s default-scheduler Successfully assigned scheduling-lab/hello-web-new to scheduling-lab-worker3
Normal Pulled 0s kubelet Container image "registry.k8s.io/pause:3.10" already present on machine
Normal Created 0s kubelet Created container: app
Normal Started 0s kubelet Started container app
同一個 UID 自己排上了 node-c,Events 依序是先前的 FailedScheduling,接著 Scheduled 與容器啟動。容量出現後,Scheduler 讓原本排不上的 Pod 重新嘗試,不需要任何人重建它。重新入列與 QueueingHint
這個結果說明了:補上符合條件的容量,是唯一沒有改變工作負載需求、Pod 也落在預期位置的做法。它沒有說明:真實環境的擴容要多久。這裡刪一個 Pod 只要一秒,雲端節點池加 Node 通常要幾分鐘,這正是 production 決策裡要和時間賽跑的部分。
五個選項的結果
| 選項 | Pod 落在 | 排程成功了嗎? | 落點是否符合原本的部署要求? |
|---|---|---|---|
| A | node-c(剩 0.5) | 是 | 位置符合,但 request 被調低 |
| B | node-c(補上容量後) | 是 | 符合 |
| C | node-b(batch 專用) | 是 | 不符合,進了專用節點 |
| D | node-d(batch label) | 是 | 不符合,離開 web 節點 |
| E | 無 | 否 | 狀態不變 |
A、C、D 都讓 Pod 取得了 Node,看起來都「修好了」。nodeName 有值只說明通過了排程,不說明這個決定是對的。 哪個改法合理,要看工作負載的需求、節點的用途和時間壓力,這部分的討論在〈Pod 建立之後,要跑在哪裡?〉第 5 節。
5. 另外三種 Pending
不是每個 Pending 的 Pod 都會有 FailedScheduling。同一個叢集裡,再建立三個 Pod:一個帶 scheduling gate、一個指定不存在的 Scheduler、一個使用不存在的 image:
kubectl apply -f - <<'EOF'
apiVersion: v1
kind: Pod
metadata:
name: gated
spec:
schedulingGates:
- name: example.com/wait-for-approval
containers:
- name: app
image: registry.k8s.io/pause:3.10
---
apiVersion: v1
kind: Pod
metadata:
name: no-scheduler
spec:
schedulerName: my-scheduler # 叢集裡沒有這個 Scheduler
containers:
- name: app
image: registry.k8s.io/pause:3.10
EOF
kubectl run bad-image --image=registry.k8s.io/pause:does-not-exist
kubectl wait --for=condition=PodScheduled=false pod/gated --timeout=60s
kubectl wait --for=jsonpath='{.status.containerStatuses[0].state.waiting.reason}'=ImagePullBackOff pod/bad-image --timeout=120s
kubectl get pod gated no-scheduler bad-image -o wide
kubectl get pod gated -o jsonpath='{.status.conditions}{"\n"}'
kubectl get pod no-scheduler -o jsonpath='conditions={.status.conditions}{"\n"}'
kubectl get events --field-selector involvedObject.name=no-scheduler
kubectl get pod bad-image \
-o jsonpath='{.status.phase}{"\n"}{.spec.nodeName}{"\n"}{range .status.conditions[?(@.type=="PodScheduled")]}{.status}{"\n"}{end}'
kubectl describe pod bad-image | sed -n '/^Events:/,$p'
$ kubectl get pod gated no-scheduler bad-image -o wide
NAME READY STATUS RESTARTS AGE IP NODE NOMINATED NODE READINESS GATES
gated 0/1 SchedulingGated 0 17s <none> <none> <none> <none>
no-scheduler 0/1 Pending 0 17s <none> <none> <none> <none>
bad-image 0/1 ImagePullBackOff 0 17s 10.244.2.3 scheduling-lab-worker4 <none> <none>
$ kubectl get pod gated -o jsonpath='{.status.conditions}'
[{"lastProbeTime":null,"lastTransitionTime":"2026-09-23T18:57:45Z","message":"Scheduling is blocked due to non-empty scheduling gates","reason":"SchedulingGated","status":"False","type":"PodScheduled"}]
$ kubectl get pod no-scheduler -o jsonpath='conditions={.status.conditions}'
conditions=
$ kubectl get events --field-selector involvedObject.name=no-scheduler
No resources found in scheduling-lab namespace.
$ kubectl get pod bad-image -o jsonpath='...'
Pending
scheduling-lab-worker4
True
$ kubectl describe pod bad-image
...
Events:
Type Reason Age From Message
---- ------ ---- ---- -------
Normal Scheduled 17s default-scheduler Successfully assigned scheduling-lab/bad-image to scheduling-lab-worker4
Normal BackOff 15s kubelet Back-off pulling image "registry.k8s.io/pause:does-not-exist"
Warning Failed 15s kubelet Error: ImagePullBackOff
Normal Pulling 1s (x2 over 16s) kubelet Pulling image "registry.k8s.io/pause:does-not-exist"
Warning Failed 0s (x2 over 16s) kubelet Failed to pull image "registry.k8s.io/pause:does-not-exist": rpc error: code = NotFound desc = failed to pull and unpack image "registry.k8s.io/pause:does-not-exist": failed to resolve reference "registry.k8s.io/pause:does-not-exist": registry.k8s.io/pause:does-not-exist: not found
Warning Failed 0s (x2 over 16s) kubelet Error: ErrImagePull
三個 Pod 的 STATUS 看起來都不像「正常」,但原因完全不同:
gated:STATUS 直接顯示SchedulingGated,PodScheduled的 reason 也是SchedulingGated。gate 被移除之前,Scheduler 不會嘗試替它選節點,所以沒有FailedScheduling。Pod Scheduling Readinessno-scheduler:schedulerName指向不存在的 Scheduler,沒有任何元件負責它。它既沒有PodScheduledcondition,也沒有任何 Events,STATUS 卻和一般排不上的 Pod 一樣是 Pending。只看 STATUS,很容易把它和第 3 節的新副本混為一談。bad-image:phase 仍是 Pending,但nodeName已經有值,PodScheduled是True。排程早就完成,卡住的是 kubelet 拉 image,和 Scheduler 已經沒有關係。
移除 gate 之後,gated 就會進入排程。Pod 建立後只允許移除 scheduling gates,不能新增:
kubectl patch pod gated --type=json -p='[{"op":"remove","path":"/spec/schedulingGates"}]'
kubectl wait --for=condition=PodScheduled pod/gated --timeout=60s
kubectl get pod gated -o wide
所以遇到 Pending,第一步先看 .spec.nodeName 和 PodScheduled:有沒有 Node、condition 存不存在、reason 是什麼,就能分出這是還沒排、被擋住、排不上,還是早已排上卻卡在啟動。
這個實驗的邊界
實驗能清楚看到的,是 Scheduler 的判斷:它依 requests 而不是使用量計算容量、每個改法如何改變候選節點、補上容量後會自動重試。
它看不到的,是 production 決策真正困難的部分:
- 執行時的影響:沒有負載,就看不到 request 調低後的 CPU 爭用,也看不到 batch 與 web 共用節點時的互相干擾。
- 組織與架構的限制:taint 和 required affinity 背後的理由,不在叢集物件裡。
- 時間:真實擴容的等待時間,是選 B 時最大的風險。
- kind 的特性:所有節點共用主機 CPU,Allocatable 等於主機 CPU 數;control-plane 也出現在訊息裡。選項 A 在 node-a 與 node-c 之間的落點取決於評分,不是固定結果。「一秒內重新排程」也只是這次的觀察,不是 Scheduler 保證的時間。
清理
刪除整個 kind 叢集與這次產生的檔案:
kind delete cluster --name scheduling-lab --kubeconfig ./scheduling-lab.kubeconfig
rm -f ./scheduling-lab.kubeconfig ./kind-config.yaml ./hello-web-new.yaml
unset KUBECONFIG