平均只用了三成 CPU,卻有三成的 period 被 throttle,p99 延遲衝到幾百毫秒。這種狀態要怎麼做出來?把 thread 調少,真的會比較好嗎? 與其用想的,不如在自己電腦上做一次。
這篇用 kind 在本機建立一次性的 Kubernetes 叢集,搭配一支自製的小程式,重現「平均使用率很低,卻被 throttle」的狀態。過程中會直接讀 container 的 cgroup 檔案和監控指標,確認 CPU limit 與 request 在 Node 上變成了什麼。
原理與判讀寫在〈CPU 才用三成,為什麼還被 throttle?〉,那篇只引用結果。這篇放完整的步驟與輸出,並說明每個實驗能證明什麼、不能證明什麼。讀過那篇會更容易理解,但沒讀也能照著做。
準備
- 需要 Docker(cgroup v2)、kind、kubectl,以及 Go 1.25 以上,用來編譯實驗程式。
- 本文的環境:Apple Silicon 的 Mac,Docker 由 OrbStack 提供(Docker Engine 29.4.0、cgroup v2、10 CPU、VM kernel 7.0);kind v0.33.0;Go 1.26.0。
- kind 的 Node 是一個 container,看得到整台電腦的 CPU。下面輸出裡的
NumCPU=10、GOMAXPROCS=10,在你的電腦上會是自己的 CPU 數。延遲也會受電腦上其他程式影響,只適合比較數量級與相對差異。 - kind 建立叢集時,預設會把
~/.kube/config的 current-context 切到新叢集。本文把 kubeconfig 存在工作目錄的tmp/,所有 kubectl 指令都經過一個小包裝,明確帶上 kubeconfig 與 context,不會誤用到公司或其他叢集。 - 第 7 節會用
docker update暫時把 worker Node 限制在一顆 CPU(第 5 號)上,結束時要還原;那一節會說明這次踩到的坑。電腦的 CPU 少於 6 顆時,請把腳本裡的5改成其他編號。 - 第 8 節同樣用
docker update,把 worker 與 control-plane 分到 CPU 0-7 與 8-9,是依 10 顆 CPU 寫的。CPU 數不同時,請一起調整腳本裡的兩個範圍與鄰居的 thread 數。 - 所有指令都在同一個空目錄執行。結束時刪除叢集、image 與這個目錄即可。
1. 實驗程式與叢集
為什麼要自己寫程式
throttle 的效果是「同樣的工作花更久」。如果用經過的時間來計量工作,例如「迴圈跑 20ms」,被暫停的時間也會被算成工作量,就看不出差別了。所以實驗程式 cpulab 先把 goroutine 鎖在一個 OS thread 上,再讀這個 thread 實際用掉的 CPU 時間(RUSAGE_THREAD),用滿指定的 CPU 時間才算做完。被 throttle 停住的時間不會算進去,只會讓完成時間變長。
cpulab 有五個子指令,完整原始碼在文末附錄:
| 子指令 | 用在 | 做什麼 |
|---|---|---|
info | 第 2、4 節 | 印出 Go 版本、GOMAXPROCS,以及 container 自己的 cpu.max 與 cpu.weight |
burst | 第 3、9 節 | 每隔一段時間讓 N 個 thread 同時做固定的 CPU 工作,印出每輪的完成時間與 cpu.stat 的變化 |
serve | 第 5、6、8 節 | HTTP 服務:/work 做固定的 CPU 工作並配置記憶體,/ping 什麼都不做,/stats 回報 cpu.stat |
load | 第 5、6、8 節 | 依 Poisson 到達送出請求的壓測程式,印出延遲分布與服務端的 throttle 統計 |
spin | 第 7、8 節 | 持續佔滿 thread,每 10 秒印一次實際用了多少 CPU |
建立兩個 image
先把附錄的程式存成 cpulab/main.go,再建立 Dockerfile 與建置腳本:
mkdir -p cpulab scripts tmp
# 把文末附錄的程式存成 cpulab/main.go
cat > cpulab/Dockerfile <<'EOF'
FROM busybox:1.37
COPY cpulab /cpulab
ENTRYPOINT ["/cpulab"]
EOF
cat > build.sh <<'EOF'
#!/usr/bin/env bash
# 建兩個映像:程式碼相同,只差 go.mod 的 go 版本行。
# cpulab:go124mod go.mod 寫 go 1.24(GOMAXPROCS 沿用舊預設)
# cpulab:go125mod go.mod 寫 go 1.25(GOMAXPROCS 參考 cgroup CPU limit)
set -euo pipefail
here=$(cd "$(dirname "$0")" && pwd)
arch=$(docker info --format '{{.Architecture}}' | sed 's/aarch64/arm64/;s/x86_64/amd64/')
work=$(mktemp -d)
trap 'rm -rf "$work"' EXIT
for v in 1.24 1.25; do
tag=go${v/./}mod
mkdir -p "$work/$tag"
cp "$here/cpulab/main.go" "$here/cpulab/Dockerfile" "$work/$tag/"
printf 'module cpulab\n\ngo %s\n' "$v" > "$work/$tag/go.mod"
(cd "$work/$tag" && CGO_ENABLED=0 GOOS=linux GOARCH=$arch GOTOOLCHAIN=local go build -o cpulab .)
docker build -q -t "cpulab:$tag" "$work/$tag"
done
go version
EOF
chmod +x build.sh
./build.sh
build.sh 會建立兩個 image,程式碼完全相同,只差 go.mod 裡的 go 版本行。第 4 節會看到,這一行決定了 Go 程式在 container 裡預設用幾個 thread 同時執行。兩個 image 都用本機的 Go toolchain 編譯(GOTOOLCHAIN=local),本文用的是 Go 1.26.0。
建立叢集
kubectl 的包裝 scripts/kx:第一個參數選叢集,其餘原樣交給 kubectl。
cat > scripts/kx <<'EOF'
#!/usr/bin/env bash
# kx <old|new> <kubectl args...>:明確指定 kubeconfig 與 context,不碰 ~/.kube/config
here=$(cd "$(dirname "$0")/.." && pwd); c=$1; shift
exec kubectl --kubeconfig "$here/tmp/kubeconfig-$c" --context "kind-cpu-lab-$c" "$@"
EOF
chmod +x scripts/kx
cat > scripts/kind-2node.yaml <<'EOF'
kind: Cluster
apiVersion: kind.x-k8s.io/v1alpha4
nodes:
- role: control-plane
- role: worker
EOF
kind create cluster --name cpu-lab-new \
--image kindest/node:v1.34.11@sha256:44e222ee2132dab25ff87301682f89eb82c7880ea3a1bf543bfe9708fd08d67d \
--config scripts/kind-2node.yaml --kubeconfig tmp/kubeconfig-new
kind load docker-image cpulab:go124mod cpulab:go125mod --name cpu-lab-new
scripts/kx new get nodes
docker exec cpu-lab-new-worker sh -c 'runc --version | head -1; stat -fc %T /sys/fs/cgroup'
一個 control-plane、一個 worker。所有實驗 Pod 都用 nodeSelector 放在 worker 上;第 5、6、8 節的壓測程式放在 control-plane,避免和服務搶 CPU。
node image 用 digest 固定,因為裡面的 runc 版本會影響結果:第 7 節要比較的 cpu.weight 公式,就是由 container runtime 決定的。最後一行應該看到 runc 1.4.3 與 cgroup2fs(cgroup v2)。
之後每一節都會先用 cat > ... <<'EOF' 存下腳本再執行。腳本開頭註解裡的 E1、E2a 等是實驗紀錄的編號,和本文的節次不同。
2. limit 被寫到哪裡?
兩個 Pod,各有一個 app(limit 2)和一個 sidecar。e1-all-limits 的 sidecar 有 limit,e1-partial 的沒有。這份 manifest 也會建立後面都會用到的 namespace cpulab:
cat > scripts/e1-limits.yaml <<'EOF'
# E1:limit 被寫到哪裡。兩個 Pod 都有 app(limit 2)與 sidecar;
# all-limits 的 sidecar 有 limit,partial 的 sidecar 沒有。
apiVersion: v1
kind: Namespace
metadata:
name: cpulab
---
apiVersion: v1
kind: Pod
metadata:
name: e1-all-limits
namespace: cpulab
spec:
nodeSelector: {kubernetes.io/hostname: WORKER}
containers:
- name: app
image: cpulab:go125mod
command: ["sh", "-c", "/cpulab info; sleep 3600"]
resources: {requests: {cpu: 500m}, limits: {cpu: "2"}}
- name: sidecar
image: cpulab:go125mod
command: ["sh", "-c", "/cpulab info; sleep 3600"]
resources: {requests: {cpu: 100m}, limits: {cpu: 500m}}
---
apiVersion: v1
kind: Pod
metadata:
name: e1-partial
namespace: cpulab
spec:
nodeSelector: {kubernetes.io/hostname: WORKER}
containers:
- name: app
image: cpulab:go125mod
command: ["sh", "-c", "/cpulab info; sleep 3600"]
resources: {requests: {cpu: 500m}, limits: {cpu: "2"}}
- name: sidecar
image: cpulab:go125mod
command: ["sh", "-c", "/cpulab info; sleep 3600"]
resources: {requests: {cpu: 100m}}
EOF
sed s/WORKER/cpu-lab-new-worker/ scripts/e1-limits.yaml | scripts/kx new apply -f -
scripts/kx new -n cpulab wait --for=condition=Ready pod --all --timeout=120s
for p in e1-all-limits e1-partial; do for c in app sidecar; do
echo "$p/$c: $(scripts/kx new -n cpulab logs $p -c $c)"
done; done
每個 container 啟動時執行 cpulab info,印出自己看到的設定:
e1-all-limits/app: go=go1.26.0 GOMAXPROCS=2 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env=""
cpu.max="200000 100000" cpu.weight="59"
e1-all-limits/sidecar: go=go1.26.0 GOMAXPROCS=2 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env=""
cpu.max="50000 100000" cpu.weight="17"
e1-partial/app: go=go1.26.0 GOMAXPROCS=2 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env=""
cpu.max="200000 100000" cpu.weight="59"
e1-partial/sidecar: go=go1.26.0 GOMAXPROCS=10 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env=""
cpu.max="max 100000" cpu.weight="17"
container 的 cpu.max 就是 limit 換算的結果:前面是 quota,後面是 period,單位是微秒。limit 2 是每 100ms 最多 200ms,500m 是 50ms;沒有 limit 的 sidecar 是 max,不限制。
同一行的 GOMAXPROCS 也值得注意:有 limit 的 container 是 2,沒有 limit 的是 10。這是 go.mod 寫 go 1.25 的 image,第 4 節再細看。cpu.weight 則由 request 換算而來,第 7 節會用到。
接著到 worker Node 裡,看 Pod 這一層與它上面的幾層。cgtree.sh 從 Pod 的 cgroup 往上走到根,印出每一層的 cpu.weight 與 cpu.max,再列出 Pod 底下的每個 container:
cat > scripts/cgtree.sh <<'EOF'
#!/bin/sh
# 在 kind node 內,印出某個 Pod UID 從 cgroup root 到各 container 的 cpu.weight/cpu.max
uid=$(echo "$1" | tr - _)
pod=$(find /sys/fs/cgroup -type d -name "*pod${uid}.slice" | head -1)
[ -n "$pod" ] || { echo "pod cgroup not found"; exit 1; }
d=$pod; chain=""
while [ "$d" != "/sys/fs/cgroup" ]; do chain="$d $chain"; d=$(dirname "$d"); done
for d in $chain; do printf '%-70s weight=%-5s max=%s\n' "${d#/sys/fs/cgroup}" "$(cat $d/cpu.weight)" "$(cat $d/cpu.max)"; done
for c in "$pod"/*/; do [ -f "$c/cpu.weight" ] && printf ' %-68s weight=%-5s max=%s\n' "$(basename $c)" "$(cat $c/cpu.weight)" "$(cat $c/cpu.max)"; done
EOF
docker cp scripts/cgtree.sh cpu-lab-new-worker:/cgtree.sh
for p in e1-all-limits e1-partial; do
echo "== $p"
docker exec cpu-lab-new-worker sh /cgtree.sh "$(scripts/kx new -n cpulab get pod $p -o jsonpath='{.metadata.uid}')"
done
== e1-all-limits
/kubelet.slice weight=100 max=max 100000
/kubelet.slice/kubelet-kubepods.slice weight=391 max=max 100000
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-burstable.slice weight=51 max=max 100000
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-burstable.slice/kubelet-kubepods-burstable-poda562a08c_3829_44ee_9bd1_c48d68973a26.slice weight=24 max=250000 100000
cri-containerd-4de8b08221dfa8a22321d110a6dec2c51f5a851ed29bdb31e39dd6941577c54a.scope weight=17 max=50000 100000
cri-containerd-507b6210e898adff9a780ab969559d4843a410cf8cc3f712859b6fa0c3f04579.scope weight=1 max=max 100000
cri-containerd-906c59210735dabc6bbd403e64c1b67feb9e8916cb64bad37ae999b397744a94.scope weight=59 max=200000 100000
== e1-partial
/kubelet.slice weight=100 max=max 100000
/kubelet.slice/kubelet-kubepods.slice weight=391 max=max 100000
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-burstable.slice weight=51 max=max 100000
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-burstable.slice/kubelet-kubepods-burstable-podee5545ac_13e0_401f_b07b_370dee16585d.slice weight=24 max=max 100000
cri-containerd-02508278f52c40262b86427ba22081210783af2fac243c4280e41d2995e4e430.scope weight=59 max=200000 100000
cri-containerd-a648dabc12e1b9f6db1451ecf4f5d42c2cc616d0ec991d5bbb1521c7229dac0e.scope weight=1 max=max 100000
cri-containerd-e4fa2ea818a16f840234c79b818d51fa79b2d6d93d1eb88691b104a64fbd6763.scope weight=17 max=max 100000
Pod 這一層(...pod<uid>.slice)的 max,在 e1-all-limits 是 250000 100000,也就是兩個 container 的 limit 相加(2 + 0.5);e1-partial 有一個 container 沒有 limit,Pod 層就是 max。weight=1 的那個 scope 是 Pod 的 pause container。
路徑最上面多了一層 /kubelet.slice,這是 kind 的設定(kubelet 的 cgroupRoot 是 /kubelet)。一般用 systemd cgroup driver 的 Node 上,Pod 會在 /kubepods.slice 底下。
看完就刪掉這兩個 Pod,讓後面的實驗在安靜的 Node 上進行:
scripts/kx new -n cpulab delete pod e1-all-limits e1-partial --wait=true
這個結果說明了:limit 最後就是 container 與 Pod 兩層 cgroup 的 cpu.max,Pod 層只在每個 container 都有 limit 時才設定。它沒有說明:kernel 怎麼執行這個上限。這是下一節的事。
3. 同一份工作,四種 limit
讓 cpulab burst 每秒做一次同樣的工作:8 個 thread 同時各算 20ms,總共 160ms 的 CPU,其餘時間閒置,平均使用率約 0.16 顆 CPU。依序在不設 limit、limit 2、1、500m 下各跑 20 輪;最後一組同樣是 limit 1,但改成 2 個 thread 各算 80ms,總工作量不變:
cat > scripts/e2a-burst.sh <<'EOF'
#!/usr/bin/env bash
# E2a:每秒一次 burst,同樣的工作量(總共 160ms CPU),在不同 limit 與 thread 數下的完成時間與 throttle。
# 用法:e2a-burst.sh <kx 指令> <worker node>;依序執行,避免互相干擾。
set -euo pipefail
KX=$1; NODE=$2
run() { # name limit threads work gomaxprocs
local name=$1 limit=$2 threads=$3 work=$4 procs=$5 res='{"requests":{"cpu":"100m"}}'
[ "$limit" != none ] && res="{\"requests\":{\"cpu\":\"100m\"},\"limits\":{\"cpu\":\"$limit\"}}"
$KX -n cpulab run "$name" --image=cpulab:go125mod --restart=Never \
--overrides="{\"spec\":{\"nodeSelector\":{\"kubernetes.io/hostname\":\"$NODE\"},\"containers\":[{\"name\":\"c\",\"image\":\"cpulab:go125mod\",\"args\":[\"burst\",\"-threads=$threads\",\"-work=$work\",\"-interval=1s\",\"-rounds=20\"],\"resources\":$res,\"env\":[{\"name\":\"GOMAXPROCS\",\"value\":\"$procs\"}]}]}}" >/dev/null
$KX -n cpulab wait --for=jsonpath='{.status.phase}'=Succeeded pod/$name --timeout=120s >/dev/null
echo "===== $name limit=$limit threads=$threads work=$work GOMAXPROCS=$procs"
$KX -n cpulab logs $name
$KX -n cpulab delete pod $name --wait=false >/dev/null
}
run e2a-nolimit none 8 20ms 8
run e2a-limit2 2 8 20ms 8
run e2a-limit1 1 8 20ms 8
run e2a-limit500m 500m 8 20ms 8
run e2a-limit1-2thr 1 2 80ms 2
EOF
chmod +x scripts/e2a-burst.sh
scripts/e2a-burst.sh "scripts/kx new" cpu-lab-new-worker
腳本用環境變數 GOMAXPROCS 把 Go 的 thread 數固定成和工作 thread 一樣多。這個 image 的 go.mod 寫 go 1.25,不固定的話,GOMAXPROCS 會跟著 limit 變(第 4 節),各組比較的就不只是 limit 了。
limit 1 的前五輪與總結:
===== e2a-limit1 limit=1 threads=8 work=20ms GOMAXPROCS=8
go=go1.26.0 GOMAXPROCS=8 NumCPU=10 GOMAXPROCS_env="8" GODEBUG_env=""
cpu.max="100000 100000" cpu.weight="17"
burst threads=8 work=20ms interval=1s rounds=20
round wall_ms d_nr_periods d_nr_throttled d_throttled_ms d_usage_ms
1 83.2 1 1 421.9 164.1
2 82.4 1 1 390.2 164.6
3 85.2 1 1 361.7 166.0
4 78.8 1 1 325.3 165.2
5 117.0 1 1 379.5 166.4
...
summary wall_ms p50=84.4 max=118.1 avg_cpu_cores=0.165 nr_periods=60 nr_throttled=20 throttled_ms=7052.5
每一輪的欄位依序是:完成時間,以及這一輪前後 cpu.stat 的差值(nr_periods、nr_throttled、throttled_usec、usage_usec,後兩者換算成毫秒)。最後一行是 20 輪的總結,avg_cpu_cores 是整段時間的平均使用率。
五組的標頭與總結:
===== e2a-nolimit limit=none threads=8 work=20ms GOMAXPROCS=8
summary wall_ms p50=34.3 max=38.4 avg_cpu_cores=0.165 nr_periods=0 nr_throttled=0 throttled_ms=0.0
===== e2a-limit2 limit=2 threads=8 work=20ms GOMAXPROCS=8
summary wall_ms p50=34.6 max=38.1 avg_cpu_cores=0.166 nr_periods=40 nr_throttled=0 throttled_ms=0.0
===== e2a-limit1 limit=1 threads=8 work=20ms GOMAXPROCS=8
summary wall_ms p50=84.4 max=118.1 avg_cpu_cores=0.165 nr_periods=60 nr_throttled=20 throttled_ms=7052.5
===== e2a-limit500m limit=500m threads=8 work=20ms GOMAXPROCS=8
summary wall_ms p50=260.4 max=263.7 avg_cpu_cores=0.166 nr_periods=100 nr_throttled=60 throttled_ms=23981.0
===== e2a-limit1-2thr limit=1 threads=2 work=80ms GOMAXPROCS=2
summary wall_ms p50=132.3 max=136.6 avg_cpu_cores=0.162 nr_periods=77 nr_throttled=17 throttled_ms=1301.6
整理成表:
| 設定 | 平均使用率 | 完成時間 p50/最長 | 20 輪累積的 nr_periods/nr_throttled |
|---|---|---|---|
| 不設 limit | 0.165 | 34.3/38.4ms | 0/0 |
| limit 2 | 0.166 | 34.6/38.1ms | 40/0 |
| limit 1 | 0.165 | 84.4/118.1ms | 60/20 |
| limit 500m | 0.166 | 260.4/263.7ms | 100/60 |
| limit 1,2 thread × 80ms | 0.162 | 132.3/136.6ms | 77/17 |
幾個值得對照的地方:
- 平均都一樣,完成時間卻差很多。 limit 2 的預算裝得下 160ms,和不設 limit 幾乎相同;limit 1 每一輪都被 throttle 一次,500m 則每一輪三次。
- limit 1 的完成時間分成兩群,約 80ms 與約 116ms。工作開始時落在 period 的哪個位置,決定了要等多久才輪到下一次發放預算。
d_throttled_ms比完成時間還長。 limit 1 每輪約 220 到 420ms,但整輪才 80 到 118ms。這是 8 個 thread 各自被停住的時間加總,不是經過的時間。nr_periods只算忙碌的 period。 limit 2 跑了 20 秒,只累積 40 個 period,平均每次 burst 算 2 個;閒置的時間不計入。每輪讀到的d_nr_periods是 0,表示這兩個 period 是在 burst 做完之後才記上去的。- thread 少,throttle 變少,但沒有比較快。 2 個 thread 各算 80ms,就算完全不受限制也要 80ms;在 limit 1 下,throttle 只有 17 次,完成時間的中位數卻是 132ms,比 8 個 thread 的 84ms 還久。第 5 節會在 HTTP 服務上看到同樣的取捨。
這個結果說明了:quota 是每個 period 內所有 thread 共用的預算,平均使用率低也可能每次都被 throttle。它沒有說明:真實服務的請求會不會這樣集中,這要看第 5 節。
4. GOMAXPROCS 看的是 go.mod
預算是所有 thread 共用的,thread 越多,預算就越早用完。Go 程式同時執行 goroutine 的 thread 數由 GOMAXPROCS 決定,過去預設等於 Node 的 CPU 數量,不管 limit 設多少。Go 1.25 改了這個預設:GOMAXPROCS 會參考 cgroup 的 CPU limit,取 limit 無條件進位、且至少為 2,並定期更新;它不看 request。Go 1.25 release notes · GOMAXPROCS 的計算
有一個容易漏掉的條件:這個行為由 go.mod 裡的 go 版本決定,不是由編譯用的 Go 版本決定;go.mod 寫的版本低於 1.25,就沿用舊行為。GODEBUG 預設值表
這一節用同一份程式碼、只差 go.mod 版本行的兩個 image,在不同 limit 下印出 GOMAXPROCS。最後兩個 Pod 用環境變數覆寫預設值:
cat > scripts/e3-gomaxprocs.sh <<'EOF'
#!/usr/bin/env bash
# E3:同一份程式碼,只差 go.mod 的 go 版本行,在不同 CPU limit 下的 GOMAXPROCS。
# 用法:e3-gomaxprocs.sh <kx 指令> <worker node>
set -euo pipefail
KX=$1; NODE=$2
run() { # name image limit env...
local name=$1 image=$2 limit=$3; shift 3
local res='{}' envs='[]'
[ "$limit" != none ] && res="{\"requests\":{\"cpu\":\"100m\"},\"limits\":{\"cpu\":\"$limit\"}}"
[ $# -gt 0 ] && envs="[$(for e in "$@"; do printf '{"name":"%s","value":"%s"},' "${e%%=*}" "${e#*=}"; done | sed 's/,$//')]"
$KX -n cpulab run "$name" --image="cpulab:$image" --restart=Never --command \
--overrides="{\"spec\":{\"nodeSelector\":{\"kubernetes.io/hostname\":\"$NODE\"},\"containers\":[{\"name\":\"c\",\"image\":\"cpulab:$image\",\"command\":[\"/cpulab\",\"info\"],\"resources\":$res,\"env\":$envs}]}}" \
-- /cpulab info >/dev/null
}
for img in go124mod go125mod; do
for lim in none 500m 1 1500m 2 2500m; do run "e3-$img-$(echo $lim | tr -d m)" $img $lim; done
done
run e3-go125mod-2-env go125mod 2 GOMAXPROCS=6
run e3-go125mod-2-godebug go125mod 2 GODEBUG=containermaxprocs=0
sleep 15
for p in $($KX -n cpulab get pods -o name | grep e3- | sort -V); do
printf '%-28s limit=%-6s %s\n' "${p#pod/}" "$($KX -n cpulab get $p -o jsonpath='{.spec.containers[0].resources.limits.cpu}')" "$($KX -n cpulab logs $p | tr '\n' ' ')"
done
EOF
chmod +x scripts/e3-gomaxprocs.sh
scripts/e3-gomaxprocs.sh "scripts/kx new" cpu-lab-new-worker
e3-go124mod-1 limit=1 go=go1.26.0 GOMAXPROCS=10 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env="" cpu.max="100000 100000" cpu.weight="17"
e3-go124mod-2 limit=2 go=go1.26.0 GOMAXPROCS=10 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env="" cpu.max="200000 100000" cpu.weight="17"
e3-go124mod-500 limit=500m go=go1.26.0 GOMAXPROCS=10 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env="" cpu.max="50000 100000" cpu.weight="17"
e3-go124mod-1500 limit=1500m go=go1.26.0 GOMAXPROCS=10 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env="" cpu.max="150000 100000" cpu.weight="17"
e3-go124mod-2500 limit=2500m go=go1.26.0 GOMAXPROCS=10 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env="" cpu.max="250000 100000" cpu.weight="17"
e3-go124mod-none limit= go=go1.26.0 GOMAXPROCS=10 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env="" cpu.max="max 100000" cpu.weight="1"
e3-go125mod-1 limit=1 go=go1.26.0 GOMAXPROCS=2 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env="" cpu.max="100000 100000" cpu.weight="17"
e3-go125mod-2 limit=2 go=go1.26.0 GOMAXPROCS=2 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env="" cpu.max="200000 100000" cpu.weight="17"
e3-go125mod-2-env limit=2 go=go1.26.0 GOMAXPROCS=6 NumCPU=10 GOMAXPROCS_env="6" GODEBUG_env="" cpu.max="200000 100000" cpu.weight="17"
e3-go125mod-2-godebug limit=2 go=go1.26.0 GOMAXPROCS=10 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env="containermaxprocs=0" cpu.max="200000 100000" cpu.weight="17"
e3-go125mod-500 limit=500m go=go1.26.0 GOMAXPROCS=2 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env="" cpu.max="50000 100000" cpu.weight="17"
e3-go125mod-1500 limit=1500m go=go1.26.0 GOMAXPROCS=2 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env="" cpu.max="150000 100000" cpu.weight="17"
e3-go125mod-2500 limit=2500m go=go1.26.0 GOMAXPROCS=3 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env="" cpu.max="250000 100000" cpu.weight="17"
e3-go125mod-none limit= go=go1.26.0 GOMAXPROCS=10 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env="" cpu.max="max 100000" cpu.weight="1"
- go.mod 寫
go 1.24的 image,不論 limit 多少,GOMAXPROCS 都是 Node 的 CPU 數 10。 - go.mod 寫
go 1.25的 image,GOMAXPROCS 是 limit 無條件進位、且至少為 2:500m、1、1.5、2 都是 2,2.5 是 3;沒有 limit 時仍是 10。 - 設定
GOMAXPROCS=6就是 6;GODEBUG=containermaxprocs=0則回到舊行為。
兩個 image 都用 Go 1.26 編譯。決定行為的是 go.mod 的版本行,不是編譯器的版本,所以只升級 toolchain 不夠,要確認 go.mod 的版本,或直接設定 GOMAXPROCS 環境變數。沒有 limit 的兩個 Pod 也沒有 request,是 BestEffort,所以 cpu.weight 是 1。這些 Pod 印完就結束,不佔 CPU,可以留到最後隨叢集一起刪除。
其他 runtime 也有類似的問題。JVM 依可用的 processor 數量決定 GC thread 與 thread pool 的預設大小,而這個數量在 container 裡是從 quota 算出來的;JDK 19 起不再參考 request 換算出的 shares。JDK-8281181 不同版本的細節不同,調整前先確認實際看到幾顆 CPU。本文沒有實測 JVM。
這個結果說明了:只升級 toolchain,GOMAXPROCS 不會跟著 limit 走。它沒有說明:GOMAXPROCS 跟著 limit 之後,延遲會不會比較好。這是下一節要測的。
5. 重現開頭的服務
實驗設計
cpulab serve 是一個小型 HTTP 服務,request 500m:
- 每個
/work請求配置 1MB 的記憶體再丟掉,並做 5ms 的 CPU 工作;程式常駐約 64MB 的存活資料,讓 GC 每次都有東西要標記。加上 HTTP 與 GC 的開銷,每個請求平均約用 6ms 的 CPU。 /ping什麼都不做,用來觀察服務忙碌時,無關的小請求會被拖累多少。
壓測程式 cpulab load 跑在 control-plane 上,以 Poisson 到達送出請求,平均每秒 100 個 /work,另外每秒 20 個 /ping。它不等回應就送下一個(open-loop),延遲從預定送出的時間開始算,服務一慢,排隊的時間就會算進延遲裡。暖機 10 秒後量測 60 秒,前後各讀一次服務的 /stats,算出這段時間的 cpu.stat 差值。
-clump 控制請求怎麼到達:1 表示一個一個來;50 表示每次 50 個一起到,平均速率不變,只是變成每秒平均兩群。每種到達方式各跑三種設定:
| 名稱 | limit | image | GOMAXPROCS |
|---|---|---|---|
nolimit-g10 | 不設 | go.mod 1.24 | 10 |
limit2-g10 | 2 | go.mod 1.24 | 10 |
limit2-g2 | 2 | go.mod 1.25 | 2 |
cat > scripts/e2b-serve.sh <<'EOF'
#!/usr/bin/env bash
# E2b:HTTP 服務(每個請求 CPU 工作 + 配置記憶體、常駐 heap 讓 GC 有工作),
# 從 control-plane 上的 load Pod 以 Poisson 到達送請求,比較 limit 與 GOMAXPROCS。
# 用法:e2b-serve.sh <kx 指令> <worker node> <cp node> <rps> <name:image:limit>...
# 中途失敗或中斷時,trap 會刪除還在跑的 Pod,避免殘留的 server 被 Service 選到、混進下一組的量測。
set -euo pipefail
KX=$1; NODE=$2; CP=$3; RPS=$4; shift 4
active=() # 目前還在跑的 Pod
cleanup() { [ ${#active[@]} -eq 0 ] || $KX -n cpulab delete pod "${active[@]}" --ignore-not-found --wait=false >/dev/null 2>&1 || true; }
trap cleanup EXIT
$KX -n cpulab get svc cpulab-server >/dev/null 2>&1 || \
$KX -n cpulab create service clusterip cpulab-server --tcp=8080:8080 >/dev/null
$KX -n cpulab patch svc cpulab-server -p '{"spec":{"selector":{"app":"cpulab-server"}}}' >/dev/null
for spec in "$@"; do
IFS=: read -r name image limit <<<"$spec"
res='{"requests":{"cpu":"500m"}}'
[ "$limit" != none ] && res="{\"requests\":{\"cpu\":\"500m\"},\"limits\":{\"cpu\":\"$limit\"}}"
LIVEMB=${LIVEMB:-64}; $KX -n cpulab run "$name" --image=cpulab:$image --restart=Never --labels=app=cpulab-server \
--overrides="{\"spec\":{\"nodeSelector\":{\"kubernetes.io/hostname\":\"$NODE\"},\"containers\":[{\"name\":\"c\",\"image\":\"cpulab:$image\",\"args\":[\"serve\",\"-work=5ms\",\"-alloc=1048576\",\"-live-mb=$LIVEMB\"],\"resources\":$res,\"ports\":[{\"containerPort\":8080}],\"readinessProbe\":{\"httpGet\":{\"path\":\"/stats\",\"port\":8080}}}]}}" >/dev/null
active=("$name")
$KX -n cpulab wait --for=condition=Ready pod/$name --timeout=120s >/dev/null
$KX -n cpulab run "load-$name" --image=cpulab:go125mod --restart=Never \
--overrides="{\"spec\":{\"nodeSelector\":{\"kubernetes.io/hostname\":\"$CP\"},\"tolerations\":[{\"operator\":\"Exists\"}],\"containers\":[{\"name\":\"c\",\"image\":\"cpulab:go125mod\",\"args\":[\"load\",\"-target=http://cpulab-server:8080\",\"-rps=$RPS\",\"-clump=${CLUMP:-1}\",\"-ping-rps=${PINGRPS:-0}\",\"-duration=60s\",\"-warmup=10s\",\"-label=$name\"]}]}}" >/dev/null
active+=("load-$name")
$KX -n cpulab wait --for=jsonpath='{.status.phase}'=Succeeded pod/load-$name --timeout=180s >/dev/null
echo "===== $name image=$image limit=$limit rps=$RPS clump=${CLUMP:-1} ping_rps=${PINGRPS:-0} live_mb=$LIVEMB"
$KX -n cpulab logs load-$name
$KX -n cpulab delete pod $name load-$name --wait=true >/dev/null
active=()
done
EOF
chmod +x scripts/e2b-serve.sh
for c in 1:steady 50:clumped; do
CLUMP=${c%%:*} PINGRPS=20 scripts/e2b-serve.sh "scripts/kx new" cpu-lab-new-worker cpu-lab-new-control-plane 100 \
e2b-${c#*:}-nolimit-g10:go124mod:none e2b-${c#*:}-limit2-g10:go124mod:2 e2b-${c#*:}-limit2-g2:go125mod:2
done
六組依序執行,每組約一分半鐘。
一個一個來:完全不被 throttle
load label=e2b-steady-nolimit-g10 gomaxprocs=10 clump=1 requests=5920 errors=0 achieved_rps=98.7
work_latency_ms p50=7.0 p90=8.9 p99=13.2 p999=18.8 max=23.3
ping_latency_ms requests=1182 p50=1.3 p90=3.1 p99=4.9 max=6.9
server avg_cpu_cores=0.600 nr_periods=0 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0 num_gc=75
load label=e2b-steady-limit2-g10 gomaxprocs=10 clump=1 requests=6149 errors=0 achieved_rps=102.5
work_latency_ms p50=7.0 p90=9.0 p99=13.4 p999=18.9 max=21.9
ping_latency_ms requests=1168 p50=1.3 p90=3.0 p99=4.6 max=12.2
server avg_cpu_cores=0.624 nr_periods=600 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0 num_gc=78
load label=e2b-steady-limit2-g2 gomaxprocs=2 clump=1 requests=6026 errors=0 achieved_rps=100.4
work_latency_ms p50=7.2 p90=9.7 p99=15.7 p999=21.9 max=26.8
ping_latency_ms requests=1236 p50=1.6 p90=3.8 p99=8.7 max=17.8
server avg_cpu_cores=0.604 nr_periods=600 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0 num_gc=78
三種設定的平均都約 0.6 顆 CPU,nr_throttled 全是 0,p99 在 13 到 16ms 之間。limit 2 的兩組各累積 600 個 period,也就是 60 秒裡每個 period 都有在跑,但每個 period 需要的 CPU 遠低於 200ms,從來沒有用完預算。不設 limit 的那組沒有 quota,nr_periods 也就不會累積。
每次 50 個一起來
load label=e2b-clumped-nolimit-g10 gomaxprocs=10 clump=50 requests=5800 errors=0 achieved_rps=96.6
work_latency_ms p50=33.1 p90=52.6 p99=74.4 p999=90.0 max=94.6
ping_latency_ms requests=1196 p50=2.2 p90=4.6 p99=13.0 max=43.0
server avg_cpu_cores=0.607 nr_periods=0 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0 num_gc=58
load label=e2b-clumped-limit2-g10 gomaxprocs=10 clump=50 requests=6100 errors=0 achieved_rps=101.7
work_latency_ms p50=70.8 p90=187.5 p99=370.1 p999=658.1 max=665.0
ping_latency_ms requests=1200 p50=2.5 p90=23.9 p99=99.4 max=289.1
server avg_cpu_cores=0.639 nr_periods=490 nr_throttled=147 throttled_ratio=0.300 throttled_ms=53589.5 num_gc=60
load label=e2b-clumped-limit2-g2 gomaxprocs=2 clump=50 requests=6250 errors=0 achieved_rps=104.2
work_latency_ms p50=110.7 p90=239.0 p99=447.1 p999=560.0 max=589.1
ping_latency_ms requests=1251 p50=3.0 p90=113.8 p99=335.8 max=434.7
server avg_cpu_cores=0.643 nr_periods=487 nr_throttled=22 throttled_ratio=0.045 throttled_ms=36.2 num_gc=69
平均使用率幾乎沒變,約 0.61 到 0.64 顆 CPU,也就是 limit 2 的三成。可是 limit 2、GOMAXPROCS 10 的那組,有 30% 的 period 被 throttle,p99 從不設 limit 時的 74ms 變成 370ms,/ping 的 p99 也從 13ms 變成 99ms。
一群 50 個請求,大約需要 300ms 的 CPU,一個 period 的 200ms 預算裝不下。10 個 thread 同時跑,約 20ms 就把預算用完,整個服務暫停到 period 結束;這段時間到達的 /ping 也只能等。
把 GOMAXPROCS 降到 2,throttle 的比例從 30% 降到 4.5%,throttled_ms 從 53 秒降到 36 毫秒,指標看起來好了很多。但 p99 變成 447ms,/ping 的 p99 變成 336ms,兩者都比 GOMAXPROCS 10 更差。2 個 thread 一個 period 內最多只用得掉 200ms,本來就很難超過預算;同樣 300ms 的工作只能兩個兩個慢慢做,/ping 排在後面等得更久。每 100ms 能做的事沒有變多,只是把「暫停」換成了「排隊」。
throttled_ms 的 53 秒是 60 秒量測裡各 CPU 被停住的時間加總,不是服務停了 53 秒。
再跑一次
延遲容易受電腦上其他程式影響,所以成群到達的三組又跑了一次:
CLUMP=50 PINGRPS=20 scripts/e2b-serve.sh "scripts/kx new" cpu-lab-new-worker cpu-lab-new-control-plane 100 \
e2b-clumped-nolimit-g10:go124mod:none e2b-clumped-limit2-g10:go124mod:2 e2b-clumped-limit2-g2:go125mod:2
load label=e2b-clumped-nolimit-g10 gomaxprocs=10 clump=50 requests=6750 errors=0 achieved_rps=112.5
work_latency_ms p50=34.0 p90=56.3 p99=86.3 p999=109.5 max=115.5
ping_latency_ms requests=1228 p50=2.1 p90=4.7 p99=18.7 max=56.1
server avg_cpu_cores=0.702 nr_periods=0 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0 num_gc=67
load label=e2b-clumped-limit2-g10 gomaxprocs=10 clump=50 requests=7000 errors=0 achieved_rps=116.8
work_latency_ms p50=63.0 p90=199.1 p99=359.9 p999=608.6 max=622.2
ping_latency_ms requests=1237 p50=2.6 p90=32.4 p99=110.7 max=326.7
server avg_cpu_cores=0.730 nr_periods=513 nr_throttled=160 throttled_ratio=0.312 throttled_ms=58780.4 num_gc=67
load label=e2b-clumped-limit2-g2 gomaxprocs=2 clump=50 requests=7050 errors=0 achieved_rps=117.2
work_latency_ms p50=114.8 p90=262.0 p99=511.2 p999=627.3 max=645.0
ping_latency_ms requests=1223 p50=3.0 p90=131.7 p99=341.2 max=542.6
server avg_cpu_cores=0.720 nr_periods=499 nr_throttled=22 throttled_ratio=0.044 throttled_ms=40.4 num_gc=78
| 設定 | 平均 CPU | throttle 的 period | /work p50/p99 | /ping p99 |
|---|---|---|---|---|
| 不設 limit,GOMAXPROCS 10 | 0.61/0.70 | — | 33/74、34/86ms | 13/19ms |
| limit 2,GOMAXPROCS 10 | 0.64/0.73 | 30.0%/31.2% | 71/370、63/360ms | 99/111ms |
| limit 2,GOMAXPROCS 2 | 0.64/0.72 | 4.5%/4.4% | 111/447、115/511ms | 336/341ms |
每格依序是第一次與第二次。第二次的實際速率稍高,約每秒 112 到 117 個請求,但三組之間的相對關係相同。
這個結果說明了:同樣的平均使用率,請求成群到達就會被 throttle;調低 GOMAXPROCS 會讓 throttle 指標下降,延遲卻變差。它沒有說明:其他形狀的負載會怎樣。例如長時間接近 limit,或 GC 比重更高的程式,thread 數的影響可能不同,本文沒有測。
6. 增加副本,還是提高 limit?
正文討論解法時,比較了兩種常見的做法:增加副本,以及 request 不變、只提高 limit。這一節用第 5 節同樣的成群流量(每秒平均 100 個 /work、每次 50 個一起到,另有每秒 20 個 /ping)比較五種設定。server 都用 go.mod 1.24 的 image(GOMAXPROCS 10),每個 Pod 的 request 都是 500m:
| 名稱 | 副本數 | 每個 Pod 的 limit | limit 總和 |
|---|---|---|---|
e7-1x-nolimit | 1 | 不設 | — |
e7-1x-limit2 | 1 | 2 | 2 |
e7-2x-limit1 | 2 | 1 | 2 |
e7-2x-limit2 | 2 | 2 | 4 |
e7-1x-limit4 | 1 | 4 | 4 |
壓測程式的 /stats 經過 Service,有兩個副本時只會打到其中一個 Pod,所以這個腳本改用 kubectl exec 在壓測前後讀取每個 server Pod 的 cpu.stat。讀取的區間包含 10 秒暖機,平均用量以整段經過的秒數計算,所以和第 5 節的 server 行不完全可比。
cat > scripts/e7-scale.sh <<'EOF'
#!/usr/bin/env bash
# E7:同樣的成群流量,比較「增加副本」與「提高 limit」。server 都用 go124mod(GOMAXPROCS 10)、request 500m。
# 用法:e7-scale.sh <kx 指令> <worker node> <cp node> <rps> <name:replicas:limit>...
# 壓測程式的 /stats 經過 Service,多副本時只會打到其中一個 Pod,所以改用 kubectl exec 在壓測前後
# 讀每個 server Pod 的 cpu.stat。區間包含 10 秒暖機,平均用量以整段經過的秒數計算。
# 中途失敗或中斷時,trap 會刪除還在跑的 Pod,避免殘留的 server 被 Service 選到、混進下一組的量測。
set -euo pipefail
KX=$1; NODE=$2; CP=$3; RPS=$4; shift 4
active=() # 目前還在跑的 Pod
cleanup() { [ ${#active[@]} -eq 0 ] || $KX -n cpulab delete pod "${active[@]}" --ignore-not-found --wait=false >/dev/null 2>&1 || true; }
trap cleanup EXIT
$KX get ns cpulab >/dev/null 2>&1 || $KX create ns cpulab >/dev/null
$KX -n cpulab get svc cpulab-server >/dev/null 2>&1 || \
$KX -n cpulab create service clusterip cpulab-server --tcp=8080:8080 >/dev/null
$KX -n cpulab patch svc cpulab-server -p '{"spec":{"selector":{"app":"cpulab-server"}}}' >/dev/null
cpustat() { $KX -n cpulab exec "$1" -- cat /sys/fs/cgroup/cpu.stat | awk '{printf "%s=%s ", $1, $2}'; }
delta() { # before after elapsed_s -> 這段時間的用量與 throttle
printf '%s\n%s\n' "$1" "$2" | awk -v t="$3" '
{ for (i = 1; i <= NF; i++) { split($i, kv, "="); v[NR, kv[1]] = kv[2] } }
END {
u = v[2, "usage_usec"] - v[1, "usage_usec"]; p = v[2, "nr_periods"] - v[1, "nr_periods"]
n = v[2, "nr_throttled"] - v[1, "nr_throttled"]; s = v[2, "throttled_usec"] - v[1, "throttled_usec"]
printf "avg_cpu_cores=%.3f nr_periods=%d nr_throttled=%d throttled_ratio=%.3f throttled_ms=%.1f\n", u / t / 1e6, p, n, (p ? n / p : 0), s / 1000
}'
}
for spec in "$@"; do
IFS=: read -r name replicas limit <<<"$spec"
res='{"requests":{"cpu":"500m"}}'
[ "$limit" != none ] && res="{\"requests\":{\"cpu\":\"500m\"},\"limits\":{\"cpu\":\"$limit\"}}"
pods=()
for i in $(seq 1 "$replicas"); do
$KX -n cpulab run "$name-$i" --image=cpulab:go124mod --restart=Never --labels=app=cpulab-server \
--overrides="{\"spec\":{\"nodeSelector\":{\"kubernetes.io/hostname\":\"$NODE\"},\"containers\":[{\"name\":\"c\",\"image\":\"cpulab:go124mod\",\"args\":[\"serve\",\"-work=5ms\",\"-alloc=1048576\",\"-live-mb=64\"],\"resources\":$res,\"ports\":[{\"containerPort\":8080}],\"readinessProbe\":{\"httpGet\":{\"path\":\"/stats\",\"port\":8080}}}]}}" >/dev/null
pods+=("$name-$i"); active=("${pods[@]}")
done
$KX -n cpulab wait --for=condition=Ready pod "${pods[@]}" --timeout=120s >/dev/null
before=()
for p in "${pods[@]}"; do before+=("$(cpustat "$p")"); done
t0=$(date +%s)
$KX -n cpulab run "load-$name" --image=cpulab:go125mod --restart=Never \
--overrides="{\"spec\":{\"nodeSelector\":{\"kubernetes.io/hostname\":\"$CP\"},\"tolerations\":[{\"operator\":\"Exists\"}],\"containers\":[{\"name\":\"c\",\"image\":\"cpulab:go125mod\",\"args\":[\"load\",\"-target=http://cpulab-server:8080\",\"-rps=$RPS\",\"-clump=${CLUMP:-1}\",\"-ping-rps=${PINGRPS:-0}\",\"-duration=60s\",\"-warmup=10s\",\"-label=$name\"]}]}}" >/dev/null
active+=("load-$name")
$KX -n cpulab wait --for=jsonpath='{.status.phase}'=Succeeded pod/load-$name --timeout=180s >/dev/null
t1=$(date +%s)
echo "===== $name replicas=$replicas limit=$limit request=500m rps=$RPS clump=${CLUMP:-1} ping_rps=${PINGRPS:-0} window_s=$((t1 - t0))"
$KX -n cpulab logs load-$name | grep -E '^(load|work_latency_ms|ping_latency_ms) '
i=0
for p in "${pods[@]}"; do
echo "pod $p $(delta "${before[$i]}" "$(cpustat "$p")" $((t1 - t0)))"
i=$((i + 1))
done
$KX -n cpulab delete pod "${pods[@]}" load-$name --wait=true >/dev/null
active=()
done
EOF
chmod +x scripts/e7-scale.sh
for round in 1 2; do
CLUMP=50 PINGRPS=20 scripts/e7-scale.sh "scripts/kx new" cpu-lab-new-worker cpu-lab-new-control-plane 100 \
e7-1x-nolimit:1:none e7-1x-limit2:1:2 e7-2x-limit1:2:1 e7-2x-limit2:2:2 e7-1x-limit4:1:4
done
五組依序執行,每組約一分半鐘,同樣跑兩次。第一次的輸出:
===== e7-1x-nolimit replicas=1 limit=none request=500m rps=100 clump=50 ping_rps=20 window_s=72
load label=e7-1x-nolimit gomaxprocs=10 clump=50 requests=5600 errors=0 achieved_rps=93.4
work_latency_ms p50=34.3 p90=55.4 p99=85.7 p999=98.0 max=101.7
ping_latency_ms requests=1213 p50=2.2 p90=4.7 p99=12.4 max=46.9
pod e7-1x-nolimit-1 avg_cpu_cores=0.604 nr_periods=0 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
===== e7-1x-limit2 replicas=1 limit=2 request=500m rps=100 clump=50 ping_rps=20 window_s=73
load label=e7-1x-limit2 gomaxprocs=10 clump=50 requests=5850 errors=0 achieved_rps=97.4
work_latency_ms p50=46.1 p90=136.1 p99=263.9 p999=379.4 max=431.0
ping_latency_ms requests=1207 p50=2.3 p90=15.4 p99=86.1 max=216.0
pod e7-1x-limit2-1 avg_cpu_cores=0.565 nr_periods=563 nr_throttled=133 throttled_ratio=0.236 throttled_ms=34491.9
===== e7-2x-limit1 replicas=2 limit=1 request=500m rps=100 clump=50 ping_rps=20 window_s=72
load label=e7-2x-limit1 gomaxprocs=10 clump=50 requests=6050 errors=0 achieved_rps=100.6
work_latency_ms p50=127.0 p90=356.6 p99=602.5 p999=961.0 max=1525.5
ping_latency_ms requests=1215 p50=3.1 p90=65.9 p99=224.6 max=702.8
pod e7-2x-limit1-1 avg_cpu_cores=0.150 nr_periods=229 nr_throttled=67 throttled_ratio=0.293 throttled_ms=24624.2
pod e7-2x-limit1-2 avg_cpu_cores=0.509 nr_periods=615 nr_throttled=325 throttled_ratio=0.528 throttled_ms=101140.6
===== e7-2x-limit2 replicas=2 limit=2 request=500m rps=100 clump=50 ping_rps=20 window_s=71
load label=e7-2x-limit2 gomaxprocs=10 clump=50 requests=6250 errors=0 achieved_rps=104.4
work_latency_ms p50=33.5 p90=82.8 p99=140.1 p999=202.9 max=220.3
ping_latency_ms requests=1226 p50=2.1 p90=5.1 p99=40.7 max=61.9
pod e7-2x-limit2-1 avg_cpu_cores=0.383 nr_periods=495 nr_throttled=51 throttled_ratio=0.103 throttled_ms=7771.2
pod e7-2x-limit2-2 avg_cpu_cores=0.258 nr_periods=357 nr_throttled=14 throttled_ratio=0.039 throttled_ms=985.2
===== e7-1x-limit4 replicas=1 limit=4 request=500m rps=100 clump=50 ping_rps=20 window_s=73
load label=e7-1x-limit4 gomaxprocs=10 clump=50 requests=6400 errors=0 achieved_rps=106.7
work_latency_ms p50=34.2 p90=58.9 p99=105.1 p999=156.1 max=182.3
ping_latency_ms requests=1227 p50=2.1 p90=4.8 p99=23.5 max=57.3
pod e7-1x-limit4-1 avg_cpu_cores=0.634 nr_periods=531 nr_throttled=21 throttled_ratio=0.040 throttled_ms=3717.7
第二次的輸出:
===== e7-1x-nolimit replicas=1 limit=none request=500m rps=100 clump=50 ping_rps=20 window_s=72
load label=e7-1x-nolimit gomaxprocs=10 clump=50 requests=6600 errors=0 achieved_rps=110.1
work_latency_ms p50=35.4 p90=59.5 p99=98.7 p999=143.1 max=150.0
ping_latency_ms requests=1124 p50=2.2 p90=4.8 p99=11.3 max=35.7
pod e7-1x-nolimit-1 avg_cpu_cores=0.689 nr_periods=0 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
===== e7-1x-limit2 replicas=1 limit=2 request=500m rps=100 clump=50 ping_rps=20 window_s=72
load label=e7-1x-limit2 gomaxprocs=10 clump=50 requests=5000 errors=0 achieved_rps=83.4
work_latency_ms p50=73.9 p90=162.8 p99=304.5 p999=520.0 max=525.2
ping_latency_ms requests=1211 p50=2.6 p90=14.5 p99=101.9 max=294.5
pod e7-1x-limit2-1 avg_cpu_cores=0.519 nr_periods=587 nr_throttled=125 throttled_ratio=0.213 throttled_ms=35555.8
===== e7-2x-limit1 replicas=2 limit=1 request=500m rps=100 clump=50 ping_rps=20 window_s=73
load label=e7-2x-limit1 gomaxprocs=10 clump=50 requests=5800 errors=0 achieved_rps=96.6
work_latency_ms p50=89.5 p90=213.2 p99=495.7 p999=689.7 max=693.2
ping_latency_ms requests=1209 p50=2.5 p90=28.2 p99=99.3 max=189.9
pod e7-2x-limit1-1 avg_cpu_cores=0.294 nr_periods=490 nr_throttled=163 throttled_ratio=0.333 throttled_ms=39898.4
pod e7-2x-limit1-2 avg_cpu_cores=0.291 nr_periods=449 nr_throttled=157 throttled_ratio=0.350 throttled_ms=39071.2
===== e7-2x-limit2 replicas=2 limit=2 request=500m rps=100 clump=50 ping_rps=20 window_s=73
load label=e7-2x-limit2 gomaxprocs=10 clump=50 requests=6400 errors=0 achieved_rps=106.7
work_latency_ms p50=34.5 p90=104.7 p99=166.5 p999=248.8 max=271.9
ping_latency_ms requests=1166 p50=2.3 p90=5.5 p99=52.1 max=159.9
pod e7-2x-limit2-1 avg_cpu_cores=0.445 nr_periods=517 nr_throttled=82 throttled_ratio=0.159 throttled_ms=17798.0
pod e7-2x-limit2-2 avg_cpu_cores=0.203 nr_periods=329 nr_throttled=12 throttled_ratio=0.036 throttled_ms=3015.5
===== e7-1x-limit4 replicas=1 limit=4 request=500m rps=100 clump=50 ping_rps=20 window_s=72
load label=e7-1x-limit4 gomaxprocs=10 clump=50 requests=5800 errors=0 achieved_rps=96.7
work_latency_ms p50=35.0 p90=58.0 p99=123.3 p999=176.2 max=192.6
ping_latency_ms requests=1200 p50=2.3 p90=4.7 p99=26.9 max=101.8
pod e7-1x-limit4-1 avg_cpu_cores=0.629 nr_periods=544 nr_throttled=9 throttled_ratio=0.017 throttled_ms=1092.7
整理成表(每格依序是第一次與第二次;兩個副本時,用量與 throttle 的比例是兩個 Pod 各自的數字):
| 設定 | 每個 Pod 的平均用量 | throttle 的 period | /work p99 | /ping p99 |
|---|---|---|---|---|
| 1 個 Pod,不設 limit | 0.60/0.69 | — | 86/99ms | 12/11ms |
| 1 個 Pod,limit 2 | 0.57/0.52 | 23.6%/21.3% | 264/305ms | 86/102ms |
| 2 個 Pod,各 limit 1 | 0.15 與 0.51/0.29 與 0.29 | 29% 與 53%/33% 與 35% | 603/496ms | 225/99ms |
| 2 個 Pod,各 limit 2 | 0.38 與 0.26/0.45 與 0.20 | 10% 與 4%/16% 與 4% | 140/167ms | 41/52ms |
| 1 個 Pod,limit 4 | 0.63/0.63 | 4.0%/1.7% | 105/123ms | 24/27ms |
- 同樣的 limit 總量拆給兩個副本,反而更差。
e7-2x-limit1的 p99 兩次都比e7-1x-limit2高。每個 Pod 分到約 25 個請求、約 150ms 的 CPU,仍然超過自己 100ms 的預算,而且一個 Pod 用不完的預算,另一個借不到。第二次兩個 Pod 的用量幾乎相同(0.29 與 0.29),結果仍然比較差,所以不只是分配不平均的問題。 - 兩個 limit 2 的副本,不如一個 limit 4 的 Pod。 limit 總和都是 4,但兩次執行中,兩個副本的用量都不平均,分到較多請求的那個 Pod 仍有 10% 到 16% 的 period 被 throttle。一個 limit 4 的 Pod,整群請求 300ms 的 CPU 在一個 period 內就裝得下。
- 提高 limit 最接近不設 limit。
e7-1x-limit4的 p99 是 105/123ms,不設 limit 是 86/99ms。 - 這一輪的
e7-1x-limit2比第 5 節好一些(p99 264/305ms,第 5 節是 370/360ms),是不同時間執行的差異。比較時看同一輪裡的相對關係。
副本之間分得不平均,推測和 Service 的分配方式有關:kube-proxy 在建立連線時選擇 Pod,壓測程式會重用連線,一群請求落在哪個 Pod,取決於它當時用到哪幾條連線。Service 的虛擬 IP 機制 實驗沒有記錄每個 Pod 的連線數,只觀察到用量不平均。
這個結果說明了:面對成群的請求,突發時單一 Pod 能用到的上限,比 limit 的總量更重要;只增加副本、不增加 limit 總量,沒有幫助。它沒有說明:Node 上沒有空閒 CPU 時會怎樣。這個實驗的 Node 有 10 顆 CPU,幾乎沒有其他工作負載,提高 limit 借得到空閒的 CPU;Node 很忙時的結果,在第 8 節。副本分散在不同 Node,或使用逐請求分配的負載平衡(例如 L7 proxy)時,結果也可能不同。
7. request 變成 cpu.weight
各層的 weight
建立幾個不同 request 的 Pod,讀出它們在各層的 cpu.weight。接著用 docker update --cpuset-cpus 把整個 worker Node 限制在一顆 CPU 上,讓這些持續忙碌的 container 互相搶 CPU,看實際各分到多少:
cat > scripts/e4-e5-weight.sh <<'EOF'
#!/usr/bin/env bash
# E4:不同 request 在各層 cgroup 的 cpu.weight。
# E5:把 worker 限在 1 顆 CPU,讓兩個忙碌的 thread 爭用,量測實際分到的比例。
# pods:兩個單一 container 的 Pod(request 100m 與 1),比較的是 Pod 層的 weight
# containers:同一個 Pod 裡的兩個 container(request 100m 與 1),比較的是 container 層的 weight
# 用法:e4-e5-weight.sh <kx 指令> <worker node>
set -euo pipefail
KX=$1; NODE=$2; here=$(cd "$(dirname "$0")/.." && pwd)
docker cp "$here/scripts/cgtree.sh" "$NODE:/cgtree.sh" >/dev/null
echo "== node runtime"; docker exec "$NODE" sh -c 'runc --version | head -1; containerd --version | cut -d" " -f3'
$KX get node "$NODE" -o jsonpath='kubelet={.status.nodeInfo.kubeletVersion} allocatable_cpu={.status.allocatable.cpu}{"\n"}'
pod() { # name spec-containers-json
$KX -n cpulab apply -f - >/dev/null <<YAML
{"apiVersion":"v1","kind":"Pod","metadata":{"name":"$1","namespace":"cpulab"},
"spec":{"nodeSelector":{"kubernetes.io/hostname":"$NODE"},"containers":$2}}
YAML
}
c() { # name request [limit]
local res="{\"requests\":{\"cpu\":\"$2\"}}"
[ $# -ge 3 ] && res="{\"requests\":{\"cpu\":\"$2\"},\"limits\":{\"cpu\":\"$3\"}}"
printf '{"name":"%s","image":"cpulab:go125mod","args":["spin","-threads=1","-report=10s"],"resources":%s}' "$1" "$res"
}
tree() { docker exec "$NODE" sh /cgtree.sh "$($KX -n cpulab get pod "$1" -o jsonpath='{.metadata.uid}')"; }
names() { # 印出 container 名稱與 cgroup scope 的對應
$KX -n cpulab get pod "$1" -o jsonpath='{range .status.containerStatuses[*]}{.name}={.containerID}{"\n"}{end}' | sed 's|containerd://\(.\{12\}\).*|cri-containerd-\1…|'
}
$KX get ns cpulab >/dev/null 2>&1 || $KX create ns cpulab >/dev/null
# 中途失敗或中斷時,由 trap 刪除這些持續運算的 Pod,並把 worker 還原為全部 CPU
# (空字串不會清除設定,需明確寫出範圍)
all="0-$(( $(docker info --format '{{.NCPU}}') - 1 ))"
restore() { docker update --cpuset-cpus "$all" "$NODE" >/dev/null; }
cleanup() {
$KX -n cpulab delete pod e4-req100m e4-req1 e4-req2 e4-req1-limit1 e5-two-containers \
--ignore-not-found --wait=false >/dev/null 2>&1 || true
restore
}
trap cleanup EXIT
# 先不限制 CPU,讀出各層 weight(spin 會跑,但 10 顆 CPU 不會爭用)
pod e4-req100m "[$(c app 100m)]"
pod e4-req1 "[$(c app 1)]"
pod e4-req2 "[$(c app 2)]"
pod e4-req1-limit1 "[$(c app 1 1)]" # 只設 CPU,沒有 memory,所以仍是 Burstable
pod e5-two-containers "[$(c small 100m),$(c big 1)]"
$KX -n cpulab wait --for=condition=Ready pod e4-req100m e4-req1 e4-req2 e4-req1-limit1 e5-two-containers --timeout=120s >/dev/null
echo "== E4 cgroup tree"
for p in e4-req100m e4-req1 e4-req2 e4-req1-limit1 e5-two-containers; do echo "-- $p"; tree $p; names $p; done
echo "-- root and kubepods siblings"
docker exec "$NODE" sh -c 'for d in /sys/fs/cgroup/*/ /sys/fs/cgroup/kubelet.slice/*/ /sys/fs/cgroup/kubelet.slice/kubelet-kubepods.slice/*/; do [ -f $d/cpu.weight ] && printf "%-90s weight=%s\n" "${d#/sys/fs/cgroup}" "$(cat $d/cpu.weight)"; done | grep -v "pod[0-9a-f_]*.slice/$"'
echo "== E5 pin worker to one CPU"
$KX -n cpulab delete pod e4-req2 e4-req1-limit1 --wait=true >/dev/null
docker update --cpuset-cpus 5 "$NODE" >/dev/null
docker exec "$NODE" cat /sys/fs/cgroup/cpuset.cpus.effective
sleep 45
echo "-- pods: e4-req100m vs e4-req1 (Pod 層比較),同時 e5-two-containers 也在跑"
for p in e4-req100m e4-req1; do echo "$p: $($KX -n cpulab logs $p --tail=3 | tr '\n' ' ')"; done
for cn in small big; do echo "e5-two-containers/$cn: $($KX -n cpulab logs e5-two-containers -c $cn --tail=3 | tr '\n' ' ')"; done
echo "-- only the two-container pod (container 層比較)"
$KX -n cpulab delete pod e4-req100m e4-req1 --wait=true >/dev/null
sleep 45
for cn in small big; do echo "e5-two-containers/$cn: $($KX -n cpulab logs e5-two-containers -c $cn --tail=3 | tr '\n' ' ')"; done
restore
$KX -n cpulab delete pod e5-two-containers --wait=false >/dev/null
echo "== restored cpuset: $(docker exec "$NODE" cat /sys/fs/cgroup/cpuset.cpus.effective)"
EOF
chmod +x scripts/e4-e5-weight.sh
scripts/e4-e5-weight.sh "scripts/kx new" cpu-lab-new-worker
節錄 request 1 的 Pod、兩個 container 的 Pod,以及 root 附近的幾層:
== node runtime
runc version 1.4.3
v2.3.4
kubelet=v1.34.11 allocatable_cpu=10
== E4 cgroup tree
...
-- e4-req1
/kubelet.slice weight=100 max=max 100000
/kubelet.slice/kubelet-kubepods.slice weight=391 max=max 100000
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-burstable.slice weight=207 max=max 100000
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-burstable.slice/kubelet-kubepods-burstable-podc54295e7_f94f_4a91_9bea_bf0393eacb82.slice weight=39 max=max 100000
cri-containerd-820a1ebf6ed44fe7d1f184221ddcc247daa6a5590210df041f60bd42d3c60650.scope weight=1 max=max 100000
cri-containerd-ccd098da4e5c8c486a18746d717a2fbdeefa7af21d56afc5c2c1bfa1e90f70e5.scope weight=100 max=max 100000
app=cri-containerd-ccd098da4e5c…
...
-- e5-two-containers
/kubelet.slice weight=100 max=max 100000
/kubelet.slice/kubelet-kubepods.slice weight=391 max=max 100000
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-burstable.slice weight=207 max=max 100000
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-burstable.slice/kubelet-kubepods-burstable-pod767f0518_7b92_487f_a740_9b8c96d55a0f.slice weight=43 max=max 100000
cri-containerd-773b31d845ecdb884077fbe73ce281f319333a73b1078b7f2a22ddc50eaeebf3.scope weight=17 max=max 100000
cri-containerd-77eff6a412b38a06aa05c58fc7019d6cee64124e0ed30575d62a44e7effb7016.scope weight=1 max=max 100000
cri-containerd-fcf7d3f159249325fadaa4634a4ef30feadbe8900561858d36faf9d7d2e968ab.scope weight=100 max=max 100000
big=cri-containerd-fcf7d3f15924…
small=cri-containerd-773b31d845ec…
-- root and kubepods siblings
/init.scope/ weight=100
/kubelet.slice/ weight=100
/kubelet/ weight=100
/sys-fs-fuse-connections.mount/ weight=100
/sys-kernel-debug.mount/ weight=100
/sys-kernel-tracing.mount/ weight=100
/system.slice/ weight=100
/kubelet.slice/kubelet-kubepods.slice/ weight=391
/kubelet.slice/kubelet.service/ weight=100
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-besteffort.slice/ weight=1
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-burstable.slice/ weight=207
Pod 層與 container 層是兩個不同的數字:
| request | Pod 層(kubelet) | container 層(runc 1.4.3) |
|---|---|---|
| 100m | 4 | 17 |
| 1 | 39 | 100 |
| 2 | 79 | 174 |
Pod、QoS 與 kubepods 這幾層由 kubelet 建立,用的是線性換算,1 顆 CPU 是 39;container 層由 runc 建立,runc 1.3.2 起改用新公式,1 顆 CPU 剛好是 cgroup v2 的預設值 100。kubepods 的 391 由 Node 的 Allocatable(10 CPU)換算而來,和它同一層的 kubelet.service 是 100。
一顆 CPU 上的實際分配
== E5 pin worker to one CPU
5
-- pods: e4-req100m vs e4-req1 (Pod 層比較),同時 e5-two-containers 也在跑
e4-req100m: 2026-09-26T06:06:26Z usage_cores=0.046 2026-09-26T06:06:36Z usage_cores=0.046 2026-09-26T06:06:46Z usage_cores=0.046
e4-req1: 2026-09-26T06:06:26Z usage_cores=0.449 2026-09-26T06:06:36Z usage_cores=0.447 2026-09-26T06:06:46Z usage_cores=0.445
e5-two-containers/small: 2026-09-26T06:06:26Z usage_cores=0.072 2026-09-26T06:06:36Z usage_cores=0.072 2026-09-26T06:06:46Z usage_cores=0.071
e5-two-containers/big: 2026-09-26T06:06:26Z usage_cores=0.424 2026-09-26T06:06:36Z usage_cores=0.421 2026-09-26T06:06:46Z usage_cores=0.420
-- only the two-container pod (container 層比較)
e5-two-containers/small: 2026-09-26T06:07:16Z usage_cores=0.144 2026-09-26T06:07:26Z usage_cores=0.144 2026-09-26T06:07:36Z usage_cores=0.144
e5-two-containers/big: 2026-09-26T06:07:16Z usage_cores=0.846 2026-09-26T06:07:26Z usage_cores=0.848 2026-09-26T06:07:36Z usage_cores=0.848
輸出分成兩段:
- 三個忙碌的 Pod 一起搶。
e4-req100m、e4-req1與兩個 container 的e5-two-containers(request 合計 1.1)分到 0.046、0.447 與 0.493(兩個 container 相加),約 4.7% : 45.3% : 50.0%,和 Pod 層的 weight 4 : 39 : 43 換算的比例相同。 - 只剩兩個 container 的 Pod。 同一個 Pod 裡,request 100m 與 1 的兩個 container 分到 0.144 : 0.847,約 14.5% : 85.5%,對應 container 層的 17 : 100,而不是 request 的 1 : 10。
每組量測前等 45 秒,讓分配穩定下來。每個 container 每 10 秒印一次用量,上面的數字是最後三次讀數的平均。
還原 CPU:這次踩到的坑
腳本最後用 docker update 把 worker 的 cpuset 改回全部 CPU。第一次執行時,還原指令傳的是空字串 --cpuset-cpus "",結果沒有生效,輸出最後一行仍是 restored cpuset: 5,worker 還是只有一顆 CPU。腳本後來改成明確寫出範圍(0-9 這種形式),並用 trap 在腳本結束時刪除持續運算的 Pod、再還原一次 cpuset,中途失敗或按 Ctrl-C 也不會讓 worker 停在一顆 CPU 上,或留下一直在跑的 Pod,就是上面的版本。
執行後確認一下;如果不是全部的 CPU,手動改回來:
docker exec cpu-lab-new-worker cat /sys/fs/cgroup/cpuset.cpus.effective
docker update --cpuset-cpus "0-$(( $(docker info --format '{{.NCPU}}') - 1 ))" cpu-lab-new-worker
這個結果說明了:request 在執行時變成各層的 cpu.weight;Pod 之間依 kubelet 的換算分配,同一個 Pod 裡的 container 之間依 runtime 的換算分配。它沒有說明:真實 Node 上 Pod 和系統服務的競爭。kind 的 Node 本身是 container,root 附近的幾層和真實主機不同。
8. Node 很忙時,提高 limit 還有用嗎?
第 6 節的 Node 幾乎沒有其他工作,提高 limit 借得到空閒的 CPU。這一節在同一個 Node 上多放一個持續用滿 CPU 的鄰居,看提高 limit 還有沒有用,以及 request 這時扮演什麼角色。
- 流量:和第 6 節相同,每秒平均 100 個
/work、每次 50 個一起到,另有每秒 20 個/ping。server 只有一個 Pod,用 go.mod 1.24 的 image。 - 鄰居:request 4、不設 limit,用
cpulab spin讓 8 個 thread 持續運算,代表 Node 上其他沒有 limit、正在全速執行的工作。 - 分開 CPU:兩個 kind Node 其實是同一台 VM 上的 container,共用 10 顆 CPU。鄰居如果把 10 顆全部用滿,control-plane 上的壓測程式也會變慢,量到的延遲就混進了壓測端的誤差。所以腳本先用
docker update --cpuset-cpus把 worker 限制在 CPU 0-7、control-plane 限制在 8-9。結束時,或中途失敗、按 Ctrl-C 時,trap會刪除還在跑的 Pod(鄰居不會一直佔著 CPU),再還原 cpuset。worker 只剩 8 顆 CPU,server 的 GOMAXPROCS 也變成 8,所以這一節只和同一輪的「Node 閒」組比較,不和第 6 節的數字比。
六種設定依序執行:
| 名稱 | server 的 request | limit | 鄰居 |
|---|---|---|---|
e8-idle-limit2 | 500m | 2 | 無 |
e8-idle-limit4 | 500m | 4 | 無 |
e8-busy-limit2 | 500m | 2 | 有 |
e8-busy-limit4 | 500m | 4 | 有 |
e8-busy-req2-limit4 | 2 | 4 | 有 |
e8-busy-req4-limit4 | 4 | 4 | 有 |
依第 7 節 kubelet 的換算,Pod 層的 weight 是 500m → 20、2 → 79、4 → 157。鄰居的 request 是 4,所以 server 的 request 從 500m 提高到 4,Node 滿載時分到的比例依序約是 20 : 157、79 : 157 與 157 : 157。
cat > scripts/e8-contention.sh <<'EOF'
#!/usr/bin/env bash
# E8:同樣的成群流量,Node 閒著與 Node 被鄰居用滿時,比較只提高 limit 與連 request 一起調高。
# 用法:e8-contention.sh <kx 指令> <worker node> <cp node> <rps> <name:request:limit:idle|busy>...
# 壓測程式和 server 在同一台 VM 上,為了不讓鄰居也拖慢壓測程式,先用 cpuset 把兩個 kind node 分開:
# worker 用 CPU 0-7,control-plane 用 8-9。busy 時,鄰居是 request 4、不設 limit、8 個 thread 持續運算的 Pod。
# 中途失敗或中斷時,trap 會刪除還在跑的 Pod(鄰居不會一直佔著 CPU),再還原 cpuset。
set -euo pipefail
KX=$1; NODE=$2; CP=$3; RPS=$4; shift 4
all="0-$(( $(docker info --format '{{.NCPU}}') - 1 ))"
restore() { docker update --cpuset-cpus "$all" "$NODE" "$CP" >/dev/null; }
active=() # 目前還在跑的 Pod
cleanup() {
[ ${#active[@]} -eq 0 ] || $KX -n cpulab delete pod "${active[@]}" --ignore-not-found --wait=false >/dev/null 2>&1 || true
restore
}
trap cleanup EXIT
docker update --cpuset-cpus 0-7 "$NODE" >/dev/null
docker update --cpuset-cpus 8-9 "$CP" >/dev/null
echo "== cpuset worker=$(docker exec "$NODE" cat /sys/fs/cgroup/cpuset.cpus.effective) control-plane=$(docker exec "$CP" cat /sys/fs/cgroup/cpuset.cpus.effective)"
$KX get ns cpulab >/dev/null 2>&1 || $KX create ns cpulab >/dev/null
$KX -n cpulab get svc cpulab-server >/dev/null 2>&1 || \
$KX -n cpulab create service clusterip cpulab-server --tcp=8080:8080 >/dev/null
$KX -n cpulab patch svc cpulab-server -p '{"spec":{"selector":{"app":"cpulab-server"}}}' >/dev/null
cpustat() { $KX -n cpulab exec "$1" -- cat /sys/fs/cgroup/cpu.stat | awk '{printf "%s=%s ", $1, $2}'; }
delta() { # before after elapsed_s -> 這段時間的用量與 throttle
printf '%s\n%s\n' "$1" "$2" | awk -v t="$3" '
{ for (i = 1; i <= NF; i++) { split($i, kv, "="); v[NR, kv[1]] = kv[2] } }
END {
u = v[2, "usage_usec"] - v[1, "usage_usec"]; p = v[2, "nr_periods"] - v[1, "nr_periods"]
n = v[2, "nr_throttled"] - v[1, "nr_throttled"]; s = v[2, "throttled_usec"] - v[1, "throttled_usec"]
printf "avg_cpu_cores=%.3f nr_periods=%d nr_throttled=%d throttled_ratio=%.3f throttled_ms=%.1f\n", u / t / 1e6, p, n, (p ? n / p : 0), s / 1000
}'
}
for spec in "$@"; do
IFS=: read -r name request limit node <<<"$spec"
res="{\"requests\":{\"cpu\":\"$request\"}}"
[ "$limit" != none ] && res="{\"requests\":{\"cpu\":\"$request\"},\"limits\":{\"cpu\":\"$limit\"}}"
pods=("$name")
$KX -n cpulab run "$name" --image=cpulab:go124mod --restart=Never --labels=app=cpulab-server \
--overrides="{\"spec\":{\"nodeSelector\":{\"kubernetes.io/hostname\":\"$NODE\"},\"containers\":[{\"name\":\"c\",\"image\":\"cpulab:go124mod\",\"args\":[\"serve\",\"-work=5ms\",\"-alloc=1048576\",\"-live-mb=64\"],\"resources\":$res,\"ports\":[{\"containerPort\":8080}],\"readinessProbe\":{\"httpGet\":{\"path\":\"/stats\",\"port\":8080}}}]}}" >/dev/null
active=("$name")
if [ "$node" = busy ]; then
$KX -n cpulab run "$name-neighbor" --image=cpulab:go124mod --restart=Never \
--overrides="{\"spec\":{\"nodeSelector\":{\"kubernetes.io/hostname\":\"$NODE\"},\"containers\":[{\"name\":\"c\",\"image\":\"cpulab:go124mod\",\"args\":[\"spin\",\"-threads=8\",\"-report=10s\"],\"resources\":{\"requests\":{\"cpu\":\"4\"}}}]}}" >/dev/null
pods+=("$name-neighbor"); active+=("$name-neighbor")
fi
$KX -n cpulab wait --for=condition=Ready pod "${pods[@]}" --timeout=120s >/dev/null
sleep 15 # 讓鄰居先把 CPU 用滿
before=()
for p in "${pods[@]}"; do before+=("$(cpustat "$p")"); done
t0=$(date +%s)
$KX -n cpulab run "load-$name" --image=cpulab:go125mod --restart=Never \
--overrides="{\"spec\":{\"nodeSelector\":{\"kubernetes.io/hostname\":\"$CP\"},\"tolerations\":[{\"operator\":\"Exists\"}],\"containers\":[{\"name\":\"c\",\"image\":\"cpulab:go125mod\",\"args\":[\"load\",\"-target=http://cpulab-server:8080\",\"-rps=$RPS\",\"-clump=${CLUMP:-1}\",\"-ping-rps=${PINGRPS:-0}\",\"-duration=60s\",\"-warmup=10s\",\"-label=$name\"]}]}}" >/dev/null
active+=("load-$name")
$KX -n cpulab wait --for=jsonpath='{.status.phase}'=Succeeded pod/load-$name --timeout=180s >/dev/null
t1=$(date +%s)
echo "===== $name request=$request limit=$limit node=$node rps=$RPS clump=${CLUMP:-1} ping_rps=${PINGRPS:-0} window_s=$((t1 - t0))"
$KX -n cpulab logs "$name" | grep -m1 '^go='
$KX -n cpulab logs load-$name | grep -E '^(load|work_latency_ms|ping_latency_ms) '
i=0
for p in "${pods[@]}"; do
echo "pod $p $(delta "${before[$i]}" "$(cpustat "$p")" $((t1 - t0)))"
i=$((i + 1))
done
$KX -n cpulab delete pod "${pods[@]}" load-$name --wait=true >/dev/null
active=()
done
restore
echo "== restored cpuset worker=$(docker exec "$NODE" cat /sys/fs/cgroup/cpuset.cpus.effective) control-plane=$(docker exec "$CP" cat /sys/fs/cgroup/cpuset.cpus.effective)"
EOF
chmod +x scripts/e8-contention.sh
for round in 1 2; do
CLUMP=50 PINGRPS=20 scripts/e8-contention.sh "scripts/kx new" cpu-lab-new-worker cpu-lab-new-control-plane 100 \
e8-idle-limit2:500m:2:idle e8-idle-limit4:500m:4:idle e8-busy-limit2:500m:2:busy e8-busy-limit4:500m:4:busy \
e8-busy-req2-limit4:2:4:busy e8-busy-req4-limit4:4:4:busy
done
每組約一分半鐘,同樣跑兩次。第一次的輸出:
== cpuset worker=0-7 control-plane=8-9
===== e8-idle-limit2 request=500m limit=2 node=idle rps=100 clump=50 ping_rps=20 window_s=72
go=go1.26.0 GOMAXPROCS=8 NumCPU=8 GOMAXPROCS_env="" GODEBUG_env=""
load label=e8-idle-limit2 gomaxprocs=8 clump=50 requests=6300 errors=0 achieved_rps=105.3
work_latency_ms p50=69.7 p90=198.2 p99=392.2 p999=616.5 max=641.1
ping_latency_ms requests=1162 p50=2.5 p90=22.6 p99=102.1 max=596.6
pod e8-idle-limit2 avg_cpu_cores=0.617 nr_periods=587 nr_throttled=153 throttled_ratio=0.261 throttled_ms=29664.4
===== e8-idle-limit4 request=500m limit=4 node=idle rps=100 clump=50 ping_rps=20 window_s=72
go=go1.26.0 GOMAXPROCS=8 NumCPU=8 GOMAXPROCS_env="" GODEBUG_env=""
load label=e8-idle-limit4 gomaxprocs=8 clump=50 requests=6000 errors=0 achieved_rps=100.1
work_latency_ms p50=38.4 p90=64.8 p99=98.7 p999=125.3 max=130.4
ping_latency_ms requests=1210 p50=2.2 p90=3.9 p99=13.7 max=41.8
pod e8-idle-limit4 avg_cpu_cores=0.620 nr_periods=560 nr_throttled=10 throttled_ratio=0.018 throttled_ms=602.9
===== e8-busy-limit2 request=500m limit=2 node=busy rps=100 clump=50 ping_rps=20 window_s=72
go=go1.26.0 GOMAXPROCS=8 NumCPU=8 GOMAXPROCS_env="" GODEBUG_env=""
load label=e8-busy-limit2 gomaxprocs=8 clump=50 requests=7048 errors=2 achieved_rps=117.1
work_latency_ms p50=637.3 p90=3372.0 p99=8358.8 p999=9917.5 max=9941.8
ping_latency_ms requests=1207 p50=45.2 p90=1954.6 p99=4257.4 max=8340.0
pod e8-busy-limit2 avg_cpu_cores=0.665 nr_periods=636 nr_throttled=19 throttled_ratio=0.030 throttled_ms=471.5
pod e8-busy-limit2-neighbor avg_cpu_cores=7.341 nr_periods=0 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
===== e8-busy-limit4 request=500m limit=4 node=busy rps=100 clump=50 ping_rps=20 window_s=72
go=go1.26.0 GOMAXPROCS=8 NumCPU=8 GOMAXPROCS_env="" GODEBUG_env=""
load label=e8-busy-limit4 gomaxprocs=8 clump=50 requests=6100 errors=0 achieved_rps=101.5
work_latency_ms p50=553.5 p90=2916.2 p99=6044.7 p999=8730.0 max=9278.3
ping_latency_ms requests=1186 p50=32.8 p90=1026.4 p99=4513.0 max=8332.0
pod e8-busy-limit4 avg_cpu_cores=0.603 nr_periods=654 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
pod e8-busy-limit4-neighbor avg_cpu_cores=7.457 nr_periods=0 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
===== e8-busy-req2-limit4 request=2 limit=4 node=busy rps=100 clump=50 ping_rps=20 window_s=72
go=go1.26.0 GOMAXPROCS=8 NumCPU=8 GOMAXPROCS_env="" GODEBUG_env=""
load label=e8-busy-req2-limit4 gomaxprocs=8 clump=50 requests=5800 errors=0 achieved_rps=96.7
work_latency_ms p50=81.5 p90=188.5 p99=471.8 p999=564.8 max=666.0
ping_latency_ms requests=1200 p50=2.3 p90=15.6 p99=119.3 max=261.4
pod e8-busy-req2-limit4 avg_cpu_cores=0.565 nr_periods=468 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
pod e8-busy-req2-limit4-neighbor avg_cpu_cores=7.355 nr_periods=0 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
===== e8-busy-req4-limit4 request=4 limit=4 node=busy rps=100 clump=50 ping_rps=20 window_s=72
go=go1.26.0 GOMAXPROCS=8 NumCPU=8 GOMAXPROCS_env="" GODEBUG_env=""
load label=e8-busy-req4-limit4 gomaxprocs=8 clump=50 requests=6400 errors=0 achieved_rps=106.7
work_latency_ms p50=70.7 p90=134.5 p99=222.3 p999=289.3 max=330.6
ping_latency_ms requests=1197 p50=2.3 p90=9.3 p99=62.0 max=119.8
pod e8-busy-req4-limit4 avg_cpu_cores=0.581 nr_periods=463 nr_throttled=1 throttled_ratio=0.002 throttled_ms=8.4
pod e8-busy-req4-limit4-neighbor avg_cpu_cores=7.354 nr_periods=0 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
== restored cpuset worker=0-9 control-plane=0-9
第二次的輸出:
== cpuset worker=0-7 control-plane=8-9
===== e8-idle-limit2 request=500m limit=2 node=idle rps=100 clump=50 ping_rps=20 window_s=72
go=go1.26.0 GOMAXPROCS=8 NumCPU=8 GOMAXPROCS_env="" GODEBUG_env=""
load label=e8-idle-limit2 gomaxprocs=8 clump=50 requests=6600 errors=0 achieved_rps=110.1
work_latency_ms p50=79.1 p90=207.3 p99=492.0 p999=680.2 max=828.7
ping_latency_ms requests=1186 p50=2.6 p90=33.0 p99=197.8 max=309.4
pod e8-idle-limit2 avg_cpu_cores=0.656 nr_periods=606 nr_throttled=158 throttled_ratio=0.261 throttled_ms=27100.3
===== e8-idle-limit4 request=500m limit=4 node=idle rps=100 clump=50 ping_rps=20 window_s=72
go=go1.26.0 GOMAXPROCS=8 NumCPU=8 GOMAXPROCS_env="" GODEBUG_env=""
load label=e8-idle-limit4 gomaxprocs=8 clump=50 requests=7100 errors=0 achieved_rps=118.6
work_latency_ms p50=38.7 p90=67.3 p99=137.3 p999=203.4 max=216.3
ping_latency_ms requests=1198 p50=2.2 p90=5.0 p99=31.6 max=51.9
pod e8-idle-limit4 avg_cpu_cores=0.694 nr_periods=565 nr_throttled=23 throttled_ratio=0.041 throttled_ms=2570.3
===== e8-busy-limit2 request=500m limit=2 node=busy rps=100 clump=50 ping_rps=20 window_s=73
go=go1.26.0 GOMAXPROCS=8 NumCPU=8 GOMAXPROCS_env="" GODEBUG_env=""
load label=e8-busy-limit2 gomaxprocs=8 clump=50 requests=6600 errors=0 achieved_rps=110.1
work_latency_ms p50=437.7 p90=1476.7 p99=3589.7 p999=5414.4 max=5518.4
ping_latency_ms requests=1229 p50=38.0 p90=574.5 p99=3033.9 max=3480.7
pod e8-busy-limit2 avg_cpu_cores=0.614 nr_periods=634 nr_throttled=15 throttled_ratio=0.024 throttled_ms=543.0
pod e8-busy-limit2-neighbor avg_cpu_cores=7.344 nr_periods=0 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
===== e8-busy-limit4 request=500m limit=4 node=busy rps=100 clump=50 ping_rps=20 window_s=73
go=go1.26.0 GOMAXPROCS=8 NumCPU=8 GOMAXPROCS_env="" GODEBUG_env=""
load label=e8-busy-limit4 gomaxprocs=8 clump=50 requests=5800 errors=0 achieved_rps=96.3
work_latency_ms p50=361.7 p90=1675.5 p99=3136.6 p999=3575.5 max=4230.2
ping_latency_ms requests=1176 p50=4.7 p90=346.6 p99=1664.3 max=3199.5
pod e8-busy-limit4 avg_cpu_cores=0.541 nr_periods=593 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
pod e8-busy-limit4-neighbor avg_cpu_cores=7.423 nr_periods=0 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
===== e8-busy-req2-limit4 request=2 limit=4 node=busy rps=100 clump=50 ping_rps=20 window_s=73
go=go1.26.0 GOMAXPROCS=8 NumCPU=8 GOMAXPROCS_env="" GODEBUG_env=""
load label=e8-busy-req2-limit4 gomaxprocs=8 clump=50 requests=6000 errors=0 achieved_rps=100.1
work_latency_ms p50=84.7 p90=188.5 p99=341.1 p999=494.6 max=598.4
ping_latency_ms requests=1238 p50=2.2 p90=13.2 p99=99.7 max=410.7
pod e8-busy-req2-limit4 avg_cpu_cores=0.589 nr_periods=460 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
pod e8-busy-req2-limit4-neighbor avg_cpu_cores=7.281 nr_periods=0 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
===== e8-busy-req4-limit4 request=4 limit=4 node=busy rps=100 clump=50 ping_rps=20 window_s=72
go=go1.26.0 GOMAXPROCS=8 NumCPU=8 GOMAXPROCS_env="" GODEBUG_env=""
load label=e8-busy-req4-limit4 gomaxprocs=8 clump=50 requests=6500 errors=0 achieved_rps=108.4
work_latency_ms p50=74.8 p90=141.7 p99=213.4 p999=267.9 max=287.2
ping_latency_ms requests=1230 p50=2.1 p90=8.0 p99=77.8 max=213.6
pod e8-busy-req4-limit4 avg_cpu_cores=0.655 nr_periods=465 nr_throttled=2 throttled_ratio=0.004 throttled_ms=14.9
pod e8-busy-req4-limit4-neighbor avg_cpu_cores=7.306 nr_periods=0 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
== restored cpuset worker=0-9 control-plane=0-9
整理成表(每格依序是第一次與第二次):
| 設定 | server 的平均用量 | throttle 的 period | /work p99 | /ping p99 |
|---|---|---|---|---|
e8-idle-limit2 | 0.62/0.66 | 26.1%/26.1% | 392/492ms | 102/198ms |
e8-idle-limit4 | 0.62/0.69 | 1.8%/4.1% | 99/137ms | 14/32ms |
e8-busy-limit2 | 0.67/0.61 | 3.0%/2.4% | 8,359/3,590ms | 4,257/3,034ms |
e8-busy-limit4 | 0.60/0.54 | 0%/0% | 6,045/3,137ms | 4,513/1,664ms |
e8-busy-req2-limit4 | 0.57/0.59 | 0%/0% | 472/341ms | 119/100ms |
e8-busy-req4-limit4 | 0.58/0.66 | 0.2%/0.4% | 222/213ms | 62/78ms |
- Node 忙時,limit 2 與 limit 4 都要好幾秒。 throttle 幾乎是 0,limit 4 完全沒有,p50 也有 360 到 640ms。server 的平均用量和 Node 閒時差不多(0.54 到 0.67 顆),但突發時分不到 CPU。依上面的 weight,Node 滿載時 server 推算只分得到 8 × 20 ÷ 177,約 0.9 顆 CPU;這是推算,腳本沒有量到突發當下實際分到多少。
- throttle 的指標看不出問題。 如果只看 throttle,
e8-busy-limit4是六組裡最好的之一,延遲卻是最差的兩組之一。 - request 調高,延遲依序改善。 request 2 的 p99 是 341/472ms,request 4 是 213/222ms。request 4 和鄰居的 weight 相同,推算在 Node 滿載時分得到約 4 顆,剛好等於 limit,實際仍比
e8-idle-limit4慢。weight 決定的是一段時間內的比例,推測突發剛開始時,server 的 thread 仍要和正在執行的鄰居輪流;這一點沒有量測。 - 鄰居一直用掉 7.3 到 7.5 顆 CPU,也就是 8 顆裡 server 沒有用到的部分。
- Node 忙的幾組,兩次差異很大(
e8-busy-limit2是 8.4 與 3.6 秒)。服務處在排隊邊緣時,幾群請求剛好靠得近一點,隊伍就拉長,延遲對到達的時間很敏感,只看數量級即可。第一次的e8-busy-limit2另有 2 個請求失敗(errors=2),不計入延遲。壓測程式的 timeout 是 10 秒,最長的成功請求是 9.9 秒,推測是逾時;當時的程式把/work與/ping的錯誤加在一起,所以不確定是哪一種請求的 p99 因此略為低估(附錄的程式已改成分開計算)。
這個結果說明了:Node 忙時,limit 只是上限。limit 調高之後 throttle 更少,延遲卻是秒級;決定能拿到多少 CPU 的,是 request 換算的 weight。它沒有說明:真實 Node 上會多忙。鄰居是刻意做出的極端情況,持續用滿 Node、request 4、不設 limit;真實 Node 很少長時間滿載,較常見的是幾個 Pod 的突發剛好重疊。結果的大小也取決於鄰居的 request。worker 只有 8 顆 CPU、壓測程式在另外 2 顆,數字和第 6 節不能直接比較。
9. 監控看到的數字
最後看監控系統實際讀到的指標。沿用第 3 節 limit 1 的工作(8 個 thread 各 20ms,每秒一次),在執行途中兩次讀取 kubelet 內建的 cAdvisor 指標,相隔約 20 秒:
cat > scripts/e6-cadvisor.sh <<'EOF'
#!/usr/bin/env bash
# E6:監控實際讀到的 cAdvisor 指標,對照 cpu.stat 與經過的時間。
# 用法:e6-cadvisor.sh <kx 指令> <worker node>
set -euo pipefail
KX=$1; NODE=$2
echo "== E6: cAdvisor metrics for a throttled burst pod (limit 1, 8 threads x 20ms, 30 rounds)"
$KX -n cpulab run e6-burst --image=cpulab:go125mod --restart=Never --overrides="{\"spec\":{\"nodeSelector\":{\"kubernetes.io/hostname\":\"$NODE\"},\"containers\":[{\"name\":\"c\",\"image\":\"cpulab:go125mod\",\"args\":[\"burst\",\"-threads=8\",\"-work=20ms\",\"-interval=1s\",\"-rounds=30\"],\"resources\":{\"requests\":{\"cpu\":\"100m\"},\"limits\":{\"cpu\":\"1\"}},\"env\":[{\"name\":\"GOMAXPROCS\",\"value\":\"8\"}]}]}}" >/dev/null
$KX -n cpulab wait --for=condition=Ready pod/e6-burst --timeout=60s >/dev/null
m() { $KX get --raw "/api/v1/nodes/$NODE/proxy/metrics/cadvisor" | grep -E '^container_cpu_(cfs_periods_total|cfs_throttled_periods_total|cfs_throttled_seconds_total|usage_seconds_total)\{' | grep 'pod="e6-burst"' | grep 'container="c"' | sed -E 's/\{[^}]*\}//'; }
sleep 3; echo "-- t0 $(date -u +%T)"; m
sleep 20; echo "-- t0+20s $(date -u +%T)"; m
$KX -n cpulab wait --for=jsonpath='{.status.phase}'=Succeeded pod/e6-burst --timeout=60s >/dev/null
$KX -n cpulab logs e6-burst | tail -1
$KX -n cpulab delete pod e6-burst --wait=false >/dev/null
EOF
chmod +x scripts/e6-cadvisor.sh
scripts/e6-cadvisor.sh "scripts/kx new" cpu-lab-new-worker
== E6: cAdvisor metrics for a throttled burst pod (limit 1, 8 threads x 20ms, 30 rounds)
-- t0 06:11:40
container_cpu_cfs_periods_total 1 1790403097312
container_cpu_cfs_throttled_periods_total 0 1790403097312
container_cpu_cfs_throttled_seconds_total 0 1790403097312
container_cpu_usage_seconds_total 0.100817 1790403097312
-- t0+20s 06:12:00
container_cpu_cfs_periods_total 38 1790403109461
container_cpu_cfs_throttled_periods_total 13 1790403109461
container_cpu_cfs_throttled_seconds_total 5.77952 1790403109461
container_cpu_usage_seconds_total 2.169029 1790403109461
summary wall_ms p50=108.2 max=118.4 avg_cpu_cores=0.165 nr_periods=90 nr_throttled=30 throttled_ms=14215.2
每個指標後面的第二個數字是 cAdvisor 取樣的時間戳(毫秒)。cAdvisor 定期收集一次,回傳的是最近一次收集的值,所以兩次取樣只相差 12.149 秒,不是腳本等待的 20 秒。換算 rate 時,分母要用時間戳的差:
| 指標 | 增加量 | 除以 12.15 秒 |
|---|---|---|
container_cpu_usage_seconds_total | 2.07 | 0.17 |
container_cpu_cfs_periods_total | 37 | — |
container_cpu_cfs_throttled_periods_total | 13 | — |
container_cpu_cfs_throttled_seconds_total | 5.78 | 0.48 |
在 Prometheus 裡,rate(container_cpu_cfs_throttled_seconds_total[1m]) 大約就是 0.48,看起來像「每秒停了將近半秒」。但這個程式每秒只停一次:最後一行的總結是完成時間 p50 108ms,不被 throttle 時約 34ms,實際每秒停頓約 75ms。0.48 是 8 個 thread 各自被停住的時間加總。
throttled_periods / periods 是 13 / 37,約 35%。分母只算忙碌的 period:12 秒裡有 120 個 period,但只計入了 37 個。
這個結果說明了:throttled_seconds 適合比較同一個服務的變化,不能直接當成暫停的比例;throttled period 的比例,分母也不是時間。它沒有說明:使用者有沒有受影響。要回答這個,得看延遲。
10. 選做:舊版 runc 的對照
第 7 節的 container 層 weight 由 runtime 換算。要看舊公式,可以用 kind v0.30.0 和它預設的 kindest/node:v1.34.0 建立第二個叢集,裡面是 runc 1.3.0。kind 的執行檔要和 node image 版本相配,所以另外下載一個舊版執行檔,放在 tmp/:
curl -Lo tmp/kind-v0.30.0 "https://kind.sigs.k8s.io/dl/v0.30.0/kind-$(uname -s | tr '[:upper:]' '[:lower:]')-$(uname -m | sed 's/x86_64/amd64/;s/aarch64/arm64/')"
chmod +x tmp/kind-v0.30.0
tmp/kind-v0.30.0 create cluster --name cpu-lab-old \
--image kindest/node:v1.34.0@sha256:7416a61b42b1662ca6ca89f02028ac133a309a2a30ba309614e8ec94d976dc5a \
--config scripts/kind-2node.yaml --kubeconfig tmp/kubeconfig-old
tmp/kind-v0.30.0 load docker-image cpulab:go124mod cpulab:go125mod --name cpu-lab-old
scripts/e4-e5-weight.sh "scripts/kx old" cpu-lab-old-worker
節錄:
== node runtime
runc version 1.3.0
v2.1.3
kubelet=v1.34.0 allocatable_cpu=10
== E4 cgroup tree
...
-- e4-req1
/kubelet.slice weight=100 max=max 100000
/kubelet.slice/kubelet-kubepods.slice weight=391 max=max 100000
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-burstable.slice weight=203 max=max 100000
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-burstable.slice/kubelet-kubepods-burstable-poda3c4cda5_4e6a_4e4d_b481_780b0582dc30.slice weight=39 max=max 100000
cri-containerd-3dc2421657aaa0a60f8c022adc82906c0a2352363caa0aed347b7f516fc53ca2.scope weight=39 max=max 100000
cri-containerd-f7531c39e1b6774cd6330e7d367c31b2ff75f2aaa3d7ed9ec64fb8d9401003d7.scope weight=1 max=max 100000
app=cri-containerd-3dc2421657aa…
...
-- e5-two-containers
/kubelet.slice weight=100 max=max 100000
/kubelet.slice/kubelet-kubepods.slice weight=391 max=max 100000
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-burstable.slice weight=203 max=max 100000
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-burstable.slice/kubelet-kubepods-burstable-pod87d2c9b4_0278_44a1_9b52_0dd55decf53a.slice weight=43 max=max 100000
cri-containerd-817b5f50511244450dc154c802676b48735eee1e4f6e8876bfd285d1ba70952b.scope weight=4 max=max 100000
cri-containerd-e9d04eb20e481bbc64c2c1dce0caab68d555013219a8243d4a14aad5156333b1.scope weight=1 max=max 100000
cri-containerd-f4e1ff96bd0ba8ba7c529a98d4f66d113ddc22218d4bc57105bd1616c136c0ea.scope weight=39 max=max 100000
big=cri-containerd-f4e1ff96bd0b…
small=cri-containerd-817b5f505112…
...
== E5 pin worker to one CPU
5
-- pods: e4-req100m vs e4-req1 (Pod 層比較),同時 e5-two-containers 也在跑
e4-req100m: 2026-09-26T06:09:25Z usage_cores=0.046 2026-09-26T06:09:35Z usage_cores=0.046 2026-09-26T06:09:45Z usage_cores=0.046
e4-req1: 2026-09-26T06:09:25Z usage_cores=0.450 2026-09-26T06:09:35Z usage_cores=0.446 2026-09-26T06:09:45Z usage_cores=0.451
e5-two-containers/small: 2026-09-26T06:09:25Z usage_cores=0.046 2026-09-26T06:09:35Z usage_cores=0.046 2026-09-26T06:09:45Z usage_cores=0.046
e5-two-containers/big: 2026-09-26T06:09:25Z usage_cores=0.450 2026-09-26T06:09:35Z usage_cores=0.446 2026-09-26T06:09:45Z usage_cores=0.451
-- only the two-container pod (container 層比較)
e5-two-containers/small: 2026-09-26T06:10:15Z usage_cores=0.092 2026-09-26T06:10:25Z usage_cores=0.092 2026-09-26T06:10:35Z usage_cores=0.093
e5-two-containers/big: 2026-09-26T06:10:15Z usage_cores=0.898 2026-09-26T06:10:25Z usage_cores=0.898 2026-09-26T06:10:35Z usage_cores=0.901
== restored cpuset: 0-9
Pod 層的 weight 和新版完全相同(節錄中的 39 與 43),三個 Pod 分到 0.046、0.449 與 0.495,也和新版幾乎一樣。container 層則和 Pod 層取相同的值:request 1 是 39,100m 是 4。同一個 Pod 裡的兩個 container 因此分到 0.092 : 0.899,約 9.3% : 90.7%。兩個叢集只差在 runtime 的換算。
這個實驗的邊界
實驗能清楚看到的,是 kernel 怎麼執行 limit:quota 每個 period 發放一次、所有 thread 共用、用完就全部暫停,以及 request 如何變成各層的 weight。
它看不到的部分:
- 流量的形狀:成群到達是刻意製造的,用來重現「平均三成、仍被 throttle」。真實服務的突發長什麼樣子,要從自己的監控看。
- 其他負載:thread 數的取捨只在這一種流量下測過。長時間接近 limit、GC 比重高的程式,本文沒有測。
- Node 的忙碌程度:第 8 節的鄰居持續用滿 Node,是刻意做出的極端情況;幾個 Pod 的突發偶爾重疊這類較常見的情況,本文沒有測。
- kind 與電腦的特性:kind 的 Node 是 container,cgroup 是巢狀的,root 附近的幾層不是真實主機的樣子;OrbStack 的 VM 和 macOS 共用 CPU,延遲的絕對值會受干擾。
- 沒測的設定:CPU Manager 的 static policy、
cpu.max.burst、自訂 CFS period,以及 JVM。
清理
刪除叢集、image 與這次產生的檔案:
kind delete cluster --name cpu-lab-new --kubeconfig tmp/kubeconfig-new
tmp/kind-v0.30.0 delete cluster --name cpu-lab-old --kubeconfig tmp/kubeconfig-old # 有做第 10 節才需要
docker rmi cpulab:go124mod cpulab:go125mod
最後刪除整個工作目錄。
附錄:cpulab 原始碼
存成 cpulab/main.go。go.mod 不用自己寫,build.sh 會產生兩個版本。
實驗跑完之後,壓測程式 load 依 review 改了兩處:/work 與 /ping 的錯誤分開計算,以及 /work 全部失敗時不再 panic。所以文中的輸出和這份程式的輸出有一點不同:文中 load 一行的 errors= 是兩種請求的錯誤合計,ping_latency_ms 一行沒有 errors=;用這份程式跑,load 一行的 errors= 只算 /work,/ping 的錯誤印在 ping_latency_ms 那一行。
// cpulab 是 CPU throttling 實驗用的小程式,只在一次性的 kind 叢集內執行。
//
// cpulab info 印出 Go 版本、GOMAXPROCS 與所在 cgroup 的 CPU 設定
// cpulab burst [flags] 每隔一段時間讓 N 個 thread 同時做固定的 CPU 工作,量測完成時間
// cpulab spin [flags] 持續佔滿 N 個 thread,定期印出自己 cgroup 的 CPU 用量
// cpulab serve [flags] HTTP 服務:每個請求做固定的 CPU 工作並配置記憶體
// cpulab load [flags] open-loop 壓測:依 Poisson 到達送出請求,印出延遲分布
//
// CPU 工作以 thread 的 CPU 時間計量(鎖定 OS thread 後讀 RUSAGE_THREAD),
// 因此被 throttle 的等待時間不會被算成工作量。
package main
import (
"bufio"
"encoding/json"
"flag"
"fmt"
"io"
"math"
"math/rand"
"net/http"
"os"
"runtime"
"sort"
"strconv"
"strings"
"sync"
"syscall"
"time"
)
const cgroupDir = "/sys/fs/cgroup"
func main() {
if len(os.Args) < 2 {
fmt.Fprintln(os.Stderr, "usage: cpulab info|burst|spin|serve|load [flags]")
os.Exit(2)
}
args := os.Args[2:]
switch os.Args[1] {
case "info":
info()
case "burst":
burst(args)
case "spin":
spin(args)
case "serve":
serve(args)
case "load":
load(args)
default:
fmt.Fprintln(os.Stderr, "unknown command:", os.Args[1])
os.Exit(2)
}
}
func readFile(name string) string {
b, err := os.ReadFile(cgroupDir + "/" + name)
if err != nil {
return "(" + err.Error() + ")"
}
return strings.TrimSpace(string(b))
}
func cpuStat() map[string]int64 {
m := map[string]int64{}
f, err := os.Open(cgroupDir + "/cpu.stat")
if err != nil {
return m
}
defer f.Close()
s := bufio.NewScanner(f)
for s.Scan() {
fields := strings.Fields(s.Text())
if len(fields) == 2 {
v, _ := strconv.ParseInt(fields[1], 10, 64)
m[fields[0]] = v
}
}
return m
}
func info() {
fmt.Printf("go=%s GOMAXPROCS=%d NumCPU=%d GOMAXPROCS_env=%q GODEBUG_env=%q\n",
runtime.Version(), runtime.GOMAXPROCS(0), runtime.NumCPU(),
os.Getenv("GOMAXPROCS"), os.Getenv("GODEBUG"))
fmt.Printf("cpu.max=%q cpu.weight=%q\n", readFile("cpu.max"), readFile("cpu.weight"))
}
// threadCPU 回傳目前 OS thread 已使用的 CPU 時間;呼叫端必須先 LockOSThread。
func threadCPU() time.Duration {
var ru syscall.Rusage
const rusageThread = 1 // RUSAGE_THREAD
if err := syscall.Getrusage(rusageThread, &ru); err != nil {
panic(err)
}
return time.Duration(ru.Utime.Nano() + ru.Stime.Nano())
}
var sink uint64
// spinCPU 在目前的 thread 上消耗 d 的 CPU 時間後返回。
func spinCPU(d time.Duration) {
runtime.LockOSThread()
defer runtime.UnlockOSThread()
start := threadCPU()
x := uint64(1)
for threadCPU()-start < d {
for i := 0; i < 20000; i++ {
x ^= x*31 + uint64(i)
}
}
sink += x
}
func burst(args []string) {
fs := flag.NewFlagSet("burst", flag.ExitOnError)
threads := fs.Int("threads", 8, "同時工作的 thread 數")
work := fs.Duration("work", 20*time.Millisecond, "每個 thread 每輪要做的 CPU 時間")
interval := fs.Duration("interval", time.Second, "每輪開始的間隔")
rounds := fs.Int("rounds", 20, "輪數")
fs.Parse(args)
info()
fmt.Printf("burst threads=%d work=%s interval=%s rounds=%d\n", *threads, *work, *interval, *rounds)
fmt.Println("round wall_ms d_nr_periods d_nr_throttled d_throttled_ms d_usage_ms")
var walls []float64
begin := cpuStat()
t0 := time.Now()
next := t0
for r := 1; r <= *rounds; r++ {
time.Sleep(time.Until(next))
next = next.Add(*interval)
before := cpuStat()
start := time.Now()
var wg sync.WaitGroup
for i := 0; i < *threads; i++ {
wg.Add(1)
go func() { defer wg.Done(); spinCPU(*work) }()
}
wg.Wait()
wall := time.Since(start)
after := cpuStat()
walls = append(walls, ms(wall))
fmt.Printf("%d %.1f %d %d %.1f %.1f\n", r, ms(wall),
after["nr_periods"]-before["nr_periods"],
after["nr_throttled"]-before["nr_throttled"],
float64(after["throttled_usec"]-before["throttled_usec"])/1000,
float64(after["usage_usec"]-before["usage_usec"])/1000)
}
time.Sleep(time.Until(next))
elapsed := time.Since(t0)
end := cpuStat()
sort.Float64s(walls)
fmt.Printf("summary wall_ms p50=%.1f max=%.1f avg_cpu_cores=%.3f nr_periods=%d nr_throttled=%d throttled_ms=%.1f\n",
pct(walls, 50), walls[len(walls)-1],
float64(end["usage_usec"]-begin["usage_usec"])/float64(elapsed.Microseconds()),
end["nr_periods"]-begin["nr_periods"], end["nr_throttled"]-begin["nr_throttled"],
float64(end["throttled_usec"]-begin["throttled_usec"])/1000)
}
func spin(args []string) {
fs := flag.NewFlagSet("spin", flag.ExitOnError)
threads := fs.Int("threads", 1, "持續忙碌的 thread 數")
report := fs.Duration("report", 10*time.Second, "印出用量的間隔")
fs.Parse(args)
info()
for i := 0; i < *threads; i++ {
go func() {
for {
spinCPU(time.Second)
}
}()
}
prev, prevT := cpuStat()["usage_usec"], time.Now()
for range time.Tick(*report) {
cur, now := cpuStat()["usage_usec"], time.Now()
fmt.Printf("%s usage_cores=%.3f\n", now.UTC().Format(time.RFC3339),
float64(cur-prev)/float64(now.Sub(prevT).Microseconds()))
prev, prevT = cur, now
}
}
// liveHeap 讓 GC 每次都有一份含指標的存活資料要標記。
type blob struct {
p *[8]byte
pad [48]byte
}
var liveHeap []*blob
func serve(args []string) {
fs := flag.NewFlagSet("serve", flag.ExitOnError)
addr := fs.String("addr", ":8080", "listen address")
work := fs.Duration("work", 5*time.Millisecond, "每個請求的 CPU 工作量")
alloc := fs.Int("alloc", 1<<20, "每個請求配置後丟棄的位元組數")
live := fs.Int("live-mb", 64, "常駐的存活 heap 大小(MB)")
fs.Parse(args)
info()
for i := 0; i < *live<<20/64; i++ {
liveHeap = append(liveHeap, &blob{p: new([8]byte)})
}
fmt.Printf("serve addr=%s work=%s alloc=%d live_mb=%d\n", *addr, *work, *alloc, *live)
http.HandleFunc("/work", func(w http.ResponseWriter, r *http.Request) {
garbage := make([][]byte, 0, *alloc/4096+1)
for n := 0; n < *alloc; n += 4096 {
garbage = append(garbage, make([]byte, 4096))
}
spinCPU(*work)
sink += uint64(len(garbage))
io.WriteString(w, "ok\n")
})
http.HandleFunc("/ping", func(w http.ResponseWriter, r *http.Request) {
io.WriteString(w, "ok\n")
})
http.HandleFunc("/stats", func(w http.ResponseWriter, r *http.Request) {
var ms runtime.MemStats
runtime.ReadMemStats(&ms)
st := cpuStat()
st["gomaxprocs"] = int64(runtime.GOMAXPROCS(0))
st["num_gc"] = int64(ms.NumGC)
json.NewEncoder(w).Encode(st)
})
if err := http.ListenAndServe(*addr, nil); err != nil {
panic(err)
}
}
func fetchStats(c *http.Client, base string) map[string]int64 {
m := map[string]int64{}
resp, err := c.Get(base + "/stats")
if err != nil {
return m
}
defer resp.Body.Close()
json.NewDecoder(resp.Body).Decode(&m)
return m
}
func load(args []string) {
fs := flag.NewFlagSet("load", flag.ExitOnError)
target := fs.String("target", "http://cpulab:8080", "服務位址")
rps := fs.Float64("rps", 60, "平均每秒 /work 請求數(Poisson 到達)")
clump := fs.Int("clump", 1, "每次到達同時送出的 /work 請求數;平均 rps 不變")
pingRPS := fs.Float64("ping-rps", 0, "另外送出的 /ping 請求數(不做工作),分開統計延遲")
duration := fs.Duration("duration", 60*time.Second, "量測時間")
warmup := fs.Duration("warmup", 10*time.Second, "暖機時間,不計入結果")
label := fs.String("label", "", "結果標籤")
fs.Parse(args)
c := &http.Client{
Timeout: 10 * time.Second,
Transport: &http.Transport{MaxIdleConnsPerHost: 1000, MaxConnsPerHost: 0},
}
// fire 以 Poisson 到達送出請求,每次到達送 n 個;延遲從預定送出時間起算。
fire := func(path string, rate float64, n int, d time.Duration, record func(time.Duration, error)) {
var wg sync.WaitGroup
end := time.Now().Add(d)
next := time.Now()
for next.Before(end) {
time.Sleep(time.Until(next))
sent := next
for i := 0; i < n; i++ {
wg.Add(1)
go func() {
defer wg.Done()
resp, err := c.Get(*target + path)
if err == nil {
io.Copy(io.Discard, resp.Body)
resp.Body.Close()
}
record(time.Since(sent), err)
}()
}
next = next.Add(time.Duration(rand.ExpFloat64() / (rate / float64(n)) * float64(time.Second)))
}
wg.Wait()
}
both := func(d time.Duration, work, ping func(time.Duration, error)) {
var wg sync.WaitGroup
wg.Add(1)
go func() { defer wg.Done(); fire("/work", *rps, *clump, d, work) }()
if *pingRPS > 0 {
wg.Add(1)
go func() { defer wg.Done(); fire("/ping", *pingRPS, 1, d, ping) }()
}
wg.Wait()
}
discard := func(time.Duration, error) {}
both(*warmup, discard, discard)
var mu sync.Mutex
var lat, pingLat []float64
var workErrs, pingErrs int
collect := func(dst *[]float64, errs *int) func(time.Duration, error) {
return func(d time.Duration, err error) {
mu.Lock()
defer mu.Unlock()
if err != nil {
*errs++
return
}
*dst = append(*dst, ms(d))
}
}
before := fetchStats(c, *target)
t0 := time.Now()
both(*duration, collect(&lat, &workErrs), collect(&pingLat, &pingErrs))
elapsed := time.Since(t0)
after := fetchStats(c, *target)
sort.Float64s(lat)
sort.Float64s(pingLat)
d := func(k string) int64 { return after[k] - before[k] }
fmt.Printf("load label=%s gomaxprocs=%d clump=%d requests=%d errors=%d achieved_rps=%.1f\n",
*label, after["gomaxprocs"], *clump, len(lat), workErrs, float64(len(lat))/elapsed.Seconds())
fmt.Printf("work_latency_ms p50=%.1f p90=%.1f p99=%.1f p999=%.1f max=%.1f\n",
pct(lat, 50), pct(lat, 90), pct(lat, 99), pct(lat, 99.9), pct(lat, 100))
if *pingRPS > 0 {
fmt.Printf("ping_latency_ms requests=%d errors=%d p50=%.1f p90=%.1f p99=%.1f max=%.1f\n",
len(pingLat), pingErrs, pct(pingLat, 50), pct(pingLat, 90), pct(pingLat, 99), pct(pingLat, 100))
}
fmt.Printf("server avg_cpu_cores=%.3f nr_periods=%d nr_throttled=%d throttled_ratio=%.3f throttled_ms=%.1f num_gc=%d\n",
float64(d("usage_usec"))/float64(elapsed.Microseconds()),
d("nr_periods"), d("nr_throttled"),
float64(d("nr_throttled"))/math.Max(1, float64(d("nr_periods"))),
float64(d("throttled_usec"))/1000, d("num_gc"))
}
func ms(d time.Duration) float64 { return float64(d.Microseconds()) / 1000 }
func pct(sorted []float64, p float64) float64 {
if len(sorted) == 0 {
return math.NaN()
}
i := int(math.Ceil(p/100*float64(len(sorted)))) - 1
if i < 0 {
i = 0
}
return sorted[i]
}