The service uses only 30% of its CPU on average, yet 30% of its periods are throttled, and p99 latency jumps to hundreds of milliseconds. How do you build this state? And does using fewer threads really help? Instead of guessing, try it on your own machine.
This article uses kind to create a throwaway Kubernetes cluster on your machine, and a small custom program to reproduce a service with low average usage that still gets throttled. Along the way, it reads the container's cgroup files and monitoring metrics directly to confirm what CPU limits and requests become on the Node.
The mechanism, and how to read the numbers, are explained in CPU is only 30% used, so why is it throttled?, which only cites the results. This article has the full steps and output, and explains what each experiment shows and what it does not. It is easier to follow if you have read that article, but you can follow the steps without it.
Before you start
- You need Docker (cgroup v2), kind, kubectl, and Go 1.25 or later to build the lab program.
- The environment used here: an Apple Silicon Mac with Docker provided by OrbStack (Docker Engine 29.4.0, cgroup v2, 10 CPUs, VM kernel 7.0); kind v0.33.0; Go 1.26.0.
- A kind Node is a container and can see every CPU on the machine.
NumCPU=10andGOMAXPROCS=10in the output below will show your own CPU count on your machine. Latency is also affected by other programs on the machine, so it is only good for comparing orders of magnitude and relative differences. - By default, kind switches the current context in
~/.kube/configto the new cluster. This article saves the kubeconfig intmp/in the working directory, and every kubectl command goes through a small wrapper that passes the kubeconfig and context explicitly, so it cannot reach a work cluster or any other cluster by mistake. - Section 7 uses
docker updateto pin the worker Node to one CPU (number 5) for a while, and restores it at the end. That section also describes a mistake made along the way. If your machine has fewer than 6 CPUs, change the5in the script to another number. - Section 8 also uses
docker update, to split the worker and the control-plane onto CPUs 0-7 and 8-9. This assumes 10 CPUs. If you have a different number, adjust both ranges and the neighbor's thread count in the script. - Run every command in the same empty directory. At the end, delete the clusters, the images, and this directory.
- The comments and progress messages in the scripts and the program were translated from the Chinese version of this article. The numbers in the output are unchanged.
1. The lab program and the cluster
Why write a custom program
Throttling makes the same work take longer. If you measure work by elapsed time, such as "loop for 20ms," the paused time is also counted as work, and you cannot see the difference. So the lab program cpulab first locks a goroutine to an OS thread, then reads the CPU time that thread has actually used (RUSAGE_THREAD), and only counts the work as done after it has used the given amount of CPU time. Time paused by throttling is not counted; it only makes the work finish later.
cpulab has five subcommands. The full source code is in the appendix at the end:
| Subcommand | Used in | What it does |
|---|---|---|
info | Sections 2, 4 | Prints the Go version, GOMAXPROCS, and the container's own cpu.max and cpu.weight |
burst | Sections 3, 9 | At a fixed interval, makes N threads do a fixed amount of CPU work at once, then prints each round's completion time and the change in cpu.stat |
serve | Sections 5, 6, 8 | HTTP service: /work does a fixed amount of CPU work and allocates memory, /ping does nothing, /stats reports cpu.stat |
load | Sections 5, 6, 8 | Load generator that sends requests with Poisson arrivals and prints the latency distribution and the server's throttling stats |
spin | Sections 7, 8 | Keeps threads busy and prints how much CPU they actually used every 10 seconds |
Build two images
First save the program from the appendix as cpulab/main.go, then create the Dockerfile and the build script:
mkdir -p cpulab scripts tmp
# Save the program from the appendix as cpulab/main.go
cat > cpulab/Dockerfile <<'EOF'
FROM busybox:1.37
COPY cpulab /cpulab
ENTRYPOINT ["/cpulab"]
EOF
cat > build.sh <<'EOF'
#!/usr/bin/env bash
# Build two images: same code, only the go line in go.mod differs.
# cpulab:go124mod go.mod says go 1.24 (GOMAXPROCS keeps the old default)
# cpulab:go125mod go.mod says go 1.25 (GOMAXPROCS follows the cgroup CPU limit)
set -euo pipefail
here=$(cd "$(dirname "$0")" && pwd)
arch=$(docker info --format '{{.Architecture}}' | sed 's/aarch64/arm64/;s/x86_64/amd64/')
work=$(mktemp -d)
trap 'rm -rf "$work"' EXIT
for v in 1.24 1.25; do
tag=go${v/./}mod
mkdir -p "$work/$tag"
cp "$here/cpulab/main.go" "$here/cpulab/Dockerfile" "$work/$tag/"
printf 'module cpulab\n\ngo %s\n' "$v" > "$work/$tag/go.mod"
(cd "$work/$tag" && CGO_ENABLED=0 GOOS=linux GOARCH=$arch GOTOOLCHAIN=local go build -o cpulab .)
docker build -q -t "cpulab:$tag" "$work/$tag"
done
go version
EOF
chmod +x build.sh
./build.sh
build.sh builds two images. The code is exactly the same; only the go version line in go.mod differs. As section 4 shows, this line decides how many threads a Go program runs at once in a container by default. Both images are built with the local Go toolchain (GOTOOLCHAIN=local), which was Go 1.26.0 for this article.
Create the cluster
The kubectl wrapper scripts/kx takes the cluster as its first argument and passes the rest to kubectl as is.
cat > scripts/kx <<'EOF'
#!/usr/bin/env bash
# kx <old|new> <kubectl args...>: pass the kubeconfig and context explicitly, never touch ~/.kube/config
here=$(cd "$(dirname "$0")/.." && pwd); c=$1; shift
exec kubectl --kubeconfig "$here/tmp/kubeconfig-$c" --context "kind-cpu-lab-$c" "$@"
EOF
chmod +x scripts/kx
cat > scripts/kind-2node.yaml <<'EOF'
kind: Cluster
apiVersion: kind.x-k8s.io/v1alpha4
nodes:
- role: control-plane
- role: worker
EOF
kind create cluster --name cpu-lab-new \
--image kindest/node:v1.34.11@sha256:44e222ee2132dab25ff87301682f89eb82c7880ea3a1bf543bfe9708fd08d67d \
--config scripts/kind-2node.yaml --kubeconfig tmp/kubeconfig-new
kind load docker-image cpulab:go124mod cpulab:go125mod --name cpu-lab-new
scripts/kx new get nodes
docker exec cpu-lab-new-worker sh -c 'runc --version | head -1; stat -fc %T /sys/fs/cgroup'
The cluster has one control-plane and one worker. Every lab Pod uses nodeSelector to run on the worker. The load generator in sections 5, 6, and 8 runs on the control-plane, so it does not compete with the service for CPU.
The node image is pinned by digest because its runc version affects the results: the cpu.weight formula compared in section 7 is decided by the container runtime. The last command should show runc 1.4.3 and cgroup2fs (cgroup v2).
From here on, each section first saves a script with cat > ... <<'EOF', then runs it. The E numbers (E1, E2a, and so on) in the script header comments are experiment log numbers, not the section numbers of this article.
2. Where is the limit written?
Two Pods each have an app container (limit 2) and a sidecar. The sidecar in e1-all-limits has a limit; the one in e1-partial does not. This manifest also creates the cpulab namespace used for the rest of the article:
cat > scripts/e1-limits.yaml <<'EOF'
# E1: where the limit is written. Both Pods have app (limit 2) and sidecar;
# the sidecar in all-limits has a limit, the sidecar in partial does not.
apiVersion: v1
kind: Namespace
metadata:
name: cpulab
---
apiVersion: v1
kind: Pod
metadata:
name: e1-all-limits
namespace: cpulab
spec:
nodeSelector: {kubernetes.io/hostname: WORKER}
containers:
- name: app
image: cpulab:go125mod
command: ["sh", "-c", "/cpulab info; sleep 3600"]
resources: {requests: {cpu: 500m}, limits: {cpu: "2"}}
- name: sidecar
image: cpulab:go125mod
command: ["sh", "-c", "/cpulab info; sleep 3600"]
resources: {requests: {cpu: 100m}, limits: {cpu: 500m}}
---
apiVersion: v1
kind: Pod
metadata:
name: e1-partial
namespace: cpulab
spec:
nodeSelector: {kubernetes.io/hostname: WORKER}
containers:
- name: app
image: cpulab:go125mod
command: ["sh", "-c", "/cpulab info; sleep 3600"]
resources: {requests: {cpu: 500m}, limits: {cpu: "2"}}
- name: sidecar
image: cpulab:go125mod
command: ["sh", "-c", "/cpulab info; sleep 3600"]
resources: {requests: {cpu: 100m}}
EOF
sed s/WORKER/cpu-lab-new-worker/ scripts/e1-limits.yaml | scripts/kx new apply -f -
scripts/kx new -n cpulab wait --for=condition=Ready pod --all --timeout=120s
for p in e1-all-limits e1-partial; do for c in app sidecar; do
echo "$p/$c: $(scripts/kx new -n cpulab logs $p -c $c)"
done; done
Each container runs cpulab info at startup and prints the settings it sees:
e1-all-limits/app: go=go1.26.0 GOMAXPROCS=2 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env=""
cpu.max="200000 100000" cpu.weight="59"
e1-all-limits/sidecar: go=go1.26.0 GOMAXPROCS=2 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env=""
cpu.max="50000 100000" cpu.weight="17"
e1-partial/app: go=go1.26.0 GOMAXPROCS=2 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env=""
cpu.max="200000 100000" cpu.weight="59"
e1-partial/sidecar: go=go1.26.0 GOMAXPROCS=10 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env=""
cpu.max="max 100000" cpu.weight="17"
A container's cpu.max is the converted limit: the first number is the quota and the second is the period, both in microseconds. Limit 2 is at most 200ms every 100ms, and 500m is 50ms. The sidecar without a limit shows max, meaning no limit.
GOMAXPROCS on the same lines is also worth noticing: it is 2 in containers with a limit and 10 in the one without. This image has go 1.25 in go.mod; section 4 looks at this in detail. cpu.weight is converted from the request, and section 7 uses it.
Next, go into the worker Node to look at the Pod level and the levels above it. cgtree.sh walks from the Pod's cgroup up to the root, prints each level's cpu.weight and cpu.max, then lists every container under the Pod:
cat > scripts/cgtree.sh <<'EOF'
#!/bin/sh
# Inside a kind node, print cpu.weight/cpu.max from the cgroup root down to each container of a Pod UID
uid=$(echo "$1" | tr - _)
pod=$(find /sys/fs/cgroup -type d -name "*pod${uid}.slice" | head -1)
[ -n "$pod" ] || { echo "pod cgroup not found"; exit 1; }
d=$pod; chain=""
while [ "$d" != "/sys/fs/cgroup" ]; do chain="$d $chain"; d=$(dirname "$d"); done
for d in $chain; do printf '%-70s weight=%-5s max=%s\n' "${d#/sys/fs/cgroup}" "$(cat $d/cpu.weight)" "$(cat $d/cpu.max)"; done
for c in "$pod"/*/; do [ -f "$c/cpu.weight" ] && printf ' %-68s weight=%-5s max=%s\n' "$(basename $c)" "$(cat $c/cpu.weight)" "$(cat $c/cpu.max)"; done
EOF
docker cp scripts/cgtree.sh cpu-lab-new-worker:/cgtree.sh
for p in e1-all-limits e1-partial; do
echo "== $p"
docker exec cpu-lab-new-worker sh /cgtree.sh "$(scripts/kx new -n cpulab get pod $p -o jsonpath='{.metadata.uid}')"
done
== e1-all-limits
/kubelet.slice weight=100 max=max 100000
/kubelet.slice/kubelet-kubepods.slice weight=391 max=max 100000
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-burstable.slice weight=51 max=max 100000
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-burstable.slice/kubelet-kubepods-burstable-poda562a08c_3829_44ee_9bd1_c48d68973a26.slice weight=24 max=250000 100000
cri-containerd-4de8b08221dfa8a22321d110a6dec2c51f5a851ed29bdb31e39dd6941577c54a.scope weight=17 max=50000 100000
cri-containerd-507b6210e898adff9a780ab969559d4843a410cf8cc3f712859b6fa0c3f04579.scope weight=1 max=max 100000
cri-containerd-906c59210735dabc6bbd403e64c1b67feb9e8916cb64bad37ae999b397744a94.scope weight=59 max=200000 100000
== e1-partial
/kubelet.slice weight=100 max=max 100000
/kubelet.slice/kubelet-kubepods.slice weight=391 max=max 100000
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-burstable.slice weight=51 max=max 100000
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-burstable.slice/kubelet-kubepods-burstable-podee5545ac_13e0_401f_b07b_370dee16585d.slice weight=24 max=max 100000
cri-containerd-02508278f52c40262b86427ba22081210783af2fac243c4280e41d2995e4e430.scope weight=59 max=200000 100000
cri-containerd-a648dabc12e1b9f6db1451ecf4f5d42c2cc616d0ec991d5bbb1521c7229dac0e.scope weight=1 max=max 100000
cri-containerd-e4fa2ea818a16f840234c79b818d51fa79b2d6d93d1eb88691b104a64fbd6763.scope weight=17 max=max 100000
At the Pod level (...pod<uid>.slice), max in e1-all-limits is 250000 100000, the two containers' limits added together (2 + 0.5). e1-partial has a container without a limit, so its Pod level is max. The scope with weight=1 is the Pod's pause container.
There is an extra /kubelet.slice level at the top of the path. This comes from kind's settings (kubelet's cgroupRoot is /kubelet). On a typical Node that uses the systemd cgroup driver, Pods live under /kubepods.slice.
Delete these two Pods when you are done, so the later experiments run on a quiet Node:
scripts/kx new -n cpulab delete pod e1-all-limits e1-partial --wait=true
What this shows: in the end, a limit becomes the cpu.max of two cgroup levels, the container and the Pod, and the Pod level is only set when every container has a limit. What it does not show: how the kernel enforces this cap. That is the next section.
3. The same job under four limits
Let cpulab burst do the same job once per second: 8 threads each compute 20ms at once, 160ms of CPU in total, and stay idle for the rest of the second, for an average of about 0.16 CPU. Run 20 rounds each with no limit, limit 2, limit 1, and limit 500m, in that order. The last run also uses limit 1, but with 2 threads computing 80ms each, so the total work stays the same:
cat > scripts/e2a-burst.sh <<'EOF'
#!/usr/bin/env bash
# E2a: one burst per second, same amount of work (160ms of CPU in total); completion time and throttling under different limits and thread counts.
# Usage: e2a-burst.sh <kx command> <worker node>; runs one at a time so the runs do not interfere.
set -euo pipefail
KX=$1; NODE=$2
run() { # name limit threads work gomaxprocs
local name=$1 limit=$2 threads=$3 work=$4 procs=$5 res='{"requests":{"cpu":"100m"}}'
[ "$limit" != none ] && res="{\"requests\":{\"cpu\":\"100m\"},\"limits\":{\"cpu\":\"$limit\"}}"
$KX -n cpulab run "$name" --image=cpulab:go125mod --restart=Never \
--overrides="{\"spec\":{\"nodeSelector\":{\"kubernetes.io/hostname\":\"$NODE\"},\"containers\":[{\"name\":\"c\",\"image\":\"cpulab:go125mod\",\"args\":[\"burst\",\"-threads=$threads\",\"-work=$work\",\"-interval=1s\",\"-rounds=20\"],\"resources\":$res,\"env\":[{\"name\":\"GOMAXPROCS\",\"value\":\"$procs\"}]}]}}" >/dev/null
$KX -n cpulab wait --for=jsonpath='{.status.phase}'=Succeeded pod/$name --timeout=120s >/dev/null
echo "===== $name limit=$limit threads=$threads work=$work GOMAXPROCS=$procs"
$KX -n cpulab logs $name
$KX -n cpulab delete pod $name --wait=false >/dev/null
}
run e2a-nolimit none 8 20ms 8
run e2a-limit2 2 8 20ms 8
run e2a-limit1 1 8 20ms 8
run e2a-limit500m 500m 8 20ms 8
run e2a-limit1-2thr 1 2 80ms 2
EOF
chmod +x scripts/e2a-burst.sh
scripts/e2a-burst.sh "scripts/kx new" cpu-lab-new-worker
The script uses the GOMAXPROCS environment variable to set Go's thread count to the number of work threads. This image has go 1.25 in go.mod. Without this setting, GOMAXPROCS would change with the limit (section 4), and the runs would differ in more than just the limit.
The first five rounds and the summary for limit 1:
===== e2a-limit1 limit=1 threads=8 work=20ms GOMAXPROCS=8
go=go1.26.0 GOMAXPROCS=8 NumCPU=10 GOMAXPROCS_env="8" GODEBUG_env=""
cpu.max="100000 100000" cpu.weight="17"
burst threads=8 work=20ms interval=1s rounds=20
round wall_ms d_nr_periods d_nr_throttled d_throttled_ms d_usage_ms
1 83.2 1 1 421.9 164.1
2 82.4 1 1 390.2 164.6
3 85.2 1 1 361.7 166.0
4 78.8 1 1 325.3 165.2
5 117.0 1 1 379.5 166.4
...
summary wall_ms p50=84.4 max=118.1 avg_cpu_cores=0.165 nr_periods=60 nr_throttled=20 throttled_ms=7052.5
The columns in each round are: the completion time, then how much cpu.stat changed during the round (nr_periods, nr_throttled, throttled_usec, usage_usec, with the last two converted to milliseconds). The last line summarizes the 20 rounds, and avg_cpu_cores is the average usage over the whole run.
The headers and summaries of all five runs:
===== e2a-nolimit limit=none threads=8 work=20ms GOMAXPROCS=8
summary wall_ms p50=34.3 max=38.4 avg_cpu_cores=0.165 nr_periods=0 nr_throttled=0 throttled_ms=0.0
===== e2a-limit2 limit=2 threads=8 work=20ms GOMAXPROCS=8
summary wall_ms p50=34.6 max=38.1 avg_cpu_cores=0.166 nr_periods=40 nr_throttled=0 throttled_ms=0.0
===== e2a-limit1 limit=1 threads=8 work=20ms GOMAXPROCS=8
summary wall_ms p50=84.4 max=118.1 avg_cpu_cores=0.165 nr_periods=60 nr_throttled=20 throttled_ms=7052.5
===== e2a-limit500m limit=500m threads=8 work=20ms GOMAXPROCS=8
summary wall_ms p50=260.4 max=263.7 avg_cpu_cores=0.166 nr_periods=100 nr_throttled=60 throttled_ms=23981.0
===== e2a-limit1-2thr limit=1 threads=2 work=80ms GOMAXPROCS=2
summary wall_ms p50=132.3 max=136.6 avg_cpu_cores=0.162 nr_periods=77 nr_throttled=17 throttled_ms=1301.6
As a table:
| Setting | Average usage | Completion time p50 / max | nr_periods / nr_throttled over 20 rounds |
|---|---|---|---|
| No limit | 0.165 | 34.3 / 38.4ms | 0 / 0 |
| limit 2 | 0.166 | 34.6 / 38.1ms | 40 / 0 |
| limit 1 | 0.165 | 84.4 / 118.1ms | 60 / 20 |
| limit 500m | 0.166 | 260.4 / 263.7ms | 100 / 60 |
| limit 1, 2 threads × 80ms | 0.162 | 132.3 / 136.6ms | 77 / 17 |
A few things are worth comparing:
- The averages are the same, but the completion times differ a lot. Limit 2's budget can hold 160ms, so it is almost the same as no limit. Limit 1 is throttled once every round, and 500m three times every round.
- Limit 1's completion times fall into two groups, about 80ms and about 116ms. Where in the period the job starts decides how long it waits for the next budget.
d_throttled_msis longer than the completion time. For limit 1 it is about 220 to 420ms per round, but a whole round takes only 80 to 118ms. It adds up the time each of the 8 threads was paused; it is not elapsed time.nr_periodsonly counts busy periods. Limit 2 ran for 20 seconds but only accumulated 40 periods, 2 per burst on average; idle time is not counted. Thed_nr_periodsread in each round is 0, which means these two periods were recorded after the burst had finished.- Fewer threads means less throttling, but not faster work. 2 threads computing 80ms each need 80ms even with no limit at all. Under limit 1, there were only 17 throttles, but the median completion time was 132ms, longer than the 84ms with 8 threads. Section 5 shows the same trade-off in an HTTP service.
What this shows: the quota is a budget shared by all threads in each period, and even a program with low average usage can be throttled every time. What it does not show: whether a real service's requests bunch up like this. That is section 5.
4. GOMAXPROCS follows go.mod
The budget is shared by all threads, so the more threads there are, the sooner it runs out. The number of threads a Go program uses to run goroutines at the same time is set by GOMAXPROCS. It used to default to the Node's CPU count, whatever the limit. Go 1.25 changed this default: GOMAXPROCS now follows the cgroup CPU limit, rounded up and at least 2, and is updated periodically. It does not look at the request. Go 1.25 release notes · How GOMAXPROCS is computed
There is an easy-to-miss condition: this behavior depends on the go version in go.mod, not on the Go version used to build the program. If go.mod says a version below 1.25, the old behavior stays. GODEBUG default table
This section prints GOMAXPROCS under different limits, using the two images with the same code and different go.mod version lines. The last two Pods override the default with environment variables:
cat > scripts/e3-gomaxprocs.sh <<'EOF'
#!/usr/bin/env bash
# E3: same code, only the go line in go.mod differs; GOMAXPROCS under different CPU limits.
# Usage: e3-gomaxprocs.sh <kx command> <worker node>
set -euo pipefail
KX=$1; NODE=$2
run() { # name image limit env...
local name=$1 image=$2 limit=$3; shift 3
local res='{}' envs='[]'
[ "$limit" != none ] && res="{\"requests\":{\"cpu\":\"100m\"},\"limits\":{\"cpu\":\"$limit\"}}"
[ $# -gt 0 ] && envs="[$(for e in "$@"; do printf '{"name":"%s","value":"%s"},' "${e%%=*}" "${e#*=}"; done | sed 's/,$//')]"
$KX -n cpulab run "$name" --image="cpulab:$image" --restart=Never --command \
--overrides="{\"spec\":{\"nodeSelector\":{\"kubernetes.io/hostname\":\"$NODE\"},\"containers\":[{\"name\":\"c\",\"image\":\"cpulab:$image\",\"command\":[\"/cpulab\",\"info\"],\"resources\":$res,\"env\":$envs}]}}" \
-- /cpulab info >/dev/null
}
for img in go124mod go125mod; do
for lim in none 500m 1 1500m 2 2500m; do run "e3-$img-$(echo $lim | tr -d m)" $img $lim; done
done
run e3-go125mod-2-env go125mod 2 GOMAXPROCS=6
run e3-go125mod-2-godebug go125mod 2 GODEBUG=containermaxprocs=0
sleep 15
for p in $($KX -n cpulab get pods -o name | grep e3- | sort -V); do
printf '%-28s limit=%-6s %s\n' "${p#pod/}" "$($KX -n cpulab get $p -o jsonpath='{.spec.containers[0].resources.limits.cpu}')" "$($KX -n cpulab logs $p | tr '\n' ' ')"
done
EOF
chmod +x scripts/e3-gomaxprocs.sh
scripts/e3-gomaxprocs.sh "scripts/kx new" cpu-lab-new-worker
e3-go124mod-1 limit=1 go=go1.26.0 GOMAXPROCS=10 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env="" cpu.max="100000 100000" cpu.weight="17"
e3-go124mod-2 limit=2 go=go1.26.0 GOMAXPROCS=10 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env="" cpu.max="200000 100000" cpu.weight="17"
e3-go124mod-500 limit=500m go=go1.26.0 GOMAXPROCS=10 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env="" cpu.max="50000 100000" cpu.weight="17"
e3-go124mod-1500 limit=1500m go=go1.26.0 GOMAXPROCS=10 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env="" cpu.max="150000 100000" cpu.weight="17"
e3-go124mod-2500 limit=2500m go=go1.26.0 GOMAXPROCS=10 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env="" cpu.max="250000 100000" cpu.weight="17"
e3-go124mod-none limit= go=go1.26.0 GOMAXPROCS=10 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env="" cpu.max="max 100000" cpu.weight="1"
e3-go125mod-1 limit=1 go=go1.26.0 GOMAXPROCS=2 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env="" cpu.max="100000 100000" cpu.weight="17"
e3-go125mod-2 limit=2 go=go1.26.0 GOMAXPROCS=2 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env="" cpu.max="200000 100000" cpu.weight="17"
e3-go125mod-2-env limit=2 go=go1.26.0 GOMAXPROCS=6 NumCPU=10 GOMAXPROCS_env="6" GODEBUG_env="" cpu.max="200000 100000" cpu.weight="17"
e3-go125mod-2-godebug limit=2 go=go1.26.0 GOMAXPROCS=10 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env="containermaxprocs=0" cpu.max="200000 100000" cpu.weight="17"
e3-go125mod-500 limit=500m go=go1.26.0 GOMAXPROCS=2 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env="" cpu.max="50000 100000" cpu.weight="17"
e3-go125mod-1500 limit=1500m go=go1.26.0 GOMAXPROCS=2 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env="" cpu.max="150000 100000" cpu.weight="17"
e3-go125mod-2500 limit=2500m go=go1.26.0 GOMAXPROCS=3 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env="" cpu.max="250000 100000" cpu.weight="17"
e3-go125mod-none limit= go=go1.26.0 GOMAXPROCS=10 NumCPU=10 GOMAXPROCS_env="" GODEBUG_env="" cpu.max="max 100000" cpu.weight="1"
- With
go 1.24in go.mod, GOMAXPROCS is the Node's CPU count, 10, whatever the limit. - With
go 1.25in go.mod, GOMAXPROCS is the limit rounded up, and at least 2: 500m, 1, 1.5, and 2 all give 2, and 2.5 gives 3. With no limit, it is still 10. - Setting
GOMAXPROCS=6gives 6, andGODEBUG=containermaxprocs=0goes back to the old behavior.
Both images were built with Go 1.26. The go.mod version line decides the behavior, not the compiler version, so upgrading only the toolchain is not enough. Check the version in go.mod, or set the GOMAXPROCS environment variable directly. The two Pods without a limit also have no request, so they are BestEffort and their cpu.weight is 1. These Pods exit after printing and use no CPU, so you can leave them until the cluster is deleted at the end.
Other runtimes have similar issues. The JVM sets the default size of its GC threads and thread pools based on the number of available processors, and inside a container that number is computed from the quota. Since JDK 19, it no longer considers the shares converted from the request. JDK-8281181 The details differ between versions, so check how many CPUs it actually sees before tuning. This article did not test the JVM.
What this shows: upgrading only the toolchain does not make GOMAXPROCS follow the limit. What it does not show: whether latency improves once GOMAXPROCS follows the limit. The next section tests that.
5. Reproduce the service from the start
Experiment design
cpulab serve is a small HTTP service with request 500m:
- Each
/workrequest allocates 1MB of memory and throws it away, and does 5ms of CPU work. The program keeps about 64MB of live data, so the GC always has something to mark. With HTTP and GC overhead, each request uses about 6ms of CPU on average. /pingdoes nothing. It shows how much unrelated small requests slow down while the service is busy.
The load generator cpulab load runs on the control-plane and sends requests with Poisson arrivals: on average 100 /work requests per second, plus 20 /ping requests per second. It sends the next request without waiting for a response (open-loop), and measures latency from the scheduled send time, so when the service slows down, the time spent in the queue counts as latency. After a 10-second warmup, it measures for 60 seconds, and reads the service's /stats before and after to compute how much cpu.stat changed.
-clump controls how requests arrive: 1 means one at a time, and 50 means 50 arrive together each time. The average rate stays the same, so with 50 there are two groups per second on average. Each arrival pattern runs with three settings:
| Name | limit | image | GOMAXPROCS |
|---|---|---|---|
nolimit-g10 | None | go.mod 1.24 | 10 |
limit2-g10 | 2 | go.mod 1.24 | 10 |
limit2-g2 | 2 | go.mod 1.25 | 2 |
cat > scripts/e2b-serve.sh <<'EOF'
#!/usr/bin/env bash
# E2b: HTTP service (each request does CPU work + allocates memory; a live heap gives the GC work),
# requests sent with Poisson arrivals from a load Pod on the control-plane; compares limit and GOMAXPROCS.
# Usage: e2b-serve.sh <kx command> <worker node> <cp node> <rps> <name:image:limit>...
# If the script fails or is interrupted, the trap deletes Pods still running, so a leftover server is not picked by the Service and mixed into the next measurement.
set -euo pipefail
KX=$1; NODE=$2; CP=$3; RPS=$4; shift 4
active=() # Pods still running
cleanup() { [ ${#active[@]} -eq 0 ] || $KX -n cpulab delete pod "${active[@]}" --ignore-not-found --wait=false >/dev/null 2>&1 || true; }
trap cleanup EXIT
$KX -n cpulab get svc cpulab-server >/dev/null 2>&1 || \
$KX -n cpulab create service clusterip cpulab-server --tcp=8080:8080 >/dev/null
$KX -n cpulab patch svc cpulab-server -p '{"spec":{"selector":{"app":"cpulab-server"}}}' >/dev/null
for spec in "$@"; do
IFS=: read -r name image limit <<<"$spec"
res='{"requests":{"cpu":"500m"}}'
[ "$limit" != none ] && res="{\"requests\":{\"cpu\":\"500m\"},\"limits\":{\"cpu\":\"$limit\"}}"
LIVEMB=${LIVEMB:-64}; $KX -n cpulab run "$name" --image=cpulab:$image --restart=Never --labels=app=cpulab-server \
--overrides="{\"spec\":{\"nodeSelector\":{\"kubernetes.io/hostname\":\"$NODE\"},\"containers\":[{\"name\":\"c\",\"image\":\"cpulab:$image\",\"args\":[\"serve\",\"-work=5ms\",\"-alloc=1048576\",\"-live-mb=$LIVEMB\"],\"resources\":$res,\"ports\":[{\"containerPort\":8080}],\"readinessProbe\":{\"httpGet\":{\"path\":\"/stats\",\"port\":8080}}}]}}" >/dev/null
active=("$name")
$KX -n cpulab wait --for=condition=Ready pod/$name --timeout=120s >/dev/null
$KX -n cpulab run "load-$name" --image=cpulab:go125mod --restart=Never \
--overrides="{\"spec\":{\"nodeSelector\":{\"kubernetes.io/hostname\":\"$CP\"},\"tolerations\":[{\"operator\":\"Exists\"}],\"containers\":[{\"name\":\"c\",\"image\":\"cpulab:go125mod\",\"args\":[\"load\",\"-target=http://cpulab-server:8080\",\"-rps=$RPS\",\"-clump=${CLUMP:-1}\",\"-ping-rps=${PINGRPS:-0}\",\"-duration=60s\",\"-warmup=10s\",\"-label=$name\"]}]}}" >/dev/null
active+=("load-$name")
$KX -n cpulab wait --for=jsonpath='{.status.phase}'=Succeeded pod/load-$name --timeout=180s >/dev/null
echo "===== $name image=$image limit=$limit rps=$RPS clump=${CLUMP:-1} ping_rps=${PINGRPS:-0} live_mb=$LIVEMB"
$KX -n cpulab logs load-$name
$KX -n cpulab delete pod $name load-$name --wait=true >/dev/null
active=()
done
EOF
chmod +x scripts/e2b-serve.sh
for c in 1:steady 50:clumped; do
CLUMP=${c%%:*} PINGRPS=20 scripts/e2b-serve.sh "scripts/kx new" cpu-lab-new-worker cpu-lab-new-control-plane 100 \
e2b-${c#*:}-nolimit-g10:go124mod:none e2b-${c#*:}-limit2-g10:go124mod:2 e2b-${c#*:}-limit2-g2:go125mod:2
done
The six runs go one after another, about a minute and a half each.
One at a time: no throttling at all
load label=e2b-steady-nolimit-g10 gomaxprocs=10 clump=1 requests=5920 errors=0 achieved_rps=98.7
work_latency_ms p50=7.0 p90=8.9 p99=13.2 p999=18.8 max=23.3
ping_latency_ms requests=1182 p50=1.3 p90=3.1 p99=4.9 max=6.9
server avg_cpu_cores=0.600 nr_periods=0 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0 num_gc=75
load label=e2b-steady-limit2-g10 gomaxprocs=10 clump=1 requests=6149 errors=0 achieved_rps=102.5
work_latency_ms p50=7.0 p90=9.0 p99=13.4 p999=18.9 max=21.9
ping_latency_ms requests=1168 p50=1.3 p90=3.0 p99=4.6 max=12.2
server avg_cpu_cores=0.624 nr_periods=600 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0 num_gc=78
load label=e2b-steady-limit2-g2 gomaxprocs=2 clump=1 requests=6026 errors=0 achieved_rps=100.4
work_latency_ms p50=7.2 p90=9.7 p99=15.7 p999=21.9 max=26.8
ping_latency_ms requests=1236 p50=1.6 p90=3.8 p99=8.7 max=17.8
server avg_cpu_cores=0.604 nr_periods=600 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0 num_gc=78
All three settings average about 0.6 CPU, nr_throttled is 0 in every case, and p99 is between 13 and 16ms. The two limit 2 runs each accumulate 600 periods, which means something ran in every period of the 60 seconds, but each period needed far less than 200ms of CPU, so the budget never ran out. The run without a limit has no quota, so nr_periods does not grow.
50 at a time
load label=e2b-clumped-nolimit-g10 gomaxprocs=10 clump=50 requests=5800 errors=0 achieved_rps=96.6
work_latency_ms p50=33.1 p90=52.6 p99=74.4 p999=90.0 max=94.6
ping_latency_ms requests=1196 p50=2.2 p90=4.6 p99=13.0 max=43.0
server avg_cpu_cores=0.607 nr_periods=0 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0 num_gc=58
load label=e2b-clumped-limit2-g10 gomaxprocs=10 clump=50 requests=6100 errors=0 achieved_rps=101.7
work_latency_ms p50=70.8 p90=187.5 p99=370.1 p999=658.1 max=665.0
ping_latency_ms requests=1200 p50=2.5 p90=23.9 p99=99.4 max=289.1
server avg_cpu_cores=0.639 nr_periods=490 nr_throttled=147 throttled_ratio=0.300 throttled_ms=53589.5 num_gc=60
load label=e2b-clumped-limit2-g2 gomaxprocs=2 clump=50 requests=6250 errors=0 achieved_rps=104.2
work_latency_ms p50=110.7 p90=239.0 p99=447.1 p999=560.0 max=589.1
ping_latency_ms requests=1251 p50=3.0 p90=113.8 p99=335.8 max=434.7
server avg_cpu_cores=0.643 nr_periods=487 nr_throttled=22 throttled_ratio=0.045 throttled_ms=36.2 num_gc=69
The average usage barely changes, about 0.61 to 0.64 CPU, or 30% of limit 2. But in the run with limit 2 and GOMAXPROCS 10, 30% of the periods are throttled, p99 goes from 74ms without a limit to 370ms, and /ping p99 goes from 13ms to 99ms.
A group of 50 requests needs about 300ms of CPU, more than one period's 200ms budget can hold. With 10 threads running at once, the budget is used up in about 20ms, and the whole service pauses until the period ends. /ping requests that arrive during this time can only wait.
Lowering GOMAXPROCS to 2 cuts the share of throttled periods from 30% to 4.5%, and throttled_ms from 53 seconds to 36 milliseconds. The metrics look much better. But p99 becomes 447ms, and /ping p99 becomes 336ms, both worse than with GOMAXPROCS 10. 2 threads can use at most 200ms in one period, so they rarely go over the budget. The same 300ms of work can only be done two requests at a time, and /ping waits longer behind it. The work done in each 100ms does not grow; "pausing" just turns into "queueing."
The 53 seconds of throttled_ms is the sum of the time each CPU was paused during the 60-second measurement. It does not mean the service stopped for 53 seconds.
Run it again
Latency is easily affected by other programs on the machine, so the three clumped runs were repeated:
CLUMP=50 PINGRPS=20 scripts/e2b-serve.sh "scripts/kx new" cpu-lab-new-worker cpu-lab-new-control-plane 100 \
e2b-clumped-nolimit-g10:go124mod:none e2b-clumped-limit2-g10:go124mod:2 e2b-clumped-limit2-g2:go125mod:2
load label=e2b-clumped-nolimit-g10 gomaxprocs=10 clump=50 requests=6750 errors=0 achieved_rps=112.5
work_latency_ms p50=34.0 p90=56.3 p99=86.3 p999=109.5 max=115.5
ping_latency_ms requests=1228 p50=2.1 p90=4.7 p99=18.7 max=56.1
server avg_cpu_cores=0.702 nr_periods=0 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0 num_gc=67
load label=e2b-clumped-limit2-g10 gomaxprocs=10 clump=50 requests=7000 errors=0 achieved_rps=116.8
work_latency_ms p50=63.0 p90=199.1 p99=359.9 p999=608.6 max=622.2
ping_latency_ms requests=1237 p50=2.6 p90=32.4 p99=110.7 max=326.7
server avg_cpu_cores=0.730 nr_periods=513 nr_throttled=160 throttled_ratio=0.312 throttled_ms=58780.4 num_gc=67
load label=e2b-clumped-limit2-g2 gomaxprocs=2 clump=50 requests=7050 errors=0 achieved_rps=117.2
work_latency_ms p50=114.8 p90=262.0 p99=511.2 p999=627.3 max=645.0
ping_latency_ms requests=1223 p50=3.0 p90=131.7 p99=341.2 max=542.6
server avg_cpu_cores=0.720 nr_periods=499 nr_throttled=22 throttled_ratio=0.044 throttled_ms=40.4 num_gc=78
| Setting | Average CPU | Throttled periods | /work p50/p99 | /ping p99 |
|---|---|---|---|---|
| No limit, GOMAXPROCS 10 | 0.61 / 0.70 | — | 33/74, 34/86ms | 13 / 19ms |
| limit 2, GOMAXPROCS 10 | 0.64 / 0.73 | 30.0% / 31.2% | 71/370, 63/360ms | 99 / 111ms |
| limit 2, GOMAXPROCS 2 | 0.64 / 0.72 | 4.5% / 4.4% | 111/447, 115/511ms | 336 / 341ms |
Each cell lists the first run, then the second; the /work column shows p50/p99 for each run. In the second run, the actual rate was a bit higher, about 112 to 117 requests per second, but the relative differences between the three settings are the same.
What this shows: with the same average usage, requests that arrive in groups get throttled. Lowering GOMAXPROCS lowers the throttling metrics but makes latency worse. What it does not show: what happens with other load shapes. For example, the thread count may matter differently for programs that stay near the limit for a long time, or programs where GC takes a larger share. This article did not test those.
6. Add replicas, or raise the limit?
When the main article discusses fixes, it compares two common approaches: adding replicas, and keeping the request while raising only the limit. This section compares five settings under the same clumped traffic as section 5 (on average 100 /work per second, 50 arriving together each time, plus 20 /ping per second). Every server uses the go.mod 1.24 image (GOMAXPROCS 10), and every Pod's request is 500m:
| Name | Replicas | Limit per Pod | Total limit |
|---|---|---|---|
e7-1x-nolimit | 1 | None | — |
e7-1x-limit2 | 1 | 2 | 2 |
e7-2x-limit1 | 2 | 1 | 2 |
e7-2x-limit2 | 2 | 2 | 4 |
e7-1x-limit4 | 1 | 4 | 4 |
The load generator reads /stats through the Service, so with two replicas it only reaches one of the Pods. This script therefore uses kubectl exec to read each server Pod's cpu.stat before and after the load test. The window includes the 10-second warmup, and the average usage is computed over the whole elapsed time, so it is not fully comparable with the server line in section 5.
cat > scripts/e7-scale.sh <<'EOF'
#!/usr/bin/env bash
# E7: same clumped traffic; compares "add replicas" with "raise the limit". Every server uses go124mod (GOMAXPROCS 10) and request 500m.
# Usage: e7-scale.sh <kx command> <worker node> <cp node> <rps> <name:replicas:limit>...
# The load generator reads /stats through the Service, which reaches only one Pod when there are several replicas, so use kubectl exec
# to read each server Pod's cpu.stat before and after the load test. The window includes the 10-second warmup; average usage uses the whole elapsed time.
# If the script fails or is interrupted, the trap deletes Pods still running, so a leftover server is not picked by the Service and mixed into the next measurement.
set -euo pipefail
KX=$1; NODE=$2; CP=$3; RPS=$4; shift 4
active=() # Pods still running
cleanup() { [ ${#active[@]} -eq 0 ] || $KX -n cpulab delete pod "${active[@]}" --ignore-not-found --wait=false >/dev/null 2>&1 || true; }
trap cleanup EXIT
$KX get ns cpulab >/dev/null 2>&1 || $KX create ns cpulab >/dev/null
$KX -n cpulab get svc cpulab-server >/dev/null 2>&1 || \
$KX -n cpulab create service clusterip cpulab-server --tcp=8080:8080 >/dev/null
$KX -n cpulab patch svc cpulab-server -p '{"spec":{"selector":{"app":"cpulab-server"}}}' >/dev/null
cpustat() { $KX -n cpulab exec "$1" -- cat /sys/fs/cgroup/cpu.stat | awk '{printf "%s=%s ", $1, $2}'; }
delta() { # before after elapsed_s -> usage and throttling over this window
printf '%s\n%s\n' "$1" "$2" | awk -v t="$3" '
{ for (i = 1; i <= NF; i++) { split($i, kv, "="); v[NR, kv[1]] = kv[2] } }
END {
u = v[2, "usage_usec"] - v[1, "usage_usec"]; p = v[2, "nr_periods"] - v[1, "nr_periods"]
n = v[2, "nr_throttled"] - v[1, "nr_throttled"]; s = v[2, "throttled_usec"] - v[1, "throttled_usec"]
printf "avg_cpu_cores=%.3f nr_periods=%d nr_throttled=%d throttled_ratio=%.3f throttled_ms=%.1f\n", u / t / 1e6, p, n, (p ? n / p : 0), s / 1000
}'
}
for spec in "$@"; do
IFS=: read -r name replicas limit <<<"$spec"
res='{"requests":{"cpu":"500m"}}'
[ "$limit" != none ] && res="{\"requests\":{\"cpu\":\"500m\"},\"limits\":{\"cpu\":\"$limit\"}}"
pods=()
for i in $(seq 1 "$replicas"); do
$KX -n cpulab run "$name-$i" --image=cpulab:go124mod --restart=Never --labels=app=cpulab-server \
--overrides="{\"spec\":{\"nodeSelector\":{\"kubernetes.io/hostname\":\"$NODE\"},\"containers\":[{\"name\":\"c\",\"image\":\"cpulab:go124mod\",\"args\":[\"serve\",\"-work=5ms\",\"-alloc=1048576\",\"-live-mb=64\"],\"resources\":$res,\"ports\":[{\"containerPort\":8080}],\"readinessProbe\":{\"httpGet\":{\"path\":\"/stats\",\"port\":8080}}}]}}" >/dev/null
pods+=("$name-$i"); active=("${pods[@]}")
done
$KX -n cpulab wait --for=condition=Ready pod "${pods[@]}" --timeout=120s >/dev/null
before=()
for p in "${pods[@]}"; do before+=("$(cpustat "$p")"); done
t0=$(date +%s)
$KX -n cpulab run "load-$name" --image=cpulab:go125mod --restart=Never \
--overrides="{\"spec\":{\"nodeSelector\":{\"kubernetes.io/hostname\":\"$CP\"},\"tolerations\":[{\"operator\":\"Exists\"}],\"containers\":[{\"name\":\"c\",\"image\":\"cpulab:go125mod\",\"args\":[\"load\",\"-target=http://cpulab-server:8080\",\"-rps=$RPS\",\"-clump=${CLUMP:-1}\",\"-ping-rps=${PINGRPS:-0}\",\"-duration=60s\",\"-warmup=10s\",\"-label=$name\"]}]}}" >/dev/null
active+=("load-$name")
$KX -n cpulab wait --for=jsonpath='{.status.phase}'=Succeeded pod/load-$name --timeout=180s >/dev/null
t1=$(date +%s)
echo "===== $name replicas=$replicas limit=$limit request=500m rps=$RPS clump=${CLUMP:-1} ping_rps=${PINGRPS:-0} window_s=$((t1 - t0))"
$KX -n cpulab logs load-$name | grep -E '^(load|work_latency_ms|ping_latency_ms) '
i=0
for p in "${pods[@]}"; do
echo "pod $p $(delta "${before[$i]}" "$(cpustat "$p")" $((t1 - t0)))"
i=$((i + 1))
done
$KX -n cpulab delete pod "${pods[@]}" load-$name --wait=true >/dev/null
active=()
done
EOF
chmod +x scripts/e7-scale.sh
for round in 1 2; do
CLUMP=50 PINGRPS=20 scripts/e7-scale.sh "scripts/kx new" cpu-lab-new-worker cpu-lab-new-control-plane 100 \
e7-1x-nolimit:1:none e7-1x-limit2:1:2 e7-2x-limit1:2:1 e7-2x-limit2:2:2 e7-1x-limit4:1:4
done
The five runs go one after another, about a minute and a half each, and the whole set was run twice. Output of the first run:
===== e7-1x-nolimit replicas=1 limit=none request=500m rps=100 clump=50 ping_rps=20 window_s=72
load label=e7-1x-nolimit gomaxprocs=10 clump=50 requests=5600 errors=0 achieved_rps=93.4
work_latency_ms p50=34.3 p90=55.4 p99=85.7 p999=98.0 max=101.7
ping_latency_ms requests=1213 p50=2.2 p90=4.7 p99=12.4 max=46.9
pod e7-1x-nolimit-1 avg_cpu_cores=0.604 nr_periods=0 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
===== e7-1x-limit2 replicas=1 limit=2 request=500m rps=100 clump=50 ping_rps=20 window_s=73
load label=e7-1x-limit2 gomaxprocs=10 clump=50 requests=5850 errors=0 achieved_rps=97.4
work_latency_ms p50=46.1 p90=136.1 p99=263.9 p999=379.4 max=431.0
ping_latency_ms requests=1207 p50=2.3 p90=15.4 p99=86.1 max=216.0
pod e7-1x-limit2-1 avg_cpu_cores=0.565 nr_periods=563 nr_throttled=133 throttled_ratio=0.236 throttled_ms=34491.9
===== e7-2x-limit1 replicas=2 limit=1 request=500m rps=100 clump=50 ping_rps=20 window_s=72
load label=e7-2x-limit1 gomaxprocs=10 clump=50 requests=6050 errors=0 achieved_rps=100.6
work_latency_ms p50=127.0 p90=356.6 p99=602.5 p999=961.0 max=1525.5
ping_latency_ms requests=1215 p50=3.1 p90=65.9 p99=224.6 max=702.8
pod e7-2x-limit1-1 avg_cpu_cores=0.150 nr_periods=229 nr_throttled=67 throttled_ratio=0.293 throttled_ms=24624.2
pod e7-2x-limit1-2 avg_cpu_cores=0.509 nr_periods=615 nr_throttled=325 throttled_ratio=0.528 throttled_ms=101140.6
===== e7-2x-limit2 replicas=2 limit=2 request=500m rps=100 clump=50 ping_rps=20 window_s=71
load label=e7-2x-limit2 gomaxprocs=10 clump=50 requests=6250 errors=0 achieved_rps=104.4
work_latency_ms p50=33.5 p90=82.8 p99=140.1 p999=202.9 max=220.3
ping_latency_ms requests=1226 p50=2.1 p90=5.1 p99=40.7 max=61.9
pod e7-2x-limit2-1 avg_cpu_cores=0.383 nr_periods=495 nr_throttled=51 throttled_ratio=0.103 throttled_ms=7771.2
pod e7-2x-limit2-2 avg_cpu_cores=0.258 nr_periods=357 nr_throttled=14 throttled_ratio=0.039 throttled_ms=985.2
===== e7-1x-limit4 replicas=1 limit=4 request=500m rps=100 clump=50 ping_rps=20 window_s=73
load label=e7-1x-limit4 gomaxprocs=10 clump=50 requests=6400 errors=0 achieved_rps=106.7
work_latency_ms p50=34.2 p90=58.9 p99=105.1 p999=156.1 max=182.3
ping_latency_ms requests=1227 p50=2.1 p90=4.8 p99=23.5 max=57.3
pod e7-1x-limit4-1 avg_cpu_cores=0.634 nr_periods=531 nr_throttled=21 throttled_ratio=0.040 throttled_ms=3717.7
Output of the second run:
===== e7-1x-nolimit replicas=1 limit=none request=500m rps=100 clump=50 ping_rps=20 window_s=72
load label=e7-1x-nolimit gomaxprocs=10 clump=50 requests=6600 errors=0 achieved_rps=110.1
work_latency_ms p50=35.4 p90=59.5 p99=98.7 p999=143.1 max=150.0
ping_latency_ms requests=1124 p50=2.2 p90=4.8 p99=11.3 max=35.7
pod e7-1x-nolimit-1 avg_cpu_cores=0.689 nr_periods=0 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
===== e7-1x-limit2 replicas=1 limit=2 request=500m rps=100 clump=50 ping_rps=20 window_s=72
load label=e7-1x-limit2 gomaxprocs=10 clump=50 requests=5000 errors=0 achieved_rps=83.4
work_latency_ms p50=73.9 p90=162.8 p99=304.5 p999=520.0 max=525.2
ping_latency_ms requests=1211 p50=2.6 p90=14.5 p99=101.9 max=294.5
pod e7-1x-limit2-1 avg_cpu_cores=0.519 nr_periods=587 nr_throttled=125 throttled_ratio=0.213 throttled_ms=35555.8
===== e7-2x-limit1 replicas=2 limit=1 request=500m rps=100 clump=50 ping_rps=20 window_s=73
load label=e7-2x-limit1 gomaxprocs=10 clump=50 requests=5800 errors=0 achieved_rps=96.6
work_latency_ms p50=89.5 p90=213.2 p99=495.7 p999=689.7 max=693.2
ping_latency_ms requests=1209 p50=2.5 p90=28.2 p99=99.3 max=189.9
pod e7-2x-limit1-1 avg_cpu_cores=0.294 nr_periods=490 nr_throttled=163 throttled_ratio=0.333 throttled_ms=39898.4
pod e7-2x-limit1-2 avg_cpu_cores=0.291 nr_periods=449 nr_throttled=157 throttled_ratio=0.350 throttled_ms=39071.2
===== e7-2x-limit2 replicas=2 limit=2 request=500m rps=100 clump=50 ping_rps=20 window_s=73
load label=e7-2x-limit2 gomaxprocs=10 clump=50 requests=6400 errors=0 achieved_rps=106.7
work_latency_ms p50=34.5 p90=104.7 p99=166.5 p999=248.8 max=271.9
ping_latency_ms requests=1166 p50=2.3 p90=5.5 p99=52.1 max=159.9
pod e7-2x-limit2-1 avg_cpu_cores=0.445 nr_periods=517 nr_throttled=82 throttled_ratio=0.159 throttled_ms=17798.0
pod e7-2x-limit2-2 avg_cpu_cores=0.203 nr_periods=329 nr_throttled=12 throttled_ratio=0.036 throttled_ms=3015.5
===== e7-1x-limit4 replicas=1 limit=4 request=500m rps=100 clump=50 ping_rps=20 window_s=72
load label=e7-1x-limit4 gomaxprocs=10 clump=50 requests=5800 errors=0 achieved_rps=96.7
work_latency_ms p50=35.0 p90=58.0 p99=123.3 p999=176.2 max=192.6
ping_latency_ms requests=1200 p50=2.3 p90=4.7 p99=26.9 max=101.8
pod e7-1x-limit4-1 avg_cpu_cores=0.629 nr_periods=544 nr_throttled=9 throttled_ratio=0.017 throttled_ms=1092.7
As a table (each cell lists the first run, then the second; with two replicas, usage and the throttled share are shown for each Pod):
| Setting | Average usage per Pod | Throttled periods | /work p99 | /ping p99 |
|---|---|---|---|---|
| 1 Pod, no limit | 0.60 / 0.69 | — | 86 / 99ms | 12 / 11ms |
| 1 Pod, limit 2 | 0.57 / 0.52 | 23.6% / 21.3% | 264 / 305ms | 86 / 102ms |
| 2 Pods, limit 1 each | 0.15 and 0.51 / 0.29 and 0.29 | 29% and 53% / 33% and 35% | 603 / 496ms | 225 / 99ms |
| 2 Pods, limit 2 each | 0.38 and 0.26 / 0.45 and 0.20 | 10% and 4% / 16% and 4% | 140 / 167ms | 41 / 52ms |
| 1 Pod, limit 4 | 0.63 / 0.63 | 4.0% / 1.7% | 105 / 123ms | 24 / 27ms |
- Splitting the same total limit between two replicas makes things worse.
e7-2x-limit1has a higher p99 thane7-1x-limit2in both runs. Each Pod gets about 25 requests, about 150ms of CPU, which is still more than its own 100ms budget, and the budget one Pod does not use cannot be borrowed by the other. In the second run, the two Pods used almost the same amount (0.29 and 0.29), and the result was still worse, so an uneven split is not the only problem. - Two replicas with limit 2 do worse than one Pod with limit 4. Both have a total limit of 4, but in both runs the two replicas' usage was uneven, and the Pod that got more requests still had 10% to 16% of its periods throttled. With one limit 4 Pod, a whole group's 300ms of CPU fits within one period.
- Raising the limit comes closest to no limit.
e7-1x-limit4has a p99 of 105 / 123ms, and no limit has 86 / 99ms. e7-1x-limit2did somewhat better in this round than in section 5 (p99 264 / 305ms, compared with 370 / 360ms in section 5). That is the difference between runs at different times. When comparing, look at the relative differences within the same round.
The uneven split between replicas probably comes from how the Service distributes traffic: kube-proxy picks a Pod when a connection is set up, the load generator reuses connections, and which Pod a group of requests lands on depends on which connections it happens to use. Service virtual IPs The lab did not record the number of connections per Pod. It only observed the uneven usage.
What this shows: for requests that arrive in groups, the cap a single Pod can use during a burst matters more than the total limit. Adding replicas without raising the total limit does not help. What it does not show: what happens when the Node has no idle CPU. The Node in this experiment has 10 CPUs and almost no other workload, so a higher limit can borrow idle CPU. The results on a busy Node are in section 8. Results may also differ when replicas are spread across Nodes, or with load balancing that picks a Pod for each request (such as an L7 proxy).
7. The request becomes cpu.weight
Weight at each level
Create a few Pods with different requests and read their cpu.weight at each level. Then use docker update --cpuset-cpus to pin the whole worker Node to one CPU, so these constantly busy containers compete for CPU, and see how much each one actually gets:
cat > scripts/e4-e5-weight.sh <<'EOF'
#!/usr/bin/env bash
# E4: cpu.weight at each cgroup level for different requests.
# E5: pin the worker to 1 CPU, let two busy threads compete, and measure the share each actually gets.
# pods: two single-container Pods (request 100m and 1); compares the Pod-level weight
# containers: two containers in the same Pod (request 100m and 1); compares the container-level weight
# Usage: e4-e5-weight.sh <kx command> <worker node>
set -euo pipefail
KX=$1; NODE=$2; here=$(cd "$(dirname "$0")/.." && pwd)
docker cp "$here/scripts/cgtree.sh" "$NODE:/cgtree.sh" >/dev/null
echo "== node runtime"; docker exec "$NODE" sh -c 'runc --version | head -1; containerd --version | cut -d" " -f3'
$KX get node "$NODE" -o jsonpath='kubelet={.status.nodeInfo.kubeletVersion} allocatable_cpu={.status.allocatable.cpu}{"\n"}'
pod() { # name spec-containers-json
$KX -n cpulab apply -f - >/dev/null <<YAML
{"apiVersion":"v1","kind":"Pod","metadata":{"name":"$1","namespace":"cpulab"},
"spec":{"nodeSelector":{"kubernetes.io/hostname":"$NODE"},"containers":$2}}
YAML
}
c() { # name request [limit]
local res="{\"requests\":{\"cpu\":\"$2\"}}"
[ $# -ge 3 ] && res="{\"requests\":{\"cpu\":\"$2\"},\"limits\":{\"cpu\":\"$3\"}}"
printf '{"name":"%s","image":"cpulab:go125mod","args":["spin","-threads=1","-report=10s"],"resources":%s}' "$1" "$res"
}
tree() { docker exec "$NODE" sh /cgtree.sh "$($KX -n cpulab get pod "$1" -o jsonpath='{.metadata.uid}')"; }
names() { # print which cgroup scope belongs to each container name
$KX -n cpulab get pod "$1" -o jsonpath='{range .status.containerStatuses[*]}{.name}={.containerID}{"\n"}{end}' | sed 's|containerd://\(.\{12\}\).*|cri-containerd-\1…|'
}
$KX get ns cpulab >/dev/null 2>&1 || $KX create ns cpulab >/dev/null
# If the script fails or is interrupted, the trap deletes these constantly busy Pods and gives the worker all CPUs again
# (an empty string does not clear the setting, so write out the range)
all="0-$(( $(docker info --format '{{.NCPU}}') - 1 ))"
restore() { docker update --cpuset-cpus "$all" "$NODE" >/dev/null; }
cleanup() {
$KX -n cpulab delete pod e4-req100m e4-req1 e4-req2 e4-req1-limit1 e5-two-containers \
--ignore-not-found --wait=false >/dev/null 2>&1 || true
restore
}
trap cleanup EXIT
# First read each level's weight without pinning the CPU (spin runs, but with 10 CPUs there is no contention)
pod e4-req100m "[$(c app 100m)]"
pod e4-req1 "[$(c app 1)]"
pod e4-req2 "[$(c app 2)]"
pod e4-req1-limit1 "[$(c app 1 1)]" # only CPU is set, no memory, so it is still Burstable
pod e5-two-containers "[$(c small 100m),$(c big 1)]"
$KX -n cpulab wait --for=condition=Ready pod e4-req100m e4-req1 e4-req2 e4-req1-limit1 e5-two-containers --timeout=120s >/dev/null
echo "== E4 cgroup tree"
for p in e4-req100m e4-req1 e4-req2 e4-req1-limit1 e5-two-containers; do echo "-- $p"; tree $p; names $p; done
echo "-- root and kubepods siblings"
docker exec "$NODE" sh -c 'for d in /sys/fs/cgroup/*/ /sys/fs/cgroup/kubelet.slice/*/ /sys/fs/cgroup/kubelet.slice/kubelet-kubepods.slice/*/; do [ -f $d/cpu.weight ] && printf "%-90s weight=%s\n" "${d#/sys/fs/cgroup}" "$(cat $d/cpu.weight)"; done | grep -v "pod[0-9a-f_]*.slice/$"'
echo "== E5 pin worker to one CPU"
$KX -n cpulab delete pod e4-req2 e4-req1-limit1 --wait=true >/dev/null
docker update --cpuset-cpus 5 "$NODE" >/dev/null
docker exec "$NODE" cat /sys/fs/cgroup/cpuset.cpus.effective
sleep 45
echo "-- pods: e4-req100m vs e4-req1 (Pod-level comparison), with e5-two-containers also running"
for p in e4-req100m e4-req1; do echo "$p: $($KX -n cpulab logs $p --tail=3 | tr '\n' ' ')"; done
for cn in small big; do echo "e5-two-containers/$cn: $($KX -n cpulab logs e5-two-containers -c $cn --tail=3 | tr '\n' ' ')"; done
echo "-- only the two-container pod (container-level comparison)"
$KX -n cpulab delete pod e4-req100m e4-req1 --wait=true >/dev/null
sleep 45
for cn in small big; do echo "e5-two-containers/$cn: $($KX -n cpulab logs e5-two-containers -c $cn --tail=3 | tr '\n' ' ')"; done
restore
$KX -n cpulab delete pod e5-two-containers --wait=false >/dev/null
echo "== restored cpuset: $(docker exec "$NODE" cat /sys/fs/cgroup/cpuset.cpus.effective)"
EOF
chmod +x scripts/e4-e5-weight.sh
scripts/e4-e5-weight.sh "scripts/kx new" cpu-lab-new-worker
An excerpt showing the request 1 Pod, the two-container Pod, and the levels near the root:
== node runtime
runc version 1.4.3
v2.3.4
kubelet=v1.34.11 allocatable_cpu=10
== E4 cgroup tree
...
-- e4-req1
/kubelet.slice weight=100 max=max 100000
/kubelet.slice/kubelet-kubepods.slice weight=391 max=max 100000
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-burstable.slice weight=207 max=max 100000
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-burstable.slice/kubelet-kubepods-burstable-podc54295e7_f94f_4a91_9bea_bf0393eacb82.slice weight=39 max=max 100000
cri-containerd-820a1ebf6ed44fe7d1f184221ddcc247daa6a5590210df041f60bd42d3c60650.scope weight=1 max=max 100000
cri-containerd-ccd098da4e5c8c486a18746d717a2fbdeefa7af21d56afc5c2c1bfa1e90f70e5.scope weight=100 max=max 100000
app=cri-containerd-ccd098da4e5c…
...
-- e5-two-containers
/kubelet.slice weight=100 max=max 100000
/kubelet.slice/kubelet-kubepods.slice weight=391 max=max 100000
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-burstable.slice weight=207 max=max 100000
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-burstable.slice/kubelet-kubepods-burstable-pod767f0518_7b92_487f_a740_9b8c96d55a0f.slice weight=43 max=max 100000
cri-containerd-773b31d845ecdb884077fbe73ce281f319333a73b1078b7f2a22ddc50eaeebf3.scope weight=17 max=max 100000
cri-containerd-77eff6a412b38a06aa05c58fc7019d6cee64124e0ed30575d62a44e7effb7016.scope weight=1 max=max 100000
cri-containerd-fcf7d3f159249325fadaa4634a4ef30feadbe8900561858d36faf9d7d2e968ab.scope weight=100 max=max 100000
big=cri-containerd-fcf7d3f15924…
small=cri-containerd-773b31d845ec…
-- root and kubepods siblings
/init.scope/ weight=100
/kubelet.slice/ weight=100
/kubelet/ weight=100
/sys-fs-fuse-connections.mount/ weight=100
/sys-kernel-debug.mount/ weight=100
/sys-kernel-tracing.mount/ weight=100
/system.slice/ weight=100
/kubelet.slice/kubelet-kubepods.slice/ weight=391
/kubelet.slice/kubelet.service/ weight=100
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-besteffort.slice/ weight=1
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-burstable.slice/ weight=207
The Pod level and the container level are two different numbers:
| request | Pod level (kubelet) | Container level (runc 1.4.3) |
|---|---|---|
| 100m | 4 | 17 |
| 1 | 39 | 100 |
| 2 | 79 | 174 |
kubelet creates the Pod, QoS, and kubepods levels with a linear conversion, where 1 CPU is 39. runc creates the container level. Starting with runc 1.3.2, it uses a new formula, where 1 CPU is exactly 100, the cgroup v2 default. The 391 for kubepods comes from the Node's Allocatable (10 CPUs), and kubelet.service at the same level is 100.
Actual shares on one CPU
== E5 pin worker to one CPU
5
-- pods: e4-req100m vs e4-req1 (Pod-level comparison), with e5-two-containers also running
e4-req100m: 2026-09-26T06:06:26Z usage_cores=0.046 2026-09-26T06:06:36Z usage_cores=0.046 2026-09-26T06:06:46Z usage_cores=0.046
e4-req1: 2026-09-26T06:06:26Z usage_cores=0.449 2026-09-26T06:06:36Z usage_cores=0.447 2026-09-26T06:06:46Z usage_cores=0.445
e5-two-containers/small: 2026-09-26T06:06:26Z usage_cores=0.072 2026-09-26T06:06:36Z usage_cores=0.072 2026-09-26T06:06:46Z usage_cores=0.071
e5-two-containers/big: 2026-09-26T06:06:26Z usage_cores=0.424 2026-09-26T06:06:36Z usage_cores=0.421 2026-09-26T06:06:46Z usage_cores=0.420
-- only the two-container pod (container-level comparison)
e5-two-containers/small: 2026-09-26T06:07:16Z usage_cores=0.144 2026-09-26T06:07:26Z usage_cores=0.144 2026-09-26T06:07:36Z usage_cores=0.144
e5-two-containers/big: 2026-09-26T06:07:16Z usage_cores=0.846 2026-09-26T06:07:26Z usage_cores=0.848 2026-09-26T06:07:36Z usage_cores=0.848
The output has two parts:
- Three busy Pods compete.
e4-req100m,e4-req1, and the two-containere5-two-containers(total request 1.1) get 0.046, 0.447, and 0.493 (both containers added together), about 4.7% : 45.3% : 50.0%. This is the same ratio as the Pod-level weights 4 : 39 : 43. - Only the two-container Pod is left. In the same Pod, the two containers with request 100m and 1 get 0.144 : 0.847, about 14.5% : 85.5%. This matches the container-level weights 17 : 100, not the request ratio of 1 : 10.
Each measurement waits 45 seconds for the shares to settle. Each container prints its usage every 10 seconds, and the numbers above are the average of the last three readings.
Restoring the CPUs: a mistake made this time
At the end, the script uses docker update to give the worker's cpuset all CPUs again. In the first run, the restore command passed an empty string, --cpuset-cpus "", which had no effect: the last line of output still said restored cpuset: 5, and the worker still had only one CPU. The script was then changed to write out the range explicitly (in the form 0-9), and to use a trap that deletes the constantly busy Pods and restores the cpuset again when the script exits. That way, a failure midway or Ctrl-C does not leave the worker on one CPU or leave Pods running forever. The script above is that version.
Check after running it. If it is not all CPUs, change it back by hand:
docker exec cpu-lab-new-worker cat /sys/fs/cgroup/cpuset.cpus.effective
docker update --cpuset-cpus "0-$(( $(docker info --format '{{.NCPU}}') - 1 ))" cpu-lab-new-worker
What this shows: at run time, the request becomes the cpu.weight at each level. Pods share CPU according to kubelet's conversion, and containers in the same Pod share it according to the runtime's conversion. What it does not show: how Pods compete with system services on a real Node. A kind Node is itself a container, so the levels near the root differ from a real host.
8. When the Node is busy, does raising the limit still help?
The Node in section 6 had almost no other work, so a higher limit could borrow idle CPU. This section adds a neighbor that keeps the CPU fully busy on the same Node, to see whether raising the limit still helps, and what role the request plays.
- Traffic: the same as section 6, on average 100
/workper second, 50 arriving together each time, plus 20/pingper second. The server is a single Pod using the go.mod 1.24 image. - Neighbor: request 4 and no limit, with
cpulab spinkeeping 8 threads busy. It stands for other work on the Node that has no limit and runs at full speed. - Separate CPUs: the two kind Nodes are really containers on the same VM, sharing 10 CPUs. If the neighbor used all 10, the load generator on the control-plane would also slow down, and the measured latency would include the load generator's own delay. So the script first uses
docker update --cpuset-cpusto pin the worker to CPUs 0-7 and the control-plane to 8-9. At the end, or when the script fails or you press Ctrl-C, atrapdeletes Pods still running (so the neighbor does not keep using the CPU) and restores the cpuset. With only 8 CPUs left on the worker, the server's GOMAXPROCS also becomes 8. So this section only compares with the idle-Node runs in the same round, not with the numbers in section 6.
The six settings run one after another:
| Name | Server request | limit | Neighbor |
|---|---|---|---|
e8-idle-limit2 | 500m | 2 | No |
e8-idle-limit4 | 500m | 4 | No |
e8-busy-limit2 | 500m | 2 | Yes |
e8-busy-limit4 | 500m | 4 | Yes |
e8-busy-req2-limit4 | 2 | 4 | Yes |
e8-busy-req4-limit4 | 4 | 4 | Yes |
By kubelet's conversion in section 7, the Pod-level weight is 20 for 500m, 79 for 2, and 157 for 4. The neighbor's request is 4, so as the server's request goes from 500m up to 4, its share on a fully loaded Node is roughly 20 : 157, 79 : 157, and 157 : 157.
cat > scripts/e8-contention.sh <<'EOF'
#!/usr/bin/env bash
# E8: same clumped traffic, on an idle Node and on a Node filled by a neighbor; compares raising only the limit with raising the request too.
# Usage: e8-contention.sh <kx command> <worker node> <cp node> <rps> <name:request:limit:idle|busy>...
# The load generator and the server share one VM. So the neighbor does not slow down the load generator too, use cpuset to separate the two kind nodes:
# the worker uses CPUs 0-7 and the control-plane uses 8-9. When busy, the neighbor is a Pod with request 4, no limit, and 8 threads computing nonstop.
# If the script fails or is interrupted, the trap deletes Pods still running (so the neighbor does not keep using the CPU), then restores the cpuset.
set -euo pipefail
KX=$1; NODE=$2; CP=$3; RPS=$4; shift 4
all="0-$(( $(docker info --format '{{.NCPU}}') - 1 ))"
restore() { docker update --cpuset-cpus "$all" "$NODE" "$CP" >/dev/null; }
active=() # Pods still running
cleanup() {
[ ${#active[@]} -eq 0 ] || $KX -n cpulab delete pod "${active[@]}" --ignore-not-found --wait=false >/dev/null 2>&1 || true
restore
}
trap cleanup EXIT
docker update --cpuset-cpus 0-7 "$NODE" >/dev/null
docker update --cpuset-cpus 8-9 "$CP" >/dev/null
echo "== cpuset worker=$(docker exec "$NODE" cat /sys/fs/cgroup/cpuset.cpus.effective) control-plane=$(docker exec "$CP" cat /sys/fs/cgroup/cpuset.cpus.effective)"
$KX get ns cpulab >/dev/null 2>&1 || $KX create ns cpulab >/dev/null
$KX -n cpulab get svc cpulab-server >/dev/null 2>&1 || \
$KX -n cpulab create service clusterip cpulab-server --tcp=8080:8080 >/dev/null
$KX -n cpulab patch svc cpulab-server -p '{"spec":{"selector":{"app":"cpulab-server"}}}' >/dev/null
cpustat() { $KX -n cpulab exec "$1" -- cat /sys/fs/cgroup/cpu.stat | awk '{printf "%s=%s ", $1, $2}'; }
delta() { # before after elapsed_s -> usage and throttling over this window
printf '%s\n%s\n' "$1" "$2" | awk -v t="$3" '
{ for (i = 1; i <= NF; i++) { split($i, kv, "="); v[NR, kv[1]] = kv[2] } }
END {
u = v[2, "usage_usec"] - v[1, "usage_usec"]; p = v[2, "nr_periods"] - v[1, "nr_periods"]
n = v[2, "nr_throttled"] - v[1, "nr_throttled"]; s = v[2, "throttled_usec"] - v[1, "throttled_usec"]
printf "avg_cpu_cores=%.3f nr_periods=%d nr_throttled=%d throttled_ratio=%.3f throttled_ms=%.1f\n", u / t / 1e6, p, n, (p ? n / p : 0), s / 1000
}'
}
for spec in "$@"; do
IFS=: read -r name request limit node <<<"$spec"
res="{\"requests\":{\"cpu\":\"$request\"}}"
[ "$limit" != none ] && res="{\"requests\":{\"cpu\":\"$request\"},\"limits\":{\"cpu\":\"$limit\"}}"
pods=("$name")
$KX -n cpulab run "$name" --image=cpulab:go124mod --restart=Never --labels=app=cpulab-server \
--overrides="{\"spec\":{\"nodeSelector\":{\"kubernetes.io/hostname\":\"$NODE\"},\"containers\":[{\"name\":\"c\",\"image\":\"cpulab:go124mod\",\"args\":[\"serve\",\"-work=5ms\",\"-alloc=1048576\",\"-live-mb=64\"],\"resources\":$res,\"ports\":[{\"containerPort\":8080}],\"readinessProbe\":{\"httpGet\":{\"path\":\"/stats\",\"port\":8080}}}]}}" >/dev/null
active=("$name")
if [ "$node" = busy ]; then
$KX -n cpulab run "$name-neighbor" --image=cpulab:go124mod --restart=Never \
--overrides="{\"spec\":{\"nodeSelector\":{\"kubernetes.io/hostname\":\"$NODE\"},\"containers\":[{\"name\":\"c\",\"image\":\"cpulab:go124mod\",\"args\":[\"spin\",\"-threads=8\",\"-report=10s\"],\"resources\":{\"requests\":{\"cpu\":\"4\"}}}]}}" >/dev/null
pods+=("$name-neighbor"); active+=("$name-neighbor")
fi
$KX -n cpulab wait --for=condition=Ready pod "${pods[@]}" --timeout=120s >/dev/null
sleep 15 # let the neighbor fill up the CPU first
before=()
for p in "${pods[@]}"; do before+=("$(cpustat "$p")"); done
t0=$(date +%s)
$KX -n cpulab run "load-$name" --image=cpulab:go125mod --restart=Never \
--overrides="{\"spec\":{\"nodeSelector\":{\"kubernetes.io/hostname\":\"$CP\"},\"tolerations\":[{\"operator\":\"Exists\"}],\"containers\":[{\"name\":\"c\",\"image\":\"cpulab:go125mod\",\"args\":[\"load\",\"-target=http://cpulab-server:8080\",\"-rps=$RPS\",\"-clump=${CLUMP:-1}\",\"-ping-rps=${PINGRPS:-0}\",\"-duration=60s\",\"-warmup=10s\",\"-label=$name\"]}]}}" >/dev/null
active+=("load-$name")
$KX -n cpulab wait --for=jsonpath='{.status.phase}'=Succeeded pod/load-$name --timeout=180s >/dev/null
t1=$(date +%s)
echo "===== $name request=$request limit=$limit node=$node rps=$RPS clump=${CLUMP:-1} ping_rps=${PINGRPS:-0} window_s=$((t1 - t0))"
$KX -n cpulab logs "$name" | grep -m1 '^go='
$KX -n cpulab logs load-$name | grep -E '^(load|work_latency_ms|ping_latency_ms) '
i=0
for p in "${pods[@]}"; do
echo "pod $p $(delta "${before[$i]}" "$(cpustat "$p")" $((t1 - t0)))"
i=$((i + 1))
done
$KX -n cpulab delete pod "${pods[@]}" load-$name --wait=true >/dev/null
active=()
done
restore
echo "== restored cpuset worker=$(docker exec "$NODE" cat /sys/fs/cgroup/cpuset.cpus.effective) control-plane=$(docker exec "$CP" cat /sys/fs/cgroup/cpuset.cpus.effective)"
EOF
chmod +x scripts/e8-contention.sh
for round in 1 2; do
CLUMP=50 PINGRPS=20 scripts/e8-contention.sh "scripts/kx new" cpu-lab-new-worker cpu-lab-new-control-plane 100 \
e8-idle-limit2:500m:2:idle e8-idle-limit4:500m:4:idle e8-busy-limit2:500m:2:busy e8-busy-limit4:500m:4:busy \
e8-busy-req2-limit4:2:4:busy e8-busy-req4-limit4:4:4:busy
done
Each run takes about a minute and a half, and the whole set was run twice. Output of the first run:
== cpuset worker=0-7 control-plane=8-9
===== e8-idle-limit2 request=500m limit=2 node=idle rps=100 clump=50 ping_rps=20 window_s=72
go=go1.26.0 GOMAXPROCS=8 NumCPU=8 GOMAXPROCS_env="" GODEBUG_env=""
load label=e8-idle-limit2 gomaxprocs=8 clump=50 requests=6300 errors=0 achieved_rps=105.3
work_latency_ms p50=69.7 p90=198.2 p99=392.2 p999=616.5 max=641.1
ping_latency_ms requests=1162 p50=2.5 p90=22.6 p99=102.1 max=596.6
pod e8-idle-limit2 avg_cpu_cores=0.617 nr_periods=587 nr_throttled=153 throttled_ratio=0.261 throttled_ms=29664.4
===== e8-idle-limit4 request=500m limit=4 node=idle rps=100 clump=50 ping_rps=20 window_s=72
go=go1.26.0 GOMAXPROCS=8 NumCPU=8 GOMAXPROCS_env="" GODEBUG_env=""
load label=e8-idle-limit4 gomaxprocs=8 clump=50 requests=6000 errors=0 achieved_rps=100.1
work_latency_ms p50=38.4 p90=64.8 p99=98.7 p999=125.3 max=130.4
ping_latency_ms requests=1210 p50=2.2 p90=3.9 p99=13.7 max=41.8
pod e8-idle-limit4 avg_cpu_cores=0.620 nr_periods=560 nr_throttled=10 throttled_ratio=0.018 throttled_ms=602.9
===== e8-busy-limit2 request=500m limit=2 node=busy rps=100 clump=50 ping_rps=20 window_s=72
go=go1.26.0 GOMAXPROCS=8 NumCPU=8 GOMAXPROCS_env="" GODEBUG_env=""
load label=e8-busy-limit2 gomaxprocs=8 clump=50 requests=7048 errors=2 achieved_rps=117.1
work_latency_ms p50=637.3 p90=3372.0 p99=8358.8 p999=9917.5 max=9941.8
ping_latency_ms requests=1207 p50=45.2 p90=1954.6 p99=4257.4 max=8340.0
pod e8-busy-limit2 avg_cpu_cores=0.665 nr_periods=636 nr_throttled=19 throttled_ratio=0.030 throttled_ms=471.5
pod e8-busy-limit2-neighbor avg_cpu_cores=7.341 nr_periods=0 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
===== e8-busy-limit4 request=500m limit=4 node=busy rps=100 clump=50 ping_rps=20 window_s=72
go=go1.26.0 GOMAXPROCS=8 NumCPU=8 GOMAXPROCS_env="" GODEBUG_env=""
load label=e8-busy-limit4 gomaxprocs=8 clump=50 requests=6100 errors=0 achieved_rps=101.5
work_latency_ms p50=553.5 p90=2916.2 p99=6044.7 p999=8730.0 max=9278.3
ping_latency_ms requests=1186 p50=32.8 p90=1026.4 p99=4513.0 max=8332.0
pod e8-busy-limit4 avg_cpu_cores=0.603 nr_periods=654 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
pod e8-busy-limit4-neighbor avg_cpu_cores=7.457 nr_periods=0 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
===== e8-busy-req2-limit4 request=2 limit=4 node=busy rps=100 clump=50 ping_rps=20 window_s=72
go=go1.26.0 GOMAXPROCS=8 NumCPU=8 GOMAXPROCS_env="" GODEBUG_env=""
load label=e8-busy-req2-limit4 gomaxprocs=8 clump=50 requests=5800 errors=0 achieved_rps=96.7
work_latency_ms p50=81.5 p90=188.5 p99=471.8 p999=564.8 max=666.0
ping_latency_ms requests=1200 p50=2.3 p90=15.6 p99=119.3 max=261.4
pod e8-busy-req2-limit4 avg_cpu_cores=0.565 nr_periods=468 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
pod e8-busy-req2-limit4-neighbor avg_cpu_cores=7.355 nr_periods=0 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
===== e8-busy-req4-limit4 request=4 limit=4 node=busy rps=100 clump=50 ping_rps=20 window_s=72
go=go1.26.0 GOMAXPROCS=8 NumCPU=8 GOMAXPROCS_env="" GODEBUG_env=""
load label=e8-busy-req4-limit4 gomaxprocs=8 clump=50 requests=6400 errors=0 achieved_rps=106.7
work_latency_ms p50=70.7 p90=134.5 p99=222.3 p999=289.3 max=330.6
ping_latency_ms requests=1197 p50=2.3 p90=9.3 p99=62.0 max=119.8
pod e8-busy-req4-limit4 avg_cpu_cores=0.581 nr_periods=463 nr_throttled=1 throttled_ratio=0.002 throttled_ms=8.4
pod e8-busy-req4-limit4-neighbor avg_cpu_cores=7.354 nr_periods=0 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
== restored cpuset worker=0-9 control-plane=0-9
Output of the second run:
== cpuset worker=0-7 control-plane=8-9
===== e8-idle-limit2 request=500m limit=2 node=idle rps=100 clump=50 ping_rps=20 window_s=72
go=go1.26.0 GOMAXPROCS=8 NumCPU=8 GOMAXPROCS_env="" GODEBUG_env=""
load label=e8-idle-limit2 gomaxprocs=8 clump=50 requests=6600 errors=0 achieved_rps=110.1
work_latency_ms p50=79.1 p90=207.3 p99=492.0 p999=680.2 max=828.7
ping_latency_ms requests=1186 p50=2.6 p90=33.0 p99=197.8 max=309.4
pod e8-idle-limit2 avg_cpu_cores=0.656 nr_periods=606 nr_throttled=158 throttled_ratio=0.261 throttled_ms=27100.3
===== e8-idle-limit4 request=500m limit=4 node=idle rps=100 clump=50 ping_rps=20 window_s=72
go=go1.26.0 GOMAXPROCS=8 NumCPU=8 GOMAXPROCS_env="" GODEBUG_env=""
load label=e8-idle-limit4 gomaxprocs=8 clump=50 requests=7100 errors=0 achieved_rps=118.6
work_latency_ms p50=38.7 p90=67.3 p99=137.3 p999=203.4 max=216.3
ping_latency_ms requests=1198 p50=2.2 p90=5.0 p99=31.6 max=51.9
pod e8-idle-limit4 avg_cpu_cores=0.694 nr_periods=565 nr_throttled=23 throttled_ratio=0.041 throttled_ms=2570.3
===== e8-busy-limit2 request=500m limit=2 node=busy rps=100 clump=50 ping_rps=20 window_s=73
go=go1.26.0 GOMAXPROCS=8 NumCPU=8 GOMAXPROCS_env="" GODEBUG_env=""
load label=e8-busy-limit2 gomaxprocs=8 clump=50 requests=6600 errors=0 achieved_rps=110.1
work_latency_ms p50=437.7 p90=1476.7 p99=3589.7 p999=5414.4 max=5518.4
ping_latency_ms requests=1229 p50=38.0 p90=574.5 p99=3033.9 max=3480.7
pod e8-busy-limit2 avg_cpu_cores=0.614 nr_periods=634 nr_throttled=15 throttled_ratio=0.024 throttled_ms=543.0
pod e8-busy-limit2-neighbor avg_cpu_cores=7.344 nr_periods=0 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
===== e8-busy-limit4 request=500m limit=4 node=busy rps=100 clump=50 ping_rps=20 window_s=73
go=go1.26.0 GOMAXPROCS=8 NumCPU=8 GOMAXPROCS_env="" GODEBUG_env=""
load label=e8-busy-limit4 gomaxprocs=8 clump=50 requests=5800 errors=0 achieved_rps=96.3
work_latency_ms p50=361.7 p90=1675.5 p99=3136.6 p999=3575.5 max=4230.2
ping_latency_ms requests=1176 p50=4.7 p90=346.6 p99=1664.3 max=3199.5
pod e8-busy-limit4 avg_cpu_cores=0.541 nr_periods=593 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
pod e8-busy-limit4-neighbor avg_cpu_cores=7.423 nr_periods=0 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
===== e8-busy-req2-limit4 request=2 limit=4 node=busy rps=100 clump=50 ping_rps=20 window_s=73
go=go1.26.0 GOMAXPROCS=8 NumCPU=8 GOMAXPROCS_env="" GODEBUG_env=""
load label=e8-busy-req2-limit4 gomaxprocs=8 clump=50 requests=6000 errors=0 achieved_rps=100.1
work_latency_ms p50=84.7 p90=188.5 p99=341.1 p999=494.6 max=598.4
ping_latency_ms requests=1238 p50=2.2 p90=13.2 p99=99.7 max=410.7
pod e8-busy-req2-limit4 avg_cpu_cores=0.589 nr_periods=460 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
pod e8-busy-req2-limit4-neighbor avg_cpu_cores=7.281 nr_periods=0 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
===== e8-busy-req4-limit4 request=4 limit=4 node=busy rps=100 clump=50 ping_rps=20 window_s=72
go=go1.26.0 GOMAXPROCS=8 NumCPU=8 GOMAXPROCS_env="" GODEBUG_env=""
load label=e8-busy-req4-limit4 gomaxprocs=8 clump=50 requests=6500 errors=0 achieved_rps=108.4
work_latency_ms p50=74.8 p90=141.7 p99=213.4 p999=267.9 max=287.2
ping_latency_ms requests=1230 p50=2.1 p90=8.0 p99=77.8 max=213.6
pod e8-busy-req4-limit4 avg_cpu_cores=0.655 nr_periods=465 nr_throttled=2 throttled_ratio=0.004 throttled_ms=14.9
pod e8-busy-req4-limit4-neighbor avg_cpu_cores=7.306 nr_periods=0 nr_throttled=0 throttled_ratio=0.000 throttled_ms=0.0
== restored cpuset worker=0-9 control-plane=0-9
As a table (each cell lists the first run, then the second):
| Setting | Server average usage | Throttled periods | /work p99 | /ping p99 |
|---|---|---|---|---|
e8-idle-limit2 | 0.62 / 0.66 | 26.1% / 26.1% | 392 / 492ms | 102 / 198ms |
e8-idle-limit4 | 0.62 / 0.69 | 1.8% / 4.1% | 99 / 137ms | 14 / 32ms |
e8-busy-limit2 | 0.67 / 0.61 | 3.0% / 2.4% | 8,359 / 3,590ms | 4,257 / 3,034ms |
e8-busy-limit4 | 0.60 / 0.54 | 0% / 0% | 6,045 / 3,137ms | 4,513 / 1,664ms |
e8-busy-req2-limit4 | 0.57 / 0.59 | 0% / 0% | 472 / 341ms | 119 / 100ms |
e8-busy-req4-limit4 | 0.58 / 0.66 | 0.2% / 0.4% | 222 / 213ms | 62 / 78ms |
- On the busy Node, both limit 2 and limit 4 take several seconds. There is almost no throttling, and none at all with limit 4, yet even p50 is 360 to 640ms. The server's average usage is about the same as on the idle Node (0.54 to 0.67 CPU), but it cannot get CPU during bursts. Based on the weights above, on a fully loaded Node the server should get only 8 × 20 ÷ 177, about 0.9 CPU. This is an estimate; the script did not measure how much it actually got during a burst.
- The throttling metrics do not show the problem. If you only looked at throttling,
e8-busy-limit4would be one of the best of the six runs, yet its latency is one of the two worst. - Each step up in the request improves latency. With request 2, p99 is 341 / 472ms; with request 4, it is 213 / 222ms. Request 4 has the same weight as the neighbor, so on a fully loaded Node the server should get about 4 CPUs, exactly its limit, yet it is still slower than
e8-idle-limit4. Weight decides the share over a period of time. Probably, at the start of a burst, the server's threads still have to take turns with the neighbor that is already running. This was not measured. - The neighbor kept using 7.3 to 7.5 CPUs, which is the part of the 8 CPUs the server did not use.
- The busy-Node runs varied a lot between the two runs (
e8-busy-limit2was 8.4 and 3.6 seconds). When a service is on the edge of queueing, a few groups arriving a bit closer together make the queue longer, so latency is very sensitive to arrival times. Read only the order of magnitude. The first run ofe8-busy-limit2also had 2 failed requests (errors=2), which are not counted in the latency. The load generator's timeout is 10 seconds, and the longest successful request took 9.9 seconds, so these were probably timeouts. The program used at the time added up the errors from/workand/ping, so it is unclear which kind of request's p99 is slightly too low because of them (the program in the appendix now counts them separately).
What this shows: on a busy Node, a limit is only a cap. With a higher limit, there is less throttling, yet latency is measured in seconds. What decides how much CPU the service gets is the weight converted from the request. What it does not show: how busy a real Node gets. The neighbor is an extreme case made on purpose: it keeps the Node full, has request 4, and has no limit. Real Nodes are rarely fully loaded for long; more often, bursts from a few Pods happen to overlap. The size of the effect also depends on the neighbor's request. The worker has only 8 CPUs and the load generator runs on the other 2, so these numbers cannot be compared directly with section 6.
9. The numbers monitoring sees
Finally, look at the metrics a monitoring system actually reads. Reuse the limit 1 job from section 3 (8 threads × 20ms, once per second), and read kubelet's built-in cAdvisor metrics twice while it runs, about 20 seconds apart:
cat > scripts/e6-cadvisor.sh <<'EOF'
#!/usr/bin/env bash
# E6: the cAdvisor metrics that monitoring actually reads, compared with cpu.stat and elapsed time.
# Usage: e6-cadvisor.sh <kx command> <worker node>
set -euo pipefail
KX=$1; NODE=$2
echo "== E6: cAdvisor metrics for a throttled burst pod (limit 1, 8 threads x 20ms, 30 rounds)"
$KX -n cpulab run e6-burst --image=cpulab:go125mod --restart=Never --overrides="{\"spec\":{\"nodeSelector\":{\"kubernetes.io/hostname\":\"$NODE\"},\"containers\":[{\"name\":\"c\",\"image\":\"cpulab:go125mod\",\"args\":[\"burst\",\"-threads=8\",\"-work=20ms\",\"-interval=1s\",\"-rounds=30\"],\"resources\":{\"requests\":{\"cpu\":\"100m\"},\"limits\":{\"cpu\":\"1\"}},\"env\":[{\"name\":\"GOMAXPROCS\",\"value\":\"8\"}]}]}}" >/dev/null
$KX -n cpulab wait --for=condition=Ready pod/e6-burst --timeout=60s >/dev/null
m() { $KX get --raw "/api/v1/nodes/$NODE/proxy/metrics/cadvisor" | grep -E '^container_cpu_(cfs_periods_total|cfs_throttled_periods_total|cfs_throttled_seconds_total|usage_seconds_total)\{' | grep 'pod="e6-burst"' | grep 'container="c"' | sed -E 's/\{[^}]*\}//'; }
sleep 3; echo "-- t0 $(date -u +%T)"; m
sleep 20; echo "-- t0+20s $(date -u +%T)"; m
$KX -n cpulab wait --for=jsonpath='{.status.phase}'=Succeeded pod/e6-burst --timeout=60s >/dev/null
$KX -n cpulab logs e6-burst | tail -1
$KX -n cpulab delete pod e6-burst --wait=false >/dev/null
EOF
chmod +x scripts/e6-cadvisor.sh
scripts/e6-cadvisor.sh "scripts/kx new" cpu-lab-new-worker
== E6: cAdvisor metrics for a throttled burst pod (limit 1, 8 threads x 20ms, 30 rounds)
-- t0 06:11:40
container_cpu_cfs_periods_total 1 1790403097312
container_cpu_cfs_throttled_periods_total 0 1790403097312
container_cpu_cfs_throttled_seconds_total 0 1790403097312
container_cpu_usage_seconds_total 0.100817 1790403097312
-- t0+20s 06:12:00
container_cpu_cfs_periods_total 38 1790403109461
container_cpu_cfs_throttled_periods_total 13 1790403109461
container_cpu_cfs_throttled_seconds_total 5.77952 1790403109461
container_cpu_usage_seconds_total 2.169029 1790403109461
summary wall_ms p50=108.2 max=118.4 avg_cpu_cores=0.165 nr_periods=90 nr_throttled=30 throttled_ms=14215.2
The second number after each metric is cAdvisor's sample timestamp (in milliseconds). cAdvisor collects values periodically and returns the latest collection, so the two samples are only 12.149 seconds apart, not the 20 seconds the script waited. When computing a rate, divide by the difference between the timestamps:
| Metric | Increase | Divided by 12.15 seconds |
|---|---|---|
container_cpu_usage_seconds_total | 2.07 | 0.17 |
container_cpu_cfs_periods_total | 37 | — |
container_cpu_cfs_throttled_periods_total | 13 | — |
container_cpu_cfs_throttled_seconds_total | 5.78 | 0.48 |
In Prometheus, rate(container_cpu_cfs_throttled_seconds_total[1m]) would be about 0.48, which looks like "paused for almost half of every second." But this program pauses only once per second. The summary on the last line shows a p50 completion time of 108ms, compared with about 34ms without throttling, so the actual pause is about 75ms per second. 0.48 is the sum of the time each of the 8 threads was paused.
throttled_periods / periods is 13 / 37, about 35%. The denominator only counts busy periods: 12 seconds contain 120 periods, but only 37 were counted.
What this shows: throttled_seconds is good for comparing changes in the same service, but cannot be read directly as the share of time spent paused. The denominator of the throttled-period share is not time either. What it does not show: whether users were affected. To answer that, look at latency.
10. Optional: compare with an older runc
The container-level weight in section 7 is converted by the runtime. To see the old formula, create a second cluster with kind v0.30.0 and its default kindest/node:v1.34.0, which has runc 1.3.0. The kind binary has to match the node image version, so download the older binary separately and put it in tmp/:
curl -Lo tmp/kind-v0.30.0 "https://kind.sigs.k8s.io/dl/v0.30.0/kind-$(uname -s | tr '[:upper:]' '[:lower:]')-$(uname -m | sed 's/x86_64/amd64/;s/aarch64/arm64/')"
chmod +x tmp/kind-v0.30.0
tmp/kind-v0.30.0 create cluster --name cpu-lab-old \
--image kindest/node:v1.34.0@sha256:7416a61b42b1662ca6ca89f02028ac133a309a2a30ba309614e8ec94d976dc5a \
--config scripts/kind-2node.yaml --kubeconfig tmp/kubeconfig-old
tmp/kind-v0.30.0 load docker-image cpulab:go124mod cpulab:go125mod --name cpu-lab-old
scripts/e4-e5-weight.sh "scripts/kx old" cpu-lab-old-worker
An excerpt:
== node runtime
runc version 1.3.0
v2.1.3
kubelet=v1.34.0 allocatable_cpu=10
== E4 cgroup tree
...
-- e4-req1
/kubelet.slice weight=100 max=max 100000
/kubelet.slice/kubelet-kubepods.slice weight=391 max=max 100000
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-burstable.slice weight=203 max=max 100000
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-burstable.slice/kubelet-kubepods-burstable-poda3c4cda5_4e6a_4e4d_b481_780b0582dc30.slice weight=39 max=max 100000
cri-containerd-3dc2421657aaa0a60f8c022adc82906c0a2352363caa0aed347b7f516fc53ca2.scope weight=39 max=max 100000
cri-containerd-f7531c39e1b6774cd6330e7d367c31b2ff75f2aaa3d7ed9ec64fb8d9401003d7.scope weight=1 max=max 100000
app=cri-containerd-3dc2421657aa…
...
-- e5-two-containers
/kubelet.slice weight=100 max=max 100000
/kubelet.slice/kubelet-kubepods.slice weight=391 max=max 100000
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-burstable.slice weight=203 max=max 100000
/kubelet.slice/kubelet-kubepods.slice/kubelet-kubepods-burstable.slice/kubelet-kubepods-burstable-pod87d2c9b4_0278_44a1_9b52_0dd55decf53a.slice weight=43 max=max 100000
cri-containerd-817b5f50511244450dc154c802676b48735eee1e4f6e8876bfd285d1ba70952b.scope weight=4 max=max 100000
cri-containerd-e9d04eb20e481bbc64c2c1dce0caab68d555013219a8243d4a14aad5156333b1.scope weight=1 max=max 100000
cri-containerd-f4e1ff96bd0ba8ba7c529a98d4f66d113ddc22218d4bc57105bd1616c136c0ea.scope weight=39 max=max 100000
big=cri-containerd-f4e1ff96bd0b…
small=cri-containerd-817b5f505112…
...
== E5 pin worker to one CPU
5
-- pods: e4-req100m vs e4-req1 (Pod-level comparison), with e5-two-containers also running
e4-req100m: 2026-09-26T06:09:25Z usage_cores=0.046 2026-09-26T06:09:35Z usage_cores=0.046 2026-09-26T06:09:45Z usage_cores=0.046
e4-req1: 2026-09-26T06:09:25Z usage_cores=0.450 2026-09-26T06:09:35Z usage_cores=0.446 2026-09-26T06:09:45Z usage_cores=0.451
e5-two-containers/small: 2026-09-26T06:09:25Z usage_cores=0.046 2026-09-26T06:09:35Z usage_cores=0.046 2026-09-26T06:09:45Z usage_cores=0.046
e5-two-containers/big: 2026-09-26T06:09:25Z usage_cores=0.450 2026-09-26T06:09:35Z usage_cores=0.446 2026-09-26T06:09:45Z usage_cores=0.451
-- only the two-container pod (container-level comparison)
e5-two-containers/small: 2026-09-26T06:10:15Z usage_cores=0.092 2026-09-26T06:10:25Z usage_cores=0.092 2026-09-26T06:10:35Z usage_cores=0.093
e5-two-containers/big: 2026-09-26T06:10:15Z usage_cores=0.898 2026-09-26T06:10:25Z usage_cores=0.898 2026-09-26T06:10:35Z usage_cores=0.901
== restored cpuset: 0-9
The Pod-level weights are exactly the same as with the new runtime (39 and 43 in the excerpt), and the three Pods get 0.046, 0.449, and 0.495, almost the same as before. The container level, however, takes the same value as the Pod level: 39 for request 1 and 4 for 100m. So the two containers in the same Pod get 0.092 : 0.899, about 9.3% : 90.7%. The only difference between the two clusters is the runtime's conversion.
What this experiment can and cannot show
What the experiment shows clearly is how the kernel enforces a limit: the quota is handed out once per period and shared by all threads, and when it runs out, everything pauses. It also shows how the request becomes the weight at each level.
What it cannot show:
- The shape of the traffic: requests arriving in groups was made on purpose, to reproduce a service that averages 30% and is still throttled. What real bursts look like has to come from your own monitoring.
- Other loads: the thread-count trade-off was only tested with this one traffic pattern. Programs that stay near the limit for a long time, or where GC takes a large share, were not tested.
- How busy the Node is: the neighbor in section 8 keeps the Node full, which is an extreme case made on purpose. The more common case, where bursts from a few Pods sometimes overlap, was not tested.
- Properties of kind and the machine: a kind Node is a container with nested cgroups, so the levels near the root do not look like a real host. The OrbStack VM shares CPUs with macOS, so other activity affects the absolute latency.
- Settings not tested: CPU Manager's static policy,
cpu.max.burst, a custom CFS period, and the JVM.
Clean up
Delete the clusters, the images, and the files created this time:
kind delete cluster --name cpu-lab-new --kubeconfig tmp/kubeconfig-new
tmp/kind-v0.30.0 delete cluster --name cpu-lab-old --kubeconfig tmp/kubeconfig-old # only needed if you did section 10
docker rmi cpulab:go124mod cpulab:go125mod
Finally, delete the whole working directory.
Appendix: cpulab source code
Save it as cpulab/main.go. You do not need to write go.mod yourself; build.sh creates both versions.
After the experiments were run, the load generator load was changed in two places after review: errors from /work and /ping are now counted separately, and it no longer panics when every /work request fails. So the output in this article differs a little from what this program prints. In the article, errors= on the load line is the total for both kinds of requests, and the ping_latency_ms line has no errors=. With this program, errors= on the load line only counts /work, and /ping errors are printed on the ping_latency_ms line.
// cpulab is a small program for the CPU throttling experiments. It only runs inside throwaway kind clusters.
//
// cpulab info print the Go version, GOMAXPROCS, and the CPU settings of its cgroup
// cpulab burst [flags] at a fixed interval, make N threads do a fixed amount of CPU work at once; measure completion time
// cpulab spin [flags] keep N threads busy; periodically print the CPU usage of its own cgroup
// cpulab serve [flags] HTTP service: each request does a fixed amount of CPU work and allocates memory
// cpulab load [flags] open-loop load test: send requests with Poisson arrivals, print the latency distribution
//
// CPU work is measured in thread CPU time (lock the OS thread, then read RUSAGE_THREAD),
// so time spent waiting while throttled is not counted as work.
package main
import (
"bufio"
"encoding/json"
"flag"
"fmt"
"io"
"math"
"math/rand"
"net/http"
"os"
"runtime"
"sort"
"strconv"
"strings"
"sync"
"syscall"
"time"
)
const cgroupDir = "/sys/fs/cgroup"
func main() {
if len(os.Args) < 2 {
fmt.Fprintln(os.Stderr, "usage: cpulab info|burst|spin|serve|load [flags]")
os.Exit(2)
}
args := os.Args[2:]
switch os.Args[1] {
case "info":
info()
case "burst":
burst(args)
case "spin":
spin(args)
case "serve":
serve(args)
case "load":
load(args)
default:
fmt.Fprintln(os.Stderr, "unknown command:", os.Args[1])
os.Exit(2)
}
}
func readFile(name string) string {
b, err := os.ReadFile(cgroupDir + "/" + name)
if err != nil {
return "(" + err.Error() + ")"
}
return strings.TrimSpace(string(b))
}
func cpuStat() map[string]int64 {
m := map[string]int64{}
f, err := os.Open(cgroupDir + "/cpu.stat")
if err != nil {
return m
}
defer f.Close()
s := bufio.NewScanner(f)
for s.Scan() {
fields := strings.Fields(s.Text())
if len(fields) == 2 {
v, _ := strconv.ParseInt(fields[1], 10, 64)
m[fields[0]] = v
}
}
return m
}
func info() {
fmt.Printf("go=%s GOMAXPROCS=%d NumCPU=%d GOMAXPROCS_env=%q GODEBUG_env=%q\n",
runtime.Version(), runtime.GOMAXPROCS(0), runtime.NumCPU(),
os.Getenv("GOMAXPROCS"), os.Getenv("GODEBUG"))
fmt.Printf("cpu.max=%q cpu.weight=%q\n", readFile("cpu.max"), readFile("cpu.weight"))
}
// threadCPU returns the CPU time used by the current OS thread; the caller must call LockOSThread first.
func threadCPU() time.Duration {
var ru syscall.Rusage
const rusageThread = 1 // RUSAGE_THREAD
if err := syscall.Getrusage(rusageThread, &ru); err != nil {
panic(err)
}
return time.Duration(ru.Utime.Nano() + ru.Stime.Nano())
}
var sink uint64
// spinCPU uses d of CPU time on the current thread, then returns.
func spinCPU(d time.Duration) {
runtime.LockOSThread()
defer runtime.UnlockOSThread()
start := threadCPU()
x := uint64(1)
for threadCPU()-start < d {
for i := 0; i < 20000; i++ {
x ^= x*31 + uint64(i)
}
}
sink += x
}
func burst(args []string) {
fs := flag.NewFlagSet("burst", flag.ExitOnError)
threads := fs.Int("threads", 8, "number of threads working at once")
work := fs.Duration("work", 20*time.Millisecond, "CPU time each thread uses per round")
interval := fs.Duration("interval", time.Second, "interval between the starts of rounds")
rounds := fs.Int("rounds", 20, "number of rounds")
fs.Parse(args)
info()
fmt.Printf("burst threads=%d work=%s interval=%s rounds=%d\n", *threads, *work, *interval, *rounds)
fmt.Println("round wall_ms d_nr_periods d_nr_throttled d_throttled_ms d_usage_ms")
var walls []float64
begin := cpuStat()
t0 := time.Now()
next := t0
for r := 1; r <= *rounds; r++ {
time.Sleep(time.Until(next))
next = next.Add(*interval)
before := cpuStat()
start := time.Now()
var wg sync.WaitGroup
for i := 0; i < *threads; i++ {
wg.Add(1)
go func() { defer wg.Done(); spinCPU(*work) }()
}
wg.Wait()
wall := time.Since(start)
after := cpuStat()
walls = append(walls, ms(wall))
fmt.Printf("%d %.1f %d %d %.1f %.1f\n", r, ms(wall),
after["nr_periods"]-before["nr_periods"],
after["nr_throttled"]-before["nr_throttled"],
float64(after["throttled_usec"]-before["throttled_usec"])/1000,
float64(after["usage_usec"]-before["usage_usec"])/1000)
}
time.Sleep(time.Until(next))
elapsed := time.Since(t0)
end := cpuStat()
sort.Float64s(walls)
fmt.Printf("summary wall_ms p50=%.1f max=%.1f avg_cpu_cores=%.3f nr_periods=%d nr_throttled=%d throttled_ms=%.1f\n",
pct(walls, 50), walls[len(walls)-1],
float64(end["usage_usec"]-begin["usage_usec"])/float64(elapsed.Microseconds()),
end["nr_periods"]-begin["nr_periods"], end["nr_throttled"]-begin["nr_throttled"],
float64(end["throttled_usec"]-begin["throttled_usec"])/1000)
}
func spin(args []string) {
fs := flag.NewFlagSet("spin", flag.ExitOnError)
threads := fs.Int("threads", 1, "number of threads kept busy")
report := fs.Duration("report", 10*time.Second, "interval between usage reports")
fs.Parse(args)
info()
for i := 0; i < *threads; i++ {
go func() {
for {
spinCPU(time.Second)
}
}()
}
prev, prevT := cpuStat()["usage_usec"], time.Now()
for range time.Tick(*report) {
cur, now := cpuStat()["usage_usec"], time.Now()
fmt.Printf("%s usage_cores=%.3f\n", now.UTC().Format(time.RFC3339),
float64(cur-prev)/float64(now.Sub(prevT).Microseconds()))
prev, prevT = cur, now
}
}
// liveHeap gives the GC a set of live data with pointers to mark in every cycle.
type blob struct {
p *[8]byte
pad [48]byte
}
var liveHeap []*blob
func serve(args []string) {
fs := flag.NewFlagSet("serve", flag.ExitOnError)
addr := fs.String("addr", ":8080", "listen address")
work := fs.Duration("work", 5*time.Millisecond, "CPU work per request")
alloc := fs.Int("alloc", 1<<20, "bytes allocated and thrown away per request")
live := fs.Int("live-mb", 64, "size of the live heap kept in memory (MB)")
fs.Parse(args)
info()
for i := 0; i < *live<<20/64; i++ {
liveHeap = append(liveHeap, &blob{p: new([8]byte)})
}
fmt.Printf("serve addr=%s work=%s alloc=%d live_mb=%d\n", *addr, *work, *alloc, *live)
http.HandleFunc("/work", func(w http.ResponseWriter, r *http.Request) {
garbage := make([][]byte, 0, *alloc/4096+1)
for n := 0; n < *alloc; n += 4096 {
garbage = append(garbage, make([]byte, 4096))
}
spinCPU(*work)
sink += uint64(len(garbage))
io.WriteString(w, "ok\n")
})
http.HandleFunc("/ping", func(w http.ResponseWriter, r *http.Request) {
io.WriteString(w, "ok\n")
})
http.HandleFunc("/stats", func(w http.ResponseWriter, r *http.Request) {
var ms runtime.MemStats
runtime.ReadMemStats(&ms)
st := cpuStat()
st["gomaxprocs"] = int64(runtime.GOMAXPROCS(0))
st["num_gc"] = int64(ms.NumGC)
json.NewEncoder(w).Encode(st)
})
if err := http.ListenAndServe(*addr, nil); err != nil {
panic(err)
}
}
func fetchStats(c *http.Client, base string) map[string]int64 {
m := map[string]int64{}
resp, err := c.Get(base + "/stats")
if err != nil {
return m
}
defer resp.Body.Close()
json.NewDecoder(resp.Body).Decode(&m)
return m
}
func load(args []string) {
fs := flag.NewFlagSet("load", flag.ExitOnError)
target := fs.String("target", "http://cpulab:8080", "service address")
rps := fs.Float64("rps", 60, "average /work requests per second (Poisson arrivals)")
clump := fs.Int("clump", 1, "number of /work requests sent together on each arrival; the average rps stays the same")
pingRPS := fs.Float64("ping-rps", 0, "extra /ping requests per second (no work), with latency counted separately")
duration := fs.Duration("duration", 60*time.Second, "measurement time")
warmup := fs.Duration("warmup", 10*time.Second, "warmup time, not counted in the results")
label := fs.String("label", "", "result label")
fs.Parse(args)
c := &http.Client{
Timeout: 10 * time.Second,
Transport: &http.Transport{MaxIdleConnsPerHost: 1000, MaxConnsPerHost: 0},
}
// fire sends requests with Poisson arrivals, n per arrival; latency is measured from the scheduled send time.
fire := func(path string, rate float64, n int, d time.Duration, record func(time.Duration, error)) {
var wg sync.WaitGroup
end := time.Now().Add(d)
next := time.Now()
for next.Before(end) {
time.Sleep(time.Until(next))
sent := next
for i := 0; i < n; i++ {
wg.Add(1)
go func() {
defer wg.Done()
resp, err := c.Get(*target + path)
if err == nil {
io.Copy(io.Discard, resp.Body)
resp.Body.Close()
}
record(time.Since(sent), err)
}()
}
next = next.Add(time.Duration(rand.ExpFloat64() / (rate / float64(n)) * float64(time.Second)))
}
wg.Wait()
}
both := func(d time.Duration, work, ping func(time.Duration, error)) {
var wg sync.WaitGroup
wg.Add(1)
go func() { defer wg.Done(); fire("/work", *rps, *clump, d, work) }()
if *pingRPS > 0 {
wg.Add(1)
go func() { defer wg.Done(); fire("/ping", *pingRPS, 1, d, ping) }()
}
wg.Wait()
}
discard := func(time.Duration, error) {}
both(*warmup, discard, discard)
var mu sync.Mutex
var lat, pingLat []float64
var workErrs, pingErrs int
collect := func(dst *[]float64, errs *int) func(time.Duration, error) {
return func(d time.Duration, err error) {
mu.Lock()
defer mu.Unlock()
if err != nil {
*errs++
return
}
*dst = append(*dst, ms(d))
}
}
before := fetchStats(c, *target)
t0 := time.Now()
both(*duration, collect(&lat, &workErrs), collect(&pingLat, &pingErrs))
elapsed := time.Since(t0)
after := fetchStats(c, *target)
sort.Float64s(lat)
sort.Float64s(pingLat)
d := func(k string) int64 { return after[k] - before[k] }
fmt.Printf("load label=%s gomaxprocs=%d clump=%d requests=%d errors=%d achieved_rps=%.1f\n",
*label, after["gomaxprocs"], *clump, len(lat), workErrs, float64(len(lat))/elapsed.Seconds())
fmt.Printf("work_latency_ms p50=%.1f p90=%.1f p99=%.1f p999=%.1f max=%.1f\n",
pct(lat, 50), pct(lat, 90), pct(lat, 99), pct(lat, 99.9), pct(lat, 100))
if *pingRPS > 0 {
fmt.Printf("ping_latency_ms requests=%d errors=%d p50=%.1f p90=%.1f p99=%.1f max=%.1f\n",
len(pingLat), pingErrs, pct(pingLat, 50), pct(pingLat, 90), pct(pingLat, 99), pct(pingLat, 100))
}
fmt.Printf("server avg_cpu_cores=%.3f nr_periods=%d nr_throttled=%d throttled_ratio=%.3f throttled_ms=%.1f num_gc=%d\n",
float64(d("usage_usec"))/float64(elapsed.Microseconds()),
d("nr_periods"), d("nr_throttled"),
float64(d("nr_throttled"))/math.Max(1, float64(d("nr_periods"))),
float64(d("throttled_usec"))/1000, d("num_gc"))
}
func ms(d time.Duration) float64 { return float64(d.Microseconds()) / 1000 }
func pct(sorted []float64, p float64) float64 {
if len(sorted) == 0 {
return math.NaN()
}
i := int(math.Ceil(p/100*float64(len(sorted)))) - 1
if i < 0 {
i = 0
}
return sorted[i]
}