A developer reports high latency, but the CPU on the dashboard looks idle. Follow the 100ms cgroup quota to see how a CPU limit is enforced on the Node, and what to do about it.
Read articleRead cpu.max and cpu.stat in a throwaway local cluster, build a service that gets throttled by groups of requests, then measure GOMAXPROCS, adding replicas, raising the limit, cpu.weight, and what changes on a busy Node.
Read articleA new replica stays Pending while the CPU sits idle. Break down the Scheduler's filtering, scoring, and assignment to find what is blocking it.
Read articleBuild the "idle CPU, but the new replica is Pending" state in a throwaway local cluster, read FailedScheduling, then see where each fix sends the Pod.
Read articleFollow the work of two controllers, from Watch and local caches to work queues and rolling updates.
Read articleFor readers who know basic Kubernetes operations and want to understand controllers, scheduling, resource management, and containers.
Read articleA node startup failure, a config-source mismatch, and the checks that helped us find it.
Read article