All posts

How Kubernetes works: from the API to containers

For readers who know basic Kubernetes operations and want to understand controllers, scheduling, resource management, and containers.

This series is for people who already know basic Kubernetes operations and want to understand how it works inside. If you can use kubectl to deploy a Deployment, check Pod status, and read logs, this is a place to take the next step.

We will follow what happens after you submit a resource: who makes decisions, who does the work, and how those decisions reach Linux containers. The aim is to connect the roles of the different components, so familiar commands make more sense.

Start with controllers

How does a Deployment become Pods? Reconcile through two controllers starts with a question: why does a deleted Pod come back? It explains the roles of Deployment and ReplicaSet controllers, then follows Watch, Informers, work queues, and cache delays.

The article stops at creating Pod objects. After reading it, try to answer: how does a controller learn about changes? Why does it check the state again after a notification? How do two controllers work together through resource objects?

Then scheduling

Where does a new Pod run? How the Scheduler decides and hands off picks up after the Pod object is created. It starts from a new replica that stays Pending while the CPU sits idle, walks through filtering, scoring, and binding, then comes back to decide what to change first. If you want to try it yourself, Hands-on: reproduce an unschedulable Pod in kind builds the same state on your machine and tests each fix.

After reading it, try to answer: what roles do requests and actual usage play in scheduling? What do Filtering and Scoring each decide? When a Pod cannot be scheduled, which changes only get it a Node, and which actually solve the problem?

Then CPU limits

CPU is only 30% used, so why is it throttled? Inside the 100ms quota of a CPU limit picks up after the Pod is placed on a Node. It starts from high latency while the dashboard shows an idle CPU, looks at the cgroup settings that CPU limits and requests become on the Node and how the kernel enforces them, then compares the effects and costs of several fixes. If you want to try it yourself, Hands-on: reproduce "only 30% CPU, yet throttled" in kind builds the same state on your machine and measures adding replicas, raising the limit, and what changes on a busy Node.

After reading it, try to answer: why can low average usage and throttling both be true? How should you read the throttling metrics? Why doesn't raising only the limit always help? For a sudden burst, what does adding replicas cost, and what does raising the limit cost?

What comes next?

These questions will guide later articles. Each topic still needs research and discussion. The topics below are planned, not finished articles. Completed posts will appear in the list at the end.

TopicQuestions to explore
Memory and resource managementHow do memory requests and limits reach cgroups? When usage goes over the limit, why is the container killed instead of slowed down like CPU?
OOM, eviction, and preemptionWho triggers OOMKilled, node-pressure eviction, and preemption? What roles do QoS, Priority, and PDBs play?
From kubelet to a containerWhat does kubelet do after scheduling? What are the roles of CRI, containerd, and runc? Where do CNI and CSI take part?
Pods and Linux isolationWhat problems do namespaces and cgroups solve? What is a Pod sandbox? What do containers in the same Pod share?
Restarting and replacingHow does a container restart differ from replacing a Pod? What can object identity and runtime state tell us?

These ideas also help with troubleshooting. When you see Pending, ContainerCreating, CPU delays, or memory problems, first identify the component responsible. Then look for evidence that supports your explanation.

Articles in this series

Read more posts