After you run
kubectl applyto create a Deployment, Kubernetes creates a ReplicaSet and Pods. If you later delete one of those Pods, a new one usually appears soon after. What is a controller, and how does it keep this process going?
0. What is a controller?
A Deployment is a Kubernetes resource object. It describes a target, such as which image to use and how many replicas to keep. A controller is a program that keeps watching related resources. It decides what to do based on the target and the state it can see.
When you change a Deployment with kubectl, the controller decides which resources to create or change. It sends requests to the API server. The API server handles those requests and stores the objects.
The Deployment Controller and ReplicaSet Controller in this article usually run inside kube-controller-manager, in the control plane. Creating ten Deployments does not start ten separate controller processes. One type of controller can handle many objects of that type.
On each check, a controller asks what still needs to change to reach the target. It then sends the required API requests. This repeated process of checking and acting is called reconcile. Official controller guide
When a Pod goes missing, who notices first? Does the controller keep asking the API server? If you change both replicas and the image, how does it decide what to do? Why do we need a ReplicaSet when the Deployment already knows the target count?
We can start with the work of the two controllers, then look at how each one makes decisions.
1. From Deployment to Pod: two controllers at work
The relationship is Deployment → ReplicaSet → Pod. Two controllers move this process forward.
Suppose we create hello-web for the first time. It asks for three replicas, and there are no existing objects to adopt:
- The Deployment Controller checks
hello-web. It finds no matching ReplicaSet, so it asks the API server to create one with the Pod template and replica target. - The ReplicaSet Controller observes that ReplicaSet and schedules a check. It sees a target of three and no Pods that belong to the ReplicaSet, so it requests new Pods.
- Later, it observes the new Pods. The next checks can then see the results.
In this diagram, dashed arrows mean watching a resource. Solid arrows mean changing a resource through the API. Both start at the controller doing the work. The diagram shows responsibilities, not a timeline.
The two controllers work together through resource state in the API. They do not call each other directly. A resource written by one controller becomes input for the other controller's next decision.
Information also flows back. Changes to a ReplicaSet's status help the Deployment Controller check the progress of an update. Each controller runs its own control loop, using the shared resource state. Deployment Controller · ReplicaSet Controller
2. Reconcile: decide from the state you see now
Suppose hello-web already has three Pods. You delete one. You do not run apply again, but a new Pod soon appears.
The work did not end when the apply request finished. The target stays in the cluster. On later checks, the controller decides what still needs to happen to meet it.
Creating three Pods once is not enough. We want the system to keep moving toward the target even when the current state changes. The controller control loop
To do this, each check must use the state it sees at that time.
A simple response to a Pod deletion might be: “When I receive a delete notification, create a new Pod.” But what if you change replicas from three to two while that notification is waiting? Creating another Pod would do too much.
Reconcile takes a different approach: When it is my turn, read the target and current state again. Then decide what needs to happen.
| Why check again? | Target and state seen on this check | What needs to change? |
|---|---|---|
| Replicas changed from 2 to 3 | Want 3, have 2 | Create 1 |
| One of 3 Pods was deleted | Want 3, have 2 | Create 1 |
| The last create request failed; this is a retry | Want 3, have 2 | Try to create 1 again |
| A Pod was deleted, and the controller has also observed the new target of 2 | Want 2, have 2 | Create none |
The first three cases can use the same logic. They all lead to one question: how many are still missing?
This diagram only shows the count check. The real code has more checks, including protection against repeated operations, which we will discuss later.
The key idea is: a notification means “check again.” It does not decide the action. Reconcile also does not have to finish all the work in one pass. It can send a request, observe the result later, and decide the next step.
3. Track changes with Watch and a local cache
How does a controller learn about changes to objects?
One option is to ask the API server every few seconds and compare Pod counts. But as the number of objects grows, fetching unchanged data again and again costs more work.
For the two controllers in this article, start with this model: get the existing objects, then receive later changes through Watch. Watch is a streaming connection from a client to the API server. When a related object is added, updated, or deleted, the change comes back through this connection. The client does not need to fetch all objects each time. We will look at where the API server gets these changes in Where do Watch notifications come from?.
When a change arrives, the first step is to update the local cache with the observed data. Later decisions can read that data.
client-go is the official Go client library for Kubernetes. It lets Go programs use the Kubernetes API. It also provides tools for controllers, such as caches, Informers, and work queues. The two controllers use these tools to observe resources and schedule work. client-go
An Informer keeps the observed object data up to date. A controller can use a Lister to read this cache. This puts the data updates in one place. Each reconcile pass does not need to fetch a whole set of objects from the API server.
There is a tradeoff: the local cache can be a little behind. The API server may have accepted a create request, but the controller may not see the new object on its next cache read. This delay matters when deciding whether to create another Pod. Informer consistency
These change notifications are also different from the Events shown by kubectl describe. A controller does not need a diagnostic Event saying “a Pod was deleted” before it can check again.
4. The queue holds objects to check
After updating the cache, the Informer sends notifications to event handlers registered by the controller. A handler is a callback function. It handles notifications about added, updated, or deleted objects and is part of the controller's code.
The handler finds the object affected by the change and adds it to a work queue. A worker is a loop that takes work from the queue and runs reconcile. The handler receives notifications and schedules checks. The worker makes the decisions for that pass.
In the v1.34.0 code used for this article, the controllers register these handlers:
| Controller | Object notifications | Object to check |
|---|---|---|
| Deployment Controller | Deployment added, updated, or deleted | That Deployment |
| Deployment Controller | ReplicaSet added, updated, or deleted | Related Deployment |
| Deployment Controller | A specific Pod deletion path | Deployment that meets the conditions |
| ReplicaSet Controller | ReplicaSet added, updated, or deleted | That ReplicaSet |
| ReplicaSet Controller | Pod added, updated, or deleted | Related ReplicaSet |
These are the registered notification types. Handlers still check ownership and other conditions. Not every notification adds work to the queue. For example, when a Pod is deleted, the ReplicaSet Controller's handler finds the ReplicaSet that controls it and schedules a check. Some handlers also update expectations, which we will cover next, to record that an expected change has been observed. Deployment handlers · ReplicaSet handlers
Imagine several Pods under the same ReplicaSet change within a short time. The handlers add that ReplicaSet to the queue for a worker to process. They do not run a full reconcile pass immediately for every notification.
The queue holds a namespace/name key, such as default/hello-web-abc123. It does not hold an action such as “create one Pod.”
When a Pod is deleted, the queued work means “check this ReplicaSet again.” When a worker takes the key, it reads the target and Pods from the cache. Only then does it decide whether to create a Pod.
The client-go queue combines duplicate keys that are waiting to be processed. If a related change arrives during processing, the key can be queued again after that pass. There is no one-to-one link between the number of notifications and the number of Pods created. Work queue code
If a create request fails, there is no new Pod to observe. A retry is needed to keep making progress. Failed work can be queued again with backoff, so an API problem does not lead to a flood of requests. Some decisions also schedule a later check. Watch is important, but it is not the only reason to reconcile again.
5. If the cache is behind, how do we avoid creating too many Pods?
The queue combines duplicate keys, but another problem remains: a request may already have been sent, while the cache has not caught up.
Suppose a ReplicaSet wants three Pods and the cache shows two. The controller successfully requests a third Pod. Before that Pod appears in the cache, another check begins. It still sees two.
If the only rule were “three minus two means create one,” it could create too many.
The ReplicaSet Controller uses expectations to handle this. It first records the creates or deletes it expects to observe. It waits for the related notifications before making more count changes. It also handles failures and checks for expired expectations, so the record does not block progress forever. Expiration alone does not guarantee that a new pass starts at that exact moment.
Comparing the target with the observed state is the core logic. Other mechanisms are needed to handle delays in an asynchronous system. Replica count and expectations code
Also, the active Pod count used by a ReplicaSet is not the Ready count. A new Pod that is not Ready does not mean a replica is missing. The controller should not create an extra Pod just for that reason. Available replicas matter when deciding whether old replicas can be reduced during an update.
6. Rolling updates make the two roles clearer
Developers still write controller behavior. They just do not need a separate script for every possible combination of events. A ReplicaSet maintains one group of replicas. A Deployment manages the targets of those groups and the update strategy.
Changing a Deployment's replicas affects ReplicaSet targets. Changing its Pod template, such as the image, starts a rollout involving the ReplicaSet for that version. It does not simply apply field differences to existing Pods. Deployment updates
If the only task were “keep three Pods,” we could imagine one controller doing all the work. A rolling update asks more questions: how many v2 Pods should we add now? Can we remove a v1 Pod before v2 is available?
The Deployment Controller coordinates these steps. It can write its decision as two ReplicaSet targets: keep two in the old group, and ask for one in the new group.
For replicas: 3, maxSurge: 1, and maxUnavailable: 0, with no extra failures, we can use these stages to explain the idea:
| Stage | v1 target | v2 target | Reason |
|---|---|---|---|
| Before the update | 3 | 0 | v1 serves the workload |
| Add one v2 | 3 | 1 | Use the one extra replica allowed, then wait for v2 to become available |
| After v2 is available | 2 | 1 | Remove one v1 and continue the update |
| Update complete | 0 | 3 | v2 serves the whole workload |
These are selected snapshots of target counts. Steps in the middle are left out. A real update may not stop at each row, and terminating Pods may still be present for a while.
Now suppose a v1 Pod is deleted halfway through the update. The ReplicaSet Controller does not need to understand the whole rollout plan. It checks the current target of the v1 ReplicaSet. If that target is now two, it does not keep trying to return to the original three.
The roles are now clearer:
- The Deployment Controller decides how many replicas each version should have now.
- The ReplicaSet Controller moves each group toward its own target.
The Deployment still cares about replica counts. It moves the update forward by changing targets at the next level. This explains the current design; it does not mean the two roles could never be implemented together. Rolling update code
7. What if you delete a Pod and immediately reduce replicas?
Return to hello-web, with no version update in progress. The target is three. You delete a Pod, then immediately change the Deployment's replicas to two.
Will Kubernetes create a third Pod first?
Either result is possible. It depends on what the controllers have observed when they act.
If the ReplicaSet Controller still sees a target of three when it handles the deletion, it may create a Pod first. Later, the Deployment Controller lowers the ReplicaSet target to two, and the ReplicaSet Controller removes the extra replica.
If the new target has already reached the ReplicaSet and its cache, it sees a target of two and two existing Pods. It does not need to create one.
Several control loops work together through the API and caches, so these intermediate states can happen. As long as later requests succeed and changes can be observed, the loops can keep moving toward the new target.
The next time a deleted Pod comes back, you can ask: which object was queued? What target did this pass read? Has the cache caught up with the earlier create request?
These questions help explain more than “if a Pod is missing, create one.”
More detail: API server Watch, reconnects, and resync
The main flow is: changes update the cache, related objects enter the queue, and a worker decides what to do from the state it reads. We can now look more closely at where notifications come from and how observation recovers.
Where do API server Watch notifications come from?
Take a Deployment change from three replicas to two. For a built-in resource with the watch cache enabled, the usual path is:
etcd provides its own Watch mechanism. The API server's storage layer uses it to receive data changes, then turns them into Kubernetes resource notifications. It does not need to compare two YAML files field by field to learn that the data changed. etcd Watch code
If the Deployment continues to match the watch filter, the update notification is MODIFIED and includes the updated object. This example only shows the message shape; other fields are left out:
{
"type": "MODIFIED",
"object": {
"metadata": {
"name": "hello-web",
"resourceVersion": "..."
},
"spec": {
"replicas": 2
}
}
}
The message does not say “replicas went down by one; delete a Pod.” The controller still uses the meaning of the resource and the state it observes to decide what to do.
The API server's watch cache holds current object state and a limited history of recent changes. It helps serve many watchers and resume notifications from a resourceVersion while that history is available. This cache is separate from the Informer's local cache inside a controller. The first serves API read and watch requests. The second supplies data for controller decisions. Watch cache code
When sending notifications, the API server may still check both the old and new objects. For example, a label change may cause an object to enter or leave a watcher's filter. That affects the event type sent to that watcher. This is a separate step from learning about the data change through etcd. Watch event conversion
How does Watch recover after a disconnect?
The client tries to reconnect and resume using resourceVersion. If the required history is no longer available, it gets the current state again and continues watching. This is often explained as List/Watch; initialization may also use a streaming list. The goal is to recover a view of the state, not to promise that every notification is kept forever. Tracking API changes
Resync is not the same as a new List
When resync is enabled, the Informer sends objects already in its cache to handlers again. This gives the controller another chance to check them. It does not mean the API server has new changes, and it does not mean fetching all objects from the server again.
Retry limits, delays, and whether resync is used depend on the code and configuration. They are not one fixed timer shared by all controllers.
No cluster experiments were run for this article. The scenarios and count tables explain concepts; they are not measured timelines. Code details were checked against Kubernetes v1.34.0 and client-go v0.34.0. Each diagram shows only the relationships needed for its section, not every handler. Retry and resync policies also differ between controllers.
See How Kubernetes works: from the API to containers for the series guide and planned topics.