Self-healing, one step at a time

You don't tell Kubernetes what to do. You tell it what you want: a Deployment with replicas: 3.

The scheduler places each on a ship with room, spreading them out so one bad ship can't sink everything.

Disaster! The middle ship crashes and goes down with its . Now 2 are running, but you wanted 3.

The ReplicaSet's clerk counts heads: want 3, have 2. It creates a brand-new (new name, new IP address) and the scheduler puts it on a ship still afloat.

This count-and-fix cycle is the reconciliation loop, and it's the entire personality of Kubernetes. It works the same way for every kind of trouble: a deleted , a lost ship, or you changing replicas from 3 to 5. A container that merely crashes is handled closer to home: the kubelet restarts it in place, in the same , and the RESTARTS column goes up.

The replacement isn't the old swimming over. It's a fresh copy with a new name and a new IP address. Anything the old one kept in memory or on its own disk is gone (chapter 5 fixes that). And it causes a problem: how do customers find a whose address keeps changing?