Kubernetes ships a new minor version every few months, and AKS supports only a
rolling window of the most recent ones. Your cluster's version is therefore
a slowly-expiring asset. Ignore it long enough and one day you're running a
version Azure no longer supports: no security backports, no help on the path you
care about, under your production traffic. Nobody decides to be there. Everyone
who's there got there by not deciding.
The cruelty is in how it compounds. Stay current and each upgrade is one small,
boring hop. Fall three versions behind and you can't leap straight to current.
You must walk the hops in sequence, each with its own deprecations and surprises,
and you're now doing all of them under time pressure because support already
lapsed. One deferred upgrade quietly becomes a forced multi-upgrade weekend.
The fix isn't heroic, just deliberate:
- Know where you are relative to the window. Make "how many versions behind
are we" a number someone watches, not a thing you rediscover during an audit.
- Upgrade regularly, in small hops. A predictable cadence (test in non-prod,
then prod) keeps every jump small and dull. Small and dull is the goal.
- Pick an automation posture. Auto-upgrade channels keep you current with
less toil; manual keeps you in control if you stay vigilant. Both work.
Drifting with no posture is the one that bites.
This is the most preventable incident in all of AKS. It announces itself months
in advance, on a public schedule, with plenty of warning. The teams who get
caught aren't unlucky. They just treated "upgrade the cluster" as something with
no deadline, right up until it had a very urgent one.