Skip to content
All posts
AKS Day-2 OpsJun 11, 20262 min read

The First Three Dashboards Every AKS Cluster Needs

You can drown a new AKS cluster in metrics on day one. Resist. Three dashboards answer the questions you'll actually have at 2 a.m. Start there, add the rest when you miss them.

A fresh AKS cluster wired to Managed Prometheus and Grafana can show you ten thousand metrics by lunch. That's the trap: more panels feel like more observability, and instead you get a wall of graphs nobody reads when it actually matters. Start with three dashboards, the ones that answer the questions you'll genuinely have when something's wrong. Add more only when you catch yourself missing them.

1. Cluster and node capacity. Am I running out of room? Node CPU and memory allocation versus capacity, node count against autoscaler limits, and memory pressure. This is the first thing you want during a scale event or a creeping slowdown. It tells you whether the cluster itself is the constraint before you go hunting through individual pods.

2. Workload health. Is anything actually broken right now? Pod restart counts, pods stuck pending, crash-looping containers, readiness across deployments. When someone says "the app is down," this dashboard says where and how. Pending pods point at capacity, restarts point at the workload, and you've saved yourself twenty minutes of guessing.

3. Requests versus actual usage. Are my numbers honest? Resource requests and limits laid against real consumption, per workload. This is the one most teams skip and the one that quietly pays for itself: it's how you right-size requests (which is how your autoscaler and your bill behave), and it surfaces the pod that requested 2 cores to use a tenth of one.

Notice the through-line: each dashboard answers a question, not a metric category. Capacity, health, honesty. That's most of what you need to operate AKS day to day, and three readable dashboards beat thirty you scroll past.

The instinct to instrument everything on day one comes from a good place and ends in noise. Build the three that map to your real 2 a.m. questions, learn what they don't tell you from actual incidents, and grow your observability from genuine gaps instead of from a dashboard library. Signal first. Volume later, if ever.

Want this distilled for your stack? Ask the assistant for the key takeaways or related reading.

Keep reading

Related notes

The Clarity Brief

Every other Tuesday · unsubscribe anytime

A short, high-signal briefing on architecture, AI, observability, and engineering leadership, written to make hard things clear.

Topics you care about

We respect your inbox. Your email is used only to send the newsletter and is never sold or shared.

Sathpal-OS