Kubernetes in Production: 15 Lessons From the On-Call
Nine years of Golang on AWS taught me one discipline: declare what you want, watch what converges, plus price the gap. Each lesson starts from a real production failure, states its assumptions up front, plus ends in math you can re-run. The bill hides in the boundary.
TL;DR: Fifteen forensic Kubernetes lessons with live diagrams, from control planes to GPU scheduling plus cost discipline. Built from production postmortems with calculators for the math.
By Mukul Kumar Mishra · Backend plus SRE course · Updated September 18, 2026
What you will be able to do
- Run etcd-backed control planes with bounded inputs, staggered deploys, plus convergence windows you measured.
- Size HPA replicas plus connection pools from RPS math, then right-size fleets instead of feeding them.
- Design zones plus cells that fail locally, tier queues that absorb spikes, plus retries that spare the mesh.
- Observe with symptom alerts plus runbooks, schedule GPUs by scarcity, manage secrets plus state, then price the month.
- Answer platform interviews with clarified scope, out-loud estimates, plus three ways you would break your own design.
How interviews test this course: scale before pods, failure domains before features, price before praise. Lesson 14 rehearses the performance.
The Cluster Is a Promise
Silent TLS connections arrive carrying nothing, then never leave. Memory climbs while the API answers slower, then stops answering at all. The workloads are healthy. The brain is full. Assumption for this lesson: every component below trusts the control plane to answer, so price what happens when it cannot.
- etcd is the cluster brain. Bound its inputs with handshake deadlines plus memory limits before trusting it.
- Unbounded waits turn connection counters into memory counters. Every wait needs a ceiling.
- Deep dive: etcd Handshake That Never Ends, plus the pool sizer for the same math.
Count your exposure with the pool sizer: peak connections times hold time equals state your servers must carry. The calculator prices the honest case. The attacker prices the missing bound.
Q1. Silent connections exhaust etcd memory. Walk through the fix live. Seen at: infra plus security loops.
Q2. Where does the control plane keep state plus what happens when it fills? Seen at: platform loops.
Brain guarded. Now the workloads it was promised. Lesson 02: Declare converge →
Declare Then Converge
A routine rolling deploy restarts 120 consumers one by one. Each restart rejoins the group. Each rejoin triggers a rebalance across roughly 40,000 partitions. The group never settles for 11 modeled minutes. Assumption: declared state plus converged state differ during every rollout, so measure the convergence window, not just the manifest.
- Stagger restarts plus prefer cooperative protocols. Eager rebalancing pauses everything the deploy touches.
- GitOps makes the desired state visible. Static membership plus canary steps make convergence survivable.
- Deep dive: Kafka Rebalance Storm, coordinator math across 40,000 partitions.
The deploy that touches one consumer at a time paused all of them continuously. Three defaults, each reasonable alone, composed into an outage while every broker dashboard stayed green. Convergence is the feature. The pause is the price.
Q1. A rolling deploy pauses all consumers. Walk through the fix. Seen at: platform plus SRE loops.
Q2. Eager versus cooperative rebalancing: when does each hurt? Seen at: messaging infra loops.
Convergence priced. Now the replicas it converges toward. Lesson 03: Requests replicas →
Count Requests, Then Replicas
Trillions of messages cross 177 database nodes that a leaner fleet of 72 could carry. Hot partitions plus quorum amplification plus JVM pauses kept the larger fleet busy looking busy. Assumption: desired replicas equal current replicas times utilization over target, so autoscaling without utilization targets is decoration.
- Size from utilization, not anxiety. Try the HPA sizer with your own live CPU numbers.
- Right-size before scaling out. Hot partitions need key redesign, not more nodes.
- Deep dive: Discord Trillions of Messages, 177 nodes down to 72.
Junior Kubernetes is a cost generator. Senior Kubernetes is a cost controller. The fleet that looks busy is not the fleet that is needed. Measure, then subtract.
Q1. Size replicas for tripled evening traffic with HPA math out loud. Seen at: Google-style loops.
Q2. When does autoscaling make an outage worse? Seen at: AWS-style plus startup loops.
Fleet trimmed. Now the doorway every request queues at. Lesson 04: Pool door →
Pool at the Door
Two million databases hide behind one pooling promise. Each backend forks a process per connection, so ten thousand friends share one door through a pooler holding 675. Assumption: pool size equals peak RPS times p99 seconds times a safety factor. Vibes do not fork processes. Physics does.
- Pool size is math, not vibes. Try the pool sizer with your own RPS plus pod counts.
- Serverless relocates caps into docs nobody reads. Read the cap before traffic does.
- Deep dive: Serverless Postgres, the pooling promise with receipts.
Watch for the pooler becoming the bottleneck itself. One proxy in front of the database is a single point with a nice dashboard. Pair it, or shard past it, before traffic proves the point.
Q1. Size the pool for 50k RPS at 20ms p99 across 40 pods. Seen at: Google-style plus Uber-style loops.
Q2. Connections exhaust at 3 AM with flat traffic. What leaked? Seen at: SRE loops.
Door managed. Now the building the door stands in. Lesson 05: Zones unit →
Zones Are the Unit
Cooling fails first. Power follows. One zone goes dark for hours while snapshots decide how much history survives. Assumption: the blast radius you plan for is the blast radius you get, so multi-AZ is a design input, not a checkbox.
- Keep traffic live in all zones. Cold standby rots while warm standby bills honestly.
- Test restores, not backups. Price recovery with the downtime calculator before the heat arrives.
- Deep dive: AWS Thermal Outage, cooling failed first.
Game days that kill one zone prove the other two carry load. Untested failover is a rumor with runbooks. Run the rumor until it becomes a drill.
Q1. One zone dies during peak checkout. Now what? Seen at: Amazon-style loops.
Q2. Price four hours of single-zone impairment out loud. Seen at: FinOps-flavored loops.
Zones survived. Now shrink the blast further. Lesson 06: Cells blast →
Cells Bound the Blast
One fault cancels two thousand flights across three days. The computers recover Tuesday. The schedules recover Friday. One unmirrored component gated a whole sky. Assumption: global shared fate fails globally, so cells with independent control planes fail locally by construction.
- Split the control plane into autonomous Raft-backed cells. One dark cell is an incident. One dark region is a disaster.
- Route around dark cells automatically. Manual failover at 3 AM is a second outage wearing a pager.
- Deep dives: Cell-Based Control Planes plus NATS Four Hours.
Cells cost duplication up front plus save multiplication later. The spreadsheet that refuses cells is pricing Tuesday while Friday sends the invoice.
Q1. Design failover customers never notice across three cells. Seen at: Google-style plus Netflix-style loops.
Q2. Cells versus zones: when does each contain what? Seen at: principal-level loops.
Blast bounded. Now the traffic the cells absorb. Lesson 07: Queues spikes →
Queues Absorb Spikes
Eighteen trillion messages a day cross one platform. Half of them could have waited. Ten thousand UPI transactions per second cannot. Assumption: urgent plus patient traffic share nothing but the logo, so tier topics before counting partitions.
- Separate urgent from patient traffic or one surge starves everything the business promises.
- Partitions buy concurrency plus sell broker load. Count them like money because brokers do.
- Deep dives: Roblox 18 Trillion plus Razorpay Kafka.
Dead letters deserve readers. Every message in the red box is a bug report from production. Teams that review them weekly find failures before customers file them.
Q1. Design ingestion for one million events per second without losing any. Seen at: Uber-style plus LinkedIn-style loops.
Q2. One lane floods. Prove the other lane stays dry. Seen at: messaging infra loops.
Spikes absorbed. Now the retries inside the mesh. Lesson 08: Retries mesh →
Retries Spare the Mesh
Agents retried failed deploys until retries cost more than features. Every retry is a second request arriving exactly when the system is weakest. Assumption: the same amplifier lives in your service mesh, so bounded retries with backoff are a capacity plan, not a courtesy.
- Exponential backoff plus jitter plus a hard stop. Miss any of the three plus retries amplify.
- Circuit breakers convert repeated failure into fast failure. Fast failure is a feature.
- Deep dives: Replit Retries plus the retry storm calculator.
Unbounded retries are a self-inflicted DDoS with a deploy log. Budget the remaining retries like any other load. The mesh forgives planned load. It punishes surprise load with interest.
Q1. Convince me your mesh retries cannot cascade. Seen at: Amazon-style loops.
Q2. Design a breaker with closed plus open plus half-open states. Seen at: Netflix-style loops.
Mesh spared. Now the senses that watch it. Lesson 09: Seeing cluster →
Seeing the Cluster
The signals existed the whole time. Nobody had drawn the dashboard that would have shown them. Detection lag set the price while the sky stayed grounded. Assumption: alert on symptoms users feel, not causes you guess, so latency plus error rate lead while CPU watches from the back.
- Sample smart. Keep every error, sample the boring middle with the sampling fitter plus burn budgets with the error budget calculator.
- Attach a runbook to every page. Heroics do not scale. Checklists do.
- Deep dive: NATS Recovery Math, where detection lag set the price.
Monitoring is not paperwork. It is the only sense organ production has. Blasts you cannot see, you cannot survive, which is why observability closes the resilience arc.
Q1. p99 triples at 2 AM with no deploy. Walk through it live. Seen at: Google-style plus Meta-style SRE loops.
Q2. Design dashboards for a payments API before launch. Seen at: Stripe-style plus fintech loops.
Senses wired. Now the workload that changed the schedule. Lesson 10: GPU schedule →
GPUs Change the Schedule
Every question gets paid for twice, retrieval plus inference, at thirty-one cents of model time against two cents of lookup. The cluster that served web traffic now serves spiky inference with expensive accelerators attached. Assumption: GPU time dominates the envelope, so scheduling plus caching decisions outrank tuning.
- Schedule GPUs as the scarce resource. Bin-pack inference plus autoscale on queue depth, not CPU.
- Cache repeated inference first. Every repeated answer should be nearly free. Try the RPS envelope calculator.
- Deep dive: Perplexity Answer Engine, the twice-paid question.
AI did not reduce the need for operators. It raised the bar on what one is worth. Spiky inference on shared clusters punishes every lesson this course skipped, with GPU pricing attached.
Q1. Schedule inference on a shared cluster with spiky load. Seen at: AI infra loops.
Q2. GPU queue grows while CPUs idle. What does the scheduler not know? Seen at: platform loops.
Schedule set. Now the invoice that grades it. Lesson 11: Bill design →
The Bill Is the Design
A preview button reinstalls the internet on every click, near a billion tokens a minute that a cache could have spared. Reviews put the waste near $2.6M a month in modeled estimates. Assumption: platform conveniences compound into invoices, so price the golden path before paving it.
- Cache build layers plus preview environments aggressively. Repeated installs are a subscription to nothing.
- Make the golden path the cheapest path. Developers follow cost gradients faster than mandates.
- Deep dive: Lovable Preview Button, the $2.6M modeled lesson.
Revisit the bill quarterly. Traffic shapes drift, prices change, plus yesterday optimization rots. The platform team that prices its own conveniences earns the right to mandate them.
Q1. Price a preview environment per developer monthly, out loud. Seen at: platform plus startup loops.
Q2. Cut platform spend 30 percent without touching velocity. Seen at: FinOps-flavored loops.
Bill read. Now the secrets riding every pod. Lesson 12: Secrets ride →
Secrets Ride Along
A candidate uploads a CV. The summary comes back with API keys inside a rendered image. No clicks. No malware. One poisoned document plus every permission the key carried. Assumption: secrets mounted into pods are one exfiltration from exposure, so scope narrows plus TTLs shorten in direct proportion to blast radius.
- Issue secrets per pod with one permission set plus short TTLs. Stolen keys should die alone.
- Keep secrets out of images plus logs plus summaries. Three leak paths, all closed by construction.
- Deep dive: AgentFlayer Zero-Click, the upload that downloaded keys.
Lesson 4 of the foundation course taught this for APIs. Clusters multiply the lesson across every pod. External secret operators plus rotation plus audit complete the pattern. The vault is the only pod allowed to know everything.
Q1. Design secret management for 200 microservices on one cluster. Seen at: platform plus fintech loops.
Q2. A key leaks from a pod log at midnight. Your first hour? Seen at: SRE plus security loops.
Secrets scoped. Now the state the pods refuse to hold. Lesson 13: State outside →
State Lives Outside
A database copies every row plus loses customer data anyway. The backup skips encoding metadata, then garbage collection deletes what the backup forgot. The backup succeeds. The restore testifies against it. Assumption: pods are cattle, so state lives in volumes with snapshots plus restore drills, never in container layers.
- Replicas protect uptime. Only verified restores protect data. Schedule the drill, not just the snapshot.
- Garbage collection must never delete what it cannot prove dead. Retention policies are data architecture.
- Deep dive: Doltgres Backup That Forgot, metadata missed plus references erased.
Operators encode this discipline: custom resources describe the desired data state, controllers converge toward it, plus backups verify themselves. Stateful on Kubernetes works when the state never trusts the pod lifecycle.
Q1. Design backup plus restore for 10 TB of stateful data on Kubernetes. Seen at: AWS-style plus fintech loops.
Q2. A volume snapshot restores empty. What was never tested? Seen at: SRE loops.
State secured. Now turn all of it into interview answers. Lesson 14: Interview cluster →
The Interview Cluster
Pastebin in 45 minutes hires. A cluster debug in 45 minutes hires faster, because fewer candidates have owned a real blast radius. The panel is not testing kubectl trivia. It tests whether you ask about scale before drawing pods, name failure domains before features, plus price the month before praising the design.
- Clarify scope first. One service on EKS is a weekend. Ten thousand pods across regions is lessons 1 through 13.
- Estimate out loud with the HPA sizer plus downtime calculator. Panels score the math.
- Close by breaking it. Name three failures plus their fixes before they ask. That is the hire signal.
Rehearse three answers this week: an HPA sizing with live numbers, a zone failure with recovery math, plus a bill defense with the envelope. Bring one incident story of your own. Borrowed failures teach. Owned failures hire.
Q1. Design checkout for 100k RPS on Kubernetes, then break it three ways. Seen at: Amazon-style plus Google-style loops.
Q2. Tell me about a production incident you owned end to end. Seen at: every senior loop.
Answers rehearsed. Now run one service for real. Lesson 15: Run service →
Run One Service
Take one service from your own work. Declare it in Git, bound its brain inputs, stagger its deploys, size its pool, spread it across zones, tier its queues, cap its retries, alert on its symptoms, plus price its month. Then page yourself on purpose with a game day and watch the drills hold.
- Assemble the eleven artifacts from lessons 1 through 11 into one service folder.
- Run one game day before calling it done. Kill a zone, watch the other two carry load, write the postmortem.
- Continue with the foundation course: System Design: 15 Lessons From Real Outages.
Course complete. Keep the cluster honest: system design course plus all courses and all case studies. New failures arrive weekly. So will new lessons.
Bring one incident story of your own. Interviewers remember operators who owned the blast radius plus priced it. This course gave you twelve borrowed ones. Earn one.
Sources plus method
This course teaches from published Buildopsy postmortems linked under each lesson: etcd handshake limits, rebalance cascades, Discord right-sizing, Supabase pooling, AWS thermal recovery, cell design, NATS blast radius, Kafka tiers, UPI payments, retry economics, detection math, inference bills, plus preview environment waste. Incident details follow those accounts. Traffic plus cost figures are modeled estimates with assumptions stated per lesson.
