SRE and Observability in Production: 15 Lessons From the On-Call
Nine years of Golang on AWS taught me one discipline: measure recovery in user-minutes, budget every signal, plus drill the fire before it schedules itself. Each lesson starts from a real production outage, states its assumptions up front, plus ends in math you can re-run. The bill hides in the boundary.
TL;DR: Fifteen forensic SRE lessons with live diagrams, from SLOs to burn rates plus tracing bills plus game days. Built from production postmortems with calculators for the math.
By Mukul Kumar Mishra · Backend plus SRE course · Updated September 26, 2026
What you will be able to do
- Set SLOs the business signs, alert on burn rate instead of symptoms, plus keep the phone quiet with two windows.
- Budget tracing plus logs plus metrics as three tiers, staff on-call that holds humans, plus command incidents with calm protocol.
- Write blameless postmortems that change topology, plan capacity with Little Law, plus rehearse fires as quarterly game days.
- Publish honest status, price paging as payroll, ship canaries with auto-rollback, plus total the telemetry month.
- Answer SRE interviews with SLO math out loud, timelines with owners, plus one debugging war story with logs.
How interviews test this course: SLO math out loud first, incident stories with timelines second, tradeoff pricing third. Lesson 14 rehearses the bench.
SLOs Are the Contract
Assumption first: 100 percent uptime is a rumor with a dashboard. A 99.9 percent SLO over thirty days allows 43.2 minutes of downtime. That allowance is a budget to spend on velocity. Define SLIs users feel, set SLOs the business signs, plus hold error budgets that freeze deploys when exhausted. Reliability without numbers is negotiation without prices.
- Measure what users feel: success rate plus latency tails. Server CPU never felt anything.
- Freeze feature deploys when the budget burns out. The freeze is the feature.
- Deep dive: NATS Four Hours, schedule-hours versus system-minutes.
Q1. What SLO fits a checkout API versus an internal dashboard? Seen at: SRE plus backend loops.
Q2. Your budget hits zero Thursday. What ships Friday? Seen at: observability plus platform loops.
Contract numbered. Now alert on burn. Lesson 02: Burn →
Burn-Rate Alerting
Threshold alerts page on symptoms. Burn-rate alerts page on budget consumption speed. A fast-burn alert fires when 2 percent of budget vanishes in an hour. A slow-burn alert tickets when the 30-day trend drifts. Two windows, two severities, one quiet phone. Alert fatigue ends where burn math begins.
- Page on 14.4 times burn over 1 hour plus 6 times over 6 hours. Textbook windows, tuned per service.
- Ticket slow burns to the backlog with a date. Slow pages train teams to ignore fast ones.
- Deep dive: AWS Thermal Postmortem, cooling failed first plus pages followed.
Q1. Size burn windows for a 99.95 percent SLO. Show the math. Seen at: SRE plus backend loops.
Q2. Your fast-burn fires during a deploy. Page or ticket? Seen at: observability plus platform loops.
Burn windowed. Now price the telemetry. Lesson 03: Tracing →
Tracing Economics
Full-fidelity tracing at 1M RPS is a second infrastructure bill wearing a trench coat. Head-based sampling decides at request start with constant cost. Tail-based sampling keeps errors plus slow traces after the fact. Sample the boring 99 percent at 1 percent, keep 100 percent of errors, plus watch the Sampling Fit calculator prove the invoice.
- Run head sampling at 1 percent for steady traffic. Flat cost, honest aggregates.
- Add tail rules for errors plus p99 latencies. Rare events deserve full fidelity.
- Deep dive: MoEngage Millisecond, write-heavy economics, priced.
Q1. Price full tracing at your RPS with modeled per-span cost. What dominates? Seen at: SRE plus backend loops.
Q2. When does tail sampling lie to you? Seen at: observability plus platform loops.
Traces budgeted. Now split the signals. Lesson 04: Signals →
Logs Versus Metrics Versus Traces
Logs explain single events at high per-GB cost. Metrics aggregate behavior at low per-series cost. Traces connect hops at sampling-dependent cost. Log errors with context, metric everything with cardinality caps, trace across boundaries. The team that logs request bodies at 1M RPS funds its vendor second yacht.
- Cap metric cardinality: no user IDs, no request IDs in labels. Cardinality is the silent invoice.
- Log structured JSON with request IDs plus error codes. Grep-friendly logs are incident-friendly logs.
- Deep dive: Cloudflare Loaded the Wrong File, when telemetry must explain a wrong file.
Q1. Your logging bill triples. What do you cut first? Seen at: SRE plus backend loops.
Q2. A debug needs one request path. Which signal answers? Seen at: observability plus platform loops.
Signals split. Now staff the night. Lesson 05: Oncall →
On-Call That Holds
On-call fails through fatigue, not through incidents. Follow-the-sun rotations with 6-person minimums keep any human under 25 percent night load. Compensation per page plus post-incident rest days price the burden honestly. The rota that never pages is designed. The hero that never sleeps quits.
- Require 6 humans per rotation with handoff notes. Five means holidays break the math.
- Pay per page plus grant rest after nights. Unpriced burden compounds into resignations.
- Deep dive: Simultaneous AI Outage, shared fate plus shared nights.
Q1. Staff a rota for a 4-person team covering 24/7. What breaks? Seen at: SRE plus backend loops.
Q2. A teammate burns out mid-quarter. What changes first? Seen at: observability plus platform loops.
Nights staffed. Now command the incident. Lesson 06: Command →
Incident Command
Incidents need three roles filled in five minutes: commander who decides, scribe who records, plus comms who speaks. The commander never debugs. The scribe timestamps every decision. Comms updates the status page every 20 minutes whether or not anything changed. Structure scales. Heroics do not.
- Declare severity by user impact within 10 minutes. Severity by vibes pages the wrong people.
- Update status every 20 minutes with known plus unknown plus next check. Silence breeds tickets.
- Deep dive: The Login Queue, support dark with the product.
Q1. Your sev-1 has no commander after 15 minutes. What do you do? Seen at: SRE plus backend loops.
Q2. Write the first status update for a login outage with unknown cause. Seen at: observability plus platform loops.
Command set. Now learn without blame. Lesson 07: Postmortems →
Blameless Postmortems
NATS repeated its 2023 shape in 2026 because the first postmortem changed paperwork, not architecture. Blameless means no names in the causal chain, every action linked to the topology that allowed it, plus owners plus dates on every fix. Postmortems that do not change topology are memoirs.
- Write timelines from logs, never from memory. Memory edits itself to look competent.
- Assign every action an owner plus a date. Actions without dates are wishes.
- Deep dive: NATS Four Hours, the repeat that indicted follow-through.
Q1. Your postmortem names an engineer. Rewrite that paragraph. Seen at: SRE plus backend loops.
Q2. Which fix ships first from your last incident? Seen at: observability plus platform loops.
Blame removed. Now plan the capacity. Lesson 08: Capacity →
Capacity Planning
Azure plus GCP both proved that recovery speed is designed, not improvised. Little Law converts RPS plus latency into concurrent slots. Headroom math converts growth forecasts into purchase dates. Plan quarterly with modeled growth plus buffers, then verify with game days. Hope is not a capacity strategy.
- Compute slots as peak RPS times p99 seconds times safety factor. The Pool Sizer calculator runs it.
- Hold 30 percent headroom on stateful tiers plus 50 on control planes. Control starvation cascades fastest.
- Deep dive: GCP us-west1, queue math at 8 times baseline.
Q1. Size connection pools for 40k peak RPS at 120ms p99 across 20 pods. Seen at: SRE plus backend loops.
Q2. Growth runs 10 percent monthly. When do you buy? Seen at: observability plus platform loops.
Capacity planned. Now rehearse the fire. Lesson 09: Gamedays →
Game Days
Game days inject failure on purpose: kill a zone, slow DNS, expire certs, withdraw routes in staging. Score detection time plus mitigation time plus comms quality. The team that drills route withdrawal in staging never meets it first in production. Schedule quarterly. Grade honestly. Fix before the real fire.
- Start with zone kills plus disk stalls. Graduate to route withdrawal plus IAM slowness.
- Freeze drills that find sev-1 gaps until fixed. Drills without follow-through are theater.
- Deep dive: Azure West US, the drill that would have caught it.
Q1. Design a 2-hour game day for route withdrawal in staging. Seen at: SRE plus backend loops.
Q2. Your drill finds a 40-minute detection gap. What ships Monday? Seen at: observability plus platform loops.
Fire rehearsed. Now speak through it. Lesson 10: Status →
Status-Page Honesty
Status pages fail in two directions: green during fires, or red with no detail. Publish components with real probes, post within 15 minutes of sev-1 declare, plus update on a clock even when nothing changed. Customers forgive outages faster than they forgive surprise. The page is the product during the fire.
- Bind each component to an independent probe. Manual page flips lag reality by definition.
- Template updates as known plus unknown plus next check. Fill it in 5 minutes, not 50.
- Deep dive: Cloudflare Wrong File, error pages replacing X plus OpenAI.
Q1. Your probe is green but users fail. What changed? Seen at: SRE plus backend loops.
Q2. Draft the 15-minute update for a regional degradation. Seen at: observability plus platform loops.
Comms honest. Now total the telemetry bill. Lesson 11: TeleBill →
The Observability Bill
Assume 1M RPS with 2 KB structured logs per request retained 30 days. That is about 5 PB per month before compression, near $115,000 at modeled $0.023 per GB. Sampling plus aggregation plus short retention cut it by 90 percent. Price yours with the sampling calculator. Vendors bill curiosity by the gigabyte.
- Retain raw logs 3 days plus sampled 30 plus aggregates 13 months. Three tiers, one sane invoice.
- Charge telemetry per team with showback. Free telemetry grows until the invoice lands.
- Deep dive: 18 Trillion Messages, tiers plus cheaper shapes.
Q1. Price your log pipeline at current volume. What dominates? Seen at: SRE plus backend loops.
Q2. Cut the bill 50 percent without losing error fidelity. How? Seen at: observability plus platform loops.
Telemetry priced. Now price the page. Lesson 12: Paging →
Paging Economics
Assume 200 pages monthly at 30 minutes each across engineers costing a modeled $150 per hour fully loaded. That is $15,000 monthly in interrupt time before context-switch tax, which research prices near 25 percent of productive hours around pages. Every alert that never needed a human is friendly fire with payroll attached. Tune or delete.
- Review every paging alert quarterly: did it need a human at 3 AM? Demote the rest to tickets.
- Measure pages per engineer per week with a cap of 2. Caps force tuning conversations.
- Deep dive: etcd Handshake, silent connections billing memory plus attention.
Q1. Your team takes 40 pages weekly. What is the monthly cost? Seen at: SRE plus backend loops.
Q2. Which three alerts do you demote first? Seen at: observability plus platform loops.
Pages priced. Now deploy without fear. Lesson 13: Deploys →
Deploys Without Fear
Cloudflare shipped a wrong internal file network-wide because generation lacked guardrails. Canary 1 percent plus watch burn plus auto-roll back on SLO breach. Progressive delivery turns deploys into experiments with kill switches. Teams that cannot roll back in minutes do not deploy, they gamble.
- Gate canaries on burn rate, not vibes. Breach means automatic rollback, no meeting required.
- Cap deploy blast with cells plus traffic shifting. Small blast, fast learning.
- Deep dive: Cloudflare Wrong File, generation without guardrails.
Q1. Design canary stages for a gateway handling 900k RPS. Seen at: SRE plus backend loops.
Q2. Your canary burns budget in 20 minutes. Who decides? Seen at: observability plus platform loops.
Deploys safe. Now face the panel. Lesson 14: Interviews →
The Interview Bench
SRE interviews test four moves: SLO math out loud, incident stories with timelines, tradeoff pricing, plus one debugging war story with logs. Burn-rate questions want window math. Capacity questions want Little Law. Postmortem questions want topology fixes, never heroics. Panels hire engineers who measure recovery in user-minutes.
- Open with user impact plus budget math. Five minutes of scoping saves forty of redesign.
- Close with what you would monitor first on day one. New hires who ask for dashboards get dashboards.
- Deep dive: System Design Course, Lesson 01 starts from outages.
Q1. Debug a 5x latency spike with only dashboards in 45 minutes. Seen at: SRE plus backend loops.
Q2. Tell me about an outage you owned plus the topology fix. Seen at: observability plus platform loops.
Bench warmed. Now run the review. Lesson 15: Capstone →
Run One Review
Start small on purpose. One service, one SLO review, one game day, one priced telemetry tier. Define SLIs, set burn alerts, run a zone-kill drill, write the blameless postmortem, price the observability month. A service you have broken on purpose is the only kind you can promise.
- Run the fourteen drills from lessons 1 through 13 against one real service. Gaps show at a glance.
- Hold the SLO review monthly with budget spend on the agenda. Reviews without budgets are book clubs.
- Close with the invoice. Telemetry plus headroom plus pages per month, labeled as estimates.
Q1. Present your SLO review to one skeptic. What do they challenge? Seen at: SRE plus backend loops.
Q2. Your drill pages at 2 AM. Was the rota fair? Seen at: observability plus platform loops.
Course complete. Price the gateway next: API Design in Production: 15 Lessons From the Gateway
Sources plus method
This course teaches from published Buildopsy postmortems linked under each lesson: NATS recovery math, AWS thermal paging, Azure route drills, GCP queue math, Cloudflare guardrails, MoEngage telemetry bills, Roblox tiers, etcd attention costs, plus the observability invoice model. Incident details follow those accounts. Cost framing is modeled from public list prices.
