[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"glossary-runbook::en":3,"gloss-cluster-runbook::en":26,"gloss-next-runbook::en":9},{"slug":4,"category":5,"name":6,"definition":7,"meta_desc":8,"faq":9,"schema_markup":9,"related":10},"runbook","cloud","Runbook","A runbook is a written procedure for handling one specific operational situation — a failing health check, a queue backing up, a certificate about to expire, a region failing over — structured so that someone who did not build the system can execute it correctly while an incident is in progress. It is not documentation about how a service works; that is an architecture document, and reaching for one during an outage is a sign the runbook is missing. The distinction that makes runbooks useful is scope. A good runbook covers one alert or one symptom, states how to confirm the situation is what you think it is, lists the diagnostic commands with their expected output, gives the remediation steps in order, and names the point at which you stop and escalate. That last part is the one most often omitted and the most valuable: an on-call engineer who does not know when to wake someone else will either escalate too early or, far worse, keep trying things on production at four in the morning. Every alert that pages a human should link to the runbook for that alert. An alert without one is asking the responder to improvise, which is where outages get extended by well-intentioned action. This is also the cheapest quality signal in an ops setup — if an alert has no runbook, either the alert is not actionable and should not page, or the procedure exists only in one person's head. Runbooks decay faster than most documentation, because they encode the current shape of the system: a command changes, a dashboard moves, a service is renamed, and the procedure silently becomes fiction. The fix is to treat the runbook as part of the change rather than as a follow-up, and to have whoever used it during an incident correct it immediately afterwards, while the gap between what it said and what actually happened is still fresh. Practical note: keep runbooks in version control next to the service they cover and review them in the postmortem of every incident where one was used. Automating a runbook's steps is usually the right end state, but write it as prose first — a procedure nobody has executed by hand is not ready to be turned into code.","A runbook is a short, executable procedure for handling one specific operational situation, written to be followed under pressure at 3am.",null,[11,14,17,20,23],{"slug":12,"name":13},"distributed-tracing","Distributed Tracing",{"slug":15,"name":16},"error-budget","Error Budget",{"slug":18,"name":19},"incident-management","Incident Management",{"slug":21,"name":22},"observability","Observability",{"slug":24,"name":25},"service-level-objective","Service Level Objective (SLO)",[27,31,34,37,41,44,47,50,53,56,59,62],{"slug":28,"category":5,"name":29,"updated_at":30},"autoscaling","Autoscaling","2026-08-24T02:46:37+00:00",{"slug":32,"category":5,"name":33,"updated_at":30},"availability-zone","Availability Zone (AZ)",{"slug":35,"category":5,"name":36,"updated_at":30},"block-storage","Block Storage",{"slug":38,"category":5,"name":39,"updated_at":40},"disaster-recovery","Disaster Recovery","2026-08-24T02:46:38+00:00",{"slug":42,"category":5,"name":43,"updated_at":40},"edge-ai","Edge AI",{"slug":45,"category":5,"name":46,"updated_at":30},"egress-fees","Egress Fees (Data Transfer Out)",{"slug":48,"category":5,"name":49,"updated_at":30},"finops","FinOps (Cloud Financial Operations)",{"slug":51,"category":5,"name":52,"updated_at":40},"immutable-infrastructure","Immutable Infrastructure",{"slug":54,"category":5,"name":55,"updated_at":40},"infrastructure-drift","Infrastructure Drift",{"slug":57,"category":5,"name":58,"updated_at":30},"managed-kubernetes","Managed Kubernetes",{"slug":60,"category":5,"name":61,"updated_at":30},"multi-region","Multi-Region",{"slug":63,"category":5,"name":64,"updated_at":40},"noisy-neighbor","Noisy Neighbor"]