Blame
|
1 | # Template — Incident Review |
||||||
| 2 | ||||||||
| 3 | **Incident date/time:** |
|||||||
| 4 | **Service(s):** |
|||||||
| 5 | **Impact:** |
|||||||
| 6 | **Status:** Resolved / Monitoring / Open |
|||||||
| 7 | ||||||||
| 8 | ## Summary |
|||||||
| 9 | ||||||||
| 10 | Short factual description of user-visible impact and duration. |
|||||||
| 11 | ||||||||
| 12 | ## Timeline |
|||||||
| 13 | ||||||||
| 14 | Use exact timestamps with time zones. Mark hypotheses as hypotheses. |
|||||||
| 15 | ||||||||
| 16 | ## Detection |
|||||||
| 17 | ||||||||
| 18 | How the issue was first detected and whether monitoring should have detected it earlier. |
|||||||
| 19 | ||||||||
| 20 | ## Root and contributing conditions |
|||||||
| 21 | ||||||||
| 22 | Explain technical/system conditions without assigning personal blame. |
|||||||
| 23 | ||||||||
| 24 | ## Recovery |
|||||||
| 25 | ||||||||
| 26 | What restored service and how restoration was verified. |
|||||||
| 27 | ||||||||
| 28 | ## What worked / what did not |
|||||||
| 29 | ||||||||
| 30 | Tools, runbooks, alerts, architecture and communication. |
|||||||
| 31 | ||||||||
| 32 | ## Actions |
|||||||
| 33 | ||||||||
| 34 | Concrete owner, action and tracking reference. Separate immediate fixes from longer-term prevention. |
|||||||
| 35 | ||||||||
| 36 | ## Documentation changes |
|||||||
| 37 | ||||||||
| 38 | List runbooks/reference pages that were created or updated because of this incident. |
|||||||