Template — Incident Review

Incident date/time:
Service(s):
Impact:
Status: Resolved / Monitoring / Open

Summary

Short factual description of user-visible impact and duration.

Timeline

Use exact timestamps with time zones. Mark hypotheses as hypotheses.

Detection

How the issue was first detected and whether monitoring should have detected it earlier.

Root and contributing conditions

Explain technical/system conditions without assigning personal blame.

Recovery

What restored service and how restoration was verified.

What worked / what did not

Tools, runbooks, alerts, architecture and communication.

Actions

Concrete owner, action and tracking reference. Separate immediate fixes from longer-term prevention.

Documentation changes

List runbooks/reference pages that were created or updated because of this incident.

0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9