Incident Response
Incidents should optimise first for safe restoration, then for understanding and prevention.
Lifecycle
- Detect and scope — identify the user-visible impact and affected systems.
- Stabilise — stop further damage; prefer a known-good rollback when appropriate.
- Recover — restore the service and verify from the user's perspective.
- Preserve evidence — keep relevant logs, timestamps, deploy IDs and configuration diffs.
- Review — document contributing conditions and follow-up actions without blame.
- Improve — implement fixes, monitoring and documentation changes.
Communication
Use exact times with time zones for material events. Distinguish confirmed facts from hypotheses. Never paste secrets into incident pages.
Use Templates/Incident-Review for the durable record and create or update a runbook when the incident revealed a repeatable recovery procedure.