Incident Response

Incidents should optimise first for safe restoration, then for understanding and prevention.

Lifecycle

  1. Detect and scope — identify the user-visible impact and affected systems.
  2. Stabilise — stop further damage; prefer a known-good rollback when appropriate.
  3. Recover — restore the service and verify from the user's perspective.
  4. Preserve evidence — keep relevant logs, timestamps, deploy IDs and configuration diffs.
  5. Review — document contributing conditions and follow-up actions without blame.
  6. Improve — implement fixes, monitoring and documentation changes.

Communication

Use exact times with time zones for material events. Distinguish confirmed facts from hypotheses. Never paste secrets into incident pages.

Use Templates/Incident-Review for the durable record and create or update a runbook when the incident revealed a repeatable recovery procedure.

0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6 7 8 9