{ }
Engineering · 14 steps · updated 2h ago

Incident Response P0

Captured from the last six sev-0 events. Follow top to bottom; every step has an owner and a hard time budget.

/playbooks/incident-response-p0?step=0

Steps

Owner · Reliability Guild

0 of 14 steps complete

Loading your saved progress…

0%
  1. Declare the incident

    Open an incident channel and set severity to P0 in the tracker.

  2. Assign incident commander

    First responder holds command until explicitly handed over.

  3. Page secondary on-call

    Escalate to the platform secondary within 5 minutes of declaration.

  4. Post the first status update

    Public status page within 10 minutes, even with no root cause yet.

  5. Open a mitigation track

    Split investigation and mitigation into two parallel workstreams.

  6. Attempt rollback first

    Prefer reverting the last deploy over forward fixes during a P0.

  7. Update stakeholders every 30 min

    Cadence holds until severity is downgraded.

  8. Confirm recovery signals

    Error rate, latency and queue depth back to baseline for 15 minutes.

  9. Downgrade severity

    Commander announces the downgrade and remaining follow-ups.

  10. Freeze related deploys

    Hold the affected service until the postmortem action list exists.

  11. Collect the timeline

    Export channel history, alerts and deploy events into the doc.

  12. Draft the postmortem

    Blameless template, due within 48 hours of resolution.

  13. Review with the guild

    Walk the timeline, agree on contributing factors.

  14. File and assign actions

    Every action item gets an owner and a due date before closing.