What to Include
A complete incident response SOP covers every phase from first alert to post-mortem completion:Incident Severity Levels (SEV1 / SEV2 / SEV3)
Incident Severity Levels (SEV1 / SEV2 / SEV3)
- SEV1 — Critical: Full service outage, data loss risk, or security breach. All customers affected. Requires immediate all-hands response.
- SEV2 — Major: Significant functionality degraded or a subset of customers affected. Core workflows impaired.
- SEV3 — Minor: Non-critical functionality affected or a workaround is available. Limited customer impact.
Detection and Alert Criteria
Detection and Alert Criteria
Incident Commander Role and Responsibilities
Incident Commander Role and Responsibilities
Communication Cadence (Internal and External)
Communication Cadence (Internal and External)
Escalation Contacts
Escalation Contacts
Mitigation and Resolution Steps
Mitigation and Resolution Steps
Post-Mortem Process
Post-Mortem Process
Building It in Seilers
Create a new SOP titled 'Incident Response'
Incident Response. Assign it to the Engineering or Operations category. In the description field, note that this SOP applies to all production incidents and is required reading for every member of the on-call rotation.Define severity levels in a Decision step
Add steps for each phase: Detect, Assess, Respond, Communicate, Resolve
Embed a Post-Mortem checklist
Assign an Engineering Lead as owner
Publish and share with the on-call rotation
#on-call or #incidents Slack channel. Consider adding it as a bookmark in your incident management tool so it’s one click away when an alert fires.Running a Post-Mortem
A post-mortem is only valuable if it’s systematic and blameless. Use the following checklist inside your Post-Mortem SOP to structure every review:Timeline Reconstruction
Timeline Reconstruction
Root Cause Analysis
Root Cause Analysis
Action Items
Action Items
Post-Mortem Review Meeting
Post-Mortem Review Meeting