Operator resource

Incident Response Checklist for SRE Teams

Use this checklist as a shared prompt under pressure. Assign an owner for each open item, record decisions as they happen, and adapt the sequence to your response policy.

Updated August 2, 2026

Declare and establish control

Make the incident legible before the response fans out.

Name the problem

Write a factual title, assign the immutable incident reference, and record the first known impact time separately from detection time.

Set initial scope

Record affected services, regions, user journeys, and the current customer-impact assessment. Mark unknowns explicitly.

Assign response roles

Name the incident lead, technical lead, and communications owner. Give every immediate action one accountable owner.

Investigate and mitigate

Separate observations, hypotheses, decisions, and actions so reviewers can follow the response later.

Preserve evidence

Attach the alert, dashboards, logs, deploys, traces, screenshots, and source links that informed the response.

Track hypotheses

Record what is being tested, the owner, the evidence for and against it, and what result would change the next action.

Make mitigation reversible

State the expected effect, rollback path, owner, start time, and validation signal before applying a risky change.

Communicate and close

Keep responders and stakeholders aligned through recovery and review.

Set the next update

Publish what is known, what remains unknown, the current action, the owner, and the time of the next update.

Verify recovery

Confirm customer-facing health and internal signals, watch for recurrence, and record who accepted the recovery evidence.

Open follow-up work

Capture unresolved risk, assign owners and due dates, and schedule the postmortem while context is still available.