This section establishes the official record of the incident. Please provide precise information to ensure accurate documentation and trend analysis for future prevention efforts.
Incident ID/Ticket Number
Incident Title/Descriptive Name
Incident Classification
Product/Service Outage
Security Breach
Data Loss/Corruption
Performance Degradation
Public Relations Crisis
Third-Party Vendor Failure
Regulatory Compliance Violation
Other
Initial Severity Level (1 = Low, 5 = Critical)
Revised Severity Level after full assessment (1 = Low, 5 = Critical)
Date and Time Incident Was First Detected
Date and Time Incident Actually Began (if different from detection)
Date and Time Incident Was Fully Resolved
Total Incident Duration in Minutes
Which primary systems, services, or business functions were directly impacted? (Select all that apply)
Customer-facing Web/Mobile Application
API/Backend Services
Database/Storage Layer
Payment Processing
Authentication/Authorization
Internal Tools/Dashboards
Cloud Infrastructure
Network/Connectivity
Supply Chain/Logistics
Customer Support Systems
Marketing/Communication Channels
Other Critical Systems
Estimated Number of Customers/Users Affected
Estimated Financial Impact (Revenue Loss, Penalties, etc.)
Did this incident generate external media coverage or public social media attention?
How was the incident initially detected?
Automated Monitoring Alert
Customer Support Ticket
Social Media Mention
Employee Report
Third-Party Vendor Notification
Media Inquiry
Security Scan
Other
Reconstruct the critical timeline by identifying key milestones. This helps identify delays and opportunities for faster response.
Critical Incident Timeline Reconstruction
Timestamp | Event/Milestone Description | Category | Primary Actor/Team | Time to Next Milestone (minutes) | |
|---|---|---|---|---|---|
6/15/2025, 2:23 PM | First automated alert fired: Database latency threshold breached | Detection | Monitoring System | 7 | |
6/15/2025, 2:30 PM | On-call engineer acknowledged alert and began initial investigation | Initial Response | Infrastructure Team | 45 | |
Was a physical location or specific geographic region relevant to this incident (e.g., data center, office, regional service)?
A thorough root-cause analysis is essential for preventing recurrence. This section dives deep into the technical, human, and procedural factors that contributed to the incident.
Has a formal Root Cause Analysis (RCA) document been completed and peer-reviewed?
What is the PRIMARY root cause category?
Technical/Systems Failure
Human Error
Process/Procedure Deficiency
Third-Party Vendor Issue
Security Vulnerability Exploit
Capacity/Planning Oversight
Code Defect
Configuration Error
Other
What CONTRIBUTING factors amplified the incident? (Select all that apply)
Inadequate Monitoring/Alerting
Insufficient Testing
Knowledge Silos
Recent Organizational Change
Staffing Shortage
Unclear Ownership/Responsibility
Tool Limitations
Documentation Gaps
Alert Fatigue
Slow Escalation
Competing Priorities
Technical Debt
None of the above
Was a recent change (code deploy, config change, infrastructure update) identified as a trigger?
Did the incident involve a third-party vendor, supplier, or external partner?
I confirm that the root cause analysis has been reviewed by at least one technical peer or architect who was not directly involved in the incident response.
Were there any attempts to obscure or delay disclosure of the true root cause during internal discussions?
Major incidents take a significant toll on team members. This section assesses the human impact to ensure proper support, recovery time, and resources are provided. Honest responses are critical for organizational learning and team health.
How many core team members were directly involved in the incident response?
Which primary roles/functions were engaged? (Select all that apply)
Software Engineering
Site Reliability Engineering (SRE)
Infrastructure/Platform
Database Administration
Network Engineering
Security/InfoSec
Product Management
Customer Support
Communications/PR
Legal/Compliance
Executive Leadership
Other
Average number of hours worked per person during the active incident period
Average hours of sleep lost per person during the incident
Personal Stress & Well-Being Assessment (Rate each dimension based on your personal experience)
Overall stress level during incident | |
Current stress level (post-incident) | |
Feeling of psychological safety to raise concerns | |
Confidence in leadership support | |
Sense of personal accomplishment | |
Physical exhaustion | |
Mental fatigue/cognitive load | |
Worry about future blame or repercussions |
Which burnout indicators have you or your team members experienced? (Select all that apply)
Emotional exhaustion
Cynicism or detachment from work
Reduced sense of personal efficacy
Physical symptoms (headaches, insomnia)
Increased irritability
Difficulty concentrating
Avoidance of work-related discussions
None observed
Do you feel you had adequate managerial and organizational support during the incident?
Rate the effectiveness of the post-incident debrief or retrospective process (1 star = Poor, 5 stars = Excellent)
Do you believe team members need dedicated recovery time or mental health support following this incident?
Were there any interpersonal conflicts or communication breakdowns within the response team during the incident?
What positive aspects of teamwork or individual heroics should be recognized and celebrated?
This section identifies where processes and communication channels failed or underperformed, creating opportunities for systemic improvement beyond technical fixes.
Was an incident response plan or runbook available for this scenario?
Rate the effectiveness of communication channels used during the incident
Completely Ineffective | Mostly Ineffective | Neutral | Mostly Effective | Completely Effective | |
|---|---|---|---|---|---|
Primary incident chat channel (e.g., Slack) | |||||
Video conferencing for war room | |||||
Email updates to stakeholders | |||||
Phone/SMS for escalations | |||||
Status page updates | |||||
Internal wiki/documentation | |||||
Ticketing system (Jira, etc.) |
How timely were stakeholder communications (executives, customers, partners)?
Far Too Late
Slightly Late
Adequate
Timely
Proactive
Which stakeholder groups experienced communication gaps or delays? (Select all that apply)
Internal Engineering Teams
Executive Leadership
Customer Support
Customers (direct)
Partners/Vendors
Media/Press
Investors
Regulatory Bodies
All communications were timely
Other
Rate the adequacy and accessibility of tools used during incident response
Monitoring and alerting tools | |
Log aggregation and analysis | |
Dashboards and visualization | |
Communication platforms | |
Incident management system | |
Documentation/knowledge base | |
Access management (VPN, credentials) |
Were there any delays due to lack of decision-making authority or unclear ownership?
Which process gaps contributed to the incident or slowed response? (Select all that apply)
No formal incident commander designated
Unclear escalation paths
Lack of regular disaster recovery drills
Inadequate change management process
Insufficient code review requirements
Missing operational readiness checks
No formal handoff between shifts/teams
Inadequate post-mortem culture
Slow patch management process
Other
Did you have access to all necessary systems, credentials, and data needed to diagnose and resolve the issue?
What single process or communication improvement would have had the biggest impact on reducing resolution time?
This final section translates lessons learned into concrete, trackable actions with clear ownership and accountability. These items should be reviewed weekly until completion.
How many distinct preventive action items have been identified from this incident?
Preventive Action Items & Accountability Matrix
Action Item Description | Priority | Owner (Name or Role) | Target Completion Date | Status | Resource Needs (Budget, Personnel, Tools) | Success Metrics/Definition of Done | Estimated Risk Reduction (1-5) | |
|---|---|---|---|---|---|---|---|---|
Implement automated database failover for primary cluster | P0 - Critical | Infrastructure Lead | 7/30/2025 | In Progress | 3 engineer-months, $15k for testing environment | Failover completes in <30 seconds with zero data loss | 0 | |
Revise on-call escalation policy to require manager notification within 15 min | P1 - High | Engineering Manager | 7/15/2025 | Not Started | Policy document update, manager training | 100% compliance in next incident | 0 | |
0 | ||||||||
0 | ||||||||
0 | ||||||||
0 | ||||||||
0 | ||||||||
0 | ||||||||
0 | ||||||||
0 |
Do you have confidence that leadership will provide adequate resources (budget, headcount, time) to complete these action items?
How will progress on these action items be tracked and reported?
Dedicated program management office (PMO)
Engineering team sprints
Monthly incident review meetings
Quarterly business reviews
Ad-hoc tracking
No formal tracking mechanism
Have success metrics and KPIs been defined to measure the effectiveness of these preventive measures?
Proposed date for first follow-up review of action items
How clear is the accountability structure for ensuring these action items are completed?
Very Unclear
Unclear
Neutral
Clear
Very Clear
What systemic organizational changes (if any) are needed to prevent similar incidents across other teams or products?
By signing below, you confirm that the information provided is accurate to the best of your knowledge and that you are committed to the preventive action items outlined above.
Primary Incident Commander/Lead Investigator Signature
Team Manager/Department Head Signature (Acknowledging Support Commitment)