This section captures foundational incident identification and chronological metadata. Accurate timestamp recording is critical for SLA compliance verification and future prevention analysis.
Primary Cloud Service Provider or SaaS Vendor
Amazon Web Services (AWS)
Microsoft Azure
Google Cloud Platform (GCP)
Oracle Cloud Infrastructure (OCI)
Salesforce
ServiceNow
Datadog
Other SaaS Provider
Official Incident Tracking ID (e.g., INC-2025-XXXX)
Incident Severity Classification
P1 - Critical (Complete Service Outage)
P2 - Major (Significant Functional Degradation)
P3 - Moderate (Partial Feature Impact)
P4 - Low (Minor Issue with Workaround)
Incident Detection Timestamp (UTC)
Incident Acknowledgment Timestamp (UTC)
Service Degradation Start Timestamp (UTC)
Mitigation Action Initiation Timestamp (UTC)
Service Restoration Timestamp (UTC)
Detection Method(s) - Select all that applied
Automated Monitoring Alert (Prometheus, CloudWatch)
Synthetic Transaction Monitoring
Customer Support Ticket
Internal User Report
Vendor Health Dashboard Notification
Security Event Monitoring
Log Anomaly Detection
No Detection - Discovered Accidentally
Affected Service Components - Select all that applied
Compute Instances (VMs, Containers)
Object Storage (S3, Blob, GCS)
Database Services (RDS, Cosmos DB, Cloud SQL)
Serverless Functions (Lambda, Azure Functions)
Load Balancers & Traffic Management
Identity & Access Management (IAM)
API Gateway & Management
Content Delivery Network (CDN)
Message Queuing & Event Streaming
Monitoring & Observability Tools
CI/CD Pipeline Infrastructure
Other Critical Infrastructure Component
Geographic Scope of Impact
Single Availability Zone
Multiple Availability Zones (Same Region)
Multiple Regions (Same Continent)
Global Impact (Multiple Continents)
Isolated to Specific IP Range or Tenant
Did this incident affect customer-facing production services?
Executive Summary of Timeline (max 250 characters)
This section documents deep technical root cause analysis and quantifies SLA compliance failures. Rigorous analysis here drives long-term reliability improvements and vendor accountability.
Primary Root Cause Category
Infrastructure Hardware Failure (Physical)
Software Bug or Defect (Vendor Code)
Configuration Change Error (Human)
Capacity or Resource Exhaustion
Third-Party Dependency Failure
Cybersecurity Incident (Attack or Breach)
Natural Disaster or Environmental Factor
Procedural or Process Failure
Monitoring/Observability Blind Spot
Unclear or Under Investigation
Contributing Factors - Select all that applied
Inadequate Testing of Changes
Insufficient Monitoring Coverage
Lack of Runbooks or Documentation
Staffing or On-Call Coverage Gaps
Alert Fatigue or Threshold Misconfiguration
Technical Debt or Legacy System Constraints
Vendor Communication Delay
Change Management Process Bypass
Insufficient Capacity Planning
Unclear Ownership or Escalation Path
Detailed Technical Root Cause Description
SLA Breach Quantification Matrix
SLA Metric Name | Contractual Target (e.g., 99.9%) | Actual Performance | Breach Duration (minutes) | Penalty Exposure (if applicable) | |
|---|---|---|---|---|---|
API Availability | 99.95% | 87.3% | 150 | $0.00 | |
Database Query Latency p99 | <100ms | 450ms | 180 | $0.00 | |
Total Error Budget Consumed by This Incident (percentage points)
Was a formal Root Cause Analysis (RCA) document already published internally?
Did this incident reveal monitoring or observability gaps?
Could this incident have been prevented with existing knowledge or prior incidents?
Have similar incidents occurred in the past 12 months?
Rate the confidence level of this RCA (1 = Speculative, 5 = Definitive with Evidence)
This section quantifies the tangible and intangible business impact across applications, revenue streams, and workforce productivity. Accurate financial accounting ensures proper risk prioritization and investment justification.
Business Application Impact Register
Application Name | Business Criticality | Number of Users Affected | Downtime Duration (minutes) | Estimated Revenue Loss | Impact Severity (1-5) | |
|---|---|---|---|---|---|---|
Customer Payment Portal | Tier 1 - Mission Critical | 12500 | 150 | $75,000.00 | ||
Internal CRM System | Tier 2 - Business Essential | 850 | 150 | $0.00 | ||
Total Direct Revenue Loss (from transaction failures, SLA penalties, etc.)
Estimated Indirect Revenue Impact (customer churn, reputational damage)
Workforce Productivity Loss (employee hours wasted * average cost)
Emergency Remediation and Vendor Support Costs
Total Financial Impact (Sum of Above - Auto-Calculated)
Were any customer-facing Service Level Agreements (SLAs) breached?
Did this incident trigger any regulatory or compliance reporting obligations?
Rate the anticipated long-term reputational impact (1 = Very Unhappy/Extremely Negative, 5 = Very Happy/Minimal Impact)
Did competitors publicly capitalize on this outage or offer migration incentives?
Summarize the most significant business impact narrative for stakeholder communication
This section evaluates the effectiveness of disaster recovery mechanisms, failover execution fidelity, and recovery verification processes. Performance against RTO/RPO targets determines architectural resilience adequacy.
Was the formal Disaster Recovery (DR) Plan activated during this incident?
Was a failover to secondary region or backup system attempted?
Did the failover execution succeed within the defined Recovery Time Objective (RTO)?
Was data integrity maintained within the Recovery Point Objective (RPO) tolerance?
Was a failback to primary environment successfully executed?
Were any DR automation scripts or Infrastructure as Code (IaC) playbooks executed?
Rate the overall effectiveness of the DR plan execution (1-5 stars)
How recently was the DR plan tested prior to this incident?
Within last 30 days
Within last 90 days
Within last 6 months
Within last 12 months
More than 12 months ago
Never tested
Key lessons learned about DR architecture and failover mechanisms
This final section ensures executive accountability, risk register updates, and formal approval of corrective action plans. Leadership sign-off mandates resource allocation and strategic risk treatment decisions.
Executive Summary for CIO and Risk Management Director (max 500 characters)
Updated Enterprise Risk Assessment Rating
Extreme Risk - Requires Board Level Review
High Risk - Requires Executive Committee Oversight
Medium Risk - Manage with Enhanced Controls
Low Risk - Accept and Monitor
Positive Opportunity - No Risk
Corrective Action Plan and Remediation Tracker
Corrective Action Description | Action Owner (Team/Individual) | Target Completion Date | Priority | Budget Required | Requires Vendor Engagement? | |
|---|---|---|---|---|---|---|
Implement additional region failover automation | Cloud Infrastructure Team | 8/15/2025 | P0 - Immediate | $50,000.00 | ||
Enhance monitoring coverage for third-party APIs | Observability Team | 7/30/2025 | P1 - High | $15,000.00 | ||
Total Budget Required for Remediation Actions
Does this incident require updating the enterprise risk register?
Is Board of Directors notification required for this incident?
Does this incident require external regulatory filing or disclosure?
Chief Information Officer (CIO) - Digital Sign-Off
Director of Risk Management - Digital Sign-Off
Final Sign-Off Completion Timestamp (UTC)