This section captures foundational incident identification and chronological metadata. Accurate timestamp recording is critical for SLA compliance verification and future prevention analysis.
Primary Cloud Service Provider or SaaS Vendor
Amazon Web Services (AWS)
Microsoft Azure
Google Cloud Platform (GCP)
Oracle Cloud Infrastructure (OCI)
Salesforce
ServiceNow
Datadog
Other SaaS Provider
Specify Other SaaS Provider Name
Official Incident Tracking ID (e.g., INC-2025-XXXX)
Incident Severity Classification
P1 - Critical (Complete Service Outage)
P2 - Major (Significant Functional Degradation)
P3 - Moderate (Partial Feature Impact)
P4 - Low (Minor Issue with Workaround)
Incident Detection Timestamp (UTC)
Incident Acknowledgment Timestamp (UTC)
Service Degradation Start Timestamp (UTC)
Mitigation Action Initiation Timestamp (UTC)
Service Restoration Timestamp (UTC)
Detection Method(s) - Select all that applied
Automated Monitoring Alert (Prometheus, CloudWatch)
Synthetic Transaction Monitoring
Customer Support Ticket
Internal User Report
Vendor Health Dashboard Notification
Security Event Monitoring
Log Anomaly Detection
No Detection - Discovered Accidentally
Affected Service Components - Select all that applied
Compute Instances (VMs, Containers)
Object Storage (S3, Blob, GCS)
Database Services (RDS, Cosmos DB, Cloud SQL)
Serverless Functions (Lambda, Azure Functions)
Load Balancers & Traffic Management
Identity & Access Management (IAM)
API Gateway & Management
Content Delivery Network (CDN)
Message Queuing & Event Streaming
Monitoring & Observability Tools
CI/CD Pipeline Infrastructure
Other Critical Infrastructure Component
Describe Other Critical Infrastructure Component(s) Affected
Geographic Scope of Impact
Single Availability Zone
Multiple Availability Zones (Same Region)
Multiple Regions (Same Continent)
Global Impact (Multiple Continents)
Isolated to Specific IP Range or Tenant
Did this incident affect customer-facing production services?
Estimated Number of External Customers Affected
Describe Internal Business Functions Impacted
Executive Summary of Timeline (max 250 characters)
This section documents deep technical root cause analysis and quantifies SLA compliance failures. Rigorous analysis here drives long-term reliability improvements and vendor accountability.
Primary Root Cause Category
Infrastructure Hardware Failure (Physical)
Software Bug or Defect (Vendor Code)
Configuration Change Error (Human)
Capacity or Resource Exhaustion
Third-Party Dependency Failure
Cybersecurity Incident (Attack or Breach)
Natural Disaster or Environmental Factor
Procedural or Process Failure
Monitoring/Observability Blind Spot
Unclear or Under Investigation
Describe Configuration Management Process Breakdown
Specify Third-Party Dependency Name
Detail Procedural Failure and Recommended Process Changes
Contributing Factors - Select all that applied
Inadequate Testing of Changes
Insufficient Monitoring Coverage
Lack of Runbooks or Documentation
Staffing or On-Call Coverage Gaps
Alert Fatigue or Threshold Misconfiguration
Technical Debt or Legacy System Constraints
Vendor Communication Delay
Change Management Process Bypass
Insufficient Capacity Planning
Unclear Ownership or Escalation Path
Detailed Technical Root Cause Description
SLA Breach Quantification Matrix
SLA Metric Name | Contractual Target (e.g., 99.9%) | Actual Performance | Breach Duration (minutes) | Penalty Exposure (if applicable) | ||
|---|---|---|---|---|---|---|
A | B | C | D | E | ||
1 | API Availability | 99.95% | 87.3% | 150 | $0.00 | |
2 | Database Query Latency p99 | <100ms | 450ms | 180 | $0.00 | |
3 | ||||||
4 | ||||||
5 | ||||||
6 | ||||||
7 | ||||||
8 | ||||||
9 | ||||||
10 |
Total Error Budget Consumed by This Incident (percentage points)
Was a formal Root Cause Analysis (RCA) document already published internally?
Provide RCA Document Link or Reference ID
Explain why RCA documentation was not produced
Did this incident reveal monitoring or observability gaps?
What monitoring improvements are required? Select all that apply
New Metric Collection Needed
Alert Threshold Tuning Required
Dashboard Enhancement
Log Aggregation Improvement
Tracing Implementation
Synthetic Monitoring Addition
Anomaly Detection ML Model Training
Could this incident have been prevented with existing knowledge or prior incidents?
Describe why prevention did not occur and identify accountability gaps
Explain what novel conditions made this incident unpreventable
Have similar incidents occurred in the past 12 months?
How many similar incidents in past 12 months?
Rate the confidence level of this RCA (1 = Speculative, 5 = Definitive with Evidence)
This section quantifies the tangible and intangible business impact across applications, revenue streams, and workforce productivity. Accurate financial accounting ensures proper risk prioritization and investment justification.
Business Application Impact Register
Application Name | Business Criticality | Number of Users Affected | Downtime Duration (minutes) | Estimated Revenue Loss | Impact Severity (1-5) | ||
|---|---|---|---|---|---|---|---|
A | B | C | D | E | F | ||
1 | Customer Payment Portal | Tier 1 - Mission Critical | 12500 | 150 | $75,000.00 | ||
2 | Internal CRM System | Tier 2 - Business Essential | 850 | 150 | $0.00 | ||
3 | |||||||
4 | |||||||
5 | |||||||
6 | |||||||
7 | |||||||
8 | |||||||
9 | |||||||
10 |
Total Direct Revenue Loss (from transaction failures, SLA penalties, etc.)
Estimated Indirect Revenue Impact (customer churn, reputational damage)
Workforce Productivity Loss (employee hours wasted * average cost)
Emergency Remediation and Vendor Support Costs
Total Financial Impact (Sum of Above - Auto-Calculated)
Were any customer-facing Service Level Agreements (SLAs) breached?
Number of enterprise customers impacted by SLA breach
Did this incident trigger any regulatory or compliance reporting obligations?
Which compliance frameworks are affected? Select all that apply
GDPR Data Availability
PCI DSS Transaction Processing
HIPAA System Availability
SOC 2 Type II Uptime Controls
ISO 27001 Service Continuity
Industry-Specific Financial Regulations
Rate the anticipated long-term reputational impact (1 = Very Unhappy/Extremely Negative, 5 = Very Happy/Minimal Impact)
Did competitors publicly capitalize on this outage or offer migration incentives?
Describe competitive threats and customer retention risks identified
Summarize the most significant business impact narrative for stakeholder communication
This section evaluates the effectiveness of disaster recovery mechanisms, failover execution fidelity, and recovery verification processes. Performance against RTO/RPO targets determines architectural resilience adequacy.
Was the formal Disaster Recovery (DR) Plan activated during this incident?
DR Plan Activation Timestamp (UTC)
Explain why DR Plan was not activated and whether it should have been
Was a failover to secondary region or backup system attempted?
Target Failover Location/Region
Justify why failover was not a viable option
Did the failover execution succeed within the defined Recovery Time Objective (RTO)?
Actual Time to Failover (minutes)
Detail failover failures, delays, or partial success scenarios
Was data integrity maintained within the Recovery Point Objective (RPO) tolerance?
Estimated Data Loss (minutes of transactions)
Was a failback to primary environment successfully executed?
Failback Completion Timestamp (UTC)
Explain why failback was not performed or remains pending
Were any DR automation scripts or Infrastructure as Code (IaC) playbooks executed?
Document script execution results, failures, or manual interventions required
Rate the overall effectiveness of the DR plan execution (1-5 stars)
How recently was the DR plan tested prior to this incident?
Within last 30 days
Within last 90 days
Within last 6 months
Within last 12 months
More than 12 months ago
Never tested
Key lessons learned about DR architecture and failover mechanisms
This final section ensures executive accountability, risk register updates, and formal approval of corrective action plans. Leadership sign-off mandates resource allocation and strategic risk treatment decisions.
Executive Summary for CIO and Risk Management Director (max 500 characters)
Updated Enterprise Risk Assessment Rating
Extreme Risk - Requires Board Level Review
High Risk - Requires Executive Committee Oversight
Medium Risk - Manage with Enhanced Controls
Low Risk - Accept and Monitor
Positive Opportunity - No Risk
Corrective Action Plan and Remediation Tracker
Corrective Action Description | Action Owner (Team/Individual) | Target Completion Date | Priority | Budget Required | Requires Vendor Engagement? | ||
|---|---|---|---|---|---|---|---|
A | B | C | D | E | F | ||
1 | Implement additional region failover automation | Cloud Infrastructure Team | 8/15/2025 | P0 - Immediate | $50,000.00 | ||
2 | Enhance monitoring coverage for third-party APIs | Observability Team | 7/30/2025 | P1 - High | $15,000.00 | ||
3 | |||||||
4 | |||||||
5 | |||||||
6 | |||||||
7 | |||||||
8 | |||||||
9 | |||||||
10 |
Total Budget Required for Remediation Actions
Does this incident require updating the enterprise risk register?
Specify risk register updates, new entries, or control modifications needed
Is Board of Directors notification required for this incident?
Board Notification Date
Does this incident require external regulatory filing or disclosure?
Specify regulatory bodies, filing deadlines, and disclosure requirements
Chief Information Officer (CIO) - Digital Sign-Off
Director of Risk Management - Digital Sign-Off
Final Sign-Off Completion Timestamp (UTC)
To configure an element, select it on the form.