Incident Response
Source: CloudOps Incident Management Process (COE 19418284214)
Priority Levels
| Priority | Definition | Response |
|---|---|---|
| P1 — Critical | Complete loss of service, no workaround available | Immediate |
| P2 — Significant | Major function degraded, workaround exists or partial outage | Urgent |
| P3 — Minor | Minor degradation, most functions working | Scheduled |
| P4 — No Impact | Informational, no user impact | As time permits |
Incident Lifecycle
graph LR
A[Identification] --> B[Containment]
B --> C[Resolution]
C --> D[Post-Incident Maintenance / RCA]
Stage 1 — Identification
- Alert comes in via DataDog / Coralogix / PagerDuty / customer report
- Determine priority (P1–P4)
- Declare the incident in
MS Teams → Cloud → Incidentschannel - Create a Jira OPS ticket:
- Type: Cloud Ticket
- Priority: P1/P2/P3 per urgency
- For after-hours P1: email
raas-email-alerts@rhapsody.pagerduty.com - Assemble the Cloud Incident Team (CIT)
Stage 2 — Containment
- Identify the scope of impact (which customers? which environments?)
- Implement immediate mitigations to contain the blast radius
- Keep the customer informed via Salesforce communication templates
- If AWS infrastructure is involved: escalate via AWS Support Center
AWS Escalation (5-minute response)
For AWS infrastructure incidents requiring fast escalation:
- AWS Support Center → Technical → Incident Detection and Response
- Select: Active Incident → Business-critical system down
- AWS contacts:
- TAM: Anushree —
nssht@amazon.com - Team alias:
aws-rhapsody-team@amazon.com
Stage 3 — Resolution
- Apply the fix
- Verify services are restored (connectivity, metrics, customer confirmation)
- Update the Jira ticket and CIT channel with resolution notes
- Close the PagerDuty incident
Stage 4 — Post-Incident / RCA
- Required after every P1 and P2 incident
- Write a Root Cause Analysis (RCA) document
- Track MTTA / MTTR / Total downtime metrics
- Identify action items to prevent recurrence
Cloud Incident Team (CIT)
The CIT is assembled for P1/P2 incidents:
| Role | Who |
|---|---|
| On-call SRE | Current PagerDuty on-call |
| Customer Support | Customer Success Manager or support lead |
| Incident Manager | Team manager / senior engineer |
| Vendor contact | If the issue involves a third-party system |
Communication
- Internal:
MS Teams → Cloud → Incidentschannel - External (customers): Use Salesforce communication templates — do not send ad-hoc messages to customers
- After hours P1 trigger:
raas-email-alerts@rhapsody.pagerduty.com
PagerDuty Services
| Service | Purpose |
|---|---|
| RaaS (PJR7X6W) | On-call alerting for RaaS environments |
| RaaS Support (P8OELI3) | Customer-facing support escalations |
For on-call schedules and urgency settings, see PagerDuty.