Skip to content

Incident Response

Source: CloudOps Incident Management Process (COE 19418284214)

Priority Levels

Priority Definition Response
P1 — Critical Complete loss of service, no workaround available Immediate
P2 — Significant Major function degraded, workaround exists or partial outage Urgent
P3 — Minor Minor degradation, most functions working Scheduled
P4 — No Impact Informational, no user impact As time permits

Incident Lifecycle

graph LR
    A[Identification] --> B[Containment]
    B --> C[Resolution]
    C --> D[Post-Incident Maintenance / RCA]

Stage 1 — Identification

  • Alert comes in via DataDog / Coralogix / PagerDuty / customer report
  • Determine priority (P1–P4)
  • Declare the incident in MS Teams → Cloud → Incidents channel
  • Create a Jira OPS ticket:
  • Type: Cloud Ticket
  • Priority: P1/P2/P3 per urgency
  • For after-hours P1: email raas-email-alerts@rhapsody.pagerduty.com
  • Assemble the Cloud Incident Team (CIT)

Stage 2 — Containment

  • Identify the scope of impact (which customers? which environments?)
  • Implement immediate mitigations to contain the blast radius
  • Keep the customer informed via Salesforce communication templates
  • If AWS infrastructure is involved: escalate via AWS Support Center

AWS Escalation (5-minute response)

For AWS infrastructure incidents requiring fast escalation:

  1. AWS Support Center → TechnicalIncident Detection and Response
  2. Select: Active IncidentBusiness-critical system down
  3. AWS contacts:
  4. TAM: Anushree — nssht@amazon.com
  5. Team alias: aws-rhapsody-team@amazon.com

Stage 3 — Resolution

  • Apply the fix
  • Verify services are restored (connectivity, metrics, customer confirmation)
  • Update the Jira ticket and CIT channel with resolution notes
  • Close the PagerDuty incident

Stage 4 — Post-Incident / RCA

  • Required after every P1 and P2 incident
  • Write a Root Cause Analysis (RCA) document
  • Track MTTA / MTTR / Total downtime metrics
  • Identify action items to prevent recurrence

Cloud Incident Team (CIT)

The CIT is assembled for P1/P2 incidents:

Role Who
On-call SRE Current PagerDuty on-call
Customer Support Customer Success Manager or support lead
Incident Manager Team manager / senior engineer
Vendor contact If the issue involves a third-party system

Communication

  • Internal: MS Teams → Cloud → Incidents channel
  • External (customers): Use Salesforce communication templates — do not send ad-hoc messages to customers
  • After hours P1 trigger: raas-email-alerts@rhapsody.pagerduty.com

PagerDuty Services

Service Purpose
RaaS (PJR7X6W) On-call alerting for RaaS environments
RaaS Support (P8OELI3) Customer-facing support escalations

For on-call schedules and urgency settings, see PagerDuty.