Skip to content

Cloud Incident Process Flow Chart

Incident Management Overview


This Cloud Incident Management (CIM) process is a standard process to identify and resolve service-impacting problems for Rhapsody cloud customers. How we react to service-impacting incidents makes all the difference in minimizing the impact of the incident and bringing service back up for our customers. The CIM process will help us swiftly respond to and resolve incidents.

The Rhapsody Cloud Operations team aims to eliminate service-impacting incidents. However, service-impacting incidents are impossible to prevent entirely, and the only thing we can do is be prepared for them. In this process document, the Rhapsody Cloud Operations team will define how to handle incidents that impact one or more of our cloud customers. This process applies to customers hosted in the Rhapsody AWS Cloud, including any customers that are managed by our managed service provider (MSP) partners such as ClearDATA and Logicworks.

Scope for this process

  • CaaS

  • RaaS

  • EMPI via LogicWorks

  • EMPI via Rhapsody (aka Zeus)

  • Envoy (Rhapsody-managed and Rhapsody-hosted in Cloud)

  • Semantic

Future scope:

  • Corepoint One via ClearDATA

  • CMI via ClearDATA

  • Emissary

  • Rapid API Gateway

  • CareIndexing

  • Insights << TBD- Not sure>>

What is a Cloud Incident?


A cloud incident is a high-impact, urgent issue that usually affects one or more Rhapsody cloud customers. A cloud incident almost always results in cloud services becoming unavailable, which causes a disturbance in our customers' operations. The Cloud Operations team will consider any problem an incident if it meets the definition of a P1 issue (see below) and involves the participation of the Cloud Ops team.

A cloud incident can result from a security threat, network, or infrastructure issue. A well-prepared Cloud Operations team will be equipped to assess cloud incidents and develop solutions or workarounds to reduce and control the impact of an incident. This incident management process enables the Cloud Operations team to minimize customer downtime and institute internal and external communications that are necessary for visibility and overall situation management.

Four Stages of Incident


Cloud incidents have four main stages:

  1. Identification

  2. Containment

  3. Resolution

  4. Maintenance

Stage 1: Identification


If the issue is identified as “Priority 1” (P1), we need to declare it as an incident. A P1 issue is defined as a “Critical Business Impact causing a complete loss of the Service to the Customer’s production environment such that work cannot reasonably continue and for which no workarounds to provide all of the functionality of the Service required under the agreement are possible or can be implemented in time to minimize the impact on the customer’s business.”

For customer-reported P2, P3, and P4 support cases and for customer-reported P1 cases that are unrelated to cloud infrastructure, network, and security, please use the product’s usual processes for engaging Support and Cloud Ops. Do not invoke this Cloud Incident Management Process for anything other than cloud-relatedP1 incidents. For more information on handling customer-reported issues and on the definitions of priority levels, see the Rhapsody Support Guide.

Only issues that meet the definition of Priority 1 (P1) and that require the involvement of the Cloud Ops team trigger this Cloud Incident Management process. P1 issues that are unrelated to cloud infrastructure, network, or security and issues of a lower priority do not trigger this process.

Incident Declaration: The Cloud Operations team allows multiple channels for internal and external teams to identify problems. These include Cloud Monitoring and Alerting systems, customer-reported problems, or security vulnerabilities that have the potential to impact our customer environments. One or more members of the following Rhapsody groups will declare the incident:

  • Cloud Operations and Engineering

  • Customer Support

  • Professional Service

To declare an incident start a new conversation on the MS Teams | Cloud | Incidents channel, announcing the potential incident and describing the observed behavior and known details of the issue.

Stage 2: Containment


Assemble the Cloud Incident Team (CIT): the individual/s who first identified and/or declared the incident then takes the following steps:

How to Request 5-Min Incident Response

  * Open the AWS Support Center Console and then select `Create a case`

  * Choose `Technical`

  * For Service, choose `Incident Detection and Response`

  * For Category, choose `Active Incident`

  * For Severity, choose `Business-critical system down`

  * In the `Additional Contact` section, enter any email address that you want to receive correspondence about this incident.

Quick Steps:

  1. Create a new conversation on the MS Teams | Cloud | Incidents channel.

  2. Create a Cloud Ops ticket in the Jira OPS project.

  3. Post the Cloud Ops ticket number/s to the MS Teams | Cloud | Incidents thread.

  4. Engage the appropriate Cloud Ops product teams (contact info below).

  5. Contact the appropriate Support product teams (contact info below).

  6. Engage appropriate Leadership personnel.

Contact Info

AWS Unified Operations (UOps) Engagement

  • AWS Unified Operations Team

    • Account Manager

    • Hayder alaraji@amazon.com

    • Solution Architect

    • Mohannad awsmo@amazon.com

    • Technical Account Manager (TAM)

    • Anushree nssht@amazon.com

    • Domain Specialist Engineer

    • Vihar kodavi@amazon.com

    • Team Alias Contact Email:

    • aws-rhapsody-team@amazon.com

  • How to Request 5-Min Incident Response

    • Open the AWS Support Center Console and then select Create a case

    • Choose Technical

    • For Service, choose Incident Detection and Response

    • For Category, choose Active Incident

    • For Severity, choose Business-critical system down

    • In the Additional Contact section, enter any email address that you want to receive correspondence about this incident.

CIT Conference Bridge: The Incident Manager will initiate a CIT Conference Bridge (via MS Teams video call); this will act as a clear, fast channel of communication to members of the CIT. The CIT Conference Bridge will remain open until the incident is resolved.

CIT members collaborate internally on the CIT Conference Bridge, established by the Incident Manager. The bridge is only for those CIT members actively involved in resolving the issue.

Ticket to Document the Incident: The Incident Manager will ensure that a ticket for the incident is created in Jira or Salesforce to track the issue and document the progress. This ticket will help drive Root Cause Analysis (RCA) to prevent similar issues from happening.

Informing Stakeholders: Progress on resolving the incident will be communicated to all key stakeholders via the MS Teams | Cloud | Incidents channel by the Incident Manager and other members of the CIT as needed.

  • The Incident Manager will notify the Director of Cloud Operations & Engineering via the incident channel.

  • The Incident Manager will inform the following stakeholders via the incident channel and provide progress updates toward the resolution:

    • Customer Support Leader

    • Account Management Leader

    • Chief Technology Officer

    • Professional Services Leader

  • The Incident Manager will provide information and updates to the customer support team members on the CIT to send out progress notifications to affected customers using the pre-drafted templates within Salesforce. The Customer Support team members will use the List Email functionality in Salesforce to send out these notifications.

Internal stakeholders stay informed of incident resolution progress by following the MS Teams | Cloud | Incidents channel.

The Incident Manager provides information and updates to members on the CIT.

Customers stay informed of incident resolution progress by communications from the Customer Support team members on the CIT.

The Support Agent communicates updates to the affected customers, when prompted by the Incident Manager.

Escalation: The Director of Cloud Operations & Engineering will decide if the incident warrants notification to the CTO and/or COO of Rhapsody. The Director of Cloud Operations & Engineering will get input from the Incident Manager and the CIT on the needs and desires to escalate further within Rhapsody.

Stage 3: Resolution


The Cloud Incident Team will work to resolve the problem that resulted in the incident. It is good practice to implement the fix for the incident as a change via Jira tickets and code commits where possible to ensure that the resolution is properly documented and implemented. Once the problem is resolved, the impacted environments will be verified and monitored.

Stage 4: Maintenance


Root Cause Analysis: After each incident, the Rhapsody Cloud Operations team will conduct a Root Cause Analysis (RCA). The Cloud Ops team manager (or their delegate) will review the completed RCA for completeness and approve it for customer release.

Documentation: Each incident will produce the following:

  • Internal Root Cause Analysis

  • Customer-facing Root Cause Analysis if requested by customer

  • Change requests/Action items with dates

Incident Metrics: After each incident, the Cloud Operations and Engineering team will measure itself in several key metrics to assess the effectiveness of the Cloud Operation team’s incident management process.

  • Mean time to acknowledge (MTTA)

  • Mean time to resolve (MTTR)

  • Total downtime

Cloud Incident Process Flow Chart


The following flow chart summarizes Rhapsody’s Cloud Incident management process.

Roles and Responsibilities


Incident Manager

The Incident Manager is a member of the Cloud Operations, Customer Support, Professional Services, or a Technology Leader who is responsible for managing an incident in the Rhapsody Cloud infrastructure. This Incident Manager’s priority is to guide an incident to its resolution as quickly and completely as possible. They are expected to manage the resource allocation and communication involved in that incident. The Incident Manager is the primary point of contact and source of truth about the cloud incident. The Incident Manager is responsible for setting up communication channels, establishing conference bridges, and inviting the appropriate people into those channels during an incident.

Primary responsibilities include the following:

  • Declaring an incident

  • Assembling the Cloud Incident Team

  • Establishing Conference Bridge

  • Communicating with stakeholders and leadership

  • Ensuring customer-facing departments are communicating with the affected customers

  • Supporting and managing the incident until it is fully resolved

  • Organizing and coordinating Root Cause Analysis of the incident

Cloud Incident Team (CIT)

The Cloud Incident Team (CIT) is assembled to analyze and resolve the incident as quickly as possible. This team will preserve all evidence such as logs and system status that will help postmortem and RCA of the incident. This team communicates with the Incident Manager to ensure the flow of information throughout the incident.

The CIT consists of the following:

  • Cloud Operations SREs and Architects for impacted systems

  • Customer Support Engineer/Team Lead/Manager

  • Incident Manager

  • Vendors or external partners

For each incident, the CIT will consist of at a minimum the following two groups:

  • Cloud Operations and Engineering

  • Customer Support

Cloud Operations Engineers/Site Reliability Engineers(SRE)

The Cloud Ops Engineers/SRE’s primary responsibilities include the following:

  • Analyzing problems reported by monitoring systems, Customer Support, or customers.

  • Escalating issues that are of high impact to the Incident Manager.

  • Troubleshoot the incident and are primarily responsible for implementing the incident resolution.

Customer Support Engineer/Lead/Manager

The Customer Support Engineer/Lead/Manager continuously stays informed of the status of the incident and communicates with customers during the identification, containment, and resolution stages.

Development Lead

The Product Development Leads**** are available to provide help when required. For example, a database expert may be called upon to help analyze database issues.

Rhapsody Leadership

Rhapsody Leadership (S/ELT members, such as COO, CTO, and VP of Customer Support) will support the Cloud Incident Team and Incident Manager to ensure information is flowing to all relevant stakeholders (internal and external). In addition, these key leaders will help determine follow-up actions as required.

Incident Communication


The key to a successful resolution and reducing the service availability impact is communication.

Once the incident is declared, the following will be initiated:

Internal Communication

  • CIT Conference Bridge: The Incident Manager will establish an MS Teams call/chat for the CIT members to communicate & collaborate with each other to resolve the incident. The Incident Manager will initiate this bridge when an incident is declared, which will stay open until all services are restored.

    • Create Teams chat and rename as follows: “CIT - [CustomerName] [brief issue statement] - [Severity] [Date ddmmmyy] [OPS Jira#]”

    • Example: “CIT - Acme Health Production Down - P1 01Jun2025 OPS-4444”

  • Incident Channel: The Incident Manager will post the following information via the MS Teams | Cloud | Incidents channel:

    • Declaration:

    • Impacted product(s)

    • Severity

    • Customers Impacted by Incident

    • Status

    • A brief description and Scope of the Incident

    • Impacted customers and/or customer cloud environment identifiers

    • Incident issue references (Jira, SalesForce Case)

    • Incident start time (Please be sure to keep consistent time zone recording of all times)

    • Incident Manager:

    • Cloud Operations Engineers:

    • Customer Support Engineers:

    • Periodic Updates

    • Resolution

    • RCA (once completed)

Health

Health

Example Declaration Message:

RAAS - P1 Incident affecting Acme Health Production
Status: Investigating
Failures on all outbound traffic from RaaS connections. This is limited to Acme Health production only - Customer ID 111111.
Incident Detected: 01/01/2023 at 09:45 AM CST
Jira Issue: OPS-4444 /// Salesforce Case: 00012345
Incident Manager: John Deer
Cloud Operations Engineer(s): Jane Doe
Customer Support Engineer(s): Helpful Andy

External Communication

It is essential that we provide proactive communications to any full outage, partial outage, or scheduled maintenance. The Rhapsody Cloud Operations team will rely on the Customer Support organization to handle all customer-facing communications using Salesforce and the customer contacts stored within the Salesforce system.

In the event that a customer notification is required, the Cloud Operations team or Incident Manager will ask the CIT Customer Support Representative to send messages utilizing the templates in Salesforce along with the information provided by the Incident Manager.

Intermediate Status

Providing regular and transparent status updates to customers during a cloud SaaS service outage is crucial for maintaining trust and managing customer expectations. Each incident will be different in nature and length. Periodic status to impacted customer should provide information that shows we are making progress. Here are some intermediate status updates can be used to provide current status of the outage:

Initial Notification: Notify customers as soon as you become aware of the issue. Include a brief description of the problem and reassure them that you are actively working on it.

Ongoing Progress Updates: Provide regular updates at predefined intervals (e.g., every 30 minutes) or when significant progress is made. These updates should include what has been accomplished and what steps are next.

Root Cause Analysis: Share any insights into the root cause of the issue, if known. This demonstrates transparency and shows customers that you're actively working on a solution.

Estimated Time to Resolution (ETR): Provide an estimated time for when you expect services to be fully restored. Be cautious with this estimate and clearly state that it may change as you learn more about the problem.

Milestones Achieved: Communicate significant milestones or progress points during the resolution process. For example, "We have identified the issue and are working on a fix."

Workarounds: If possible, suggest temporary workarounds that customers can use to minimize the impact on their operations while the issue is being resolved.

Service Health Dashboard: If you have a service health dashboard, encourage customers to check it for real-time updates on the status of different components of the service.

Support Channels: Remind customers of the available support channels (phone, email, chat) for those who need immediate assistance or have specific questions.

Impact Assessment: Describe the scope and impact of the issue, including affected services, regions, or user groups.

Testing and Validation: Share information about the testing and validation processes that are being conducted before fully restoring services. Highlight the importance of ensuring that services are stable before returning to normal operation.

Customer Feedback Collection: Encourage customers to provide feedback on their experience during the incident. This can help you improve your incident response in the future.

Thank You and Apology: Express gratitude for their patience and apologize for any inconvenience caused. A sincere apology can go a long way in maintaining customer goodwill.

Final Resolution: When services are fully restored, notify customers that the issue has been resolved and provide a brief summary of what caused the problem and the steps taken to fix it.

Post-Incident Report: After the incident is resolved, provide a comprehensive post-incident report that includes a detailed analysis of the root cause, actions taken, and steps to prevent a recurrence. Share this with customers to build trust and demonstrate your commitment to improving the service.

Follow-Up: After service restoration, follow up with customers to ensure everything is working as expected and inquire if they have any lingering concerns or questions.

Communication Plan


During any service interruption, it is critical to communicate internally and externally to our customers. It is hard to write clear and eloquent updates when things are on fire and the person responsible for communication is under pressure. For this reason, we have created pre-written communication templates in Salesforce to help achieve the results we need. Fast and accurate communication will lessen incoming queries and prevent the erosion of trust with our customers. The following communications should be exercised as required:

| Communication| Who| What| When
---|---|---|---|---
1| Incident Declaration (Support): If Support comes to know about a service interruption, they will inform Cloud Ops via the process outlined in Containment Stage.| Customer Support| PagerDuty to reach CloudOps| Upon Incident Declaration
2| Incident Declaration (Cloud Operations): If Cloud Ops/SRE comes to know about a service interruption, they will inform Support via the process outlined in Containment Stage.| CloudOps| PagerDuty to reach Customer Support| Upon Incident Declaration
3| Internal Communication: Incident Manager will create communication channel(s) and inform stakeholders.| Incident Manager| MS Teams | Cloud | Incidents channel| Upon Incident Declaration
4| Internal communication: Incident Manager: Will inform the Customer Support Representative to inform impacted customers.| Incident Manager| MS Teams | Cloud | Incidents channel| Upon Incident Declaration
5| Customer Initial Communication| Customer Support| “New Incident” template in Salesforce| Incident Manager tells Support to inform Customer
6| Periodic Customer Communication| Customer Support| “Incident Update” template in Salesforce| Incident Manager tells Support to inform Customer - typically 30 minute intervals depending on severity
7| Final “All Clear” communication to Customer| Customer Support| “Incident Resolved” template in Salesforce| Incident Manager tells Support to inform Customer
8| Internal: Incident Manager sends email internally to stakeholders, CIT, and support that incident is resolved.| Incident Manager| MS Teams | Cloud | Incidents channel| Incident is Resolved

Customer Problem Classification


Each issue reported by the customer or detected internally by the monitoring and alerting system will be classified and prioritized in these four Cloud Fault levels.

Cloud Fault Level Definition Example
Priority 1 / Level 1
Incidents Critical Business Impact means an incident that causes a complete loss of the Service to the Customer’s production environment such that work cannot reasonably continue and no workarounds to provide all of the functionality of the Service required under the agreement are possible or can be implemented in time to minimize the impact on the customer’s business.
  • Customer unable to access Rhapsody’s Cloud hosted application
  • No traffic is flowing to or from the system
  • Loss of access to all users

Priority 2 / Level 2| Significant Business Impact means an issue that results in a significant loss of service to the customer’s production environment such that processing can proceed in a restricted fashion but performance is significantly reduced and/or operation of the service is considered severely limited, and no workaround to provide the affected functionality is possible or can be implemented in time to minimize the impact on the customer’s business.|

  • Imminent outage or severely degraded performance
  • Loss of access to subset of users

Priority 3 / Level 3| Minor Business Impact means an issue that results in minimal loss of service thereof to the customer’s production environment such that the impact of the incident is minor or an inconvenience, such as requiring a manual bypass to restore product functionality.|

  • Component level issues (interface, comm point, filter, etc.)
  • Application or configuration regressions (break-fix)
  • Defects (bugs) which impede non-critical functionality and a workaround is available

Priority 4 / Level 4| No Business Impact means an issue that causes no loss of service and in no way impedes the customer’s use of service in accordance with the agreement.|

  • Technical or general questions
  • Requests for components/interfaces in development
  • Defects (bugs) which do not impede functionality
  • Enhancement requests

Synced from Confluence: CloudOps Incident Management Process | Last modified: 2026-03-03 | Author: Noel Gutierrez