๐Ÿšจ On-Call Triage Decision Tool

Validate, investigate, and route incidents to the correct team

What type of incident is being reported?
Start by identifying the nature of the issue. You'll validate and gather information before escalating to a development team.
Create Squadcast Incident
You've validated the issue and gathered the necessary information. Now create a Squadcast incident to alert the development team.

๐Ÿ“‹ Squadcast Incident Creation

1. Navigate to: Squadcast Incidents Page โ†’

2. Verify you're on the correct team (top left dropdown)

3. Click "Create New Incident"

4. Fill in the form with the information you've gathered

โœ… Required Information

Always include in your incident description:

  • How many users affected? (Single user / Multiple users / Widespread)
  • When did this issue start? (Timestamp - be as specific as possible)
  • What steps have been taken so far for troubleshooting?

๐Ÿ“ Incident Form Fields

Service: Each team has their own service - select the appropriate one from the dropdown

Priority: Use your discretion based on impact (P0 = Widespread/Revenue-affecting, P1 = Multiple users, P2 = Limited impact)

Description: Include your findings, expected vs current behavior, and contact information for follow-up questions

๐Ÿ”ฅ CRITICAL - FULL OUTAGE

Emergency Response - Everything Down

๐Ÿ“ž Immediate Actions

1. Alert Multiple Teams via SquadCast:

  • API Team - Core services and integrations
  • Core Team - Database and infrastructure layer

2. Navigate to Squadcast:

https://app.squadcast.com/incident?tab=open โ†’

๐Ÿ“‹ SquadCast Message Template

โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ” ๐Ÿšจ CRITICAL OUTAGE - FULL SYSTEM DOWN โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ” ๐Ÿ“‹ INCIDENT SUMMARY Issue: Complete service outage - Multiple systems down Severity: P0 - CRITICAL Teams: API + Core Reporter: [Your Name] Time Reported: [Current Time] โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ” ๐Ÿ” HOW MANY USERS AFFECTED? Widespread - All users unable to access services โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ” โฐ WHEN DID THIS ISSUE START? [Timestamp from first report/detection] โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ” ๐Ÿ”ง TROUBLESHOOTING STEPS TAKEN SO FAR: โ€ข Confirmed multiple systems down โ€ข Checked: [List systems checked - Shop2, Portal, Office, etc.] โ€ข Datadog API Dashboard: [Observations] โ€ข AWS Status: [Any known issues] โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ” โš ๏ธ AFFECTED SYSTEMS [List all confirmed down systems] Examples: โ€ข Shop2 storefront โ€ข Portal/CRM โ€ข Office tools โ€ข API endpoints โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ” ๐ŸŒ MARKETS AFFECTED [List all affected markets] โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ” ๐Ÿ’ฅ BUSINESS IMPACT โ€ข Complete service disruption โ€ข All user functions unavailable โ€ข [Add specific impacts if known] โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ” ๐Ÿ“ž CONTACT FOR QUESTIONS [Your Name] - [Your Contact Method] โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”โ”

๐Ÿ“… Business Hours Ticket

This issue does not meet the criteria for after-hours emergency response.

Action: Create a ticket for the appropriate team to address during normal business hours.

If the situation changes or you learn that more users are affected, you can re-evaluate using this tool.