-
Assume the Incident Commander role for all P1 and P2 incidents: own the bridge call, drive resolution, and coordinate cross-functional resolver groups.
Establish roles on incident bridges (scribe, technical lead, communications lead) and enforce time-boxed troubleshooting with 30-minute checkpoints to avoid stagnation.
Make escalation decisions: page on-call engineers, engage leadership or vendors, notify compliance, or execute rollback/failover/service-degradation procedures.
Send initial stakeholder notifications within the SLA and maintain a consistent cadence: technical details for engineering; business impact for leadership.
Coordinate with Compliance and Regulatory teams for impact notifications and initiate customer-facing communications (app banners, status page updates) when required.
Enforce monitoring coverage requirements for all Caesars Digital production services, ensure no service goes to production without adequate monitoring and alerting in place.
Continuously tune alert thresholds based on feedback from Digital System Support Engineers, reducing false positives and alert fatigue while driving the alert signal-to-noise ratio above 80% actionable.
Implement alert deduplication, correlation, and suppression rules to ensure Engineers receive clean, actionable signals rather than noise that degrades response effectiveness.
Define and maintain alert severity standards that clearly distinguish P1 vs. P2 vs. P3 vs. informational alerts, ensuring consistent classification across all services.
Review "missed detection" findings from the Major Incident Manager's Post-Incident Reviews and build new monitoring coverage to prevent recurrence of undetected issues.
Ensure every alert in the ecosystem links to a documented runbook with clear response procedures that Engineers can execute independently.
Build and own end-to-end customer journey monitoring covering the critical user flows: Registration Deposit Bet Placement Bet Settlement Withdrawal, with defined thresholds for success rates and drop-off alerts.
Design and implement real-time revenue monitoring dashboards tracking deposit/withdrawal volumes, payment gateway health, and transaction success rates with anomaly detection against expected baselines.
Build revenue impact calculation models for use during major incidents, enabling the team to quantify business impact in dollar terms.
Define and maintain business KPI dashboards monitoring operational metrics including handle, active users, concurrent sessions, bet volume per minute, and Caesars Rewards pipeline health.
Create executive-visible business health dashboards that provide real-time situational awareness during high-revenue events and peak traffic periods.
Conduct monthly monitoring audits to assess coverage completeness, alert quality, and identify stale or orphaned alerts for decommissioned services.
Build synthetic monitoring for critical customer journeys to validate service availability and performance from the customer's perspective.
Collaborate with the Problem Manager to support event readiness by building enhanced dashboards, lowering detection thresholds, and adding event-specific synthetic monitors per the readiness plan.
Report on noisy alert sources monthly and drive engineering teams to fix the root causes generating non-actionable alerts.
Partner with product and engineering teams to agree on monitoring thresholds and ensure observability is built into the development lifecycle.