We are seeking an experienced PagerDuty Engineer with 8–10 years of experience in incident management, alert orchestration, monitoring integrations, and PagerDuty platform administration. The ideal candidate will have strong hands-on expertise configuring PagerDuty services, escalation policies, event rules, on-call schedules, alert routing, and integrations with enterprise monitoring and ITSM platforms.
Key Responsibilities
- Own end-to-end PagerDuty platform configuration, including services, escalation policies, on-call schedules, notification profiles, event rules, and integrations.
- Design and implement alert urgency, routing, and escalation models based on business criticality.
- Configure time-based routing to prevent non-critical alerts from triggering unnecessary after-hours pages.
- Implement intelligent alert grouping, deduplication, suppression, and event orchestration.
- Drive alert-noise reduction across monitoring and infrastructure platforms.
- Configure event rules for suppression, priority downgrades, deduplication, and alert consolidation.
- Implement flap detection and consolidation for repetitive CI/service alerts.
- Configure maintenance-window integrations for automated alert suppression during planned changes.
- Maintain transparency, reporting, and governance around suppressed alerts.
- Design and manage on-call schedules, including follow-the-sun coverage and rotation policies.
- Audit and optimize alert routing to ensure incidents reach the appropriate first responder.
- Develop and maintain an incident runbook library with diagnostic and remediation procedures.
- Manage integrations with APM, infrastructure, cloud, scheduler, webhook, and other monitoring platforms.
- Collaborate with monitoring SMEs supporting tools such as Dynatrace and SevOne.
- Manage PagerDuty–ServiceNow integration, including incident creation, updates, and lifecycle synchronization.
- Troubleshoot alerting and integration issues and perform root-cause analysis.
Required Skills
- 8–10 years of relevant experience in PagerDuty, incident management, monitoring, or SRE/operations engineering.
- Strong hands-on experience with PagerDuty administration and configuration.
- Expertise in:
- Services and escalation policies
- On-call schedules
- Event rules and event orchestration
- Alert routing and urgency
- Alert suppression and deduplication
- Alert grouping and noise reduction
- Maintenance windows
- Incident management and runbooks
- Experience integrating PagerDuty with ServiceNow.
- Experience with monitoring/APM platforms such as Dynatrace, SevOne, or similar tools.
- Strong understanding of webhooks, APIs, integrations, and event pipelines.
- Excellent troubleshooting and incident-management skills.
- Strong communication and collaboration skills.