Senior Site Reliability Engineer

PB consulting

Charlotte, NC / Detroit, MI

Posted On: Aug 25, 2026

Posted On: Aug 25, 2026

Job Overview

Experience

10 - 15 Years

Salary

Depends on Experience

Work Arrangement

On-Site

Travel Requirement

0%

Required Skills

  • SRE/DevOps
  • AWS cloud services
  • Kubernetes
  • Docker
  • Amazon EKS
  • Amazon ECS
Job Description
Job Summary

We are seeking a Senior SRE / DevSecOps Engineer with strong experience in Kubernetes, AWS, container platforms, observability, infrastructure automation, and AI-assisted troubleshooting. The role will focus on platform reliability, incident management, SLO/SLI governance, and automation across cloud-native environments.

Roles and Responsibilities
  • Design, maintain, and troubleshoot Kubernetes and container platforms including EKS, ECS, and Docker.
  • Manage AWS services including EC2, ECS, EKS, Lambda, RDS, S3, and IAM.
  • Implement infrastructure automation using Terraform.
  • Monitor and improve platform reliability using Dynatrace, Splunk, and Grafana.
  • Develop automation and troubleshooting scripts using Python and Bash.
  • Lead incident response, root cause analysis, and production issue resolution.
  • Define and govern SLOs, SLIs, and reliability standards.
  • Apply DevSecOps practices across cloud and containerized environments.
  • Leverage AI/LLM tools, including Claude AI, for pipeline troubleshooting and operational problem solving.
  • Improve system availability, performance, scalability, security, and operational efficiency.
Required Skills & Experience
  • 10+ years of experience in SRE/DevOps.
  • Strong experience with AWS cloud services, including ECS, EKS, EC2, Lambda, RDS, S3, and IAM.
  • Hands-on experience with Kubernetes, Docker, EKS, and ECS.
  • Strong experience with Terraform and Infrastructure as Code.
  • Experience with Dynatrace, Splunk, and Grafana.
  • Strong Python and Bash scripting skills.
  • Must have experience using Claude AI or similar AI/LLM tools for pipeline troubleshooting.
  • Strong incident management and production support experience.
  • Experience with SLO/SLI governance and site reliability practices.
  • Strong understanding of DevSecOps, cloud security, automation, and CI/CD.

Job ID: PC522147


Posted By

Naincy

Technical Recruiter