Partner with Security, SRE, Production Engineering, Operations, and Platform Engineering teams to improve operational readiness and service reliability.
Develop, maintain, and standardize operational documentation, including runbooks, SOPs, playbooks, and technical procedures.
Build and manage knowledge management frameworks to ensure operational knowledge is accessible, accurate, and up to date.
Capture institutional and tribal knowledge, transforming it into scalable, repeatable operational processes.
Collaborate with engineering and operations teams to document production workflows, incident response procedures, and service recovery processes.
Support incident management activities by documenting response procedures, escalation paths, and post-incident improvements.
Identify process improvement opportunities to enhance operational efficiency and reduce operational risk.
Ensure documentation aligns with security, compliance, and operational best practices.
Facilitate cross-functional communication between technical and non-technical stakeholders.
Required Skills & Qualifications
Experience supporting Site Reliability Engineering (SRE), Production Engineering, Operations, or Platform Engineering organizations.
Strong experience developing and maintaining:
Operational Runbooks
Standard Operating Procedures (SOPs)
Playbooks
Knowledge Management Frameworks
Proven ability to document operational processes and translate complex technical concepts into clear, user-friendly documentation.
Familiarity with:
Incident Management
Production Support
Monitoring and Alerting
Service Reliability and Operational Excellence practices
Demonstrated ability to capture and operationalize tribal knowledge into standardized, repeatable processes.
Strong analytical, organizational, and documentation skills.
Excellent communication and stakeholder management abilities.
Preferred Qualifications
Experience working in enterprise security or cloud infrastructure environments.
Knowledge of ITIL, DevOps, SRE, or operational excellence methodologies.
Experience with documentation and knowledge management platforms such as Confluence, SharePoint, or similar tools.
Familiarity with incident management tools such as ServiceNow, Jira, PagerDuty, or Opsgenie.
Preferred Experience
8+ years of experience in Program Management, Security Operations, Technical Operations, SRE, Platform Engineering, or related technical environments.
Proven track record of driving operational improvements, documentation standards, and cross-functional process optimization in large-scale enterprise environments.