We are seeking an experienced Kubernetes Engineer with strong expertise in Kubernetes, OpenShift, cloud platforms, and observability. The ideal candidate will be responsible for Kubernetes administration, platform reliability, monitoring, troubleshooting, and infrastructure automation while driving strong engineering practices.
Required Skills
- Kubernetes administration and troubleshooting
- OpenShift or equivalent Kubernetes platform experience
- Strong Observability experience
- Prometheus and Splunk
- Grafana, Alertmanager, ELK/OpenSearch or similar tools
- Kubernetes networking and container runtimes
- Infrastructure automation using Terraform, Helm, or Ansible
- GitOps tools such as Argo CD or Flux
Roles & Responsibilities
- Administer, configure, troubleshoot, and maintain Kubernetes clusters.
- Manage OpenShift or equivalent Kubernetes platforms in public cloud environments.
- Manage container runtimes and Kubernetes workloads.
- Implement and maintain observability solutions using Prometheus, Grafana, Alertmanager, Splunk, ELK/OpenSearch, or similar tools.
- Configure and monitor metrics, logs, alerts, and distributed tracing where applicable.
- Troubleshoot Kubernetes networking, including Ingress, Services, CNI plugins, and related components.
- Work with service mesh technologies such as Istio or Linkerd as a plus.
- Automate infrastructure and application deployments using Terraform, Helm, Ansible, and GitOps tools such as Argo CD or Flux.
- Monitor platform health, performance, availability, and reliability.
- Identify and resolve infrastructure and application issues proactively.
- Promote best practices for reliability, scalability, security, and operational excellence.
- Collaborate with engineering teams, provide technical guidance, and coach team members on Kubernetes and cloud-native practices.
- Drive continuous improvements in platform stability, automation, and observability.