We are seeking a skilled GPUaaS Kubernetes Platform Engineer to design, operate, and support scalable GPU-enabled cloud infrastructure. The role focuses on managing Kubernetes and OpenShift platforms, enabling AI/ML workloads, optimizing GPU resource utilization, and ensuring the reliability and performance of high-performance computing environments.
Roles and Responsibilities
- Operate and maintain Kubernetes and OpenShift GPU platforms supporting AI, ML, and high-performance computing workloads.
- Configure, deploy, and manage GPU-enabled infrastructure, including accelerator resources, GPU scheduling, and workload optimization.
- Enable and support AI/ML workloads through scalable GPU-as-a-Service (GPUaaS) platforms.
- Implement and maintain CI/CD pipelines for GPU-based applications and infrastructure deployments.
- Manage platform scalability, availability, and performance to ensure reliable GPU service delivery.
- Perform Kubernetes/OpenShift administration, including cluster configuration, upgrades, monitoring, and troubleshooting.
- Develop automation for GPU platform operations, provisioning, and lifecycle management.
- Monitor GPU utilization, capacity planning, and resource optimization across environments.
- Create and maintain operational dashboards, scaling reports, and platform performance metrics.
- Develop and maintain technical documentation, runbooks, troubleshooting guides, and operational procedures.
- Collaborate with AI/ML teams, infrastructure teams, and DevOps engineers to deliver enterprise-grade GPU services.
Required Skills & Experience
- Strong experience with Kubernetes administration and container orchestration.
- Hands-on experience managing OpenShift environments.
- Experience operating GPU-enabled Kubernetes platforms and accelerator workloads.
- Knowledge of GPU infrastructure concepts, scheduling, and resource management.
- Experience with CI/CD tools and DevOps practices.
- Strong understanding of Linux systems, networking, storage, and cloud infrastructure.
- Experience with monitoring, troubleshooting, and performance optimization of distributed platforms.
- Ability to create operational documentation, runbooks, and support procedures.
Preferred Skills
- Experience with AI/ML infrastructure platforms and MLOps environments.
- Knowledge of GPU technologies such as NVIDIA GPU architectures, CUDA, and GPU operators.
- Experience with infrastructure automation tools such as Terraform, Ansible, or similar technologies.
- Familiarity with cloud platforms and hybrid infrastructure environments.
- Experience supporting large-scale production Kubernetes platforms.