We are seeking a Linux Systems Administrator with experience supporting High-Performance Computing (HPC) environments, AI clusters, and research computing infrastructure. The ideal candidate will have strong Linux administration, cluster management, troubleshooting, and performance-tuning skills. Experience supporting research or scientific computing environments, particularly grant-funded research environments, is highly preferred.
Roles and Responsibilities
- Administer, maintain, and support Linux server environments across enterprise and HPC infrastructure.
- Support and maintain HPC and computational computing clusters, ensuring system availability and performance.
- Install, configure, upgrade, patch, and maintain Linux operating systems, applications, and system software.
- Support and administer HPC job scheduling environments, preferably Slurm.
- Monitor system health, resource utilization, performance, and capacity across Linux and HPC environments.
- Troubleshoot complex operating system, application, networking, storage, and hardware issues.
- Perform system and application performance tuning to optimize compute workloads.
- Support storage infrastructure, server hardware, networking, and high-performance computing components.
- Assist with deployment, configuration, and maintenance of AI and GPU-based computing infrastructure.
- Collaborate with research, engineering, infrastructure, and application teams to support computational workloads.
- Maintain system documentation, operational procedures, and technical standards.
- Support upgrades, migrations, and infrastructure improvements while minimizing service disruption.
- Identify opportunities to improve system reliability, performance, automation, and operational efficiency.
Required Qualifications
- Hands-on experience administering Linux server environments.
- Experience supporting HPC or computational computing clusters.
- Experience with HPC job schedulers, preferably Slurm.
- Experience installing, upgrading, patching, configuring, and maintaining Linux systems and applications.
- Strong knowledge of storage infrastructure, networking, and server hardware.
- Strong troubleshooting, systems analysis, and performance-tuning skills.
- Ability to diagnose and resolve complex infrastructure and system-level issues.
Preferred Qualifications
- Experience supporting AI clusters, GPU infrastructure, or AI-focused computing environments.
- Experience with parallel file systems.
- Experience supporting research, scientific, academic, or computational computing environments.
- Experience supporting grant-funded research infrastructure.
- Knowledge of SAN, InfiniBand, and high-performance networking technologies.
- Familiarity with automation, scripting, monitoring, and infrastructure management tools.
- Experience working in large-scale distributed computing environments.