- Design, implement, operate, and own robust and dependable infrastructure for HPC and ML/AI workloads in a cloud environment (AWS/GCP)
- Lead containerization, deployment, and operation of user- and admin-facing HPC platforms (Slurm, Open On Demand, Prometheus/Grafana)
- Translate stakeholder input into robust, high-performance, scalable, cost-effective computing platforms
- Partner with HPC specialists to capture institutional knowledge and manual processes in IaC workflows
- Develop and maintain infrastructure automation using IaC tools like Terraform and CloudFormation
- Create reusable Terraform modules and enforce standards
- Operationalize containerized solutions using Docker and Kubernetes
- Own the full lifecycle of infrastructure management, from provisioning to operations, support, updating, and teardown of production computing platforms
- Develop and maintain monitoring, logging, and alerting for the infrastructure (e.g., CloudWatch, Prometheus/Grafana)
Requirements
- B.S. in computer science, life science, data science or similar fields
- 6+ years of experience in cloud infrastructure engineering with a proven track record of developing and supporting robust IaC deployments
- Experience managing scientific computing workloads in an enterprise environment
- Advanced experience with at least one of AWS and GCP, including knowledge of core compute and storage services relevant to HPC
- Solid understanding of cloud networking, identity, and security controls
- Prior experience with HPC deployment utilities including AWS ParallelCluster, AWS Parallel Computing Services, and Google Cloud Cluster Toolkit (preferred)
- Proficiency with distributed computing environments, especially EKS/GKE/Kubernetes (preferred)
- Familiarity with HPC environments, job schedulers (Slurm), HPC application containers (Docker, Singularity, Apptainer) and NVIDIA GPU computing (preferred)
Core Competencies
Demonstrates expertise in designing and managing cloud infrastructure for HPC and ML/AI workloads, utilizing IaC tools like Terraform and CloudFormation. Proficient in containerization and orchestration technologies such as Docker and Kubernetes, with a strong focus on operational efficiency and performance optimization.
Highest-signal resume keywords
- Cloud Infrastructure Engineering
- Infrastructure As Code (IaC)
- HPC Deployment Utilities
- Containerization (Docker, Kubernetes)
- Monitoring and Logging (CloudWatch, Prometheus)
ATS Optimization Keywords
Hard Skills
- Infrastructure Management
- Cloud Networking
- AWS
- GCP
- Terraform
- CloudFormation
- Slurm
- Docker
- Kubernetes
- NVIDIA GPU Computing
Industry Keywords
- High-Performance Computing (HPC)
- Machine Learning (ML)
- Artificial Intelligence (AI)
- Distributed Computing
- Enterprise Environment
Tools & Technologies
- AWS ParallelCluster
- AWS Parallel Computing Services
- Google Cloud Cluster Toolkit
- Prometheus
- Grafana