- Collaborate closely with our AI and ML research teams to understand their infrastructure needs and obstacles
- Monitor and optimize the performance of our infrastructure ensuring high availability, scalability, and efficient resource utilization
- Help define and improve important measures of AI researcher efficiency, ensuring that our actions are in line with measurable results
- Collaborate with diverse teams, including researchers, data engineers, and DevOps professionals, to build a seamless and coordinated AI/ML infrastructure ecosystem
- Stay on top of the latest advancements in AI/ML technologies, frameworks, and effective strategies, and promote their implementation within the company
Requirements
- Recent graduate with a MS, PhD or equivalent experience in Computer Science or related field
- Proven experience in AI/ML and HPC workloads and infrastructure
- Hands-on experience in using or operating High Performance Computing (HPC) grade infrastructure
- In-depth knowledge of accelerated computing (e.g., GPU, custom silicon)
- Storage (e.g., Lustre, GPFS, BeeGFS)
- Scheduling & orchestration (e.g., Slurm, Kubernetes, LSF)
- High-speed networking (e.g., Infiniband, RoCE, Amazon EFA)
- Containers technologies (Docker, Enroot)
- Expertise in running and optimizing large-scale distributed training workloads using PyTorch (DDP, FSDP), NeMo, or JAX
- Deep understanding of AI/ML workflows, encompassing data processing, model training, and inference pipelines
- Proficiency in programming & scripting languages such as Python, Go, Bash
- Familiarity with cloud computing platforms (e.g., AWS, GCP, Azure)
- Experience with parallel computing frameworks and paradigms.
- Passion for continual learning and keeping abreast of new technologies and effective approaches in the AI/ML infrastructure field.
- Excellent communication and collaboration skills
Core Competencies
Demonstrates expertise in AI/ML infrastructure, including High Performance Computing (HPC) and accelerated computing technologies. Proficient in optimizing large-scale distributed training workloads and collaborating with diverse teams to enhance AI researcher efficiency.
Highest-signal resume keywords
- AI/ML Infrastructure
- High Performance Computing (HPC)
- Distributed Training Workloads Optimization
- Programming Languages (Python, Go, Bash)
- Cloud Computing Platforms (AWS, GCP, Azure)
ATS Optimization Keywords
Hard Skills
- AI/ML Workloads
- Accelerated Computing
- Storage Technologies (Lustre, GPFS, BeeGFS)
- Scheduling & Orchestration (Slurm, Kubernetes, LSF)
- High-Speed Networking (Infiniband, RoCE, Amazon EFA)
- Container Technologies (Docker, Enroot)
- Parallel Computing Frameworks
- Data Processing
- Model Training
- Inference Pipelines
Soft Skills
- Excellent Communication
- Collaboration Skills
- Passion for Learning
Certifications & Qualifications
- MS or PhD in Computer Science or Related Field