- Design, develop, and maintain robust AI/ML infrastructure solutions supporting training and deployment of large-scale AI models using Kubernetes and Python on AWS cloud
- Implement and improve distributed training frameworks leveraging GPUs to improve performance and scalability
- Improve resiliency, elasticity, data loading, and out-of-the-box support for FSDP and model parallelism
- Improve orchestration and scheduling to train better models
- Scale the number of jobs and enable faster experimentation with AutoML and similar tools
- Collaborate with data scientists and ML researchers to streamline model training pipelines and ensure efficient resource utilization
- Drive innovation in infrastructure practices supporting machine learning research and development
Requirements
- PhD or Master’s in computer science or related field and 5+ years of hands-on industry experience
- Proven proficiency with Python and developing systems, frameworks and SDKs
- Experience with infrastructure and understanding of model serving, training, orchestration, and management of GPU resources
- Experience with machine learning and distributed PyTorch
- Strong critical thinking, analytical and quantitative problem-solving ability
- Excellent communication, relationship skills and a strong teammate
- Experience with KubeFlow, MLFlow, Ray, SageMaker, or similar (added plus)
- Experience with PyTorch distributed, MPI, Megatron, Horovod and other AI training frameworks (added plus)
Core Competencies
Demonstrates expertise in designing and maintaining AI/ML infrastructure solutions, with a strong focus on Python, Kubernetes, and AWS. Proven ability to enhance distributed training frameworks and optimize resource utilization for large-scale AI model deployment.
Highest-signal resume keywords
- Python Proficiency
- Kubernetes Experience
- Distributed PyTorch Knowledge
- AI/ML Infrastructure Development
- GPU Resource Management
ATS Optimization Keywords
Hard Skills
- AI/ML Infrastructure Solutions
- Distributed Training Frameworks
- Model Serving
- Orchestration and Scheduling
- AutoML Tools
- Python Development
- Critical Thinking
- Analytical Problem-Solving
- Quantitative Analysis
- Machine Learning
Soft Skills
- Excellent Communication
- Relationship Skills
- Team Collaboration
Certifications & Qualifications
- PhD in Computer Science
- Master’s in Computer Science
Industry Keywords
- AI Models
- Machine Learning Research
- Infrastructure Practices
- Resource Utilization
- Scalability
Tools & Technologies
- KubeFlow
- MLFlow
- Ray
- SageMaker
- PyTorch Distributed
- MPI
- Megatron
- Horovod
Staff Machine Learning Engineer – ML Frameworks in california at Unknown Company
This position is listed as full time and onsite.