Unknown Company

Staff Machine Learning Engineer – ML Frameworks

california, mo • Posted 2 days ago
Onsite Full Time Electrical & Energy Engineering

  • Design, develop, and maintain robust AI/ML infrastructure solutions supporting training and deployment of large-scale AI models using Kubernetes and Python on AWS cloud
  • Implement and improve distributed training frameworks leveraging GPUs to improve performance and scalability
  • Improve resiliency, elasticity, data loading, and out-of-the-box support for FSDP and model parallelism
  • Improve orchestration and scheduling to train better models
  • Scale the number of jobs and enable faster experimentation with AutoML and similar tools
  • Collaborate with data scientists and ML researchers to streamline model training pipelines and ensure efficient resource utilization
  • Drive innovation in infrastructure practices supporting machine learning research and development

Requirements

  • PhD or Master’s in computer science or related field and 5+ years of hands-on industry experience
  • Proven proficiency with Python and developing systems, frameworks and SDKs
  • Experience with infrastructure and understanding of model serving, training, orchestration, and management of GPU resources
  • Experience with machine learning and distributed PyTorch
  • Strong critical thinking, analytical and quantitative problem-solving ability
  • Excellent communication, relationship skills and a strong teammate
  • Experience with KubeFlow, MLFlow, Ray, SageMaker, or similar (added plus)
  • Experience with PyTorch distributed, MPI, Megatron, Horovod and other AI training frameworks (added plus)

Core Competencies

Demonstrates expertise in designing and maintaining AI/ML infrastructure solutions, with a strong focus on Python, Kubernetes, and AWS. Proven ability to enhance distributed training frameworks and optimize resource utilization for large-scale AI model deployment.

Highest-signal resume keywords

  • Python Proficiency
  • Kubernetes Experience
  • Distributed PyTorch Knowledge
  • AI/ML Infrastructure Development
  • GPU Resource Management

ATS Optimization Keywords

Hard Skills

  • AI/ML Infrastructure Solutions
  • Distributed Training Frameworks
  • Model Serving
  • Orchestration and Scheduling
  • AutoML Tools
  • Python Development
  • Critical Thinking
  • Analytical Problem-Solving
  • Quantitative Analysis
  • Machine Learning

Soft Skills

  • Excellent Communication
  • Relationship Skills
  • Team Collaboration

Certifications & Qualifications

  • PhD in Computer Science
  • Master’s in Computer Science

Industry Keywords

  • AI Models
  • Machine Learning Research
  • Infrastructure Practices
  • Resource Utilization
  • Scalability

Tools & Technologies

  • KubeFlow
  • MLFlow
  • Ray
  • SageMaker
  • PyTorch Distributed
  • MPI
  • Megatron
  • Horovod

#J-18808-Ljbffr

Staff Machine Learning Engineer – ML Frameworks in california at Unknown Company

This position is listed as full time and onsite.

Back to Job Search