Unknown Company

Machine Learning Infrastructure Engineer (Modeling)

san francisco, ca • Posted 2 weeks ago
Onsite Full Time Electrical & Energy Engineering

  • The ML Infrastructure team supports and accelerates PI’s core modeling efforts by building the systems that make large-scale training reliable, reproducible, and fast. The team works closely with research, data, and platform engineers to ensure models can scale from prototype to production-grade training runs
  • Own training/inference infrastructure: Design, implement, and maintain systems for large-scale model training, including scheduling, job management, checkpointing, and metrics/logging
  • Scale distributed training: Work with researchers to scale JAX-based training across TPU and GPU clusters with minimal friction
  • Optimize performance: Profile and improve memory usage, device utilization, throughput, and distributed synchronization
  • Enable rapid iteration: Build abstractions for launching, monitoring, debugging, and reproducing experiments
  • Partner with researchers: Translate research needs into infra capabilities and guide best practices for training at scale
  • Contribute to core training code: Evolve JAX model and training code to support new architectures, modalities, and evaluation metrics

#J-18808-Ljbffr

Machine Learning Infrastructure Engineer (Modeling) in san francisco at Unknown Company

This position is listed as full time and onsite.

Back to Job Search