Responsibilities
- Architect and optimize large-scale distributed pipelines
- Scale up inference throughput across the pipeline
- Design systems that reliably store, index, and serve billions of data points
- Apply deep expertise in distributed systems and frameworks
- Partner with modeling teams to understand data that improves training outcomes
- Operate as a hands-on technical leader
Requirements
- 10+ years of experience in data engineering, ML infrastructure, or distributed systems
- Strong software engineering background
- Proficiency in Python and strong experience in C++, Rust, Go, or Java
- Deep knowledge of databases and storage systems at scale
- Strong ML background, particularly expertise in optimizing GPU inference pipelines
- Experience with data curation for model training
- Bachelor's, Master's, or Ph.D. in Computer Science, Engineering, Machine Learning, or a related field
Core Competencies
Demonstrates extensive expertise in architecting and optimizing large-scale distributed systems and pipelines, with a strong focus on data engineering and machine learning infrastructure. Proficient in Python and experienced in C++, Rust, Go, or Java, with a solid understanding of databases and storage systems at scale.
Certifications & Qualifications
- Bachelor's Degree
- Master's Degree
- Ph.D. in Computer Science
- Ph.D. in Engineering
- Ph.D. in Machine Learning
Soft Skills
- Technical Leadership