Lead the design and development of a production inference platform while defining the technical roadmap for inference infrastructure and model serving. Build scalable systems for LLM serving and partner with ML engineers to productionize models.
Requirements
Candidates should have significant experience with production AI inference systems and leadership in LLM serving infrastructure. A strong background in distributed systems and proficiency in programming languages such as Python, Go, Rust, or C++ is also required.
Key Skills
- AI Inference
- Production Systems
- Model Serving
- Optimization
- Monitoring
- Distributed Systems
- Backend Infrastructure
- High-Performance Engineering
- GPU Inference
- Python
- Go
- Rust
- C++
- Technical Architecture
- Communication
- Cross-Team Influence
Benefits
- Health Insurance
- Equity
Staff Software Engineer, AI Inference in new york at Unknown Company
This position is listed as full time and onsite.