An AI Unicorn startup is hiring a Senior Machine Learning Inference Engineer for a full-time role.
You will be responsible for improving efficiency for AI-native infrastructure powered by generative and multimodal models. The ideal candidate has over 3 years of professional experience and a strong understanding of GPU infrastructure, Python, and PyTorch. This is a highly autonomous role with significant ownership across inference systems and model performance in production.
This role is hybrid in San Francisco Bay Area and offers full benefits and equity.
Experience:
- Building AI applications at scale from the ground up
- Strong understanding of GPU infrastructure including Triton, TensorRT, or vLLM frameworks
- Hands-on experience with Python and PyTorch
- Building model-serving Microservices
- Diffusion and Multimodal model experience is a plus
- Equity
- $401k matching
- Medical coverage
Machine Learning Inference Engineer in san francisco at Unknown Company
This position is listed as full time and hybrid.