We're looking for a Mid-Level ML Infrastructure Engineer to build and scale the platforms and tooling that our data scientists and ML engineers rely on to train, deploy, and monitor models. You'll focus less on the models themselves and more on making the systems around them fast, reliable, and easy to use.
What You'll Do
- Design and maintain training and inference infrastructure (pipelines, orchestration, compute scheduling)
- Build internal tooling for experiment tracking, feature stores, and model versioning
- Optimize model serving for latency, throughput, and cost at scale
- Set up and maintain CI/CD pipelines for ML workflows
- Manage GPU/compute resource allocation and cluster infrastructure (e.g., Kubernetes, Ray, Slurm)
- Implement monitoring and alerting for model performance, data drift, and system health
- Partner with ML engineers and data scientists to understand their workflows and remove friction
- Contribute to platform architecture decisions as the ML org scales
What We're Looking For
- 2–5 years of experience in infrastructure, platform, or backend engineering, ideally supporting ML workloads
- Strong proficiency in Python and/or Go
- Experience with containerization and orchestration (Docker, Kubernetes)
- Familiarity with ML-specific tooling: MLflow, Kubeflow, Ray, SageMaker, Vertex AI, or similar
- Experience with cloud infrastructure (AWS, GCP, or Azure) and infrastructure-as-code (Terraform, Pulumi)
- Understanding of distributed systems and data pipeline design
- Comfort working with GPUs and understanding of training/serving performance trade-offs
- Strong communication skills and ability to work cross-functionally with ML practitioners
Nice to Have
- Experience with high-performance model serving (Triton, TorchServe, vLLM)
- Familiarity with data versioning tools (DVC, LakeFS) or feature stores (Feast, Tecton)
- Experience scaling distributed training (Horovod, DeepSpeed, PyTorch DDP)
- Background in SRE or DevOps practices applied to ML systems
What We Offer
- Competitive salary and equity
- Health, dental, and vision insurance
Machine Learning Infrastructure Engineer in new york at Unknown Company
This position is listed as full time and onsite.