Unknown Company

ML Infrastructure Engineer

san mateo, ca • Posted 3 days ago
Onsite Full Time General

Inference And Model-Serving Infrastructure EngineerYou'll own our inference and model-serving infrastructure end to end. This isn't a research role. It's a build role: you're setting up and scaling the systems that let our agents actually run in production, fast and reliably, at increasing concurrency.You report to Sofus and work closely with our ML and infra teams.What You'll OwnSet up and scale inference/Ray Serve for ML and LLM model serving, integrated with our data analysis and agent workflowsScale agent GPU infrastructure for concurrency and efficiency across multiple agent workloadsOptimize and improve the engine builder and model server that power scalable agent orchestrationWhat You NeedProven ability to build scalable ML/AI platforms from scratch, end-to-end, for production use cases.

You've owned a zero-to-one build before, or can show you're capable of itDeep understanding of the inference stack: vLLM, KV cache, and the optimization layers underneath model servingExperience building distributed systems for AI/ML workloads at scale, connecting them to real product or vertical integrations3+ years of relevant experience. We care about capability, not tenureNice to HaveRay / Ray Serve experienceFamiliarity with AIBrixWhy JoinA rare chance to shape both company and product direction as an early team engineerWork alongside engineers and researchers from LinkedIn, Visa, Meta, and BranchOnsite culture in San Mateo, built for deep collaboration and high-velocity buildingFull benefits (medical, dental, vision, 401k)We sponsor H-1B visas and assist with immigrationWe value builders over résumés. If this role excites you but you don't check every box, we still want to hear from you.

zaimler is an equal opportunity employer.

Back to Job Search