We are working with a fast-growing AI startup building the operating brain for the supply chain. They’ve grown 10x in the last year with a small engineering team and are now building out the model layer underneath their production AI systems.
They’re looking for their first dedicated ML Engineer to own models end-to-end, from raw data through to production. You’ll work with years of real-world operational data across 500k+ SKUs , building systems that directly impact how the business operates.
This is not a research role , and it’s not an LLM-wrapper role. They’re looking for someone who can build, deploy and operate production ML systems - and take ownership when reality changes.
What you'll own
- Build production forecasting models across messy, intermittent and seasonal demand, including cold-start SKUs, promotions, perishability and long-tail demand
- Build datasets and fine-tune models using LoRA / PEFT , with rigorous evaluations determining what actually ships to production
- Build the representation layer that allows AI systems to reason across inconsistent products, vendors, pack sizes and units of measure
- Own the infrastructure around those models, including deployment, versioning, monitoring, drift detection and automated retraining
- Build large-scale ML and data workloads using Python + Spark
- Work with AWS SageMaker, S3, Glue + Step Functions
- Build production inference and evaluation infrastructure
- Use MLflow, Kubeflow or equivalent MLOps tooling
- Contribute outside the model layer when needed, including enough TypeScript/React to work across the wider product
There are no handoffs . You’ll build the model, put it into production, monitor it and fix it when reality changes.
What we're looking for
- 5-7 years of experience building production ML systems
- Experience building and maintaining time-series forecasting models serving production traffic
- Hands-on experience with AWS SageMaker
- Experience fine-tuning LLMs using LoRA or PEFT on real datasets
- Experience building systems backed by ontologies or knowledge graphs
- Strong experience engineering large-scale data pipelines with Spark
- Experience owning production models through deployment, monitoring, drift detection and retraining
- Strong architecture skills, with the ability to explain and defend technical decisions in detail
- Comfortable working across the full ML lifecycle rather than owning just one part of the process
They’re looking for someone who can talk about what happened after the model shipped - when it degraded, how you detected it, what it got wrong and what you changed.
#J-18808-LjbffrSenior AI/ML Engineer in san francisco at Unknown Company
This position is listed as full time and onsite.