Unknown Company

Senior Software Engineer, ML Infrastructure

cupertino, ca • Posted 2 days ago
Onsite Full Time Software Companies

The RoleWe're looking for a Senior Software Engineer to join our ML Infrastructure team and support the foundational infrastructure that powers Gridmatic. Our platform challenges are shaped by the nature of energy markets: forecasts and trading decisions run on tight schedules, battery dispatch commands must execute reliably in real time, and ML models need to train and deploy continuously as new data arrives.What You'll DoDesign and build the compute platform that runs Gridmatic's production services, batch jobs, and ML training workloadsHelp own our production workflow orchestration end to end, from the cluster and node pools it runs on to the abstractions our teams build on top of itMake our ML iteration cycle faster and cheaper by profiling workflows, cutting latency, and improving how we use computeBuild the observability and developer tooling that lets engineers understand and troubleshoot their own workloads, from post-run cost reporting to monitoring and alertingImprove the reliability and resiliency of production workloads, including how we handle capacity constraints and multi-region routingManage GPU and accelerator capacity: node pools, drivers, spot vs. on-demand tradeoffs, and scheduling and quota so training and batch jobs get the compute they need without overspendingOwn autoscaling and quota management so the platform scales up under load and down to zero when idleWork closely with the ML team to find pain points and quickly ship solutionsMake architectural decisions that shape how we build software as we growWhat We're Looking ForSignificant experience building and operating production infrastructure on a public cloud platform (we run on GCP, but AWS or Azure experience translates well)Strong distributed systems and infrastructure skills: standing up services, scaling and debugging Kubernetes (we run on GKE and it's foundational to our platform), writing Terraform, and comfort with cloud networking, IAM, and secrets managementHands-on experience with workflow orchestration tools like Flyte, Temporal, or AirflowProficiency in Python, and either already know Go or have experience with a similar systems language (C++, Java, Rust) and are excited to work in Python and Go day-to-dayExperience working closely with ML engineers or data scientists as your customers, and genuine excitement to keep doing itClear communication, whether writing a design doc, reviewing code, or explaining a complex system to someone new to itNice to HaveExperience running GPU or accelerator workloads on Kubernetes (node pools, drivers, scheduling, quota) is strongly preferred for this roleFamiliarity with observability tooling (we use Grafana + Google Cloud Monitoring)Experience building internal developer platforms or tooling that other engineers rely onPrior work in domains where latency and reliability have direct business consequencesTaking care of you today:- Continuing Education Opportunities- Flexible PTO- Medical, Dental and Vision plans with competitive employer contributions- Pre-Tax commuter benefits- $1500/year non profit donation matching program through Millie- Home Office StipendProtecting your future for you and your family:- 401K contribution match up to 4%- Company-paid parental leave- Company Paid Life Insurance- Stock Option Loan ProgramPay range $209,000—$270,000 USD

Senior Software Engineer, ML Infrastructure in cupertino at Unknown Company

This position is listed as full time and onsite.

Back to Job Search