Unknown Company

Engineering Manager, ML Performance

sunnyvale, ca • Posted 4 days ago
Onsite Full Time IT & Technology

Benefits

  • Health, dental, vision, life, disability insurance
  • Retirement Benefits: 401(k) with company match
  • Paid Time Off: 20 days of vacation per year, accruing at a rate of 6.15 hours per pay period for the first five years of employment
  • Sick Time: 40 hours/year (increased to 69 hours/year for Seattle) including 5 discretionary sick days per instance
  • Maternity Leave (Short-Term Disability + Baby Bonding): 28-30 weeks
  • Baby Bonding Leave: 18 weeks
  • Holidays: 13 paid days per year

Note: By applying to this position you will have an opportunity to share your preferred working location from the following: Sunnyvale, CA, USA; Kirkland, WA, USA .

Requirements

  • Bachelor’s degree or equivalent practical experience.
  • 8 years of experience in software development.
  • 5 years of experience leading ML design and optimizing ML infrastructure (e.g., model deployment, model evaluation, data processing, debugging, fine tuning).
  • 3 years of experience in a technical leadership role.
  • 2 years of experience in a people management or team leadership role.
  • Experience with ML performance analysis, benchmarking, and computer architecture.

Preferred Qualifications

  • Master’s degree or PhD in Engineering, Computer Science, or a related technical field.
  • 3 years of experience working in a complex, matrixed organization involving cross‑functional or cross‑business projects.
  • Experience in ML accelerators (GPUs, TPUs) and low‑level kernel programming/tuning using tools like CUDA, Triton, or Pallas.
  • Experience with compiler optimization (MLIR, OpenXLA) and integrating frameworks/serving libraries (PyTorch, JAX, vLLM) to maximize hardware efficiency.
  • Ability to adapt ML models to specific hardware strengths and use performance benchmarking to guide both optimization and future hardware design.

Responsibilities

  • Lead a team of software engineers focused on identifying and maintaining ML training and serving benchmarks that are representative to Google production and the broader ML industry.
  • Achieve performance for customer launches, and for engaged benchmark submissions (ML Commons, InferenceX, etc.).
  • Use benchmarks to identify performance opportunities and drive both near‑term state‑of‑the‑art kernel development and out‑of‑the‑box performance (compiler/runtime optimizations, agentic tooling, auto‑sharding) directly and in collaboration with partner teams.
  • Participate in algorithmic innovations exploiting new TPU hardware features and model‑preserving optimizations (speculative decoding, sparsity, quantization, LoRA, etc.).
  • Co‑design models that are TPU‑friendly to showcase model quality and performance advancement for open‑source models typically designed on GPUs.

Compensation

US: $207,000 - $301,000 (USD) + 20% bonus target + equity + benefits.

Equal Opportunity Employer

Google is proud to be an equal opportunity and affirmative action employer. We are committed to building a workforce that is representative of the users we serve, creating a culture of belonging, and providing an equal employment opportunity regardless of race, creed, color, religion, gender, sexual orientation, gender identity or expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition (including breastfeeding), expecting or parents‑to‑be, criminal histories consistent with legal requirements, or any other basis protected by law. See also Google’s EEO Policy, Know your rights: workplace discrimination is illegal, Belonging at Google, and How we hire.

Google is a global company and, in order to facilitate efficient collaboration and communication globally, English proficiency is a requirement for all roles unless stated otherwise in the job posting.

#J-18808-Ljbffr
Back to Job Search