Unknown Company

GPU Reliability Engineer

san francisco, ca • Posted 1 weeks ago
Onsite Full Time Engineering

EVONA is partnering with a fast-growing space technology company developing high-performance compute infrastructure for Low Earth Orbit.

They’re looking for an experienced GPU RAS / Hardware Validation Engineer to own reliability validation across advanced GPU server platforms.

When hardware is operating hundreds of kilometres above Earth, you can’t simply replace a failed component. These systems need to detect, classify, contain and recover from faults autonomously.

What You’ll Do

  • Develop fault injection, detection and recovery testing
  • Debug failures across GPU, CPU, DDR, HBM, PCIe and NVLink
  • Characterise fault propagation across hardware, firmware and OS layers
  • Validate ECC, MCA/MCI and system recovery mechanisms
  • Work with BMC, IPMI, Redfish and MCTP/PLDM
  • Partner directly with silicon vendors on failure analysis and root cause
  • Build Python automation for testing and log analysis

What We’re Looking For

  • 5+ years in hardware validation, silicon validation or platform reliability
  • Strong CPU/GPU architecture knowledge
  • Experience with DDR/HBM and server-class compute systems
  • Deep understanding of RAS, ECC, fault containment and recovery
  • Experience with PCIe, NVLink and/or XGMI
  • Strong Python or equivalent scripting skills

Why Join?

You’ll be applying cutting-edge GPU and server reliability expertise to an entirely different environment: high-performance computing in space .

US export control requirements apply to this position.

#J-18808-Ljbffr

GPU Reliability Engineer in san francisco at Unknown Company

This position is listed as full time and onsite.

Back to Job Search