EVONA is partnering with a fast-growing space technology company developing high-performance compute infrastructure for Low Earth Orbit.
They’re looking for an experienced GPU RAS / Hardware Validation Engineer to own reliability validation across advanced GPU server platforms.
When hardware is operating hundreds of kilometres above Earth, you can’t simply replace a failed component. These systems need to detect, classify, contain and recover from faults autonomously.
What You’ll Do
- Develop fault injection, detection and recovery testing
- Debug failures across GPU, CPU, DDR, HBM, PCIe and NVLink
- Characterise fault propagation across hardware, firmware and OS layers
- Validate ECC, MCA/MCI and system recovery mechanisms
- Work with BMC, IPMI, Redfish and MCTP/PLDM
- Partner directly with silicon vendors on failure analysis and root cause
- Build Python automation for testing and log analysis
What We’re Looking For
- 5+ years in hardware validation, silicon validation or platform reliability
- Strong CPU/GPU architecture knowledge
- Experience with DDR/HBM and server-class compute systems
- Deep understanding of RAS, ECC, fault containment and recovery
- Experience with PCIe, NVLink and/or XGMI
- Strong Python or equivalent scripting skills
Why Join?
You’ll be applying cutting-edge GPU and server reliability expertise to an entirely different environment: high-performance computing in space .
US export control requirements apply to this position.
#J-18808-LjbffrGPU Reliability Engineer in san francisco at Unknown Company
This position is listed as full time and onsite.