E
GPU Reliability Engineer
Accepting applicationsEVONA · San Francisco Bay Area
Full-Time Mid_senior DDRPCIePython
Posted
3d ago
Category
Verification
Experience
Mid_senior
Country
United States
EVONA is partnering with a fast-growing space technology company developing high-performance compute infrastructure for Low Earth Orbit.
They’re looking for an experienced GPU RAS / Hardware Validation Engineer to own reliability validation across advanced GPU server platforms.
When hardware is operating hundreds of kilometres above Earth, you can’t simply replace a failed component. These systems need to detect, classify, contain and recover from faults autonomously.
What You’ll Do
• Lead RAS validation across GPU server platforms
• Develop fault injection, detection and recovery testing
• Debug failures across GPU, CPU, DDR, HBM, PCIe and NVLink
• Characterise fault propagation across hardware, firmware and OS layers
• Validate ECC, MCA/MCI and system recovery mechanisms
• Work with BMC, IPMI, Redfish and MCTP/PLDM
• Partner directly with silicon vendors on failure analysis and root cause
• Build Python automation for testing and log analysis
What We’re Looking For
• 5+ years in hardware validation, silicon validation or platform reliability
• Strong CPU/GPU architecture knowledge
• Experience with DDR/HBM and server-class compute systems
• Deep understanding of RAS, ECC, fault containment and recovery
• Hands-on fault injection experience
• Experience with PCIe, NVLink and/or XGMI
• Strong Python or equivalent scripting skills
Why Join?
You’ll be applying cutting-edge GPU and server reliability expertise to an entirely different environment: high-performance computing in space.
US export control requirements apply to this position.
Interested? Apply today or contact EVONA for a confidential conversation
Show more Show less
They’re looking for an experienced GPU RAS / Hardware Validation Engineer to own reliability validation across advanced GPU server platforms.
When hardware is operating hundreds of kilometres above Earth, you can’t simply replace a failed component. These systems need to detect, classify, contain and recover from faults autonomously.
What You’ll Do
• Lead RAS validation across GPU server platforms
• Develop fault injection, detection and recovery testing
• Debug failures across GPU, CPU, DDR, HBM, PCIe and NVLink
• Characterise fault propagation across hardware, firmware and OS layers
• Validate ECC, MCA/MCI and system recovery mechanisms
• Work with BMC, IPMI, Redfish and MCTP/PLDM
• Partner directly with silicon vendors on failure analysis and root cause
• Build Python automation for testing and log analysis
What We’re Looking For
• 5+ years in hardware validation, silicon validation or platform reliability
• Strong CPU/GPU architecture knowledge
• Experience with DDR/HBM and server-class compute systems
• Deep understanding of RAS, ECC, fault containment and recovery
• Hands-on fault injection experience
• Experience with PCIe, NVLink and/or XGMI
• Strong Python or equivalent scripting skills
Why Join?
You’ll be applying cutting-edge GPU and server reliability expertise to an entirely different environment: high-performance computing in space.
US export control requirements apply to this position.
Interested? Apply today or contact EVONA for a confidential conversation
Show more Show less