E

GPU Reliability Engineer

Accepting applications

EVONA · San Francisco Bay Area

Full-Time Mid_senior DDRPCIePython
Posted
3d ago
Category
Verification
Experience
Mid_senior
Country
United States
EVONA is partnering with a fast-growing space technology company developing high-performance compute infrastructure for Low Earth Orbit.

They’re looking for an experienced GPU RAS / Hardware Validation Engineer to own reliability validation across advanced GPU server platforms.

When hardware is operating hundreds of kilometres above Earth, you can’t simply replace a failed component. These systems need to detect, classify, contain and recover from faults autonomously.

What You’ll Do
• Lead RAS validation across GPU server platforms
• Develop fault injection, detection and recovery testing
• Debug failures across GPU, CPU, DDR, HBM, PCIe and NVLink
• Characterise fault propagation across hardware, firmware and OS layers
• Validate ECC, MCA/MCI and system recovery mechanisms
• Work with BMC, IPMI, Redfish and MCTP/PLDM
• Partner directly with silicon vendors on failure analysis and root cause
• Build Python automation for testing and log analysis

What We’re Looking For
• 5+ years in hardware validation, silicon validation or platform reliability
• Strong CPU/GPU architecture knowledge
• Experience with DDR/HBM and server-class compute systems
• Deep understanding of RAS, ECC, fault containment and recovery
• Hands-on fault injection experience
• Experience with PCIe, NVLink and/or XGMI
• Strong Python or equivalent scripting skills

Why Join?
You’ll be applying cutting-edge GPU and server reliability expertise to an entirely different environment: high-performance computing in space.

US export control requirements apply to this position.

Interested? Apply today or contact EVONA for a confidential conversation
Show more Show less