OG
ML Compiler Engineer
Accepting applicationsOho Group · San Jose, CA
Full-Time Mid_senior AISoC
Posted
18 Jun
Category
Design
Experience
Mid_senior
Country
United States
ML Compiler Engineer – Full Stack
Join a stealth-stage AI hardware startup building a custom SoC and the full compiler and inference stack alongside it.
What You'll Do
You'll own the full stack from PyTorch and Triton down to efficient machine IR on leading-edge GPUs and novel accelerator targets
You'll build and optimise ML compiler infrastructure across every layer, not just one slice
You'll drive GPU optimisation across CUDA and ROCm with real, measurable performance outcomes
You'll collaborate directly with hardware engineers on ISA and microarchitecture tradeoffs that shape the silicon itself
You'll tackle compiler problems on XPU and novel accelerator targets where no established playbook exists
What We're Looking For
3+ years in compilers, ML systems, or equivalent depth
Strong compiler fundamentals with a track record of delivered performance improvements
Hands-on experience with PyTorch, Triton, and low-level code generation
GPU optimisation depth in CUDA and/or ROCm
HW-SW co-design experience or strong evidence of meaningful hardware collaboration
Show more Show less
Join a stealth-stage AI hardware startup building a custom SoC and the full compiler and inference stack alongside it.
What You'll Do
You'll own the full stack from PyTorch and Triton down to efficient machine IR on leading-edge GPUs and novel accelerator targets
You'll build and optimise ML compiler infrastructure across every layer, not just one slice
You'll drive GPU optimisation across CUDA and ROCm with real, measurable performance outcomes
You'll collaborate directly with hardware engineers on ISA and microarchitecture tradeoffs that shape the silicon itself
You'll tackle compiler problems on XPU and novel accelerator targets where no established playbook exists
What We're Looking For
3+ years in compilers, ML systems, or equivalent depth
Strong compiler fundamentals with a track record of delivered performance improvements
Hands-on experience with PyTorch, Triton, and low-level code generation
GPU optimisation depth in CUDA and/or ROCm
HW-SW co-design experience or strong evidence of meaningful hardware collaboration
Show more Show less
Similar Jobs
M
PRINCIPAL RTL DESIGN
Micron · Boise, United States, North America
M
Principal Engineer, High‑Speed RTL Design
Micron · Bangalore, India, Asia
N
ASIC TOP Floorplan Design Engineer
NVIDIA · Shanghai, China, Asia
N
Senior System Software Engineer, SoC Power and Performance - O-RAN Infrastructure
NVIDIA · 3 Locations