IA

Member of Technical Staff - Hardware Bringup

Accepting applications

Infinity Artificial Intelligence Institute · San Francisco Bay Area

Full-Time Mid_senior AIC++PythonRISC-VRTL
Posted
19h ago
Category
Design
Experience
Mid_senior
Country
United States
Member of Technical Staff - Hardware Bringup

San Francisco, CA · On-site · Full-time

About Infinity
Infinity is an early-stage AI infrastructure research company building the software layer that makes non-NVIDIA chips competitive for AI inference. Rather than relying on scarce human kernel engineers, we use AI to automatically generate, test, and optimize the low-level code that determines how efficiently a chip runs AI models. We've signed or are negotiating design partnerships with d-Matrix, AMD, AWS Trainium, Microsoft (Maia and Nexus), Qualcomm, and others. Founded by Jeremy Nixon (former Google Brain; co-founder of AGI House with Andrej Karpathy), Infinity has raised $15M from investors including the founder of Intercom, the VP of AI at AMD, and the founder of MLCommons. We're headquartered in San Francisco.

About The Role
Bringing a new AI accelerator from bare firmware to running state-of-the-art open-source LLMs takes months to years today. It spans firmware, kernel drivers, toolchains, compilers, kernels, runtime, and the inference stack — almost all of it written by hand, per chip.
We're building Ignition: an agentic system that compresses that to under a day, on any accelerator architecture. A coding-agent controller autonomously bootstraps every layer of the stack, validates each layer against the one below it through a strict test ladder, and exposes a standards-compatible inference interface (vLLM plugin / OpenAI-compatible API) once every gate passes. Ignition took the d-Matrix Corsair — a chip with a proprietary ISA and no existing inference ecosystem — from first hardware access to tensor-parallel matmuls across all 32 compute units in 10 hours, and to three frontier models running end-to-end in 10 days. That work is now live as the Infinity d-Matrix Cloud.
As one of the first engineers building this, you'll own one or more layers of the stack and the agent that generates them. Your work will be foundational to how every new chip in the industry comes online.

Responsibilities
Depending on your strengths, you'll own one or more of the following:
Build the agent controller — the coding-agent loop that consumes hardware specs, fuzzer results, and the current test gate; writes code; runs it on the target; reads results; and iterates — plus the monitoring, escalation, and dependency logic that keeps dozens of components progressing in parallel
Develop hardware characterization probes and structured ISA fuzzing that build a behavioral model of a chip with no complete spec, populating an execution-model-neutral hardware schema
Bootstrap toolchains by wrapping existing tools, generating LLVM backends, or hand-writing raw instruction encoders and ELF/flat-binary loaders when nothing else exists
Implement compiler and codegen strategies that branch on the chip's actual execution model (warp-based, scalar tile mesh, dataflow, flat SIMD, analog MAC)
Generate and validate the kernel library — matmul, attention (MHA/GQA/MLA, flash), normalization, RoPE, MoE, and collectives — against reference implementations and measured-peak performance gates
Handle parallelism and interconnect: topology discovery, collective-algorithm selection by latency regime, and clean degradation to a single device
Build the runtime and serving interface: model loading, paged-attention KV cache, continuous batching, and the vLLM / OpenAI-compatible layer
Maintain the test harness — the generic, per-layer test suite every gate is written against, running across multiple hardware platforms and simulation
You May Be a Good Fit If You
Have real low-level systems experience across at least two of: firmware/bare-metal, Linux kernel/driver development (ioctl, mmap, DMA, interrupts, IOMMU), compilers/codegen (LLVM, MLIR, TableGen), GPU/accelerator kernels (CUDA, ROCm/HIP, Triton, Metal), or ML inference internals (vLLM, attention kernels, KV cache, quantization)
Are comfortable working from incomplete or wrong information — reverse-engineering undocumented behavior, fuzzing an ISA, reading a datasheet that doesn't match the silicon, and building a validated model anyway
Have a test-first instinct, and find it satisfying rather than tedious that every layer must be provably correct against a reference before the layer above is attempted
Have hands-on experience building with coding agents / LLMs — prompting, tool-use loops, evaluating and constraining model output, and designing systems where the model writes the code and tests catch its mistakes
Are fluent in Python and at least one systems language (Rust, C, or C++)
Are drawn to the hardest part of the problem and comfortable being the person who figures out what's actually happening at the lowest level
Strong Candidates May Also Have Experience With
Bringing up a new accelerator, board, or ISA before — vendor-side or from the outside
Contributing to LLVM, MLIR, vLLM, TVM, or a hardware vendor's compiler/runtime stack
Non-GPU execution models (dataflow, wafer-scale, in-memory/analog compute, RISC-V mesh)
RTL simulation (Verilator, Icarus Verilog) for validating against a chip model before hardware exists
Distributed training/inference (NCCL, Megatron-LM, DeepSpeed) and collective-communication internals
Logistics
Deadline to apply: None. Applications are reviewed on a rolling basis.
Location: San Francisco, CA. This role is on-site.
Compensation: $200k - $420k
Show more Show less