TC

Global Inference Library Engineer

Accepting applications

Touring Capital · San Francisco Bay Area

Full-Time Mid AIC++JavaMachine LearningPython
Posted
2d ago
Category
Design
Experience
Mid
Country
United States
Infinity Artificial Intelligence Institute San Francisco Bay Area

Global Inference Library Engineer

Infinity Artificial Intelligence Institute San Francisco Bay Area

1 week ago 59 applicants

See who Infinity Artificial Intelligence Institute has hired for this role

Save

Report this job

Global Inference Library - Member of Technical Staff

Company: Infinity

Team: Systems / AI Infrastructure
Location: San Francisco (on-site)
Type: Full-time The Mission Every accelerator that comes online needs an inference stack, and today that stack is hand-built per chip, per model, per optimization - a permanent, growing backlog of engineering work that scales linearly with the number of chips and models in the world.

We're building Infy: a global inference library, generated and continuously maintained by AI, that targets every major chip - NVIDIA, AMD, Trainium, TPU, Maia, Cerebras, Groq, Tenstorrent, and dozens more - across every model category, from LLMs to vision, audio, and robotics. Think vLLM, but for every accelerator on the market, kept automatically up to date as chips, models, and optimization techniques change.

What You'll Work On

Depending on your strengths, you'll own one or more layers of the system:

Optimization agents - the strategy → analyze → generate → curate loop that turns a chip + model pair into an optimized kernel or serving path, and the orchestrator that runs this across many chips and models in parallel.
Kernel optimization loop - the hierarchical agent system that iterates on individual kernels (matmul, attention variants, normalization, RoPE, MoE, collectives) against reference implementations and measured-peak performance gates, drawing on and growing a registry of thousands of known kernels.
Library generator - the system that takes a validated set of optimized components and assembles them into a real, installable inference library per chip, with a standards-compatible serving interface.
Hardware probe - agents that build a structured representation of a chip's memory hierarchy, compute regions, and inter-component communication characteristics, so optimization strategy can branch correctly on the chip's actual architecture.
Coverage tracking - the enablement and optimization tables that track, per chip and per model category, what's implemented and how fast it runs, and that drive prioritization of what to build next.
SOTA paper replication - agents that read inference-optimization papers and automatically implement and validate the techniques they describe, so the library's optimizations keep pace with published research rather than lagging behind it.
Continuous tracking - the pipeline that watches for new papers, new hardware, and new models, and feeds that into what Infy builds and rebuilds next.
Benchmark harness - the metrics layer (TTFT, TPOT, ITL, E2EL, goodput, and model-category-specific equivalents like image-generation latency) that every optimization is measured against.
Inference service - the gateway, router, and billing layers that turn the library into a running, AI-optimized inference service people can actually call.

What we're looking for

We care more about depth and range than a specific checklist, but strong candidates will have most of:

Real experience with ML inference internals - vLLM or similar serving stacks, attention kernel variants, KV cache management, quantization, continuous batching.
Comfort working across heterogeneous hardware - GPU kernels (CUDA, ROCm/HIP, Triton), and ideally exposure to non-GPU execution models (systolic arrays, dataflow, wafer-scale, in-memory/analog compute).
Hands-on experience building with coding agents / LLMs - prompting, tool-use loops, evaluating and constraining model output, and designing systems where the model writes and optimizes code while tests and benchmarks catch regressions.
Fluency in Python and at least one systems language (Rust, C, or C++).
Comfort reading and implementing ideas directly from research papers, not just from existing open-source code.

Nice to have

Contributed to vLLM, TVM, LLVM, MLIR, or a hardware vendor's compiler/runtime stack.
Experience with distributed inference or training internals (NCCL, Megatron-LM, DeepSpeed) and collective-communication algorithms.
Familiarity with model categories beyond LLMs - vision, audio, multimodal, or robotics inference.
Experience building or maintaining a benchmark suite or performance regression system at scale.
Background reproducing results from ML systems papers (KernelBench-style benchmarks, evolutionary/search-based kernel optimization).

Who you are

You want to build the thing that makes every chip's inference performance a solved problem instead of a standing engineering project. You think in terms of systems that scale across hardware and models rather than one-off implementations, you're energized by AI agents doing the generation and optimization work under a tight benchmark harness, and you want to help set the global standard for how inference libraries get built and maintained.

Infinity is an early-stage AI infrastructure research company building the software layer that makes non-NVIDIA chips competitive for AI inference. Rather than relying on scarce human kernel engineers, we use AI to automatically generate, test, and optimize the low-level code that determines how efficiently a chip runs AI models. We've signed or are negotiating design partnerships with d-Matrix, AMD, AWS Trainium, Microsoft (Maia and Nexus), Qualcomm, and others. Founded by Jeremy Nixon (former Google Brain; co-founder of AGI House with Andrej Karpathy), Infinity has raised $15M from investors including the founder of Intercom, the VP of AI at AMD, and the founder of MLCommons. We're headquartered in San Francisco.

Seniority level Entry level
Employment type Full-time
Job function Engineering and Information Technology
Industries Software Development

Referrals increase your chances of interviewing at Infinity Artificial Intelligence Institute by 2x

See who you know

Get notified about new Engineer jobs in San Francisco Bay Area.

Sign in to create job alert

Similar jobs

AI Inference Performance Engineer

AI Inference Performance Engineer

NVIDIA

Santa Clara, CA 3 weeks ago

Post-Training Platform Infrastructure Engineer

Post-Training Platform Infrastructure Engineer

AMD

San Jose, CA $204,000 - $306,000 1 week ago

Member of Technical Staff - Model Serving / API Backend Engineer

Member of Technical Staff - Model Serving / API Backend Engineer

Black Forest Labs

San Francisco, CA 4 days ago

LLM Inference Engineer

LLM Inference Engineer

NEAR AI

San Francisco Bay Area 2 weeks ago

Member of Technical Staff, Inference

Member of Technical Staff, Inference

Inferact

San Francisco, CA $200,000 - $400,000 1 month ago

Senior Software Engineer, AI Inference Systems

Senior Software Engineer, AI Inference Systems

NVIDIA

Santa Clara, CA 12 hours ago

Systems Generalist, GPT Infrastructure

Systems Generalist, GPT Infrastructure

OpenAI

San Francisco, CA

$293,000.00

$445,000.00

6 days ago

Software Engineer, Model Inference

Software Engineer, Model Inference

OpenAI

San Francisco, CA

$295,000.00

$555,000.00

2 weeks ago

Platform / Inference Optimization Engineer

Platform / Inference Optimization Engineer

Kaon (prev. FlowGPT)

San Francisco Bay Area 5 months ago

Member of Technical Staff — Inference Palo Alto, CA

Member of Technical Staff — Inference Palo Alto, CA

RadixArk

Palo Alto, CA 2 weeks ago

LLM Inference Engineer

LLM Inference Engineer

Hippocratic AI

Menlo Park, CA 2 weeks ago

Member of Technical Staff - Inference

Member of Technical Staff - Inference

Prime Intellect

San Francisco, CA 2 months ago

Runtime Engineer

Runtime Engineer

MatX

Mountain View, CA 2 weeks ago

Applied AI Inference Engineer

Applied AI Inference Engineer

Crusoe

San Francisco, CA 2 days ago

LLM Inference Engineer

LLM Inference Engineer

Majestic Labs ai

Los Altos, CA 1 month ago

AI/ML Platform Engineer

AI/ML Platform Engineer

AMD

Santa Clara, CA

$204,000.00

$306,000.00

1 day ago

Software Engineer, Inference

Software Engineer, Inference

Luma

San Francisco Bay Area 14 hours ago

Member of Technical Staff, Model Efficiency

Member of Technical Staff, Model Efficiency

Cohere

San Francisco, CA 1 day ago

Senior Software Engineer - Performance

Senior Software Engineer - Performance

Microsoft

Mountain View, CA 2 weeks ago

Member of Technical Staff, Performance and Scale

Member of Technical Staff, Performance and Scale

Inferact

San Francisco, CA

$200,000.00

$400,000.00

6 months ago

Inference Engineer

Inference Engineer

Cartesia

San Francisco, CA 2 weeks ago

Sr. Software Engineer, AI Infrastructure

Sr. Software Engineer, AI Infrastructure

LinkedIn

Sunnyvale, CA 3 days ago

Lead Member of Technical Staff, Inference Infrastructure

Lead Member of Technical Staff, Inference Infrastructure

Cohere

San Francisco, CA 3 weeks ago

Distributed LLM Inference Engineer

Distributed LLM Inference Engineer

Anyscale

San Francisco, CA 1 week ago

Principal Software Engineer - Performance

Principal Software Engineer - Performance

Microsoft

Mountain View, CA 2 weeks ago

Staff Engineer, Inference Optimizations

Staff Engineer, Inference Optimizations

DigitalOcean

San Francisco, CA 7 hours ago

Senior ML Infrastructure Engineer (Compute)

Senior ML Infrastructure Engineer (Compute)

General Motors

Sunnyvale, CA 1 week ago

People also viewed

Cloud Inference Engineer

Cloud Inference Engineer

Luminal

San Francisco, CA $150,000 - $250,000 9 months ago

AI/Backend Engineer | San Francisco | $250k + equity

AI/Backend Engineer | San Francisco | $250k + equity

Harrison Clarke

San Francisco Bay Area 3 weeks ago

Member of Technical Staff - ML Systems & Inference

Member of Technical Staff - ML Systems & Inference

Gimlet Labs

San Francisco, CA $150,000 - $350,000 2 months ago

Staff+ Software Engineer, Inference Runtime

Staff+ Software Engineer, Inference Runtime

Anthropic

San Francisco Bay Area 2 weeks ago

Member of Technical Staff - ML Infrastructure & Performance

Member of Technical Staff - ML Infrastructure & Performance

Embedding VC

San Mateo, CA 7 months ago

Performance Engineer

Performance Engineer

Anthropic

San Francisco, CA 2 weeks ago

Senior ML Infrastructure Engineer (Compute)

Senior ML Infrastructure Engineer (Compute)

General Motors

Mountain View, CA 1 week ago

Principal LLM Inference Engineer

Principal LLM Inference Engineer

d-Matrix

Santa Clara, CA 1 day ago

LLM Inference Frameworks and Optimization Engineer

LLM Inference Frameworks and Optimization Engineer

Together AI

San Francisco, CA 1 week ago

Software Engineer, Inference - Multi Modal

Software Engineer, Inference - Multi Modal

OpenAI

San Francisco, CA $295,000 - $555,000 2 weeks ago

Similar Searches

Site Civil Engineer jobs

33,448 open jobs

Data Engineer jobs

192,126 open jobs

Software Engineer jobs

300,699 open jobs

Network Security Officer jobs

5,280 open jobs

Java Software Engineer jobs

36,180 open jobs

Senior Reservoir Engineer jobs

246 open jobs

Civil Engineer jobs

43,143 open jobs

Electrical Software Engineer jobs

223,433 open jobs

Developer jobs

258,935 open jobs

Senior Data Engineer jobs

78,076 open jobs

Senior Software Engineer jobs

78,145 open jobs

Full Stack Engineer jobs

38,546 open jobs

Graduate Civil Engineer jobs

13,551 open jobs

Site Reliability Engineer jobs

169,128 open jobs

Associate Software Engineer jobs

223,979 open jobs

Machine Learning Engineer jobs

148,937 open jobs

Mechanical Engineer jobs

46,392 open jobs

Project Test Engineer jobs

8,924 open jobs

Python Developer jobs

46,642 open jobs

Roads Engineer jobs

17,373 open jobs

Big Data Developer jobs

26,392 open jobs

Network Engineer jobs

173,828 open jobs

Senior Network Engineer jobs

26,545 open jobs

Engineer in Training jobs

143,522 open jobs

Design Verification Engineer jobs

5,299 open jobs

Explore top content on LinkedIn

Find curated posts and insights for relevant topics all in one place.

View top content
Show more Show less