TC
Global Inference Library Engineer
Accepting applicationsTouring Capital · San Francisco Bay Area
Full-Time Mid AIC++JavaMachine LearningPython
Posted
2d ago
Category
Design
Experience
Mid
Country
United States
Infinity Artificial Intelligence Institute San Francisco Bay Area
Global Inference Library Engineer
Infinity Artificial Intelligence Institute San Francisco Bay Area
1 week ago 59 applicants
See who Infinity Artificial Intelligence Institute has hired for this role
Save
Report this job
Global Inference Library - Member of Technical Staff
Company: Infinity
Team: Systems / AI Infrastructure
Location: San Francisco (on-site)
Type: Full-time The Mission Every accelerator that comes online needs an inference stack, and today that stack is hand-built per chip, per model, per optimization - a permanent, growing backlog of engineering work that scales linearly with the number of chips and models in the world.
We're building Infy: a global inference library, generated and continuously maintained by AI, that targets every major chip - NVIDIA, AMD, Trainium, TPU, Maia, Cerebras, Groq, Tenstorrent, and dozens more - across every model category, from LLMs to vision, audio, and robotics. Think vLLM, but for every accelerator on the market, kept automatically up to date as chips, models, and optimization techniques change.
What You'll Work On
Depending on your strengths, you'll own one or more layers of the system:
Optimization agents - the strategy → analyze → generate → curate loop that turns a chip + model pair into an optimized kernel or serving path, and the orchestrator that runs this across many chips and models in parallel.
Kernel optimization loop - the hierarchical agent system that iterates on individual kernels (matmul, attention variants, normalization, RoPE, MoE, collectives) against reference implementations and measured-peak performance gates, drawing on and growing a registry of thousands of known kernels.
Library generator - the system that takes a validated set of optimized components and assembles them into a real, installable inference library per chip, with a standards-compatible serving interface.
Hardware probe - agents that build a structured representation of a chip's memory hierarchy, compute regions, and inter-component communication characteristics, so optimization strategy can branch correctly on the chip's actual architecture.
Coverage tracking - the enablement and optimization tables that track, per chip and per model category, what's implemented and how fast it runs, and that drive prioritization of what to build next.
SOTA paper replication - agents that read inference-optimization papers and automatically implement and validate the techniques they describe, so the library's optimizations keep pace with published research rather than lagging behind it.
Continuous tracking - the pipeline that watches for new papers, new hardware, and new models, and feeds that into what Infy builds and rebuilds next.
Benchmark harness - the metrics layer (TTFT, TPOT, ITL, E2EL, goodput, and model-category-specific equivalents like image-generation latency) that every optimization is measured against.
Inference service - the gateway, router, and billing layers that turn the library into a running, AI-optimized inference service people can actually call.
What we're looking for
We care more about depth and range than a specific checklist, but strong candidates will have most of:
Real experience with ML inference internals - vLLM or similar serving stacks, attention kernel variants, KV cache management, quantization, continuous batching.
Comfort working across heterogeneous hardware - GPU kernels (CUDA, ROCm/HIP, Triton), and ideally exposure to non-GPU execution models (systolic arrays, dataflow, wafer-scale, in-memory/analog compute).
Hands-on experience building with coding agents / LLMs - prompting, tool-use loops, evaluating and constraining model output, and designing systems where the model writes and optimizes code while tests and benchmarks catch regressions.
Fluency in Python and at least one systems language (Rust, C, or C++).
Comfort reading and implementing ideas directly from research papers, not just from existing open-source code.
Nice to have
Contributed to vLLM, TVM, LLVM, MLIR, or a hardware vendor's compiler/runtime stack.
Experience with distributed inference or training internals (NCCL, Megatron-LM, DeepSpeed) and collective-communication algorithms.
Familiarity with model categories beyond LLMs - vision, audio, multimodal, or robotics inference.
Experience building or maintaining a benchmark suite or performance regression system at scale.
Background reproducing results from ML systems papers (KernelBench-style benchmarks, evolutionary/search-based kernel optimization).
Who you are
You want to build the thing that makes every chip's inference performance a solved problem instead of a standing engineering project. You think in terms of systems that scale across hardware and models rather than one-off implementations, you're energized by AI agents doing the generation and optimization work under a tight benchmark harness, and you want to help set the global standard for how inference libraries get built and maintained.
Infinity is an early-stage AI infrastructure research company building the software layer that makes non-NVIDIA chips competitive for AI inference. Rather than relying on scarce human kernel engineers, we use AI to automatically generate, test, and optimize the low-level code that determines how efficiently a chip runs AI models. We've signed or are negotiating design partnerships with d-Matrix, AMD, AWS Trainium, Microsoft (Maia and Nexus), Qualcomm, and others. Founded by Jeremy Nixon (former Google Brain; co-founder of AGI House with Andrej Karpathy), Infinity has raised $15M from investors including the founder of Intercom, the VP of AI at AMD, and the founder of MLCommons. We're headquartered in San Francisco.
Seniority level Entry level
Employment type Full-time
Job function Engineering and Information Technology
Industries Software Development
Referrals increase your chances of interviewing at Infinity Artificial Intelligence Institute by 2x
See who you know
Get notified about new Engineer jobs in San Francisco Bay Area.
Sign in to create job alert
Similar jobs
AI Inference Performance Engineer
AI Inference Performance Engineer
NVIDIA
Santa Clara, CA 3 weeks ago
Post-Training Platform Infrastructure Engineer
Post-Training Platform Infrastructure Engineer
AMD
San Jose, CA $204,000 - $306,000 1 week ago
Member of Technical Staff - Model Serving / API Backend Engineer
Member of Technical Staff - Model Serving / API Backend Engineer
Black Forest Labs
San Francisco, CA 4 days ago
LLM Inference Engineer
LLM Inference Engineer
NEAR AI
San Francisco Bay Area 2 weeks ago
Member of Technical Staff, Inference
Member of Technical Staff, Inference
Inferact
San Francisco, CA $200,000 - $400,000 1 month ago
Senior Software Engineer, AI Inference Systems
Senior Software Engineer, AI Inference Systems
NVIDIA
Santa Clara, CA 12 hours ago
Systems Generalist, GPT Infrastructure
Systems Generalist, GPT Infrastructure
OpenAI
San Francisco, CA
$293,000.00
$445,000.00
6 days ago
Software Engineer, Model Inference
Software Engineer, Model Inference
OpenAI
San Francisco, CA
$295,000.00
$555,000.00
2 weeks ago
Platform / Inference Optimization Engineer
Platform / Inference Optimization Engineer
Kaon (prev. FlowGPT)
San Francisco Bay Area 5 months ago
Member of Technical Staff — Inference Palo Alto, CA
Member of Technical Staff — Inference Palo Alto, CA
RadixArk
Palo Alto, CA 2 weeks ago
LLM Inference Engineer
LLM Inference Engineer
Hippocratic AI
Menlo Park, CA 2 weeks ago
Member of Technical Staff - Inference
Member of Technical Staff - Inference
Prime Intellect
San Francisco, CA 2 months ago
Runtime Engineer
Runtime Engineer
MatX
Mountain View, CA 2 weeks ago
Applied AI Inference Engineer
Applied AI Inference Engineer
Crusoe
San Francisco, CA 2 days ago
LLM Inference Engineer
LLM Inference Engineer
Majestic Labs ai
Los Altos, CA 1 month ago
AI/ML Platform Engineer
AI/ML Platform Engineer
AMD
Santa Clara, CA
$204,000.00
$306,000.00
1 day ago
Software Engineer, Inference
Software Engineer, Inference
Luma
San Francisco Bay Area 14 hours ago
Member of Technical Staff, Model Efficiency
Member of Technical Staff, Model Efficiency
Cohere
San Francisco, CA 1 day ago
Senior Software Engineer - Performance
Senior Software Engineer - Performance
Microsoft
Mountain View, CA 2 weeks ago
Member of Technical Staff, Performance and Scale
Member of Technical Staff, Performance and Scale
Inferact
San Francisco, CA
$200,000.00
$400,000.00
6 months ago
Inference Engineer
Inference Engineer
Cartesia
San Francisco, CA 2 weeks ago
Sr. Software Engineer, AI Infrastructure
Sr. Software Engineer, AI Infrastructure
LinkedIn
Sunnyvale, CA 3 days ago
Lead Member of Technical Staff, Inference Infrastructure
Lead Member of Technical Staff, Inference Infrastructure
Cohere
San Francisco, CA 3 weeks ago
Distributed LLM Inference Engineer
Distributed LLM Inference Engineer
Anyscale
San Francisco, CA 1 week ago
Principal Software Engineer - Performance
Principal Software Engineer - Performance
Microsoft
Mountain View, CA 2 weeks ago
Staff Engineer, Inference Optimizations
Staff Engineer, Inference Optimizations
DigitalOcean
San Francisco, CA 7 hours ago
Senior ML Infrastructure Engineer (Compute)
Senior ML Infrastructure Engineer (Compute)
General Motors
Sunnyvale, CA 1 week ago
People also viewed
Cloud Inference Engineer
Cloud Inference Engineer
Luminal
San Francisco, CA $150,000 - $250,000 9 months ago
AI/Backend Engineer | San Francisco | $250k + equity
AI/Backend Engineer | San Francisco | $250k + equity
Harrison Clarke
San Francisco Bay Area 3 weeks ago
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference
Gimlet Labs
San Francisco, CA $150,000 - $350,000 2 months ago
Staff+ Software Engineer, Inference Runtime
Staff+ Software Engineer, Inference Runtime
Anthropic
San Francisco Bay Area 2 weeks ago
Member of Technical Staff - ML Infrastructure & Performance
Member of Technical Staff - ML Infrastructure & Performance
Embedding VC
San Mateo, CA 7 months ago
Performance Engineer
Performance Engineer
Anthropic
San Francisco, CA 2 weeks ago
Senior ML Infrastructure Engineer (Compute)
Senior ML Infrastructure Engineer (Compute)
General Motors
Mountain View, CA 1 week ago
Principal LLM Inference Engineer
Principal LLM Inference Engineer
d-Matrix
Santa Clara, CA 1 day ago
LLM Inference Frameworks and Optimization Engineer
LLM Inference Frameworks and Optimization Engineer
Together AI
San Francisco, CA 1 week ago
Software Engineer, Inference - Multi Modal
Software Engineer, Inference - Multi Modal
OpenAI
San Francisco, CA $295,000 - $555,000 2 weeks ago
Similar Searches
Site Civil Engineer jobs
33,448 open jobs
Data Engineer jobs
192,126 open jobs
Software Engineer jobs
300,699 open jobs
Network Security Officer jobs
5,280 open jobs
Java Software Engineer jobs
36,180 open jobs
Senior Reservoir Engineer jobs
246 open jobs
Civil Engineer jobs
43,143 open jobs
Electrical Software Engineer jobs
223,433 open jobs
Developer jobs
258,935 open jobs
Senior Data Engineer jobs
78,076 open jobs
Senior Software Engineer jobs
78,145 open jobs
Full Stack Engineer jobs
38,546 open jobs
Graduate Civil Engineer jobs
13,551 open jobs
Site Reliability Engineer jobs
169,128 open jobs
Associate Software Engineer jobs
223,979 open jobs
Machine Learning Engineer jobs
148,937 open jobs
Mechanical Engineer jobs
46,392 open jobs
Project Test Engineer jobs
8,924 open jobs
Python Developer jobs
46,642 open jobs
Roads Engineer jobs
17,373 open jobs
Big Data Developer jobs
26,392 open jobs
Network Engineer jobs
173,828 open jobs
Senior Network Engineer jobs
26,545 open jobs
Engineer in Training jobs
143,522 open jobs
Design Verification Engineer jobs
5,299 open jobs
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content
Show more Show less
Global Inference Library Engineer
Infinity Artificial Intelligence Institute San Francisco Bay Area
1 week ago 59 applicants
See who Infinity Artificial Intelligence Institute has hired for this role
Save
Report this job
Global Inference Library - Member of Technical Staff
Company: Infinity
Team: Systems / AI Infrastructure
Location: San Francisco (on-site)
Type: Full-time The Mission Every accelerator that comes online needs an inference stack, and today that stack is hand-built per chip, per model, per optimization - a permanent, growing backlog of engineering work that scales linearly with the number of chips and models in the world.
We're building Infy: a global inference library, generated and continuously maintained by AI, that targets every major chip - NVIDIA, AMD, Trainium, TPU, Maia, Cerebras, Groq, Tenstorrent, and dozens more - across every model category, from LLMs to vision, audio, and robotics. Think vLLM, but for every accelerator on the market, kept automatically up to date as chips, models, and optimization techniques change.
What You'll Work On
Depending on your strengths, you'll own one or more layers of the system:
Optimization agents - the strategy → analyze → generate → curate loop that turns a chip + model pair into an optimized kernel or serving path, and the orchestrator that runs this across many chips and models in parallel.
Kernel optimization loop - the hierarchical agent system that iterates on individual kernels (matmul, attention variants, normalization, RoPE, MoE, collectives) against reference implementations and measured-peak performance gates, drawing on and growing a registry of thousands of known kernels.
Library generator - the system that takes a validated set of optimized components and assembles them into a real, installable inference library per chip, with a standards-compatible serving interface.
Hardware probe - agents that build a structured representation of a chip's memory hierarchy, compute regions, and inter-component communication characteristics, so optimization strategy can branch correctly on the chip's actual architecture.
Coverage tracking - the enablement and optimization tables that track, per chip and per model category, what's implemented and how fast it runs, and that drive prioritization of what to build next.
SOTA paper replication - agents that read inference-optimization papers and automatically implement and validate the techniques they describe, so the library's optimizations keep pace with published research rather than lagging behind it.
Continuous tracking - the pipeline that watches for new papers, new hardware, and new models, and feeds that into what Infy builds and rebuilds next.
Benchmark harness - the metrics layer (TTFT, TPOT, ITL, E2EL, goodput, and model-category-specific equivalents like image-generation latency) that every optimization is measured against.
Inference service - the gateway, router, and billing layers that turn the library into a running, AI-optimized inference service people can actually call.
What we're looking for
We care more about depth and range than a specific checklist, but strong candidates will have most of:
Real experience with ML inference internals - vLLM or similar serving stacks, attention kernel variants, KV cache management, quantization, continuous batching.
Comfort working across heterogeneous hardware - GPU kernels (CUDA, ROCm/HIP, Triton), and ideally exposure to non-GPU execution models (systolic arrays, dataflow, wafer-scale, in-memory/analog compute).
Hands-on experience building with coding agents / LLMs - prompting, tool-use loops, evaluating and constraining model output, and designing systems where the model writes and optimizes code while tests and benchmarks catch regressions.
Fluency in Python and at least one systems language (Rust, C, or C++).
Comfort reading and implementing ideas directly from research papers, not just from existing open-source code.
Nice to have
Contributed to vLLM, TVM, LLVM, MLIR, or a hardware vendor's compiler/runtime stack.
Experience with distributed inference or training internals (NCCL, Megatron-LM, DeepSpeed) and collective-communication algorithms.
Familiarity with model categories beyond LLMs - vision, audio, multimodal, or robotics inference.
Experience building or maintaining a benchmark suite or performance regression system at scale.
Background reproducing results from ML systems papers (KernelBench-style benchmarks, evolutionary/search-based kernel optimization).
Who you are
You want to build the thing that makes every chip's inference performance a solved problem instead of a standing engineering project. You think in terms of systems that scale across hardware and models rather than one-off implementations, you're energized by AI agents doing the generation and optimization work under a tight benchmark harness, and you want to help set the global standard for how inference libraries get built and maintained.
Infinity is an early-stage AI infrastructure research company building the software layer that makes non-NVIDIA chips competitive for AI inference. Rather than relying on scarce human kernel engineers, we use AI to automatically generate, test, and optimize the low-level code that determines how efficiently a chip runs AI models. We've signed or are negotiating design partnerships with d-Matrix, AMD, AWS Trainium, Microsoft (Maia and Nexus), Qualcomm, and others. Founded by Jeremy Nixon (former Google Brain; co-founder of AGI House with Andrej Karpathy), Infinity has raised $15M from investors including the founder of Intercom, the VP of AI at AMD, and the founder of MLCommons. We're headquartered in San Francisco.
Seniority level Entry level
Employment type Full-time
Job function Engineering and Information Technology
Industries Software Development
Referrals increase your chances of interviewing at Infinity Artificial Intelligence Institute by 2x
See who you know
Get notified about new Engineer jobs in San Francisco Bay Area.
Sign in to create job alert
Similar jobs
AI Inference Performance Engineer
AI Inference Performance Engineer
NVIDIA
Santa Clara, CA 3 weeks ago
Post-Training Platform Infrastructure Engineer
Post-Training Platform Infrastructure Engineer
AMD
San Jose, CA $204,000 - $306,000 1 week ago
Member of Technical Staff - Model Serving / API Backend Engineer
Member of Technical Staff - Model Serving / API Backend Engineer
Black Forest Labs
San Francisco, CA 4 days ago
LLM Inference Engineer
LLM Inference Engineer
NEAR AI
San Francisco Bay Area 2 weeks ago
Member of Technical Staff, Inference
Member of Technical Staff, Inference
Inferact
San Francisco, CA $200,000 - $400,000 1 month ago
Senior Software Engineer, AI Inference Systems
Senior Software Engineer, AI Inference Systems
NVIDIA
Santa Clara, CA 12 hours ago
Systems Generalist, GPT Infrastructure
Systems Generalist, GPT Infrastructure
OpenAI
San Francisco, CA
$293,000.00
$445,000.00
6 days ago
Software Engineer, Model Inference
Software Engineer, Model Inference
OpenAI
San Francisco, CA
$295,000.00
$555,000.00
2 weeks ago
Platform / Inference Optimization Engineer
Platform / Inference Optimization Engineer
Kaon (prev. FlowGPT)
San Francisco Bay Area 5 months ago
Member of Technical Staff — Inference Palo Alto, CA
Member of Technical Staff — Inference Palo Alto, CA
RadixArk
Palo Alto, CA 2 weeks ago
LLM Inference Engineer
LLM Inference Engineer
Hippocratic AI
Menlo Park, CA 2 weeks ago
Member of Technical Staff - Inference
Member of Technical Staff - Inference
Prime Intellect
San Francisco, CA 2 months ago
Runtime Engineer
Runtime Engineer
MatX
Mountain View, CA 2 weeks ago
Applied AI Inference Engineer
Applied AI Inference Engineer
Crusoe
San Francisco, CA 2 days ago
LLM Inference Engineer
LLM Inference Engineer
Majestic Labs ai
Los Altos, CA 1 month ago
AI/ML Platform Engineer
AI/ML Platform Engineer
AMD
Santa Clara, CA
$204,000.00
$306,000.00
1 day ago
Software Engineer, Inference
Software Engineer, Inference
Luma
San Francisco Bay Area 14 hours ago
Member of Technical Staff, Model Efficiency
Member of Technical Staff, Model Efficiency
Cohere
San Francisco, CA 1 day ago
Senior Software Engineer - Performance
Senior Software Engineer - Performance
Microsoft
Mountain View, CA 2 weeks ago
Member of Technical Staff, Performance and Scale
Member of Technical Staff, Performance and Scale
Inferact
San Francisco, CA
$200,000.00
$400,000.00
6 months ago
Inference Engineer
Inference Engineer
Cartesia
San Francisco, CA 2 weeks ago
Sr. Software Engineer, AI Infrastructure
Sr. Software Engineer, AI Infrastructure
Sunnyvale, CA 3 days ago
Lead Member of Technical Staff, Inference Infrastructure
Lead Member of Technical Staff, Inference Infrastructure
Cohere
San Francisco, CA 3 weeks ago
Distributed LLM Inference Engineer
Distributed LLM Inference Engineer
Anyscale
San Francisco, CA 1 week ago
Principal Software Engineer - Performance
Principal Software Engineer - Performance
Microsoft
Mountain View, CA 2 weeks ago
Staff Engineer, Inference Optimizations
Staff Engineer, Inference Optimizations
DigitalOcean
San Francisco, CA 7 hours ago
Senior ML Infrastructure Engineer (Compute)
Senior ML Infrastructure Engineer (Compute)
General Motors
Sunnyvale, CA 1 week ago
People also viewed
Cloud Inference Engineer
Cloud Inference Engineer
Luminal
San Francisco, CA $150,000 - $250,000 9 months ago
AI/Backend Engineer | San Francisco | $250k + equity
AI/Backend Engineer | San Francisco | $250k + equity
Harrison Clarke
San Francisco Bay Area 3 weeks ago
Member of Technical Staff - ML Systems & Inference
Member of Technical Staff - ML Systems & Inference
Gimlet Labs
San Francisco, CA $150,000 - $350,000 2 months ago
Staff+ Software Engineer, Inference Runtime
Staff+ Software Engineer, Inference Runtime
Anthropic
San Francisco Bay Area 2 weeks ago
Member of Technical Staff - ML Infrastructure & Performance
Member of Technical Staff - ML Infrastructure & Performance
Embedding VC
San Mateo, CA 7 months ago
Performance Engineer
Performance Engineer
Anthropic
San Francisco, CA 2 weeks ago
Senior ML Infrastructure Engineer (Compute)
Senior ML Infrastructure Engineer (Compute)
General Motors
Mountain View, CA 1 week ago
Principal LLM Inference Engineer
Principal LLM Inference Engineer
d-Matrix
Santa Clara, CA 1 day ago
LLM Inference Frameworks and Optimization Engineer
LLM Inference Frameworks and Optimization Engineer
Together AI
San Francisco, CA 1 week ago
Software Engineer, Inference - Multi Modal
Software Engineer, Inference - Multi Modal
OpenAI
San Francisco, CA $295,000 - $555,000 2 weeks ago
Similar Searches
Site Civil Engineer jobs
33,448 open jobs
Data Engineer jobs
192,126 open jobs
Software Engineer jobs
300,699 open jobs
Network Security Officer jobs
5,280 open jobs
Java Software Engineer jobs
36,180 open jobs
Senior Reservoir Engineer jobs
246 open jobs
Civil Engineer jobs
43,143 open jobs
Electrical Software Engineer jobs
223,433 open jobs
Developer jobs
258,935 open jobs
Senior Data Engineer jobs
78,076 open jobs
Senior Software Engineer jobs
78,145 open jobs
Full Stack Engineer jobs
38,546 open jobs
Graduate Civil Engineer jobs
13,551 open jobs
Site Reliability Engineer jobs
169,128 open jobs
Associate Software Engineer jobs
223,979 open jobs
Machine Learning Engineer jobs
148,937 open jobs
Mechanical Engineer jobs
46,392 open jobs
Project Test Engineer jobs
8,924 open jobs
Python Developer jobs
46,642 open jobs
Roads Engineer jobs
17,373 open jobs
Big Data Developer jobs
26,392 open jobs
Network Engineer jobs
173,828 open jobs
Senior Network Engineer jobs
26,545 open jobs
Engineer in Training jobs
143,522 open jobs
Design Verification Engineer jobs
5,299 open jobs
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content
Show more Show less