TC
Chip Performance Profiling Engineer - Member of Technical Staff
Accepting applicationsTouring Capital · San Francisco Bay Area
Full-Time Senior AIASICArmC++Deep Learning
Posted
2d ago
Category
Design
Experience
Senior
Country
United States
Infinity Artificial Intelligence Institute San Francisco Bay Area
Chip Performance Profiling Engineer - Member of Technical Staff
Infinity Artificial Intelligence Institute San Francisco Bay Area
1 week ago 60 applicants
See who Infinity Artificial Intelligence Institute has hired for this role
Save
Report this job
Company: Infinity
Team: Systems / AI Infrastructure Location: San Francisco (on-site)
Type: Full-time
The Mission
You can't optimize what you can't measure, and on a fresh accelerator there is usually nothing to measure with - no Nsight, no rocprof, no performance counters anyone has documented how to read. Visibility today is a per-vendor artifact, hand-built by the people who shipped the silicon, so every chip without a mature profiler leaves engineers optimizing in the dark until someone ports one over by hand. And even where a profiler exists, most of them hand you data instead of an answer: a thousand numbers that never say which one is the bottleneck.
We're building the agent that generates that visibility. Give it a supported chip and it produces a profiler for that chip - one that attributes runtime to every operation and sub-operation on an inference pass, fine-grained enough to show which step inside an attention kernel is the problem rather than just that attention is slow. It reimplements the surface area engineers already expect from nvprof and Nsight, so moving to a new accelerator doesn't cost you the tools you profile with, and it runs on the chip itself with no simulator in the loop. The same machinery profiles anything the chip does, not only inference.
Measurement sits upstream of everything else this stack does. The optimizer can't move a number it can't see, and an engineer can't fix a bottleneck no tool will name. Making that visibility something we generate rather than something each vendor hand-builds means every accelerator we support arrives already observable - and the profiler points at the specific thing standing between the current code and more performance instead of leaving you to find it in the noise.
What You'll Work On
You'll build the instrument the rest of the stack reads from. Depending on your strengths, you'll own one or more parts of the system:
Profiler-generation agent - the system that takes a supported chip and builds a working profiler for it, so visibility on a new accelerator is something you generate rather than something someone ports by hand every time.
Per-operation attribution - breaking an inference pass down until every operation and sub-operation carries its own runtime, fine enough to see which step inside a kernel is the one costing you rather than just which kernel is slow.
The nvprof and Nsight surface area, per chip - reimplementing the features engineers already lean on to profile, on accelerators that never shipped a profiler of their own, so the tooling doesn't reset every time the hardware does.
Counter and telemetry discovery - finding and validating the signals on parts where the performance-monitoring unit is undocumented or has to be inferred from behavior, since everything downstream depends on trusting what those counters report.
Faithful instrumentation - keeping the act of measuring from changing the timing it measures, because a profiler that perturbs the numbers it reports is worse than no profiler at all.
Bottleneck surfacing - turning a wall of measurements into the specific thing standing between the current code and more performance, so both the optimizer and the engineer know where to push.
What We're Looking For
We care more about depth and range than a specific checklist, but strong candidates will have most of:
Real performance-analysis experience - you've profiled hard problems and know the difference between a number and a number you can trust.
Low-level systems background - hardware counters, tracing, sampling, and instrumentation.
Comfort building measurement tools when the documentation for what you're measuring is thin or absent.
Statistical care - you worry about the observer effect, noise, and sample size before you report a result.
Fluency in Python and at least one systems language (Rust, C, or C++).
Nice to have
Built a profiler, tracer, or telemetry pipeline that other people relied on.
Know the internals of perf, VTune, or Nsight rather than just their front ends.
Worked directly with hardware performance counters and their sharp edges.
Hands-on experience building with coding agents.
Who We Are
Infinity is an early-stage AI infrastructure research company building the software layer that makes non-NVIDIA chips competitive for AI inference. Rather than relying on scarce human kernel engineers, we use AI to automatically generate, test, and optimize the low-level code that determines how efficiently a chip runs AI models. We've signed or are negotiating design partnerships with d-Matrix, AMD, AWS Trainium, Microsoft (Maia and Nexus), Qualcomm, and others. Founded by Jeremy Nixon (former Google Brain; co-founder of AGI House with Andrej Karpathy), Infinity has raised $15M from investors including the founder of Intercom, the VP of AI at AMD, and the founder of MLCommons. We're headquartered in San Francisco.
Seniority level Mid-Senior level
Employment type Full-time
Job function Engineering and Information Technology
Industries Software Development
Referrals increase your chances of interviewing at Infinity Artificial Intelligence Institute by 2x
See who you know
Get notified about new Member of Technical Staff jobs in San Francisco Bay Area.
Sign in to create job alert
Similar jobs
Senior Performance Verification Engineer
Senior Performance Verification Engineer
NVIDIA
Santa Clara, CA 2 weeks ago
CPU Performance Analysis Engineer
CPU Performance Analysis Engineer
Qualcomm
Santa Clara, CA 4 days ago
CPU Performance Analysis Engineer (Multiple Locations- San Diego, Santa Clara, Austin)
CPU Performance Analysis Engineer (Multiple Locations- San Diego, Santa Clara, Austin)
Qualcomm
Santa Clara, CA 4 days ago
Senior Power and Performance Engineer
Senior Power and Performance Engineer
Intel
Santa Clara, CA 1 week ago
CPU Performance Architect
CPU Performance Architect
Google
Mountain View, CA 2 days ago
Performance Modeling Engineer
Performance Modeling Engineer
MediaTek
San Jose, CA 1 day ago
Senior Data Center Performance Engineer - Benchmarking and Optimization
Senior Data Center Performance Engineer - Benchmarking and Optimization
NVIDIA
Santa Clara, CA 3 weeks ago
Principal Performance Architect
Principal Performance Architect
Microsoft
Mountain View, CA 17 hours ago
Workload Porting & Performance Engineer
Workload Porting & Performance Engineer
OpenAI
San Francisco, CA
$293,000.00
$385,000.00
2 weeks ago
Principal Performance Modeling Engineer
Principal Performance Modeling Engineer
AMD
Santa Clara, CA
$188,160.00
$282,240.00
1 day ago
Senior Performance Engineer
Senior Performance Engineer
Samsung Semiconductor
San Jose, CA 1 day ago
System Performance Modeling Engineer
System Performance Modeling Engineer
AMD
Santa Clara, CA
$189,600.00
$284,400.00
4 days ago
System Performance Engineer, Consumer Devices
System Performance Engineer, Consumer Devices
OpenAI
San Francisco, CA
$293,000.00
$325,000.00
2 weeks ago
CPU & Microarchitecture Performance Engineer
CPU & Microarchitecture Performance Engineer
Aramas AI
San Francisco Bay Area 1 month ago
Modeling Engineer
Modeling Engineer
Arm
San Jose, CA 6 hours ago
Performance Engineer, Deep Learning and HPC
Performance Engineer, Deep Learning and HPC
NVIDIA AI
Santa Clara, CA 2 days ago
Performance Modeling Engineer
Performance Modeling Engineer
Etched
San Jose, CA
$175,000.00
$275,000.00
1 week ago
Performance Research Engineer (multiple levels)
Performance Research Engineer (multiple levels)
Efficient Computer
San Jose, CA 2 weeks ago
Power Instrumentation Engineer
Power Instrumentation Engineer
Intel
Santa Clara, CA 1 week ago
System Architect
System Architect
Micron Technology
San Jose, CA 1 week ago
Performance Engineer, Inference Systems
Performance Engineer, Inference Systems
Anthropic
San Francisco, CA 2 weeks ago
Workload / Performance Model Lead
Workload / Performance Model Lead
SiFive
Santa Clara, CA 3 months ago
Sr Performance Validation Engineer
Sr Performance Validation Engineer
Amazon Lab126
Sunnyvale, CA 3 days ago
Member of Technical Staff — Performance Palo Alto, CA
Member of Technical Staff — Performance Palo Alto, CA
RadixArk
Palo Alto, CA 2 weeks ago
Workload / Performance Model Lead
Workload / Performance Model Lead
SiFive
Berkeley, CA 3 months ago
ASIC Engineer, Architecture
ASIC Engineer, Architecture
Meta
Sunnyvale, CA
$178,000.00
$250,000.00
2 weeks ago
SoC Performance Architect
SoC Performance Architect
Samsung Electronics America
San Jose, CA 3 months ago
People also viewed
Senior System Performance Engineer
Senior System Performance Engineer
General Motors
Sunnyvale, CA 2 weeks ago
Senior System Performance Engineer
Senior System Performance Engineer
General Motors
Mountain View, CA 2 weeks ago
Senior Engineer, Performance Architecture
Senior Engineer, Performance Architecture
Samsung Semiconductor
San Jose, CA 3 days ago
Lead CPU Performance Analysis Engineer
Lead CPU Performance Analysis Engineer
Qualcomm
Santa Clara, CA 5 days ago
CPU Performance Research Engineer
CPU Performance Research Engineer
Qualcomm
Santa Clara, CA 6 days ago
Performance Engineer, Deep Learning and HPC
Performance Engineer, Deep Learning and HPC
NVIDIA
Santa Clara, CA 3 days ago
Senior CPU Performance Architect
Senior CPU Performance Architect
NVIDIA
Santa Clara, CA 2 weeks ago
Senior CPU Performance Architect
Senior CPU Performance Architect
NVIDIA
Santa Clara, CA 2 weeks ago
CPU Performance Modeling Engineer (Multiple Levels)
CPU Performance Modeling Engineer (Multiple Levels)
Qualcomm
Santa Clara, CA 1 week ago
Performance Modeling Engineer
Performance Modeling Engineer
OpenAI
San Francisco, CA $293,000 - $385,000 2 weeks ago
Similar Searches
Senior Member of Technical Staff jobs
8,067 open jobs
Member Technical jobs
62,993 open jobs
Senior Wireless Engineer jobs
35,731 open jobs
Physical Design Engineer jobs
6,753 open jobs
Staff Test Engineer jobs
2,677 open jobs
Principal Firmware Engineer jobs
1,780 open jobs
Line Technician jobs
109,890 open jobs
Lead Infrastructure Engineer jobs
16,227 open jobs
Support Team Manager jobs
83,600 open jobs
Vice President Software jobs
49,146 open jobs
Senior Lead Software Engineer jobs
49,381 open jobs
Principal Researcher jobs
4,530 open jobs
Switch Engineer jobs
9,391 open jobs
Staff Software Engineer jobs
64,945 open jobs
Lead Quality Engineer jobs
10,548 open jobs
Control Coordinator jobs
39,876 open jobs
Market Maker jobs
1,432 open jobs
Yield Engineer jobs
9,445 open jobs
Computer Scientist jobs
49,477 open jobs
Lead Test Engineer jobs
13,921 open jobs
House Supervisor jobs
29,485 open jobs
Core Engineer jobs
33,936 open jobs
Cable Technician jobs
11,124 open jobs
Principal Software Engineer jobs
73,845 open jobs
Logic Design Engineer jobs
1,858 open jobs
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content
Show more Show less
Chip Performance Profiling Engineer - Member of Technical Staff
Infinity Artificial Intelligence Institute San Francisco Bay Area
1 week ago 60 applicants
See who Infinity Artificial Intelligence Institute has hired for this role
Save
Report this job
Company: Infinity
Team: Systems / AI Infrastructure Location: San Francisco (on-site)
Type: Full-time
The Mission
You can't optimize what you can't measure, and on a fresh accelerator there is usually nothing to measure with - no Nsight, no rocprof, no performance counters anyone has documented how to read. Visibility today is a per-vendor artifact, hand-built by the people who shipped the silicon, so every chip without a mature profiler leaves engineers optimizing in the dark until someone ports one over by hand. And even where a profiler exists, most of them hand you data instead of an answer: a thousand numbers that never say which one is the bottleneck.
We're building the agent that generates that visibility. Give it a supported chip and it produces a profiler for that chip - one that attributes runtime to every operation and sub-operation on an inference pass, fine-grained enough to show which step inside an attention kernel is the problem rather than just that attention is slow. It reimplements the surface area engineers already expect from nvprof and Nsight, so moving to a new accelerator doesn't cost you the tools you profile with, and it runs on the chip itself with no simulator in the loop. The same machinery profiles anything the chip does, not only inference.
Measurement sits upstream of everything else this stack does. The optimizer can't move a number it can't see, and an engineer can't fix a bottleneck no tool will name. Making that visibility something we generate rather than something each vendor hand-builds means every accelerator we support arrives already observable - and the profiler points at the specific thing standing between the current code and more performance instead of leaving you to find it in the noise.
What You'll Work On
You'll build the instrument the rest of the stack reads from. Depending on your strengths, you'll own one or more parts of the system:
Profiler-generation agent - the system that takes a supported chip and builds a working profiler for it, so visibility on a new accelerator is something you generate rather than something someone ports by hand every time.
Per-operation attribution - breaking an inference pass down until every operation and sub-operation carries its own runtime, fine enough to see which step inside a kernel is the one costing you rather than just which kernel is slow.
The nvprof and Nsight surface area, per chip - reimplementing the features engineers already lean on to profile, on accelerators that never shipped a profiler of their own, so the tooling doesn't reset every time the hardware does.
Counter and telemetry discovery - finding and validating the signals on parts where the performance-monitoring unit is undocumented or has to be inferred from behavior, since everything downstream depends on trusting what those counters report.
Faithful instrumentation - keeping the act of measuring from changing the timing it measures, because a profiler that perturbs the numbers it reports is worse than no profiler at all.
Bottleneck surfacing - turning a wall of measurements into the specific thing standing between the current code and more performance, so both the optimizer and the engineer know where to push.
What We're Looking For
We care more about depth and range than a specific checklist, but strong candidates will have most of:
Real performance-analysis experience - you've profiled hard problems and know the difference between a number and a number you can trust.
Low-level systems background - hardware counters, tracing, sampling, and instrumentation.
Comfort building measurement tools when the documentation for what you're measuring is thin or absent.
Statistical care - you worry about the observer effect, noise, and sample size before you report a result.
Fluency in Python and at least one systems language (Rust, C, or C++).
Nice to have
Built a profiler, tracer, or telemetry pipeline that other people relied on.
Know the internals of perf, VTune, or Nsight rather than just their front ends.
Worked directly with hardware performance counters and their sharp edges.
Hands-on experience building with coding agents.
Who We Are
Infinity is an early-stage AI infrastructure research company building the software layer that makes non-NVIDIA chips competitive for AI inference. Rather than relying on scarce human kernel engineers, we use AI to automatically generate, test, and optimize the low-level code that determines how efficiently a chip runs AI models. We've signed or are negotiating design partnerships with d-Matrix, AMD, AWS Trainium, Microsoft (Maia and Nexus), Qualcomm, and others. Founded by Jeremy Nixon (former Google Brain; co-founder of AGI House with Andrej Karpathy), Infinity has raised $15M from investors including the founder of Intercom, the VP of AI at AMD, and the founder of MLCommons. We're headquartered in San Francisco.
Seniority level Mid-Senior level
Employment type Full-time
Job function Engineering and Information Technology
Industries Software Development
Referrals increase your chances of interviewing at Infinity Artificial Intelligence Institute by 2x
See who you know
Get notified about new Member of Technical Staff jobs in San Francisco Bay Area.
Sign in to create job alert
Similar jobs
Senior Performance Verification Engineer
Senior Performance Verification Engineer
NVIDIA
Santa Clara, CA 2 weeks ago
CPU Performance Analysis Engineer
CPU Performance Analysis Engineer
Qualcomm
Santa Clara, CA 4 days ago
CPU Performance Analysis Engineer (Multiple Locations- San Diego, Santa Clara, Austin)
CPU Performance Analysis Engineer (Multiple Locations- San Diego, Santa Clara, Austin)
Qualcomm
Santa Clara, CA 4 days ago
Senior Power and Performance Engineer
Senior Power and Performance Engineer
Intel
Santa Clara, CA 1 week ago
CPU Performance Architect
CPU Performance Architect
Mountain View, CA 2 days ago
Performance Modeling Engineer
Performance Modeling Engineer
MediaTek
San Jose, CA 1 day ago
Senior Data Center Performance Engineer - Benchmarking and Optimization
Senior Data Center Performance Engineer - Benchmarking and Optimization
NVIDIA
Santa Clara, CA 3 weeks ago
Principal Performance Architect
Principal Performance Architect
Microsoft
Mountain View, CA 17 hours ago
Workload Porting & Performance Engineer
Workload Porting & Performance Engineer
OpenAI
San Francisco, CA
$293,000.00
$385,000.00
2 weeks ago
Principal Performance Modeling Engineer
Principal Performance Modeling Engineer
AMD
Santa Clara, CA
$188,160.00
$282,240.00
1 day ago
Senior Performance Engineer
Senior Performance Engineer
Samsung Semiconductor
San Jose, CA 1 day ago
System Performance Modeling Engineer
System Performance Modeling Engineer
AMD
Santa Clara, CA
$189,600.00
$284,400.00
4 days ago
System Performance Engineer, Consumer Devices
System Performance Engineer, Consumer Devices
OpenAI
San Francisco, CA
$293,000.00
$325,000.00
2 weeks ago
CPU & Microarchitecture Performance Engineer
CPU & Microarchitecture Performance Engineer
Aramas AI
San Francisco Bay Area 1 month ago
Modeling Engineer
Modeling Engineer
Arm
San Jose, CA 6 hours ago
Performance Engineer, Deep Learning and HPC
Performance Engineer, Deep Learning and HPC
NVIDIA AI
Santa Clara, CA 2 days ago
Performance Modeling Engineer
Performance Modeling Engineer
Etched
San Jose, CA
$175,000.00
$275,000.00
1 week ago
Performance Research Engineer (multiple levels)
Performance Research Engineer (multiple levels)
Efficient Computer
San Jose, CA 2 weeks ago
Power Instrumentation Engineer
Power Instrumentation Engineer
Intel
Santa Clara, CA 1 week ago
System Architect
System Architect
Micron Technology
San Jose, CA 1 week ago
Performance Engineer, Inference Systems
Performance Engineer, Inference Systems
Anthropic
San Francisco, CA 2 weeks ago
Workload / Performance Model Lead
Workload / Performance Model Lead
SiFive
Santa Clara, CA 3 months ago
Sr Performance Validation Engineer
Sr Performance Validation Engineer
Amazon Lab126
Sunnyvale, CA 3 days ago
Member of Technical Staff — Performance Palo Alto, CA
Member of Technical Staff — Performance Palo Alto, CA
RadixArk
Palo Alto, CA 2 weeks ago
Workload / Performance Model Lead
Workload / Performance Model Lead
SiFive
Berkeley, CA 3 months ago
ASIC Engineer, Architecture
ASIC Engineer, Architecture
Meta
Sunnyvale, CA
$178,000.00
$250,000.00
2 weeks ago
SoC Performance Architect
SoC Performance Architect
Samsung Electronics America
San Jose, CA 3 months ago
People also viewed
Senior System Performance Engineer
Senior System Performance Engineer
General Motors
Sunnyvale, CA 2 weeks ago
Senior System Performance Engineer
Senior System Performance Engineer
General Motors
Mountain View, CA 2 weeks ago
Senior Engineer, Performance Architecture
Senior Engineer, Performance Architecture
Samsung Semiconductor
San Jose, CA 3 days ago
Lead CPU Performance Analysis Engineer
Lead CPU Performance Analysis Engineer
Qualcomm
Santa Clara, CA 5 days ago
CPU Performance Research Engineer
CPU Performance Research Engineer
Qualcomm
Santa Clara, CA 6 days ago
Performance Engineer, Deep Learning and HPC
Performance Engineer, Deep Learning and HPC
NVIDIA
Santa Clara, CA 3 days ago
Senior CPU Performance Architect
Senior CPU Performance Architect
NVIDIA
Santa Clara, CA 2 weeks ago
Senior CPU Performance Architect
Senior CPU Performance Architect
NVIDIA
Santa Clara, CA 2 weeks ago
CPU Performance Modeling Engineer (Multiple Levels)
CPU Performance Modeling Engineer (Multiple Levels)
Qualcomm
Santa Clara, CA 1 week ago
Performance Modeling Engineer
Performance Modeling Engineer
OpenAI
San Francisco, CA $293,000 - $385,000 2 weeks ago
Similar Searches
Senior Member of Technical Staff jobs
8,067 open jobs
Member Technical jobs
62,993 open jobs
Senior Wireless Engineer jobs
35,731 open jobs
Physical Design Engineer jobs
6,753 open jobs
Staff Test Engineer jobs
2,677 open jobs
Principal Firmware Engineer jobs
1,780 open jobs
Line Technician jobs
109,890 open jobs
Lead Infrastructure Engineer jobs
16,227 open jobs
Support Team Manager jobs
83,600 open jobs
Vice President Software jobs
49,146 open jobs
Senior Lead Software Engineer jobs
49,381 open jobs
Principal Researcher jobs
4,530 open jobs
Switch Engineer jobs
9,391 open jobs
Staff Software Engineer jobs
64,945 open jobs
Lead Quality Engineer jobs
10,548 open jobs
Control Coordinator jobs
39,876 open jobs
Market Maker jobs
1,432 open jobs
Yield Engineer jobs
9,445 open jobs
Computer Scientist jobs
49,477 open jobs
Lead Test Engineer jobs
13,921 open jobs
House Supervisor jobs
29,485 open jobs
Core Engineer jobs
33,936 open jobs
Cable Technician jobs
11,124 open jobs
Principal Software Engineer jobs
73,845 open jobs
Logic Design Engineer jobs
1,858 open jobs
Explore top content on LinkedIn
Find curated posts and insights for relevant topics all in one place.
View top content
Show more Show less