AT

Principal Microarchitect

Accepting applications

Acceler8 Talent · Santa Clara, CA

Full-Time Mid_senior AIRTLSoCSystemVerilogVerilog
Posted
31 Jul
Category
Design
Experience
Mid_senior
Country
United States
Senior/Principal Microarchitect — Compute Engine

Santa Clara, CA


About the Role

An early-stage semiconductor company is seeking a Senior or Principal Microarchitect to help design the compute engines for a next-generation high-performance accelerator.

This role spans microarchitecture and RTL development, with ownership from early architectural definition through implementation, verification, first silicon, and post-silicon optimization.

You will work closely with architecture, compiler, kernel, verification, and performance teams to define execution pipelines, data movement, memory structures, and programming-model tradeoffs for demanding compute workloads.


What You’ll Do

Define and implement compute-engine microarchitecture.
Design execution pipelines, scheduling mechanisms, control logic, and issue structures.
Architect datapaths, register files, scratchpads, and local memory hierarchies.
Support modern floating-point and reduced-precision numerical formats.
Optimize designs across performance, utilization, power, area, and implementation complexity.
Balance compute throughput against memory bandwidth, latency, and broader SoC constraints.
Analyze representative workloads to guide architectural decisions.
Evaluate instruction-set and programming-model tradeoffs.
Use performance modeling, benchmarking, and profiling to validate design choices.
Lead the initial RTL implementation of major compute-engine blocks.
Support ongoing RTL development, design verification, synthesis, and timing closure.
Participate in emulation, debug, silicon bring-up, and post-silicon performance tuning.
Write detailed microarchitecture specifications and implementation plans.
Provide technical leadership across architecture and design teams.


What We’re Looking For

Strong experience designing AI compute engines, GPUs, vector processors, matrix engines, DSPs, or other specialized accelerator hardware.
Proven experience designing complex hardware units composed of multiple interacting blocks.
Deep understanding of:
Matrix and vector execution pipelines
Floating-point arithmetic
Reduced-precision computation
Quantization
Scheduling and control
Datapath design
Register files and local memories
Performance-per-watt optimization
Strong RTL development experience using Verilog or SystemVerilog.
Experience taking complex designs from architecture through implementation and verification.
Experience delivering production silicon.
Ability to make tradeoffs across performance, power, area, bandwidth, schedule, and verification risk.
Strong collaboration skills across architecture, compiler, kernel, RTL, verification, and physical-design teams.


Preferred Experience

AI accelerators, GPUs, NPUs, or custom compute silicon.
FP16, BF16, FP8, FP4, integer, or mixed-precision datapaths.
GEMM, tensor, vector, or matrix-processing hardware.
ISA design or programming-model development.
Performance modeling and workload analysis.
Compiler or kernel interaction with accelerator hardware.
Emulation and pre-silicon validation.
Post-silicon bring-up and performance tuning.
Early-stage semiconductor development.
Advanced process-node experience.
Master’s degree, PhD, or equivalent practical experience.


Relevant Keywords
Microarchitecture, Microarchitect, Compute Engine, AI Accelerator, NPU, GPU, Tensor Processor, Matrix Engine, Vector Processor, RTL Design, SystemVerilog, Verilog, Execution Pipeline, Instruction Scheduling, Control Logic, Datapath, Register File, Scratchpad Memory, Local Memory, Memory Hierarchy, Matrix Pipeline, Vector Pipeline, Floating Point, FP16, BF16, FP8, FP4, Mixed Precision, Quantization, GEMM, Tensor Operations, Throughput Optimization, Performance per Watt, PPA, Power Optimization, Area Optimization, Memory Bandwidth, ISA Design, Programming Model, Performance Modeling, Workload Analysis, Hardware-Software Co-Design, Compiler Co-Design, Kernel Optimization, Synthesis, Timing Closure, Emulation, Silicon Bring-Up, Post-Silicon Validation, First Silicon
Show more Show less