I’m working with a well-funded AI compute company building a programmable accelerator platform and the complete software stack required to run modern AI workloads efficiently.
They are hiring a senior architect to shape the full path from models and frameworks through graph capture, compiler, runtime and driver layers, down to kernels and hardware execution. This is a rare role for someone who can connect high-level workload behaviour with low-level GPU architecture and code generation.
You will work across:
- PyTorch, JAX and ONNX integration
- Graph capture, partitioning and optimization
- MLIR/LLVM lowering and backend code generation
- Runtime, scheduling and driver architecture
- SIMT execution, occupancy and memory hierarchy
- CUDA and Triton programming models
- Kernel fusion, tiling and custom operator lowering
- Hardware/software co-design and platform validation
- Performance profiling from models down to instructions
They are looking for:
- 12+ years across AI, compiler, runtime or systems software
- Strong C++ with deep LLVM and MLIR experience
- Expertise in GPU/SIMT execution and CUDA programming
- Understanding of modern AI models at graph and operator level
- Experience connecting frameworks, compilers, runtimes and hardware
- A hands-on, measurement-led approach to architecture
- Experience guiding senior engineers through design and code reviews
Experience with custom accelerators, PyTorch compilation, XLA, Triton, quantization, performance modelling or pre-silicon software development would be particularly valuable.
This is an opportunity to define the software architecture for a new compute platform as it moves from a working stack into large-scale production.
#J-18808-LjbffrPrincipal Software Architect in san francisco at Unknown Company
This position is listed as full time and onsite.