Unknown Company

Inference Runtime Performance Engineer — GPU Kernels

san francisco, ca • Posted 3 days ago
Onsite Full Time IT & Technology

iFrame Corporation is hiring for a role on the runtime team to own end-to-end performance for multiple model families on our managed inference runtime. You will write fused CUDA/Triton kernels and work across tokenizer through KV cache and decoding, targeting H100/H200 and other accelerators.

With 5+ years in systems-level performance and a track record of speed-ups, you will design long-context primitives and publish external write-ups quarterly while sharing on-call duties with SRE and

#J-18808-Ljbffr
Back to Job Search