Oho Group in the San Francisco area is seeking an experienced inference specialist to optimize the full stack—from transformer models through execution engines, compilers, runtimes and kernels—to multi-accelerator systems for AI workloads.
You will architect high-performance transformer inference on a new compute platform, improve latency and throughput, and contribute hands-on low-level work in C++, CUDA or Triton to push the frontier of AI inference.
#J-18808-LjbffrPrincipal Transformer Inference Engineer - High Performance in san francisco at Unknown Company
This position is listed as full time and onsite.