Unknown Company

GenAI Inference Engineer — High-Performance ML Systems

san francisco, ca • Posted 6 days ago
Onsite Full Time IT & Technology

Databricks is seeking a software engineer for GenAI inference to design, develop, and optimize the inference engine powering our Foundation Model API. You will work at the intersection of research and production, ensuring fast, scalable LLM serving across GPUs and accelerators.

You’ll collaborate with researchers to bring new model architectures into the engine, optimize latency and memory usage, and build tooling for profiling, routing, and fault tolerance.

#J-18808-Ljbffr
Back to Job Search