Databricks is seeking a software engineer for GenAI inference to design, develop, and optimize the inference engine powering our Foundation Model API. You will work at the intersection of research and production, ensuring fast, scalable LLM serving across GPUs and accelerators.
You’ll collaborate with researchers to bring new model architectures into the engine, optimize latency and memory usage, and build tooling for profiling, routing, and fault tolerance.
#J-18808-Ljbffr