Baseten is seeking a Software Engineer for the Inference Stack to own distributed infrastructure for large-scale LLM inference. You will navigate the stack from developer-facing features to low-level systems, enabling robust, scalable model deployments.
You’ll work on routing, autoscaling, observability, and runtime management while collaborating with Model Performance engineers to push optimizations broadly to customers. This role emphasizes reliability, performance, and developer experience.
#J-18808-Ljbffr