Nebius is building a full-stack AI cloud platform addressing data and model training to production deployment. We seek an engineer to own the reliability, performance, and observability of the entire inference stack in a high-demand environment.
Responsibilities include designing telemetry pipelines, tuning autoscalers, and crafting infrastructure-as-code modules. You will collaborate with software engineers to deliver self-healing, scalable systems that meet aggressive cost and reliability
#J-18808-Ljbffr