Baseten is seeking a backend-focused engineer for the Model Performance team to own low-latency APIs powering hosted model endpoints. You’ll work across distributed systems, model serving, and developer tooling to optimize throughput, reliability, and cost efficiency.
You’ll design and deploy benchmarking, profiling, and observability, implement platform fundamentals like authentication and quotas, and collaborate with multiple squads to deliver robust model serving experiences.
#J-18808-Ljbffr