Crusoe seeks a hands-on engineer to optimize large-model inference, owning the inference stack end-to-end from profiling to production deployment. You will design and optimize serving architectures for demanding models and work with customers to tailor deployments that meet latency and cost targets.
You will use frameworks like vLLM and SGLang, dive into CUDA kernels, and collaborate with AI/ML teams to ship robust, well-monitored production services.
#J-18808-LjbffrSenior AI Inference Engineer - Production-Ready in san francisco at Unknown Company
This position is listed as full time and onsite.