Crusoe is seeking an engineer to accelerate large language models in production. You will own the end-to-end inference stack—profiling, optimizing, and deploying fast, cost-effective models for customer workloads.
Expect hands-on work across Python and C++, CUDA, and frameworks like vLLM/SGLang. You’ll collaborate with customer teams, push performance improvements into production, and shape product requirements with engineering and solutions teams.
#J-18808-LjbffrSenior AI Inference Architect for Production LLMs in san francisco at Unknown Company
This position is listed as full time and onsite.