Unknown Company

Lead Production-Scale LLM Inference Engineer

san francisco, ca • Posted 1 weeks ago
Onsite Full Time IT & Technology

Crusoe is seeking an engineer to make large language models run faster and cheaper in production. You will own the inference stack end to end, profiling time and cost, and bringing optimization techniques into real deployments.

This hands-on role blends systems work with customer collaboration, requiring coding, profiling, and low-level optimization across frameworks and kernels. You will shape deployments for diverse models and traffic while shipping reliable results.

#J-18808-Ljbffr
Back to Job Search