Crusoe is seeking an engineer to make large language models run faster and cheaper in production. You will own the inference stack end to end, profiling time and cost, and bringing optimization techniques into real deployments.
This hands-on role blends systems work with customer collaboration, requiring coding, profiling, and low-level optimization across frameworks and kernels. You will shape deployments for diverse models and traffic while shipping reliable results.
#J-18808-Ljbffr