Inferact is seeking an inference runtime engineer to push the boundaries of LLM and diffusion model serving. You will optimize how models execute across diverse hardware and architectures, contributing to the core of vLLM and related projects.
This fully remote role offers salary plus equity, visa sponsorship on a case-by-case basis, and comprehensive benefits. Regular overlap with Pacific Time is expected for critical syncs.
#J-18808-LjbffrRemote Inference Engine Engineer - LLMs & Diffusion in Location not specified at Unknown Company
This position is listed as full time and able to be worked remotely.