ByteDance's Inference Infrastructure team is seeking engineers to design, build, and operate cloud-native GPU-accelerated ML infrastructure at scale, including vLLM, SGLang, and TensorRT-LLM work.
You will join a world-class team within Core Compute Infrastructure, contributing to open-source ecosystems, scheduling, and orchestration across multi-cloud environments. A PhD and deep expertise in distributed systems help you drive high-performance inference at scale.
#J-18808-LjbffrGraduate Software Engineer - Cloud-Native GPU Inference in san jose at Unknown Company
This position is listed as full time and onsite.