Leading a high-performing AI Inference team, the full-time Tech Lead - AI Inference will bridge complex research and production engineering, mentoring developers while managing the core inference infrastructure and driving execution on high-performance LLM serving systems in a remote environment. Key responsibilities: Take end-to-end ownership of core inference infrastructure and drive technical decisions Guide engineers through design, implementation, and delivery of high-throughput LLM inference systems Mentor and coach engineers, fostering a culture of collaboration and technical excellence Required qualifications: 5+ years of experience in software engineering with a focus on AI/ML infrastructure or high-performance computing Hands-on expertise with LLM serving systems and familiarity with frameworks like vLLM and LMCache Strong programming skills in Python and C++, with knowledge of CUDA and GPU memory management Experience deploying and scaling GPU workloads on Kubernetes Proven ability to mentor engineers and support their career growth
Tech Lead - AI Inference in workfromhome at Unknown Company
This position is listed as full time and able to be worked remotely.