Seeking a remote Inference Infrastructure Architect, the full-time position will focus on optimizing a global fleet for AI-agent traffic, ensuring efficient throughput and latency while managing serverless inference deployments for both internal and external clients. Key responsibilities Operate and expand the GPU fleet to maximize inference throughput and minimize costs Develop serverless serving pools and dedicated model deployments for enterprise clients Implement observability and capacity planning strategies to enhance performance and reliability Required qualifications Proven experience with production LLM serving under traffic and latency constraints Hands-on expertise with Kubernetes on GPU fleets, including GPU Operator and topology-aware scheduling Deep operational knowledge of vLLM or SGLang, including deployment and tuning Strong background in performance engineering and system-level metrics analysis Proficiency in Python and Go for automation, with a solid understanding of Linux and networking
Inference Infrastructure Architect in workfromhome at Unknown Company
This position is listed as full time and able to be worked remotely.