Amazon is seeking a Senior Inference Engineer in Sunnyvale, CA to own the end-to-end inference stack for real-time multimodal conversational AI. You will shape model architectures, build low-latency streaming serving paths, and develop offline training and RL/evaluation infrastructure.
Expect cross-team collaboration with scientists and hardware partners to ensure scalable, cost-effective performance. You will drive architecture co-design, optimize large-scale inference, and implement
#J-18808-LjbffrSenior Real-Time Multimodal Inference Architect in sunnyvale at Unknown Company
This position is listed as full time and onsite.