Amazon in Seattle is seeking a Senior Inference Engineer to own end-to-end real-time inference for multimodal models, spanning research to production. You will shape architectures for servability, build streaming runtimes, and develop offline systems for RL and post-training rollout.
You will collaborate with scientists and hardware partners to achieve sub-second latency on real-time workloads while controlling cost and ensuring scalability across distributed systems and multiple GPUs.
#J-18808-LjbffrSenior Real-Time Multimodal Inference Engineer in seattle at Unknown Company
This position is listed as full time and onsite.