Arm in Seattle is seeking a Software Engineer for the AI Inference Runtime team to set technical direction for distributed inference runtime powering SOTA AI models. You will lead hands-on work across scheduling, batching, KV-cache management, memory allocation, distributed execution, kernel development and optimization, shaping how efficiently models use compute.
You will partner with AI Infrastructure, compute and product teams to raise performance and energy efficiency of Arm’s AI platform,
#J-18808-LjbffrStaff AI Inference Runtime Engineer in seattle at Unknown Company
This position is listed as full time and onsite.