What You Do At Amd Changes EverythingAt Amd, our mission is to build great products that accelerate next-generation computing experiences—from Ai and data centers, to Pcs, gaming and embedded systems. Grounded in a culture of innovation and collaboration, we believe real progress comes from bold ideas, human ingenuity and a shared passion to create something extraordinary. When you join Amd, you'll discover the real differentiator is our culture.
We push the limits of innovation to solve the world's most important challenges—striving for execution excellence, while being direct, humble, collaborative, and inclusive of diverse perspectives. Join us as we shape the future of Ai and beyond.The RoleWe are looking for a systems-minded engineer who lives at the intersection of large-scale model inference, distributed systems, and performance optimization. This role focuses on post-training and inference infrastructure, with particular emphasis on P/D disaggregation, KV cache lifecycle management, and efficient offloading mechanisms across both inference and reinforcement learning (RL) systems.The PersonYou enjoy reverse-engineering modern ML infrastructure, reasoning about memory and compute tradeoffs, and turning research insights into production-grade features.
You are comfortable diving into unfamiliar frameworks, understanding their architectural choices, identifying bottlenecks, and improving them through principled engineering.Key ResponsibilitiesResearch and deeply understand modern LLM inference frameworks, including:Architecture and design tradeoffs of P/D (prefill / decode) disaggregationKV cache lifecycle, memory layout, eviction strategies, and reuseKV cache offloading mechanisms across GPU, CPU, and storage backendsAnalyze and compare inference execution paths to identify:Performance bottlenecks (latency, throughput, memory pressure)Inefficiencies in scheduling, cache management, and resource utilizationDevelop and implement infrastructure-level features to:Improve inference latency, throughput, and memory efficiencyOptimize KV cache management and offloading strategiesEnhance scalability across multi-GPU and multi-node deploymentsApply the same research-driven approach to RL frameworks:Study post-training and RL systems (e.g., policy rollout, inference-heavy loops)Debug performance and correctness issues in distributed RL pipelinesOptimize inference, rollout efficiency, and memory usage during trainingCollaborate with research and applied ML teams to:Translate model-level requirements into infrastructure capabilitiesValidate performance gains with benchmarks and real workloadsDocument findings, architectural insights, and best practices to guide future system designPreferred ExperienceStrong background in systems engineering, distributed systems, or ML infrastructureHands-on experience with GPU-accelerated workloads and memory-constrained systemsSolid understanding of:LLM inference workflows (prefill vs decode)Attention mechanisms and KV cache behaviorMulti-process / multi-GPU execution modelsProficiency in Python and C++ (or similar systems languages)Experience debugging performance issues using profiling tools (GPU, CPU, memory)Ability to read, understand, and modify complex open-source codebasesStrong analytical skills and comfort working in research-heavy, ambiguous problem spacesDirect experience with LLM inference frameworks or serving stacksFamiliarity with:GPU memory hierarchies (HBM, pinned memory, NUMA considerations)KV cache compression, paging, or eviction strategiesStorage-backed offloading (NVMe, object stores, distributed file system)Experience with distributed RL or post-training pipelinesKnowledge of scheduling systems, async execution, or actor-based runtimesContributions to open-source ML or systems projectsExperience designing benchmarking suites or performance evaluation frameworksAcademic Credentials:Bachelor's or master's degree in computer science, computer engineering, electrical engineering, or equivalentLocationSan Jose, CA (Hybrid). May consider other US locations.#Li-Cj3#Hybrid