Job TitleLLM Serving Engineer (Cloud AI Engineering)CompanyQualcomm Technologies, Inc.Job AreaEngineering Group, Engineering Group > Machine Learning EngineeringGeneral SummaryQualcomm is utilizing its traditional strengths in digital wireless technologies to play a central role in the evolution of Cloud AI. We are investing in several supporting technologies including Deep Learning. The Qualcomm Cloud AI team is developing hardware and software solutions for Inference Acceleration.Role DescriptionWe are hiring LLM Serving Engineers at multiple levels to join our dynamic, collaborative team.
This role spans the full product lifecycle—from cutting-edge research and development to commercial deployment—and demands strategic thinking, strong execution, and excellent communication skills.Key ActivitiesBuilding a scalable LLM inference platform using inference techniques.Contribute to the development of LLM Serving packages.Work closely with customers to drive solutions by collaborating with internal compiler, firmware and platform teams.Work at the forefront of Gen AI by understanding advanced algorithms and numerics to identify new optimization opportunities.Drive efficient serving through smart autoscaling, load balancing and routing.Engage with open-source serving communities to evolve the framework.Candidates for this position will demonstrate the following:Hands-on experience in one or more of the following LLM serving/Orchestration packages.Deep understanding of foundational LLMs, VLMs, SLMs, transformer-based architectures.Strong experience in developing language models using PyTorch.Strong computer science fundamentals—algorithms, data structures, parallel and distributed programming.Understanding of computer architecture, ML accelerators, in-memory processing and distributed systems.Strong Python development skills for large-scale projects with passion for software engineering.Experience in analyzing, profiling, and optimizing deep learning workloads.Proactive learning about the latest inference optimization techniques.Excellent communication and problem-solving skills, with the ability to thrive in a fast-paced and collaborative environment.MS in Computer Science, Machine Learning, Computer Engineering or Electrical Engineering.Bonus SkillsOpen-source contribution to any GenAI package.Experience architecting and developing large-scale distributed systems.High-level kernel design experience (PyTorch, CUDA, Triton).Knowledge of torch.compile or torchDynamo.PhD in Computer Science, Computer Engineering or Machine Learning.Minimum Qualifications• Bachelor's degree in Computer Science, Engineering, Information Systems, or related field and 4+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience. OR Master's degree in Computer Science, Engineering, Information Systems, or related field and 3+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience. OR PhD in Computer Science, Engineering, Information Systems, or related field and 2+ years of Hardware Engineering, Software Engineering, Systems Engineering, or related work experience.Pay Range and Other Compensation & Benefits$158,400.00 - $237,600.00The above pay scale reflects the broad, minimum to maximum, pay scale for this job code for the location for which it has been posted.
Even more importantly, please note that salary is only one component of total compensation at Qualcomm. We also offer a competitive annual discretionary bonus program and opportunity for annual RSU grants (employees on sales-incentive plans are not eligible for our annual bonus). In addition, our highly competitive benefits package is designed to support your success at work, at home, and at play.
Your recruiter will be happy to discuss all that Qualcomm has to offer – and you can review more details about our US benefits at this link.