Sr. Staff Software Engineer, Systems InfrastructureThis role will be based in Sunnyvale or Mountain View, CA.At LinkedIn, our approach to flexible work is centered on trust and optimized for culture, connection, clarity, and the evolving needs of our business. The work location of this role is hybrid, meaning it will be performed both from home and from a LinkedIn office on select days, as determined by the business needs of the team.LinkedIn's AI Infrastructure organization is responsible for building the foundational platforms that power AI across LinkedIn. The LLM Serving team builds the critical infrastructure that enables efficient, reliable, and large-scale deployment of large language models and other advanced AI models in production.This team sits at the center of LinkedIn's AI platform, owning the layer between model training and production serving.
The work focuses on making large-scale models run faster, cheaper, and more efficiently on GPUs at LinkedIn scale. The team builds and extends high-performance serving infrastructure and contributes to leading open-source technologies such as SGLang, vLLM, and related model serving frameworks.We are looking for a Senior Staff Software Engineer with deep expertise at the intersection of systems, machine learning, GPU infrastructure, and large-scale inference. This is a highly technical, high-leverage role for someone who enjoys going deep into how models interact with runtimes, compilers, and hardware, and who wants to drive meaningful improvements in performance, cost, latency, and scalability across LinkedIn's AI systems.ResponsibilitiesLead the design, development, and optimization of LinkedIn's large-scale LLM serving infrastructureDrive performance improvements across AI inference systems, including latency, throughput, GPU utilization, and cost efficiencyBuild and scale online and offline inference systems for LLMs and other AI modelsOptimize model execution across the full stack, including model architecture, runtime, compiler, kernel, and hardware layersDrive model optimization techniques such as quantization, pruning, compression, batching, and memory optimizationImprove GPU efficiency through low-level systems work, including kernel-level optimization, runtime tuning, and hardware-aware performance improvementsPartner closely with ML, infrastructure, and product teams to identify serving bottlenecks and improve end-to-end model performanceContribute to and/or extend open-source LLM serving frameworks such as SGLang, vLLM, Triton, or similar technologiesSet technical direction for model serving, inference performance, and next-generation AI infrastructure designMentor engineers and influence technical strategy across AI InfrastructureBasic QualificationsBA/BS degree in Computer Science or related technical field, or equivalent practical experience8+ years of experience in software engineering, distributed systems, infrastructure, or machine learning systemsExperience building or optimizing large-scale production ML systems, model serving platforms, or AI infrastructureExperience with GPU-based systems, CUDA, kernel optimization, or hardware-aware performance tuningExperience with large-scale inference systems, including latency, throughput, reliability, and cost optimizationExperience with deep learning frameworks such as PyTorch, TensorFlow, or similarExperience programming in one or more systems languages such as C++, Go, Python, or JavaPreferred QualificationsDeep experience with LLM serving infrastructure, AI inference platforms, or large-scale model deployment systemsFamiliarity with or contributions to open-source serving frameworks such as vLLM, SGLang, Triton, TensorRT, Ray, or similar technologiesExperience with ML compilers, runtimes, or graph optimization frameworks such as XLA, TVM, TensorRT, Triton, or similarAn understanding of model optimization techniques such as quantization, pruning, compression, batching, caching, and memory optimizationExperience improving GPU utilization and cost/performance efficiency for large-scale ML workloadsExperience building high-performance online or offline inference pipelinesAn understanding of distributed systems, scheduling, resource management, and large-scale infrastructure operationsExperience operating across the stack from model-level optimization to runtime, compiler, kernel, and hardware-level performance improvementsExperience influencing technical direction across teams and partnering effectively with ML researchers, infrastructure engineers, and product teamsSuggested SkillsAI/ML Systems and InfrastructureGPU and Performance OptimizationModel Serving and Inference SystemsDistributed SystemsTechnical LeadershipLinkedIn is committed to fair and equitable compensation practices.The pay range for this role is $198,000 to $326,000. Actual compensation packages are based on several factors that are unique to each candidate, including but not limited to skill set, depth of experience, certifications, and specific work location.
This may be different in other locations due to differences in the cost of labor.The total compensation package for this position may also include annual performance bonus, stock, benefits and/or other applicable incentive compensation plans. For more information, visit InformationEqual Opportunity StatementWe seek candidates with a wide range of perspectives and backgrounds and we are proud to be an equal opportunity employer. LinkedIn considers qualified applicants without regard to race, color, religion, creed, gender, national origin, age, disability, veteran status, marital status, pregnancy, sex, gender expression or identity, sexual orientation, citizenship, or any other legally protected class.LinkedIn is committed to offering an inclusive and accessible experience for all job seekers, including individuals with disabilities. Our goal is to foster an inclusive and accessible workplace where everyone has the opportunity to be successful.If you need a Reasonable Accommodation to search for a job opening, apply for a position, or participate in the interview process, connect with us and describe the specific Accommodation requested for a disability-related limitation.
Fill out an Accommodation request here: accommodations are modifications or adjustments to the application or hiring process that would enable you to fully participate in that process. Examples of reasonable accommodations include but are not limited to:Documents in alternate formats or read aloud to youHaving interviews in an accessible locationBeing accompanied by a service dogHaving a sign language interpreter present for the interviewA request for an accommodation will be responded to within three business days. However, non-disability related requests, such as following up on an application, will not receive a response.San Francisco Fair Chance OrdinancePursuant to the San Francisco Fair Chance Ordinance, LinkedIn will consider for employment qualified applicants with arrest and conviction records.Pay Transparency Policy StatementAs a federal contractor, LinkedIn follows the Pay Transparency and non-discrimination provisions described at this link: Data Privacy Notice and Compliance Posters for Job CandidatesPlease use this link to access documents that provide information about how LinkedIn handles the personal data of employees and job applicants, as well as the E-Verify Participation Notice and the Department of Justice Immigrant and Employee Rights Section Right to Work posters:
Sr. Staff Software Engineer, Systems Infrastructure in mountain view at Unknown Company
This position is listed as contract and hybrid.