Senior Software Engineer, Infrastructure Software for AI (Centralized AI Data Centers & Distributed AI-RAN Environments)
We challenge conventional limits by building transformative products that fully exploit our state-of-the-art (SOTA) infrastructure—including NVIDIA GB200 (e.g., liquid‑cooled NVL72 rack‑scale systems with 72 Blackwell GPUs and 36 Grace CPUs unified via massive NVLink domains), MGX modular architectures, and DGX Grace Hopper platforms—combined with cloud‑native software. These solutions power massive centralized AI data centers and pioneering distributed AI Radio Access Network (AI‑RAN) deployments, where AI is embedded directly into radio networks for superior efficiency, edge intelligence, and new capabilities.
We're seeking experienced senior engineers passionate about innovation, ready to shape foundational AI infrastructure software on world‑class GPU systems.
Minimum Qualifications
- Bachelor's degree in Computer Science, Electrical Engineering, or a related technical field.
- 5+ years of experience in software engineering, hardware platforms, distributed systems, or infrastructure development.
- 2+ years in technical lead roles, owning high‑impact projects and leading teams.
- Proven hands‑on experience building systems software, AI frameworks, or applied AI systems.
Preferred Qualifications
- Master's degree in a relevant technical discipline (e.g., CS, Systems Engineering).
- Direct experience with Kubernetes and container orchestration at scale.
- Hands‑on work with GPU‑accelerated systems and high‑performance computing (HPC) environments.
- Familiarity with AI developer frameworks, MLOps tools, automation pipelines, and CI/CD systems.
Role Overview
Join our infrastructure team as a key contributor building foundational systems software atop next‑generation GPU platforms to support large‑scale AI workloads (training, fine‑tuning, and serving). Own significant portions of our new AI infrastructure stack, with a strong emphasis on Kubernetes orchestration and GPU resource management. Drive architectural innovation in systems software and automation to achieve maximum efficiency and utilization. As a Senior Software Engineer, collaborate closely with Engineering Leads, Product Management, and Program teams to execute toward commercialization and market impact.
Key Responsibilities
- Design and develop robust systems software to enable AI workloads on large‑scale GPU clusters (e.g., GB200 NVL72 rack‑scale systems).
- Deliver critical control plane components for workload scheduling, orchestration, and resource management; build management plane for underlying hardware platforms.
- Create northbound APIs and interfaces for customer portals and self‑service access to the infrastructure.
- Contribute to product requirements documents (PRDs), sprint planning, and agile program execution.
- Help attract, mentor, and grow top engineering talent.
- Exemplify and cultivate a culture of humility, bold innovation, and disciplined delivery to bring products to market.
We are pushing the boundaries of AI by challenging conventional approaches and building groundbreaking products that fully harness our state‑of‑the‑art (SOTA) infrastructure—including NVIDIA GB200 NVL72 (liquid‑cooled, 72‑GPU rack‑scale systems with massive NVLink domains for exascale performance), MGX modular architectures, and DGX Grace Hopper platforms. These products power both massive centralized AI data centers and innovative distributed AI Radio Access Network (AI‑RAN) deployments, where AI is natively integrated into radio networks for enhanced efficiency, edge intelligence, and new revenue opportunities.
We're seeking passionate, experienced practitioners eager to drive innovation and deliver transformative AI products on world‑class hardware.
Minimum Qualifications
- Bachelor's degree in Computer Science, Electrical Engineering, Mathematics, Statistics, or a related technical field.
- 3+ years of hands‑on experience in machine learning, deep learning, and software engineering.
- Proficiency in Python; experience with C/C++.
- Strong working knowledge of major AI/ML frameworks (PyTorch, TensorFlow, JAX, or similar).
- Solid foundation in data structures, algorithms, and software design principles.
Preferred Qualifications
- Master's or PhD in Computer Science, AI/ML, or a related discipline.
- Experience with Large Language Models (LLMs), Generative AI, or Computer Vision.
- Familiarity with distributed training frameworks and techniques (e.g., Ray, DeepSpeed, Megatron‑LM).
- Proven expertise optimizing models for GPU inference (e.g., TensorRT, Triton Inference Server).
- Knowledge of MLOps tools and practices (Kubeflow, MLflow, etc.).
Role Overview
As a core member of our AI engineering team, you will design, develop, and optimize cutting‑edge AI models and workloads that run natively on our high‑performance GPU clusters. Leverage our SOTA infrastructure to train, fine‑tune, and serve massive‑scale models at unprecedented efficiency. Collaborate across infrastructure, product, and research teams to align hardware capabilities with real‑world AI demands, driving breakthroughs in performance, scalability, and innovation.
Key Responsibilities
- Design, implement, and train state‑of‑the‑art ML models for high‑impact applications (e.g., NLP, Computer Vision, Network Optimization).
- Optimize AI workloads for extreme performance and scalability on large‑scale GPU systems like GB200 NVL72, using tools such as Dynamo, vLLM, and advanced inference engines.
- Partner with cross‑functional teams to co‑design hardware‑software solutions that maximize AI processing efficiency.
- Build robust tools, data pipelines, evaluation frameworks, and deployment systems.
- Track and incorporate the latest AI research and technological advancements.
- Contribute to product requirements (PRDs) and agile execution (sprint planning and delivery).
- Champion a culture of humility, bold innovation, and high‑velocity product delivery.