Senior AI Infrastructure Engineer - Model TrainingMountain View, CAKodiak Robotics, Inc. was founded in 2018 and has become a leader in autonomous ground transportation committed to a safer and more efficient future for all. The company has developed an artificial intelligence (AI) powered technology stack purpose-built for commercial trucking and the public sector. The company delivers freight daily for its customers across the southern United States using its autonomous technology.
In 2024, Kodiak became the first known company to publicly announce delivering a driverless semi-truck to a customer. Kodiak is also leveraging its commercial self-driving software to develop, test and deploy autonomous capabilities for the U.S. Department of Defense.Kodiak's AI is only as good as the speed at which we can train it. Every improvement to our models – from GigaFusionNet to large-scale world models – depends on infrastructure that turns thousands of hours of multimodal driving data into training throughput.
We are looking for engineers who make model training fast: streaming massive camera, LiDAR, and radar datasets without stalling a single GPU, sharding data and models efficiently across nodes, and extracting every FLOP from the latest hardware. If you measure your impact in tokens per second and GPU utilization, this role is for you. In this role, you will:Design high-throughput data loading and streaming systems for multimodal sensor data (camera, LiDAR, radar), including dataset formats, sharding strategies, and prefetching pipelines that keep GPUs saturatedBuild and optimize distributed training infrastructure across multi-node GPU clusters, applying data, tensor, pipeline, and fully sharded (FSDP/ZeRO) parallelism to models that don't fit on a single deviceMaximize utilization of modern accelerators such as NVIDIA B200s through mixed-precision training (BF16/FP8), fused kernels, memory optimization, and communication/computation overlapProfile end-to-end training pipelines to find and eliminate bottlenecks across storage, network, CPU preprocessing, and GPU computeDevelop scalable dataset construction pipelines that convert petabytes of raw driving logs into training-ready, streamable formatsPartner with ML teams to scale new architectures from prototype to full-cluster training runs efficiently and reliablyWhat you'll bring:BS, MS, or PhD in Computer Science or a related field, and at least 2-3 years of industry experience in ML systems or infrastructureHands-on experience with distributed training frameworks and techniques (PyTorch DDP/FSDP, DeepSpeed, Megatron, NCCL) and a strong grasp of parallelism trade-offsExperience building high-performance data pipelines for large-scale training, including streaming dataset formats (WebDataset, MosaicML Streaming/MDS, or similar), sharding, and storage/network-aware loadingDeep understanding of GPU performance: mixed precision, memory hierarchy, kernel fusion, profiling tools (Nsight, PyTorch Profiler), and interconnects (NVLink, InfiniBand)Strong Python skills and proficiency in PyTorch internals; systems-level experience (C++/CUDA/Triton) a plusPassion for building the infrastructure that lets AI for the physical world train faster, scale further, and improve continuouslyWhat we offer:Competitive compensation package including equity and annual bonusesExcellent Medical, Dental, and Vision plans through Kaiser Permanente, Cigna, and MetLife (including a medical plan with infertility benefits)MetLife Legal Services, Identity & Fraud Protection, Hospital Indemnity Insurance, Accident Insurance, & Critical Illness InsuranceFlexible PTO, 10 paid holidays, and generous parental leave policiesOur office is centrally located in Mountain View, CAOffice perks: dog-friendly, free catered lunch, a fully stocked kitchen, and free EV chargingLong Term Disability, Short Term Disability, Life InsuranceWellbeing Benefits - Headspace through Cigna, Calm through Kaiser, One Medical, Gympass, Spring Health through Cigna, Rula (mental health navigation)Fidelity 401(k)Commuter, FSA, Dependent Care FSA, HSAVarious incentive programs (referral bonuses, patent bonuses, etc.)The pay range listed below reflects the base salary in our SF/Silicon Valley location, across several internal levels. Actual starting pay will be based on job-related factors including: work location, experience, relevant training, education, skill level and performance during interview.
Total compensation at Kodiak includes base pay, equity, bonus and a competitive benefits packageCalifornia Pay Range$190,000 - $260,000 USD
Senior AI Infrastructure Engineer - Model Training in mountain view at Unknown Company
This position is listed as full time and onsite.