Research EngineerAt Mind Robotics, we're building generalized physical AI—robotic systems capable of dexterous, adaptive, and reasoning-intensive work in real-world industrial environments. Our ability to iterate quickly on large-scale models depends on world-class ML infrastructure.We're looking for a Research Engineer to build the core systems that enable fast, reliable, and scalable model training—powering everything from experimentation to production deployment.ResponsibilitiesDesign and implement scalable systems for training large ML modelsEnable efficient workflows for data ingestion, training, and iterationDevelop and optimize distributed training systems across hundreds of GPUsImplement strategies for parallelization, sharding, and efficient compute utilizationImprove training efficiency through techniques such as attention optimizations, kernel fusion, and memory managementPartner closely with modeling teams to accelerate iteration speed and reduce training costsBuild internal tools for experiment tracking, monitoring, and debuggingImplement systems for tracking training performance, failures, and resource utilizationDebug and resolve bottlenecks across the training stackProvide lightweight infrastructure support for deploying and running models in production environmentsOptimize inference performance and reliability where neededSupport core cloud infrastructure needs for training workloads (without heavy DevOps overhead)Manage compute resources efficiently across training jobsQualificationsStrong experience building infrastructure for large-scale ML trainingDeep understanding of how modern LLM/VLM systems are trained and scaledProven experience setting up and scaling distributed training across hundreds of GPUsStrong understanding of parallelization strategies (data, model, pipeline parallelism)Strong proficiency in Python programmingExpert-level proficiency in PyTorch and/or JAXStrong understanding of techniques like attention optimization, kernel fusion, and efficient memory usageNice to HaveExperience supporting inference systems in productionFamiliarity with robotics or embodied AI workloadsExperience building tools for experiment management and researcher productivity