Unknown Company

Senior Research Software Engineer

cambridge, ma • Posted 3 days ago
Remote Full Time General

Senior Research Software EngineerDesign, plan, and implement software and data services that support and enrich research productivity and reliability. Develop software and data services with researchers to ensure that modern standards of reproducible research are kept.Job-Specific Responsibilities:The Harvard Data Science Initiative (HDSI) is hiring a Senior Research Software Engineer (RSE) to support a portfolio of faculty-led research projects under the HDSI–AWS Impact Computing Alliance. This role is designed for an engineer who thrives in research settings and enjoys translating scientific goals into robust, efficient, and reproducible AI/ML systems.Rather than being tied to a single lab, the RSE will provide shared, cross-project engineering support—helping multiple teams accelerate discovery by building and optimizing machine learning infrastructure, improving performance on modern hardware (including AI accelerators), and enabling scalable execution in AWS and HPC environments.

Projects may span domains such as climate and environmental science, global health, and other areas aligned with the alliance's mission to deliver measurable social and environmental impact.This is a hands-on role with strong collaboration expectations: you'll work directly with researchers, HDSI technical leadership, and the alliance team to deliver production-grade research software and reusable technical patterns that benefit multiple projects across the Impact Computing umbrella.This position is a benefits-eligible, two-year term appointment through June 30, 2028.Core Responsibilities:Design, build, and maintain ML/AI systems and research software in Python and C/C++Develop and optimize machine learning training and inference pipelines for accelerator-based systemsApply systems- and compiler-level optimizations, including: Loop transformations, vectorization, parallelization, and hardware-specific tuning (e.g., SIMD)Implement and optimize kernels using CUDA, OpenMP, OpenCL, or accelerator-specific programming modelsContribute to or integrate with compiler and IR frameworks such as MLIR, LLVM, XLA, IREE, TVM, or HalideAnalyze and improve performance using profiling and diagnostics focused on: Latency, memory bandwidth, I/O throughput, and compute utilizationSupport execution in AWS cloud and HPC environments, including large-scale model training, profiling, debugging, scaling, cost/performance tuning, reliability, CI/testing, packaging, deployment, reproducibility engineeringFollow and promote modern ML and scientific software best practices: Experiment tracking, reproducibility, version control, testing, packaging, and documentationCollaborate closely with faculty, researchers, and AWS consulting partners on systems engineering, performance optimization, ML infrastructure, compilers/framework integration, cloud/HPC execution.Communicate technical findings, tradeoffs, and progress clearly to research stakeholders (including documentation and handoff-ready tooling)Working Conditions:Occasionally required to work outside of normal business hours, and may be contacted during off-hoursHybrid / primarily remote within approved payroll statesQualifications:Basic Qualifications are the minimum threshold a candidate must meet in order to be considered for this role.Minimum of seven years' post-secondary education or relevant work experienceAdditional Qualifications and Skills:BS or MS (or equivalent practical experience) in Computer Science, Computer Engineering, Data Science, or a closely related fieldStrong programming skills in Python/C/C++Experience working with ML frameworks such as PyTorch, TensorFlow, JAX, XLA, Triton, ONNX, Caffe2, or TensorRTProven experience in deep learning at scale, familiarity with the "alphabet soup" of distributed computing (DP, TP, SP, CP, EP)Experience with production environments, including Git-based workflowsExperience working in AWS cloud or HPC environments used for large-scale computationPrior experience in a research or research-adjacent environment, with an understanding of the scientific software lifecycleStrong communication skills and a collaborative working styleContributed to compiler infrastructures and optimization frameworks (MLIR, LLVM, XLA, TVM, IREE, Halide)Experience developing or optimizing high-performance with libraries or kernels (e.g., cuBLAS, cuDNN, CUTLASS, HIP, ROCm, or similar)Experience with distributed AI/ML training and performance optimization (e.g., PyTorch DDP, FSDP, DeepSpeed)Experience building tooling for runtime analysis, profiling, and performance diagnosticsExperience with secure or privacy-constrained data environments (e.g., HIPAA-aware engineering practices)Experience working in interdisciplinary research areas such as climate, environment, health, or astrophysicsCertificates and Licenses:Completion of Harvard IT Academy specified foundational courses (or external equivalent) preferredAdditional Information:Appointment End Date: This is a two-year term position, expected to end on 06/30/2028Standard Hours/Schedule: 35 hours per weekVisa Sponsorship: Harvard University is unable to provide visa sponsorship for this positionPre-Employment Screening: Harvard University requires pre-employment reference and background screenings: Identity and EducationOther Information:Position Type: Full-time, benefited, two-year term appointmentThis position will have a 3-month orientation and review period.

Back to Job Search