Unknown Company

Machine Learning Research Engineer

new york, ny • Posted 3 days ago
Onsite Full Time Electrical & Energy Engineering

You will focus on expanding our ML research platform to benchmark, rapidly prototype, and stress-test both software and hardware layers across our entire distributed ML stack. By leveraging AI agents and auto-research capabilities, you will push our systems to their limits, identify bottlenecks, and create a frictionless environment to test novel machine learning models on realistic, large-scale data.

Responsibilities

  • Platform Validation & Infrastructure Benchmarking:
    • Serve as the primary feedback loop for the entire ML stack.
    • Actively run complex models through our full ML pipeline to comprehensively test both the training and inference environments.
    • Validate the central infrastructure in practice, seeing exactly how new research ideas fare and identifying system bottlenecks before broader rollout to research teams.
  • Streamline Rapid Prototyping for ML Research:
    • Build high-level abstractions that allow users to bypass setup friction.
    • Integrate our core ML tooling directly with our underlying simulation and data frameworks, providing a unified entry point to access our full tech stack.
    • Enable rapid iteration on real-world data and seamless distributed training via Ray.
  • Agentic Workflows for ML Research:
    • Leverage AI agents and auto-research workflows to autonomously generate experiments, stress-test our distributed clusters, and provide data-driven, actionable feedback on what infrastructure needs to be optimized or built next.
  • Research Platform Feedback & Insights Sharing:
    • Act as the critical bridge between infrastructure builders and ML researchers.
    • Be the first to exhaustively test new models and push the platform's limits.
    • Document and publish empirical findings on system capabilities and hardware performance.
    • Take your validated insights to assist engineering teams with platform improvements and advise researchers on how to best leverage the stack.

Qualifications

  • Strong Software Engineering Foundation:
    • Deep proficiency in Python and software design principles.
    • Ability to build clean, scalable APIs and abstractions that other developers and researchers are enthusiastic about using.
  • Applied Machine Learning:
    • Hands-on experience with modern frameworks (PyTorch, TensorFlow, etc.)
    • Strong practical understanding of how to train, evaluate, and deploy models at scale.
  • Distributed Compute:
    • Experience scaling ML workloads across GPUs and multi-node clusters using frameworks like Ray, Dask, or PyTorch Distributed.
  • AI Agent Workflows:
    • Familiarity with LLM tooling, agentic frameworks, and using AI to automate coding, research, or testing tasks.
  • System Profiling & Optimization:
    • Ability to debug and identify bottlenecks across hardware and software layers (e.g., memory limits, GPU utilization, data pipeline latency).

Comp: $200-300K + Bonus

#J-18808-Ljbffr
Back to Job Search