We are seeking an experienced Machine Learning Engineer to join the Machine Learning Operations team in Client's R&D AI Center of Excellence. In this role, you will build and operate the infrastructure that turns Data Scientists' models into production systems - from scalable training and inference pipelines to optimized on-device deployment on edge hardware. The Machine Learning Operations team supports all projects in the Center of Excellence, including Client's inVue Dx's computer vision and microscopy-based diagnostics, and you will contribute to high impact problems.
Job Description:
- Partner with Data Scientists to productionize machine learning and computer vision models, and implement scalable, efficient training and inference pipelines
- Optimize models for performance and scalability, including GPU/CUDA-level tuning and edge/embedded inference optimization (e.g., TensorRT, DLA)
- Work with software engineers and Data Scientists to integrate ML models into production systems on the cloud and at the edge
- Collaborate with Data Scientists to identify and analyze data requirements
- Develop and implement data preprocessing and feature engineering pipelines
- Build and maintain infrastructure-as-code for ML platforms and deployments
- Stay up to date with the latest developments in machine learning, computer vision, and MLOps
Required Skillsets:
- Previous experience in machine learning as evidenced by delivering solutions into production
- Excellent software engineering skills including a test-driven mindset
- Strong programming skills in Python and Spark
- Experience with distributed computing frameworks such as Ray, and ML frameworks such as PyTorch and modern foundation/vision models
- Experience with ML tech stacks, including Databricks (including Databricks Asset Bundles), MLflow, Amazon AWS (EC2, S3, Lambda), containerization technologies such as Docker, and infrastructure-as-code tools including Terraform and CloudFormation
- Experience optimizing and deploying models for GPU or edge inference (CUDA, TensorRT); familiarity with NVIDIA edge platforms (Jetson Orin/Thor) and tooling (Nsight, DeepStream SDK) a plus
- Familiarity with computer vision and/or microscopy/scientific imaging data a plus
- Experience creating and deploying applications to AWS Lambda
- High level of comfort with Linux
- Experience with SQL
- Strong problem-solving and analytical skills
- Excellent communication and collaboration skills