- Design, build, and maintain ML Ops pipelines supporting model training, validation, and deployment across AWS environments.
- Implement automation for model packaging, testing, deployment, and monitoring using CI/CD best practices.
- Collaborate with data engineers and data scientists to operationalize ML workloads within the data lakehouse ecosystem.
- Develop and maintain integrations between data ingestion, feature stores, and model repositories.
- Apply infrastructure-as-code (Terraform, AWS CDK, CloudFormation) to automate ML pipeline infrastructure.
- Implement and manage model versioning, reproducibility, and lineage tracking using tools such as MLflow or SageMaker Model Registry.
- Define and automate monitoring, alerting, and retraining strategies for deployed models.
- Ensure all ML infrastructure and pipelines meet enterprise security, compliance, and governance standards.
- Participate in code reviews, knowledge sharing, and continuous improvement of ML Ops practices.
- Mentor junior engineers and contribute to documentation, standards, and best practices for ML Ops across teams.
Requirements
- Bachelor's Degree in Computer Science, Data Engineering, or a related technical field or equivalent experience.
- 4+ years of professional experience in software, data, or ML engineering.
- 2+ years of direct experience implementing and maintaining ML pipelines in production.
- Strong proficiency in Python and familiarity with ML frameworks such as PyTorch, TensorFlow, or Scikit-learn.
- Hands-on experience with AWS services (SageMaker, Step Functions, Lambda, ECR, S3, Glue, IAM).
- Solid understanding of CI/CD, containerization (Docker).
- Experience with building CI/CD pipelines (Jenkins, Github Actions, etc.).
- Experience with infrastructure-as-code and automation (Terraform, AWS CDK, or CloudFormation).
- Strong understanding of data pipelines, ETL/ELT concepts, and feature engineering in a lakehouse environment.
- Proven ability to apply software engineering practices to machine learning workflows.
- Strong communication and collaboration skills across multidisciplinary teams.
Core Competencies
Demonstrates expertise in designing and maintaining ML Ops pipelines, utilizing AWS services and infrastructure-as-code tools to automate processes. Strong proficiency in Python and experience with ML frameworks, alongside a solid understanding of CI/CD practices and data engineering principles.
Highest-signal resume keywords
- ML Ops Pipeline Development
- AWS Services (SageMaker, Lambda, S3)
- Python Programming
- Infrastructure-as-Code (Terraform, CloudFormation)
- CI/CD Best Practices
ATS Optimization Keywords
Hard Skills
- ML Pipeline Implementation
- Python
- AWS CDK
- Terraform
- CI/CD
- Docker
- ETL/ELT Concepts
- Feature Engineering
- ML Frameworks (PyTorch, TensorFlow, Scikit-learn)
- Model Versioning
Soft Skills
- Communication
- Collaboration
- Mentoring
Certifications & Qualifications
- Bachelor's Degree in Computer Science or Related Field
Industry Keywords
- Machine Learning
- Data Engineering
- Model Deployment
- Automation
- Compliance
Tools & Technologies
- MLflow
- SageMaker Model Registry
- Jenkins
- Github Actions
- Data Lakehouse