About The Role
The role owns the full lifecycle of machine learning systems, from experimental research and prototyping to reliable, low-latency production deployments.
The engineering team builds robust ML infrastructure and production-grade models that scale to support core business operations and millions of daily queries.
Key Responsibilities
- Design, train, and validate machine learning models using PyTorch and Python for complex production use cases
- Build high-throughput data ingestion and feature engineering pipelines using Apache Spark and distributed computing frameworks
- Deploy, monitor, and scale models on AWS using SageMaker, Docker, and Kubernetes with strict latency and reliability guarantees
- Implement comprehensive monitoring systems to track data drift, concept drift, and prediction latency in production
- Collaborate with backend engineers to integrate model endpoints into scalable microservices and APIs
- Conduct code reviews, write unit and integration tests, and establish best practices for machine learning
What We Are Looking For
- 3-6 years of professional experience in software engineering, with at least 3 years focused on machine learning engineering in production
- Expert proficiency in Python and deep familiarity with PyTorch or TensorFlow
- Hands-on experience with cloud infrastructure and MLOps tools such as AWS SageMaker, MLflow, Docker, and Kubernetes
- Solid understanding of distributed systems, data structures, and algorithms
- BS or MS in Computer Science, Machine Learning, or a related quantitative field
- Bonus: Experience with large language models, vector databases, or real-time streaming architectures using Kafka
Machine Learning Engineer in phoenix at Unknown Company
This position is listed as full time and onsite.