- Lead AI System Design & Development
- Develop real-time and batch inference pipelines integrated with streaming data platforms.
- Design feature engineering pipelines leveraging high-volume behavioral and content metadata.
- Implement end-to-end ML workflows from data ingestion to model serving.
- Build AI-Powered Data Products
- Develop production-grade AI services that power user-facing and internal data products.
- Design APIs and services to expose AI capabilities to downstream applications and platforms.
- Ensure tight integration between AI systems and the core data platform.
- Architect Scalable ML Infrastructure
- Define architecture for model training, evaluation, deployment, and monitoring.
- Build and optimize feature stores, model registries, and inference services.
- Design systems that support low-latency, high-throughput model serving.
- Establish best practices for reproducibility, versioning, and lifecycle management.
- Production Reliability & Model Performance
- Monitor and optimize model performance, latency, and system reliability in production.
- Implement observability for data quality, feature drift, and model degradation.
- Establish automated testing, validation, and deployment pipelines for ML systems.
- Ensure scalability and cost efficiency across AI workloads.
- Cross-Functional Collaboration
- Partner with Data Engineers to integrate AI pipelines with real-time and batch data systems.
- Collaborate with Product Managers to define AI-driven product capabilities and roadmap.
- Work with Software Engineers to integrate AI services into user-facing applications.
- Align with analytics and experimentation teams to measure model impact.
- Technical Leadership
- Lead architectural decisions for AI/ML systems and data-driven applications.
- Mentor engineers in machine learning engineering, system design, and best practices.
- Establish standards for model development, deployment, and operational excellence.
- Drive innovation in applied AI across streaming and content platforms.
Requirements
- Strong experience building and deploying machine learning models in production.
- Expertise in recommendation systems, personalization, ranking models, or NLP.
- Experience with model training frameworks (e.g., TensorFlow, PyTorch, or similar).
- Understanding of feature engineering, model evaluation, and experimentation frameworks.
- Experience designing large-scale feature pipelines using batch and streaming data.
- Strong knowledge of data modeling and transformation for ML use cases.
- Familiarity with feature stores and real-time feature serving architectures.
- Experience integrating ML systems with real-time data platforms (e.g., Kafka, Pub/Sub).
- Understanding of event-driven architectures and low-latency processing patterns.
- Ability to design real-time inference and decisioning systems.
- Strong experience with cloud-native architectures (GCP preferred).
- Experience deploying ML systems in Kubernetes-based environments.
- Understanding of distributed systems, scalability, and fault tolerance.
- Proficiency in Python, Java, or similar languages for production systems.
- Experience building microservices and APIs for model serving.
- Strong software engineering fundamentals, including testing, CI/CD, and observability.
- Strong foundation in machine learning engineering, data systems, and distributed architecture.
- Proven track record of building and scaling AI/ML systems in production environments.
- Experience working with real-time data platforms and high-scale user-facing systems.
- Ability to balance long-term architecture with rapid product delivery.
- Excellent leadership, problem-solving, and cross-functional collaboration skills.
- Self-motivated, quality-driven, and focused on delivering measurable impact through AI.
Core Competencies
Demonstrates expertise in building and deploying machine learning models in production, with a strong focus on real-time inference systems and scalable AI architectures. Proficient in integrating AI capabilities with data platforms and ensuring model performance and reliability.
Highest-signal resume keywords
- Machine Learning Model Deployment
- Real-Time Inference Systems
- Feature Engineering Pipelines
- Cloud-Native Architectures
- Cross-Functional Collaboration
ATS Optimization Keywords
Hard Skills
- Machine Learning Engineering
- Feature Engineering
- Model Evaluation
- Data Modeling
- Python
- Java
- TensorFlow
- PyTorch
- Kubernetes
- APIs
Soft Skills
- Leadership
- Problem-Solving
- Collaboration
- Self-Motivated
- Quality-Driven
Industry Keywords
- AI Systems
- Real-Time Data Platforms
- Scalable ML Infrastructure
- Event-Driven Architectures
- Distributed Systems
Tools & Technologies
- Kafka
- Pub/Sub
- CI/CD
- Observability
- Microservices