Responsibilities
- build dynamic troubleshooting agents that understand networks
- solve unstructured production log data complexities
- optimize hardware utilization for data collection
- automate synthetic datasets creation
- architect data infrastructure for real-time network failures analysis
- design and scale automated pipelines transforming raw production logs into insights
- develop systems generating synthetic data for edge cases learning
- tackle unique network complexity problems
- optimize data collection and hardware utilization
Requirements
- Bachelor's degree in STEM and 5+ years of relevant experience
- Master's degree in STEM and 3+ years of relevant experience
- PhD in STEM +0 years of relevant experience or equivalent related work experience
- 5+ years of experience in data engineering, machine learning engineering, or related roles
- Data Pipeline experience, designing and scaling data pipelines for unstructured or semi-structured data, including ingestion, cleansing, and auditing
- ML Infrastructure experience working with ML data workflows, including dataset creation, labeling, and evaluation
- Experience with Python and data processing frameworks (e.g., Spark, Beam, Ray)
- Experience with ML systems and tools, such as training pipelines and model evaluation frameworks
- Experience with human-in-the-loop ML systems, active learning, weak supervision or self-evolving agents (preferred)
- Exposure large language models, computer vision, or speech datasets (preferred)
- Experience building internal tools or platforms used by annotation or operations teams (preferred)
Hard Skills
- Data Engineering
- Machine Learning Engineering
- Data Pipeline
- Data Processing
- Synthetic Data Creation
- Real-Time Analysis
- Network Troubleshooting
- Data Cleansing
- Model Evaluation
- Active Learning
#J-18808-Ljbffr