Wirestock is one of the leading data platforms for ethically sourced multimodal data. We serve some of the world’s top AI labs, including several foundation models, by providing high-quality, fully licensed training datasets. As the AI data landscape undergoes a major shift, we are scaling rapidly to meet rising demand for curated visual data.
About the Role
We're hiring a Principal Machine Learning Engineer to lead the data-science function powering our content platform - the curation, quality control, and enrichment of a petabyte-scale library using computer vision and AI.
This role sits at the intersection of data science, data engineering, and computer vision. The ideal candidate possesses deep algorithmic expertise and the proficiency to manage massive data infrastructures. You will define how content is understood across our library, building upon a robust foundation and leading a high-performing team.
You'll work closely with our CTO and existing data team and play a central role in building out the data science function in San Francisco.
What You'll Own
Technical Ownership
- Multimodal Curation & QC: Manage curation and quality control for one of the market's largest content collections, ensuring library integrity at petabyte scale.
- Automated Enrichment: Deploy CV and LLM models for classification, object detection, and metadata enrichment to enhance content discoverability and value.
- Standards & Criteria Design: Define and automate grading standards tailored to various content types, building consistent and scalable evaluation models.
- Big Data Infrastructure: Execute complex algorithms across AWS and on-site lakehouse environments, focusing on video understanding and multi-modal classifiers.
Leadership & Management
- Team Building & Mentorship: Recruit, train, and supervise a growing ML team, fostering professional development and maximizing productivity through performance data.
- Strategic Alignment: Partner with Research, Product, and Engineering teams to refine Trust & Safety strategies and ensure the success of project SLAs.
- Executive Reporting: Report directly to the CTO, providing effective communication on risks, mitigation, and the evaluation of scalable tools and processes.
What You'll Bring
- 6+ years of experience in ML/Data Science with a core specialty in computer vision (6+ years in machine learning/data science, with computer vision (classification, object detection, content enrichment) as a core specialty.
- Proven track record of operating on large-scale data , demonstrating fluency in both AI algorithm depth and big-data engineering.
- The rare hybrid of CV/AI algorithm depth and big-data engineering ability .
- Experience in leading and mentoring data science teams within a fast-paced environment.
- Video/multimedia expertise and a background in content marketplaces or moderation platforms are significant advantages.
What You Won't Own
Production model training at scale from scratch. Your intellectual energy goes into applying computer vision and AI to curate, classify, and enrich content at volume — and into leading the data-science team that does it. That said, understanding how models are built and running small training experiments to validate the dataset and the enrichment quality are genuine advantages here, not disqualifiers.
Our Stack
Petabyte-scale image and video content
MongoDB as a Vector DB
dbt
Spark
Trino
Argo Workflows and Argo Events
Ness
We value the rare combination of CV/AI algorithm depth and big-data engineering fluency over familiarity with any single tool.
#J-18808-LjbffrPrincipal Machine Learning Engineer in san francisco at Unknown Company
This position is listed as full time and hybrid.