We are hiring an Applied Data Scientist / Applied Scientist to improve how we perform entity resolution at scale. You will develop and test new ML, embedding, and LLM-based approaches for deciding when messy records from different sources refer to the same real-world business.
The work is centered on model quality, experimentation, and evaluation; engineering partners will help productionize successful approaches. A key part of the role is figuring out how far we can push newer foundation-model techniques while keeping the system practical for hundreds of millions of entities.
Key Responsibilities
Develop better ways to match company records
- Build new ML, embedding, and LLM-based approaches for matching entities
- Improve how the system handles messy data, including name variations, aliases, domains, websites, firmographic attributes, multilingual records, and data hierarchies.
- Develop scoring and ranking approaches that help separate true matches from lookalikes, duplicates, or unrelated entities.
- Evaluate and implement AI and machine learning techniques that improve entity resolution while balancing accuracy, scalability, and cost.
- Design approaches with scale in mind. We are not looking for solutions that only work by calling an expensive model on every possible pair of records.
Improve evaluation, experimentation, and match quality
- Lead the development of practical ways to measure match quality, including precision, recall, false positives, false negatives, confidence, coverage, and manual review burden.
- Assist in Building trusted benchmark sets that allow us to compare new models against the current matching engine before production rollout.
- Use LLM-assisted review and validation as both a matching technique and a performance ceiling for cheaper, more scalable approaches.
- Turn ambiguous matching problems into clear hypotheses, experiments, metrics, and recommendations.
Partner with engineering to bring successful ideas into production
- Work closely with data engineering and software engineering teams to turn promising prototypes into production-ready matching logic.
- Provide clear model specifications, expected behavior, evaluation results, edge cases, and rollout criteria for engineering partners.
- Help decide which matching techniques should be used for different data tiers, confidence levels, and cost profiles.
- Support iterative improvements to the matching engine by measuring impact, diagnosing regressions, and recommending model or logic changes.
- Communicate tradeoffs clearly, especially around match quality, scale, cost, latency, explainability, and operational risk.
Qualifications
Required Skills
- Strong applied ML fundamentals, with hands-on experience building and evaluating models on real data.
- Excellent Python and SQL skills.
- Practical experience with embeddings, semantic similarity, LLMs, or related AI techniques.
- Hands-on experience training supervised and unsupervised models, including classification and NLP tasks.
- Working knowledge of neural network and transformer architectures.
- Proficiency with common ML frameworks such as TensorFlow, PyTorch, and PyCaret.
- Experience retraining a taxonomy classifier or maintaining classification models in production.
- Experimental judgment: able to define baselines, metrics, test sets, and error analysis that show whether quality improved.
- Ability to explain model behavior, tradeoffs, and edge cases clearly to engineering and business partners.
Preferred Skills
- Experience with entity resolution, record linkage, deduplication, or similar matching problems.
- Experience with ranking, similarity scoring, retrieval, clustering, or candidate generation.
- Experience applying LLMs or embeddings to business problems where cost and scale matter.
- Exposure to large-scale data platforms such as Spark, Snowflake, Databricks, or BigQuery.
- Familiarity with company, domain, website, firmographic, or other business‑entity data.
Data Scientist in town of poland at Unknown Company
This position is listed as full time and onsite.