Unknown Company

Data Scientist

new york, ny • Posted 2 weeks ago
Onsite Full Time IT & Technology

About Sunset

About SunsetAt its core, Sunset was founded to help founders. We started by supporting startups through shutting down, but we have since expanded into unlocking a new revenue stream for all types of businesses.In 2025, we had a unique insight: the data every company generates each day through collaboration, communication, and building is some of the most valuable training data in the world. Public and synthetic data can only get frontier models so far, so the next generation of model progress depends on real, proprietary data grounded in how actual businesses operate. We are a primary source of it, partnering directly with the frontier AI labs building what comes next.

Why Join Sunset Now

We have scaled from $0 to a multi-eight-figure run rate in a matter of monthsWe have raised from top-tier investors, including Floodgate, Afore, Ludlow, and Hustle FundWe are small enough that you will carry outsized responsibility and grow as quickly as the company doesYou will partner with and build for some of the fastest and most important companies in the worldYou will help build a massive, category-defining business from the ground floor

The Role

Sunset turns sensitive internal enterprise data into de-identified datasets without destroying the structure and meaning that make the data valuable. That creates a difficult measurement problem. A system can improve aggregate F1 while missing a high-risk slice, remove more sensitive information while also destroying useful context, or pass one stage while defects escape somewhere else in the pipeline.As Sunset's first Data Scientist focused on evaluation, you will establish how we know whether that data is actually getting better. You will build the datasets, experiments, quality measures, and feedback loops that expose hidden failures, accelerate model and pipeline improvement, and give the team confidence in what it delivers.This is a hands-on, zero-to-one role at the intersection of data science, AI, and a real production system. You will write Python and SQL, construct evaluation corpora, study failure patterns, design comparisons, calibrate human and model-based judgments, and turn the result into a clear decision. The questions are scientifically difficult, but the output must be practical enough to change what the team builds and ships.You will work closely with Machine Learning, Product Engineering, Data Engineering, Security, Quality, domain experts, and the team making delivery decisions. Machine Learning Engineers own changing model behavior. You own the credibility of the evidence used to decide whether a model, pipeline, or delivery change actually made the data safer or more useful.

Questions You Might Answer

  • Did a higher NER or entity-resolution score actually reduce sensitive misses across the messages, documents, tables, and providers that matter?
  • Is a new model finding more sensitive information, or simply removing more of the useful structure our customers need?
  • Can we trust a golden dataset, a human review process, or an LLM judge enough to use it for a release decision?
  • Which customer, modality, entity, language, or format slices are hidden by a strong aggregate result?
  • Where did a quality loss enter between source data, processing, de-identification, review, and delivery?
  • What is the smallest credible experiment that would tell us whether to ship, revise, or stop a change?

What You’ll Do

  • Define what high-quality and safe-to-deliver data mean across de-identification, structure preservation, semantic coherence, and customer utility
  • Design representative samples and build golden, adversarial, replay, and production-like corpora with explicit provenance, labeling policy, agreement, adjudication, and versioning
  • Turn ambiguous concepts such as “useful,” “clean,” or “s afte …

#J-18808-Ljbffr

Data Scientist in new york at Unknown Company

This position is listed as full time and onsite.

Back to Job Search