Join the team supporting one of the largest federal enterprises in history—an initiative originally rooted in President Roosevelt’s 1935 vision to secure the nation’s most important humanitarian program. This mission today remains the largest humanitarian program ever created, serving tens of millions of Americans and operating at a scale few data platforms in the world can match.
As a Senior Data Scientist, you will work with some of the most complex, high-volume datasets in the federal space. Your work will directly improve the accuracy, integrity, and accessibility of data powering critical national services. This role is ideal for a highly skilled technical expert who thrives in environments where the work truly matters.
Key Required Skills
- Deep, practical experience with NLP, text processing, and information extraction (NER, blocking/indexing, string distance metrics, TF-IDF/cosine similarity, phonetic encoding, address standardization)
- Strong Python, Regex, and SQL development skills
- Excellent communication and collaboration abilities
Position Description
- Design Advanced Analytics Solutions: Build and refine large-scale data pipelines, entity-resolution systems, and analytical workflows using Python and SQL across enterprise-grade data platforms.
- Elevate Data Quality: Clean, standardize, and manage massive datasets from diverse and complex sources, ensuring unmatched accuracy, consistency, and security.
- Drive Performance: Optimize high-complexity SQL and data operations to support rapid processing and scalable analytics workloads.
- Champion Engineering Excellence: Participate in code reviews, enforce version control best practices, and uphold rigorous standards for privacy, reproducibility, and maintainability.
- Deliver End-to-End Solutions: Support validation, testing, deployment, and monitoring in a fast-paced, mission-driven environment.
Basic Qualifications
- Bachelor’s degree in Statistics, Applied Mathematics, Computer Science, Information Science, or related fields
- 10+ years overall IT experience
- Hands-on experience with NLP, text processing, Python, SQL, Regex, and relevant libraries
Required Skills
- Candidate must be able to obtain and maintain a public trust clearance
- Candidate must be willing to work on-site in Woodlawn, MD, 5 days per week
- Master’s + 10 years, Bachelor’s + 12 years, or 18 years in lieu of degree
- Deep experience with NLP concepts, including NER, TF-IDF/cosine similarity, phonetic encoding, blocking/indexing, and advanced text cleaning
- Strong Python development for analytics and pipeline design
- Advanced SQL expertise for complex queries and optimization
- Proficient with Regex for text extraction and cleansing
- Experience with libraries such as spaCy, Scikit-Learn, and linkage frameworks (Splink, FastLink, Dedupe, recordlinkage)
- Strong grounding in engineering best practices, version control, and data protection
Desired Skills
- Prior experience with federal, state, or local government data projects
- Ability to independently architect and deliver large-scale data pipelines
- Experience with legacy and distributed data systems, including PostgreSQL, DB2, Oracle, SQL Server, Hadoop, and unstructured files
- Experience with Jenkins for CI/CD of data workflows
- Familiarity with automation tools for orchestrating large data cleansing operations
- Ability to translate complex algorithms (e.g., probabilistic matching thresholds) into clear business explanations for leadership
- Strong problem-solving skills and cross-functional communication abilities
Data Scientist in baltimore at Unknown Company
This position is listed as full time and onsite.