Sr. Data Scientist role supporting a federal healthcare insurance organization, delivering end-to-end data science projects with statistical modeling, ML, and AI in a mostly remote setup.
Responsibilities
- Drive solutions to complex business problems using statistical modeling , machine learning , AI , operations research , and data mining
- Own high-impact data science projects end-to-end, from inception through production deployment
- Lead and coordinate other data scientists working on shared project efforts
- Build, train, and deploy ML solutions on AWS , using SageMaker , Bedrock , Kendra , and Lambda
- Develop production-ready code with documentation and unit tests
- Apply advanced ML approaches including k-nearest neighbors , random forests , and ensemble methods
- Partner with product teams and customers to explain and educate on AI , ML , and statistical models
- Collaborate with Product teams to refine requirements and interface with stakeholders across departments
Requirements
- 8+ years of experience as a Data Scientist with both model building and deployment experience
- Advanced Python and SQL proficiency
- Advanced proficiency in Python plus Spark/Scala for statistical analysis, data modeling, ML, and ETL
- Intermediate to advanced data visualization skills using Python
- Solid hands-on knowledge of AWS SageMaker , AWS Bedrock , AWS Kendra , and AWS Lambda (required)
- Ability to write production-ready code including documentation and unit tests
- Hands‑on experience with ML methods including k-nearest neighbors , random forests , and ensemble methods
- Strong AI/ML expertise across Machine Learning , Deep Learning , Decision Trees , Random Forest , Neural Networks , Supervised/Unsupervised Learning , Forecasting , Predictive Modeling , and Clustering
- Deep knowledge of machine learning fundamentals , data mining , and statistical predictive modeling
- Proficiency with Python ML and data pre-processing libraries: Scikit-Learn , NumPy , Pandas
- Strong software prototyping and engineering skills across Python , R , and Spark/Scala
- Ability to initiate and drive projects to completion with minimal guidance
- Strong communication skills for presenting analysis results clearly and effectively
- Degree preferred; 4 additional years of experience may substitute for a degree
Location
- Mostly Remote
- Onsite requirements: 1x/month in Reston, VA (requirement also stated as 1–2 times/month in the DC Metro area)
- Candidates must reside in DC, MD, or VA
Project Details
- Mostly remote role with onsite requirements 1–2 times per month in the DC Metro area
- Travel expenses will not be covered
- Candidates must reside in the DC, MD, or VA area or a touching state
Technologies
- Python, SQL, Spark, Scala
- AWS SageMaker, AWS Bedrock, AWS Kendra, AWS Lambda
- Scikit-Learn, NumPy, Pandas
- R
- k-nearest neighbors, random forests, ensemble methods
- Deep learning, decision trees, neural networks
Nice-to-Have Skills
- Experience with Agentic AI
- Familiarity with statistical packages such as R , MATLAB , SPSS , SAS , or Stata
- Proficiency with healthcare analytics and data structures
- Experience with big data technologies, ETL, statistics, causal inference, and simulation
- Experience with large data sets and distributed computing (Hive/Hadoop)
- Prior experience leading data science projects or teams independently