- You will help design and implement high impact AI and ML systems
- We work in cross-functional teams collaborating with Machine Learning Engineers, Data Engineers and Product
- Focus on a variety of content classification use cases, leveraging everything from traditional NLP to sophisticated LLMs and generative models
- Investigate methods of solving our most challenging problems at Scribd, at scale
- Collaborate with other Data Scientists, Machine Learning Engineers and ML Data Engineers on cross-functional projects
- Leverage any algorithm at your disposal: from classical Scikit-learn and NumPy models to custom Neural Networks in PyTorch to third party LLM APIs
- Process massive amounts of data with Python, SQL and Spark
- Align with stakeholders through written and verbal communications methods on the approaches and results of projects, while writing detailed, accurate and concise project documentation
Benefits
- Healthcare Insurance Coverage (Medical/Dental/Vision): 100% paid for employees
- 12 weeks paid parental leave
- Short-term/long-term disability plans
- 401k/RSP matching
- Tuition Reimbursement
- Learning & Development programs
- Quarterly stipend for Wellness, Connectivity & Comfort
- Mental Health support & resources
- Free subscription to Scribd + gift memberships for friends & family
- Referral Bonuses
- Book Benefit
- Sabbaticals
- Company wide events
- Team engagement budgets
- Vacation & Personal Days
- Paid Holidays (+ winter break)
- Flexible Sick Time
- Volunteer Day
- Company-wide Diversity, Equity, & Inclusion programs
Hands-on experience building ML pipelines and working with distributed data processing frameworks like Apache Spark, Databricks, or similarBachelors or Masters in relevant quantitative discipline including but not limited to Statistics, Computer Science, Data Science, Artificial Intelligence or another field with a strong quantitative focusWe are seeking a curious and collaborative individual with an eye for simplicity, end-end visibility and impact and that is excited about building models using massive amounts of data, using language models and deploying modelsIntermediate level or greater experience with SQL or PySparkIntermediate level in at least three of these fields: classification algorithms, natural language processing, search, information retrieval, named entity recognition, deep learning, generative models3+ years of post qualification experience developing machine learning models, working with systems at scale and deploying to production environmentsWe are seeking a Data Scientist II with experience developing and deploying machine learning modelsProficiency in Python
#J-18808-LjbffrData Scientist in dallas at Unknown Company
This position is listed as full time and onsite.