Unknown Company

Data Scientist

dallas, tx • Posted 2 days ago
Remote Full Time IT & Technology

Vytwo is hiring a Data Scientist to build end-to-end machine learning and NLP solutions across both structured and unstructured data. The role is based in Dallas, TX (onsite) , with the India team also open to local consultants eligible for INDIA . You will work on forecasting and predictive modeling as well as low-latency, production-ready NLP and LLM systems.

This position sits at the intersection of traditional analytics and modern generative AI, with an emphasis on deploying models and maintaining them in production using Databricks and Azure. Expect a blend of large-scale feature engineering, model lifecycle work, and real-time inference for text understanding and semantic retrieval.

What you’ll do

  • Build, deploy, and optimize ML models for predictive analytics, forecasting, classification, and regression .
  • Perform large-scale feature engineering using PySpark and big data tools.
  • Develop batch pipelines , including model versioning and experiment tracking .
  • Create cost estimation and risk/likelihood models using statistical and machine learning techniques.
  • Build NLP pipelines using deep learning frameworks such as PyTorch and TensorFlow (or similar).
  • Develop real-time, low-latency inference for use cases including text classification, embeddings, semantic search, summarization, and retrieval.
  • Design prompts, context graphs, and agentic workflows for LLM-based systems.
  • Apply prompt engineering , context engineering , and autonomous agent frameworks in production systems.
  • Use Databricks for ETL, feature engineering, model training, and orchestration.
  • Use Azure services for model deployment, data pipelines, and infrastructure.
  • Collaborate using Git-based workflows and leverage AI coding tools such as GitHub Copilot and Claude Code .
  • Implement model monitoring, observability, drift detection , and performance tracking.

Key focus areas

  • Structured Data (80–90%) : predictive analytics, forecasting, cost estimation, likelihood modeling, and batch-oriented machine learning pipelines.
  • Text / Unstructured Data (NLP & GenAI) : low‑latency real‑time systems using deep learning, LLMs, prompt engineering, and agentic AI frameworks.

What you bring

  • 3+ years of data science experience.
  • Strong hands‑on experience with Databricks (including Delta Lake , MLflow , and Job Orchestration ).
  • Excellent PySpark skills for large-scale distributed data processing.
  • Proficiency in Azure services, including ADF , Azure ML , AKS , and Databricks on Azure .
  • Strong understanding of ML algorithms , statistical methods , and data analysis.
  • Deep learning experience with PyTorch and TensorFlow .
  • Experience with Transformers via HuggingFace .
  • Experience with model monitoring and ML observability .
  • Ability to write clean, optimized code and leverage AI code assistants.
  • Prompt engineering experience involving task prompts, chain of thought, tool calling, and retrieval prompts.
  • Context engineering experience including retrieval pipelines, RAG , memory management, and context structuring.
  • Knowledge of LLM-based agentic frameworks such as LangChain , Semantic Kernel , CrewAI , and AutoGen .
  • Vector databases and embedding models are plus .

Technologies you may use

  • PySpark , PyTorch , TensorFlow , Transformers (HuggingFace) , Databricks
  • Delta Lake , MLflow , Job Orchestration , Azure
  • ADF , Azure ML , AKS , Databricks on Azure
  • Git , GitHub Copilot , Claude Code , LangChain , Semantic Kernel , CrewAI , AutoGen
  • Vector databases, embedding models, Docker , Kubernetes
  • Azure DevOps , GitHub Actions , Kafka , EventHub , Spark Streaming
  • REST APIs

Flexible work option

  • Flexible work from home options available .

Good to have

  • Experience with containerization (Docker , Kubernetes , AKS ).
  • Experience deploying models to production (real‑time endpoints and REST APIs ).
  • Knowledge of streaming technologies (Kafka , EventHub , Spark Streaming ).
  • Understanding of CI/CD for ML using Azure DevOps and/or GitHub Actions .

#J-18808-Ljbffr
Back to Job Search