Design and implement data models across relational, graph, and lakehouse systems using AWS RDS, AWS Neptune, and Databricks
Develop warehouse/lakehouse schemas in Databricks Delta Lake for analytics and ML
Implement graph-based retrieval and integrate with vector search for AI use cases
Generate embeddings and manage hybrid retrieval pipelines (graph + vector)
Build ETL/ELT workflows using Databricks (PySpark), Airflow, or AWS Glue
Ensure data quality, consistency, and freshness across systems
Tune queries and optimize storage for Neptune, RDS, and Delta Lake
Work with data scientists and AI engineers to support knowledge graph and RAG workflows
Document data models and maintain governance standards
Requirements
Master’s or Bachelor’s Degree in Computer Science or engineering
7–10 years of experience in data analysis, data model design, and implementation to deliver scalable, enterprise-grade data solutions in a fast-paced environment
Manage/mentor team members, and provide training as required
Excellent analytical and interpersonal skills
Excellent verbal, written communication and presentation skills
Very strong practical experience in implementing large mature incident and change management processes
Team player, Ethical, Self-driven and self‑motivated, goal and results oriented
Experience in building, managing diverse geographically located teams and championing customer satisfaction
Experience in support or implementation of IAM products
Experience in implementing SDLC engagements projects, that must include activities such as requirements gathering, analysis, design, development, testing, deployment and application support
3–5 years of experience in data engineering or similar roles
Strong SQL and schema design for AWS RDS (PostgreSQL/MySQL)
Hands‑on experience with AWS Neptune (Gremlin/SPARQL) and graph modeling
Proficiency in Databricks, PySpark, and Delta Lake
Familiarity with GraphRAG concepts, embeddings, and vector search
Programming in Python for data pipelines and API integration