Unknown Company

Senior Data Engineer with Databricks Exp. - 100% Remote

workfromhome • Posted Today
Remote Full Time Computer and Mathematical Occupations

Job Title: Senior Databricks Engineer

Location: 100% Remote

Interview Process: Video (2 Rounds)

Job Description:
Job Description Lead/Senior Databricks Data Engineer

Position: Lead Databricks Data Engineer / Databricks Architect

Experience: 10+ Years

Location: Remote

Job Summary

We are looking for an experienced Lead Databricks Data Engineer / Databricks Architect with 10+ years of overall experience in Data Engineering and strong hands-on expertise in Databricks, Apache Spark, PySpark, SQL, Delta Lake, and cloud-based data platforms .

The candidate will be responsible for designing and implementing scalable Lakehouse architectures , enterprise data pipelines, data integration solutions, governance frameworks, and high-performance analytics platforms using Databricks.

Key Responsibilities
  • Design and develop scalable data engineering solutions using Databricks and Lakehouse architecture .

  • Build robust ETL/ELT pipelines using PySpark, Spark SQL, Python, and SQL .

  • Design and implement Bronze, Silver, and Gold/Medallion architecture .

  • Develop and optimize Delta Lake tables, including MERGE, schema evolution, Change Data Feed, and incremental processing.

  • Build batch and real-time/streaming pipelines using Structured Streaming, Auto Loader, and Lakeflow .

  • Develop and manage Databricks Jobs/Workflows for pipeline orchestration, scheduling, dependencies, retries, and monitoring.

  • Implement enterprise data governance using Unity Catalog , including access control, data lineage, auditing, catalogs, schemas, and external locations. Unity Catalog provides centralized governance, access control, lineage, and auditing across Databricks data and AI assets. ()

  • Perform Spark and Databricks performance tuning , including cluster configuration, partitioning, caching, query optimization, Photon, and workload optimization.

  • Design data models supporting Data Warehousing, BI, Analytics, and AI/ML workloads .

  • Integrate Databricks with cloud platforms such as AWS, Azure, or Google Cloud Platform .

  • Work with cloud services such as AWS S3, Azure ADLS Gen2, Azure Data Factory, AWS Glue, Synapse, Event Hubs/Kafka/Kinesis , as applicable.

  • Implement CI/CD pipelines using Git, Azure DevOps/GitHub/Jenkins and Databricks deployment capabilities.

  • Work with Terraform/IaC for infrastructure provisioning and automation.

  • Troubleshoot production pipeline failures, performance issues, data-quality problems, and Spark/cluster issues.

  • Establish data quality, monitoring, logging, and observability practices.

  • Provide technical leadership, code reviews, architecture guidance, and mentorship to junior/mid-level engineers.

  • Collaborate with Data Architects, Data Scientists, Business Analysts, DevOps teams, and application teams.

Required Technical Skills
Databricks
  • Databricks Lakehouse Platform

  • Delta Lake

  • Unity Catalog

  • Databricks Workflows/Jobs

  • Lakeflow / Delta Live Tables

  • Auto Loader

  • Databricks SQL

  • Databricks notebooks

  • Databricks Asset Bundles

  • Photon

  • Cluster/workload optimization

Big Data
  • Apache Spark

  • PySpark

  • Spark SQL

  • Structured Streaming

  • Kafka

  • Batch and real-time data processing

Programming
  • Python

  • SQL

  • PySpark

  • Scala good to have

Cloud Strong experience in at least one
  • AWS: S3, Glue, EMR, Lambda, Redshift, IAM, Kinesis

  • Azure: ADLS Gen2, ADF, Synapse, Azure DevOps, Event Hubs, Key Vault

  • Google Cloud Platform: GCS, BigQuery, Dataflow, Pub/Sub

Data Engineering
  • ETL/ELT

  • Data Warehousing

  • Dimensional Modeling

  • Data Lake/Lakehouse

  • Medallion Architecture

  • CDC

  • Data Quality

  • Data Governance

  • Metadata and Data Lineage

DevOps / CI-CD
  • Git

  • Azure DevOps / GitHub

  • Jenkins

  • Terraform

  • CI/CD automation

  • Infrastructure as Code

Preferred / Nice-to-Have Skills
  • MLflow

  • Databricks Machine Learning

  • Feature Store

  • Mosaic AI / GenAI

  • dbt

  • Apache Airflow

  • Power BI / Tableau

  • Delta Sharing

  • Lakehouse Federation

  • Liquid Clustering

  • Data security and PII masking

MLflow is particularly useful if the role touches ML/AI, as Databricks supports model tracking, lifecycle management, and deployment workflows alongside governed data. ()

Qualifications
  • Bachelor's degree in Computer Science, Engineering, Information Technology, or related field.

  • 10+ years of experience in Data Engineering / Big Data / Analytics.

  • 4+ years of hands-on Databricks experience preferred.

  • Strong experience designing enterprise-scale data platforms.

  • Demonstrated experience leading technical projects and mentoring engineers.

  • Strong communication and stakeholder-management skills.

#J-18808-Ljbffr

Senior Data Engineer with Databricks Exp. - 100% Remote in workfromhome at Unknown Company

This position is listed as full time and able to be worked remotely.

Back to Job Search