Unknown Company

Data Engineer Lead

washington, dc • Posted 3 days ago
Onsite Full Time General

Data EngineerA mission-driven organization supporting a vital medical professional community is seeking a Data Engineer to join its Data Governance and Analytics team. This role focuses on building and optimizing scalable data pipelines, enabling advanced analytics and machine learning models, and supporting enterprise data governance standards. You'll work on high-impact initiatives that shape data strategy and architecture across the organization.The ideal candidate is a hands-on engineer with expertise in Databricks, Azure, and CI/CD automation, combined with strong Python and SQL skills.

You'll lead technical best practices, mentor peers, and drive innovation in data engineering and DevOps.Minimum Qualifications:Bachelor's degree in Computer Science, Engineering, or related field7+ years building enterprise-level data solutions on cloud platformsProven expertise in Databricks, Azure, Python, SQL, and Apache SparkStrong experience with CI/CD automation and DevOps practicesFamiliarity with data governance frameworks and security complianceExcellent collaboration and communication skillsNice To Have: Experience in Non-Profit, Associations, or Mission Driven OrganizationsResponsibilities:Design and optimize scalable data pipelines for analytics, AI/ML, and BI workloadsImplement ETL/ELT solutions for multi-source data ingestion and transformationDrive DevOps adoption with CI/CD automation and Infrastructure as CodeManage and optimize Azure cloud data storage for cost and performanceEnsure data quality, integrity, and compliance with governance standardsCollaborate with architects, analysts, and business stakeholders to deliver data-driven solutionsMentor junior engineers and foster continuous improvementEvaluate emerging technologies to enhance scalability and innovationDesired Skills:Databricks – Advanced experience with Databricks Workflows for orchestrating production-grade pipelinesAzure Cloud – Expertise in Azure Data Factory, Data Lake, Synapse, Key Vault, and MonitorCI/CD & DevOps – Automated deployments, Infrastructure as Code (Terraform), GitHub Actions, Azure DevOpsProgramming – Strong Python and SQL for ETL/ELT, analytics, and ML workflowsApache Spark (PySpark) – Distributed data processing and real-time analyticsData Governance – Data quality, lineage, cataloging, and compliance frameworksETL/ELT Design – Structured streaming, batch ingestion, orchestration (Airflow, Kafka)RESTful APIs – Integration with OAuth authentication, JSON parsing, and scalable pipelinesMentorship – Ability to guide junior engineers and promote best practices

Back to Job Search