Lead Data Engineer
Seize your opportunity to make a personal impact supporting the Case Management Modernization (CMM) Program. The CMM program is an initiative to support the Administrative Office of the US Courts (AO) in developing a modern cloud-based solution to support all 204+ federal courts across the United States. GDIT is your place to make meaningful contributions to challenging projects and grow a rewarding career. The Lead Data Engineer will work as part of the CMM Data Modernization and Governance team responsible for delivering an integrated data governance, engineering, data platform, reporting, analytics, and Artificial Intelligence (AI)/Machine Learning (ML) capabilities that support operational decision-making and fulfill AO's data and analytics objectives in support of the CMM program.
The Lead Data Engineer partners closely with the Data Architecture, Data Governance, and Migration teams to ensure engineering work is consistent, auditable, and aligned with CMM's cloud-based CI/CD standards.
Responsibilities
Data Platform & Architecture
- Develop and implement a scalable, secure, cloud-based data platform supporting operational data, reporting, and analytics delivering cloud-based architecture (data lake, lakehouse, or data warehouse).
- Ensure alignment with federal security requirements, judiciary architecture standards, data governance policies, and application modernization initiatives.
- Engineer multi-tenant, cloud-based environments supporting hybrid/on-premises systems, enabling SQL, NoSQL, IaaS, PaaS, distributed SQL, multi-modal, and event-driven/streaming databases.
- Design and implement auditable data integration patterns across Judiciary systems and external platforms, including the legacy CM/ECF system during the coexistence period, with pipelines integrated into Government-provided CI/CD processes.
- Implement and maintain data quality controls, validation rules, and governance-aligned data structures, coordinating with the Data Quality/Validation Engineer to sustain required accuracy and field-completeness thresholds on an ongoing basis.
- Drive continuous database and query performance monitoring and optimization, including automated performance tuning, query optimization, and indexing strategies.
- Coordinate with the Cloud Data Architect/Data Modeler to ensure engineering implementation stays aligned with the Data Architecture Blueprint and evolving data models.
- Document data platform architecture, integrations, data quality metrics, and service statistics, updating this documentation each Program Increment (PI).
- Partner with DevSecOps/CI-CD engineering to ensure data pipelines are built, tested, and deployed through Government-provided CI/CD tooling.
- Escalate and help resolve technical risks, defects, and dependencies affecting data engineering delivery across the CMM program.
- Support development of the Data and Technology Enabler Roadmap across data engineering, reporting/analytics, and AI/ML use cases.
Data Engineering & Integration
- Design and develop data ingestion, ETL/ELT pipelines for integrating data from multiple enterprise sources, and transformation logic for near-real-time and batch workloads.
- Lead the implementation and ongoing maintenance of data management solutions, including operational databases, document storage services, data models, schemas, and data access APIs.
- Provide technical direction and day-to-day oversight to the Data Engineers (ETL/ELT Pipelines) team, reviewing designs and ensuring consistent engineering standards across pipelines.
- Implement Infrastructure as Code (IaC) for database provisioning, configuration, and management to ensure consistency, repeatability, and auditability.
- Implement robust, reusable data services supporting analytics, reporting, and downstream data marts and / or gold layers.
- Collaborate with architects to implement logical and physical data models in cloud-based platforms (e.g., Snowflake, Databricks, or AWS Redshift)
- Develop and maintain high-quality, testable code using secure coding standards and best practices.
- Support the operationalization CMM Data Classification Standards, including workshops, security controls, metadata requirements, and integration into system design, procurement, and training.
- Configure metadata structures, workflows, automation, and security settings; develop ingestion processes and attribute definitions for catalog entries.
- Create training, job aids, and communications to support user adoption.
- Integrate data pipelines with cloud services, messaging, and storage components.
- Implement data quality checks, validations, and error-handling mechanisms.
- Optimize pipeline performance, scalability, and cost efficiency.
- Support CI/CD-enabled deployments, including automated testing and promotion across environments.
- Support incident resolution and root cause analysis for data pipeline failures.
- Produce and maintain technical documentation, runbooks, and workflows.
- Ensure deliverables meet federal security, governance, and audit requirements.
Data Migration & Archiving
- Partners closely with the Data Architecture, Data Governance, and Migration teams to execute the data migration, archival, and disaster recovery (DR) strategy.
- Deliver migration documentation for scripts, transformations, and validation results.
- Support the Failover Testing and DR drills engineering activities.
- Implement operational databases supporting data migration, archival, DR, document storage, indexing services, schemas, and secure data APIs using IaC for provisioning and configuration.
Qualifications
- MA/MS degree with 7+ years or BS/BA degree with 9+ years of general experience in information systems and 7+ years of specialized experience.
- Experience may be considered in lieu of degree as follows: HS (16+ years), AA/AS (14+ years), BA/BS (12+ years), Doctorate Degree/Ph.D. (9+ years).
- Experience in data engineering and design data architecture.
- Experience and understanding of best practices regarding system security measures.
- Experience in conducting research for advanced technologies to determine how IT can support business needs by leveraging software, hardware, or infrastructure.
- Experience with AWS data and compute services.
- Experience with data orchestration tools.
- Proven track record in software and data engineering roles.
- Hands-on experience building enterprise-scale data pipelines.
- Strong proficiency with SQL, Python and data transformation techniques.
- Experience developing in cloud-based data platforms.
- Familiarity with Agile/Scrum delivery environments.
- Hands-on experience with data platforms, data analytics, and AI/ML solutions.
- Proficiency with ETL/ELT frameworks.
- Experience with streaming or near-real-time data ingestion.
- Familiarity with data governance, metadata, and classification standards.
- Experience in leading and mentoring data engineers.
Communication & Organizational
- Excellent presentation and communication (oral and written) skills.
- Consultant mindset with the ability to work with high level customer stakeholders and build excellent customer relationships.
- Experience identifying and applying industry tools, solutions, methods best practices, and emerging technologies.
- Strong analytical skills and problem-solving skills with the ability to formulate and communicate recommendations for improvement.
- Demonstrated ability to work effectively, independently, and as part of a team.
Certifications (Preferred)
- Certified Data Management Professional (CDMP)
- Snowflake SnowPro Core
- Databricks Certified Data Professional
- AWS Certified Data Analytics - Specialty
- AWS Certified Solutions Architect - Professional
- AWS Certified Data Engineer
The likely salary range for this position is $128,039 - $173,229. This is not, however, a guarantee of compensation or salary. Rather, salary will be set based on experience, geographic location and possibly contractual requirements and could fall outside of this range. Scheduled Weekly Hours: 40 Travel Required: None Telecommuting Options: Remote Work Location: Any Location / Remote
Lead Data Engineer in Remote at Unknown Company
This position is listed as contract and able to be worked remotely.