Unknown Company

Data Engineering Lead

suitland, md • Posted 4 days ago
Onsite Contract General

Data Engineering LeadThe Data Engineering Lead is responsible for designing and implementing modern, scalable data architectures to support migration of legacy, file-based analytical systems to AWS Cloud Native environments.This role leads the transformation of legacy SAS-based data storage models—including flat files, batch outputs, and subsystem-specific data artifacts—into structured, governed, and scalable data models optimized for cloud-native processing.The Data Engineering Lead will ensure data integrity, performance, and visibility across a system-of-systems modernization initiative, while providing technical leadership for data modeling, ingestion patterns, validation frameworks, and transparency reporting.Expert-level proficiency in Python and strong experience designing AWS-based data architectures are required.Key ResponsibilitiesLegacy Data Discovery & Data Model TransformationParticipate in structured system inventory efforts to document:Legacy file-based storage structuresSAS dataset dependenciesSubsystem data flowsManual gating and handoff processesAnalyze legacy storage models and design target-state data models aligned to AWS Cloud Native architecture.Replace file-driven batch dependencies with:API-based ingestionEvent-driven workflowsDatabase-backed storage (e.g., Aurora/Postgres)Define canonical data schemas and transformation standards.Cloud-Native Data Architecture DesignArchitect scalable AWS data pipelines using services such as:S3GlueLambdaEventBridgeSNS/SQSAurora/PostgresBatchAthenaDesign data ingestion, staging, transformation, and validation workflows.Establish schema management, versioning, and data lineage practices.Optimize data storage for performance, scalability, and cost efficiency.Support serverless and containerized data processing architectures.Expert Python-Based Data EngineeringDevelop advanced Python-based data transformation and validation pipelines.Implement modular, reusable data processing components.Optimize large-scale data manipulation for distributed execution.Develop high-performance ETL/ELT frameworks.Embed automated validation checks directly into data pipelines.Expert-level Python proficiency is required, particularly for:High-volume data processingData validation logicModular data engineering frameworksData Accuracy, Validation & VisibilityDesign and implement automated data validation frameworks to ensure:Functional equivalence during migrationRecord-level and aggregate-level consistencyDownstream compatibility across subsystemsDevelop dashboards and reporting mechanisms providing:Data accuracy metricsPipeline health indicatorsVariance detection summariesEnable transparency into data transformation impacts across modernization phases.Support regression validation through golden datasets and automated comparisons.System-of-Systems Data CoordinationCoordinate with Senior Developers and Requirements Engineers to align data models with application modernization.Ensure upstream/downstream data contract stability.Prevent data thrashing during phased migration.Support orchestration of gated workflows through automated triggers rather than manual file exchanges.Collaborate across workstreams to establish shared data standards.DevSecOps & Governance AlignmentIntegrate data pipelines into CI/CD frameworks.Support infrastructure-as-code alignment (Terraform/CloudFormation collaboration).Ensure compliance with security controls (IAM, encryption, key management).Produce documentation supporting:Architecture review boardsInterface control documentsData flow diagramsSupport ATO-related data validation evidence.

Back to Job Search