Data Engineer -- 100% RemoteRate: $54/Hr on C2CEST or CST time zoneInterview Process: 1-2 roundsCandidate Requirements:• The role is for a mid-level person with 3+ years of experience…doesn't need to be super senior, just someone with a decent amount of AWS experience.• Eastern or Central time zone is ideal.Technical Requirements:• Experience with Steaming Data Tools - AWS Kinesis or AWS DMS for real-time streaming data is a nice to have, but not a requirement.• Data governance experience is relevant, but no specific tools are required.• OvalEdge is a nice to have but not required and he would be surprised if we can find someone with that specific experiencePrimary System/Tool:• Needs someone with AWS experience, as all infrastructure and architecture lives there.• Not looking for an Azure resource.• Biggest tools: Glue, Python, PI Spark.• PI Spark is adjacent to Python, but specific to Glue and data engineering systems within AWS.• Data engineer/cloud engineer can be used interchangeably.Responsibilities and Tasks:• Want the person to start going through the backlog of data sources, creating API connections, and pulling data into the data lake.• There are already two data engineers that can take the data through the rest of the way.• Hoping the person has experience with APIs, AWS, Glue, and other tools listed in the JD.• Will help support the OvalEdge implementation… The OvalEdge team handles most of the technical lift, but there is still some internal data engineering support needed. (OvalEdge is a data governance software system that will read all of the source data)• Candidates should be comfortable working with small teams and have agency in decision-making.• Main tasks include integrating six source systems into the data lake and supporting Oval Edge implementation.ROLE OVERVIEW:We are seeking a hands-on Data Engineer contractor whose core strength is building AWS Glue ingestion pipelines against REST APIs. The primary focus of this role is onboarding new source systems into our AWS data lake, moving data downstream through our medallion architecture, and helping establish CDC and real-time streaming capabilities.
The role also includes supporting the rollout of OvalEdge as our data governance and cataloging platform. This is an execution-focused role with real ownership of production infrastructure from day one.REQUIRED SKILLS & EXPERIENCE:• 3+ years of hands-on data engineering experience in a cloud environment• Strong proficiency building AWS Glue jobs in Python/PySpark for production data ingestion• Deep experience integrating with REST APIs — authentication patterns (OAuth, API keys, JWT), pagination, rate limiting, and error handling• Proven ability to independently onboard new SaaS data sources end-to-end, from API exploration to a running production pipeline• Experience with Apache Airflow for pipeline orchestration and monitoring• Working knowledge of AWS data lake patterns and services (S3, Glue, IAM, CloudWatch)• Familiarity with Apache Iceberg or similar open table formats• Comfort working independently in a remote, asynchronous environmentPREFERRED:• Hands-on experience with CDC tooling (AWS DMS, Debezium, Kinesis, or similar)• Experience designing or operating real-time or near-real-time streaming pipelines• Experience ingesting from platforms such as Zendesk, Jira, Lattice, or similar SaaS tools• Knowledge of SCD Type 2 patterns in PySpark• Familiarity with medallion architecture (Bronze/Silver/Gold) in a production lakehouse environment• Experience with data governance or cataloging tools (OvalEdge, Alation, Collibra, or similar)NICE TO HAVE:• SQL proficiency for light data validation and Gold layer support• Experience with dbt Cloud for transformation layer development• Exposure to Microsoft Fabric or Power BI connectivity• Background in professional services, compliance, or B2B SaaS data environments