Job TitleThis position may be offered to a candidate authorized to work in the US for his/her/their stated employer, without any restrictions which would prevent the candidate from working on the proposed assignment for the duration of the assignment period.Must Haves: Object oriented programming experience using Python. SQL. API experience is mandatory with a preference for familiarity with Boto3.
Experience with Pyspark with a solid understanding of big data.Your Work Falls Into Two Primary Categories:Strategy Development and ImplementationDevelop data filtering, transformational and loading requirementsDefine and execute ETLs using Apache Sparks on Hadoop among other Data technologiesDetermine appropriate translations and validations between source data and target databasesImplement business logic to cleanse & transform dataDesign and implement appropriate error handling proceduresDevelop project, documentation and storage standards in conjunction with data architectsMonitor performance, troubleshoot and tune ETL processes as appropriate using tools like in the AWS ecosystem.Create and automate ETL mappings to consume loan level data source applications to target applicationsExecution of end to end implementation of underlying data ingestion workflow.Operations and TechnologyLeverage and align work to appropriate resources across the team to ensure work is completed in the most efficient and impactful wayUnderstand capabilities of and current trends in Data Engineering domainQualificationsAt least 5 years of experience developing in Python, SQL (postgres/snowflake preferred) Bachelor’s degree with equivalent work experience in computer science, data science or a related field.
Experience working with different Databases and understanding of data concepts (including data warehousing, data lake patterns, structured and unstructured data) 3+ years’ experience of Data Storage/Hadoop platform implementation, including 3+ years of hands-on experience in implementation and performance tuning Hadoop/Spark implementations. Implementation and tuning experience specifically using Amazon Elastic Map Reduce (EMR). Implementing AWS services in a variety of distributed computing, enterprise environments.
Experience writing automated unit, integration, regression, performance and acceptance tests Solid understanding of software design principlesKey to Success in This RoleStrong consultation and communication skills Ability to work with and collaborate across the team and where silos exist Deep curiosity to learn about new trends and how to do things better Ability to use data to help inform strategy and directionTop Personal Competencies to PossessSeek and Embrace Change – Continuously improve work processes rather than accepting the status quo Growth and Development – Know or learn what is needed to deliver results and successfully competePreferred SkillsUnderstanding of Apache Hadoop and the Hadoop ecosystem.
Experience with one or more relevant tools (Sqoop, Flume, Kafka, Oozie, Hue, Zookeeper, HCatalog, Solr, Avro). Deep knowledge on Extract, Transform, Load (ETL) and distributed processing techniques such as Map-Reduce Experience with Columnar databases like Snowflake, Redshift Experience in building and deploying applications in AWS (EC2, S3, Hive, Glue, EMR, RDS, ELB, Lambda, etc.