Lead Data EngineerThe Lead Data Engineer will report to the Credit Card team. The candidate will be responsible for identifying relevant data and utilizing data tools, technologies and processes to develop continuous, data driven and automated customer communications for the Credit Card team.RequirementsExperience in AWS GlueExperience in Apache ParquetProficient in AWS S3 and data lakeKnowledge of SnowflakeUnderstanding of file-based ingestion best practices.Scripting language - Python & pysparkCore ResponsibilitiesCreate and manage cloud resources in AWSData ingestion from different data sources which exposes data using different technologies, such as: RDBMS, REST HTTP API, flat files, Streams, and Time series data based on various proprietary systems. Implement data ingestion and processing with the help of Big Data technologiesData processing/transformation using various technologies such as Spark and Cloud Services.
You will need to understand your part of business logic and implement it using the language supported by the base data platformDevelop automated data quality check to make sure right data enters the platform and verifying the results of the calculationsDevelop an infrastructure to collect, transform, combine and publish/distribute customer data.Define process improvement opportunities to optimize data collection, insights and displays.Ensure data and results are accessible, scalable, efficient, accurate, complete and flexibleIdentify and interpret trends and patterns from complex data setsConstruct a framework utilizing data visualization tools and techniques to present consolidated analytical and actionable results to relevant stakeholders.Key participant in regular Scrum ceremonies with the agile teamsProficient at developing queries, writing reports and presenting findingsMentor junior members and bring best industry practicesQualifications5-7+ years’ experience as data engineer in consumer finance or equivalent industry (consumer loans, collections, servicing, optional product, and insurance sales)Strong background in math, statistics, computer science, data science or related disciplineAdvanced knowledge one of language: Java, Scala, Python, C#Production experience with: HDFS, YARN, Hive, Spark, Kafka, Oozie / Airflow, Amazon Web Services (AWS), Docker / Kubernetes, SnowflakeProficient with: Data mining/programming tools (e.g. SAS, SQL, R, Python)Database technologies (e.g. PostgreSQL, Redshift, Snowflake.
and Greenplum)Data visualization (e.g. Tableau, Looker, MicroStrategy)Comfortable learning about and deploying new technologies and tools.Organizational skills and the ability to handle multiple projects and priorities simultaneously and meet established deadlines.Good written and oral communication skills and ability to present results to non-technical audiencesKnowledge of business intelligence and analytical tools, technologies and techniques.Familiarity and experience in the following is a plus: AWS certificationSpark StreamingKafka Streaming / Kafka ConnectELK StackCassandra / MongoDBCI/CD: Jenkins, GitLab, Jira, Confluence other related toolsRequired Skills: Amazon Web Services (AWS) SNS/SQS, EC2, S3, IAM Services & PySpark/Spark, and Parquet Basic Qualification: Additional Skills: Background Check: Yes Drug Screen: Yes Notes: The candidate needs to be onsite in Charlotte, NC - 2-3 days a week.