Unknown Company

Lead Data Engineer-PySpark, RedShift, Airflow, AWS

jericho, ny • Posted 4 days ago
Onsite Full Time General

Lead Data EngineerCandidate should have 12+ years of experience in Data Engineering. Must have strong work experience with onshore-offshore modelDesigning, creating, testing and maintaining the complete data management & processing systems.Candidate need to have in depth understanding of how data pipelines are builtTypical challenges with fetching data from various sources. How incremental/CDC data flows are handled.How do you ensure data qualityHow do you do Data profilingHands-on experience with PySpark, Redshift (SQL) and Airflow at minimumStrong hands-on with required tech skills, flexible, right attitude to play the lead roleShould be able to design and document data model at various levelsWorking closely with the stakeholdersBuilding highly scalable, robust & fault-tolerant systemsKnowledge of Hadoop ecosystem and different frameworks inside it – HDFS, YARN, MapReduce, Apache Pig, Hive, Flume, Sqoop, ZooKeeper, Oozie, Impala and KafkaMust have experience on SQL-based technologies (e.g.

MySQL/ Oracle DB) and NoSQL technologies (e.g. Cassandra and MongoDB)Should have Python/Scala/Java Programming skillsDiscovering data acquisitions opportunitiesFinding ways & methods to find value out of existing dataImproving data quality, reliability & efficiency of the individual components & the complete system

Back to Job Search