Lead Data EngineerWe are looking for senior candidates with at least 12+ years experience. We are looking for tech experts who can lead the project as well with good communication and stakeholder management skills. Must have onshore-offshore coordination experience.Here are the key expectations from the tech perspective: Primary skills are PySpark, RedShift, Airflow, AWSLocation: 777 Old Saw Mill River Rd, Tarrytown, NY 10591 (100% onsite role) – Please share profile local to NY/NJ only.Duration: 6+ MonthsJob Description:Candidate should have 12+ years of experience in Data Engineering.
Must have strong work experience with onshore-offshore modelDesigning, creating, testing and maintaining the complete data management & processing systems.Candidate need to have in depth understanding of how data pipelines are builtTypical challenges with fetching data from various sources. How incremental/CDC data flows are handled.How do you ensure data qualityHow do you do Data profilingHands-on experience with PySpark, Redshift (SQL) and Airflow at minimumStrong hands-on with required tech skills, flexible, right attitude to play the lead roleShould be able to design and document data model at various levelsWorking closely with the stakeholders.Building highly scalable, robust & fault-tolerant systems.Knowledge of Hadoop ecosystem and different frameworks inside it – HDFS, YARN, MapReduce, Apache Pig, Hive, Flume, Sqoop, ZooKeeper, Oozie, Impala and KafkaMust have experience on SQL-based technologies (e.g. MySQL/ Oracle DB) and NoSQL technologies (e.g.
Cassandra and MongoDB)Should have Python/Scala/Java Programming skillsDiscovering data acquisitions opportunitiesFinding ways & methods to find value out of existing data.Improving data quality, reliability & efficiency of the individual components & the complete system.