Senior Data EngineerLocation: Mclean VADuration: 12+ months contractResponsibilities include:Cleanse, manipulate and analyze large datasets (Structured and Unstructured data – XMLs, JSONs, PDFs) using Hadoop platform.Develop Python, PySpark, Spark scripts to filter/cleanse/map/aggregate data.Be able to build Dashboards in R/Shiny for end user consumptionManage and implement data processes (Data Quality reports)Develop data profiling, deduping logic, matching logic for analysisProgramming Languages experience in Python, PySpark and Spark for data ingestionProgramming experience in BigData platform using Hadoop platformPresent ideas and recommendations on Hadoop and other technologies best use to managementQualifications:5+ years of experience in processing large volumes and variety of data (Structured and unstructured data, writing code for parallel processing, XMLS, JSONs, PDFs) - Mandatory3+ years of programming experience in Python, Spark for data processing and analysis. - MandatoryStrong SQL experience is a must - Mandatory3+ years of experience – using Hadoop platform and performing analysis. Familiarity with Hadoop cluster environment and configurations for resource management for analysis work2+ years of experience with containerization and orchestration.Hands on experience with AWS, Kubernetes, Kubeflow, Docker etc.
- MandatoryDetail oriented. Excellent communication skills (verbal and written)Must be able to manage multiple priorities and meet deadlinesDegree in Statistics, Economics, Business, Mathematics, Computer Science or related field