Data EngineerLocation: Hybrid Onsite (3 days onsite in Mclean, VA)Duration: 6-12+ months ContractMust Haves:PythonSparksJsonXMLSnowflakeSQLYour Impact:You will be able to apply advanced Data Engineering and Machine learning skills to solve real world challenges in building applications that help the company build better models and do advanced reporting.Cleanse, manipulate and analyze large datasets (Structured and Unstructured data – XMLs, JSONs, PDFs) using Hadoop platform.Develop Python, PySpark, Spark scripts to filter/cleanse/map/aggregate data.Manage and implement data processes (Data Quality reports)Develop data profiling, deduping logic, matching logic for analysisProgramming Languages experience in Python, PySpark and Spark for data ingestion.Programming experience in BigData platform using Hadoop platform.Present ideas and recommendations on Hadoop and other technologies best use to management.Qualifications:Bachelor’s degree in Computer Science, Statistics, Data science or a related quantitative field.5+ years of experience in processing large volumes and variety of data (Structured and unstructured data, writing code for parallel processing, XMLS, JSONs, PDFs)5+ years of programming experience in Hadoop, Spark, Python for data processing and analysis.Strong SQL experience is a must5+ years of experience using Hadoop platform and performing analysis. Familiarity with Hadoop cluster environment and configurations for resource management for analysis work2+ Prior experience working in Cloud platforms AWS. Kubernetes experience is highly desirableHands on Work experience using technologies for manipulating structured and unstructured big data.
Big data technologies may include—but are not limited to—Hadoop, Hive, Spark, relational databases, and NoSQL.Desired prior experience with MPP Databases like SnowFlake.Detail oriented and superb communication and written skills