Sr. Data EngineerMountain View, CA / Denver, CO / RemoteFull-TimePosition Summary:We are looking for a Sr. Data Engineer to join our growing Data Platform and Engineering teams.
The ideal candidate has significant experience in building scalable data platforms that enable business intelligence, analytics, data science and data products. They must have strong, hands-on technical expertise in a variety of technologies and the proven ability to fashion robust scalable solutions. They must be at ease working in an agile environment with little supervision.
The ability to work across teams with product managers, data scientists and business stakeholders to translate sometimes vague business requirements into working code will be critical to success in this role. This person should embody a passion for continuous improvement and data quality.Responsibilities:Design and implement data processing pipelinesIntegrate data from multiple data sources, develop cross-platform ETL processesData Validation and VerificationAnalyze data, solve problems, and implement solutions for ensuring data quality and deliveryCreate systems for data acquisition and wranglingDevelop new tools and processes for managing our data workflows and data infrastructureCollaborate with our Engineering and Data Science teams on building, maintaining, and monitoring the database infrastructureCollaborate with product managers, data scientists, business users and other engineers to define requirements and design solutionsClient and analyze data from the web (census, open data, commercial vendors)Skills/Requirements:Expert in reporting, analytics, and databasesData ingestion, ETL and storageInterest in pulling data from many sourcesExperience in big data, data mining and statistical analysisCloud computing, especially AWS technologies (S3, EC2, etc.)Comfortable choosing technologies that fit the application (e.g., MySQL versus PostgreSQL, Hadoop versus Cassandra)More than 5 years of experience in object-oriented development with PythonOther languages like Scala, C++, Java, or similar are a plusExperience with sparkExpertise with SQLFamiliarity with DockerMachine Learning libraries and frameworks like scikit-learn, TensorFlow, Pytorch a plusDeploying algorithms at scale