Unknown Company

Information Technology_USA - USA_Data Scientist

atlanta, ga • Posted 1 weeks ago
Onsite Full Time General

Senior Data Engineer (Databricks, PySpark)Location: Atlanta, GADuration: 6 monthsLead data engineering initiatives using Databricks and Spark ecosystemDesign and build scalable big data pipelines for enterprise use casesAct as a senior technical contributor and mentor junior engineersCollaborate with data science and BI teams for analytics solutionsDrive data architecture and engineering best practicesWork closely with business stakeholders to define data requirementsDevelop and optimize ETL and data ingestion frameworksEnsure high data quality, governance, and compliance standardsSupport real-time and batch data processing systemsProvide strategic insights and recommendations using data analysisGather, analyze, and define data requirements for structured and unstructured dataBuild and maintain scalable data pipelines using PySpark and DatabricksDevelop data ingestion, transformation, and aggregation workflowsImplement data validation, cleansing, and normalization processesDesign and manage Delta Lake architectures and data modelsWork with streaming technologies like Kafka and Spark StreamingOptimize data pipelines for performance and scalabilityDevelop and enforce data governance, retention, and anonymization policiesCollaborate with cross-functional teams for data integration solutionsSupport machine learning and analytics teams with data preparationUse version control tools like Git for code managementLead and mentor junior team members and support multiple projectsStrong hands-on experience with Databricks platformDeep understanding of Apache Spark and PySparkProficiency in Python for data engineering and analyticsExperience with Delta Lake and Databricks ecosystem toolsStrong SQL skills for data querying and transformationExperience with big data tools like Hadoop, Hive, KafkaKnowledge of cloud platforms like AWS or GCPExperience in building ETL pipelines and data architecturesFamiliarity with streaming data processing frameworksStrong problem-solving and analytical skillsDatabricks and Spark ecosystemPySpark and Spark SQLData pipeline development and ETL processesBig data architecture and distributed systemsDelta Lake and data lakehouse conceptsStreaming and batch processing (Kafka, Spark Streaming)Cloud data engineering (AWS/GCP)Data modeling and transformationVersion control using GitAgile methodologies (Scrum, Kanban, SAFe)Experience Required: 6-8 years or above

Back to Job Search