Unknown Company

Spark job migration specialist

ca • Posted 5 days ago
Onsite Full Time General

Spark Job Migration SpecialistA Spark job migration specialist migrates data pipelines, JAR tasks, and analytics workloads from legacy systems (like Hadoop/CDH or AWS EMR) to ACOS modern platforms. This involves refactoring code (e.g., Hive to PySpark), performance testing, and updating Spark 2.x to 3.x.Key Job ResponsibilitiesWorkload Migration: Migrate JVM workloads and Spark-Submit tasks to Databricks JAR tasks or Notebook tasks.Pipeline Re-engineering: Convert existing HiveQL scripts and Oozie workflows into optimized Spark SQL or PySpark applications.Refactoring: Adapt data pipelines from Azure Synapse to any cloud platform, including updating library dependencies and notebook references.Performance Optimization: Implement Adaptive Query Execution (AQE) in Spark 3 to improve shuffle performance and fix skew joins.Testing & Validation: Perform regression testing to ensure output consistency between old and new systems using validation scripts.Job Customization: Use spark.sparkContext.setJobDescription() to label, monitor, and troubleshoot specific Spark tasks in the UI.Job Description/ProfileRole: Big Data Migration Engineer (Spark)Experience: 5+ years experience with Apache Spark (PySpark/Scala) and Cloud platforms (Azure/AWS).Requirements:Strong experience with HDFS, Hadoop ecosystem (Hive, Spark, HBase, MapReduce).Experience in data migration to cloud / enterprise data platforms.Knowledge of: Data ingestion tools (Sqoop, Kafka, NiFi, etc.

Back to Job Search