Data EngineerDesigning and building ETL pipeline using Sqoop, Hive, Map Reduce and Spark on on-prem and cloud environments. Functional programming using Python and Scala for complex data transformations and in-memory computations. Using Erwin for logical/physical data modeling and dimensional data modeling.
Designing and developing UNIX/Linux scripts for handling complex file formats and structures. Orchestration of workflows and jobs using Airflow and Automic, creating multiple Kafka producers and consumers for data transferring. Performing continuous integration and deployment (CI/CD) using tools like GIT, Jenkins to run test cases and build applications with code coverage using Scala test.
Analyzing data using SQL, Big Query monitoring the cluster performance, setting up alerts, documenting the designs, workflow. Providing production support, troubleshooting and fixing the issues by tracking the status of running applications to perform system administrator tasks.