Unknown Company

Associate Performance Engineer

md • Posted 3 days ago
Onsite Full Time IT & Technology

Responsibilities

  • Assist with Performance Tuning: collaborate with a senior engineer to analyze and tune Spark/PySpark and AWS Glue job performance, including partitioning, clustering, batch sizing, and file format optimization such as Parquet, ORC, Iceberg.
  • Support Query Optimization: optimize query performance across Athena, Trino, PostgreSQL, Redshift, and semantic-layer objects, contributing to efficient execution of user queries, reports, and dashboards.
  • Collaborative Troubleshooting: identify, diagnose, and resolve failed production jobs and performance bottlenecks.
  • Support Performance Testing: participate in periodic load and performance tests; define benchmarks, document and analyze results, and draft test reports with recommendations for system improvements.
  • Pipeline Monitoring: monitor pipeline health using CloudWatch, ETL metadata, log analysis, and dashboards; harden and optimize error handling and automated job notifications.
  • Team Collaboration: coordinate with developers, testers, and stakeholders to resolve technical issues; participate actively in agile ceremonies and Program Increment planning events.
  • Production Operations Support: work as part of a team to support production operations across a large data pipeline portfolio.

Requirements

  • Bachelor's degree in computer science, Information Systems, Engineering, or related technical discipline
  • In lieu of a degree, four additional years of related, specialized experience is required
  • Minimum of 4 years of experience in performance engineering, data engineering, or quality engineering for large‑scale, data‑centric systems
  • Hands‑on experience tuning Apache Spark/PySpark workloads and AWS Glue jobs
  • Experience with workflow orchestration using Airflow/MWAA or similar tools
  • Experience optimizing queries and storage across S3‑based data lakes (Iceberg/Parquet), query engines (Athena or Trino), and relational databases (PostgreSQL, Redshift, or Oracle)
  • Foundational proficiency in Python and SQL
  • Experience with load/performance testing and production troubleshooting of automated data pipelines at scale
  • Strong analytical, root‑cause analysis, and communication skills
  • Ability to obtain and maintain a Moderate Risk Public Trust clearance; residing in the United States

Core Competencies

Demonstrates expertise in performance tuning and optimization of data pipelines using Spark/PySpark and AWS Glue, along with strong analytical skills for troubleshooting and enhancing query performance across various data storage and processing technologies.

#J-18808-Ljbffr
Back to Job Search