Job Requirements
- Design, build, and own large-scale data pipelines, ETL/ELT workflows, and data warehouse architecture (primarily AWS and Snowflake)
- Lead dbt implementation including modeling conventions, testing frameworks, and documentation standards
- Architect and deliver high-performance Python APIs and microservices connecting the data platform to downstream applications
- Build backend services supporting both real‑time and batch access patterns
- Design and manage production data workflows across AWS, Kubernetes, Docker, and Airflow using infrastructure‑as‑code (Terraform/CloudFormation)
- Establish and enforce data quality, governance, and coding standards across internal and external data products
- Lead AI/LLM‑enabled data access initiatives to make complex datasets more discoverable and usable
- Drive cross‑functional, multi‑team data initiatives from architecture through production delivery
- Serve as the senior technical voice on platform strategy, observability, and infrastructure decisions
- Mentor engineers and potentially carry direct management responsibilities
Job Qualifications
Required
- 10+ years of experience in data or software engineering with a proven track record on production‑grade systems at scale
- Expert‑level SQL — complex transformations, query tuning, and performance optimization
- Expert Python skills — ETL/ELT frameworks, FastAPI or Flask, clean modular code
- Proven dbt architecture experience — advanced modeling, testing, and documentation
- Deep hands‑on experience with a cloud data warehouse (Snowflake, BigQuery, or Redshift), including security and performance tuning
- Production pipeline orchestration experience using Airflow, Dagster, or equivalent
- Core engineering fundamentals: system design, Git, CI/CD, Docker, Terraform
- Solid AWS experience: S3, IAM, API integrations, security configuration
- Strong communication and technical leadership skills in ambiguous, fast‑paced environments
Preferred
- Experience with biological or bioinformatics data (PCR, genomic sequencing, wastewater surveillance) or working alongside lab teams
- Snowflake production ownership: RBAC, authentication patterns, cost controls
- Observability stack experience: Grafana, Datadog, or similar
- Experience with data ingestion tooling: Airbyte, Fivetran, or custom solutions
- Data governance and cataloging exposure
- BI tool data modeling experience (Tableau, Looker, Metabase)
- Kubernetes familiarity for container orchestration
- Comfort with AI-assisted development tools (Copilot, Cursor)
#J-18808-Ljbffr