Unknown Company

Data Engineer Lead

los angeles, ca • Posted Today
Onsite Contract Database

Capital Group Companies is hiring a senior, hands-on Data Engineering Lead to set technical direction for the CSGT data platform and data products in Los Angeles.

Responsibilities

  • Own data engineering strategy and roadmap for CSGT, including Lakehouse architecture on Databricks and AWS ; make and explain decisions on scalability , security , reliability , and cost
  • Shape standards and practices across adjacent teams and the broader Capital Group technology organization, contributing to centers of excellence and firm-wide engineering practices
  • Design and build ingestion, transformation, and serving pipelines using Databricks , PySpark , Delta Lake , dbt , and Airflow ; establish reusable patterns and frameworks for consistent, maintainable data products
  • Evaluate new structured and unstructured datasets at the business-capability level, and integrate them into platform data domains and subject areas
  • Lead complex, cross-team data initiatives from requirements through production support, including estimates , work breakdown , sequencing , dependencies , and cost ; surface risks early while balancing durable platform needs with near-term business priorities
  • Advance AI-first engineering by translating business outcomes into specifications, engineering business and architectural context for agents, and directing agents to plan, build, test, and document changes in small, reviewable increments
  • Develop reusable agent workflows and skills for data engineering work including profiling, source-to-target mapping, pipeline and test generation, schema-change analysis, and incident investigation
  • Define which agent actions can be executed autonomously versus require human approval , and specify how agent activity is reviewed and traced
  • Improve AI-readiness by delivering governed, understandable data through Unity Catalog metadata, lineage, business definitions, semantic models, and access controls; connect datasets to Databricks Genie and other AI applications used by investment professionals
  • Define evaluation datasets and acceptance criteria for agents and AI-generated SQL and code
  • Set the testing strategy across platform layers (including performance , stability , and availability ) and review/approve quality metrics before release
  • Build data quality, reconciliation, freshness, observability, and recovery controls into automated testing and CI/CD , with security and policy checks embedded from the start
  • Partner with investment professionals and product managers to drive shared product vision and ownership of business outcomes, demonstrating how data and AI support portfolio construction and research at scale
  • Raise the engineering bar through design and code reviews, direct day-to-day engineering execution for initiatives, and guide through the most complex data and performance challenges
  • Teach engineers to inspect and challenge AI-generated work; share reusable patterns and context through internal and external forums; help managers identify strengths and development needs

Requirements

  • 10+ years of experience in data or software engineering, including technical leadership of complex production data platforms delivered across multiple teams
  • Strong hands-on Python and SQL skills, sound software design judgment, and deep understanding of distributed data processing , query performance , and automated testing
  • Production experience with Databricks on AWS , including PySpark , Delta Lake , Unity Catalog , Databricks Jobs , Databricks SQL , and Databricks Asset Bundles , including security, access, and cost implications
  • Experience orchestrating production pipelines with Apache Airflow (including Astronomer ) and building tested transformations with dbt , with reliable retries, backfills, and dependency management
  • Strong data modeling and governance experience, including dimensional and time-series models, slowly changing dimensions, bi-temporal history, data contracts, lineage, and semantic metadata
  • Implemented data quality, observability, and CI/CD for data platforms using tools such as Deequ , dbt tests , Lakehouse Monitoring , Datadog , Terraform , and Harness
  • Experience preparing governed data for AI using natural-language-to-SQL tools such as Databricks Genie , semantic metadata, or other governed data-access patterns
  • Use AI coding agents beyond code completion: write specifications, supply context, run tests, and review generated changes via source control
  • Evaluate AI-generated output with representative test cases, regression tests, execution traces, and human review; distinguish plausible output from verified results
  • Understand prompt injection, sensitive data handling, and least-privilege access; design approval boundaries and audit trails for agents operating against enterprise systems
  • Lead architecture discussions, influence without formal authority, develop other engineers, and clearly explain technical choices and trade-offs to investment professionals and technology leaders
  • Act as an agent of change with urgency: question how work gets done and remove or automate low-value steps while respecting existing implementations
  • Bachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience

Technologies

  • Databricks
  • AWS
  • Python
  • SQL
  • PySpark
  • Delta Lake
  • Unity Catalog
  • Databricks Jobs
  • Databricks SQL
  • Databricks Asset Bundles
  • Apache Airflow
  • Astronomer
  • dbt
  • Airflow
  • Deequ
  • Lakehouse Monitoring
  • Datadog
  • Terraform
  • Harness
  • Databricks Genie
  • CI/CD

Preferred Qualifications

  • Experience with investment management data such as portfolios, positions, returns, exposures, benchmarks, and attribution, or with multi-asset portfolio construction
  • Experience building LLM applications or agent workflows that call tools and APIs, including context management, retrieval, state, and error handling
  • Familiarity with Model Context Protocol (MCP)
  • Experience with PostgreSQL, SQL Server, or Lakebase, or modernizing legacy data platforms onto a Lakehouse

Benefits

  • Generous time-away and health benefits from day one, with opportunity for flexible work options
  • 2-for-1 matching gifts for charitable contributions
  • Opportunity to secure annual grants for organizations you love
  • Access on-demand professional development resources
  • Competitive salary, bonuses and benefits
  • Company-funded retirement contribution
  • Individual annual performance bonus
  • Capital’s annual profitability bonus
  • Retirement plan where Capital contributes 15% of eligible earnings

Location and Salary

  • Location: Los Angeles, CA (onsite)
  • Salary: USD 201,683 - 342,072 per yearly
  • Southern California base salary range: $201,683-$322,693
  • New York base salary range: $213,795-$342,072
#J-18808-Ljbffr

Data Engineer Lead in los angeles at Unknown Company

This position is listed as contract and onsite.

Back to Job Search