Capital Group Companies is hiring a senior, hands-on Data Engineering Lead to set technical direction for the CSGT data platform and data products in Los Angeles.
Responsibilities
- Own data engineering strategy and roadmap for CSGT, including Lakehouse architecture on Databricks and AWS ; make and explain decisions on scalability , security , reliability , and cost
- Shape standards and practices across adjacent teams and the broader Capital Group technology organization, contributing to centers of excellence and firm-wide engineering practices
- Design and build ingestion, transformation, and serving pipelines using Databricks , PySpark , Delta Lake , dbt , and Airflow ; establish reusable patterns and frameworks for consistent, maintainable data products
- Evaluate new structured and unstructured datasets at the business-capability level, and integrate them into platform data domains and subject areas
- Lead complex, cross-team data initiatives from requirements through production support, including estimates , work breakdown , sequencing , dependencies , and cost ; surface risks early while balancing durable platform needs with near-term business priorities
- Advance AI-first engineering by translating business outcomes into specifications, engineering business and architectural context for agents, and directing agents to plan, build, test, and document changes in small, reviewable increments
- Develop reusable agent workflows and skills for data engineering work including profiling, source-to-target mapping, pipeline and test generation, schema-change analysis, and incident investigation
- Define which agent actions can be executed autonomously versus require human approval , and specify how agent activity is reviewed and traced
- Improve AI-readiness by delivering governed, understandable data through Unity Catalog metadata, lineage, business definitions, semantic models, and access controls; connect datasets to Databricks Genie and other AI applications used by investment professionals
- Define evaluation datasets and acceptance criteria for agents and AI-generated SQL and code
- Set the testing strategy across platform layers (including performance , stability , and availability ) and review/approve quality metrics before release
- Build data quality, reconciliation, freshness, observability, and recovery controls into automated testing and CI/CD , with security and policy checks embedded from the start
- Partner with investment professionals and product managers to drive shared product vision and ownership of business outcomes, demonstrating how data and AI support portfolio construction and research at scale
- Raise the engineering bar through design and code reviews, direct day-to-day engineering execution for initiatives, and guide through the most complex data and performance challenges
- Teach engineers to inspect and challenge AI-generated work; share reusable patterns and context through internal and external forums; help managers identify strengths and development needs
Requirements
- 10+ years of experience in data or software engineering, including technical leadership of complex production data platforms delivered across multiple teams
- Strong hands-on Python and SQL skills, sound software design judgment, and deep understanding of distributed data processing , query performance , and automated testing
- Production experience with Databricks on AWS , including PySpark , Delta Lake , Unity Catalog , Databricks Jobs , Databricks SQL , and Databricks Asset Bundles , including security, access, and cost implications
- Experience orchestrating production pipelines with Apache Airflow (including Astronomer ) and building tested transformations with dbt , with reliable retries, backfills, and dependency management
- Strong data modeling and governance experience, including dimensional and time-series models, slowly changing dimensions, bi-temporal history, data contracts, lineage, and semantic metadata
- Implemented data quality, observability, and CI/CD for data platforms using tools such as Deequ , dbt tests , Lakehouse Monitoring , Datadog , Terraform , and Harness
- Experience preparing governed data for AI using natural-language-to-SQL tools such as Databricks Genie , semantic metadata, or other governed data-access patterns
- Use AI coding agents beyond code completion: write specifications, supply context, run tests, and review generated changes via source control
- Evaluate AI-generated output with representative test cases, regression tests, execution traces, and human review; distinguish plausible output from verified results
- Understand prompt injection, sensitive data handling, and least-privilege access; design approval boundaries and audit trails for agents operating against enterprise systems
- Lead architecture discussions, influence without formal authority, develop other engineers, and clearly explain technical choices and trade-offs to investment professionals and technology leaders
- Act as an agent of change with urgency: question how work gets done and remove or automate low-value steps while respecting existing implementations
- Bachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience
Technologies
- Databricks
- AWS
- Python
- SQL
- PySpark
- Delta Lake
- Unity Catalog
- Databricks Jobs
- Databricks SQL
- Databricks Asset Bundles
- Apache Airflow
- Astronomer
- dbt
- Airflow
- Deequ
- Lakehouse Monitoring
- Datadog
- Terraform
- Harness
- Databricks Genie
- CI/CD
Preferred Qualifications
- Experience with investment management data such as portfolios, positions, returns, exposures, benchmarks, and attribution, or with multi-asset portfolio construction
- Experience building LLM applications or agent workflows that call tools and APIs, including context management, retrieval, state, and error handling
- Familiarity with Model Context Protocol (MCP)
- Experience with PostgreSQL, SQL Server, or Lakebase, or modernizing legacy data platforms onto a Lakehouse
Benefits
- Generous time-away and health benefits from day one, with opportunity for flexible work options
- 2-for-1 matching gifts for charitable contributions
- Opportunity to secure annual grants for organizations you love
- Access on-demand professional development resources
- Competitive salary, bonuses and benefits
- Company-funded retirement contribution
- Individual annual performance bonus
- Capital’s annual profitability bonus
- Retirement plan where Capital contributes 15% of eligible earnings
Location and Salary
- Location: Los Angeles, CA (onsite)
- Salary: USD 201,683 - 342,072 per yearly
- Southern California base salary range: $201,683-$322,693
- New York base salary range: $213,795-$342,072
Data Engineer Lead in los angeles at Unknown Company
This position is listed as contract and onsite.