Unknown Company

ML Data Infrastructure Engineer

sunnyvale, ca • Posted 4 days ago
Onsite Full Time General

Data EngineerKey Responsibilities:Design and implement scalable data processing pipelines for ML training and validationBuild and maintain feature stores with support for both batch and real-time featuresDevelop data quality monitoring, validation, and testing frameworksCreate systems for dataset versioning, lineage tracking, and reproducibilityImplement automated data documentation and discovery toolsDesign efficient data storage and access patterns for ML workloadsPartner with data scientists to optimize data preparation workflowsTechnical Requirements:7+ years of software engineering experience, with 3+ years in data infrastructureStrong expertise in GCP's data and ML infrastructure:BigQuery for data warehousingDataflow for data processingCloud Storage for data lakesVertex AI Feature StoreCloud Composer (managed Airflow)Dataproc for Spark workloadsDeep expertise in data processing frameworks (Spark, Beam, Flink)Experience with feature stores (Feast, Tecton) and data versioning toolsProficiency in Python and SQLExperience with data quality and testing frameworksKnowledge of data pipeline orchestration (Airflow, Dagster)Nice to Have:Experience with streaming systems (Kafka, Kinesis)Experience with GCP-specific security and IAM best practicesKnowledge of Cloud Logging and Cloud Monitoring for data pipelinesFamiliarity with Cloud Build and Cloud Deploy for CI/CDExperience with streaming systems (Pub/Sub, Dataflow)Knowledge of ML metadata management systemsFamiliarity with data governance and security requirementsExperience with dbt or similar data transformation tools

Back to Job Search