Unknown Company

Principal Software Engineer, AI Platform Engineering

milpitas, ca • Posted 4 days ago
Onsite Contract General

Job TitleYou set the architectural direction for how training data flows, evolves, and is governed across the AI Platform. You define the standards ML engineers and scientists build on, and ensure every training signal is tenant-isolated, PII-free, and traceable from source to model.Responsibilities include:AI Data Lake on GCS: bucket layout, raw ? silver ?

gold tier separation, CMEK encryption, lifecycle rulesBatch pipelines: Spark on Dataproc for TB-scale feature backfills, Iceberg compaction, and daily S3?GCS incremental syncStreaming pipelines: Apache Beam on Dataflow for sub-5-min CDC ingestion with exactly-once semantics and PII assertion gatesSchema registry: Avro / Protobuf schema versioning, compatibility modes, and migration playbooks for safe schema evolutionOrchestration: Flyte as primary DAG layer — task authoring standards, domain isolation, retry policies, DataCatalog memoization; evaluate Kubeflow Pipelines where relevantMulti-tenancy: strict per-tenant GCS prefix isolation, quota policies, and cross-tenant contamination validationData Anonymizer and Data Labeler microservices: strip PII and attach ML labels before signals leave each customer environmentFeature store: Feast offline (GCS Parquet) and online (Redis) with point-in-time correctness and < 0.1% consistency SLAVector database: operate Pgvector (Cloud SQL) for POC and Qdrant on GKE for production-scale embedding storage; design index strategies (IVFFlat, HNSW) and manage ANN query latency SLAsRAG data pipeline: build embedding generation pipelines that chunk, encode, and upsert document embeddings into the vector store; own the data refresh cadence and staleness SLAs for retrieval contextService APIs: expose data platform services (feature serving, embedding upsert, schema validation) over HTTPS with mTLS and gRPC where low-latency streaming is requiredSynthetic data pipelines for dev/staging where real customer data is not permittedData quality gates: Great Expectations / dbt checks as Flyte tasks, blocking on schema and PII-absence failuresYou'll thrive here if you have:1+ years of experience as a Principal SWE at a SaaS companyDemonstrated principal impact: platform standards you defined adopted org-wide, or major cross-team pipeline/schema migrations you ledData lake ownership (essential): you have designed and operated a production data lake end-to-end — storage layout, partitioning strategy, tiered retention (hot/warm/cold), table format (Iceberg or Delta Lake), compaction, and access control; not just consumed oneDeep Spark (PySpark / Scala): executor tuning, shuffle diagnosis, Iceberg table maintenanceHands-on Beam / Dataflow: windowing, exactly-once, side inputs, autoscalingSchema registry experience: Protobuf / Avro compatibility rules, breaking-change migrations in productionOrchestration at scale: Flyte, Kubeflow Pipelines, Airflow, or Prefect — operated in production, ideally benchmarked twoMulti-tenant data architecture: per-tenant isolation as a hard requirement, not a post-hoc concernFeature store operations: Feast or Tecton, point-in-time joins, online/offline consistencyVector databases: Pgvector or Qdrant in production — index tuning, ANN search, embedding upsert pipelinesRAG data fundamentals: chunking strategies, embedding model selection, retrieval quality evaluation, and context freshness managementAPI transport: gRPC and HTTPS/mTLS for service-to-service communication; comfortable defining proto contracts and managing certificate lifecycleBachelor's degree in Computer Science, Engineering, or a related field, or equivalent practical experience or equivalent military experienceNice to have:Differential privacy or k-anonymity for ML training datasetsOpen source contributions: Feast, Great Expectations, Apache Beam, or dbtFamiliarity with IAM / access governance data: entitlements, provisioning events, access graphsIceberg or Delta Lake at petabyte scaleWhy join Saviynt:Work on a large-scale, Kubernetes-based SaaS platformSolve challenging cloud and reliability problems at scaleCollaborate with strong engineers in a reliability-focused cultureCompetitive compensation, benefits, and growth opportunitiesSecurity & Compliance:This role requires adherence to Saviynt's information security and privacy policies, including annual security training.$274,000 - $304,000 a year We offer you a competitive total rewards package, learning and tremendous opportunities to grow and advance in your career. At Saviynt, it is not typical for an individual to be hired at or near the top of the range for their role and final compensation decisions are dependent on many factors including but are not limited to location; skill sets; experience and training; licensure and certifications; and other relevant business and organizational needs. A reasonable estimate of the current range is $240,000 - $260,000 annually.

Back to Job Search