Own the design, build and operation of the data lake and ingestion platform end-to-end, from architecture through production reliability.
Build low-latency batch and streaming pipelines that ingest signals from internal and external sources, normalize them to a common schema, enrich them with context and serve model-ready data to the layers above.
Make adding a new data source a routine task rather than a project, so our view of risk keeps widening over time.
Establish data quality, freshness, completeness, lineage and observability so the platform is trustworthy enough to automate on top of.
Build data pipelines that ground generative AI, including unstructured text and threat intelligence processing, embedding generation, vector storage and retrieval.
Own deployment, CI/CD and operational reliability of the platform on Kubernetes.
Partner with data science, product and architecture to turn the platform into a shared foundation across 360 Fraud Protection.
Requirements
Extensive experience building and operating large-scale data platforms and data lakes, with comfort working at high data volumes.
Deep, hands‑on expertise with Apache Spark, Apache Flink and modern big‑data systems.
Proven command of best practices for building and maintaining data pipelines in both batch and streaming modes.
Strong production engineering skills across the full delivery lifecycle, including Kubernetes and CI/CD tooling, with the ability to ship end‑to‑end.
A track record of owning data infrastructure end‑to‑end with limited supervision.