Data Architect (Open Source)Location: Pittsburgh, PADuration: FulltimeJob Description:Skills Desired:Open-Source Technologies: Hadoop, Spark, Kafka, Airflow, dbt, Talend.Cloud Platforms: AWS, Azure, or GCP with hands-on experience in data services.Data Modeling: Dimensional, Data Vault, and normalization techniques.Programming: Python, Scala, SQL.Architecture Patterns: Microservices, event-driven, and distributed systems.Strong understanding of data security, lineage, and governance frameworks.Experience with containerization (Docker, Kubernetes).Familiarity with MLOps and AI/ML integration.Knowledge of CI/CD pipelines for data solutions.BFSI domain experience (or relevant industry exposure).Roles & Responsibilities:Define and maintain data product architecture frameworks using open-source tools.Design data pipelines, ingestion workflows, and storage layers for structured/unstructured data.Implement data mesh principles and domain-driven design for data products.Collaborate with engineering teams to deliver cloud-native and hybrid architectures.Establish data governance, lineage, and security standards for all data products.Optimize performance for big data platforms and ensure cost efficiency.Drive AI/ML enablement within data products for predictive and real-time insights. Knowledge of Spark, PySpark, Scala.Experience leading CoEs or data science accelerators.Open-source contributions or published research.If you are interested, please send me your updated resume ASAP with below details: