Master’s degree in Data Science, Statistics, or a related field
Sound knowledge of feature engineering/model evaluation/validation & on-chain patterns/risk-analysis/threat-detection methodologies
In-depth understanding of blockchain/distributed ledger data structures & analytics
Strong ability to apply machine-learning & statistical modeling techniques to large-scale datasets
Expertise in analyzing graph/text-based or transactional data
Familiar with cloud platforms (AWS/Azure/GCP) & Spark-based distributed-computing systems (e.g., Databricks)
Proficient in Python, SQL (PostgreSQL/MySQL/NoSQL) & ETL tools (Apache Airflow)
What the job involves
The primary responsibility of this role is to build/maintain ETL pipelines & process large datasets from APIs/databases/third-party platforms to enable real-time team analytics and automate data preprocessing (cleaning/normalization/validation) for client accounts using rule-based logic/statistical checks to ensure data quality & prepare analysis-ready datasets for modeling/reporting
Analyze large-scale blockchain/transactional/social-media datasets to identify patterns/trends/anomalies/risk indicators