Unknown Company

AI Data Engineer

new york, ny • Posted 4 days ago
Onsite Full Time General

AI Data EngineerNew York, New York, United StatesSchonfeld Strategic Advisors is seeking an experienced AI Data Engineer to join our Data Engineering team. In this role, you will be responsible for designing, building, and maintaining robust data pipelines that power SchonAI, our firm's internal AI platform. You will work at the intersection of data engineering and AI, ensuring that high-quality, timely, and relevant data flows seamlessly to our AI systems to support investment professionals across the firm.Key ResponsibilitiesData Pipeline DevelopmentDesign and build scalable, reliable data pipelines to ingest, transform, and deliver structured and unstructured data to SchonAI using Prefect.Develop ETL/ELT processes for diverse data sources including market data, research documents, internal databases, and third-party APIs.Implement real-time and batch data processing workflows to meet varying latency requirements.Ensure data quality, consistency, and integrity across all pipelines.AI Data InfrastructureBuild and maintain data infrastructure optimized for AI/ML workloads, including vector databases and semantic search systems.Design data schemas and storage solutions that support efficient retrieval and processing for LLM applications.Implement data versioning, lineage tracking, and observability for AI training and inference pipelines.Optimize data delivery for low-latency AI interactions and high-throughput batch processing.Integration & CollaborationPartner with AI engineers, software developers, and data scientists to understand data requirements.Integrate with existing firm systems including risk platforms, trading systems, portfolio management tools, and research databases.Collaborate with infrastructure teams on cloud architecture, security, and compliance requirements.Work closely with business stakeholders to prioritize data sources and pipeline enhancements.Data Governance & SecurityImplement appropriate data access controls, encryption, and compliance measures.Ensure adherence to data governance policies and regulatory requirements.Monitor and maintain data pipeline performance, reliability, and cost efficiency.Document data flows, transformations, and dependencies.Required QualificationsTechnical SkillsStrong proficiency in Python; experience with SQL and at least one other language (e.g.

Java, Scala, Go, Rust)5+ years of experience building production data pipelines using tools like Apache Airflow, Prefect, Dagster, or similarHands-on experience with distributed computing frameworks (Spark, Flink) and modern data platformsProficiency with AWS services (S3, Kubernetes) or equivalent GCP servicesExperience with both SQL (PostgreSQL, MySQL) and NoSQL databases (MongoDB, DynamoDB, Elasticsearch)Understanding of data requirements for ML/AI systems, including experience with vector databases (Pinecone, Weaviate, Qdrant) and embedding pipelinesProfessional SkillsBachelor's or Master's degree in Computer Science, Data Engineering, or related technical fieldStrong problem-solving skills and attention to detailExcellent communication skills with ability to translate technical concepts for non-technical stakeholdersExperience working in fast-paced, collaborative environmentsSelf-motivated with ability to manage multiple prioritiesThe base pay for this role is expected to be between $225k and $275k. This role may also be eligible for other forms of compensation such as a performance bonus and a competitive benefits package. Actual compensation for the successful candidate will be determined based on a variety of factors such as skills, qualifications, and experience.

Back to Job Search