Data EngineerThe data engineer on the Nebula team plays a critical role in building and evolving the data foundation that powers analytics, reporting, AI development, and operational decision-making across the organization. This role is responsible for designing, building, and maintaining reliable, scalable, and flexible data systems that support a wide range of internal and external use cases.Working across data ingestion, transformation, storage, modeling, and delivery, this individual partners closely with product, engineering, AI, analytics, and domain subject matter experts (smes) to translate complex business processes and data needs into production-ready data pipelines and platforms.This role contributes to the development and evolution of core data capabilities, including batch and real-time pipelines, operational and analytical data stores, semantic models, and bi-ready datasets. Success requires strong technical depth across modern data tooling, sound systems thinking, and the ability to build reliable solutions in a cloud-based, regulated, high-stakes environment.The data engineer is expected to operate effectively in a modern engineering environment, using automation, observability, and infrastructure-as-code practices to deploy, manage, and improve data pipelines and data platforms. In parallel, this individual will help enable downstream analytics, reporting, product capabilities, and ai systems by ensuring that data is trustworthy, accessible, and fit for purpose.This is a fully remote position that offers a competitive salary range of $140,000 to $190,000 usd, plus an annual bonus. You'll also receive our excellent benefits package, which includes medical coverage starting on day one and a company-matched 401(k).
Compensation may vary based on experience, location, and other job-related factors.ResponsibilitiesDesign, build, and maintain robust data pipelines for a wide variety of input and output sources, including internal systems, third-party platforms, files, apis, event streams, and databasesDevelop scalable etl and elt workflows for both batch and real-time processingEnsure pipelines are reliable, testable, observable, and easy to extend as business needs evolveBuild reusable data integration patterns that support growing volumes, new source systems, and downstream consumers across analytics, applications, and ai initiativesDesign and manage data architectures that support oltp, olap, and reporting workloads across operational and analytical environmentsBuild and optimize data models, warehouse schemas, and curated datasets for analytics and bi use casesContribute to the design and operation of modern data platforms, including warehouses, lakehouses, streaming systems, and supporting orchestration frameworksHelp define patterns for data storage, partitioning, performance optimization, retention, and lifecycle managementDeploy, operate, and improve data pipelines and data stores on major cloud platforms such as aws, gcp, or azureUse infrastructure-as-code, ci/cd, and automation practices to improve deployment speed, consistency, and reliabilityMonitor production data systems using logging, alerting, and observability tooling to proactively identify and resolve issuesSupport secure, resilient, and cost-conscious operation of cloud-based data infrastructureImplement data quality checks, validation rules, reconciliation processes, and monitoring to ensure trustworthy data across systemsEstablish and maintain standards for lineage, documentation, metadata, schema evolution, and operational runbooksPartner with stakeholders to improve data accessibility, consistency, and usability while maintaining appropriate controls and governanceContribute to practices that support security, privacy, auditability, and compliance in a regulated environmentPartner closely with product, engineering, and business stakeholders to understand data needs, workflows, and constraintsTranslate business and operational requirements into clean, scalable, and maintainable data solutionsSupport downstream consumers of data, including analysts, researchers, product teams, and operational usersCommunicate clearly with both technical and non-technical stakeholders about data availability, quality, tradeoffs, and delivery timelinesContinuously improve pipeline performance, reliability, scalability, and developer productivityIdentify opportunities to simplify architecture, reduce operational toil, and improve data platform leverage across teamsOperate with a strong bias toward action and iterative delivery, moving quickly from problem definition to implementation and improvementHelp raise the bar on engineering quality through thoughtful design, testing, documentation, and operational disciplineQualifications2-4+ years of experience building and operating production-grade data pipelines and data systemsStrong experience with industry-standard tools and platforms for etl/elt, orchestration, data warehousing, streaming, and biExperience working with both oltp and olap systems, with a strong understanding of the tradeoffs between transactional and analytical workloadsExperience building flexible data pipelines that integrate with many different source and destination types, including databases, apis, files, message queues, saas platforms, and event streamsExperience supporting both batch and real-time data processing patternsExperience deploying and operating data infrastructure on major cloud platforms such as aws, gcp, or azureStrong sql skills and experience with data modeling, transformation frameworks, and performance optimizationExperience building ai-powered capabilities on top of llms, including orchestration, evaluation, and data integration patternsExperience with modern programming languages commonly used in data engineering, such as python, java, scala, or goComfort working with ci/cd, infrastructure-as-code, observability, and production operations for data systemsStrong judgment in ambiguous environments where requirements evolve and systems must balance speed, reliability, and flexibilityClear communication skills with both technical and non-technical teammatesPreferred ExperienceExperience with modern orchestration and transformation tools such as airflow, dagster, dbt, or similar platformsExperience with cloud-native data warehouses or lakehouse platforms such as snowflake, bigquery, redshift, databricks, or equivalent technologiesExperience with streaming and real-time data platforms such as kafka, kinesis, sqs, or similar systemsExperience enabling bi and self-service analytics through curated datasets, semantic layers, and reporting platforms such as looker, power bi, tableau, or similar toolsExperience in fintech, mortgage, lending, payments, insurance, or other regulated domainsExperience building data platforms that support ai, machine learning, or decisioning workflowsExperience improving data quality, reliability, cost efficiency, and platform scalability as a system growsA note to candidatesYou do not need prior fintech or finance experience to succeed in this role. If you are a strong data engineer with solid technical judgment, a systems mindset, and excitement for solving complex data problems, we would love to hear from you.If your background does not line up perfectly with every bullet, but this role feels like the kind of work you want to do, please apply.Bayview is an equal employment opportunity employer. All aspects of consideration for employment and employment with the company are governed on the basis of merit, competence, and qualifications without regard to race, color, religion, sex, national origin, age, disability, veteran status, sexual orientation, or any other category protected by federal, state, or local law.#li-remote