Unknown Company

Senior Software Engineer, Data Platform

emeryville, ca • Posted 4 days ago
Hybrid Full Time General

Senior Software Engineer, Data PlatformEmeryville, California, United States; Hybrid (2-3 days on-site)Profluent is an AI-first protein design company. Founded in 2022, we develop deep generative models to design and validate novel, functional proteins to revolutionize biomedicine. Based in Emeryville, CA, we are backed by leading investors including Altimeter Capital, Bezos Expeditions, Spark Capital, Insight Partners, Air Street Capital, AIX Ventures, and Convergent Ventures, and have raised over $150M to date.We're looking for a Senior Software Engineer to help design, build, and scale Profluent's data platform.

This platform houses data from protein engineering campaigns, including protein designs, experimental results, partner datasets, analytical outputs, and model-ready training data. It enables rapid machine learning, biological discovery, and secure collaboration across internal and external programs.This role is ideal for an engineer who enjoys building robust data systems: secure ingestion pipelines, well-structured warehouses, reliable data models, access controls, auditability, and infrastructure that makes complex scientific data usable at scale. You will work closely with ML, bioinformatics, and program teams to ensure Profluent's data is organized, governed, accessible, and protected.ResponsibilitiesDesign, build, and maintain scalable data infrastructure for protein engineering campaigns, including ingestion, transformation, validation, storage, and retrieval of large scientific datasetsDevelop secure data pipelines for internal and partner-generated data, with strong attention to access control, data siloing, provenance, auditability, and compliance with data use restrictionsOwn core components of Profluent's data warehouse and data platform, using Python, GCP, PostgreSQL, BigQuery, and related cloud-native technologiesBuild systems that transform raw experimental, computational, and partner data into structured, reliable, analysis-ready and model-ready datasetsEstablish best practices for data modeling, metadata management, data quality checks, schema evolution, versioning, and documentationCollaborate with ML engineers, computational biologists, data scientists, and program stakeholders to understand data requirements and translate them into scalable technical systemsImprove engineering quality through thoughtful system design, code review, testing, CI/CD, observability, and maintainable development workflowsContribute to architectural decisions for how Profluent stores, secures, organizes, and uses data across programs and partnershipsQualifications5+ years of software engineering, data engineering, or data platform experienceStrong proficiency in Python and modern software development practices, including git, testing, code review, CI/CD, and production deploymentExperience designing and operating production data pipelines, data warehouses, and data models at scaleHands-on experience with cloud platforms, preferably GCP, and technologies such as BigQuery, PostgreSQL, object storage, workflow orchestration, and containerized servicesStrong understanding of data security, access control, data partitioning or siloing, audit logging, and managing sensitive or restricted datasetsExperience working with complex, heterogeneous datasets and building systems that make them reliable, discoverable, and usableAbility to work independently, make sound technical decisions, and drive projects from ambiguous requirements to production systemsBS, MS, or PhD in Computer Science, Engineering, Data Science, Bioinformatics, or a related technical field, or equivalent practical experiencePreferencesExperience with scientific, biological, clinical, genomic, laboratory, or high-throughput experimental dataExperience managing external partner, customer, or restricted-access datasetsFamiliarity with data governance, lineage, metadata systems, schema registries, or data catalogsExperience with research data systems, LIMS, ELNs, Benchling, or adjacent scientific platformsBackground working with ML, data science, computational biology, or cross-disciplinary technical teamsInterest in learning biology, gene editing, protein design, or machine learning conceptsWhat We OfferHigh-growth opportunity with meaningful impact on the future of protein designCompetitive compensation package with equity participation401(k) with a strong employer matchComprehensive benefits including health/dental/vision insuranceGenerous PTO policy and commitment to work-life balanceProfessional development opportunities in a cutting-edge field at the intersection of AI and biologyProfluent Bio, Inc is an equal opportunity employer promoting diversity and inclusion in the workspace.

We do not discriminate on the basis of race, color, religion, marital status, age, national origin, ancestry, physical or mental disability, medical conditions, veteran status, sexual orientation, gender (including gender identity and gender expression), sex (which includes pregnancy, childbirth, and breastfeeding), genetic information, taking or requesting statutorily protected leave, or any other basis protected by law.Applicants must have ongoing work authorization in the United States that does not require employer sponsorship. Sponsorship will not be provided now or at any time in the future for this position.Legal authorization to work in the United States is required. In compliance with federal law, all persons hired must verify their identity and work eligibility and complete the required employment verification form upon hire.Hiring Salary Range $170,000 - $220,000 USD

Back to Job Search