Unknown Company

Machine Learning Engineer, Infra, AI for Drug Discovery

california, mo • Posted 1 weeks ago
Onsite Full Time Electrical & Energy Engineering

  • Design, implement, ship, and operate scalable model-serving infrastructure for machine learning, scientific, LLM, and agentic workloads.
  • Help evolve our internal model deployment platform into a reliable, self-service platform for teams across the organization.
  • Improve platform scalability and reliability, including scale-to-zero, faster model startup, workload isolation, traffic management, and reduction of request failures and latency bottlenecks.
  • Build observability and operational tooling for model usage, latency, reliability, resource consumption, inference cost, bottlenecks, and service-level indicators.
  • Improve the usability of model deployment by developing validated configuration interfaces, reusable deployment patterns, APIs, command-line tools, and documentation.
  • Help converge real-time and batch inference workflows onto shared platform capabilities where appropriate.
  • Contribute to model lifecycle management infrastructure, including model registration and versioning, evaluation, promotion and release gates, monitoring, environment progression, and rollback.
  • Build event-driven integrations that connect model publication, evaluation, promotion, deployment, and retraining workflows.
  • Build consistent metrics and evaluation signals for understanding model cost, quality, reliability, and fitness for downstream workflows.
  • Partner with machine learning, data, scientific, and platform teams to translate requirements into maintainable solutions and remove infrastructure bottlenecks.
  • Own workstreams from design through implementation and production support, using strong software-engineering practices including testing, reviews, documentation, and incremental delivery.

Requirements

  • BS or MS in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
  • 3+ years of relevant industry experience in software engineering, infrastructure engineering, platform engineering, DevOps, MLOps, or a related area.
  • Strong Python programming skills and experience building and shipping maintainable production software, services, automation, or developer tooling.
  • Experience designing, deploying, or operating cloud systems (preferably on AWS) using services such as EKS, EC2, S3, IAM, SQS, SNS, and CloudWatch.
  • Experience with containers, Kubernetes, Helm, and IaC tools such as Terraform or Pulumi.
  • Experience with CI/CD, Git-based development workflows, automated testing, and software release practices.
  • Ability to troubleshoot complex systems using metrics, logs, traces, events, and observability tools such as Datadog, Prometheus, Grafana, or OpenTelemetry.
  • Understanding of distributed-systems concepts such as concurrency, queuing, retries, timeouts, idempotency, backpressure, and failure recovery.
  • Ability to gather requirements, communicate technical tradeoffs, and document systems for users and engineers with varied infrastructure experience.
  • Demonstrated ability to independently deliver practical, incremental solutions while considering immediate needs and longer-term platform direction.
  • Familiarity with model-serving or workflow-orchestration frameworks such as KServe, Triton, vLLM, Ray Serve, Prefect, or Dagster.
  • Experience optimizing model startup time, request throughput, batching, autoscaling, or GPU utilization.
  • Familiarity with model registries, experiment tracking, model evaluation, promotion workflows, or MLOps platforms.
  • Experience building event-driven systems using queues, event buses, or workflow orchestrators.
  • Familiarity with online and offline model evaluation, model-quality monitoring, data drift, or regression analysis.
  • Experience supporting scientific computing, high-performance computing, distributed training, or large-scale data processing.
  • Strong interest in the life sciences and drug discovery.

Core Competencies

Demonstrates expertise in building and operating scalable model-serving infrastructure, with strong proficiency in Python programming and cloud systems, particularly on AWS. Capable of improving platform reliability and usability through effective model lifecycle management and observability practices.

#J-18808-Ljbffr

Machine Learning Engineer, Infra, AI for Drug Discovery in california at Unknown Company

This position is listed as full time and onsite.

Back to Job Search