- Design, implement, ship, and operate scalable model-serving infrastructure for machine learning, scientific, LLM, and agentic workloads.
- Help evolve our internal model deployment platform into a reliable, self-service platform for teams across the organization.
- Improve platform scalability and reliability, including scale-to-zero, faster model startup, workload isolation, traffic management, and reduction of request failures and latency bottlenecks.
- Build observability and operational tooling for model usage, latency, reliability, resource consumption, inference cost, bottlenecks, and service-level indicators.
- Improve the usability of model deployment by developing validated configuration interfaces, reusable deployment patterns, APIs, command-line tools, and documentation.
- Help converge real-time and batch inference workflows onto shared platform capabilities where appropriate.
- Contribute to model lifecycle management infrastructure, including model registration and versioning, evaluation, promotion and release gates, monitoring, environment progression, and rollback.
- Build event-driven integrations that connect model publication, evaluation, promotion, deployment, and retraining workflows.
- Build consistent metrics and evaluation signals for understanding model cost, quality, reliability, and fitness for downstream workflows.
- Partner with machine learning, data, scientific, and platform teams to translate requirements into maintainable solutions and remove infrastructure bottlenecks.
- Own workstreams from design through implementation and production support, using strong software-engineering practices including testing, reviews, documentation, and incremental delivery.
Requirements
- BS or MS in Computer Science, Engineering, or a related technical field, or equivalent practical experience.
- 3+ years of relevant industry experience in software engineering, infrastructure engineering, platform engineering, DevOps, MLOps, or a related area.
- Strong Python programming skills and experience building and shipping maintainable production software, services, automation, or developer tooling.
- Experience designing, deploying, or operating cloud systems (preferably on AWS) using services such as EKS, EC2, S3, IAM, SQS, SNS, and CloudWatch.
- Experience with containers, Kubernetes, Helm, and IaC tools such as Terraform or Pulumi.
- Experience with CI/CD, Git-based development workflows, automated testing, and software release practices.
- Ability to troubleshoot complex systems using metrics, logs, traces, events, and observability tools such as Datadog, Prometheus, Grafana, or OpenTelemetry.
- Understanding of distributed-systems concepts such as concurrency, queuing, retries, timeouts, idempotency, backpressure, and failure recovery.
- Ability to gather requirements, communicate technical tradeoffs, and document systems for users and engineers with varied infrastructure experience.
- Demonstrated ability to independently deliver practical, incremental solutions while considering immediate needs and longer-term platform direction.
- Familiarity with model-serving or workflow-orchestration frameworks such as KServe, Triton, vLLM, Ray Serve, Prefect, or Dagster.
- Experience optimizing model startup time, request throughput, batching, autoscaling, or GPU utilization.
- Familiarity with model registries, experiment tracking, model evaluation, promotion workflows, or MLOps platforms.
- Experience building event-driven systems using queues, event buses, or workflow orchestrators.
- Familiarity with online and offline model evaluation, model-quality monitoring, data drift, or regression analysis.
- Experience supporting scientific computing, high-performance computing, distributed training, or large-scale data processing.
- Strong interest in the life sciences and drug discovery.
Core Competencies
Demonstrates expertise in building and operating scalable model-serving infrastructure, with strong proficiency in Python programming and cloud systems, particularly on AWS. Capable of improving platform reliability and usability through effective model lifecycle management and observability practices.
#J-18808-LjbffrMachine Learning Engineer, Infra, AI for Drug Discovery in california at Unknown Company
This position is listed as full time and onsite.