Unknown Company

Site Reliability Engineer

california, mo • Posted 3 days ago
Remote Full Time Engineering

SENIOR SITE RELIABILITY ENGINEER | FULLY REMOTE (US)

Location: Fully remote - US-based, working a US timezone

Package: $170,000 - $220,000

Overvie w

We're supporting a specialist AI infrastructure company - a UK sovereign AI cloud powered by renewable energy - that builds and operates large-scale compute on regenerated industrial and energy sites.

Having recently secured a major US customer and taken on their entire cluster, the business is standing up a US-based operations team and is looking for Senior Site Reliability Engineers to build that capability from the groundup.

This is a Platform/SRE role with a strong automation and software-engineering bias — not an AI model-building role. You'll turn runbooks, alerts and operational workflows into safe, auditable automation that improves reliability across the platform.

  • Get in at the ground floor of a brand-new US operations function and shape how it runs at scale
  • Work on critical infrastructure powering the next generation of AI
  • Automation-first culture - reduce toil and build tooling, rather than fire fight
  • Real autonomy and high visibility with leadership
  • Fully remote, on a US timezone (West-coast preferred)

What you'll be doing

  • Building Python-based automation for incident triage, runbook execution and routine operational tasks
  • Integrating observability, ITSM and infrastructure APIs to enrich alerts and automate workflows
  • Improving monitoring signal quality through correlation, enrichment, suppression and deduplication
  • Building internal tools and self-service capabilities - CLI utilities, ChatOps integrations and dashboards
  • Maintaining version-controlled runbook-as-code and automation libraries
  • Turning post-incident learnings into better tooling, automation and operational standards

We're keen to speak with candidates who have

  • Experience in SRE, Platform Engineering or production infrastructure operations
  • Hands-on experience with observability/monitoring tooling (Prometheus, Grafana or similar)
  • Exposure to incident management / on-call, and converting manual runbooks into automation
  • Strong Python for automation, APIs and integrations

Nice to have:

  • GPU, datacentre or colocation infrastructure experience
  • ITSM integrations (ServiceNow, Halo, Jira Service Management or similar)
  • ChatOps tooling (Slack or Microsoft Teams bots)
  • OpenTelemetry, logging or distributed tracing experience
  • DCIM, IPAM or hypervisor-control-plane integrations
  • Experience with LLM-assisted or agent-based operational automation

#J-18808-Ljbffr
Back to Job Search