SENIOR SITE RELIABILITY ENGINEER | FULLY REMOTE (US)
Location: Fully remote - US-based, working a US timezone
Package: $170,000 - $220,000
Overvie w
We're supporting a specialist AI infrastructure company - a UK sovereign AI cloud powered by renewable energy - that builds and operates large-scale compute on regenerated industrial and energy sites.
Having recently secured a major US customer and taken on their entire cluster, the business is standing up a US-based operations team and is looking for Senior Site Reliability Engineers to build that capability from the groundup.
This is a Platform/SRE role with a strong automation and software-engineering bias — not an AI model-building role. You'll turn runbooks, alerts and operational workflows into safe, auditable automation that improves reliability across the platform.
- Get in at the ground floor of a brand-new US operations function and shape how it runs at scale
- Work on critical infrastructure powering the next generation of AI
- Automation-first culture - reduce toil and build tooling, rather than fire fight
- Real autonomy and high visibility with leadership
- Fully remote, on a US timezone (West-coast preferred)
What you'll be doing
- Building Python-based automation for incident triage, runbook execution and routine operational tasks
- Integrating observability, ITSM and infrastructure APIs to enrich alerts and automate workflows
- Improving monitoring signal quality through correlation, enrichment, suppression and deduplication
- Building internal tools and self-service capabilities - CLI utilities, ChatOps integrations and dashboards
- Maintaining version-controlled runbook-as-code and automation libraries
- Turning post-incident learnings into better tooling, automation and operational standards
We're keen to speak with candidates who have
- Experience in SRE, Platform Engineering or production infrastructure operations
- Hands-on experience with observability/monitoring tooling (Prometheus, Grafana or similar)
- Exposure to incident management / on-call, and converting manual runbooks into automation
- Strong Python for automation, APIs and integrations
Nice to have:
- GPU, datacentre or colocation infrastructure experience
- ITSM integrations (ServiceNow, Halo, Jira Service Management or similar)
- ChatOps tooling (Slack or Microsoft Teams bots)
- OpenTelemetry, logging or distributed tracing experience
- DCIM, IPAM or hypervisor-control-plane integrations
- Experience with LLM-assisted or agent-based operational automation