To support a growing Site Reliability Engineering practice, the full-time Senior Network Reliability Engineer will engineer network reliability through automation, define SLOs, and enhance observability while working remotely. Key responsibilities Define SLOs and error budgets for network performance, leading postmortems focused on permanent remediation Implement Infrastructure as Code (IaC) using Terraform and Ansible to automate network configurations and improve operational efficiency Enhance telemetry and observability by building structured alerts and dashboards that reflect user experience and system health Required qualifications Deep understanding of TCP/IP, BGP, OSPF, VPNs, and SD-WAN architecture Proven experience with Terraform and Ansible in a production environment, along with proficiency in Python for automation Hands-on experience with security platforms such as Cloudflare and Zscaler Experience configuring monitoring tools like Datadog, Prometheus, and Grafana Strong documentation skills and a mindset focused on reducing operational toil through automation
Senior Network Reliability Engineer in workfromhome at Unknown Company
This position is listed as full time and able to be worked remotely.