Requirements
Must have:
- We require a bachelors degree with 12 years of infrastructure or cloud engineering experience, or equivalent practical experience.
- We seek 5 to 10 years of engineering experience with strong expertise in Linux and Windows systems.
- We need hands‑on experience with Kubernetes and container platforms.
- We require experience working in cloud infrastructure environments.
- We need proficiency in scripting languages such as Python and Go.
- We require practical knowledge of Terraform and other automation tools.
- We seek familiarity with monitoring platforms and incident management practices.
- We require experience designing and managing CI/CD pipelines.
- We prefer Kubernetes certifications.
- We prefer AWS or Azure certifications.
- We prefer DevOps certifications.
- ITIL certification is preferred.
- We require an active Secret clearance and the ability to obtain DEA suitability.
Responsibilities
- We design and support highly available production environments.
- We define, track, and manage SLIs, SLOs, and error budgets.
- We automate operational work and reduce manual intervention.
- We build monitoring, alerting, and observability capabilities.
- We improve system performance, capacity, and resilience.
- We lead incident response efforts and perform root cause analysis.
- We implement disaster recovery and business continuity plans.
- We collaborate with development teams to strengthen application reliability.
Company
We are GovCIO, hiring a Site Reliability Engineer for a hybrid or remote role based in Arlington, VA. In this position, we combine software engineering practices with infrastructure operations expertise to improve the reliability, scalability, performance, and availability of mission‑critical systems. The posted salary range is USD $230,000 to $250,000 per year.
#J-18808-Ljbffr