Site Reliability Engineer/Devops EngineerHonorVet Technologies. We're a veteran-owned IT staffing firm, ISO 9001, and ISO certified, working with federal agencies, state governments, and Fortune 500 enterprise clients across the US. What makes us different isn't a tagline; it's the way we work.
We don't forward resumes and hope for the best. We take the time to understand where a professional like you is headed and only reach out when we genuinely believe there's a fit worth exploring. Role: Site Reliability Engineer/Devops Engineer Location: Austin, TX 78701 (Hybrid – 3 days onsite per week) Engagement type: Contract – 6 MonthsSite Reliability Engineer will be responsible for ensuring the reliability, availability, performance, and scalability of production systems by applying software engineering practices to infrastructure and operations.
Partners with development teams to build resilient, observable, and automated platforms that meet defined service level objectives (SLOs).Required Experience:8 years of experience in systems engineering, DevOps, or site reliability engineering roles8 years of Strong experience with Linux/Unix systems and system internals8 years of Proficiency in one or more programming/scripting languages (Python, Go, Java, Bash)8 years of Experience designing and operating available, distributed systems.8 years of Strong knowledge of cloud platforms (AWS, or GCP) and cloud-native services8 years of Experience with containerization and orchestration (Docker, Kubernetes)8 years of Strong understanding of monitoring, alerting, and logging concepts8 years of Experience defining and managing SLIs, SLOs, and error budgets.8 years of Familiarity with incident management, root cause analysis (RCA), and postmortems8 years of Experience integrating security and compliance into operational workflows.Preferred Experience:4 years of Familiarity with observability tools (Prometheus, Grafana, Application Insights, Datadog, Splunk)4 years of Experience operating 24×7 production environments with on-call rotations4 years of Experience with chaos engineering and resiliency testing4 years of Experience with feature flags, canary deployments, and progressive delivery4 years of Strong documentation skills for runbooks, dashboards, and operational standards