Description
nEssential Functions:
n- n
- n
Partner with software developers, platform engineers, and IT staff to improve system design, operability, deployment safety, and production support readiness.
n n - n
Define and maintain operational standards, runbooks, support procedures, escalation paths, and service-level objectives.
n n - n
Evaluate system architecture and changes to ensure they balance functional requirements, service quality, reliability, security, and compliance needs.
n n - n
Drive continuous improvement in platform stability, maintenance, and availability.
n n - n
Provide advanced technical support and troubleshooting for complex platform and service issues affecting internal users and stakeholders.
n n
Experience and Skills Required:
n- n
- n
8+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, Systems Engineering, or related infrastructure roles supporting production services.
n n - n
Strong experience with Linux systems administration and troubleshooting in enterprise environments.
n n - n
Strong experience operating and maintaining on-prem Kubernetes platforms and all related components including CRI, CNI, and CSI plugins.
n n - n
Experience deploying and maintaining applications on Kubernetes using Helm, Kustomize, and similar tooling.
n n - n
Experience supporting DevOps tooling such as GitLab, Artifactory, Jira, Confluence.
n n - n
Experience with GitOps tools such as FluxCD or ArgoCD.
n n - n
Proficiency scripting with at least one of Python, Go, or Bash.
n n - n
Strong experience designing, maintaining, and maturing observability tooling including monitoring, dashboards, logging and tracing, and supporting SLOs.
n n - n
Strong understanding of reliability engineering concepts:
n n - n
Service health indicators
n n - n
High availability design, failure reduction, and testing
n n - n
Operational readiness practices, including developing documentation, runbooks, and architectural descriptions
n n - n
Incident response, root cause analysis, remediation/recovery
n n - n
Ability to obtain a security clearance, which includes U.S. citizenship.
n n
Preferred:
n- n
- n
Experience with multiple Linux distributions including Ubuntu.
n n - n
Experience with at least one of the following: Tanzu Kubernetes, Nutanix Kubernetes Platform, Canonical Kubernetes.
n n - n
Experience with cloud platforms such as AWS and Azure.
n n - n
Experience with infrastructure automation and configuration management.
n n - n
Experience managing AI tooling on Kubernetes including MCP Servers, LLM platforms (vLLM, Ollama), Kubeflow.
n n - n
Experience with security and compliance considerations in regulated environments.
n n - n
DoD experience.
n n - n
Active or inactive Secret Security Clearance.
n n
Education:
n- n
- Bachelor's degree in CS, Software Engineering or other IT-related field or equivalent experience n
REMOTE WORK NOTICE: This position may be performed fully remote, hybrid, or onsite at an ARA office. Preference will be given to candidates located onsite in the Albuquerque area.
nEqual Opportunity Employer/Protected Veterans/Individuals with Disabilities
nThis employer is required to notify all applicants of their rights pursuant to federal employment laws.
nFor further information, please review the Know Your Rights ( notice from the Department of Labor.
Senior Site Reliability Engineer in albuquerque at Unknown Company
This position is listed as full time and able to be worked remotely.