Must Have Technical/Functional Skills
For an SRE / AWS DevOps Engineer with Arize Observability, the candidate should have knowledge on AWS Cloud, DevOps, Reliability Engineering, Monitoring, MLOps, and AI Observability.
SRE Lead with 8+ years of experience in building and managing scalable, secure, and highly available AWS cloud platforms leveraging DevOps, Kubernetes, and Infrastructure-as-Code practices.
For an SRE / AWS DevOps Engineer with Arize Observability, the candidate should have knowledge on AWS Cloud, DevOps, Reliability Engineering, Monitoring, MLOps, and AI Observability.
- AWS Cloud Services - EC2, EKS, ECS, Lambda, S3, RDS, Redshift, CloudWatch, IAM, VPC.
- Infrastructure as Code (IaC) - Terraform, AWS CloudFormation, Ansible.
- CI/CD Automation - Jenkins, GitHub Actions, GitLab CI/CD, AWS CodePipeline.
- Containerization & Orchestration - Docker, Kubernetes (EKS), Helm.
- Site Reliability Engineering (SRE) - SLI/SLO/SLA management, incident response, root cause analysis (RCA), reliability engineering.
- Monitoring & Observability - Arize AI, Prometheus, Grafana, Datadog, ELK Stack, OpenTelemetry, CloudWatch.
- MLOps & AI Observability - Arize platform, model monitoring, drift detection, model performance tracking, data quality monitoring, LLM observability.
- Programming & Scripting - Python, Bash, PowerShell, SQL.
- Security & DevSecOps - IAM, Secrets Manager, AWS Security Hub, vulnerability scanning, policy enforcement.
- Collaboration & Agile Delivery - Scrum, Jira, stakeholder communication, cross-functional incident management, technical documentation.
SRE Lead with 8+ years of experience in building and managing scalable, secure, and highly available AWS cloud platforms leveraging DevOps, Kubernetes, and Infrastructure-as-Code practices.
- Experienced in implementing CI/CD pipelines, driving SRE best practices, and ensuring platform reliability through proactive monitoring, incident management, and performance optimization using CloudWatch, Prometheus, Grafana, and Arize.
- Strong collaborator with engineering, data, and ML teams, enabling MLOps, AI model observability, drift detection, and reliable deployment of production-grade AI/GenAI solutions.
- Strong leadership and stakeholder management skills, with experience leading cross-functional teams and driving delivery excellence.
- Effective communication with business and technical stakeholders.
SRE AWS DevOps with Arize in malvern at Unknown Company
This position is listed as full time and onsite.