Job Title: SRE Operations Lead - AWS, GitLab & AIOps
Location: Charlotte, NC (Onsite)
Job Type:- Full Time Permanent
Job Description
Must Have Technical/Functional Skills
- SRE / Application Reliability Engineering (ARE) and 24x7 production operations
- Major Incident Management (P1/P2), ServiceNow, and Incident / Problem / Change Management
- SLI, SLO, SLA governance; MTTR reduction; service reliability KPIs
- AWS: CloudWatch, Route 53, S3, CloudFront, Lambda, ECR, and Bedrock
- Observability: Dynatrace, Grafana, Splunk, and CloudWatch
- Control-M and enterprise batch operations
- GitLab CI/CD, REST APIs, Personal Access Tokens, security scanning, and DevSecOps
- Python GitLab library and API-based automation
- Automation, AIOps, event correlation, self-healing, and stakeholder management
- Lead 24x7 SRE operations and coordinate P1/P2 major incident response through service restoration and follow-up.
- Own reliability measures including SLIs, SLOs, SLAs, MTTR improvement, and service reliability KPIs.
- Drive end-to-end observability using Dynatrace, Grafana, CloudWatch, and Splunk.
- Lead automation, AIOps, self-healing, event correlation, and AI-driven operations initiatives.
- Oversee AWS platform operations, batch processing, and Control-M environments.
- Integrate Claude AI on AWS Bedrock with GitLab using APIs, PATs, and custom workflows.
- Develop AI-driven analysis of GitLab project data, vulnerabilities, pipelines, and security findings.
- Design GitLab API automation, custom workflows, DevSecOps controls, and CI/CD pipeline improvements.
- Build and maintain Python-based GitLab integrations and REST API solutions.
- Support vulnerability remediation and onboarding/configuration of security scanning tools.
- Manage AWS Lambda, ECR, and Bedrock for deployment and automation; optimize Lambda configuration, concurrency, and scaling.
- Design and support resilient multi-region AWS architectures and containerized deployments using Docker and Amazon ECR.
Generic Managerial Skills
- Lead geographically distributed operations teams and coordinate effectively during critical incidents.
- Communicate reliability risks, service performance, and remediation plans to technical and business stakeholders.
- Drive governance, prioritization, continuous improvement, and cross-team collaboration.
- Mentor engineers and promote automation-first, blameless, and reliability-focused ways of working.
SRE Operations Lead - AWS, GitLab & AIOps in charlotte at Unknown Company
This position is listed as full time and onsite.