Unknown Company

AWS SRE / Observability Engineer focused on AI Operations

charlotte, nc • Posted 1 weeks ago
Onsite Full Time Architecture and Engineering Occupations

Job Description

n

Insight Global is seeking an AWS Site Reliability Engineer (SRE) / Observability Engineer for a leading technology-focused client. This individual will play a critical role in maintaining the reliability, scalability, and operational intelligence of cloud-based applications and infrastructure while supporting emerging AI-enabled operational workflows.

n

This position sits at the intersection of AWS cloud operations, observability engineering, incident management, and AI Operations. The ideal candidate is a highly analytical problem solver who can quickly dive into monitoring platforms, investigate logs, analyze metrics and alerts, identify root causes, and drive rapid issue resolution. They will leverage tools such as AWS, Terraform, Datadog, Splunk, Dynatrace, GitLab, and modern observability platforms to improve system health, operational visibility, and service reliability.

n

In addition to traditional SRE responsibilities, this individual will help support a growing AI Operations ecosystem by assisting with agentic AI workflows, participating in human-in-the-loop validation processes, evaluating AI-generated outputs, and refining AI prompts to improve operational effectiveness. The team is looking for someone who can thrive in a fast-paced environment, rapidly understand complex systems, determine operational impact, and continuously improve observability, monitoring, and AI-assisted operational processes.

n

This is an exciting opportunity for an engineer who enjoys solving production challenges, strengthening observability practices, driving incident response initiatives, and helping organizations scale next-generation AI-powered operational capabilities.

n

-Monitor and support AWS application and infrastructure environments

n

-Plan and execute application and infrastructure configuration changes

n

-Respond to production incidents, critical outages, and operational emergencies

n

-Triage application and infrastructure issues using monitoring and observability platforms

n

-Analyze logs, metrics, traces, and alerts to identify root cause and operational impact

n

-Determine whether issues are system-related, application-related, infrastructure-related, or AI workflow-related

n

-Lead incident response efforts and coordinate cross-functional resolution activities

n

-Develop and maintain postmortems, operational documentation, and technical runbooks

n

-Partner with software engineering teams to improve reliability, resiliency, and performance

n

-Enhance monitoring, logging, alerting, and observability capabilities across the environment

n

-Drive faster issue detection and resolution through monitoring best practices

n

-Support emerging agentic AI solutions being deployed across operational workflows

n

-Perform human-in-the-loop validation of AI-generated recommendations and outputs

n

-Review, refine, and optimize AI prompts to improve accuracy and operational effectiveness

n

-Collaborate with AI, development, and operations teams to improve operational intelligence

n

-Develop automation solutions that reduce manual effort and improve operational efficiency

n

-Participate in system design reviews, capacity planning initiatives, and architectural discussions

n

-Drive continuous improvement efforts focused on service reliability, observability maturity, and operational excellence

n

We are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to learn more about how we collect, keep, and process your private information, please review Insight Global's Workforce Privacy Policy:

n

Skills and Requirements

n

Terraform expertise for Infrastructure as Code (IaC)

n

Strong AWS experience, particularly with EC2 and cloud-native services

n

Experience supporting AWS environments, including API Gateway technologies

n

Strong experience with Datadog monitoring and observability

n

Experience with enterprise monitoring and observability platforms such as Datadog, Dynatrace, and Splunk

n

GitLab and source control experience using Git

n

Strong troubleshooting, incident response, and production support experience

n

Strong log analysis experience within enterprise monitoring environments

n

Ability to quickly triage production issues and determine root cause using observability tools

n

Observability engineering experience, including identifying performance bottlenecks and system issues

n

Experience leading incident response efforts and supporting application teams

n

Experience creating and facilitating blameless postmortems

n

Ability to improve logging strategies, operational runbooks, and reduce technical debt

n

Experience supporting high-availability production applications and infrastructure

n

Experience supporting AI-powered operational workflows with human oversight and validation

n

Strong analytical mindset with the ability to rapidly understand system behavior and operational impact

n

Strong collaboration and communication skills when working across development, operations, and platform teams Experience with Agentic AI systems or AI-enabled operational platforms

n

Prompt engineering experience, including creating and optimizing AI prompts

n

Experience working with Anthropic Claude or comparable enterprise AI models

n

Experience evaluating and validating AI-generated outputs

n

Experience supporting AI observability and AI Operations initiatives

n

API development experience

n

Experience building and supporting microservices architectures

n

Front-end development experience

n

Automation and platform engineering experience

n

Capacity planning and system design consulting experience

n

Experience promoting observability best practices across engineering organizations

n

Previous experience in large-scale cloud environments

Back to Job Search