Job Description
nInsight Global is seeking an AWS Site Reliability Engineer (SRE) / Observability Engineer for a leading technology-focused client. This individual will play a critical role in maintaining the reliability, scalability, and operational intelligence of cloud-based applications and infrastructure while supporting emerging AI-enabled operational workflows.
nThis position sits at the intersection of AWS cloud operations, observability engineering, incident management, and AI Operations. The ideal candidate is a highly analytical problem solver who can quickly dive into monitoring platforms, investigate logs, analyze metrics and alerts, identify root causes, and drive rapid issue resolution. They will leverage tools such as AWS, Terraform, Datadog, Splunk, Dynatrace, GitLab, and modern observability platforms to improve system health, operational visibility, and service reliability.
nIn addition to traditional SRE responsibilities, this individual will help support a growing AI Operations ecosystem by assisting with agentic AI workflows, participating in human-in-the-loop validation processes, evaluating AI-generated outputs, and refining AI prompts to improve operational effectiveness. The team is looking for someone who can thrive in a fast-paced environment, rapidly understand complex systems, determine operational impact, and continuously improve observability, monitoring, and AI-assisted operational processes.
nThis is an exciting opportunity for an engineer who enjoys solving production challenges, strengthening observability practices, driving incident response initiatives, and helping organizations scale next-generation AI-powered operational capabilities.
n-Monitor and support AWS application and infrastructure environments
n-Plan and execute application and infrastructure configuration changes
n-Respond to production incidents, critical outages, and operational emergencies
n-Triage application and infrastructure issues using monitoring and observability platforms
n-Analyze logs, metrics, traces, and alerts to identify root cause and operational impact
n-Determine whether issues are system-related, application-related, infrastructure-related, or AI workflow-related
n-Lead incident response efforts and coordinate cross-functional resolution activities
n-Develop and maintain postmortems, operational documentation, and technical runbooks
n-Partner with software engineering teams to improve reliability, resiliency, and performance
n-Enhance monitoring, logging, alerting, and observability capabilities across the environment
n-Drive faster issue detection and resolution through monitoring best practices
n-Support emerging agentic AI solutions being deployed across operational workflows
n-Perform human-in-the-loop validation of AI-generated recommendations and outputs
n-Review, refine, and optimize AI prompts to improve accuracy and operational effectiveness
n-Collaborate with AI, development, and operations teams to improve operational intelligence
n-Develop automation solutions that reduce manual effort and improve operational efficiency
n-Participate in system design reviews, capacity planning initiatives, and architectural discussions
n-Drive continuous improvement efforts focused on service reliability, observability maturity, and operational excellence
nWe are a company committed to creating diverse and inclusive environments where people can bring their full, authentic selves to work every day. We are an equal opportunity/affirmative action employer that believes everyone matters. Qualified candidates will receive consideration for employment regardless of their race, color, ethnicity, religion, sex (including pregnancy), sexual orientation, gender identity and expression, marital status, national origin, ancestry, genetic factors, age, disability, protected veteran status, military or uniformed service member status, or any other status or characteristic protected by applicable laws, regulations, and ordinances. If you need assistance and/or a reasonable accommodation due to a disability during the application or recruiting process, please send a request to learn more about how we collect, keep, and process your private information, please review Insight Global's Workforce Privacy Policy:
nSkills and Requirements
nTerraform expertise for Infrastructure as Code (IaC)
nStrong AWS experience, particularly with EC2 and cloud-native services
nExperience supporting AWS environments, including API Gateway technologies
nStrong experience with Datadog monitoring and observability
nExperience with enterprise monitoring and observability platforms such as Datadog, Dynatrace, and Splunk
nGitLab and source control experience using Git
nStrong troubleshooting, incident response, and production support experience
nStrong log analysis experience within enterprise monitoring environments
nAbility to quickly triage production issues and determine root cause using observability tools
nObservability engineering experience, including identifying performance bottlenecks and system issues
nExperience leading incident response efforts and supporting application teams
nExperience creating and facilitating blameless postmortems
nAbility to improve logging strategies, operational runbooks, and reduce technical debt
nExperience supporting high-availability production applications and infrastructure
nExperience supporting AI-powered operational workflows with human oversight and validation
nStrong analytical mindset with the ability to rapidly understand system behavior and operational impact
nStrong collaboration and communication skills when working across development, operations, and platform teams Experience with Agentic AI systems or AI-enabled operational platforms
nPrompt engineering experience, including creating and optimizing AI prompts
nExperience working with Anthropic Claude or comparable enterprise AI models
nExperience evaluating and validating AI-generated outputs
nExperience supporting AI observability and AI Operations initiatives
nAPI development experience
nExperience building and supporting microservices architectures
nFront-end development experience
nAutomation and platform engineering experience
nCapacity planning and system design consulting experience
nExperience promoting observability best practices across engineering organizations
nPrevious experience in large-scale cloud environments