Atos
Phoenix, AZ • Posted Today
Onsite Full Time IT Services and IT Consulting

Job Details:

  • Design, build, and ship LLM-powered and agentic product features that enhance the team efforts and outcomes.
  • Build agentic AI systems that reason over context, invoke tools, take real actions, and recover gracefully from failure.
  • Work on integrating the existing AI tools and should know major AI frameworks and libraries.
  • Own service reliability and operational governance by defining SLA's, managing error budgets, and reporting reliability (MTTD, MTTR) to leadership for prioritization, risk decisions and planning.
  • Architect and continuously optimize the observability of platform using Kibana/Elastic (ELF) along with other observability tools like Prometheus, Grafana (dashboards, metrics, alert lifecycle), improving detection quality, reducing noise/toil, and enabling faster triage and measurable uptime improvements.
  • Engineer advance alerting and automation capabilities with Kibana alerting and anomaly detections and integrating response workflows (routing, runbooks, remediation scripts) to standardize on-call execution and accelerate restoration of services.
  • Lead incident response for customer-impacting issues across teams-coordination, communications, service restoration, and blameless RCA-then corrective actions that prevent recurrence and reduce operational risk.
  • Design, automate and validate Disaster Recovery and failover for critical services/journeys (RTO/RPO alignment, DR Drills), ensuring resiliency under failure scenarios and improving recovery.
  • Consult and partner with application teams by providing production readiness inputs (Resiliency patterns, availability, performance/capacity considerations) and driving platform enhancements that improve stability while optimizing infrastructure and observability spend.
  • Cross-team coordination, incident triage and resolution, leadership and stakeholder management.
  • Perform root cause analysis, identify recurring failure patterns, track corrective/preventive actions, and drive permanent fixes.
  • Review production changes, support deployments, perform pre/post-change validations, monitor critical releases, and support rollback/recovery when required.
  • Drive EMIM bridges for application impacts, troubleshoot issues, coordinate dependent teams, provide technical updates, and support faster service restoration.


Primary Skills:


  • Java
  • SRE
  • Lang-chain, Langraph, RAG, MCP
  • Experience with working on LLM's and integrating with the existing applications
  • Python - FastAPI
  • Cache - Redis


Secondary Skills:


  • Observability - ELK (Elastic/Kibana), Prometheus, Grafana
  • Software and automation - Java, Python/Shell/Bash, Rest-SOAP API, docker containerization, Kubernetes, Kafka
  • Reliability and DR engineering - Distributed architecture and distributed system fundamentals, micro services and event-driven architecture

Java SRE in Phoenix at Atos

This position is listed as full time and onsite. It was posted today.

See all Atos jobs on LocalWork →

Back to Job Search