Job Details:
- Design, build, and ship LLM-powered and agentic product features that enhance the team efforts and outcomes.
- Build agentic AI systems that reason over context, invoke tools, take real actions, and recover gracefully from failure.
- Work on integrating the existing AI tools and should know major AI frameworks and libraries.
- Own service reliability and operational governance by defining SLA's, managing error budgets, and reporting reliability (MTTD, MTTR) to leadership for prioritization, risk decisions and planning.
- Architect and continuously optimize the observability of platform using Kibana/Elastic (ELF) along with other observability tools like Prometheus, Grafana (dashboards, metrics, alert lifecycle), improving detection quality, reducing noise/toil, and enabling faster triage and measurable uptime improvements.
- Engineer advance alerting and automation capabilities with Kibana alerting and anomaly detections and integrating response workflows (routing, runbooks, remediation scripts) to standardize on-call execution and accelerate restoration of services.
- Lead incident response for customer-impacting issues across teams-coordination, communications, service restoration, and blameless RCA-then corrective actions that prevent recurrence and reduce operational risk.
- Design, automate and validate Disaster Recovery and failover for critical services/journeys (RTO/RPO alignment, DR Drills), ensuring resiliency under failure scenarios and improving recovery.
- Consult and partner with application teams by providing production readiness inputs (Resiliency patterns, availability, performance/capacity considerations) and driving platform enhancements that improve stability while optimizing infrastructure and observability spend.
- Cross-team coordination, incident triage and resolution, leadership and stakeholder management.
- Perform root cause analysis, identify recurring failure patterns, track corrective/preventive actions, and drive permanent fixes.
- Review production changes, support deployments, perform pre/post-change validations, monitor critical releases, and support rollback/recovery when required.
- Drive EMIM bridges for application impacts, troubleshoot issues, coordinate dependent teams, provide technical updates, and support faster service restoration.
Primary Skills:
- Java
- SRE
- Lang-chain, Langraph, RAG, MCP
- Experience with working on LLM's and integrating with the existing applications
- Python - FastAPI
- Cache - Redis
Secondary Skills:
- Observability - ELK (Elastic/Kibana), Prometheus, Grafana
- Software and automation - Java, Python/Shell/Bash, Rest-SOAP API, docker containerization, Kubernetes, Kafka
- Reliability and DR engineering - Distributed architecture and distributed system fundamentals, micro services and event-driven architecture
Java SRE in Phoenix at Atos
This position is listed as full time and onsite. It was posted today.