Site Reliability Engineering - Java, Kafka, Node.js & AI
Location
., ., .
Job Type
55.00 - 60.00
Deadline
Not specified
Education
Bachelor's Degree
Experience
Executive (10+ years)
About the Role
We are seeking an experienced Site Reliability Engineering (SRE) Lead with strong expertise in Java, Apache Kafka, and Node.js along with awareness of AI/Generative AI technologies . The ideal candidate will lead reliability initiatives, enhance platform performance, drive automation, and ensure highly available and scalable systems in a cloud-native environment.
Key Responsibilities
- Lead SRE initiatives to improve system reliability, scalability, and performance.
- Design, implement, and support highly available distributed applications.
- Develop and maintain services using Java and Node.js .
- Build and manage event-driven architectures using Apache Kafka .
- Establish observability, monitoring, logging, and alerting frameworks.
- Drive automation for deployments, incident response, and operational processes.
- Collaborate with development, infrastructure, and DevOps teams to improve system resilience.
- Conduct root cause analysis and implement preventive measures.
- Support CI/CD pipelines and infrastructure automation.
- Stay informed on emerging AI and Generative AI technologies and identify opportunities for operational improvements.
Requirements
Required Skills
- 12+ years of IT experience with strong SRE or Production Engineering background.
- Hands-on experience with Java and Node.js development.
- Strong expertise with Apache Kafka and event-driven architectures.
- Experience with monitoring and observability tools such as Splunk, Prometheus, Grafana, Datadog, or ELK.
- Knowledge of cloud platforms (AWS, Azure, or GCP).
- Experience with Docker, Kubernetes, and containerized environments.
- Strong understanding of CI/CD pipelines and automation tools.
- Excellent troubleshooting and incident management skills.
- Exposure to AI/ML or Generative AI concepts is highly desirable.
Preferred Qualifications
- Experience leading SRE or platform engineering teams.
- Knowledge of Infrastructure as Code tools such as Terraform or Ansible.
- Familiarity with DevSecOps practices.
Openings: 4
Location: Phoenix, AZ (Hybrid) – Local candidates preferred
Skills
AI and Generative AI technologies Apache Java Node.js
#J-18808-LjbffrSite Reliability Engineering - Java, Kafka, Node.js & AI in phoenix at Unknown Company
This position is listed as full time and hybrid.