Unknown Company

Site Reliability Engineering - Java, Kafka, Node.js & AI

phoenix, az • Posted Yesterday
Hybrid Full Time IT & Technology

Site Reliability Engineering - Java, Kafka, Node.js & AI

Location

., ., .

Job Type

55.00 - 60.00

Deadline

Not specified

Education

Bachelor's Degree

Experience

Executive (10+ years)

About the Role

We are seeking an experienced Site Reliability Engineering (SRE) Lead with strong expertise in Java, Apache Kafka, and Node.js along with awareness of AI/Generative AI technologies . The ideal candidate will lead reliability initiatives, enhance platform performance, drive automation, and ensure highly available and scalable systems in a cloud-native environment.

Key Responsibilities

  • Lead SRE initiatives to improve system reliability, scalability, and performance.
  • Design, implement, and support highly available distributed applications.
  • Develop and maintain services using Java and Node.js .
  • Build and manage event-driven architectures using Apache Kafka .
  • Establish observability, monitoring, logging, and alerting frameworks.
  • Drive automation for deployments, incident response, and operational processes.
  • Collaborate with development, infrastructure, and DevOps teams to improve system resilience.
  • Conduct root cause analysis and implement preventive measures.
  • Support CI/CD pipelines and infrastructure automation.
  • Stay informed on emerging AI and Generative AI technologies and identify opportunities for operational improvements.

Requirements

Required Skills

  • 12+ years of IT experience with strong SRE or Production Engineering background.
  • Hands-on experience with Java and Node.js development.
  • Strong expertise with Apache Kafka and event-driven architectures.
  • Experience with monitoring and observability tools such as Splunk, Prometheus, Grafana, Datadog, or ELK.
  • Knowledge of cloud platforms (AWS, Azure, or GCP).
  • Experience with Docker, Kubernetes, and containerized environments.
  • Strong understanding of CI/CD pipelines and automation tools.
  • Excellent troubleshooting and incident management skills.
  • Exposure to AI/ML or Generative AI concepts is highly desirable.

Preferred Qualifications

  • Experience leading SRE or platform engineering teams.
  • Knowledge of Infrastructure as Code tools such as Terraform or Ansible.
  • Familiarity with DevSecOps practices.

Openings: 4

Location: Phoenix, AZ (Hybrid) – Local candidates preferred

Skills

AI and Generative AI technologies Apache Java Node.js

#J-18808-Ljbffr

Site Reliability Engineering - Java, Kafka, Node.js & AI in phoenix at Unknown Company

This position is listed as full time and hybrid.

Back to Job Search