Unknown Company

Site Reliability Engineer

mount laurel township, nj • Posted 5 days ago
Onsite Contract Engineering

Site Reliability Engineer III (AI Platform)

Location: Mount Laurel, NJ (Onsite)
Duration: Contract
Experience: 4+ years

About the Role

We are seeking a Site Reliability Engineer (SRE) III to support a cutting-edge AI Platform Engineering team responsible for building and maintaining the infrastructure behind enterprise AI and machine learning applications. This is an exciting opportunity to work on large-scale distributed systems, Kubernetes environments, and cloud-native platforms that power next-generation AI solutions.

The ideal candidate has a strong background in cloud infrastructure, Kubernetes, Infrastructure as Code, observability, and automation. Experience supporting production environments at scale is essential.

Responsibilities

  • Design, implement, and support highly available, scalable, and secure cloud infrastructure.
  • Maintain and improve Kubernetes-based production environments.
  • Deploy and support AI/ML applications across cloud platforms.
  • Automate infrastructure provisioning and operational processes using Terraform and scripting.
  • Build and maintain CI/CD pipelines using GitHub Actions.
  • Monitor system health and performance using Prometheus, Grafana, Datadog, CloudWatch, Elasticsearch, and logging tools.
  • Troubleshoot production issues across distributed systems, cloud infrastructure, networking, and containerized applications.
  • Partner with development teams to improve application reliability, scalability, and performance.
  • Participate in incident response, root cause analysis, and continuous service improvements.
  • Support large-scale Kubernetes clusters and cloud-native workloads.

Required Qualifications

  • Bachelor's degree in Computer Science or a related technical field (or equivalent experience).
  • 4+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Cloud Infrastructure.
  • Strong understanding of distributed systems, software design, algorithms, and system performance.
  • Experience programming or scripting with one or more of the following:
    • Python
    • Bash
    • Go
    • Java
    • C/C++
  • Strong troubleshooting and problem-solving skills.
  • Excellent communication and collaboration abilities.

Required Technical Skills

Cloud Platforms

  • AWS (required)
  • Azure or GCP

Containers & Orchestration

  • Kubernetes (required)
  • Docker
  • Amazon EKS
  • Google GKE/GKS

Infrastructure as Code

  • Terraform (required)

CI/CD

  • GitHub Actions
  • CI/CD pipeline automation

Monitoring & Observability

  • Prometheus (required)
  • Grafana
  • Datadog
  • CloudWatch
  • Elasticsearch
  • Centralized logging solutions

Additional Technologies

  • MySQL
  • Kafka

Preferred Qualifications

  • Experience supporting AI/ML platforms or machine learning infrastructure.
  • Experience with GPU-based workloads.
  • Knowledge of AI model deployment or MLOps.
  • Experience automating operational processes.
  • Familiarity with enterprise-scale Kubernetes environments.

Nice to Have

  • Python or Go development experience.
  • Experience with ETL workflows.
  • Exposure to AI/ML technologies, LLMs, or agentic AI platforms.

Keywords

Site Reliability Engineer | SRE | DevOps | Platform Engineer | Kubernetes | AWS | Terraform | Docker | EKS | GKE | GitHub Actions | Prometheus | Grafana | Datadog | CloudWatch | Elasticsearch | Kafka | Python | Bash | Go | Infrastructure as Code | Observability | AI Platform | Machine Learning | Cloud Infrastructure | Distributed Systems

GCS is acting as an Employment Business in relation to this vacancy.

#J-18808-Ljbffr
Back to Job Search