Site Reliability Engineer III (AI Platform)
Location: Mount Laurel, NJ (Onsite)
Duration: Contract
Experience: 4+ years
About the Role
We are seeking a Site Reliability Engineer (SRE) III to support a cutting-edge AI Platform Engineering team responsible for building and maintaining the infrastructure behind enterprise AI and machine learning applications. This is an exciting opportunity to work on large-scale distributed systems, Kubernetes environments, and cloud-native platforms that power next-generation AI solutions.
The ideal candidate has a strong background in cloud infrastructure, Kubernetes, Infrastructure as Code, observability, and automation. Experience supporting production environments at scale is essential.
Responsibilities
- Design, implement, and support highly available, scalable, and secure cloud infrastructure.
- Maintain and improve Kubernetes-based production environments.
- Deploy and support AI/ML applications across cloud platforms.
- Automate infrastructure provisioning and operational processes using Terraform and scripting.
- Build and maintain CI/CD pipelines using GitHub Actions.
- Monitor system health and performance using Prometheus, Grafana, Datadog, CloudWatch, Elasticsearch, and logging tools.
- Troubleshoot production issues across distributed systems, cloud infrastructure, networking, and containerized applications.
- Partner with development teams to improve application reliability, scalability, and performance.
- Participate in incident response, root cause analysis, and continuous service improvements.
- Support large-scale Kubernetes clusters and cloud-native workloads.
Required Qualifications
- Bachelor's degree in Computer Science or a related technical field (or equivalent experience).
- 4+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or Cloud Infrastructure.
- Strong understanding of distributed systems, software design, algorithms, and system performance.
- Experience programming or scripting with one or more of the following:
- Python
- Bash
- Go
- Java
- C/C++
- Strong troubleshooting and problem-solving skills.
- Excellent communication and collaboration abilities.
Required Technical Skills
Cloud Platforms
- AWS (required)
- Azure or GCP
Containers & Orchestration
- Kubernetes (required)
- Docker
- Amazon EKS
- Google GKE/GKS
Infrastructure as Code
- Terraform (required)
CI/CD
- GitHub Actions
- CI/CD pipeline automation
Monitoring & Observability
- Prometheus (required)
- Grafana
- Datadog
- CloudWatch
- Elasticsearch
- Centralized logging solutions
Additional Technologies
- MySQL
- Kafka
Preferred Qualifications
- Experience supporting AI/ML platforms or machine learning infrastructure.
- Experience with GPU-based workloads.
- Knowledge of AI model deployment or MLOps.
- Experience automating operational processes.
- Familiarity with enterprise-scale Kubernetes environments.
Nice to Have
- Python or Go development experience.
- Experience with ETL workflows.
- Exposure to AI/ML technologies, LLMs, or agentic AI platforms.
Keywords
Site Reliability Engineer | SRE | DevOps | Platform Engineer | Kubernetes | AWS | Terraform | Docker | EKS | GKE | GitHub Actions | Prometheus | Grafana | Datadog | CloudWatch | Elasticsearch | Kafka | Python | Bash | Go | Infrastructure as Code | Observability | AI Platform | Machine Learning | Cloud Infrastructure | Distributed Systems
GCS is acting as an Employment Business in relation to this vacancy.
#J-18808-Ljbffr