Unknown Company

Site Reliability Engineer

wood ridge, nj • Posted 1 weeks ago
Onsite Full Time Architecture and Engineering Occupations
Job Title: Site Reliability Engineer (SRE)
Location: Wood Ridge, NJ
Duration: 6 Months
Position Overview
We are seeking a highly skilled Site Reliability Engineer (SRE) to join our digital engineering and operations team. The ideal candidate will bring a blend of SRE best practices , DevOps mindset , and IBM WebSphere Commerce Suite (WCS) expertise to ensure the reliability, scalability, and performance of our digital platforms.
This role is responsible for building and maintaining automated systems that monitor, measure, and enhance service uptime and performance. The engineer will collaborate closely with development, infrastructure, and operations teams to drive continuous improvement across the environment.
Key Responsibilities
  • Design, implement, and maintain scalable, reliable, and highly available systems to support critical business applications.
  • Manage and optimize IBM WebSphere Commerce Suite (WCS) environments for high performance and fault tolerance.
  • Develop and implement automation tools and scripts for deployment, monitoring, and maintenance tasks.
  • Proactively monitor system performance, identify bottlenecks, and recommend improvements to increase efficiency and resiliency.
  • Collaborate with software development teams to define service-level objectives (SLOs) and ensure service-level indicators (SLIs) meet operational goals.
  • Perform root cause analysis of production issues and drive permanent corrective actions.
  • Implement and manage DevOps pipelines for continuous integration and deployment (CI/CD).
  • Use monitoring tools (e.g., Prometheus, Grafana, ELK, Splunk, AppDynamics, New Relic) to maintain service health visibility.
  • Apply chaos engineering principles to test system resilience and recovery.
  • Participate in on-call rotations and incident response, ensuring quick restoration of service.
Required Skills & Qualifications
  • 5+ years of hands-on experience as a Site Reliability Engineer, DevOps Engineer, or Systems Engineer .
  • Proven experience with IBM WebSphere Commerce Suite (WCS) deployment, configuration, and optimization.
  • Strong knowledge of SRE principles - monitoring, automation, incident management, reliability metrics (SLO/SLI).
  • Proficiency in scripting languages such as Python, Shell, or Groovy.
  • Experience with CI/CD tools (Jenkins, GitLab CI, or similar).
  • Familiarity with containerization technologies (Docker, Kubernetes).
  • Hands-on experience with cloud platforms (AWS, Azure, or GCP).
  • Excellent troubleshooting and problem-solving skills.
  • Strong collaboration and communication abilities with cross-functional teams.
Preferred Qualifications
  • Experience implementing chaos engineering and resiliency testing frameworks .
  • Background in application performance monitoring (APM) and tuning.
  • Exposure to infrastructure-as-code tools (Terraform, Ansible, or CloudFormation).
  • Certification in AWS/Azure/GCP or related SRE/DevOps technologies.

Site Reliability Engineer in wood ridge at Unknown Company

This position is listed as full time and onsite.

Back to Job Search