Unknown Company

Site Reliability Engineer

seattle, washington • Posted 4 days ago
Onsite Full Time General

Job Title: Site Reliability Engineer (SRE) - Middleware API

Key Skills: SRE, Middleware API, Scripting and Automation tools

Experience: 10+ Years’ experience

Location: Seattle, Washington & Atlanta, Georgia

We at Coforge are hiring experienced professionals with strong knowledge of Site Reliability Engineering (SRE), middleware APIs, scripting (Python/Bash/Ansible), cloud platforms (AWS/Azure/GCP), Kubernetes, monitoring tools (Prometheus, Grafana, ELK), and incident management.

Key Responsibilities:

  • Drive site reliability engineering (SRE) practices to ensure high availability, scalability, and performance of systems.
  • Design, develop, and maintain middleware APIs and integrations.
  • Build and manage CI/CD pipelines with automation using tools and scripting (Python, Bash, Ansible).
  • Deploy, manage, and optimize applications on cloud platforms (AWS/Azure/GCP).
  • Work with Kubernetes for container orchestration, deployment, and scaling.
  • Implement and maintain monitoring, logging, and alerting solutions (Prometheus, Grafana, ELK).
  • Lead incident management, including troubleshooting, root cause analysis (RCA), and resolution.
  • Ensure system reliability, uptime, and performance through proactive measures and automation.
  • Collaborate with development teams to improve system design, resiliency, and observability.
  • Enforce best practices for security, compliance, and operational excellence.

Required Skills

  • Design, implement, and maintain robust Middleware API solutions to ensure high availability and performance.
  • Monitor and troubleshoot system performance, identifying and resolving issues proactively.
  • Collaborate with development teams to integrate SRE practices into the software development lifecycle.
  • Develop and maintain automation tools for deployment, monitoring, and incident response.
  • Establish and enforce best practices for API design, security, and documentation.
  • Participate in on-call rotations and incident response efforts to ensure system reliability.
  • Conduct post-incident reviews and implement improvements based on findings.
  • Provide technical guidance and mentorship to junior team members.

Preferred Qualifications

  • Strong SRE knowledge, middleware API experience, scripting (Python/Bash/Ansible), cloud (AWS/Azure/GCP), Kubernetes, monitoring (Prometheus/Grafana/ELK), and incident management expertise.

#J-18808-Ljbffr
Back to Job Search