Job Summary
Site Reliability Engineer (SRE) with strong background in Google Cloud Platform and RedHat OpenShift administration. Responsible for ensuring reliability, performance, and scalability of on‑premise and cloud‑based systems and reducing costs for Google Cloud.
Responsibilities
- Ensure reliability and uptime of critical services and infrastructure.
- Design, implement, and manage cloud infrastructure using Google Cloud services.
- Develop and maintain automation scripts and tools to improve system efficiency and reduce manual intervention.
- Implement monitoring solutions and respond to incidents to minimize downtime and ensure quick recovery.
- Collaborate with development and operations teams to improve system reliability and performance.
- Conduct capacity planning and performance tuning to ensure systems can handle future growth.
- Create and maintain comprehensive documentation for system configurations, processes, and procedures.
Qualifications
- Bachelor’s degree in Computer Science, Engineering, or related field.
- 10+ years of experience in site reliability engineering or a similar role.
Skills
- Proficiency in Google Cloud services (Compute Engine, Kubernetes Engine, Cloud Storage, BigQuery, Pub/Sub, etc.).
- Familiarity with Google BI and AI/ML tools (Looker, BigQuery ML, Vertex AI, etc.).
- Experience with automation tools (Terraform, Ansible, Puppet).
- Familiarity with CI/CD pipelines and tools (Azure pipelines, Jenkins, GitLab CI, etc.).
- Strong scripting skills (Python, Bash, etc.).
- Knowledge of networking concepts and protocols.
- Experience with monitoring tools (Prometheus, Grafana, etc.).
Preferred Certifications
- Google Cloud Professional DevOps Engineer
- Google Cloud Professional Cloud Architect
- Red Hat Certified Engineer (RHCE) or similar Linux certification
Site Reliability Engineer in san jose at Unknown Company
This position is listed as full time and onsite.