Unknown Company

Observability Engineer / Site Reliability Engineer

california, mo • Posted 2 days ago
Hybrid Full Time IT & Technology

  • Architect, optimize, and maintain observability frameworks across cloud environments, with a specific focus on implementing Google Cloud Platform (GCP) observability tools (Cloud Logging, Cloud Monitoring, Trace, and Profiler).
  • Design, deploy, and maintain robust observability stacks across hybrid ecosystems, utilizing Prometheus, Grafana, and cloud-native integrations.
  • Drive infrastructure-as-code (IaC) initiatives using Terraform and Ansible to ensure consistent, automated deployments of infrastructure and observability tooling.
  • Build, maintain, and optimize deployment workflows within Kubernetes and Google Kubernetes Engine (GKE) / OpenShift environments using GitHub, Harness, and other CI/CD pipelines.
  • Deeply analyze Linux/Unix system administration architectures, optimizing compute resource metrics and performance tuning across complex, distributed environments.
  • Implement SRE best practices, establishing meaningful Service Level Indicators (SLIs), Service Level Objectives (SLOs), and Error Budgets to ensure platform reliability.

Requirements

  • Proven engineering experience within Google Cloud Platform (GCP) environments, particularly managing cloud-native monitoring and compute resources.
  • Hands‑on experience with Grafana, Prometheus, and Google Cloud Observability suites.
  • Expert‑level knowledge of Linux/Unix operating systems paired with strong shell scripting skills for automation and system management.
  • Professional coding proficiency in at least one modern language (Python, Go, Java, Perl, or advanced Shell).
  • Hands‑on experience managing containerized applications on Kubernetes, GKE, and/or Red Hat OpenShift.

Core Competencies

Demonstrates expertise in architecting and optimizing observability frameworks within Google Cloud Platform (GCP) environments, utilizing tools such as Grafana and Prometheus. Proficient in implementing infrastructure-as-code (IaC) practices with Terraform and Ansible, and managing containerized applications in Kubernetes and GKE.

Highest-signal resume keywords

  • Google Cloud Platform (GCP)
  • Grafana
  • Prometheus
  • Infrastructure-as-Code (IaC)
  • Kubernetes

ATS Optimization Keywords

Hard Skills

  • Linux/Unix System Administration
  • Shell Scripting
  • Python
  • Go
  • Java
  • Perl
  • Terraform
  • Ansible
  • Cloud Logging
  • Cloud Monitoring

Industry Keywords

  • Observability Frameworks
  • Service Level Indicators (SLIs)
  • Service Level Objectives (SLOs)
  • Error Budgets
  • Cloud-Native Integrations

Tools & Technologies

  • Google Kubernetes Engine (GKE)
  • OpenShift
  • CI/CD Pipelines
  • GitHub
  • Harness

#J-18808-Ljbffr
Back to Job Search