Job DescriptionJob DescriptionSite Reliability Engineer Onsite- Bay Area, CA n
Skills
nRelevant Skills and Experience
nWhat You’ll Do (Day-to-Day)
n- n
- n
Own and manage our cloud infrastructure (GCP or AWS, on-prem).
n n - n
Build, maintain, and optimize Kubernetes clusters (including GPU-backed clusters).
n n - n
Implement and improve CI/CD pipelines (GitHub Actions).
n n - n
Write and maintain Infrastructure as Code (Terraform).
n n - n
Monitor system health and performance using Grafana and other observability tools.
n n - n
Ensure high availability, reliability, and uptime across platforms.
n n - n
Handle infrastructure maintenance, upgrades, and scaling.
n n - n
Administer and improve our platform architecture and apply general security best practices across the stack.
n n
Note: This is an internal-facing role — no customer interaction.
nMust-Have:
n- n
- n
4+ years in SRE, DevOps, or Infrastructure Engineering
n n - n
Solid experience with GCP or AWS (hybrid/on-prem a plus)
n n - n
Experience with Kubernetes cluster management (GPU experience a bonus)
n n - n
Hands-on with Terraform and CI/CD (GitHub)
n n - n
Experience with monitoring/observability (Grafana, etc.)
n n - n
Strong understanding of high availability and infrastructure reliability
n n - n
Familiarity with platform/cluster architecture and administration
n n - n
Security mindset and ability to apply best practice
n n
Nice-to-Have:
n- n
- n
Startup experience (you enjoy building, not just maintaining)
n n - n
Experience with scalable GPU infrastructure for AI/ML
n n
Site Reliability Engineer in Mountain View in mountain view at Unknown Company
This position is listed as full time and hybrid.