This role will report to Cloud Engineering leadership and will be based in San Jose, CA .
Key Responsibilities
- Design, implement, and maintain scalable cloud and hybrid infrastructure supporting enterprise applications, production workloads, AI/data platforms, and CI/CD ecosystems.
- Architect highly available, resilient, secure, and cost-optimized solutions across cloud platforms, with deep focus on AWS.
- Lead adoption and operationalization of Infrastructure-as-Code using Terraform, CloudFormation, and related automation frameworks.
- Develop cloud platform standards, reusable infrastructure modules, engineering patterns, and best practices to improve consistency, scalability, and operational efficiency.
- Build and mature observability capabilities using metrics, logs, traces, dashboards, and automated alerting to improve operational visibility and incident response.
- Drive adoption of SRE practices including service level objectives, error budgets, incident management, operational reviews, capacity planning, and reliability improvement.
- Lead production readiness reviews, root cause analysis, performance optimization, resiliency planning, and operational risk assessments for critical systems.
- Design, build, and optimize CI/CD pipelines using modern DevOps, automation, and GitOps practices.
- Integrate security, compliance, and operational controls into infrastructure provisioning and deployment workflows.
- Implement automated remediation, rollback, guardrails, and self-healing infrastructure patterns to improve reliability and reduce operational risk.
- Establish and enforce operational standards for monitoring, patching, change management, disaster recovery, and production support.
- Partner closely with Security, Compliance, Infrastructure, Application, and Engineering teams to align cloud platforms with enterprise security, hardening, governance, and regulatory requirements.
- Evaluate emerging cloud, automation, AI/ML infrastructure, and platform engineering capabilities to support Bloom Energy’s modernization and scalability goals.
- Mentor cloud, DevOps, and infrastructure engineers while promoting engineering excellence, ownership, documentation, and continuous improvement.
- Participate in on-call rotations and incident escalation processes for critical production systems.
Requirements
- Bachelor’s degree in Computer Science, Information Systems, Engineering, or a related field. Master’s degree preferred.
- At least 10 years of experience in cloud, infrastructure, platform engineering, or DevOps , including 3 or more years in a senior, staff, principal, or technical leadership role .
- Deep hands-on expertise in AWS cloud services , cloud architecture , cloud networking , security , resiliency , and production operations .
- Experience with Azure , Google Cloud , Oracle Cloud , or multi-cloud environments is strongly preferred.
- Proven experience designing and operating mission-critical production environments supporting enterprise applications and services.
- Strong experience in cloud networking and network architecture, including VPC , Transit Gateway , Cloud WAN , Direct Connect , firewalls , routing , and BGP .
- Strong experience with Infrastructure-as-Code tools , including Terraform and CloudFormation .
- Experience with CI/CD and DevOps platforms such as Jenkins , GitLab , GitHub Actions , ArgoCD , or similar tools.
- Experience with Kubernetes , containers , and modern platform engineering patterns.
- Experience with monitoring and observability platforms such as Prometheus , Grafana , Datadog , ELK , CloudWatch , or similar technologies.
- Strong scripting and automation skills using Python , Bash , PowerShell , or similar languages.
- Strong understanding of IAM , cloud security , compliance controls , encryption , resiliency engineering , and operational governance .
- Experience implementing SRE practices , incident management , runbooks , operational reviews , and reliability improvement programs .
- Strong troubleshooting , analytical , problem-solving , and communication skills .
- Demonstrated ability to lead large-scale, cross-functional technical initiatives from strategy through execution.
- Experience supporting high-availability SaaS , AI/ML , manufacturing , energy , industrial technology , or large-scale enterprise environments.
- Experience with FinOps , cloud cost optimization , tagging governance , chargeback/showback , and multi-cloud cost management .
- Experience with AI/ML infrastructure , GPU workloads , data platforms , or large-scale analytics environments.
- Experience developing reusable cloud platform modules , golden patterns , landing zones , or internal developer platforms .
- Experience with cloud governance , compliance automation , security guardrails , and policy-as-code .
- Cloud or platform certifications such as AWS Solutions Architect Professional , AWS DevOps Professional , Kubernetes certifications , Terraform certification , or related credentials.
- Bloom Energy is committed to fair and equitable compensation practices.
- FULL TIME ROLE ONLY: The total compensation for this position includes standard company benefits and is based on various factors including, but not limited to, relevant skills and experience.