Unknown Company
hackajob is collaborating with Oracle to connect them with exceptional professionals for this role.
About the team
This role leads high-impact GPU cluster health strategy and transformation programs that span customers, repair execution teams, SDE/SRE tooling, partner operations, spares/RMA strategy, and executive governance. This role defines KPIs, shapes scalable operating models, leads complex cross-functional programs, resolves escalated issues, and influences process improvement across multiple teams.
Description
The GPU Cluster Health Repair Strategy team is accountable for driving repair strategy, operational execution, and cross-functional governance to keep customer GPU cluster availability at or above 97.5% at all times, where a customer cluster is defined as a shape/location combination. The team establishes KPIs, monitors performance, identifies gaps, launches corrective programs, and partners with engineering, SDE/SRE, data center operations, GSL, CHS, CPV, Warminator, TRS, customer-facing teams, and hardware partners such as NVIDIA and AMD.
This role will lead complex, high-impact program improvements at organizational scale. Defines KPIs, shapes automation, drives senior-leadership alignment, and serves as an escalation point for multi-team issues.
Responsibilities
Strategic Program Leadership
* Own strategic repair-transformation programs that materially improve GPU cluster availability, repair velocity, partner accountability, and operational efficiency.
* Define end-to-end program strategy for high-risk availability areas, including long-running repair reduction, proactive risk detection, RMA loop improvement, spare placement optimization, healthy-node custody policy, partner hardware quality feedback, or AI-driven repair insights.
* Translate customer availability risk into strategic priorities, measurable goals, and execution plans.
Availability and Repair Governance
* Drive alignment across customer vertical TPMs, horizontal program TPMs, SDE/SRE, engineering, partner teams, and operations.
* Lead executive-level reviews for high-risk customers, high-risk shapes, or high-risk locations.
* Define governance mechanisms to keep every shape/location cluster above 97.5% availability.
* Advise leadership on SLA/SLO performance, repair productivity, operational bottlenecks, and resourcing gaps.
KPI and Data Framework Ownership
* Define the KPI framework for the team, including business metrics, operational metrics, leading indicators, lagging indicators, and escalation thresholds.
* Build the measurement model for shape/location health, cluster-level availability, repair aging, throughput, RMA loop time, spares sufficiency, hardware failure rate, partner blockers, and tooling efficiency.
* Shape advanced reporting and forecasting requirements with SDE/SRE and data teams.
* Drive consistent use of data to identify risk before customer escalation.
Process and Tooling Transformation
* Lead efforts to reduce manual repair execution through automation, AI-assisted triage, workflow instrumentation, and repair agent tooling.
* Partner with SDE/SRE teams in India, Morocco, and Mexico to prioritize tooling investments that accelerate repair diagnosis, status visibility, escalation, and closure.
* Drive adoption of standardized SOPs, playbooks, partner handoffs, and executive escalation paths.
Partner and Ecosystem Influence
* Lead complex engagements with partners such as NVIDIA and AMD where hardware quality, spare availability, RMA turnaround, or repair playbooks affect customer availability.
* Establish data-backed partner accountability mechanisms.
* Drive corrective action plans where partner performance affects cluster health.
Escalation and Risk Management
* Serve as escalation lead for complex availability risks that span multiple organizations.
* Drive root-cause analysis for recurring repair failures, long-tail repair issues, spare shortages, or process breakdowns.
* Make trade-offs between customer impact, business risk, technical constraints, and operational feasibility.
Qualifications
Disclaimer:
Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.
Range and benefit information provided in this posting are specific to the stated locations only
US: Hiring Range in USD from: $90,100 to $209,500 per annum. May be eligible for bonus and equity.
Oracle maintains broad salary ranges for its roles in order to
About the team
This role leads high-impact GPU cluster health strategy and transformation programs that span customers, repair execution teams, SDE/SRE tooling, partner operations, spares/RMA strategy, and executive governance. This role defines KPIs, shapes scalable operating models, leads complex cross-functional programs, resolves escalated issues, and influences process improvement across multiple teams.
Description
The GPU Cluster Health Repair Strategy team is accountable for driving repair strategy, operational execution, and cross-functional governance to keep customer GPU cluster availability at or above 97.5% at all times, where a customer cluster is defined as a shape/location combination. The team establishes KPIs, monitors performance, identifies gaps, launches corrective programs, and partners with engineering, SDE/SRE, data center operations, GSL, CHS, CPV, Warminator, TRS, customer-facing teams, and hardware partners such as NVIDIA and AMD.
This role will lead complex, high-impact program improvements at organizational scale. Defines KPIs, shapes automation, drives senior-leadership alignment, and serves as an escalation point for multi-team issues.
Responsibilities
Strategic Program Leadership
* Own strategic repair-transformation programs that materially improve GPU cluster availability, repair velocity, partner accountability, and operational efficiency.
* Define end-to-end program strategy for high-risk availability areas, including long-running repair reduction, proactive risk detection, RMA loop improvement, spare placement optimization, healthy-node custody policy, partner hardware quality feedback, or AI-driven repair insights.
* Translate customer availability risk into strategic priorities, measurable goals, and execution plans.
Availability and Repair Governance
* Drive alignment across customer vertical TPMs, horizontal program TPMs, SDE/SRE, engineering, partner teams, and operations.
* Lead executive-level reviews for high-risk customers, high-risk shapes, or high-risk locations.
* Define governance mechanisms to keep every shape/location cluster above 97.5% availability.
* Advise leadership on SLA/SLO performance, repair productivity, operational bottlenecks, and resourcing gaps.
KPI and Data Framework Ownership
* Define the KPI framework for the team, including business metrics, operational metrics, leading indicators, lagging indicators, and escalation thresholds.
* Build the measurement model for shape/location health, cluster-level availability, repair aging, throughput, RMA loop time, spares sufficiency, hardware failure rate, partner blockers, and tooling efficiency.
* Shape advanced reporting and forecasting requirements with SDE/SRE and data teams.
* Drive consistent use of data to identify risk before customer escalation.
Process and Tooling Transformation
* Lead efforts to reduce manual repair execution through automation, AI-assisted triage, workflow instrumentation, and repair agent tooling.
* Partner with SDE/SRE teams in India, Morocco, and Mexico to prioritize tooling investments that accelerate repair diagnosis, status visibility, escalation, and closure.
* Drive adoption of standardized SOPs, playbooks, partner handoffs, and executive escalation paths.
Partner and Ecosystem Influence
* Lead complex engagements with partners such as NVIDIA and AMD where hardware quality, spare availability, RMA turnaround, or repair playbooks affect customer availability.
* Establish data-backed partner accountability mechanisms.
* Drive corrective action plans where partner performance affects cluster health.
Escalation and Risk Management
* Serve as escalation lead for complex availability risks that span multiple organizations.
* Drive root-cause analysis for recurring repair failures, long-tail repair issues, spare shortages, or process breakdowns.
* Make trade-offs between customer impact, business risk, technical constraints, and operational feasibility.
Qualifications
Disclaimer:
Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.
Range and benefit information provided in this posting are specific to the stated locations only
US: Hiring Range in USD from: $90,100 to $209,500 per annum. May be eligible for bonus and equity.
Oracle maintains broad salary ranges for its roles in order to
Program Manager 4 (Computer-based Tools, Implementing Program Strategies) in nashville at Unknown Company
This position is listed as full time and onsite.