Unknown Company
hackajob is collaborating with Oracle to connect them with exceptional professionals for this role.
About the team
Leads development and architecting of GPU operation with customer awareness
Description
We are seeking a technical operations leader to join the AI Infra Operations team as a Principal Core Infrastructure Engineer supporting GPU infrastructure in Oracle Cloud Infrastructure (OCI). In this role, you will lead the development and maintenance of automation and operational tooling for GPU fleets across multiple regions, ensuring high availability. You will collaborate closely with engineering and operations teams to build robust automation, observability, and reliability solutions while continuously improving GPU operations.
Responsibilities
* Build and maintain automation and operational tooling for OCI GPU infrastructure across multiple geographic regions.
* Drive collaboration with software engineers, hardware teams, and operations partners to maintain a highly available GPU fleet.
* Build and improve monitoring, alerting, and diagnostics for GPU fleet health, performance, capacity, and utilization using tools such as Grafana.
* Serve as the senior escalation point for complex GPU host and repair issues
* Participate in incident response and root-cause analysis to remove blockers affecting GPU capacity, availability, and regional deployments.
* Continuously improve AI2 Ops processes, GPU fleet automation, and OCI region build readiness.
* Participate in on-call rotations and provide support for critical infrastructure issues.
* Document operational procedures, automation workflows, troubleshooting guides, and runbooks.
* Build and improve AI agents, ensuring safe rollout, execution and monitoring.
* Mentor and guide junior engineers in operational best practices, provide senior technical support, and drive constant improvement.
Qualifications
Disclaimer:
Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.
Range and benefit information provided in this posting are specific to the stated locations only
US: Hiring Range in USD from: $114,600 to $234,600 per annum. May be eligible for bonus, equity, and compensation deferral.
Oracle maintains broad salary ranges for its roles in order to
About the team
Leads development and architecting of GPU operation with customer awareness
Description
We are seeking a technical operations leader to join the AI Infra Operations team as a Principal Core Infrastructure Engineer supporting GPU infrastructure in Oracle Cloud Infrastructure (OCI). In this role, you will lead the development and maintenance of automation and operational tooling for GPU fleets across multiple regions, ensuring high availability. You will collaborate closely with engineering and operations teams to build robust automation, observability, and reliability solutions while continuously improving GPU operations.
Responsibilities
* Build and maintain automation and operational tooling for OCI GPU infrastructure across multiple geographic regions.
* Drive collaboration with software engineers, hardware teams, and operations partners to maintain a highly available GPU fleet.
* Build and improve monitoring, alerting, and diagnostics for GPU fleet health, performance, capacity, and utilization using tools such as Grafana.
* Serve as the senior escalation point for complex GPU host and repair issues
* Participate in incident response and root-cause analysis to remove blockers affecting GPU capacity, availability, and regional deployments.
* Continuously improve AI2 Ops processes, GPU fleet automation, and OCI region build readiness.
* Participate in on-call rotations and provide support for critical infrastructure issues.
* Document operational procedures, automation workflows, troubleshooting guides, and runbooks.
* Build and improve AI agents, ensuring safe rollout, execution and monitoring.
* Mentor and guide junior engineers in operational best practices, provide senior technical support, and drive constant improvement.
Qualifications
Disclaimer:
Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.
Range and benefit information provided in this posting are specific to the stated locations only
US: Hiring Range in USD from: $114,600 to $234,600 per annum. May be eligible for bonus, equity, and compensation deferral.
Oracle maintains broad salary ranges for its roles in order to
Principal Core Infrastructure Engineer (Hardware Accelerator, High Availability Systems) in nashville at Unknown Company
This position is listed as full time and onsite.