Program Manager 4 (Computer-based Tools, Implementing Program Strategies)
Archer is matching candidates to this role at Oracle. Create a free profile and Archer will check you against this role and every other live role, showing you exactly where you match.
About the team
This role leads high-impact GPU cluster health strategy and transformation programs that span customers, repair execution teams, SDE/SRE tooling, partner operations, spares/RMA strategy, and executive governance. This role defines KPIs, shapes scalable operating models, leads complex cross-functional programs, resolves escalated issues, and influences process improvement across multiple teams.
Description
The GPU Cluster Health Repair Strategy team is accountable for driving repair strategy, operational execution, and cross-functional governance to keep customer GPU cluster availability at or above 97.5% at all times, where a customer cluster is defined as a shape/location combination. The team establishes KPIs, monitors performance, identifies gaps, launches corrective programs, and partners with engineering, SDE/SRE, data center operations, GSL, CHS, CPV, Warminator, TRS, customer-facing teams, and hardware partners such as NVIDIA and AMD.
This role will lead complex, high-impact program improvements at organizational scale. Defines KPIs, shapes automation, drives senior-leadership alignment, and serves as an escalation point for multi-team issues.
Responsibilities
Strategic Program Leadership
Own strategic repair-transformation programs that materially improve GPU cluster availability, repair velocity, partner accountability, and operational efficiency.
Define end-to-end program strategy for high-risk availability areas, including long-running repair reduction, proactive risk detection, RMA loop improvement, spare placement optimization, healthy-node custody policy, partner hardware quality feedback, or AI-driven repair insights.
Translate customer availability risk into strategic priorities, measurable goals, and execution plans.
Availability and Repair Governance
Drive alignment across customer vertical TPMs, horizontal program TPMs, SDE/SRE, engineering, partner teams, and operations.
Lead executive-level reviews for high-risk customers, high-risk shapes, or high-risk locations.
Define governance mechanisms to keep every shape/location cluster above 97.5% availability.
Advise leadership on SLA/SLO performance, repair productivity, operational bottlenecks, and resourcing gaps.
KPI and Data Framework Ownership
Define the KPI framework for the team, including business metrics, operational metrics, leading indicators, lagging indicators, and escalation thresholds.
Build the measurement model for shape/location health, cluster-level availability, repair aging, throughput, RMA loop time, spares sufficiency, hardware failure rate, partner blockers, and tooling efficiency.
Shape advanced reporting and forecasting requirements with SDE/SRE and data teams.
Drive consistent use of data to identify risk before customer escalation.
Process and Tooling Transformation
Lead efforts to reduce manual repair execution through automation, AI-assisted triage, workflow instrumentation, and repair agent tooling.
Partner with SDE/SRE teams in India, Morocco, and Mexico to prioritize tooling investments that accelerate repair diagnosis, status visibility, escalation, and closure.
Drive adoption of standardized SOPs, playbooks, partner handoffs, and executive escalation paths.
Partner and Ecosystem Influence
Lead complex engagements with partners such as NVIDIA and AMD where hardware quality, spare availability, RMA turnaround, or repair playbooks affect customer availability.
Establish data-backed partner accountability mechanisms.
Drive corrective action plans where partner performance affects cluster health.
Escalation and Risk Management
Serve as escalation lead for complex availability risks that span multiple organizations.
Drive root-cause analysis for recurring repair failures, long-tail repair issues, spare shortages, or process breakdowns.
Make trade-offs between customer impact, business risk, technical constraints, and operational feasibility.
Qualifications
Disclaimer:
Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.
Range and benefit information provided in this posting are specific to the stated locations only
US: Hiring Range in USD from: $90,100 to $209,500 per annum. May be eligible for bonus and equity.
Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.
Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.
Oracle US offers a comprehensive benefits package which includes the following:
1. Medical, dental, and vision insurance, including expert medical opinion
2. Short term disability and long term disability
3. Life insurance and AD&D
4. Supplemental life insurance (Employee/Spouse/Child)
5. Health care and dependent care Flexible Spending Accounts
6. Pre-tax commuter and parking benefits
7. 401(k) Savings and Investment Plan with company match
8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
9. 11 paid holidays
10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
11. Paid parental leave
12. Adoption assistance
13. Employee Stock Purchase Plan
14. Financial planning and group legal
15. Voluntary benefits including auto, homeowner and pet insurance
The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.
As part of Oracle's onboarding process and consistent with applicable law, US-based employees are required to complete identity verification, which involves the collection and processing of their biometric information. Accommodations to this requirement may be granted following an individualized assessment.
Apply knowing you're qualified
One free profile is all it takes. Archer checks you against this role and every other live role we list, and shows you exactly which requirements you meet before you apply.
More roles like this
See all Operations Manager jobs- Posted 30+ days agoElectrical EngineerMechanical EngineerCyber Security Engineer +5
- Posted 5 days ago
Operations AnalystProject ManagerTechnical Program Manager +5 - Posted 22 days agoProgram ManagerOperations ManagerPMO Analyst +3
Not quite the right role?
Archer scans thousands of live roles and surfaces the ones you genuinely match, each with a clear explanation of why. It keeps working after you apply, so you hear about roles you would never have found by searching.
Create your free profile