AI Systems Engineer (OCI/AI Infrastructure)
hackajob is partnering with Oracle to fill this position. Create a free profile and Archer will check you against this role and every other live role, showing you exactly where you match.
Description
Oracle Hardware Platform Development Engineering is seeking a highly driven AI Systems Engineer to evaluate and characterize next-generation GPU and AI accelerator platforms for Oracle Cloud Infrastructure (OCI). This is a hands-on engineering role focused on bringing up new hardware platforms, enabling AI training and inference software stacks, running representative workloads, and analyzing system performance under real operating conditions.
The engineer will identify whether workloads are HBM/memory-bandwidth, compute, scale-up, or scale-out bound, while characterizing power, thermals, memory behavior, utilization, scaling, and performance efficiency. Working directly in the lab, you will debug hardware/software integration issues, design and execute experiments, and develop data-driven insights that explain system behavior beyond benchmark results.
A key part of the role is comparative architecture analysis across GPUs and emerging AI accelerators. You will evaluate architectural tradeoffs and translate performance findings into clear, actionable recommendations on which platforms are best suited for specific AI training and inference workloads. You will work closely with internal hardware and software teams as well as technology partners to help shape Oracle’s next generation of high-performance AI infrastructure.
Position Overview:
This position is ideal for someone who loves deep systems engineering, debugging complex hardware–software interactions, and optimizing performance at every layer of the ML stack. You will play a pivotal role in enabling the training and deployment of next-generation LLMs and generative AI models.
Responsibilities
Required Qualifications
- Solid knowledge of AI / GPU or/and AI/CPU platform architecture and their capabilities.
- Experience with the architecture, design, and implementation of modern server platforms consisting of multiple architectures and vendors, including x86 and ARM server architectures.
- Strong communications skills and ability to clearly communicate complex technical issue across engineering disciplines as well as clearly and succinctly articulate issues for executives.
Experience and understanding of the latest high-speed busses and interconnect used in modern Compute and AI platforms. Familiarity with their startup connectivity and operational robustness as well as performance metrics.
Debugging & Reliability: Troubleshoot complex hardware–software interaction issues, including vLLM compilation failures on ROCm, CUDA memory leaks, distributed runtime failures, and kernel-level inconsistencies.
Profiling & Performance Analysis: Conduct detailed profiling of compilation graphs, training workloads, and runtime execution to optimize performance and eliminate bottlenecks.
Preferred Qualifications
Minimum of 8+ years of experience in developing software infrastructure for large scale AI systems.
Bachelor's degree or higher in Computer Science or a related technical field (or equivalent experience).
Strong debugging skills and experience in analyzing and triaging AI applications from the application level to the hardware level.
Hands-on experience maintaining or building ML training stacks involving CUDA, ROCm, NCCL, XLA, or similar technologies.
Experience in benchmarking AI workloads across different architectures.
Background in working with the large scale clusters
Good understanding on DL frameworks internal PyTorch, TensorFlow, JAX, and Ray
Qualifications
Disclaimer:
Certain U.S. based or U.S. customer or client-facing roles may be required to comply with applicable requirements, such as immunization/occupational health mandates, and/or drug testing requirements.
Range and benefit information provided in this posting are specific to the stated locations only
US: Hiring Range in USD from: $96,800 to $306,400 per annum. May be eligible for bonus, equity, and compensation deferral.
Oracle maintains broad salary ranges for its roles in order to account for variations in knowledge, skills, experience, market conditions and locations, as well as reflect Oracle's differing products, industries and lines of business.
Candidates are typically placed into the range based on the preceding factors as well as internal peer equity.
Oracle US offers a comprehensive benefits package which includes the following:
1. Medical, dental, and vision insurance, including expert medical opinion
2. Short term disability and long term disability
3. Life insurance and AD&D
4. Supplemental life insurance (Employee/Spouse/Child)
5. Health care and dependent care Flexible Spending Accounts
6. Pre-tax commuter and parking benefits
7. 401(k) Savings and Investment Plan with company match
8. Paid time off: Flexible Vacation is provided to all eligible employees assigned to a salaried (non-overtime eligible) position. Accrued Vacation is provided to all other employees eligible for vacation benefits. For employees working at least 35 hours per week, the vacation accrual rate is 13 days annually for the first three years of employment and 18 days annually for subsequent years of employment. Vacation accrual is prorated for employees working between 20 and 34 hours per week. Employees working fewer than 20 hours per week are not eligible for vacation.
9. 11 paid holidays
10. Paid sick leave: 72 hours of paid sick leave upon date of hire. Refreshes each calendar year. Unused balance will carry over each year up to a maximum cap of 112 hours.
11. Paid parental leave
12. Adoption assistance
13. Employee Stock Purchase Plan
14. Financial planning and group legal
15. Voluntary benefits including auto, homeowner and pet insurance
The role will generally accept applications for at least three calendar days from the posting date or as long as the job remains posted.
Career Level - IC5
Apply knowing you're qualified
One free profile is all it takes. Archer checks you against this role and every other live role on hackajob, and shows you exactly which requirements you meet before you apply.
More roles like this
See all Infrastructure Engineer jobs- Posted 19 days agoOperations ManagerData Engineering ManagerFacilities Manager +6
- Posted 19 days ago
Building Automation Engineer
Oracle
Embedded EngineerControl Systems EngineerSystems Engineer +5 - Posted 19 days agoElectrical EngineerMechanical EngineerCyber Security Engineer +5
Not quite the right role?
Archer scans thousands of live roles and surfaces the ones you genuinely match, each with a clear explanation of why. It keeps working after you apply, so you hear about roles you would never have found by searching.
Create your free profile