For AI workloads, GPU rental is the bill, and everything else is rounding. A single eight GPU H100 node runs into six figures a year at list on any cloud, a modest training cluster runs into millions, and at that scale a 20 percent rate difference between platforms is not a procurement detail, it is headcount. OCI has positioned itself aggressively here: its GPU list rates undercut the equivalent AWS and Azure instances, its cluster networking carries no separate premium, and its egress pricing means trained models and datasets are not held hostage. This article puts numbers on the comparison, explains where the totals diverge beyond the headline rate, and sets out how to buy GPU capacity on OCI without overcommitting.
This article is part of our complete guide to OCI pricing and TCO. The general compute model it builds on is covered in OCI compute pricing, and the commercial mechanics of committing spend are in negotiating Oracle Universal Credits.
The headline rates
The comparison below uses list prices for on demand capacity, normalized to a single GPU hour, as of early 2026. Treat every figure as orientation: list prices move, negotiated prices move further, and availability often matters more than either. The pattern, however, has been stable for several years: OCI lists below AWS and Azure for the same silicon, often by 20 to 40 percent.
| Capacity | OCI, per GPU hour | AWS, per GPU hour | Azure, per GPU hour |
|---|---|---|---|
| H100 80 GB, 8 GPU node | Around $10 | Around $12 to $12.50 on p5 | Around $12 on ND H100 v5 |
| A100 80 GB, 8 GPU node | Around $4 | Around $5 on p4de | Around $4.90 on ND A100 v4 |
| A100 40 GB | Around $3 | Around $4.10 on p4d | Comparable band |
| Smaller inference GPUs | A10 around $2 | A10G around $1.60 to $2 on g5 | A10 v5 comparable |
Two notes on reading the table honestly. First, at the small end the platforms converge, and AWS g5 instances can beat OCI A10 rates for light inference, so OCI's advantage concentrates at the training and heavy inference end. Second, nobody pays pure list at scale on any cloud: committed spend discounts apply everywhere, and the negotiated number is the real comparison. On OCI that negotiation happens through Universal Credits, which is why the sizing and discount mechanics in our credits article matter more for GPU buyers than for anyone else.
Where the totals diverge beyond the rate
Cluster networking
Training across nodes needs RDMA class networking, and the platforms treat it differently. OCI's cluster networks for GPU shapes carry no separate charge: the RDMA fabric is part of the shape. On AWS, EFA is bundled but the instance families that carry it price accordingly, and on Azure the InfiniBand enabled ND series carries the premium in the instance rate. The practical effect is that OCI's per GPU gap at list understates the full cluster gap slightly, because there is no networking adder on top.
Egress and data gravity
AI work moves big artifacts: datasets in, checkpoints and models out. OCI's 10 TB of free monthly egress and roughly $0.0085 per GB after that, against the nine cents per GB band on AWS and Azure, means shipping a multi terabyte model archive out of OCI costs tens of dollars where the same movement elsewhere costs hundreds or thousands. For teams running multicloud AI pipelines, this line alone can decide where training lives. The full comparison is in OCI egress costs vs AWS and Azure.
Storage feeding the GPUs
Training clusters need fast scratch and large object stores, and idle GPUs waiting on slow storage are the most expensive idle resource in computing. The storage rates themselves, covered in OCI storage pricing, are a small line next to the GPUs, which is exactly why under provisioning storage performance to save pennies while starving ten dollar per hour GPUs is a false economy we see often.
Availability is part of the price
The quoted rate is meaningless if the capacity does not exist when you need it. High end GPU capacity on every cloud is quota gated and frequently reserved ahead by large buyers, and OCI is no exception: H100 and newer capacity generally requires engaging with capacity reservations or committed agreements rather than expecting on demand availability. This changes the buying conversation from a price comparison into a commercial negotiation, where the questions are how much capacity is guaranteed, in which region, from when, and what happens to the commitment if the project's GPU appetite changes. Treating GPU procurement as a click through purchase is how teams end up either blocked without capacity or locked into commitments for capacity they no longer need, the same oversizing failure we document across the wider credits portfolio in the expiring credits problem.
A worked example: a 90 day training campaign
Put the rates into a realistic shape. A team plans a 90 day training campaign on 16 nodes of 8 H100 GPUs each, 128 GPUs in total. At OCI list rates around ten dollars per GPU hour, the raw compute for the campaign runs to roughly $2.8 million if the cluster runs around the clock. The same campaign at AWS or Azure list rates in the twelve dollar band runs $3.4 to $3.5 million. The half million dollar gap is before negotiation, before the egress difference on shipping checkpoints and final weights to other environments, and before the storage and orchestration lines, which are small by comparison on every platform. Now apply the utilization lens: if the cluster idles 30 percent of the campaign waiting on data pipelines, debugging, or scheduling gaps, the waste on either platform dwarfs the rate difference between them. The cheapest decision in the whole project is the engineering effort that keeps the GPUs fed.
Inference economics: utilization beats rate
Training campaigns end, inference runs forever, and the economics differ accordingly. Production inference is a steady state cost that scales with traffic, which makes it the natural home for committed capacity at negotiated rates, and it rewards engineering that raises throughput per GPU: batching, quantization, right sized model variants, and honest latency budgets. A team serving a model on four A10s at 30 percent utilization is paying for one busy GPU and three expensive heaters, and the fix is consolidation before procurement. The smaller GPU classes matter here too: not every inference workload earns H100 silicon, and the A10 class at around two dollars per GPU hour serves a large share of production traffic at a fraction of the flagship rate. The right pattern for most estates is a small committed inference floor plus burst capacity, the same structure as the Annual Flex versus pay as you go decision in our commitment comparison.
A framework for buying GPU capacity on OCI
- Size from the workload, not the ambition. Estimate training and inference hours from actual model plans, benchmark runs, and growth assumptions you would defend to a CFO, not from the largest cluster the team can imagine wanting.
- Separate steady inference from bursty training. Inference that runs all day suits committed capacity, training campaigns suit reservations with defined windows. Mixing them in one commitment oversizes both.
- Compare negotiated, not list. Get the real Universal Credits number for the OCI side and the equivalent committed discounts from the incumbents before declaring a winner.
- Price the data path. Add storage, cluster scratch, and the egress profile of your pipeline to the model. For multicloud pipelines the egress line can outweigh the rate difference.
- Negotiate capacity, not just price. Region, start date, guaranteed quota, and flexibility to rebalance between shapes belong in the agreement alongside the discount percentage.
- Instrument utilization from day one. Track GPU hours used against hours paid weekly. The 40 percent average spend reduction we find across OCI estates has its GPU equivalent in idle cluster time nobody was watching.
Commitments, reservations, and the credits angle
GPU spend changes the shape of an OCI commercial agreement. A modest infrastructure estate negotiates Universal Credits in the hundreds of thousands, while a serious AI program can put GPU consumption an order of magnitude above everything else in the tenancy, which moves the negotiation into territory where Oracle will discuss custom terms: deeper discounts, guaranteed capacity in named regions, and start dates aligned to hardware delivery. Three practical points matter. First, keep training campaigns and steady consumption in view separately when sizing the commitment, because a credits number sized to a temporary campaign becomes an expiry problem the quarter after the campaign ends. Second, capacity reservations bill whether used or not, so a reservation is a utilization promise your delivery plan has to keep. Third, the BYOL conversations that dominate database negotiations have no GPU equivalent, but the same independent scrutiny of contract terms pays for itself at these spend levels, and walking into the negotiation with a defensible consumption model is the single strongest lever you hold.
When OCI wins and when it does not
OCI is strongest for sustained training and heavy inference on big NVIDIA silicon, where its list rates, free RDMA fabric, and egress pricing compound into a total that the other platforms struggle to match, and Oracle has been willing to negotiate hard for exactly these workloads. It is weakest where a team is deeply invested in a competitor's managed AI stack and the GPU rental is secondary to the platform services around it, or where tiny intermittent inference makes the cheapest small instance on any cloud the right answer. The honest comparison is always the full pipeline cost at negotiated rates against the integration cost of the move, which is precisely the modeling exercise we run, on a fixed project fee, inside an OCI assessment. For estates already running GPUs on OCI, our optimization practice reviews utilization and commitment fit with the fee paid only on verified savings.
Bringing it together
On the published numbers, OCI rents H100 and A100 capacity for 20 to 40 percent less than AWS and Azure, charges nothing extra for the cluster fabric, and lets the data leave almost free. Those are real advantages and they survive contact with negotiation. What the rate card does not protect you from is buying the wrong amount: GPU spend punishes oversizing harder than any other cloud line because the unit prices are so high. Measure the demand, commit deliberately, watch utilization, and the OCI rates turn into the savings they promise. Our team plans, builds, and runs GPU infrastructure on OCI through our implementation practice, from quota strategy to cluster go live.
Part of a series
This guide is part of Data & AI on OCI — our complete pillar guide on the topic.
Moving Oracle workloads to OCI, or already running on OCI and not sure the architecture or the spend is right? Most teams bring in a specialist before they commit to a region, a shape, or a Universal Credits number. OCISpecialists.com plans the landing zone, runs the migration, and manages the estate after go live, on a fixed project fee, a managed monthly retainer, or a cost optimization fee paid only on verified savings.