There is a reason the optimization pass deserves its own line in the migration business case rather than being treated as an afterthought: it is where the business case actually comes true. The migration itself rarely saves money on day one. Sizing decisions get made months before go live, against incomplete data, by people whose careers depend on the system not being slow. The result is an estate provisioned for the worst day anyone could imagine, running on its quietest days at a fraction of capacity. The optimization pass is the deliberate, evidence based correction of those fear driven decisions, and this article catalogs what it typically finds, in rough order of money. It belongs to our series on what an OCI migration really costs, because a budget that ends at go live is a budget that overstates the long run cost of OCI by nearly half.
Why the pass waits three to six months
Optimizing too early is as wasteful as not optimizing at all. The estate needs to live through real load before its true shape is knowable: at least one month end close, ideally one quarter end, every batch calendar the business runs, and enough ordinary weeks to separate signal from migration noise. Three months is the practical minimum, six is better for estates with strong seasonality. The waiting period is not idle: it is when monitoring baselines accumulate, and the quality of the pass is exactly the quality of the data it stands on. Estates that skipped observability at landing zone time pay for it here, which is one more argument made in what a production landing zone costs for building it properly the first time.
The findings catalog, in order of money
Oversized compute and database shapes. The single largest finding, nearly every time. Databases sized at on premises core counts despite faster OCI processors, application tiers carried over one to one when half the fleet was already idle at the source, and headroom multipliers stacked on headroom multipliers. OCI makes the correction unusually cheap: flexible shapes resize OCPUs and memory independently, and scaling down is a configuration change, not a procurement cycle. Typical recovery: 10 to 20 percent of total spend on its own.
Leftover dual running. Migration machinery that never got switched off: replication instances still consuming, bastions and staging hosts still up, rehearsal environments still provisioned, and occasionally an entire source system still billing in a forgotten contract. We covered the prevention in budgeting dual running; the optimization pass is where the cure happens for estates that did not prevent it.
Nonproduction running like production. Development and test environments at production scale, running 24 hours when they are used for 10, on paid service tiers nobody chose deliberately. Scheduled stop and start alone reclaims most of the difference, and OCI bills stopped OCPUs on most compute shapes at zero.
Storage on the wrong tier. Block volumes at performance levels nobody tuned after cutover, backups retained on hot storage that belong in archive tiers, object storage buckets growing without lifecycle policies, and boot volumes orphaned from instances deleted months ago. Individually small, collectively 5 to 10 percent.
License and edition mismatches. BYOL entitlements applied where license included pricing was cheaper, or the reverse, Enterprise Edition features paid for where Standard would serve, and unused option packs still attached to consolidated databases. This finding is the one that crosses from infrastructure into contract territory, and it is where an independent licensing review earns its place alongside the technical pass.
Commitment shape mismatch. Universal Credits commitments negotiated before anyone knew the real run rate: estates underconsuming an annual commitment and donating the difference, or overconsuming into on demand rates. The corrected consumption profile from the passes above becomes the evidence for the next commercial negotiation, which is the right order, fix the estate first, then fix the contract.
| Finding | Typical share of savings | Effort to fix | Risk of fixing |
|---|---|---|---|
| Oversized shapes | 25 to 40 percent | Low, resize with monitoring guardrails | Low, reversible in minutes |
| Leftover dual running | 10 to 25 percent | Low, inventory and switch off | Low, with retention checks first |
| Nonproduction sprawl | 10 to 20 percent | Low to medium, scheduling and rightsizing | Minimal |
| Storage tiers and orphans | 10 to 15 percent | Medium, lifecycle policies and cleanup | Low, with restore tests |
| License and edition fit | 10 to 20 percent | Medium, contract analysis required | Compliance sensitive, use specialists |
| Commitment reshaping | 5 to 15 percent | Negotiation, not engineering | Commercial only |
The evidence base: what the pass stands on
Every finding in the catalog is only as defensible as the data behind it, and the difference between an optimization pass and an opinion is the evidence pack. Four datasets do the work. The cost and usage reports supply the money side: every resource, every hour, every tag, exportable from the tenancy and joinable to whatever finance wants to see. The monitoring service supplies the load side: CPU, memory where agents report it, IOPS, and network throughput, ideally at one minute granularity for the systems under review. The database services supply their own layer: AWR and performance hub history for Oracle workloads, which answer the question infrastructure metrics cannot, whether the database was busy doing useful work or busy waiting. And the resource inventory itself, pulled tenancy wide, supplies the denominator: what exists, where, tagged how, owned by whom.
Two habits make the evidence persuasive to the people who approve changes. Percentiles, not averages: a database averaging 20 percent CPU may peak at 95 percent every close cycle, and the pass that misses this gets exactly one production incident before losing its mandate. And seasonality coverage: the window must include the busiest period the business has, which is why the pass waits for a quarter end rather than extrapolating around one. Where monitoring coverage turns out to be thin, fixing it becomes finding zero of the pass, because every subsequent finding inherits its credibility from the telemetry underneath it.
Guardrails: optimizing without breaking anything
The fastest way to discredit an optimization program is a performance incident wearing its name, so the pass operates under rules. Changes move from least to most consequential: orphan cleanup and nonproduction scheduling first, where the blast radius is zero, then storage tiering, then production resizing, with the revenue critical systems last and slowest. Every resize carries a stated rollback, which on OCI flexible shapes is the same console action in reverse, and a monitoring watch window with named thresholds before the next change proceeds. Nothing gets deleted on the first sweep: suspected orphans get stopped, tagged for deletion with a date, and removed only after the estate has lived without them for a full business cycle including a close. The discipline costs a few weeks of calendar and buys the program its reputation, which is the asset everything later depends on.
The guardrails also define what the pass refuses to touch without specialist cover. License affecting changes, edition downgrades, option removal, BYOL conversions, can save the most and cost the most if mishandled, because the savings arrive monthly and a compliance finding arrives all at once. Those moves get planned with independent licensing advice and executed with the contract evidence already assembled.
Running the pass: a seven step framework
- Freeze a baseline month. Pick a representative billing month, export the full cost report, and tag every line to a system and an owner. Savings only count against a baseline both sides accept.
- Rank by spend, not by interest. Sort resources by monthly cost and work from the top. The top twenty resources usually hold more than half the opportunity.
- Pull 90 days of utilization for the top spenders. CPU, memory, IOPS, and throughput percentiles, not averages. Size to the 95th percentile plus agreed headroom, and write the headroom number down.
- Sweep for the undead. Unattached volumes, stopped but provisioned databases, idle load balancers, forgotten rehearsal environments, and anything tagged with the migration project code that still exists.
- Fix nonproduction on a clock. Schedules for stop and start, smaller shapes, lower tiers, and a standing rule that nonproduction never exceeds half of production cost without written justification.
- Hand licensing to specialists. BYOL versus license included, edition fit, and option usage get reviewed against the actual contracts, not against assumptions. Verified entitlement positions also strengthen the renewal negotiation that follows.
- Verify, then institutionalize. Compare the post change month against the baseline, attribute every delta, and convert one off fixes into standing policy: budgets, quotas, tagging rules, and a quarterly mini pass so the estate never drifts this far again.
A worked example: where 40 percent hides
A composite from real engagements makes the catalog concrete. An estate lands on OCI after a nine month migration: two Exadata racks, thirty Base Database systems, two hundred compute instances, and the usual long tail of storage, load balancers, and forgotten experiments. Monthly spend settles at 100 units. The pass begins in month five, after one quarter end has passed through the monitoring history.
The utilization review finds the application tier sized at source core counts: the two hundred instances run a 95th percentile CPU under 30 percent, and resizing the worst eighty of them recovers 9 units. The database review finds twelve of the thirty Base systems carrying OCPU counts justified by nothing in their AWR history; stepping them down recovers 8 units. The undead sweep finds the migration's own residue, rehearsal environments, two replication hosts, four hundred unattached block volumes, worth 7 units. Nonproduction scheduling, stopping development and test outside working hours, recovers 6 units. Storage tiering and backup retention fixes add 4. The licensing review converts three databases from license included to BYOL against entitlements already owned and drops an options pack nobody had enabled, worth 6 units. Total: 40 units off a 100 unit baseline, with no architecture changed and no application touched. The number is not a coincidence; it is what fear based sizing plus migration residue reliably adds up to, and it is why the 40 percent average across our engagements has stayed stable for years.
Equally instructive is what the example does not include: no reserved capacity gymnastics, no exotic re platforming, no risk. The savings are administrative competence applied with evidence, which is also why they stick, the corrected estate is simpler to operate than the bloated one was.
What the pass is worth, commercially
The arithmetic that matters to a CFO: on an estate spending consistently each month, a 40 percent reduction repays a structured optimization engagement many times over within the first year, and the savings recur every month afterward. This is also the engagement where pricing model alignment is cleanest. A fixed project fee suits a bounded, one time pass with a defined report and implementation window. A managed monthly retainer suits estates that want the pass institutionalized, with 24/7/365 monitoring feeding a continuous optimization loop instead of an annual cleanup. And the optimization fee model, a percentage of verified savings with no savings, no fee, removes the buyer's risk entirely: if the pass finds nothing, it costs nothing. Across 500+ OCI engagements and 20+ years of combined Oracle experience, we have yet to run a pass on a freshly migrated estate that found nothing. The honest uncertainty is never whether there are savings; it is which of the six catalog entries dominates.
The pass also closes the loop on the migration business case. The ROI model built before the move, covered in when OCI migration ROI turns positive, almost always assumes a post migration correction. Skipping the pass does not make the business case wrong; it makes it permanently unfinished, with the estate paying a fear premium every month for sizing decisions everyone already knows were conservative. Our OCI optimization practice exists for exactly this engagement, and the verified savings model means the burden of proof sits where it belongs, on us.
Land safe, then trim with evidence. That is the whole doctrine. The 40 percent is not magic and it is not waste in the shameful sense; it is the predictable, correctable residue of migrating responsibly. The only mistake is leaving it on the invoice.
Part of a series
This guide is part of OCI Migration — our complete pillar guide on the topic.
Moving Oracle workloads to OCI, or already running on OCI and not sure the architecture or the spend is right? Most teams bring in a specialist before they commit to a region, a shape, or a Universal Credits number. OCISpecialists.com plans the landing zone, runs the migration, and manages the estate after go live, on a fixed project fee, a managed monthly retainer, or a cost optimization fee paid only on verified savings.