Home  /  Journal  /  Advanced OCI Operations
Advanced OCI Operations

Advanced OCI Operations: Upgrades, Governance, and Scale

Migration gets the headlines, the budget, and the steering committee. Then go live arrives, the project team disbands, and the estate begins the part that actually determines whether OCI was a good decision: the next five years. Operating an OCI estate well is a different discipline from building one, with its own tooling, its own failure modes, and its own compounding returns. This pillar maps the whole territory.

Published Jun 6, 2026 · By Morten Andersen · 16 min read · Independent OCI advisory
Network switch with connected cables in a data center aisle

There is a moment, somewhere between three and twelve months after go live, when every OCI estate chooses its trajectory without anyone formally deciding anything. Either the practices that keep an estate healthy get built, patching on a calendar, costs allocated to owners, capacity watched ahead of demand, recovery rehearsed before it is needed, or they do not, and the estate begins the slow accumulation of debt that ends in a panicked upgrade project, an audit finding, or an invoice nobody can explain. The difference between those trajectories is not talent or budget. It is whether operations was treated as a designed system or as whatever was left over when the migration finished.

This article is the pillar for our advanced operations series: sixteen deep dives, each owning one practice area, organized here into the five domains that together make up a complete operating model. It is written for the team that already runs workloads on OCI and suspects, correctly, that running them is not the same as operating them. For an honest answer to where your estate sits today, the companion piece, an operations maturity model for OCI estates, turns this map into a scored assessment.

Day two is the real project

The economics explain why operations deserves this much attention. A migration is a one time cost; operations is a permanent one. Over a typical five year horizon, the run cost of an estate exceeds its migration cost several times over, and the variance is enormous: two estates with identical workloads can differ by 40 percent in spend purely on operational discipline, which is precisely the average reduction our optimization engagements verify. The same multiplier logic applies to risk. A weak migration produces one bad weekend. A weak patching practice produces an exposure window that renews itself every quarter, and a backup regime that was never tested produces its bill exactly once, all at once.

Cloud operations also punishes a specific habit imported from on premises estates: the belief that infrastructure, once built, stays built. On OCI, the estate changes under you. Services gain capabilities quarterly, database versions enter and leave support, compliance baselines move, and every team with console access is one click away from creating something ungoverned. Operations on OCI is therefore less like maintaining a building and more like gardening: continuous, seasonal, and unforgiving of neglect. The five domains below are the beds that need tending.

Domain one: lifecycle, upgrades, and patching

Everything running on the estate is on a clock. Database versions age toward the end of support, operating systems accumulate vulnerabilities, and Exadata and Base Database services publish quarterly update trains that estates either ride or fall behind. The teams that handle this well share one habit: they treat lifecycle as a standing calendar rather than a sequence of emergencies.

The largest single item on that calendar for most Oracle estates right now is the move to Oracle Database 23ai, the long support release that the whole estate eventually standardizes on. The planning question is not whether but when and in what order, and the answer differs sharply by service: what is nearly automatic on Autonomous Database is a planned project on Base Database and Exadata, and a genuine migration for anything still sitting on older releases. We cover the sequencing, the prerequisite checks, and the rollback planning in planning the Oracle 23ai upgrade on OCI.

Beneath the version question sits the recurring one: patching. A defensible practice needs three properties, a defined cadence tied to Oracle's quarterly critical patch updates, an environment progression that proves every patch in nonproduction before production sees it, and records good enough to survive an auditor's sampling. The design that achieves all three without consuming a team is laid out in a patching strategy for OCI that survives audits. For the operating system layer at fleet scale, OCI ships a purpose built service that turns per server patching into policy driven group operations, covered in OS Management Hub: fleet patching for compute. And for estates running many databases, the same industrialization logic applies one level up: grouping, standardizing, and patching databases as fleets rather than individuals, the subject of managing a database fleet on OCI with fleet ops tools.

Domain two: structure, compartments, and tenancies

Every governance control on OCI, policy, budget, quota, tag default, attaches to a structural boundary, which means estates with weak structure cannot govern no matter how much tooling they buy. The foundational unit is the compartment, and compartment design is one of the few decisions on OCI that is genuinely expensive to change later: resources can move, but policies, budgets, and habits calcify around the original tree. The patterns that survive growth, by environment, by workload, by team, and the hybrids that real estates actually use, are compared in compartment design patterns for growing tenancies.

Past a certain size, one tenancy stops being enough. Mergers, regulated subsidiaries, true production isolation, and chargeback boundaries that finance refuses to blur all push estates toward multiple tenancies under a single commercial umbrella, which OCI handles through Organizations, parent and child tenancies sharing governance and, where wanted, a single Universal Credits pool. When to split, what to centralize, and what the second tenancy really costs in duplicated operational surface are the subject of organizations and child tenancies: structuring at scale.

An estate that cannot say who owns a resource, what it costs, and when it was last patched is not operating its cloud. It is hosting an incident that has not happened yet.

Domain three: financial governance

Cloud spend is the most measurable thing in IT and still the least managed, because measurement without ownership produces reports nobody acts on. The chain that fixes this has four links, and each has its own article in this series. The first link is identity: every resource carries tags that say what it is, who owns it, and which cost center pays, enforced by tag defaults at compartment level rather than by memo, because memos do not survive contact with a Friday afternoon. The enforcement mechanics are in tagging governance on OCI: cost allocation that works.

The second link is attribution: turning tagged consumption into numbers each team sees and answers for, showback first, chargeback when the organization is ready, covered in cost allocation on OCI: making teams own their spend. The third is alarm: budgets at compartment and tag level with alert thresholds that fire while an overrun is still a conversation rather than a quarterly surprise, the configuration discipline described in OCI budgets and alerts: catching overruns early. The fourth is automation that removes the waste humans never get around to: scheduled stop and start for everything that does not need to run nights and weekends, which on flexible compute shapes converts directly into reclaimed spend, the use cases and limits of which are covered in OCI Resource Scheduler: stopping idle compute on a clock.

Domain four: headroom, quotas, and capacity

Two opposite failures live in this domain, and estates regularly manage to commit both at once. The first is hitting a ceiling nobody knew existed: service limits that block a scale out during a traffic spike, or a database provisioning request that fails at the worst possible hour because the tenancy quietly ran out of allowance. The defense is treating limits as monitored resources with owners and headroom thresholds, plus compartment quotas used deliberately as a governance brake rather than discovered accidentally as an outage. Both halves are covered in service limits and quotas on OCI: managing headroom.

The second failure is the opposite: paying for headroom that never gets used, the permanent fear premium of sizing every shape for the worst day of the year. The discipline that replaces fear with arithmetic, percentile based utilization baselines, growth trending, burstable and flexible shapes for spiky loads, and a review cadence that resizes against evidence, is the subject of capacity planning on OCI: shapes, limits, and burst. Together the two articles describe one practice seen from both sides: knowing how much room the estate has, in both directions, before the day it matters.

Domain five: resilience and recovery

Everything above optimizes the estate's good days. This domain decides what its worst day costs. The foundation is backup, and the standard to hold an estate to is unglamorous: one policy framework covering every data bearing service, retention matched to actual recovery and compliance needs rather than defaults, cross region copies for anything the business cannot lose with a region, and restores that get tested on a calendar, because an untested backup is a hypothesis with a deadline. The estate wide design is laid out in an estate wide backup strategy on OCI.

The second pillar is the human layer: what happens at 03:00 when the alarm fires and the engineer on call has never seen this failure before. The answer is either a runbook or improvisation, and only one of those has a predictable outcome. Runbooks that actually get used share a shape, short, accurate, tested in game days, and owned like code, described in incident runbooks for OCI: templates that get used.

The third pillar is truth: an estate defined in Terraform is only as trustworthy as the gap between the code and the console, and that gap, drift, grows silently every time a human fixes something by hand. Detecting drift continuously, deciding what to reconcile and what to codify, and keeping the pipeline authoritative are covered in drift detection on OCI: keeping Terraform honest. Resilience, in the end, is these three pillars agreeing with each other: backups that restore, runbooks that match reality, and code that matches the estate.

The substrate: monitoring and observability

One capability sits underneath all five domains without belonging to any of them: telemetry. Every practice in this series consumes it. Capacity planning is arithmetic over utilization history. Budgets alert on consumption streams. Patching verifies against fleet state. Runbooks navigate by dashboards. Drift detection compares observed against declared. An estate with thin monitoring can adopt every practice on this page and still operate blind, because each practice will be reasoning from data that does not exist.

The OCI native stack, Monitoring, Logging, Events, Alarms, and the database performance tooling above them, covers more of the need than most teams assume, and the gaps that matter are usually configuration rather than product: agents not deployed to half the fleet, retention windows too short to cover a quarterly cycle, alarms tuned once at go live and never against real incident history. The operational standard worth holding: every production resource emits metrics and logs somewhere queryable, retention covers at least one full business cycle, and every alarm that fires either drives an action or gets retired. Alert fatigue is not a monitoring failure, it is a governance failure wearing monitoring's clothes, and the quarterly review that prunes dead alarms is as much a part of the operating model as the patch calendar.

A year in the life of an operated estate

The domains become concrete when laid on a calendar, because an operating model is ultimately a schedule that keeps happening. Quarterly, anchored to Oracle's critical patch cycle: the patch waves move through their environment progression, the exception register gets re signed or closed, service limits get reviewed against growth, and the operations review re scores a slice of the maturity model. Monthly: cost allocation reports land with named owners, budget variances get explained rather than just noticed, the untagged resource report gets driven back to zero, and one restore test executes against a rotating schedule of systems. Weekly: drift reports get reconciled, capacity dashboards get a human glance, and the on call handover reviews whatever the alarms said. Annually: a full maturity re score, a game day that exercises the runbooks against a simulated failure, the backup policy review against changed retention obligations, and the version upgrade planning that keeps the estate ahead of support clocks rather than behind them.

Written out, the calendar looks heavy. Run properly, it is the opposite: each item is small precisely because it recurs, and the recurrence is what keeps any single instance from mattering much. The estates that find operations crushing are the ones doing all of it as recovery instead of routine, twelve quarters of deferred patching as one emergency program, three years of unallocated cost as one forensic project. The calendar is not the burden. The calendar is what the burden looks like after it has been divided by twelve and given owners.

The failure patterns worth naming

Across hundreds of estate assessments, the same operational failure modes recur often enough to deserve names. The migration hangover: the estate still shaped by go live decisions years later, rehearsal environments running, fear sized shapes unrevisited, migration era tags fossilizing in the cost report. The hero dependency: operations that work only because one person knows everything, undocumented, untested, and one resignation away from being an incident. The console drift spiral: infrastructure as code abandoned in practice because reconciling drift got harder than clicking, so every month the code describes less of the estate and the next automation project starts from zero. The governance theater: policies written, standards published, and nothing enforced by the platform, so compliance decays at the speed of staff turnover. And the watermelon dashboard: green on the outside, red inside, metrics chosen because they stay green rather than because they predict failure.

Every one of these has a structural cure somewhere in this series, and none of the cures is heroic. Defaults instead of memos, calendars instead of campaigns, fleet policies instead of per server effort, and evidence instead of confidence. The pattern behind the patterns: operations fails when it depends on sustained human virtue, and succeeds when the platform makes the right behavior the path of least resistance.

The five domains at a glance

DomainThe ad hoc estateThe operated estateAnchor practices
Lifecycle and patchingPatches when frightened, upgrades when forcedQuarterly cadence, staged progression, audit ready records23ai planning, CPU calendar, OS Management Hub, fleet operations
StructureOne compartment, everything in itCompartment tree mirroring ownership, organizations at scaleCompartment patterns, child tenancies
Financial governanceOne invoice, no owners, quarterly surprisesEnforced tags, team level allocation, budgets that alert earlyTagging, cost allocation, budgets, scheduled stop and start
HeadroomLimits discovered during incidentsMonitored limits, deliberate quotas, evidence based sizingQuota governance, capacity baselines
ResilienceDefault backups, tribal knowledge, console driftTested restores, owned runbooks, authoritative TerraformEstate backup policy, runbooks, drift detection

Operations and the spend curve

The financial argument for the operating model deserves its own paragraph, because it is usually understated. An unoperated estate's spend curve only bends one way: resources accumulate, sizes never shrink, nonproduction never sleeps, and the invoice grows by accretion regardless of what the business underneath it is doing. An operated estate's curve tracks the workload, because the mechanisms that bend it downward, schedules, rightsizing reviews, lifecycle policies on storage, allocation pressure on owners, run continuously instead of episodically. The difference compounds. Two estates that start at the same monthly number diverge by double digit percentages within a year, and the gap is pure operations, no architecture, no negotiation, no migration required.

This is also why the optimization conversation and the operations conversation are one conversation a quarter apart. A one time optimization pass on an unoperated estate produces a satisfying cliff in the spend chart followed by a slow climb back, because the habits that created the waste are still running. The same pass on top of a working operating model produces a step down that holds, because the budgets, tags, schedules, and reviews catch regression while it is still small. The 40 percent average reduction our optimization work verifies is really two numbers: the cliff, which anyone can produce once, and the hold, which only the operating model produces. Buy the first without the second and the estate rents its savings instead of owning them.

The same compounding logic applies to risk, with worse arithmetic. Deferred patching, untested restores, and tribal runbooks do not accumulate linearly; they multiply, because real incidents exercise several weaknesses at once. The estate that is four patch cycles behind, with a restore last tested at go live and a runbook in one engineer's head, is not carrying three risks. It is carrying one large one with three fuses, and the operating model is the only instrument that defuses all three on a schedule.

Building the operating model: an eight step sequence

The sixteen practices in this series are not equally urgent, and trying to adopt all of them at once is the most reliable way to adopt none. The sequence below reflects how we build operating models in managed engagements: each step makes the next one cheaper.

  1. Score the estate against the maturity model. Thirty minutes with the maturity model produces the gap list everything else works from.
  2. Fix structure first. Compartment design and, where relevant, tenancy structure, because every later control attaches to these boundaries.
  3. Enforce tagging before measuring anything. Tag defaults at compartment level, a minimal required set, and a weekly untagged resource report driven to zero.
  4. Switch on the financial guardrails. Budgets with staged alerts on every compartment, cost allocation reports to named owners, and schedules that stop idle compute.
  5. Put lifecycle on a calendar. The quarterly patch cadence, the OS fleet policy, and the 23ai upgrade sequence, booked as recurring work rather than projects.
  6. Make recovery provable. Backup policies across every service, cross region for the systems that justify it, and a restore test calendar with named owners.
  7. Write the runbooks the incidents will need. Start with the five failure modes the estate has already had, test them in a game day, and keep them in version control.
  8. Close the loop with drift detection and quarterly review. Scheduled drift checks against Terraform, monitored service limits, and a quarterly operations review that re scores the maturity model and re plans the calendar.

Who runs all this

The honest staffing answer is that a complete operating model is more than most internal teams can carry alongside their actual jobs, and less than a full team's worth of work once it is built and automated. That mismatch is why the managed model exists. Some estates build the capability internally, usually the ones with platform engineering already in place. Many split the difference: the internal team owns decisions and direction, while a specialist partner carries the 24/7/365 watch, the patch calendar, the restore tests, and the quarterly reviews on a managed monthly retainer. Our OCI managed services practice runs exactly that model, built on 500+ OCI engagements and 20+ years of combined Oracle experience, with the optimization loop included because the same telemetry that operates the estate also finds the waste in it.

However the work is staffed, the test of an operating model is the same. Pick any resource in the tenancy at random and ask: who owns it, what does it cost per month, when was it last patched, when was its data last proven restorable, and does the code that defines it match what is actually running. An estate that answers all five in minutes is operated. An estate that answers none of them is hosting its next incident, and the only open question is the date. The sixteen articles in this series exist to make the first kind of estate the normal kind.

About the author

Morten Andersen, Co-founder of OCI Specialists — 20 years of enterprise IT experience in OCI migration, security, networking, and 24/7 operations. Full profile · LinkedIn

Moving Oracle workloads to OCI, or already running on OCI and not sure the architecture or the spend is right? Most teams bring in a specialist before they commit to a region, a shape, or a Universal Credits number. OCISpecialists.com plans the landing zone, runs the migration, and manages the estate after go live, on a fixed project fee, a managed monthly retainer, or a cost optimization fee paid only on verified savings.