Ask a team how good their OCI operations are and you will get adjectives. Ask them when a restore was last tested, how many resources carry no owner tag, and how many patch cycles behind production sits, and you will get the truth. The distance between the adjectives and the truth is what a maturity model measures, and the measurement matters because operational weakness is invisible until the day it is not. An estate can run for years on missing backups, dead runbooks, and a console drifted estate, and feel fine, because none of those gaps announces itself before the incident that exercises it.
This article gives you the scoring instrument for everything in our advanced OCI operations series: a model with four levels and five domains, designed to be scored in an afternoon with evidence rather than in a workshop with opinions. We use a version of it at the start of every managed engagement, and the pattern after 500+ OCI engagements is consistent enough to be a finding in itself: most estates believe they are a level higher than they are, and the overestimate concentrates exactly where it is most expensive, in resilience.
The four levels
The levels describe how work happens, not how much tooling exists, because tooling without practice scores zero in this model. Level one, reactive: work happens when something breaks or someone shouts, patching is episodic, cost is one number nobody owns, and recovery is a hope. Level two, managed: the basics have calendars and owners, patches flow quarterly, budgets alert, backups run, but practices live in individual heads and survive on individual virtue. Level three, systematic: the practices are enforced by the platform rather than by memory, tag defaults instead of memos, fleet policies instead of per server effort, drift detection instead of trust, and the calendar runs whether or not its inventor is on holiday. Level four, optimizing: the system measures itself, maturity is re scored on a cadence, incidents feed runbook revisions, utilization feeds rightsizing, and the operating model improves the way products do, by iteration against evidence.
Two properties of the ladder matter more than the labels. Levels cannot be skipped, because each one is the precondition of the next; an estate cannot automate practices it does not have, and cannot optimize a system that does not exist. And the step that changes an estate's character is the second one, from managed to systematic, because that is where operations stops depending on sustained human virtue, the single least reliable component in any architecture.
The five domains, and what each level looks like
The domains are the same five the pillar maps, and the table gives the scoring anchors for the two levels where most estates actually sit.
| Domain | Level two looks like | Level three looks like |
|---|---|---|
| Lifecycle and patching | Quarterly intent, frequent slips, records assembled on request | Patch waves on maintenance windows, prechecks automated, evidence falls out of the process |
| Structure | A compartment tree exists, exceptions accumulate | Structure mirrors ownership, policy and budgets attach cleanly, new work lands in the right place by default |
| Financial governance | Budgets alert, tagging mostly happens, reports are read | Tags enforced by defaults, team level allocation with named owners, idle compute stopped on schedules |
| Headroom | Limits checked when something fails, sizing by fear | Limits monitored with thresholds, sizing reviewed against percentile evidence on a cadence |
| Resilience | Backups run, restores assumed, runbooks exist somewhere | Restore tests on a calendar, runbooks tested in game days, drift detection keeping code authoritative |
Scoring needs evidence, and the right evidence is a short list of artifacts: the last patch cycle's dates against its plan, the untagged resource count, the most recent restore test report, the drift report and what happened to its findings, the budget alerts from last quarter and what each one triggered. A domain scores the level its artifacts prove, not the level its slide deck claims. Where the artifacts do not exist, the score is level one by definition, which is the model's bluntest and most useful rule.
Running the assessment
- Collect artifacts before opinions. One person gathers the evidence list above for each domain, half a day, no meetings.
- Score each domain to its proven level. Be harsh at the boundaries: a practice that depends on one person is level two no matter how well it runs.
- Write the gap sentence for every domain below three. What fails, who feels it, and roughly what it costs when it does, one sentence each, because ranked pain beats abstract scores.
- Pick the two cheapest level raising moves. Usually tag defaults and a restore test calendar; both are days of work and both convert level two domains to demonstrable progress.
- Book the re score. Quarterly, same artifacts, same harshness, trend recorded, because a maturity score measured once is trivia and measured eight times is a management instrument.
Reading your profile
The pattern of scores carries more information than the average. A profile that is strong in financial governance and weak in resilience, the most common shape we find, describes an estate that optimized what the CFO can see and deferred what only an incident reveals; the correction order is laid out in an estate wide backup strategy on OCI and incident runbooks for OCI. The opposite profile, careful resilience over chaotic cost, usually marks an estate run by infrastructure veterans who have not yet been handed the invoice. A flat level two everywhere is the signature of a capable team at its capacity ceiling, doing everything by hand and one resignation away from level one, and the cure is the enforcement step, not more effort: the platform mechanics in a patching strategy that survives audits and drift detection on OCI are where hands become systems.
The substrate deserves its own glance during any reading: telemetry. A domain cannot score above level two if the data that would prove level three does not exist, and thin monitoring quietly caps every other score on the sheet. An estate that cannot produce utilization history cannot do evidence based sizing, an estate without complete audit logs cannot prove its patch records, and an estate whose alarms were tuned once at go live is navigating its incidents by folklore. When several domains feel stuck at the same level, the shared cause is usually underneath them, in monitoring coverage and retention, and fixing the substrate moves more scores than attacking any single domain head on.
One more reading rule: weight the domains by what your estate actually is. A reporting estate that could tolerate a day of downtime can live with level two resilience longer than a trading platform can survive level two anything. The model is deliberately generic; the ranking of its gaps is deliberately yours.
The scoring mistakes that flatter an estate
Self assessments fail in predictable ways, and naming the failure modes in advance is the cheapest calibration available. The first is scoring intentions: the patching policy that was approved last quarter counts for nothing until a patch cycle has actually run against it, and a model that accepts plans as evidence will report level three estates that have never tested a restore. The second is scoring the best example: every estate has one well run application whose artifacts are reached for in every review, and the honest score is the one the worst governed compartment earns, because incidents do not sample politely from the showcase. The third is letting the same person score and own the domain, which produces grades that drift upward a level per year without any underlying change; the fix is as simple as swapping scorers across domains, or letting the platform numbers, untagged counts, patch lag, alert volumes, do the arguing.
The fourth mistake is subtler: treating the score as the deliverable. A maturity number that goes into a slide and stops there has consumed an afternoon and produced decoration. The deliverable is the gap list with owners and dates, and the test of whether the assessment worked is visible one quarter later, in whether the named gaps moved. Estates that re score quarterly and publish the trend get a useful side effect for free, a defensible answer to the budget question of what the operations investment actually bought, expressed in levels climbed rather than adjectives.
The climb, honestly priced
Moving from level one to two is mostly calendars and ownership, weeks of effort, and it removes the most embarrassing failure modes. Moving from two to three is the real project, a quarter or two of converting habits into platform enforcement, and it is where the compounding starts: enforced practices stop decaying, which means every later investment builds on ground that holds. Moving from three to four is not a project at all but a cadence, the quarterly re score, the game days, the review loop, run indefinitely. The economics favor the climb at every step, the same arithmetic the pillar lays out: double digit spend differences and multiplied risk between operated and unoperated estates, with the 40 percent average reduction our optimization work verifies sitting mostly in the gap between levels one and three.
The staffing question has the usual honest answer. Reaching level three is well within a good internal team's ability if the time is protected, and protecting the time is precisely what fails most often. The split that works in practice: the internal team owns the decisions, the standards, and the scores, while a specialist partner carries the recurring load, the 24/7/365 watch, the patch trains, the restore calendar, the quarterly re score, on a managed monthly retainer. That is the operating model our OCI managed services practice runs, and the maturity assessment above is, not coincidentally, the first artifact of every engagement. Score yourself first either way. The afternoon it costs is the cheapest operational investment available this quarter, and the gap list it produces will be more specific than any general article, this one included, could ever be.
Part of a series
This guide is part of OCI Operations & Observability — our complete pillar guide on the topic.
Moving Oracle workloads to OCI, or already running on OCI and not sure the architecture or the spend is right? Most teams bring in a specialist before they commit to a region, a shape, or a Universal Credits number. OCISpecialists.com plans the landing zone, runs the migration, and manages the estate after go live, on a fixed project fee, a managed monthly retainer, or a cost optimization fee paid only on verified savings.