A managed services SLA is a strange document. It is negotiated when everything is friendly and invoked when everything is on fire, and the gap between those two moods is where most of its problems live. During the sales cycle the SLA reads like reassurance: response times in minutes, availability with several nines, a credit regime that sounds like teeth. During a Severity 1 incident the same document reads very differently, and phrases that seemed harmless in procurement turn out to carry most of the weight. Whether the provider owes you an engineer in fifteen minutes or a voicemail by Tuesday often comes down to a single definition on page nine.
This article is part of our complete guide to hiring an OCI partner, and it deals with one document: the service level agreement that sits inside an OCI managed services contract. The goal is not to teach you to write an SLA from scratch. It is to show you what a fair one contains, where the traps usually sit, and what to push on before you sign, because after signature the leverage belongs entirely to the other side.
What an OCI managed services SLA actually covers
The first confusion to clear up is the boundary between Oracle's responsibilities and your provider's. Oracle publishes its own SLAs for OCI covering availability, manageability, and performance of the platform itself: the regions, the compute, the storage, the network, the Autonomous services. Those are Oracle's commitments, they pay out in Oracle credits, and your managed services provider neither controls them nor underwrites them. What the provider commits to is everything that happens on top of the platform: monitoring your estate, responding when something breaks, patching, backups, recovery, capacity, security operations, and the routine changes that keep a production environment alive.
A surprising number of provider SLAs blur this line on purpose. They quote Oracle's platform availability figures in their own marketing, then write their contractual commitments so that any incident touching Oracle infrastructure is excluded from their obligations. The result is a document that sounds like a guarantee of your application being up and is actually a guarantee of nothing in particular. A well drafted SLA separates the two layers explicitly: Oracle covers the platform, the provider covers operations on the platform, and the provider's commitments are written in terms of its own actions, response, restoration effort, escalation, communication, because those are the things it can actually control.
So the core of a usable OCI managed services SLA is a short list: severity definitions, response and resolution commitments per severity, coverage hours, an operational availability or restoration target for the services the provider manages, measurement and reporting rules, credits and remedies, and an escalation path with named roles. Everything else is decoration. If any item on that list is missing or vague, that is where your first incident will go wrong.
Severity definitions decide everything downstream
Every commitment in the SLA hangs off the severity ladder, so the definitions of Severity 1 through Severity 4 are the most important paragraphs in the document. A Severity 1 response commitment of fifteen minutes is worthless if the provider alone decides what qualifies as Severity 1.
Good definitions are written in terms of business impact, not technical symptoms. Severity 1 should mean a production service is down or unusable for a material part of the business, full stop. Severity 2 should mean production is degraded or a critical function is impaired with no acceptable workaround. Severity 3 covers degraded non production or minor production issues with workarounds, and Severity 4 covers questions and routine requests. The two clauses to insist on: you, the customer, set the initial severity when you raise the ticket, and any downgrade by the provider requires your agreement, with the clock running at the original severity until you agree. Without those two sentences, every awkward incident becomes a negotiation about classification while your application stays down.
Response time vs resolution time, and what is fair to ask
Response and resolution are different promises and providers price them very differently. Response time is how quickly a qualified human engages with the incident. It is fully within the provider's control, so it should be a hard commitment with credits attached. Resolution time is how quickly the incident is fixed, and because root causes can sit with Oracle, with your application vendor, or inside your own code, almost no serious provider will guarantee resolution as a hard SLA. What you can and should get is a resolution objective per severity, continuous effort language for Severity 1, meaning the provider works the incident around the clock without being asked, and a commitment to formal escalation when the objective is breached.
Here is the shape of a fair severity matrix for a production OCI estate under a 24/7/365 managed service. Treat it as a benchmark to negotiate against rather than a template to copy.
| Severity | Definition | Response commitment | Resolution objective | Coverage |
|---|---|---|---|---|
| Severity 1 | Production down or unusable, material business impact | 15 minutes, engineer engaged | 4 hours, continuous effort until restored | 24/7/365 |
| Severity 2 | Production degraded, critical function impaired, no workaround | 30 minutes | 8 business hours | 24/7/365 |
| Severity 3 | Minor production impact or non production issue, workaround exists | 4 business hours | 3 business days | Business hours |
| Severity 4 | Questions, requests, cosmetic issues | 1 business day | Scheduled by agreement | Business hours |
Notice the coverage column. Whether Severity 1 and 2 are covered around the clock or only during business hours is the single biggest price lever in the whole agreement, and it is also where the cheapest quotes quietly cut. The support tier you buy determines the matrix you get, and we break the tiers down in detail in OCI support tiers explained. If your estate runs revenue generating workloads, the around the clock tier is rarely optional, and a provider who prices it suspiciously low is usually planning to staff it with a pager and a prayer.
Availability targets on top of Oracle's platform
Providers love to promise availability percentages because they sound strong and measure weak. Be precise about what the number means. Oracle already commits to platform availability for most OCI services, often at 99.9 percent or better depending on the service and configuration. Your provider cannot sensibly re sell that number, because Oracle outages are outside its control. What a provider can commit to is operational availability of the things it manages: that monitoring is running and alerting correctly, that backups complete and are tested, that patching happens inside agreed windows, and that when something under its control fails, restoration starts within the committed response time.
The honest construction is therefore layered. Architect the estate for the availability your business needs, using multiple availability domains or regions where justified, then hold Oracle to its platform SLAs and the provider to its operational SLAs. If a provider offers a single end to end availability number for your application, read the exclusions, because the list of carve outs usually hands back everything the headline promised. The exception is where the provider also designed and built the architecture under a fixed fee project. In that case it owns the design assumptions, and a stronger end to end commitment is reasonable to demand.
Measurement, reporting, and the percentile problem
How performance is measured matters as much as what is promised. Insist on three things. First, the measurement source must be the ticket system and monitoring data you can see yourself, not a quarterly summary the provider compiles privately. Second, reporting must be monthly, in writing, against every committed metric, with raw data available on request. Third, and most important, percentile measurement rather than averages.
The average trick deserves a paragraph because it is everywhere. A provider that answers nine incidents in five minutes and one in five hours reports an average response of about 34 minutes, which passes a 60 minute SLA easily, while the one incident that mattered, the Saturday night Severity 1, was effectively abandoned. Averages let a provider be excellent when it is convenient and absent when it is not. The fix is to write commitments at the 95th percentile or higher: 95 percent of Severity 1 incidents get a response within 15 minutes, measured monthly, with every breach listed in the report. Percentiles make the worst cases count, and the worst cases are the entire point of the agreement.
Credits, remedies, and what teeth actually look like
Service credits are the standard remedy and they are weaker than they look. Typical regimes cap credits at 10 to 30 percent of the monthly fee, which on a managed monthly retainer means a bad month of service might earn you back a few thousand dollars against an outage that cost you fifty times that. Credits are not compensation. They are a signaling mechanism that makes the provider's failures visible and slightly painful.
So negotiate credits, but negotiate the structure around them harder. Credits should be applied automatically when the monthly report shows a breach, not claimed by you within some short window, because claim deadlines exist to make credits expire quietly. Repeated failure should trigger escalating remedies: a remediation plan after one bad month, executive review after two, and a termination right for persistent breach after three, with no early exit penalty. That last clause is the one with real teeth, and it connects directly to the exit terms we cover in exit clauses and vendor lock in for OCI service contracts. A provider who resists a termination right for chronic SLA failure is telling you how confident it is in its own operation.
Common provider tricks to catch before signature
Most SLA games are variations on three moves, and once you know them they are easy to spot.
Vague severity definitions
Watch for definitions that hinge on words like significant, substantial, or as determined by the provider. Each one moves classification power to the other side of the table. Watch equally for definitions written in technical terms, a host is unreachable, a service returns errors, because real incidents are messy and rarely match a checklist, which gives the provider room to argue every one down a level.
Business hours buried in the fine print
The proposal says 24/7 monitoring on page two and the SLA defines response commitments against business hours on page fourteen, sometimes in a time zone you do not operate in. Monitoring around the clock with humans available nine to five means alarms ring all night and nobody answers until morning. Check the coverage column of the severity matrix against the definitions section, then check which public holidays apply.
Averages instead of percentiles
Covered above, and worth repeating in one line: any SLA measured on mean response or mean resolution time is built to be passed while failing you when it matters. Demand percentiles.
A provider that runs these plays in the contract will run similar plays in delivery, and the contract is your earliest, cheapest signal. We catalogue the broader warning signs in OCI partner red flags, and a structured procurement process is your best defense, because tricks that survive a sales conversation rarely survive a written question with a deadline.
A negotiation framework that works
Run the negotiation in this order, and settle each step before moving to the next, because every later term depends on an earlier one.
- Classify your workloads first. Decide internally which systems justify Severity 1 treatment and around the clock coverage. Buying the top tier for everything is the most common form of overspend on a retainer.
- Fix the severity definitions. Business impact language, customer sets initial severity, downgrades require your agreement. Do not discuss times until definitions are closed.
- Set response commitments per severity. Hard numbers, measured at the 95th percentile from ticket data you can access, credits attached.
- Set resolution objectives and escalation. Continuous effort on Severity 1, formal escalation with named roles and time triggers when objectives slip.
- Agree measurement and reporting. Monthly written reports against every metric, your access to raw data, automatic credit application.
- Negotiate remedies with an endpoint. Escalating consequences for repeated breach, ending in termination for persistent failure without penalty.
- Schedule an annual review. Estates change. Build in a yearly recalibration of severities, coverage, and targets, with pricing adjusted both directions.
How the commercial model shapes the SLA
SLA terms and pricing model are two views of the same risk allocation, which is why they should be negotiated together rather than in sequence. On a managed monthly retainer, the SLA is the product: you are paying a fixed monthly fee precisely for the response commitments and coverage hours described above, and the retainer price should move visibly when the severity matrix moves. What a fair retainer costs, and how scope drives the number, is covered in OCI managed services pricing. On a fixed project fee, for the migration or landing zone build that usually precedes steady state operations, the equivalent of the SLA is the acceptance criteria and the warranty period, and the discipline is the same: define done in measurable terms before the price is agreed. And on an optimization model where the fee is a percentage of verified savings, the SLA question becomes the verification method, because no savings means no fee, so the measurement rules are the contract. Across our own engagements that model has produced a 40 percent average reduction in OCI spend, and it works because the measurement was agreed before the work started, which is exactly the habit this whole article is recommending.
One closing observation from 500+ OCI engagements: the providers most willing to sign sharp SLAs are usually the ones who least often trigger them. Precise commitments require a real 24/7/365 operation, honest monitoring, and confidence built on experience, and a provider who has all three has little reason to hide behind vague definitions. That is the standard our own OCI managed services practice is built to, with 20+ years of combined Oracle experience behind the escalation path. Use the framework above, hold every bidder to the same matrix, and the SLA stops being fine print and starts being what it should have been all along: the clearest possible description of what you are actually buying.
Free white paper
Go deeper on this topic with The OCI Pricing Decoder, Universal Credits, Support Rewards, and the discounts Oracle does not volunteer. An independent analyst style report with comparison tables and recommendations, free with a work email. Prefer a monthly summary instead? The OCI Brief delivers one practical OCI briefing a month.
Part of a series
This guide is part of OCI Operations & Observability — our complete pillar guide on the topic.
Moving Oracle workloads to OCI, or already running on OCI and not sure the architecture or the spend is right? Most teams bring in a specialist before they commit to a region, a shape, or a Universal Credits number. OCISpecialists.com plans the landing zone, runs the migration, and manages the estate after go live, on a fixed project fee, a managed monthly retainer, or a cost optimization fee paid only on verified savings.