Data estates do not fail loudly. They fail through small, accumulating uncertainty: nobody is sure which of three customer tables is current, the column called STATUS_2 means something only its creator remembers, and the quarterly revenue figure traces back through four systems to a spreadsheet. Each uncertainty costs a meeting, a Slack thread, or a wrong decision, and the costs compound as the estate grows. The fix is not a bigger warehouse. It is metadata: a searchable, trustworthy answer to what data exists, what it means, where it came from, and who owns it.
OCI Data Catalog is the metadata layer of the OCI data platform, and it is included with an OCI tenancy at no separate charge, which changes the adoption calculus completely. The catalogs that fail in the market are usually six figure platforms bought before anyone defined who would maintain them. With Data Catalog the platform cost is zero and the entire investment is organisational discipline, which is clarifying: there is no licence to justify, only the work itself. This article covers what the service does, where its edges are, and a rollout sequence that has survived contact with real organisations.
What the catalog actually holds
Data Catalog manages three layers of knowledge. The first is technical metadata, harvested automatically: schemas, tables, views, columns, datatypes, and files, pulled from data assets you register. Harvesters cover the OCI native estate, Autonomous Database, Object Storage, MySQL and HeatWave, plus on premises Oracle databases, Kafka schemas, and other common sources through connectors. Harvesting runs on a schedule, so the catalog tracks the estate as it changes rather than documenting a moment that has already passed.
The second layer is business meaning. A glossary holds the organisation's terms, customer, active subscription, net revenue, each with a definition, an owner, and a status. Terms link to the technical entities that implement them, which is the connection that makes a catalog useful: the analyst who finds the term finds the table, and the engineer who finds the table finds what it is supposed to mean. Custom properties extend entities with whatever your governance needs, data owner, sensitivity classification, retention class, and those properties become searchable facets.
The third layer is context: search across everything, tags for informal organisation, and lineage showing how data moves between assets, populated through integration with OCI services and through the API for tools the harvesters do not reach.
Lineage, honestly described
Lineage is where every catalog conversation needs honesty. The promise, every number traceable to its sources through every transformation, is only ever delivered in proportion to how much of the pipeline estate reports its movements. Data Catalog captures lineage from integrated OCI services, and OCI Data Integration publishes task level lineage as pipelines run, which covers the managed ETL layer well. Movement that happens outside instrumented tools, hand written scripts, third party ETL, database links, needs to be registered through the catalog's API or it simply does not appear.
The practical consequence: lineage coverage is an architectural choice, not a product feature. Estates that standardise movement on instrumented services get lineage nearly free. Estates with a script swamp get a lineage swamp, which is one more argument for the consolidation we recommend in the ETL layer anyway.
Where Data Catalog sits among the options
| Approach | Cost shape | Coverage | Fails when |
|---|---|---|---|
| OCI Data Catalog | Included with the tenancy | OCI native estate plus common sources | Nobody owns the glossary |
| Enterprise catalog platform | Six figure licence plus implementation | Broad multicloud connectors | Bought before the operating model exists |
| Wiki and spreadsheets | Free, allegedly | Whatever someone last updated | Immediately, silently |
| No catalog | Paid in meetings and rework | Tribal memory | The first departure or audit |
The independent read: for an estate centred on OCI, start with Data Catalog because it is included and harvests the platform natively, and graduate to an enterprise platform only when a genuinely multicloud catalog requirement materialises and is funded with an operating model attached. The wiki option is listed because every organisation has tried it, and the failure mode is always the same: documentation that decays the day it is written, because nothing reconciles it against reality. Harvesting is the feature that separates a catalog from a wiki.
The metadata that pays first
Not all metadata earns equally, and rollouts that try to document everything document nothing well. Three uses pay fastest. Sensitive data location: tagging which columns hold personal or regulated data, because every privacy request and audit starts with that map, and it is also the input the lakehouse security model needs, as we cover in building a lakehouse on OCI. Certified sources: marking the blessed table for each core entity so the four customer tables question has an official answer. And impact analysis: when a schema change is proposed, lineage answers who breaks downstream, which converts migration planning from archaeology into a query. If the catalog also feeds discovery for data products shared outward through Delta Sharing, the same metadata does double duty as a product description.
A six step rollout that avoids shelfware
- Register and harvest the core estate first. Autonomous instances, the lakehouse buckets, HeatWave, the warehouse. Scheduled harvests, not one off scans. This populates the catalog before anyone is asked to contribute.
- Appoint owners before authors. Every glossary domain gets a named owner whose job includes keeping it current. A catalog without owners is a wiki with a better logo.
- Start the glossary with twenty terms. The entities executives argue about: customer, order, revenue, active. Define, link to implementing tables, mark certified. Resist the urge to boil the dictionary.
- Tag sensitive data systematically. Work through the harvested schemas and classify personal and regulated columns with custom properties. This is the audit asset, build it deliberately.
- Instrument lineage where movement happens. Standardise pipelines on services that publish lineage, and register the stragglers through the API. Accept that coverage follows consolidation.
- Wire the catalog into onboarding. New analysts start in the catalog, not in a handover document. If onboarding does not use it, nothing else will keep it alive.
Limits worth naming
The service is a metadata catalog, not a governance suite. It records and publishes; it does not enforce. Access control stays in IAM and the databases, quality rules stay in the pipelines, and masking stays in the data platform, with the catalog telling you where to point each of them. Harvester coverage is strongest for the Oracle and OCI native estate, and a deeply multicloud organisation will find sources that need API work or a broader platform. And the glossary, like every glossary since the beginning of time, decays without an owner. None of this is a defect, but each one belongs in the plan rather than in the postmortem.
Bringing it together
OCI Data Catalog makes the cheapest high leverage move available to most OCI estates: it turns the questions every team loses hours to, what exists, what it means, where it came from, into queries against a maintained service that costs nothing extra to run. The platform work is days; the organisational work, owners, glossary, classification, is the real project and the real payoff. Our consulting practice runs catalog and governance rollouts as fixed scope engagements, usually alongside lakehouse builds where the metadata layer belongs in the design from day one, and an assessment is the fastest way to see whether your estate's problem is missing platform or missing discipline.
Free white paper
Go deeper on this topic with The Exadata Cloud Decision Guide, Database Service vs Cloud@Customer vs Autonomous, and how to choose. An independent analyst style report with comparison tables and recommendations, free with a work email. Prefer a monthly summary instead? The OCI Brief delivers one practical OCI briefing a month.
Part of a series
This guide is part of Data & AI on OCI — our complete pillar guide on the topic.
Moving Oracle workloads to OCI, or already running on OCI and not sure the architecture or the spend is right? Most teams bring in a specialist before they commit to a region, a shape, or a Universal Credits number. OCISpecialists.com plans the landing zone, runs the migration, and manages the estate after go live, on a fixed project fee, a managed monthly retainer, or a cost optimization fee paid only on verified savings.