Home  /  Journal  /  OCI Data Platform  /  OCI Data Catalog
OCI Data Platform

OCI Data Catalog: Metadata and Lineage for the Estate

Every data estate eventually fails the same simple test: a new analyst asks where the revenue number comes from, and four people give four answers. OCI Data Catalog exists to make that question answerable. Here is what the service harvests, how glossaries and lineage actually work, and how to roll it out so it becomes infrastructure instead of shelfware.

Published Jun 7, 2026 · By Fredrik Filipsson · 9 min read · Independent OCI advisory
Engineers collaborating at computer screens in an office

Data estates do not fail loudly. They fail through small, accumulating uncertainty: nobody is sure which of three customer tables is current, the column called STATUS_2 means something only its creator remembers, and the quarterly revenue figure traces back through four systems to a spreadsheet. Each uncertainty costs a meeting, a Slack thread, or a wrong decision, and the costs compound as the estate grows. The fix is not a bigger warehouse. It is metadata: a searchable, trustworthy answer to what data exists, what it means, where it came from, and who owns it.

OCI Data Catalog is the metadata layer of the OCI data platform, and it is included with an OCI tenancy at no separate charge, which changes the adoption calculus completely. The catalogs that fail in the market are usually six figure platforms bought before anyone defined who would maintain them. With Data Catalog the platform cost is zero and the entire investment is organisational discipline, which is clarifying: there is no licence to justify, only the work itself. This article covers what the service does, where its edges are, and a rollout sequence that has survived contact with real organisations.

What the catalog actually holds

Data Catalog manages three layers of knowledge. The first is technical metadata, harvested automatically: schemas, tables, views, columns, datatypes, and files, pulled from data assets you register. Harvesters cover the OCI native estate, Autonomous Database, Object Storage, MySQL and HeatWave, plus on premises Oracle databases, Kafka schemas, and other common sources through connectors. Harvesting runs on a schedule, so the catalog tracks the estate as it changes rather than documenting a moment that has already passed.

The second layer is business meaning. A glossary holds the organisation's terms, customer, active subscription, net revenue, each with a definition, an owner, and a status. Terms link to the technical entities that implement them, which is the connection that makes a catalog useful: the analyst who finds the term finds the table, and the engineer who finds the table finds what it is supposed to mean. Custom properties extend entities with whatever your governance needs, data owner, sensitivity classification, retention class, and those properties become searchable facets.

The third layer is context: search across everything, tags for informal organisation, and lineage showing how data moves between assets, populated through integration with OCI services and through the API for tools the harvesters do not reach.

Lineage, honestly described

Lineage is where every catalog conversation needs honesty. The promise, every number traceable to its sources through every transformation, is only ever delivered in proportion to how much of the pipeline estate reports its movements. Data Catalog captures lineage from integrated OCI services, and OCI Data Integration publishes task level lineage as pipelines run, which covers the managed ETL layer well. Movement that happens outside instrumented tools, hand written scripts, third party ETL, database links, needs to be registered through the catalog's API or it simply does not appear.

A catalog does not create truth, it publishes it. Harvesting keeps the technical layer honest automatically, but meaning and lineage are only as good as the process that maintains them.

The practical consequence: lineage coverage is an architectural choice, not a product feature. Estates that standardise movement on instrumented services get lineage nearly free. Estates with a script swamp get a lineage swamp, which is one more argument for the consolidation we recommend in the ETL layer anyway.

Where Data Catalog sits among the options

ApproachCost shapeCoverageFails when
OCI Data CatalogIncluded with the tenancyOCI native estate plus common sourcesNobody owns the glossary
Enterprise catalog platformSix figure licence plus implementationBroad multicloud connectorsBought before the operating model exists
Wiki and spreadsheetsFree, allegedlyWhatever someone last updatedImmediately, silently
No catalogPaid in meetings and reworkTribal memoryThe first departure or audit

The independent read: for an estate centred on OCI, start with Data Catalog because it is included and harvests the platform natively, and graduate to an enterprise platform only when a genuinely multicloud catalog requirement materialises and is funded with an operating model attached. The wiki option is listed because every organisation has tried it, and the failure mode is always the same: documentation that decays the day it is written, because nothing reconciles it against reality. Harvesting is the feature that separates a catalog from a wiki.

The metadata that pays first

Not all metadata earns equally, and rollouts that try to document everything document nothing well. Three uses pay fastest. Sensitive data location: tagging which columns hold personal or regulated data, because every privacy request and audit starts with that map, and it is also the input the lakehouse security model needs, as we cover in building a lakehouse on OCI. Certified sources: marking the blessed table for each core entity so the four customer tables question has an official answer. And impact analysis: when a schema change is proposed, lineage answers who breaks downstream, which converts migration planning from archaeology into a query. If the catalog also feeds discovery for data products shared outward through Delta Sharing, the same metadata does double duty as a product description.

A six step rollout that avoids shelfware

  1. Register and harvest the core estate first. Autonomous instances, the lakehouse buckets, HeatWave, the warehouse. Scheduled harvests, not one off scans. This populates the catalog before anyone is asked to contribute.
  2. Appoint owners before authors. Every glossary domain gets a named owner whose job includes keeping it current. A catalog without owners is a wiki with a better logo.
  3. Start the glossary with twenty terms. The entities executives argue about: customer, order, revenue, active. Define, link to implementing tables, mark certified. Resist the urge to boil the dictionary.
  4. Tag sensitive data systematically. Work through the harvested schemas and classify personal and regulated columns with custom properties. This is the audit asset, build it deliberately.
  5. Instrument lineage where movement happens. Standardise pipelines on services that publish lineage, and register the stragglers through the API. Accept that coverage follows consolidation.
  6. Wire the catalog into onboarding. New analysts start in the catalog, not in a handover document. If onboarding does not use it, nothing else will keep it alive.

Limits worth naming

The service is a metadata catalog, not a governance suite. It records and publishes; it does not enforce. Access control stays in IAM and the databases, quality rules stay in the pipelines, and masking stays in the data platform, with the catalog telling you where to point each of them. Harvester coverage is strongest for the Oracle and OCI native estate, and a deeply multicloud organisation will find sources that need API work or a broader platform. And the glossary, like every glossary since the beginning of time, decays without an owner. None of this is a defect, but each one belongs in the plan rather than in the postmortem.

Bringing it together

OCI Data Catalog makes the cheapest high leverage move available to most OCI estates: it turns the questions every team loses hours to, what exists, what it means, where it came from, into queries against a maintained service that costs nothing extra to run. The platform work is days; the organisational work, owners, glossary, classification, is the real project and the real payoff. Our consulting practice runs catalog and governance rollouts as fixed scope engagements, usually alongside lakehouse builds where the metadata layer belongs in the design from day one, and an assessment is the fastest way to see whether your estate's problem is missing platform or missing discipline.

Free white paper

Go deeper on this topic with The Exadata Cloud Decision Guide, Database Service vs Cloud@Customer vs Autonomous, and how to choose. An independent analyst style report with comparison tables and recommendations, free with a work email. Prefer a monthly summary instead? The OCI Brief delivers one practical OCI briefing a month.

Part of a series
This guide is part of Data & AI on OCI — our complete pillar guide on the topic.

About the author

Fredrik Filipsson, Co-founder of OCI Specialists — 20 years of enterprise IT experience in Oracle Database, OCI cost optimization, licensing, and data platforms. Full profile · LinkedIn

Moving Oracle workloads to OCI, or already running on OCI and not sure the architecture or the spend is right? Most teams bring in a specialist before they commit to a region, a shape, or a Universal Credits number. OCISpecialists.com plans the landing zone, runs the migration, and manages the estate after go live, on a fixed project fee, a managed monthly retainer, or a cost optimization fee paid only on verified savings.