Logiciel Contact Us
Success Stories Tech News Contact Us
framework

80% Of Governance Programmes Fail By 2027. The Catalog Decides That, Not The Format.

Teams spend three months arguing about Apache Iceberg against Delta Lake, then lose a year to a catalog they picked in an afternoon. This guide is the scoring session: nine weighted criteria with a straight choose-this-if on each side, a worksheet that totals to 110 per format, and the call for the three situations most estates are in.

In depth

A Forty-Row Feature Matrix Settles Nothing. The Catalog You Skipped Settles Everything.

01

The trap most data platform teams walked into: run the comparison as a feature grid, half of it true on both sides, add a benchmark published by a vendor selling one side and a 50GB proof of concept where everything looks fine, then find out eight months later that the catalog nobody held a meeting about owns table identity, permissions and who may commit, and that replacing it is far harder than replacing a file format.

02

What the teams who get this right do: count the engines that will read and write these tables in production over the next two years rather than the ones on the architecture diagram, choose the catalog first because it is the heavier and less reversible decision, price the exit in weeks before committing to anything, and score the nine criteria on the connector versions they actually run instead of on what the specification allows.

The detail

What Separates A Scorecard From A Feature Grid.

Zone · 01

Weights Set Before Scores

Set the weights as a group before anyone scores a format. That is where the real disagreement lives, and settling it separately stops the exercise becoming a vote on a preference somebody already held. Nine criteria, weights of one to three, twenty-two weight points, a ceiling of 110 per format. Argue about the weights out loud, then the scoring itself takes twenty minutes.

Zone · 02

The Gap

Not The Total

Totals invite a winner. The gap tells you what to do. Under ten points is a dead heat and the most common outcome, so take the format your team can operate today and spend the saved time on catalog design. Ten to twenty-five is a real preference worth acting on. Over twenty-five, name the two criteria driving it and check they are facts about your estate.

Zone · 03

Reversal Cost In Weeks

Write the exit cost down before you score anything. One number, in weeks, counting every dbt target, BI dataset, notebook, streaming checkpoint, permission grant, external partner and job that hardcodes a location. The data layer is Parquet either way, so converting metadata is hours to a few days. Rewiring consumers is the bill. If nobody can produce that number, fix the dependency graph first.

By the numbers

The figures that make it a board-level conversation.

80%
of data and analytics governance initiatives forecast to fail by 2027, for want of a real or manufactured crisis to force the issue
56%
of data practitioners name data quality as a key concern, which is a contract and catalog problem rather than a file format one
6
weeks of parallel run before the old tables retire
Inside the report

What you'll take away.

01

price the exit before anyone scores

One number, in weeks. Count every dbt target, BI dataset, notebook, streaming checkpoint, permission grant, external partner and job with a hardcoded path. The data platform lead owns the list and the consumer teams confirm it. If nobody can produce the number, stop the session and fix the dependency graph first.

02

set the nine weights as a group

Weights of one to three, agreed before either format is scored. We usually land on 3 for engine mix, catalog, write patterns, interoperability and exit cost, 2 for partition evolution, vendor model and operational load, and 1 for time travel. Twenty-two weight points, 110 per format. Change them out loud if your situation differs.

03

score the versions you actually run

One to five on fit, both formats, on the connector versions in production. A capability that needs an unscheduled engine upgrade scores 2, not 5. Check write and MERGE support in the connector notes rather than read support on a vendor page, and have someone demonstrate a second engine committing through your catalog.

04

read the gap and write the memo

One page: the decision, the two criteria with the largest weighted gap, the reversal cost in weeks, and the date you will review it. Send it to every team that reads these tables. If nobody objects within a week the argument is closed. Then name one source of truth per table and a date the bridge comes down.

Questions

Frequently asked.

Does the feature gap between Iceberg and Delta still matter?

Rarely, and less with each release. Both write Parquet, both do ACID commits, safe schema evolution, time travel and merge-on-read with deletion vectors.
Deletion vectors, variant types and row lineage all landed on one side and then the other within two releases. A gap with a merged pull request behind it is a timing question, not a decision input.

We already run Databricks, so is this exercise pointless?

No, but the call is probably Delta. If Spark is your primary writer and reader you would pay for a full migration to reach parity you already have. Turn on Iceberg-compatible metadata for outside readers, treat it as a read bridge rather than a two-way sync, and revisit when a second engine needs to write.

Why is the catalog a heavier decision than the format?

Because it owns identity. The catalog holds table names, permissions and commit coordination, so replacing it later touches every job, every grant and every consumer. A format you can leave in a quarter deserves less debate than a catalog you cannot. Choose the catalog first and the format argument tends to resolve itself.

Can we not run both formats and decide later?

You can, and it costs. Dual-write is two pipelines, roughly 60% to 100% more ingest compute and twice the maintenance, and the second one gets neglected first. Fund the reconciliation before the second pipeline: row counts and a per-partition checksum compared daily, with a named person who acts when the comparison disagrees.

Should we run a proof of concept before we decide?

Only a real one. Everything is fast on 50GB. Take one production table, write to it from your primary engine, read it from every other engine and the BI tool, then repeat while a compaction job runs. Failures show up as stale snapshots, missing rows after a delete, or a reader handing back rows you deleted.

Who is this guide for?

VPs of Data and platform leads who have to close this argument. It assumes you have tables in object storage, more than one engine reading them, a catalog somebody picked without a meeting, and a decision due that two teams already hold strong opinions about.

Get the framework

Have it emailed to you.

Drop your details and we'll send 80% Of Governance Programmes Fail By 2027. The Catalog Decides That, Not The Format. straight to your inbox - no spam, unsubscribe anytime.

Download framework
Next step

Put this into practice.

Book a 30-minute lakehouse review (logiciel.io)

Download the guide