Skip to main content

CryspIQ® vs Databricks

CryspIQ® runs on top of Databricks rather than instead of it. Databricks is a lakehouse platform: storage, distributed compute, notebooks, ML and the medallion pattern for organising data as it is refined. CryspIQ® is the enterprise model that platform expects you to supply. The question is not which to buy — it is where business meaning gets attached, and how many times.

What Databricks gives you

Unified storage and compute over open table formats, first-class support for machine learning and unstructured data, and a processing model that scales to whatever you can throw at it. For data science and large-scale transformation it is exceptional, and CryspIQ® depends on that rather than reproducing it.

Where the medallion architecture puts meaning

Bronze holds data as it landed. Silver cleans and conforms it. Gold makes it business-ready for a particular audience.

Meaning therefore accumulates gradually, and it accumulates in the pipeline rather than in the data. That has three consequences worth naming, and they are structural rather than a matter of doing medallion badly:

  • Every layer is a place meaning can be applied differently, which is how two gold tables built by two teams disagree
  • Every layer needs building and maintaining, which is where the engineering time goes
  • Meaning defined for one gold table is available to that consumer only, so the next one starts again

The gold layer is where the enterprise model should be, and there is usually more than one of them.

What CryspIQ® changes

CryspIQ® attaches meaning at entry. The fact type carries the business definition and the security classification, and quality is measured as data loads. Once meaning is already attached, the layers whose job was to add it later have nothing left to do.

On a lakehouseWith CryspIQ®
Where business meaning livesThe gold layer, per consumerOn the fact type, once
Enterprise data modelYou design and build itPre-defined, patented, ready on day one
Master dataYou choose an approach and maintain itA single instance per party, enforced by the structure
Data qualityApplied per pipeline, downstreamAssessed as data loads, scored organisation-wide
LineageReconstructed from the notebooks and jobsAutomatic, through the link key
Self-service analyticsRequires a gold table to be built firstNo modelling step between question and answer

Where Databricks is stronger

Everything CryspIQ® does not attempt. Training and serving models, processing unstructured data at volume, arbitrary transformation in Python or Scala, and exploratory work where the shape of the answer is not known in advance. CryspIQ® is an opinionated model for governed business facts, not a compute platform, and it would be the wrong tool for a feature engineering pipeline.

There is also a real trade in flexibility. A lakehouse will model anything you can write code for. CryspIQ® asks your data to fit seven fact types and a fixed set of dimensions — fixed at the level where organisations are alike, and open at the level where they differ.

How they fit together

CryspIQ® maps from the bronze or silver layer you already populate — events, extracts and streamed data — into standardised business entities, independent of how each source produced them. Databricks keeps ingesting, transforming and serving ML. The governed model then feeds the reporting and self-service side, which is the part that most often stalls waiting on a gold table.

Adoption does not require a migration: both paths run at the same time, fed from the same place, and the first source mapped becomes queryable while every existing job carries on untouched. See co-existence.