CryspIQ® vs Databricks
CryspIQ® runs on top of Databricks rather than instead of it. Databricks is a lakehouse platform: storage, distributed compute, notebooks, ML and the medallion pattern for organising data as it is refined. CryspIQ® is the enterprise model that platform expects you to supply. The question is not which to buy — it is where business meaning gets attached, and how many times.
What Databricks gives you
Unified storage and compute over open table formats, first-class support for machine learning and unstructured data, and a processing model that scales to whatever you can throw at it. For data science and large-scale transformation it is exceptional, and CryspIQ® depends on that rather than reproducing it.
Where the medallion architecture puts meaning
Bronze holds data as it landed. Silver cleans and conforms it. Gold makes it business-ready for a particular audience.
Meaning therefore accumulates gradually, and it accumulates in the pipeline rather than in the data. That has three consequences worth naming, and they are structural rather than a matter of doing medallion badly:
- Every layer is a place meaning can be applied differently, which is how two gold tables built by two teams disagree
- Every layer needs building and maintaining, which is where the engineering time goes
- Meaning defined for one gold table is available to that consumer only, so the next one starts again
The gold layer is where the enterprise model should be, and there is usually more than one of them.
What CryspIQ® changes
CryspIQ® attaches meaning at entry. The fact type carries the business definition and the security classification, and quality is measured as data loads. Once meaning is already attached, the layers whose job was to add it later have nothing left to do.
| On a lakehouse | With CryspIQ® | |
|---|---|---|
| Where business meaning lives | The gold layer, per consumer | On the fact type, once |
| Enterprise data model | You design and build it | Pre-defined, patented, ready on day one |
| Master data | You choose an approach and maintain it | A single instance per party, enforced by the structure |
| Data quality | Applied per pipeline, downstream | Assessed as data loads, scored organisation-wide |
| Lineage | Reconstructed from the notebooks and jobs | Automatic, through the link key |
| Self-service analytics | Requires a gold table to be built first | No modelling step between question and answer |
Where Databricks is stronger
Everything CryspIQ® does not attempt. Training and serving models, processing unstructured data at volume, arbitrary transformation in Python or Scala, and exploratory work where the shape of the answer is not known in advance. CryspIQ® is an opinionated model for governed business facts, not a compute platform, and it would be the wrong tool for a feature engineering pipeline.
There is also a real trade in flexibility. A lakehouse will model anything you can write code for. CryspIQ® asks your data to fit seven fact types and a fixed set of dimensions — fixed at the level where organisations are alike, and open at the level where they differ.
How they fit together
CryspIQ® maps from the bronze or silver layer you already populate — events, extracts and streamed data — into standardised business entities, independent of how each source produced them. Databricks keeps ingesting, transforming and serving ML. The governed model then feeds the reporting and self-service side, which is the part that most often stalls waiting on a gold table.
Adoption does not require a migration: both paths run at the same time, fed from the same place, and the first source mapped becomes queryable while every existing job carries on untouched. See co-existence.
Related
- All comparisons
- AI and machine learning — governed datasets for AI initiatives
- AI-ready data needs transactions, not copies