Skip to main content

Co-Existence

Using your existing Raw or Staging layer, you can map the required Data from the existing message structures into the CryspIQ® Data model. This enables you to leverage the flexibility of Data model, measure Data Quality and use Natural Language to access your Data. This will provide immediate productivity benefits to your business and enable you to reduce costs.

Two changes, and only one of them is visible

Adopting CryspIQ® is really two separate things happening on different timescales, and confusing them is what turns it into a two-year programme in people's heads.

What the business experiences changes almost immediately. As soon as one source is mapped, people can ask questions of it and get answers against governed definitions. There is nothing to wait for — no migration to finish, no cutover date, no phase two. That is the change users actually feel, and it does not depend on any of the technical work below being complete.

What happens underneath changes slowly, and nobody in the business notices. Layers that existed only to add meaning get retired as things stop reading them. No user sees that happen; their reports keep working throughout. It surfaces months later as a smaller bill, not as a better day at work.

Keeping those apart matters because they answer to different people. A head of finance is buying the first one. A CIO is planning the second. Neither has to wait for the other.

What people notice

This is the part with a human on the other end of it, and it arrives first.

When each benefit arrivesBenefits mapped to the adoption stages. Productivity arrives first, from stage two. Time to value follows at stage three. The cloud storage and compute saving arrives last, at stage four, when duplicated storage stops being held. Each continues once it starts.1 · Today2 · Run both3 · Retire what is unused4 · Connect direct50%Productivity75%Time to value80%Cloud storage and compute

Each benefit continues once it starts — by the final stage all three are in effect. Figures are the outcomes reported on the productivity, time to value and cloud cost pages.

Productivity comes first because it needs nothing but the model being in use — people answering their own questions against agreed definitions instead of queuing for a report. Time to value follows once the layers that only added meaning are gone, so a new question no longer waits on a modelling and pipeline cycle. Both are felt directly by the people doing the work.

The cost saving comes last, and no user will ever notice it. Until duplicated storage stops being held, the bill has not changed — it appears as an IT line item, long after the business stopped waiting for reports.

That order is worth planning around. A business case expecting the cloud saving in the first quarter will look like a failure at exactly the point productivity is improving.

What happens underneath

The sequence below is an IT concern. It is worth understanding, because it is what makes adoption low-risk, but it is not what anyone outside the data team will experience.

Moving from a traditional warehouse to CryspIQ, in four stagesSources and the consuming reports are identical in all four stages. CryspIQ first runs in parallel with the existing transformation and serving layers, both fed from the staging layer that already exists. Those layers are then retired once nothing depends on them, and finally sources connect directly to the model so the staging layer is no longer needed either.1 · TodaySourcesRaw or stagingTransform → serveReports & AI
Meaning is added across several layers before anything is reportable.
2 · Run bothSourcesRaw or stagingTransform → serveCryspIQ®Reports & AI
CryspIQ® reads the staging layer you already populate. Map one source and it is queryable — nothing is switched off.
3 · Retire what is unusedSourcesRaw or stagingTransform → serveCryspIQ®Reports & AI
Layers built only to add meaning are removed once nothing reads them.
4 · Connect directSourcesRaw or stagingCryspIQ®Reports & AI
Sources feed the model directly and the staging layer goes with it. This is where the storage and compute saving lands.

Dashed means retired only once nothing reads it. No stage requires a cutover, and stopping at any of them is a valid place to stop.

Adoption is usually imagined as a migration, which is what makes it a board decision. It does not have to be one. CryspIQ® reads the raw or staging layer you already populate, so both paths run at the same time, fed from the same place. The first source you map becomes queryable while every existing report carries on untouched.

Nothing is switched off to begin with, and each layer is retired only when nothing reads it — not before. Stopping at any stage is a valid place to stop, and the business value described above is already banked by stage two regardless of how far the technical sequence ever runs.

Using Existing Raw or Staging Layers

Most organisations already ingest data into a Raw or Staging layer within modern cloud data platforms such as:

  • AWS Data Lake (Amazon S3 with Glue / Athena)
  • Snowflake (raw schemas, landing tables, VARIANT data)
  • Azure Data Lake / Synapse
  • Databricks Lakehouse

These layers typically store data as it arrives, often in its original source structure (for example JSON messages, CSV extracts, or event streams). While this approach works well for ingestion and retention, the data is often:

  • Difficult for business users to access
  • Inconsistent across systems
  • Lacking clear business meaning
  • Dependent on specialist SQL or engineering knowledge

CryspIQ® builds on this existing foundation rather than replacing it.


Mapping into the CryspIQ® Data Model

Using the existing Raw or Staging layer, data is mapped from source-specific message structures into the CryspIQ® enterprise data model.

Examples include:

  • JSON events stored in Amazon S3 or Snowflake VARIANT columns
  • ERP and finance extracts landed in Snowflake staging tables
  • Operational data ingested via AWS Glue pipelines

These source structures are mapped into standardised business entities such as Customer, Account, Transaction, or Product—independent of how the data was originally produced.

This approach ensures that:

  • Source systems can change without breaking analytics
  • Multiple systems can feed the same business entity
  • Data definitions remain consistent across the organisation

Flexibility of the Data Model

Unlike rigid star schemas or hand-built data marts, the CryspIQ® data model is designed to be extensible and adaptable.

This enables organisations to:

  • Add new attributes without redesigning pipelines
  • Support multiple source systems for the same business entity
  • Evolve business definitions over time

For example:

  • A new data feed landed in Snowflake can be mapped into existing entities without reprocessing historical data
  • New event types added to an AWS data lake can be incorporated with minimal effort

This flexibility is particularly valuable in cloud environments where data volumes, sources, and use cases evolve rapidly.


Measuring and Managing Data Quality

Once data is aligned to a common enterprise model, data quality can be measured consistently across all sources.

Typical data quality measures include:

  • Completeness of critical attributes (e.g. missing customer identifiers)
  • Consistency between systems (e.g. mismatched account statuses)
  • Timeliness of ingestion from upstream platforms
  • Validity checks on formats and values

In platforms such as Snowflake or AWS, this replaces ad-hoc SQL checks and manual dashboards with standardised, reusable data quality metrics that are directly aligned to business entities.


Natural Language Access to Data

Because the data is structured around a business-aligned model, users no longer need to understand:

  • Source system schemas
  • Complex table joins
  • Platform-specific SQL logic

Instead, business users can query data using natural language, for example:

“Show me customers with incomplete onboarding data last quarter”

CryspIQ® translates these requests into the appropriate queries across platforms such as Snowflake, AWS Athena, or Databricks, while enforcing consistent definitions and governance.


Immediate Productivity and Cost Benefits

By leveraging existing cloud platforms rather than rebuilding them, organisations achieve rapid value:

  • Reduced engineering effort by avoiding bespoke data marts
  • Lower compute and storage costs by minimising data duplication
  • Faster access to insights through self-service analytics
  • Improved trust in data through visible and consistent quality metrics

This delivers immediate productivity improvements and measurable cost reductions, particularly in cloud environments where inefficient data usage directly impacts spend.