Skip to main content

6 posts tagged with "Data Lake"

Reference to Data Lake Warehousing Methodology.

View All Tags

How to Protect Your Data from AI Agents

· 12 min read
Dan Peacock
Chief Hustler

This is the public text of the white paper AI Agents on Data Lakes. It answers the question people are now asking search engines and assistants at the same time: how do I protect my data from AI agents?

You protect it by attaching the permission to the data, not to the tool. A bucket policy, a folder grant or a catalog role decides whether an identity can open an object. It does not decide whether that request should see a particular value inside it. An agent is usually given one standing service account, so every question it answers inherits that account's access, whoever is actually asking.

The Medicare data breach is why this stopped being a theoretical gap. In June 2026 an OpenAI agent gained unauthorised access to a public-facing Medicare statistics portal run by Services Australia, reaching public and non-public files. The Australian Government has said no personal information is believed to have been accessed. It has also called the incident unacceptable, and ministers have said the law will change if today's framework cannot control an autonomous agent. That is governments calling for control of AI. On a data lake, the control that holds is the same one: the agent must not be able to see more than the person it is acting for, and that has to be a property of the data.

SaaS Isn't Dead. We've Been Valuing the Wrong Half of It.

· 5 min read
Dan Peacock
Chief Hustler

Someone told me recently that SaaS is dead. I disagree — strongly — but not because the argument behind it is silly. It isn't. Three things in it are true.

Building software is getting dramatically cheaper. Per-seat pricing is under genuine pressure once software is doing the work people used to do. And the case for buying twenty modules from one vendor is weaker than it was, because integration difficulty was the reason that bundle existed, and integration is getting easier.

All fair. It just adds up to a different conclusion than the one people are drawing.

AI-Ready Data Needs Transactions, Not Copies

· 7 min read
Dan Peacock
Chief Hustler

In the previous post we argued that AI-ready data starts with master data — that until an organisation has one definition of "customer", nothing built on top of it can be trusted.

That is the foundation. It is not the whole building.

Master data tells an AI system who your customers are. It cannot tell it what they did. Answers that mean anything to a business — which accounts are growing, which are quietly churning, where the cross-sell actually is — come from transactional data. And transactional data is only useful to AI if it is linked to the master data that gives it meaning.

Data — Asset vs Liability?

· 4 min read
Dan Peacock
Chief Hustler

At its core, an asset is something that creates value and drives growth—directly or indirectly. It should strengthen resilience, fuel innovation, and deliver competitive advantage.

Now consider your data: Is your Enterprise Data Platform positioned as an asset, or is it on track to become a liability?

As your business expands, the foundations you set for your data become critical. A true data asset reduces reliance on ever-changing applications and removes the need for specialist technical skills just to interpret the numbers.

Unpacking Buzzwords

· 8 min read
Dan Peacock
Chief Hustler

"Lakehouses", "Lakebases", "Meshes", and "Medallion Architectures". Regardless of the "buzzword" being used, it's essential to understand the underlying methodology, as these ones all follow the same foundational Data Lake methodology — yet the core business question often remains unanswered or simply assumed. Before committing to a data journey that typically spans 3–5 years and costs in excess $25 million — an approach frequently promoted by industry quadrants, shaping strategic architecture decisions — maybe worth considering the following points: