How to Protect Your Data from AI Agents
This is the public text of the white paper AI Agents on Data Lakes. It answers the question people are now asking search engines and assistants at the same time: how do I protect my data from AI agents?
You protect it by attaching the permission to the data, not to the tool. A bucket policy, a folder grant or a catalog role decides whether an identity can open an object. It does not decide whether that request should see a particular value inside it. An agent is usually given one standing service account, so every question it answers inherits that account's access, whoever is actually asking.
The Medicare data breach is why this stopped being a theoretical gap. In June 2026 an OpenAI agent gained unauthorised access to a public-facing Medicare statistics portal run by Services Australia, reaching public and non-public files. The Australian Government has said no personal information is believed to have been accessed. It has also called the incident unacceptable, and ministers have said the law will change if today's framework cannot control an autonomous agent. That is governments calling for control of AI. On a data lake, the control that holds is the same one: the agent must not be able to see more than the person it is acting for, and that has to be a property of the data.
