Insights

Your data policy has to run where the data runs

Employees already paste company data into public AI, and agents now query systems of record on their behalf. A data policy keeps up only when it runs in the pipeline.

Paul Turner

VP, Market Strategy, Tray.ai

Most companies have a data policy that says what may and may not go into an AI tool. Most of their employees have broken it, or never knew it existed.

Sensitive data already reaches models; that part is settled. What’s open is where the control sits when it does. My answer: in the path the data takes, at the moment it moves, written as code. Today it’s usually a paragraph in a handbook.

Start with the people

One of the largest independent studies of how people use AI at work came from the University of Melbourne and KPMG in 2025. They surveyed 48,340 people in 47 countries. About half admitted uploading sensitive company information into public AI tools, and only two in five said their employer had a policy on AI use at all.1

That survey was fielded at the turn of 2025, so it’s about 18 months old. Nothing published since suggests the behaviour has eased, and the tools have become far easier to reach in the meantime. Read it as a floor.

Two things stand out. The first is that the leak is mostly ordinary: someone pastes a customer list into a chat window because they want a summary by lunch. The second is that a policy three in five employees can’t recall is failing at its one job. We made the wider version of this argument in regulators and insurers now ask for evidence: a control you can’t show running is a control you’re hoping for. Data is where that hope is most expensive.

Why the old tools miss it

Classic data loss prevention was built for a particular picture of risk: files moving across a network, on devices the company manages, through gateways the company can inspect. Block the attachment, flag the upload, quarantine the USB stick.

Gartner’s AI governance checklists are blunt about why that picture no longer holds. Their warning is that data can be exposed by routes no perimeter sees, down to a screenshot or someone reading a record aloud, and that organizations should not assume perimeter defences or access limits will deal with AI risk on their own.2 An agent that queries your CRM through an API on a user’s behalf never moves a file. A model that receives a prompt full of account numbers never touches a corporate laptop. The data still left.

The problem splits two ways

In practice the risk shows up in two different shapes, and they need different answers.

Structured data lives in systems of record: the CRM, the ERP, the HR system, the warehouse. It has a schema, so you know which column holds the salary and which holds the Social Security number. The risk here is scope. An agent that queries a system of record usually does so with a service account, and a service account typically sees far more than the person who asked the question. A sales rep asks for a summary of their territory; the agent pulls every account, because it can. Then the sensitive fields flow onward into the prompt, the response, and every log line in between.

Unstructured data is contracts, support tickets, PDFs and email. It has no schema, so there is no column to mask. A bank account number sits in paragraph four of a supplier agreement. A customer’s medical condition sits in the third reply of a support thread. You can’t apply a field rule to a field that doesn’t exist yet, so the first job is to classify the content and extract what is in it. Only then can the same rules that protect a database column protect a sentence.

Teams that treat these as one problem tend to build controls for the first and assume they cover the second. A masking rule on the ssn column does nothing for the same number pasted into a ticket.

Where the control goes

Gartner’s data governance research from July 2026 makes the structural point directly: governance scales when it operates where the data runs, at runtime, embedded in the pipelines and platforms the data passes through, without waiting on people to escalate by hand. Their policy workflow ends with IT applying technical controls such as encryption, redaction and tokenization, which in effect turns the policy into code.3

That gives the control a location and an order. Every piece of data on its way to a model or an agent passes through the same sequence.

  1. 1

    Classify and extract

    Find the sensitive fields, including the ones hiding in free text. Documents get read and their contents labelled, so a clause in a contract is treated the same way as a column in a table.

  2. 2

    Mask or tokenize

    Replace sensitive values before anything reaches a model. The model works on a token where the account number was; the real value stays in the system that owns it.

  3. 3

    Enforce access at call time

    Check what the person asking may see at the moment the request runs, so the agent returns their territory and nothing beyond it.

  4. 4

    Keep sensitive fields out of the logs

    Apply the same rules to prompts, responses and traces that you applied to the data itself.

The order matters: each step depends on the one before it, and the last one protects the record of all the others.

The last step is the one most teams forget, and it deserves its own sentence. When an AI feature goes into production, the first thing a responsible team adds is observability: full prompts, full responses, full traces, kept for audit. If nothing masked them, those logs now hold every sensitive value the system ever touched, in one searchable place, often with looser access than the source system. The observability you added for audit becomes the leak.

One policy, enforced differently per person

Gartner’s research gives an example worth borrowing for any internal debate about this. A support team may see personal data under specific criteria, because resolving a case requires it. Marketing never sees it without going through an approval process.3 That is one policy. It produces different results for different people asking the same question of the same data.

Notice what that requires. The pipeline has to know who is asking, at the moment they ask. If an agent calls the CRM as a shared service account, the request arrives with no person attached, and the only policy you can enforce is the one that applies to the service account. Either everyone gets the support view or everyone gets the marketing view. Per-person policy only works if identity travels with the request, all the way from the person to the system of record and back.

This is also why the control has to live in the path. A document can say that marketing doesn’t see personal data. The path the data takes is what makes that true on every request, for every agent, including the ones built last week by someone who never read the document.

Where Tray fits

Tray puts this in the platform, once, for every app. Credentials live in the Tray platform, so neither builders nor the models working for them ever hold a raw secret: a Tray Helix app references a named auth alias, and Helix resolves it at runtime (we covered how in builders never hold a credential). Data stays in the region you choose, with residency in the US, EU and APAC.

Helix deploys, runs and governs the apps your teams build in Claude Code, Codex and Cursor. Tray also has a full iPaaS on the same foundation, and that’s where the pipelines that move and prepare data for apps and agents run. The detail on how apps are secured is on App Security.

Footnotes

  1. University of Melbourne and KPMG, “Trust, attitudes and use of AI: A global study 2025,” n=48,340 across 47 countries, fielded late 2024 into early 2025. Academic co-author; no vendor sponsor. Self-reported behaviour. Back

  2. Gartner, “5 AI Governance Checklists That Maximize Potential and Minimize Risk,” G00851967, Andrews, Kornutick, 31 July 2026. Gartner does not endorse any vendor, product or service depicted in its research publications. Back

  3. Gartner, “Building Blocks for Effective Data Governance in the Age of AI,” G00853417, Amy Bickel, 13 July 2026. Back Back

  • governance
  • data governance
  • ai agents
  • data protection

Helix is the governed runtime for AI-built apps

Deploy what your teams build, put SSO in front of it, connect it with managed credentials, and give every app a named owner.

Want to talk to someone first? Contact us