Egress control
Where does your data actually go when staff use AI
Chadi Abi Fadel, Founder & CTO, AI Engineer
Somewhere in your business today, someone pasted something into an AI tool. A draft email, a customer complaint, a spreadsheet of figures, a paragraph of a contract. It took four seconds and it felt like nothing, because the interface is a text box and text boxes feel private.
It was not nothing. It was a document leaving your building.
A prompt is a document
If a staff member printed a customer's file, walked it out of the office and handed it to a stranger who promised to be helpful, you would have a word for that, and probably a policy. The same file pasted into a chat window is the same act with better ergonomics.
This is the useful reframe, and it is the whole subject. "Egress" is just the question of what leaves your organisation, where it goes, and who can see it once it has gone. Nothing about it requires you to be technical. It requires you to keep treating information as information after it changes shape.
The route a prompt takes
Follow one message from the keyboard outward and the picture stops being abstract.
First it goes to the vendor whose product your staff member is using. That vendor runs software in a place, under the laws of that place, and keeps records of what passes through.
Then, usually, it goes onward. Most AI products do not run their own models. They send your text to a model provider, which is a second company, often in a second country, with its own retention policy and its own terms. Your vendor's promises do not automatically bind their suppliers.
Then there are the parties nobody mentions in the sales call: the hosting provider underneath, the analytics service watching usage, the logging pipeline, the subcontractor who handles support tickets and can see conversations when debugging. Each hop is a party, a jurisdiction and a retention decision, and each one happened without anyone in your business choosing it.
None of this is sinister. It is how modern software is assembled. But "where does our data go" has a real answer with names in it, and most organisations using AI today could not produce that answer for a single tool they use.
Retention, training and telemetry are three different questions
These get blurred together, and vendors benefit from the blur, so keep them separate.
Retention is how long a copy of your prompt exists after the conversation ends. "We do not store your data" often means "we do not store it in the product you can see", while logs live elsewhere for thirty or ninety days, or indefinitely, for debugging.
Training is whether your text is used to improve someone's model. If it is, a version of your information has been absorbed into an asset that other customers benefit from and that you can never recall. Training defaults differ wildly between free and paid tiers of the same product, which is worth sitting with for a moment: the free tool your staff quietly adopted almost certainly has the worse default.
Telemetry is everything that leaves without anyone typing it. Usage data, metadata, which features were used on which documents at what time. Individually dull, collectively a fairly detailed picture of how your business operates.
A vendor can be excellent on one of these and silent on the other two. Ask about all three, by name.
The copies you do not think about
Even with good answers on the main flow, copies accumulate in the corners. Backups hold what the live system deleted. Support staff at the vendor can often view conversations, because that is how support works. A browser extension someone installed can read what is on the screen. A staff member's personal account, used on a work laptop because the work tool was slower, holds a private archive of your business correspondence that you do not know exists.
The pattern is that data does not leak through the front door of the product. It leaks through the fact that the product is surrounded by other software and other people.
What control would actually mean
Control is not a feeling of safety and it is not a clause that says "enterprise-grade security". It is the ability to answer, in writing, a short list of questions:
Which parties see our data, by name. In which countries it is processed and stored. How long each party keeps it. Whether any of it trains any model. Who at the vendor can view it, and under what circumstances. And what happens when we ask for it to be deleted.
If you can answer those, you have control, whatever tools you use. If you cannot, you do not, whatever the security page says. The honest position for most organisations is that they cannot, not because anyone did anything wrong, but because nobody was ever asked to keep the list.
Where to start
Not with a ban. Bans push usage onto personal phones and personal accounts, which is the same egress with less visibility and no contract.
Start with an inventory. Write down which AI tools your staff actually use, including the unofficial ones, which is most of them. Then ask each vendor the questions above and write down the answers, including the refusals, because a refusal is an answer.
You will end up with a one-page document that says where your information goes. Most businesses have never seen theirs. It tends to be a motivating read.
The prompt your staff member typed this morning is already somewhere. The point of doing this work is that the next one goes only where you decided it could.
RELATED READING
What sovereign AI actually means for a business
Sovereignty is the balance between how much an AI system can do for you and how much you can trust it to do it. Here is what that balance looks like in a vendor conversation, and which answers should worry you.
ReadWhy this blog exists
What we write about here, why the posts live in the repository rather than in a database, and what you can expect to find.
Read