Book a free AI Reality Check
How it works Industries Case studies Resources What you could ask Where to start Book a free AI Reality Check

Plain English

What actually happens to what you type into an AI

Most of the advice on this subject is written for people who already know the answer. This is the version for everyone else - what happens to a sentence after you press enter, which parts of it matter, and a test you can apply in about five seconds.

Four things can happen, and they are not the same thing

What happens after you press enter

  1. Your promptand anything pasted in
  2. Readby the model and automated screeningAlwaysNot covered
  3. Keptfor the vendor's retention periodUsuallyNot covered
  4. Read by a personif flagged, required or sampledSometimesNot covered
  5. Used for trainingdepending on product and planSometimesCovered

When you paste a contract into an AI assistant and ask what is wrong with clause 14, four separate things may happen to that contract. People tend to worry about the last one and overlook the first three, which is the wrong way round.

It gets read. This one always happens, and it has to. A model cannot tell you what is wrong with clause 14 without processing clause 14. On most hosted services, automated screening reads it at the same moment, checking it for misuse against rules the vendor sets - and that screening generally keeps running even on plans that switch off storage and human review. There is no version of this where the text is checked without being looked at. If the service runs on someone else's computers, then at the moment of answering, your contract is on someone else's computers.

It gets kept. Usually yes, at least for a while. Most services retain what you send for some period - to show you your history, to investigate abuse, to meet their own legal obligations. Retention periods vary enormously and are often adjustable, sometimes only on business plans.

It gets reviewed by people. Sometimes. A conversation the screening flags can be passed to a person - the vendor's staff or contractors - to confirm the finding. People may also read conversations to meet a legal obligation, such as a court order or a lawful request from authorities, or as part of a sample taken to improve the product. What happens to a finding is decided by the vendor, not by you: what counts as a flag, how long a flagged conversation is kept, who sees it and whether it is reported onward. You are typically told only if the vendor concludes there is a problem, and on some consumer services a reviewed conversation is kept for years, even after you delete your own history. Business plans usually narrow human review, and removing it altogether typically requires a separate approval. On workplace plans, your own organisation may also be able to audit what staff type - often a compliance requirement in its own right, and one more set of readers to account for.

It gets used to improve the model. Sometimes. This varies by product and by plan, and it is the one thing vendors compete on publicly. It is also the one that gets all the attention.

Why “we don't train on your data” is a smaller promise than it sounds

It is a real commitment and worth having. It answers the last question. It says nothing at all about the first three.

A vendor can honestly say your data never trains a model while its model and its screening systems still read every word, it is stored for thirty days, authorised staff can review anything flagged, and it is processed in a country whose authorities can compel access. All of that can be true at once, and none of it is a scandal - it is just how a hosted service works.

Encryption does not close the gap either. Data is protected while it is stored and while it moves. To be understood, it has to be decrypted and processed. That third state - data in use - is where AI lives, and it is the one most procurement checklists never ask about. We wrote about that separately in Enterprise AI has a data-in-use problem.

When does any of this actually matter?

Not always. A great deal of work is genuinely fine on a public model, and pretending otherwise wastes money and makes people ignore the warnings that count. Four kinds of material are where it stops being fine:

The five-second test

Before pasting something in, ask: would I email this to a supplier I have never met?

If yes, a public model is very likely fine, and you should use one, because it will be quick and cheap and good. If you hesitate, that hesitation is the finding. It does not mean AI is off the table - it means that particular material needs somewhere else to run.

So where else can it run?

Two realistic alternatives, neither of which is better in the abstract.

Your own cloud tenant. The work happens inside an account you control, in a region you choose, and your material is not used to train public models. There is no hardware to buy and no ceiling on how capable the model can be. The processing is still performed as a service by the cloud provider, under that provider's contract. The provider generally still screens that traffic automatically, under rules it sets, and on some services authorised staff can review what gets flagged unless you apply to have that switched off. It is a genuine step up from a consumer product, and a smaller step than it is sometimes sold as.

Your own infrastructure. The model runs on machines you control, in your own cloud or your own building. Nobody else reads it, your own law applies, and if you stopped paying anyone tomorrow it would keep working. If you want screening, you set the rules and you see the findings. It costs more effort to stand up, and for a great deal of work it is more than you need.

Who sets the rules

  • Public AI assistantThe vendor's rules
  • Your own cloud tenantThe provider's rules, under your contract
  • Your own infrastructureYour rules

Which of these a given job needs is a question about the material, not about technology, and it is worth answering before anyone builds anything. We have written more about how those options rank against each other in Whose rules does your AI answer to?

The one thing worth doing this week

Find out what your team is already using. Not what was approved - what is actually open in their browsers, and what has quietly appeared inside software you already pay for, switched on by default. Almost every organisation we talk to discovers at least one thing it did not know about, and it is nearly always a default rather than a decision.

That is not a failure of policy. It is what happens when useful features arrive faster than anyone can review them.

Where would your own work actually need to run?

Five questions, about a minute. You will get a plain-English starting point for your team and the lightest deployment tier your data allows. It runs entirely in your browser - there is no form and nothing to submit.