Analysis
Most organisations still assess AI with a checklist written for cloud storage. It asks where the data lives and whether it is encrypted. Both are reasonable questions. Neither one describes what an AI system actually does.
Security teams have spent fifteen years getting good at two questions. Is the data encrypted at rest, sitting in storage? And is it encrypted in transit, moving between systems? Answer both and most cloud procurement is satisfied.
Those controls are mature, well understood and largely solved. They are also, for AI, insufficient — because neither of them describes the moment the work happens.
To answer a question about a contract, a model has to read the contract. To check a report against a policy, it has to process both. There is no version of useful AI that does not, at some point, hold your material in memory and compute over it.
That moment is a third state: data in use. The idea is not new — confidential computing has been circling it for years — but AI is the first mainstream workload where the third state is the entire point. Encryption at rest protects a file nobody is reading. Encryption in transit protects a file in motion. Neither protects a file being understood.
The useful question is no longer where is our data stored? It is who processes it, where, under what contract, and what happens to it afterwards?
Four assumptions worth testing
This is the assurance most often quoted to us, and it is usually true — while answering a different question from the one asked. A commitment not to train on your data is a commitment about one downstream use. It says nothing about whether your material is read, cached, logged, retained for diagnostics, inspected for abuse, or held to satisfy an audit obligation.
All of those can be entirely legitimate, and several are requirements the vendor could not waive if it wanted to. The point is narrower: we do not train on your data and nobody else processes your data are different sentences, and only one of them is usually on offer.
Residency commitments generally describe storage. Inference — the moment the model actually runs — can be governed separately, and sometimes is. A service can hold your files in your region and still compute over them somewhere else.
There is a live, published example of exactly this pattern — including a mainstream platform on which the behaviour is switched on by default for new tenants in some regions. We keep it dated and sourced alongside the comparison table rather than repeating it here, so there is one version to maintain: see the data-exposure comparison →
De-identification is a genuinely useful control. For person-linked records it can be the difference between needing a private deployment and not needing one.
It has a hard limit: it removes what is personal, not what is valuable. A pricing model contains no personal information. Neither does a set of engineering drawings, an unreleased design, a source file or a tender response. Strip every name out of those and the commercially sensitive part is entirely intact, because the sensitivity was never in the identity.
So de-identification is a real lever for one class of material and no help at all for another. Knowing which class you are holding is the decision; the technique is the easy part.
Product families share a name, a login and a bill. They do not necessarily share an architecture. Within one vendor’s ecosystem you can find services with different processing locations, different retention behaviour, different contractual terms and different administrative controls.
That is not deception. It is what happens when a portfolio grows by acquisition and iteration. But it does mean we are already on that platform is not an answer to where does this particular workload get processed? The unit of assessment is the service, not the supplier.
None of this is a case for putting everything behind a private deployment. Most organisations hold a mix, and the mix is the whole point:
The work is matching each job to the lightest arrangement its data actually allows. That requires knowing, per workload, what the material is and who touches it in the course of answering. Not per vendor. Per workload.
Which is also why the honest answer to is this safe? is usually another question: safe for what, holding what, processed by whom?
A one-minute check maps the work your team actually does to a deployment tier your data allows — and tells you plainly where a lighter, cheaper arrangement would do. It runs in your browser; there is nothing to submit.
Wondering how this applies to your team? Five questions, about a minute — you’ll get a starting point for your own work and the deployment tier your data allows. It runs in your browser; there is nothing to submit.
Take the 1-minute check →