What Do AI Assistant Data Terms Actually Say?

“Is my data private?” is the wrong question to ask about an AI assistant, because it has no single answer. The useful version breaks into four separate questions, and the answers depend on which tier you’re on, not which company you’re with. A consumer free account and a business seat at the same vendor can be governed by completely different commitments. That’s the single most important thing in this article.

Here’s what to look for, in the order that matters.

1. Is my input used to train models?

The headline question, and the one where the free-versus-paid split is sharpest as of mid-2026. The rough industry pattern:

  • Consumer tiers (free and often personal paid) may use your conversations to improve models, frequently with an opt-out setting somewhere in preferences. Defaults vary and change.
  • Business, team, and enterprise tiers generally commit contractually to not training on your content. This is table stakes for selling to companies.
  • API and developer access generally isn’t used for training by default at the major vendors, which surprises people — the developer path is often the more conservative one.

What to actually do: find the setting in your own account and look at it rather than assuming. And note that an opt-out toggle is a product setting, which a vendor can redesign; a contractual commitment on a business tier is a different kind of promise. If it matters, you want the second kind.

Watch for the words. “We don’t sell your data” is not “we don’t train on your data.” “We don’t use your data to train third-party models” is not “we don’t train on it.” “Your data is yours” is a sentiment, not a term.

2. How long is it retained, and can I delete it?

Separate from training, and frequently missed. Even a vendor that never trains on your input will usually store it — for chat history, for abuse monitoring, and for a legally-motivated retention period.

Things to check:

  • Whether deleting a conversation deletes it from the backend, or just from your view. These are not the same and the interface rarely tells you.
  • Whether there’s a short retention window for abuse monitoring that applies even to zero-retention arrangements. Some vendors offer genuine zero-retention on request for eligible customers; most keep something briefly.
  • What happens after you close the account. Deletion is usually a process with a timeline, not an event.
  • Whether an admin can read your history on a business tier. Often yes, and that’s a feature for the organisation, not a bug — but staff should know.

3. Can a human read it?

Almost always yes, in narrow circumstances: abuse investigations, safety review, debugging, legal process. This is normal and every vendor does some version of it. What varies is how narrowly the circumstances are drawn and whether the vendor says so plainly.

The practical takeaway isn’t alarm, it’s calibration: treat an assistant conversation as being roughly as private as an email to a company you have an account with. Not published, not sealed. If you wouldn’t email it to a support desk, don’t paste it into a chat window.

4. Where does it physically go?

Only matters if you’re subject to rules that make it matter — data-residency requirements, sector-specific regulation, or a customer contract that specifies it. If you are, this is usually an enterprise-tier feature with an explicit region commitment, sometimes with a surcharge, and it’s a procurement conversation rather than a settings toggle. If you aren’t, don’t let it distract you from questions 1 and 2, which affect everyone.

Related: subprocessors. Most vendors run on someone else’s cloud and use third-party services for parts of the pipeline. There’s usually a published list. It’s rarely a problem; it’s occasionally the thing that fails a compliance review, and it’s better to find out before you’ve rolled anything out.

The questions vendors don’t answer well

Being fair to vendors: some of these are genuinely unresolved rather than evasive.

“Is my data in the model already?” If your content was used in training before you opted out, it is not extractable in any straightforward sense, and it is also not removable. Retroactive deletion from a trained model isn’t a thing anyone can offer. The honest position is that opt-outs are forward-looking.

“What about the things I attached?” Uploaded documents, images, connected files. Terms often speak in terms of “conversations” while your exposure is mostly in attachments. Worth reading carefully; the treatment isn’t always identical.

“What did the integration see?” Connect an assistant to your mail or drive and the scope of what it may read is defined by the permissions you granted, which are usually broader than the task at hand. This is where most real-world surprise lives, and it’s a permissions question you control rather than a terms question the vendor controls.

How to decide, by situation

Ordinary personal use, nothing sensitive. Check the training toggle once, get on with your life. This is genuinely fine and the anxiety industry around it is overblown.

Personal use with confidential material — client work, medical, legal, unpublished writing you care about. Either use a business tier with a no-training commitment, or don’t paste the sensitive part. Redaction is underrated: most tasks need the shape of the document, not the names.

A hard constraint that no contract satisfies — a duty of confidentiality, an air-gapped requirement, a client who has said no third parties. Then the answer is categorical rather than contractual: a model running on hardware you control. See do you need a local model for privacy for whether that’s really your situation, and can I run a ChatGPT alternative locally for what it costs you.

Buying for a team. Business tier, read the actual agreement rather than the marketing page, write a one-page internal rule about what may be pasted where, and audit file permissions before connecting anything to your document store. See does it matter which AI assistant your team uses.

The one rule to remember

Compare tiers, not brands. Every general claim of the form “vendor X is private and vendor Y isn’t” is almost certainly comparing a business tier at one company to a consumer tier at another. The distribution of good and bad practice runs within companies, across their product lines, far more than it does between them — and it changes. Read the terms attached to the tier you’re actually paying for, check the settings in your own account, and re-check after any major product change.

For the wider decision about which assistant to use, see our framework for choosing an alternative.