Documents and uploads

Private files need more than good intentions.

Your biggest AI training risk is often not your website. It is the PDF, spreadsheet, passport scan, medical letter, customer email or internal slide deck someone uploads to a consumer AI account without checking the rules.

File checklist

Before uploading any private file to AI.

CV or resume: remove phone, address, birth date, signature and references unless they are truly needed.

Medical, legal or finance: treat as high risk. Summarize manually or redact heavily before using AI.

School and family documents: check whether the file contains data about children, parents, teachers or other people.

Client or employer files: get permission and use approved tools. Customer names, prices, contracts and internal strategy are not casual prompt material.

Multi-model tools: tools such as MultipleChat AI can help separate projects and compare models, but the same privacy rules still apply.

The real risk

Staff pasting company data into consumer AI.

For most small teams, the biggest exposure is not a crawler reaching the website. It is an employee pasting a contract, a customer list or an internal deck into a free consumer AI account to save time. This is usually well-intentioned, which is exactly why a clear policy matters more than blame.

Why it happens

Consumer AI tools are fast, familiar and free, so people reach for them under deadline pressure. They may not know that some consumer tiers can use inputs to improve models, or that customer and employee data may be covered by contracts and data-protection rules.

What can go wrong

Confidential prices, source code, personal data about customers or colleagues, and unreleased strategy can leave your control. Even where a provider does not train on the data, sending it to an unapproved tool can breach client contracts, internal policy or data-protection obligations.

This is educational information, not legal advice. Data-protection obligations depend on your jurisdiction, your contracts and the specific data involved, so confirm your situation with a qualified professional.

Policy and tiers

Set a policy and choose the right tier.

Write a short internal AI policy

A usable policy is short enough that people read it. State which tools are approved, what data must never be pasted into any AI, when redaction or placeholders are required, and who to ask when unsure. Give a safe default answer for grey areas rather than assuming people will guess correctly.

Choose enterprise or API tiers with no-training terms

Where staff genuinely need AI for work data, an enterprise, business or API tier with written no-training and retention terms is usually safer than a free consumer login. Confirm the commitment in the contract and terms rather than inferring it from marketing, and check whether it applies by default or must be enabled.

Contracts and regions

DPA and data-processing considerations (hedged).

When a provider processes personal data on your behalf, a data-processing agreement (DPA) typically sets out what they may do with it, how long it is kept, whether sub-processors are used and where data is stored. Regions matter because different jurisdictions have different rules, and cross-border transfers may need extra safeguards. The points below are orientation, not legal advice.

Data-processing terms

Look for whether the provider trains on your data, the retention period, deletion rights and how sub-processors are disclosed. These belong in the contract and DPA, not in a chat window.

Regions and transfers

Where data is stored and processed can affect which rules apply. Some providers offer regional hosting or transfer safeguards; confirm what is actually available for your plan.

Get qualified advice

Obligations under laws such as GDPR and other regional regimes depend on your specifics. For anything material, confirm with a qualified data-protection or legal professional rather than relying on a general page.

Small-team checklist

A practical checklist for small teams.

Name approved tools: list the AI tools and tiers people may use for work data, and say what is off-limits.

Define red-line data: customer personal data, credentials, source code, contracts and unreleased plans should never go into unapproved tools.

Prefer no-training tiers: route real work data through enterprise or API plans with written no-training and retention terms.

Redact by default: require placeholders for names, IDs and figures whenever the full value is not needed for the task.

Check the paperwork: confirm DPA, retention, sub-processors and region for any tool handling personal or client data.

Train and review: brief staff on the policy, give them a person to ask, and revisit the list as tools and settings change.

Tier comparison

Consumer, Team or Enterprise, and API defaults (hedged).

Data handling often differs by tier, but defaults change and marketing pages are not contracts. Treat the table below as an orientation for the questions to ask, then confirm the specifics for the exact product and plan in its current terms.

ConsiderationConsumer / freeTeam or EnterpriseAPI
Training on inputs by defaultMore likely to use inputs to improve models unless a setting is changed; varies by providerOften advertises no-training defaults for business data; confirm in the contractFrequently described as not used for training by default; verify current terms
Data controls availableUsually limited to account-level settings and history controlsTypically adds admin controls, retention settings and policy optionsControlled through developer settings and the agreement rather than a chat UI
Contract and DPAStandard consumer terms; a data-processing agreement is uncommonBusiness terms and a DPA are more commonly availableCommercial terms and a DPA are usually available for review
Retention and deletionDefined by the consumer policy; options may be limitedOften configurable, with defined retention windows to confirmOften configurable through the agreement and settings
Best fitLow-risk, non-sensitive tasks with redactionTeam and company data where written commitments are neededBuilding products or workflows on real data with defined terms

This comparison is educational, not legal advice, and defaults differ between providers and change over time. Confirm training, retention, sub-processor and region details in the current official terms for the exact plan before relying on them.

Procurement checklist

Vendor questions to ask before you commit.

Before routing real work data through any AI tool, get clear written answers rather than inferring them from a marketing page. The questions below are a starting point; adapt them to your data, your contracts and your jurisdiction.

Training and reuse

Ask whether the provider trains on your inputs and outputs, whether that is the default or must be disabled, and whether any exceptions apply. Confirm the answer appears in the contract or terms, not only in a support article that can change.

Retention and deletion

Ask how long data is kept, whether retention is configurable, how deletion requests work and how quickly they take effect, and whether deletion covers backups and any derived data. Note that opt-out and deletion are usually separate processes.

Regions and sub-processors

Ask where data is stored and processed, whether regional hosting is available for your plan, which sub-processors are used and how changes are disclosed. Cross-border transfers may need extra safeguards depending on your jurisdiction.

These prompts are educational and not legal advice. Data-protection obligations depend on your jurisdiction, contracts and the specific data involved, so confirm anything material with a qualified professional and verify the provider's official documentation.

Related reading

Next steps.

ModelVersus modelversus.com

Compare AI models before choosing a workflow.

ChatGPT Alternatives chatgpt-alternatives.com

Find neutral alternatives when one AI provider is not enough.

AI Subscription Guide aisubscriptionguide.com

Learn which AI subscription fits your real work.

OpenAI crawler docs platform.openai.com/docs/gptbot

Check the current OpenAI crawler and user-agent guidance.

Google crawler docs developers.google.com/search/docs/crawling-indexing/overview-google-crawlers

Verify Google crawler names, indexing behavior and AI-related controls.

Anthropic support support.anthropic.com/

Check current Claude account, data and crawler guidance.

FAQ

Common doubts.

Can I stop AI training on my personal photos?

You can reduce exposure, but you cannot guarantee full control once photos are public. Keep sensitive photos private, review platform settings, avoid public high-resolution uploads and use crawler controls for your own website.

Should I paste my CV or resume into an AI chatbot?

Only after removing details you do not need for the task. Names, addresses, phone numbers, birth dates, signatures, employer secrets and references can often be replaced with placeholders.

Can I upload family documents to AI?

Be careful. Medical letters, passports, school documents, legal papers, bank files and family records can contain sensitive data about other people. Use approved privacy settings or do not upload them.

Do private prompts become AI training data?

It depends on the product, account type and settings. Some providers offer controls, team plans or API terms that treat data differently. Check the current settings before typing sensitive information.

Can I completely stop AI companies from training on my website?

Not with one technical switch. robots.txt can request that compliant crawlers stay away, but it is voluntary and does not remove copies that already exist elsewhere.

Does robots.txt legally block AI training?

robots.txt is a technical convention, not a contract by itself. It can support your policy position, but legal questions depend on jurisdiction, terms, copyright, contracts and enforcement.