For private people

Personal AI privacy starts before you upload.

Most AI privacy mistakes happen quietly: a scanned passport in a prompt, a child’s photo on a public page, a CV pasted into a chatbot, or private notes uploaded to summarize. This page helps you decide what is okay, what needs redaction and what should stay offline.

Personal risk map

Most people do not need a legal theory first. They need upload rules.

AI privacy often fails in ordinary moments: asking a chatbot to rewrite a complaint letter with your full address, uploading a medical PDF to summarize it, pasting a client email, or turning a child’s photo into an AI image. The question is not only “will this train a model?” It is “should this data leave my hands at all?”

  1. Green: public facts, generic examples, text you already planned to publish, anonymized drafts.
  2. Yellow: CVs, school work, personal letters, photos without sensitive context, work drafts after redaction.
  3. Red: IDs, bank documents, medical records, legal disputes, passwords, private photos of children, customer data and anything you do not have permission to share.

The best block is not uploading.

Crawler rules help public pages. They do not protect a private file you voluntarily paste into a tool. For private people, the strongest control is choosing what stays local.

Practical examples

How to use AI without oversharing.

Resume or CV

Remove phone, street address, birthday, signature and references. Ask for structure, tone and wording using placeholders.

Medical letter

Do not upload the full file casually. If you need plain-language help, remove names, dates, clinic numbers and identifying details first.

Family photo

Do not upload faces, children, home interiors or location-rich photos unless you are comfortable with the provider and product use.

Work email

Replace client names, project names, prices and internal conflict details. Keep the question generic where possible.

What counts as personal data

Personal data reaches AI systems through more paths than most people expect.

When people picture AI training, they usually imagine a company deliberately copying a document. In practice, personal information tends to arrive through many small, ordinary channels at once. Understanding those channels helps you decide where your attention is actually worth spending. None of this is legal advice, and the exact behaviour of each provider changes over time, so treat the descriptions below as a general map rather than a guarantee.

Prompts and chat history

Text you type into a chatbot, including follow-up questions, is data. Depending on the product, account type and settings, some of it may be retained, reviewed by humans for quality, or used to improve models. Consumer defaults and business or API terms often differ. Verify the current settings for the specific tool you use.

Uploads and attachments

Files you attach, such as PDFs, spreadsheets, screenshots and images, are processed by the provider. A single upload can contain far more than you intend, including metadata, hidden columns, comments and location tags embedded in photos.

Public posts and pages

Anything visible on the open web, including social profiles set to public, forum comments, reviews and personal websites, can be fetched by crawlers. Once content is public and copied, later deletion does not reliably remove every cached copy elsewhere.

Scraped social profiles

Public photos, bios, captions and connections on social platforms can be collected at scale. Platform settings and terms vary, and enforcement against unwanted scraping is inconsistent, so a "public" setting is best treated as "potentially collected".

Data brokers

Brokers assemble profiles from public records, purchases, app permissions and other sources, then resell them. This data can feed downstream systems. You often did not choose to be listed, which is why opt-out requests matter even though they are partial and slow.

Third-party sharing

Other people can expose your data too: a friend tags a photo, a colleague pastes an email thread that includes your address, or a family member uploads a shared document. You cannot fully control what others do, but you can reduce what you contribute.

A useful mental model.

Your exposure is roughly the sum of what you upload, what you publish, and what others publish about you. You have the most control over the first, meaningful control over the second, and only partial influence over the third. Spend your effort in that order.

Reduce your exposure

A practical checklist you can work through in an afternoon.

You do not need to do everything at once. Treat the list below as a menu, not a mandate. Each item lowers exposure a little; none of them makes you invisible, and that is an honest limitation rather than a failure of the approach.

Consumer defaults differ

How consumer AI training defaults vary by provider, and why you must verify.

Providers take different positions on whether consumer inputs are used to improve models, and those positions shift with policy updates, product tiers and regions. The table below describes the kinds of differences you may encounter. It is deliberately general and is not a current statement about any specific provider. Always confirm against the official documentation and in-product settings before relying on any of it, because this is exactly the sort of detail that changes without much notice.

FactorWhat it can meanHow to check
Account typeFree, personal paid, team and enterprise plans can be governed by different data terms.Read the terms tied to your exact plan, not a general marketing page.
Training defaultSome products use consumer inputs to improve models by default; others do not, or require opt-in.Look for a data-controls or privacy setting inside the app.
Opt-out availabilityAn opt-out may exist, may be limited to certain regions, or may not apply retroactively.Confirm whether the setting affects past data or only future inputs.
API vs appAPI and developer terms frequently treat data differently from the consumer chat app.Check the API data-usage policy separately from the app policy.
Human reviewInputs may be reviewed by people for safety or quality, distinct from model training.Read how the provider describes review, retention and staff access.

Verify, do not assume.

The safest assumption is that anything you would not want retained should not be typed or uploaded in the first place. Settings help, but they are only as reliable as your understanding of them on the day you use the tool.

Redact before you paste

Anonymizing content is the single highest-value habit.

Redaction is not about paranoia; it is about giving the tool only what the task requires. A well-anonymized prompt usually produces an equally good answer, because the model rarely needs your real address to fix a sentence. Below is a simple before-and-after pattern you can adapt.

Before:
"Rewrite this: Dear Dr. Meier, regarding my son Leon Muller,
patient ID 44-8821 at Zurich Clinic, appointment on 12 March..."

After:
"Rewrite this in a polite, formal tone: Dear [DOCTOR], regarding
my child [NAME], patient ID [ID] at [CLINIC], appointment on [DATE]..."

Names and relationships

Replace real names with [NAME] or a role such as [CHILD] or [MANAGER]. The wording help you want almost never depends on the actual name.

Numbers and IDs

Swap account numbers, patient IDs, phone numbers and dates of birth for placeholders. If the exact number matters for a calculation, consider doing that part offline.

Files and images

Crop out signatures, letterheads and reference numbers. For photos, remove location metadata and avoid backgrounds that reveal where you live or work.

None of this guarantees anonymity, since context can sometimes re-identify a person even without a name. It does, however, remove the most obvious and most damaging identifiers, which is a reasonable and achievable goal for everyday use.

Old content and opt-outs

Managing what is already public: search results and data brokers.

Reducing future exposure is easier than cleaning up the past, but the past still matters. This section covers realistic, if imperfect, ways to reduce what is already circulating. Expect partial results, delays and the occasional reappearance of removed content, because that is the honest nature of this work.

  1. Inventory: search your name, usernames and email addresses to see what is publicly indexed. Note the highest-risk items first, such as home address, phone number or sensitive photos.
  2. Source first: where possible, remove or privatize content at the original source, such as an old profile or a forum post, rather than only chasing search results.
  3. Search removal: major search engines offer removal request tools for certain sensitive personal information. These affect search visibility, not the underlying page, and criteria vary by region.
  4. Broker opt-outs: submit removal requests to the larger data brokers and people-search sites. Many require repeat requests over time, and some charge or make the process deliberately tedious.
  5. Maintain: re-check every few months. Opt-outs can lapse, and profiles can be rebuilt from fresh public data.

An important limitation.

Even a perfect removal effort cannot recall copies that were already cached, archived or scraped before you acted. Opt-outs and robots.txt are voluntary conventions that many, but not all, actors honour. Treat them as harm reduction, not as deletion.

Myths vs reality

Clearing up common misunderstandings.

Confident-sounding advice circulates quickly online, and some of it is misleading. The pairs below separate a few common myths from a more careful reading. As always, this is educational guidance rather than legal advice.

Myth: deleting a post removes it everywhere

Reality: deletion can remove the original, but copies may persist in caches, archives and datasets collected earlier. Deleting sooner is still better than later, since it limits how long content was available to copy.

Myth: a single setting stops all AI training

Reality: there is no universal switch. App settings, website crawler controls and broker opt-outs are separate layers, each partial. Reducing exposure means combining several of them.

Myth: private means invisible

Reality: marking an account private helps, but screenshots, re-shares and past public periods can leave traces. Treat "private" as reducing, not eliminating, exposure.

Myth: robots.txt legally forbids training

Reality: robots.txt is a voluntary technical convention. It can express your preference and support a policy position, but it is not, by itself, a contract, and compliance is not universal. Legal questions depend on jurisdiction and specific facts.

Myth: only what you upload matters

Reality: others can publish your data too. Managing your own uploads is necessary but not sufficient, which is why source removal and broker opt-outs also belong in the plan.

Myth: paid tools never train on you

Reality: some paid or business tiers do offer stronger data terms, but this is not automatic. Confirm the terms for your specific plan rather than assuming that paying changes the default.

Related reading

Next steps.

ModelVersus modelversus.com

Compare AI models before choosing a workflow.

ChatGPT Alternatives chatgpt-alternatives.com

Find neutral alternatives when one AI provider is not enough.

AI Subscription Guide aisubscriptionguide.com

Learn which AI subscription fits your real work.

OpenAI crawler docs platform.openai.com/docs/gptbot

Check the current OpenAI crawler and user-agent guidance.

Google crawler docs developers.google.com/search/docs/crawling-indexing/overview-google-crawlers

Verify Google crawler names, indexing behavior and AI-related controls.

Anthropic support support.anthropic.com/

Check current Claude account, data and crawler guidance.

FAQ

Common doubts.

Can I stop AI training on my personal photos?

You can reduce exposure, but you cannot guarantee full control once photos are public. Keep sensitive photos private, review platform settings, avoid public high-resolution uploads and use crawler controls for your own website.

Should I paste my CV or resume into an AI chatbot?

Only after removing details you do not need for the task. Names, addresses, phone numbers, birth dates, signatures, employer secrets and references can often be replaced with placeholders.

Can I upload family documents to AI?

Be careful. Medical letters, passports, school documents, legal papers, bank files and family records can contain sensitive data about other people. Use approved privacy settings or do not upload them.

Do private prompts become AI training data?

It depends on the product, account type and settings. Some providers offer controls, team plans or API terms that treat data differently. Check the current settings before typing sensitive information.

Can I completely stop AI companies from training on my website?

Not with one technical switch. robots.txt can request that compliant crawlers stay away, but it is voluntary and does not remove copies that already exist elsewhere.

Does robots.txt legally block AI training?

robots.txt is a technical convention, not a contract by itself. It can support your policy position, but legal questions depend on jurisdiction, terms, copyright, contracts and enforcement.