Anonymize a Word document for AI — in your browser

Meeting notes, interview transcripts, and case write-ups live in Word. Data Alias extracts the body text of a .docx entirely in your browser, turns each line into a reviewable row, and suggests in-place aliases for the emails, phone numbers, and other contact details scattered through the prose — so the document's meaning survives while the identities don't travel.

How it works, step by step

01

Upload your DOCX

Drop the .docx into the workspace (a Pro format, free during the beta). The body text is extracted locally in your browser — the document is never sent to a server.

02

Review what was detected

Detected emails, phones, and IPs are suggested for in-place aliasing with the surrounding sentences preserved; credential patterns default to removal. Optional on-device name detection (English and Korean) can suggest person names as aliases. Everything is reviewable before it applies.

03

Download the .safe.txt extract for the AI tool

You get an AI-ready .safe.txt extract of the document — not a rebuilt .docx. Hand that to the AI tool for summarizing or analysis, and keep the mapping locally to de-alias the response.

Example: session-notes.docx → session-notes.safe.txt

Why anonymize before you upload?

Because the answer depends on which tool you use, which account you are signed into, and what its defaults happen to be this month. A copy that never contained real identities does not depend on any of that. Here is what each vendor documents about its own handling.

ChatGPT (OpenAI)

Opting out of training is not the same as no retention

Consumer ChatGPT lets you opt out of model training via Data Controls, but OpenAI's own documentation notes retention for abuse monitoring — up to 30 days even for content you delete. An opt-out changes what training sees, not whether your pasted data reaches and briefly lives on OpenAI's servers.

Source: OpenAI Data Controls FAQ

Business-tier defaults are not consumer defaults

OpenAI documents that business tiers default to not training on customer data — but much real-world pasting happens in personal accounts, where the defaults and retention rules differ. A copy that never contained real identities doesn't depend on which account type someone used.

Source: OpenAI enterprise privacy

Claude (Anthropic)

The consumer training default changed in 2025

In August 2025 Anthropic updated its consumer terms, presenting the data-for-training choice pre-checked in favor of sharing — consumer users had to actively opt out to keep chats out of training. Defaults like this can change again; a file that only ever contained aliases is indifferent to them.

Source: Anthropic consumer terms update

Settings are per-account, your data policy isn't

Whether a given Claude account opted out is invisible to you when a file gets shared onward or pasted by a colleague. Anonymizing before upload is the control that travels with the file rather than with the account settings.

Source: Anthropic consumer terms update

Gemini (Google)

Human reviewers can read consumer conversations

Google documents that conversations in the consumer Gemini apps may be read by human reviewers to improve the service. That is a documented process, not a leak — and it is exactly why pasting a customer list into a consumer chat deserves an anonymized copy instead of the original.

Source: Google Gemini Apps privacy notice

A reviewed conversation should contain aliases, not identities

If a conversation can be sampled for review, the safest version of your file is one where the reviewer would see PERSON_001 and EMAIL_002 rather than real people. The analysis quality is the same; the exposure is not.

Source: Google Gemini Apps privacy notice

These are each vendor's own documented behaviors as of our research (July 2026). Policies and defaults change — an anonymized copy is the control that doesn't depend on them.

What to know about word (docx) files

Body text extraction, stated plainly

The extract is the document's body text. Formatting, comments, and embedded objects are not part of the output — the preview shows exactly what was extracted, so you can check nothing important was left behind before you rely on it.

Consistent aliases across a set of documents

With the Pro alias vault, the same participant keeps the same alias (PERSON_003) across every transcript in a project — the mapping is encrypted with your passphrase and stored only in your browser. Useful when a project spans many interview files.

Common questions

What parts of the document are extracted?
The body text. Headers, comments, and formatting are not included in the extract. The review table shows the extracted lines before anything is transformed, so you can verify what will be in the safe copy.
Why is the output a .txt file?
Data Alias produces an AI-ready text extract rather than rebuilding a Word file. Reconstructing a .docx would risk silent content changes; a plain extract is honest about what it is and pastes cleanly into any AI tool.
I anonymize interview transcripts — anything for researchers?
Yes. Consistent PERSON_00x aliases across interview files (via the encrypted alias vault), an on-device processing report, and a locally kept reversible mapping are built for qualitative research workflows. See the research page for the full mapping to IRB-style de-identification needs.

An honest note on detection

Data Alias reduces exposure risk but cannot catch every sensitive value, and no automated detector can. That is why nothing is applied without your review: check the safe copy before you share it with an AI tool or anyone else.

Anonymize a Word document now

No account needed. Files stay in your browser — verify it yourself.

Related guides

Every format is covered in the full guide index.