Anonymize a PDF for ChatGPT — in your browser
Reports, transcripts, and exported documents often exist only as PDFs. Data Alias extracts the PDF's text layer inside your browser — the parser is loaded on demand and runs locally, so the document is never uploaded. Each line becomes a reviewable row, detected contact details are aliased in place, and you download an AI-ready text extract.
How it works, step by step
Upload a text-layer PDF
Drop the PDF into the workspace (a Pro format, free during the beta). The text layer is extracted in your browser; scanned, image-only PDFs are declined with a clear error because Data Alias does not do OCR — it will not pretend to read what it can't.
Review the in-place replacements
Detected emails, phone numbers, and IP addresses are suggested for in-place aliasing — the sentence around each value stays readable — while credential patterns like API keys default to removal. You review every suggestion before it is applied.
Download the .safe.txt extract for ChatGPT
The output is an AI-ready .safe.txt extract of the document's text, not a rebuilt PDF. Paste or upload it to ChatGPT; the alias mapping stays on your machine so you can restore real names in the answer.
Example: interview-04.pdf → interview-04.safe.txt
Why anonymize before ChatGPT?
Opting out of training is not the same as no retention
Consumer ChatGPT lets you opt out of model training via Data Controls, but OpenAI's own documentation notes retention for abuse monitoring — up to 30 days even for content you delete. An opt-out changes what training sees, not whether your pasted data reaches and briefly lives on OpenAI's servers.
Source: OpenAI Data Controls FAQ ↗Business-tier defaults are not consumer defaults
OpenAI documents that business tiers default to not training on customer data — but much real-world pasting happens in personal accounts, where the defaults and retention rules differ. A copy that never contained real identities doesn't depend on which account type someone used.
Source: OpenAI enterprise privacy ↗These are OpenAI's own documented behaviors as of our research (July 2026). Policies and defaults change — an anonymized copy is the control that doesn't depend on them.
What to know about pdf files
Text layer only — no OCR
Data Alias reads the embedded text layer. Scanned pages without one produce an honest error instead of a silently empty extract. If your PDF came from a scanner, run OCR elsewhere first or work from the source document.
The extract preserves reading context
Each line of extracted text becomes a row you can review. In-place aliasing keeps the surrounding sentences intact, so summaries and thematic analysis still work on the safe extract.
Common questions
- Does it work with scanned PDFs?
- No. Data Alias extracts the text layer only and does not perform OCR. A scanned, image-only PDF is rejected with a clear message rather than producing an empty "safe" file.
- What does the output look like?
- A .safe.txt file — an AI-ready extract of the document's text with detected values replaced in place. Data Alias does not reassemble a PDF, and doesn't claim to.
- How are names inside the text handled?
- Emails, phones, and IPs are detected by pattern and aliased in place. Person names can be detected with an optional on-device model (English and Korean) that runs in a browser worker — you start the model download explicitly, and detection results are suggestions to review — detection can miss values, which is why the review step exists.
An honest note on detection
Data Alias reduces exposure risk but cannot catch every sensitive value, and no automated detector can. That is why nothing is applied without your review: check the safe copy before you share it with ChatGPT or anyone else.
No account needed. Files stay in your browser — verify it yourself.