Account type, settings, approvals, and identifiers to check before any file goes out
Why this question has no single answer
The ticket export is open in one browser tab and ChatGPT is open in the next. The analysis is due Thursday — inquiry categories, repeat contacts, a first look at where response times slip — and every row of the file carries a customer's name, an email address, and a free-text message describing whatever went wrong. Attaching the file would take a few seconds. Deciding whether to attach it takes longer, because the answer depends on several things that vary independently: which account the tab is signed into, what that account's settings say about training and retention, whether your organization has approved the service for customer data, what your contracts allow, and what is actually in the rows.
No single setting resolves all of those, and no provider policy can either, because the questions span your organization, your contracts, and your file as well as the service. The title says ChatGPT because that is the tab most people have open, but the same questions apply to Claude, Gemini, or any other AI service — and since each provider writes its own terms, the answers have to be checked per provider and per account. This post works through the questions in order and ends with the checklist version.
What organizations restrict — and what employees do anyway
Organizations across jurisdictions write different rules, but a recurring operational control is to keep identifiable customer data out of an unapproved AI service. The canonical example is also the one that shows why the control exists: in May 2023, after engineers pasted internal source code into ChatGPT, Samsung Electronics banned generative AI on company devices and networks and told employees not to submit company information through personal devices either (Bloomberg, CNBC). The reported reasoning is the same worry a support lead has about a ticket export: once data reaches an external system, retrieving or deleting it is no longer yours to do. Bans of that shape have since appeared in internal policies in many countries, tightening or loosening as providers changed their terms. How closely employees follow them is a separate question, and a measurable one.
(Updated August 2026) When KPMG and the University of Melbourne surveyed more than 48,000 people across 47 countries in 2025, almost half of the employees among them — 48% — admitted to using AI in ways that contravene their company's policy (KPMG).
Netskope, which measures enterprise network traffic rather than asking people, fills in where that use goes: in 2025 it found 72% of generative-AI use inside companies running through personal accounts (Netskope) — shadow AI in the textbook sense, invisible to whoever wrote the ban. Its 2026 report shows the pressure still building: generative-AI data-policy violations doubled year over year, with the average organization seeing 223 attempts per month to put sensitive data into a prompt (Netskope 2026). Keep the scope in mind — both figures come from Netskope's own enterprise network telemetry, meaning companies that already bought monitoring. The unmonitored ones are unlikely to look better.
What goes in is also getting more sensitive. Cyberhaven — measuring its own DLP customers, a self-selected sample, so treat the absolute numbers with some care — put the sensitive share of what employees pasted into ChatGPT at around 11% in early 2023 (Cyberhaven); by its 2025 report, now counting corporate data flowing into AI tools generally, that share stood at 34.8% (2025 AI Adoption & Risk Report).
Read together, the numbers describe a gap that written rules alone have not closed. When approved channels are blocked or missing, work moves to personal accounts, and personal accounts are exactly where none of the controls in the next two sections apply. So the account is the first thing worth examining.
Consumer accounts vs. business workspaces
The same chat window can sit on top of two different agreements. A consumer account — free or individually paid, created with any email address — operates under the provider's consumer terms. A business or enterprise workspace is procured by the organization, administered centrally, and covered by business terms, which can include negotiated commitments about training and data handling and, where agreed, a data-processing agreement. Which one an upload happens in changes what the provider has promised about the file.
The split shows up in each provider's own documents. OpenAI publishes separate privacy commitments for its business products (OpenAI Enterprise Privacy). Consumer terms also move over time: in August 2025, an update to Anthropic's consumer terms presented users with a choice screen on which training consent was pre-checked (Anthropic, TechCrunch) — a reminder that a setting verified once is not verified permanently. And Google's help pages note that Gemini consumer conversations may be subject to human review for quality purposes (Google). Each of the three writes its own terms on its own revision schedule, so a setting confirmed in one product tells you nothing about another, and little about the same product a year later.
The practical consequence follows directly: uploading a customer file through a personal account routes around whatever the organization negotiated.
Controls that are separate questions
It is tempting to collapse all of this into one switch, usually "training is off, so it's fine." The controls are independent, and each needs its own answer for the specific account that will receive the file.
- Training. Whether the provider may use your prompts and files to improve its models. Providers document opt-outs and per-tier commitments, but what the default is, which tier it applies to, and what counts as an exception vary by provider and shift with terms updates. Check the current documentation for your account rather than a remembered setting or a screenshot from a blog post — including this one.
- Retention. How long the data exists on the provider's systems, which is a separate question from training. OpenAI's developer documentation states that abuse-monitoring logs "are generated for all API feature usage and retained for up to 30 days, unless longer retention is required by law, or is reasonably necessary to protect our services or any third party from harm" (Your data). That page describes the API; the consumer product is governed by its own Data Controls FAQ, which has stated a comparable 30-day window for deleted conversations. The two are documented separately and can change separately, which is the point — a training opt-out governs what a model learns from, and it does not by itself decide how long the data is stored, or under which product's policy.
- Human review. Whether people at the provider may read conversations for quality or abuse purposes — the Gemini note above is one published example. Model training and human readers are different exposure paths, and a setting that addresses one may say nothing about the other.
- Data-processing agreement. Whether a DPA or equivalent contract covers this category of data. Business workspaces can carry one; a personal account runs on standard consumer terms.
- Organizational approval. Whether your organization has cleared this service for customer-derived data. A service can be configured carefully and still be unapproved.
- Account type. The previous section's distinction, restated as a control: every answer above depends on which workspace the file actually enters.
None of these answers implies another. A no-training setting does not create a DPA, a DPA does not constitute internal approval, and approval does not configure retention.
What a pseudonymized copy changes — and what it does not
Suppose the service is approved, the workspace is the right one, and the settings check out. The file itself is the remaining variable. Preparing a pseudonymized copy — replacing names, emails, phone numbers, and other identifying values with consistent aliases before anything is attached — changes what the upload contains. Consistency is what keeps the copy useful: when the same customer is PERSON_001 in every row, the analysis can still count repeat contacts and trace journeys, and an answer about PERSON_001 can be translated back to the real account afterwards, using a mapping you keep on your machine.
Privacy law recognizes the technique, within limits. The GDPR names pseudonymization as a recognized safeguard (Art. 4(5), Recital 28), and other regimes carve out narrow permissions on the same basis: Korea's Personal Information Protection Act, Article 28-2, allows processing pseudonymized information without individual consent for statistics, scientific research, and archiving in the public interest (PIPC overview, in English; the statute text is Korean-language). Recognition as a safeguard is not permission to upload. Pseudonymized data generally remains personal data for whoever holds the mapping; the European guidance and case law are summarized in Pseudonymized data after EDPS v SRB: what changed in Europe.
Equally important is what the copy leaves untouched. It does not make an unapproved service approved, put a contract in place where none exists, alter a retention setting, or decide who sees the output. Those questions keep the answers they had in the sections above; the copy changes one item on the checklist — what the file contains when it goes out. Preparing the copy is also the right moment to cut scope: columns and rows the analysis does not need add exposure without adding anything to the result, so leave them out. Plan for the output as well, because once an AI answer is translated back to real names it is customer data again, and it belongs under the same access controls as the source file.
Direct identifiers are only the first pass. Before upload, scan for combinations that narrow a row to one person: precise timestamps joined to locations, rare titles, long account histories, free-text descriptions of one-off incidents. If a row would let a colleague guess the customer, treat it as identifiable no matter how it is labeled.
Find and replace can produce a copy like this in principle, though consistency collapses quickly once identifiers repeat across columns and hide inside free text. If you use a tool to produce the copy, check that it processes the file in the browser rather than on its own server. How browser-local file processing works—and how to verify it documents the offline test and the network check. For a worked example on a ticket export, see How to prepare support tickets for AI analysis without sending direct identifiers.
Data Alias is one browser-local implementation of this preparation workflow: it detects names, emails, and phone numbers in a spreadsheet, applies consistent aliases, and shows every proposed change for review before export.
The pre-upload checklist
The sections above, compressed to the order worth checking. A customer file — original or pseudonymized copy — goes to an AI service only after all seven have answers you could repeat to your team.
- Is this AI service approved? Approved by your organization for customer-derived data, not merely popular or configured with care.
- Which account or workspace am I using? A business workspace and a personal login on the same product operate under different terms.
- What are the training and retention settings? Check both, in the account that will receive the file, against the provider's current documentation.
- Do contracts or internal policies restrict this use? Customer agreements, DPAs, and internal data policies can each rule out an upload that the settings would allow.
- Have direct and indirect identifiers been reviewed? Automatic detection narrows this work; it does not finish it.
- Is the minimum necessary data being shared? Columns and rows the analysis does not need should not travel with it.
- Is the output handled under the same controls? An answer about PERSON_001 becomes customer data again the moment it is translated back.