How to Safely Upload Company Data to AI (Checklist)

A checklist to run before you upload company data to AI: sort the input first, check which account it lands in, strip what does not need to be there, and know what to do if you already pasted it.

Published: Aug 23, 2026

8-12 mins

By Grace

Most advice about how to upload company data to AI safely starts with the wrong noun. It asks which tool is safe, when the thing that decides the outcome is what you are about to paste and which account it lands in. Same document, two accounts, completely different exposure.

This is the short version, meant to be run in the thirty seconds before you hit paste. It uses OpenAI as the worked example because OpenAI publishes its retention policy in enough detail to check, not because the others are safer.


Sort the data before you sort the tool

Ask one question about the input: if this ended up somewhere you did not choose, would anyone need to be told?

That question sorts almost everything. A blog draft, a published pricing page, a layout bug your customers can already see: no. A customer list, an unsigned contract, a colleague’s performance review, an unpatched vulnerability with the reproduction steps attached, anything with a name attached to a problem: yes. The second category is where the rest of this checklist applies, and it is smaller than people expect once they stop treating all work data as one bucket.

Do the sorting when the document is created, not when you are already in the chat box with a deadline. By then you will decide in your own favour.


Never paste credentials or client data into an AI tool

A handful of inputs are not a judgment call, because no retention setting fixes them:

  • Credentials of any kind. API keys, tokens, passwords, connection strings. Rotating a leaked key is cheap. Finding out it leaked is not.
  • Government identifiers. National ID numbers, passport numbers, tax IDs.
  • Payment details. Card numbers, bank account details.
  • Health records, or anything under a specific legal duty of confidence, unless you are working inside a system approved to hold it. A general-purpose assistant is not that system by default, and a HIPAA agreement in particular covers only certain products on certain plans.
  • Somebody else’s confidential material that you hold under an NDA or a client agreement. Permission to hold it does not automatically extend to handing it to a third party, although some agreements do allow approved subprocessors. Read the clause rather than guessing in either direction.

Everything else on this page is about reducing risk. This list is about not creating it, and the two entries with an “unless” attached are the ones where the answer lives in a contract rather than in a settings screen.


Account first, prompt second

The single largest variable is which account receives the text, and the differences are not marketing. On OpenAI’s own published terms, business tiers are not used for training by default, while consumer plans put that switch in your settings. Deleting a chat on Free, Plus or Pro removes it from your account immediately and schedules permanent deletion within 30 days, with an exception if OpenAI is “required to retain it for legal or security reasons.”

There is a second difference people discover late, and it is about the route rather than the reach. In ChatGPT Business, workspace admins “can view, access, export, and delete end user conversations in the workspace” directly. In ChatGPT Enterprise and Edu it runs through the Compliance Platform, which feeds eDiscovery, DLP and SIEM tools instead of a browsing screen, and conversation message content is among the data it can carry, behind a permission only a workspace owner can grant. So Enterprise adds a gate and an audit trail rather than putting your text out of reach. Neither arrangement is a scandal, and both are worth knowing before you type something personal into a workspace your employer pays for.

If you want the plan-by-plan comparison of what each vendor does with your text, ChatGPT vs Gemini vs Claude covers it properly. The point here is narrower: decide the account before you decide the prompt.

Decision path before you upload company data to AI, sorting the input by sensitivity and then routing it to the right account or stripping it first

Two decisions, in this order. Most people make the second one and skip the first.


How to strip company data before you paste it

For the large middle category, sensitive but useful, the fix is boring and effective: send the shape of the problem, not the record.

You rarely need the real customer name to get a better reply. “Draft a response to a customer who was charged twice and has emailed three times” works as well as the version with the account number in it, often better, because the model is not distracted by irrelevant detail. Replace identifiers with placeholders. Keep the structure, drop the specifics.

I do this with numbers rather than names. Real figures get rounded or shifted before they go in, and the answers stay useful, because what I am asking for is a direction rather than a computed result. Whatever comes back gets checked against the real figures before I act on it, which is my job either way. So the precision I would have been handing over was never doing any work.

This habit is also the one control on the list that works when nothing else is in place, and that used to be a wider gap than it is now. DLP grew up around files and email, so a paste into a chat box slipped past the older generation of it. Current endpoint tooling can see the paste: Microsoft Purview, for one, can warn on or block sensitive text pasted into third-party AI sites in Edge, Chrome and Firefox. If your organisation has none of that, and no gateway or approved tool either, stripping the input is most of what you have left. It covers a lot of everyday work. It does not cover re-identification from surrounding context, a trade secret that survives having the names taken out, or a contract you were never permitted to share the material under.


Upload company data to AI when you have no budget for tooling

Small teams read enterprise advice and conclude the safe path costs money. Three things cost nothing:

  1. Write down what may not be pasted. Not a policy document, a short specific list, closer to the one above than to a compliance framework. “Do not paste client PII, credentials, or unsigned contracts into any AI tool” gives someone a decision they can make in a second.
  2. Name one approved tool and make it the easy one. Bans push people toward personal accounts where you have no visibility at all, which is worse than the thing you banned. Shadow AI at work covers what that costs when it goes wrong.
  3. Turn off training on the accounts you already use. In ChatGPT the setting is called “Improve the model for everyone” and sits under Settings, then Data controls. It takes a minute per account and it is the one setting with no tradeoff.
The Improve the model for everyone toggle switched off in ChatGPT's Data controls settings

One minute, one account, no tradeoff. This is the setting people mean when they say they turned training off.


What to do if you already pasted company data into AI

Most people reading this have already done it once, and the useful response is a sequence rather than a feeling.

Start with what it was, using the “never” list above. If it was a credential, rotate it now and treat the rest as secondary. If it was somebody else’s confidential material, the obligation may be contractual, and that is a conversation rather than a settings change.

ChatGPT's Clear your chat history confirmation dialog, which warns that it will delete all chats but says nothing about how long deletion takes

The most destructive button in the settings, and it mentions neither the 30 days nor the legal exception. It does say memories are handled separately, which is the part people miss.

Then delete the conversation, understanding what deletion does and does not do. On consumer plans it leaves your account immediately and is scheduled for removal within 30 days, with the legal and security exception noted above. It is a real reduction in exposure, not a guarantee, and it does not undo any onward use that already happened.

Memory is a separate place, and this is the step people skip. Clearing chats does not clear what the assistant saved to memory, so if the thing you pasted was interesting enough to be remembered, deleting the conversation leaves it exactly where it was. ChatGPT’s own delete dialog points you at a different settings screen for that.

Then turn training off, if it was on, so the next paste does not add to this one. And then leave it there. For ordinary work material, one paste into a mainstream assistant with training off and the chat deleted is a small event, and treating it as a catastrophe mostly teaches people to hide the next one. It stops being small when the material was regulated, contractual, or somebody else’s, because those obligations do not move with a settings toggle. That is why the sorting at the top of this page matters more than the cleanup at the bottom.


FAQ

Q

Is OpenAI legally required to keep my deleted ChatGPT chats?

The blanket order has ended, and the belief that it is still in force keeps circulating anyway. A court order in the New York Times litigation did force OpenAI to retain consumer ChatGPT and API content, including deleted chats, but OpenAI says its “obligations under the earlier order ended on September 26, 2025” and that deleted conversations and Temporary Chats are again removed within 30 days. Two things survive that. At the Times’ continuing demand, OpenAI still stores a specific set of April to September 2025 user data, which it says is locked down and has not been turned over to anyone. And the ordinary policy has always allowed longer retention where OpenAI is required to keep something for legal or security reasons, so “deleted within 30 days” is the normal case rather than a guarantee.
Q

Can my workspace admin read my ChatGPT conversations?

Assume yes, and treat the useful question as how rather than whether. On ChatGPT Business, workspace admins can view, access, export and delete end user conversations directly. On ChatGPT Enterprise and Edu it runs through the Compliance Platform into eDiscovery, DLP or SIEM tooling, where conversation message content is one of the permissions available, grantable only by a workspace owner. Enterprise therefore adds a gate and an audit trail rather than putting your text beyond reach. A work account is a work record.
Q

Why are lawyers more cautious about ChatGPT retention than about Google Workspace or Microsoft 365?

Mostly because the obligation runs to the client, not to the tool. Firms typically have long-standing agreements and established practice covering their existing document platforms, whereas an AI assistant is a newer third party that may not be named in a client engagement letter at all. The retention window matters less than whether disclosing the material to that provider was permitted in the first place.
Q

Is company data leaking into AI tools a real problem or mostly theoretical?

It is measured rather than theoretical, and the numbers are the subject of our shadow AI guide rather than this checklist. The practical framing for a small team is that the common case is not a breach, it is an ordinary paste into a personal account that nobody logged, which is why the fix is a habit rather than a product.

Sources

Retention and admin-visibility details are OpenAI’s own published policy as of August 2026 and are used here as a worked example, not as an industry norm. Other providers differ; check each one’s current terms. Practices described for small teams reflect ongoing practitioner discussion rather than a formal survey.

Written by Grace

I test AI tools and agents in my own workflow, and write down what I find, including the settings that surprised me. About Grace and how posts are verified

Leave a Comment