ForgeAI / TOFU / shadow AI

Can Employees Use ChatGPT Without Exposing Company Data?

What happens to business data pasted into AI tools, why shadow AI is hard to control, and what changes when AI runs on hardware you own.

  • Shadow AI
  • Data privacy
  • Private AI
  • ForgeAI

On the right account, with the right training, mostly yes. But most employees have a personal ChatGPT account open, unmanaged and unmonitored. In that case, the honest answer is that you don’t know what’s leaving your building, and neither do they.

That’s not a hypothetical risk. It’s already happened at companies with security budgets far larger than most: Samsung engineers pasted proprietary source code into ChatGPT to debug it, uploaded equipment-defect data to optimize a process, and fed an internal meeting recording into the tool to generate minutes. All of this happened within a single month of the company lifting its ban. Samsung capped input length as an emergency fix, then banned the tool outright. Apple, JPMorgan, Bank of America, Verizon, and several others restricted it soon after.

What Actually Happens to Data Typed Into ChatGPT

The tier matters more than almost anything else here. On Free and Plus accounts, OpenAI may use conversations to improve its models unless the user manually opts out in settings. Most employees never touch that control because they do not know it exists. On Team, Enterprise, and API access, business data isn’t used for training by default.

Even with training turned off, data still gets processed and retained for a period, typically up to 30 days, for abuse monitoring before deletion. And “OpenAI will delete it” turned out to be a more conditional promise than it looked. In 2025, a court order in unrelated New York Times litigation required OpenAI to preserve consumer ChatGPT and API logs that would otherwise have been deleted on schedule. It overrode both the standard retention window and users’ own deletion requests, though ChatGPT Enterprise and zero-data-retention API accounts were excluded from the order. The preservation requirement was later lifted, but it’s a real example of a vendor’s deletion policy bending under external pressure that had nothing to do with your company.

None of this makes ChatGPT unsafe to use. It means the safety of any given conversation depends entirely on which account it happened in and what your employee decided to type. No IT policy fully controls those variables once the tool is open in a browser tab.

Shadow AI, in Numbers

“Shadow AI” is employee use of AI tools, models, or embedded features without IT’s knowledge or approval. It is the default state at most companies now, not the exception.

A 2025 SAP/WalkMe workplace survey found 78% of employees use AI tools their employer never provided, and 51% report getting conflicting guidance about what’s actually allowed. Only about 7.5% received substantial AI training; nearly a quarter got none at all.

Cisco’s 2025 Data Privacy Benchmark Study, surveying 2,600 privacy and security professionals across a dozen countries, found 64% worry about accidentally sharing sensitive information through AI tools. Yet close to half of respondents admitted doing exactly that. Broken down: 63% had input public company information, 60% internal process details, 46% employee names or personal data, and 42% non-public company information into a generative AI tool.

The financial exposure isn’t abstract either. IBM’s 2025 Cost of a Data Breach report, based on 600 organizations across 16 countries, found breaches involving shadow AI cost an average of $4.63 million, roughly $670,000 more than a typical breach. One in five breached organizations in the study had been compromised through shadow AI specifically, and 97% of those lacked proper AI access controls at the time.

Why It’s Hard to Catch

Most data-loss-prevention tools were built to flag file transfers and email attachments, not conversational text typed into a browser over an encrypted connection to a legitimate domain. An employee pasting a client contract into ChatGPT looks, to most security tooling, identical to someone reading the news. That’s part of why the Cyberhaven analysis of 1.6 million knowledge workers found 11% of everything pasted into ChatGPT was confidential. More telling, fewer than 1% of employees accounted for 80% of all confidential leaks. This usually isn’t malicious. It’s someone trying to work faster, without a clear sense of where the line is.

Building a Policy That Employees Actually Follow

A blanket “never put company data into AI” policy sounds safe and works badly. It drives usage underground instead of preventing it, since employees under deadline pressure will use whatever tool gets the job done fastest. A workable acceptable-use policy needs a few specific things instead of a single blunt rule.

Start with data classification: define what counts as public, internal, confidential, and regulated, and be explicit about which tiers are safe for which category. Publish an approved-tool list, including a real process for requesting new ones, so employees aren’t stuck choosing between “use nothing” and “use whatever I found.” Add monitoring that’s transparent rather than covert. Usage audits and network visibility, disclosed up front, work better than a surveillance program employees discover after the fact. Build a low-friction incident reporting path for the moment someone realizes they pasted something they shouldn’t have, because that moment will happen, and how fast it gets reported matters more than whether it happens at all. Train people specifically. The WalkMe numbers above make clear this is the single biggest gap most companies have. Review the policy at least quarterly, since AI tool capabilities are shifting faster than most companies update anything else in their compliance calendar.

Most existing shadow-AI policies also miss the tools that don’t look like AI tools, such as the Copilot feature quietly embedded in Microsoft 365 or the AI summarizer built into a note-taking app. Scope your policy to capabilities, not just standalone chatbots, or you’ll write a policy that misses where most of the actual exposure lives.

Cloud AI With a Data Agreement vs. AI That Never Leaves

A cloud LLM with a data processing agreement or BAA in place is faster to stand up and scales easily, but the protection is contractual. Data physically leaves your environment, and your safeguard is a vendor’s configuration and legal promise, which, as the 2025 court order showed, can be overridden by circumstances that have nothing to do with your company.

A private, on-premise deployment inverts that. Data never leaves your controlled environment in the first place, which satisfies data residency requirements by default, supports fully air-gapped setups for the most sensitive work, and puts complete audit logging and access control inside systems your own team runs. The tradeoff is real: higher upfront hardware cost, more operational responsibility, and a longer setup than signing up for a cloud account. For companies handling regulated, proprietary, or client-confidential data as a matter of routine, the tradeoff increasingly reads as the cost of actually controlling the answer to “where did that data go,” instead of trusting a policy document to answer it for you.

FAQ

What percentage of employees use AI tools without approval? Recent industry surveys put it around 78 to 80%, depending on the study and industry. That means shadow AI is the default condition at most companies, not an edge case.

Does ChatGPT train on data my employees paste in? It depends on the account tier. Free and Plus may use conversations to improve models unless the user opts out manually. Team, Enterprise, and API access don’t use business data for training by default.

Is a signed data agreement with an AI vendor enough to protect confidential information? It reduces risk but doesn’t eliminate it. A vendor agreement governs what happens after data reaches the vendor’s systems. It does nothing to stop an employee from pasting confidential material into the wrong account in the first place.

What’s the fastest way to reduce shadow AI risk without banning AI outright? Publish a specific, tiered acceptable-use policy with an approved-tool list and genuine training, rather than a blanket ban. Bans without alternatives tend to push usage further out of sight rather than stopping it.

Bottom Line

The data your employees paste into AI tools is exposed the moment it leaves your environment, regardless of what a vendor’s policy says about training or retention. Shadow AI use is already the norm at most companies, whether or not anyone has written a policy for it. A tiered acceptable-use policy with real training closes most of the gap. Keeping the AI itself inside your own infrastructure closes the rest of it, because at that point there’s no third party’s promise to depend on in the first place.


Sources and further reading

ForgeAI / NEXT STEP

Keep the boundary where you can see it.

See ForgeAI

KEEP READING