Skip to main content

What happens to your data in ChatGPT, Claude, Gemini, Copilot, DeepSeek and Qwen?

· 13 min read

A common warning about generative AI says that anything typed into a chatbot becomes part of the model. That is too crude to be useful. What actually happens depends on the provider, the product, whether the account is consumer or enterprise, the settings in force, whether feedback is submitted and whether the organisation has negotiated different retention terms.

For Māori organisations and GIS teams, this matters because the difference between a public chatbot and an enterprise or local system can be substantial. A planning report, shapefile attribute table, field note, Treaty research document or set of candidate place names may be handled very differently depending on the service used.

The first technical point is simple: entering a prompt normally causes inference. The already-trained model processes the current context and produces a response. The model is not usually changing its weights while the user types. After that inference, however, the service may retain the interaction, and in some consumer products the retained material may be eligible for later model improvement. Those are separate stages.

This article looks at current published policies as at 4 September 2026. They change, so the exact product terms should always be checked before important Māori or organisational data are used.

Three questions, not one

When assessing an AI service, ask three separate questions.

First, what happens during the current request? The prompt, files, retrieved passages and other context have to be processed somewhere so the model can generate an answer.

Second, what is retained after the request? The provider may keep chat history, logs, uploaded files, safety records or audit material for a period of time.

Third, can retained material be used for future model training, evaluation or product improvement?

A provider can answer “no” to the third question while still answering “yes” to the first two. This is why a statement such as “our data is not used to train the model” is important but incomplete.

For Māori data sovereignty, the location and authority around processing still matter even where training is excluded.

Personal ChatGPT

For personal ChatGPT services, OpenAI currently allows users to control whether new conversations are used to improve models. In personal Free, Plus and Pro workspaces, the setting commonly described as Improve the model for everyone controls that use. When it is disabled, new conversations are not used to train OpenAI’s models through the ordinary model-improvement process.

That does not mean the prompt never reaches OpenAI infrastructure. The service still needs to process the conversation to provide the answer. Chat history may remain available to the user, subject to the product’s retention and deletion behaviour.

Temporary Chat offers another option. OpenAI says Temporary Chats are not used to train models, do not normally appear in history and do not create memories. A copy may still be retained for a limited period for safety purposes.

There is also an important feedback exception. OpenAI documentation says that when a user deliberately provides feedback on a response, the conversation associated with that feedback may be used to improve models. This is another example of why product settings and user actions matter.

See OpenAI’s current data controls guidance, Temporary Chat guidance and how your data is used to improve model performance.

ChatGPT Business, Enterprise and the API

OpenAI’s business products are materially different from ordinary personal ChatGPT.

OpenAI says it does not train its models on customer business data from ChatGPT Business, Enterprise, Edu and its API Platform by default. Organisations can opt in to share some data, but the default business position is different from the ordinary personal product.

This distinction is significant for an iwi, trust, council or public-sector organisation. It means a governance discussion should not simply say “ChatGPT trains on everything”. It should identify the exact plan and contract being used.

However, no-training terms do not mean zero retention. The API can retain information for abuse monitoring under standard arrangements, generally for a limited period, while eligible customers can obtain stronger controls including zero-data-retention arrangements for supported use cases.

For organisational GIS, this means an enterprise/API architecture can sometimes be appropriate for ordinary internal material where the organisation has completed the necessary privacy, security, procurement and Māori data-governance assessment. More sensitive material may still justify a local or disconnected architecture.

See OpenAI’s business data privacy commitments and enterprise privacy.

Claude consumer accounts

Anthropic also separates consumer and commercial use.

For Claude Free, Pro and Max, Anthropic’s current privacy material explains circumstances in which conversations may contribute to model improvement, including where the user has chosen to allow model improvement or where material is involved in certain safety-review processes. The exact controls should be checked in the current account settings.

Anthropic has also described longer retention for data that enters its model-improvement pipeline than for ordinary conversations. This matters because a user may think they are making a simple choice about whether to help improve Claude when the practical consequence can include much longer retention of selected material.

For Māori information, the practical position is straightforward: do not rely on a generic statement that Claude is either private or not private. Check whether the account is a consumer account, what model-improvement setting is enabled and what the current retention terms say.

See Anthropic’s consumer model-training explanation and privacy updates.

Claude for Work and API

Anthropic’s commercial products again operate differently. Anthropic says inputs and outputs from its commercial products, including the API and Claude for Work, are not used to train models by default.

Standard API retention is generally limited, with stronger zero-data-retention arrangements available for some approved customers and use cases.

That makes Claude API or commercial Claude a different proposition from a personal Claude account. A Māori organisation evaluating the service should still examine jurisdiction, subprocessors, access controls, retention, connectors and the authority attached to the source information, but the shared-model training question can be addressed more directly under commercial terms.

See Anthropic’s commercial data training policy.

Gemini consumer accounts

Google’s consumer Gemini environment has a particularly visible relationship between account activity settings and model improvement.

Google says that where Gemini Apps Activity or equivalent activity retention is enabled, chats and material shared with Gemini can be used to provide and improve Google services, including generative AI models. Some material may be reviewed by humans under controlled processes. Google advises users not to enter confidential information they would not want reviewed or used in that way.

Where the relevant activity setting is disabled, future chats are not used for model training through the ordinary process, although Google may still retain conversations for a shorter period for service and safety purposes. Temporary-chat options provide another separation from ordinary activity history.

For GIS, this means attaching an internal report, image or spreadsheet to a personal Gemini account can create a substantially different data-governance position from using Gemini through an approved Workspace environment.

See Google’s Gemini Apps privacy hub.

Google Workspace with Gemini

Google Workspace provides enterprise protections that are different from the consumer service. Google says Workspace customer content, prompts and generated responses are not used to train generative models outside the customer’s domain without permission.

The practical Māori data-sovereignty assessment still needs to examine where the processing occurs, contractual arrangements, administrators, data residency where relevant and whether the proposed use is authorised. But again, the statement “Gemini trains on whatever we type” is too broad to describe Workspace correctly.

See the Generative AI in Google Workspace Privacy Hub.

Microsoft 365 Copilot

Microsoft 365 Copilot is a useful example of why training and logging should be separated.

Microsoft says prompts, responses and organisational information accessed through Microsoft Graph under Microsoft 365 Copilot’s enterprise protections are not used to train foundation models. Existing Microsoft 365 permissions remain important because Copilot can surface material a user is already entitled to access.

At the same time, Copilot interactions can be logged and retained within enterprise compliance systems. Administrators may have audit, eDiscovery and Purview controls over that material. If web grounding is enabled, search queries derived from a user’s request may also be sent to Bing.

So “Microsoft does not train the foundation model on our Copilot prompts” can be true while “the interaction is retained and governed inside our Microsoft environment” is also true.

For Māori organisations, that means permissions design becomes extremely important. AI can make existing information much easier to discover. A staff member who technically had access to 20,000 files before Copilot may suddenly be able to ask a conversational question that retrieves information from across that collection in seconds.

See Microsoft’s enterprise data protection and Microsoft 365 Copilot privacy guidance.

DeepSeek

DeepSeek’s consumer service sits under a different jurisdictional and corporate environment from the major US providers, but the technical questions remain familiar.

DeepSeek’s current privacy and terms material says the service can collect prompts, text or voice inputs, uploaded files, photos, feedback and chat history. Its terms provide mechanisms for users to opt out of some use of their data for model training or technology optimisation.

DeepSeek also describes limited use of de-identified inputs and outputs for service and technology improvement under its terms. The company is based in China, and its privacy material describes processing within its corporate group for functions including storage, research and development and model optimisation.

For public or low-risk experimentation, DeepSeek may offer useful capability. For important Māori organisational information, the jurisdictional, contractual and retention questions should be explicitly considered rather than focusing only on whether the model is technically strong.

See DeepSeek’s current privacy policy and terms of use.

Qwen consumer services

Alibaba’s Qwen consumer service also distinguishes between using a model and operating under an enterprise cloud contract.

Qwen’s consumer terms provide Alibaba with broad rights to process User Content in order to provide and improve the service. They also describe circumstances in which non-personal content can contribute to development and improvement of machine-learning and AI technologies. The service may involve processing or storage outside the user’s jurisdiction.

Qwen’s own training-data disclosure says user interactions can be used in some circumstances and describes opt-out mechanisms.

That does not make Qwen uniquely problematic. It means the consumer product should be treated as a consumer cloud AI service, not as a private local model merely because open-weight Qwen models also exist.

See Qwen Terms of Service and the Qwen training-data summary.

Alibaba Cloud Model Studio

Alibaba Cloud’s enterprise Model Studio is a different product. Alibaba says customer business data is not used to develop or improve models without explicit consent.

This distinction is exactly why product names are not enough. “Qwen” can refer to a public consumer service, an open-weight model running on a local computer, or a model provided through an enterprise cloud platform. Those three architectures have very different data pathways.

See Alibaba Cloud’s Qwen and Wan training-data disclosure.

Open-weight models run locally

A downloadable model such as Llama, Qwen, Gemma, Mistral or another open-weight model can operate without sending prompts back to the original developer, provided the software around it is genuinely configured for local use.

This is a major distinction.

If the model file, embeddings, document parser, vector index and inference runtime all operate locally, an iwi report does not need to be transmitted to Meta, Alibaba, Google or another model developer simply because their model weights are being used.

But the surrounding application matters. A local model can still sit inside software that uses web search, cloud embeddings, remote plugins, telemetry or an online geocoder. The data-flow diagram must cover the entire system, not only the language model.

A GIS example

Imagine an iwi GIS team with environmental reports, a GeoPackage of monitoring sites and a collection of field notes.

Using a consumer chatbot, the team might upload a report and ask for candidate locations. The report leaves the organisation and is processed under that consumer service’s terms.

Using an enterprise cloud model, the same report still leaves the local environment for inference, but shared-model training may be contractually excluded and retention more tightly controlled.

Using local RAG, the report can be indexed locally. The model can help retrieve references to erosion, access and historical place names while the source files and search index remain on local infrastructure.

The GIS geometry does not need to be given to the model at all. The team can keep authoritative points and polygons in QGIS or a spatial database and transfer only approved text into the AI workflow.

These are not minor configuration differences. They are different sovereignty architectures.

The practical rule

Before entering Māori or organisational information into an AI system, identify the exact product and account type. Then check whether the material is used only for inference, what is retained, whether the user can disable model improvement, whether enterprise no-training terms apply, where processing occurs, what tools or connectors may transmit additional material and how deletion works.

Do not ask only “Is ChatGPT safe?” or “Does Claude train on our data?” Those questions are too broad.

Ask instead:

What happens to this information in this product, under this account, with these settings, for this particular use?

That question produces decisions that can actually be implemented.

The next article widens the view: American and Chinese AI models: different companies, surprisingly similar data pipelines.

Further reading

OpenAI: Business data

OpenAI: Data Controls FAQ

Anthropic: Is my data used for model training?

Google: Gemini Apps Privacy Hub

Google Workspace: Generative AI Privacy Hub

Microsoft: Microsoft 365 Copilot privacy

DeepSeek: Privacy policy

Qwen: Terms of Service