Skip to main content

AI training data, privacy and Māori GIS

This page is a practical reference for people who need to understand where large AI models get training data, what Māori material is demonstrably or plausibly present in global training pipelines, what happens when information is entered into current AI services, and how those questions connect to mapping and Māori Data Sovereignty.

It should be read with AI basics for Māori GIS, Artificial intelligence and Māori GIS and the Local AI guides.

AI provider policies change. The service notes below were reviewed on 4 September 2026. Check current provider documentation before using important organisational or Māori information.

Four different data questions

The most useful starting point is to keep four processes separate.

Foundation-model training is the large historical process that creates or changes model weights. Public web data, books, code, licensed material, human examples, images and other data may contribute depending on the developer.

Inference is what happens when a trained model processes a current prompt and generates a response. Entering a prompt normally causes inference; it does not automatically retrain the model in real time.

Retention and model improvement concern what the service does after the interaction. A provider may keep chats or logs and, depending on the product and settings, some material may later enter training or evaluation pipelines.

RAG or document retrieval keeps a document collection separate from the foundation model, retrieves relevant passages and supplies them as context at inference time. RAG does not normally require retraining the foundation model on the documents.

These distinctions are essential for Māori data-governance decisions because each process creates a different data pathway.

Evidence levels for training data

When discussing whether a Māori source is “in AI”, use an evidence level rather than certainty that cannot be supported.

Evidence levelWhat it meansExample
DemonstratedThe material can be seen in a known training or pre-training corpusTe Ara and other Māori-related pages visible in the mC4 Māori subset
High likelihoodThe source was ordinary crawlable public-web material and the developer discloses large public-web trainingPublic webpages on normal indexed sites
PlausibleThe source was publicly available but in a format or repository whose ingestion cannot be demonstratedPublic PDFs such as reports and plans
UnknownNo public evidence establishes whether the material was available or usedRestricted archives, private repositories and non-public collections

Do not convert “plausible” into “proven”. Current frontier developers generally do not publish complete URL-level source manifests.

mC4 Māori subset

The multilingual C4 dataset derived from Common Crawl contains a Māori-language mi subset with roughly 101,000 training records. It includes genuine Māori-language and Māori-related webpages, alongside language-classification errors and low-quality material. Identifiable content includes Te Ara, Dictionary of New Zealand Biography material, Māori media, cultural organisations and iwi-related webpages.

The number should not be described as 101,000 clean Māori documents. It is better understood as about 101,000 web records automatically classified as Māori, some of which are demonstrably genuine Māori sources.

Inspect the mC4 Māori dataset viewer.

Te Ara and DNZB

Te Ara contains a large public corpus of Māori-focused material. Its development included around 170 Māori-focused entries and 34 iwi entries, and the combined te reo Māori content with the Dictionary of New Zealand Biography grew beyond one million words. The Dictionary contains thousands of biographies, including hundreds of Māori subjects with te reo Māori versions.

Some Te Ara and DNZB pages are demonstrably present in mC4. See Te Ara’s history.

Niupepa

The historic Niupepa collection contains more than 17,000 pages from 34 Māori periodicals published between 1842 and 1932. Around 70 percent is solely Māori, with much of the remainder bilingual. Public digitisation makes this a major online te reo Māori corpus, although inclusion in any particular current frontier model is not established.

See the Niupepa digital collection.

Waitangi Tribunal material

Tribunal reports and inquiry document collections are publicly searchable and can contain extensive Māori-related information, including place names, historical relationships, whenua, environmental evidence, maps and submissions. Much of this material is English-language and therefore would not be counted inside a Māori-language corpus.

Modern model-training pipelines are increasingly capable of extracting public PDF content, so public PDF availability should be considered when assessing likely exposure. This does not prove that a named Tribunal document was used by a named model.

See Tribunal reports and reports and documents.

Māori Land Court records

Historic Māori Land Court records contain rich land, succession, probate and whakapapa information, but much of the collection is not ordinary crawlable public-web material. Many records remain in physical, microfilm or controlled archival systems. Do not assume that foundation models contain the complete minute-book corpus.

See the Archives New Zealand Māori Land Court research guide.

Māori data is larger than te reo Māori data

A language label is not a data-sovereignty label. Large quantities of Māori-related information are written in English. Waitangi Tribunal reports, settlement documents, environmental management plans, academic research, council documents, court material and historical publications may contain detailed Māori information while being classified as English by a training pipeline.

For that reason, the visible mi portion of a dataset should not be treated as the total amount of Māori-related material available to a model.

Major model training disclosures

Developer/model familyPublicly disclosed broad training sourcesUseful scale indicator
OpenAIPublicly available internet information, partner/third-party material, user/trainer/researcher generated material, synthetic dataGPT-3 published hundreds of billions of training tokens; current frontier source inventories are not public
Anthropic ClaudePublic internet information, non-public third-party data, contractor/labelling data, internally generated data and some opted-in user dataCurrent models do not publish complete source manifests
Google GeminiWeb documents, books, code, images, audio, video and other multimodal sourcesCurrent model-family training is multimodal and large-scale
Meta Llama 3More than 15 trillion tokens from publicly available sourcesMore than 5% described as high-quality non-English data across 30+ languages
Alibaba Qwen3Web and other sources including PDF-like documents; multilingual and synthetic processingAbout 36 trillion tokens across 119 languages and dialects
DeepSeek-V3Public internet and licensed third-party data with extensive post-training14.8 trillion pre-training tokens

The table shows why model nationality alone is not a training-data explanation. Major US and Chinese developers all use industrial-scale mixtures of public, curated, licensed and synthetic information.

What happens to current prompts

The following is a practical summary, not a substitute for the provider’s current terms.

Product typeCurrent model-improvement positionImportant additional point
Personal ChatGPTPersonal users can disable model improvement; Temporary Chat is not used for trainingProcessing and limited safety retention still occur; feedback can create separate improvement pathways
ChatGPT Business/Enterprise/APICustomer business data is not used for training by defaultAPI and service retention still depend on product and contract; stronger retention controls may be available
Claude Free/Pro/MaxConsumer model-improvement use depends on settings and specified safety/feedback pathwaysMaterial entering improvement pipelines may have different retention from ordinary chats
Claude for Work/APICommercial inputs and outputs are not used for training by defaultStandard retention and eligible zero-data-retention arrangements should be checked
Gemini consumerActivity settings affect whether future chats are used for model improvementShort service/safety retention can remain even where ordinary training use is disabled
Gemini in WorkspaceWorkspace customer content is not used to train models outside the domain without permissionEnterprise processing, admin and contractual controls still apply
Microsoft 365 CopilotPrompts, responses and Graph data are not used to train foundation modelsInteractions can be retained for enterprise audit/eDiscovery; web grounding can create Bing queries
DeepSeek consumerTerms provide model-improvement and opt-out pathwaysChinese jurisdiction and current privacy/retention terms should be explicitly assessed
Qwen consumerConsumer terms allow broad processing and some AI-improvement use with opt-out mechanismsDo not confuse the cloud consumer product with locally run open-weight Qwen
Alibaba Cloud Model StudioCustomer business data is not used for model improvement without explicit consentEnterprise contractual and hosting terms still need review
Local open-weight modelThe model developer need not receive prompts if the entire workflow is localApplication telemetry, web search, plugins, remote embeddings and other tools can still create network paths

OpenAI: Business data, Data Controls FAQ, Temporary Chat

Anthropic: consumer training privacy, commercial training privacy

Google: Gemini Apps Privacy Hub, Workspace generative AI privacy

Microsoft: Microsoft 365 Copilot privacy

DeepSeek: privacy policy, terms

Qwen: terms, training-data summary, Alibaba Cloud training-data disclosure

The Southland AI factory

Datagrid NZ has resource consent for a hyperscale data centre at Makarewa, Southland. Southland District Council describes six data halls and up to 240 MW of IT capacity for AI training, data processing and storage. Datagrid markets the wider campus as 280 MW of hyperscale capacity and says it is purpose-built for AI training, inference and high-performance computing.

The project includes major electricity and telecommunications infrastructure. Datagrid and Mercury announced a 140 MW long-term power purchase option agreement in March 2026. The Tasman Ring Network is intended to provide high-capacity international connectivity through the South Island.

The central interpretation point is:

A New Zealand AI training data centre tells us where computation occurs. It does not tell us where the training data came from or who has authority over those data.

A globally sourced model could be trained in Southland. Conversely, the same infrastructure could potentially support New Zealand-based private inference, sovereign-cloud services or locally controlled open-weight models if suitable technical and governance arrangements are established.

See Southland District Council’s Datagrid resource consent page and Datagrid’s project site.

GIS-specific data-flow questions

Before AI receives a GIS project, identify what it genuinely needs. A project folder can contain much more sensitive information than the final map.

Ask where the authoritative geometry lives, whether the AI needs coordinates, whether imagery is being uploaded, whether field photographs contain location clues, whether candidate locations are being inferred from text, where GIS attributes are processed, and what happens to derived layers.

In many workflows the safest and most reproducible design is to keep authoritative geometry inside QGIS, ArcGIS, a GeoPackage or a spatial database while AI works only with approved text or selected attributes.

When AI creates candidate spatial information, retain provenance fields such as source document, page, source wording, model or method, review status, geometry source, uncertainty and reviewer notes.

A practical architecture ladder

Public cloud AI with public information can be appropriate for ordinary learning, public research and generic GIS assistance. Check the provider’s current consumer settings.

Enterprise cloud AI can be appropriate for organisational work where contracts, security, retention and Māori data-governance requirements have been assessed. No-training terms do not remove every sovereignty question.

Local documents with a cloud model can reduce the amount of material sent externally, but retrieved passages still leave the organisation if the model is remote.

Local RAG with a local model can keep source files, parsed text, embeddings, vector indexes and model inference on locally controlled infrastructure. Internal authority, access and cultural appropriateness still matter.

Shared Māori-controlled infrastructure could extend local AI from individual workstations into organisation-scale compute, particularly as domestic GPU and sovereign-cloud capacity grows.

Five checks before using a model

  1. What information will the system receive, including files, coordinates, imagery and derived information?
  2. Where does each component run: model, embeddings, document store, vector index, plugins, web search and spatial database?
  3. What is retained, for how long, and can any material enter future training or evaluation?
  4. Who has authority over the information and the proposed use, and what new information could AI make easier to infer?
  5. How will AI output be checked before it becomes GIS data, a public map or a decision input?

For more detailed governance questions use Before you use AI with Māori information.

Six-article research series

  1. How AI actually learns: training, inference, RAG and what happens to your data
  2. What Māori material is already in AI training data?
  3. What happens to your data in ChatGPT, Claude, Gemini, Copilot, DeepSeek and Qwen?
  4. American and Chinese AI models: different companies, surprisingly similar data pipelines
  5. Southland’s AI factory: what it could mean for AI, data sovereignty and Māori
  6. What this means for Māori GIS: from protecting data to controlling the AI architecture

The central idea

Māori data sovereignty in the AI era is not one question about whether a chatbot is safe. It spans historic training-data provenance, present-day inference, service retention, local and cloud infrastructure, derived information, spatial inference, model choice and the authority to interpret and act on outputs.

The most practical approach is to make those layers visible. Follow the information through the system, preserve the provenance of important findings, and choose architecture according to the kaupapa rather than according to whichever AI product is most familiar.

Last reviewed: 4 September 2026