Skip to main content

What Māori material is already in AI training data?

· 13 min read

When people ask whether ChatGPT, Claude, Gemini, Llama, Qwen or DeepSeek has been trained on Māori material, the honest answer is more complicated than yes or no. The developers of current frontier models do not publish complete lists of every webpage, PDF, book, image or other item used during training. That means we cannot point to a current model and prove that it learned from a particular Waitangi Tribunal report or iwi website unless the developer has disclosed it.

But that does not mean we know nothing. We can inspect some of the large public web corpora that have fed the modern AI ecosystem. When we do that, Māori material is plainly there.

The most useful evidence comes from mC4, the multilingual version of Google’s C4 corpus built from Common Crawl. The publicly accessible Māori-language subset contains roughly 101,000 training records labelled mi. Those records include genuine Māori-language and Māori-related pages from Te Ara, Dictionary of New Zealand Biography material, Māori media, arts organisations and at least one iwi website. They also include classification errors and rubbish, so the number must not be misrepresented as 101,000 clean Māori documents. What it does prove is that Māori online material has already been swept into web-scale corpora used by the AI research community.

Common Crawl matters

Common Crawl is one of the most important pieces of infrastructure behind modern language models. It has collected web pages for more than a decade and now contains hundreds of billions of pages. Each month it adds billions more. Researchers and AI companies can take those crawls, filter them, remove some duplication and low-quality content, classify languages and use the result to train models.

This is important because Māori material did not need to be deliberately selected by somebody building an AI model. If a page was publicly accessible, linked on the web and crawlable, it could be collected automatically as part of a general crawl.

That creates a very different history from somebody approaching an iwi and requesting permission to create a Māori AI corpus. The public-web pipeline is broad, automated and global. A crawler does not know why a page was published, what relationships sit behind the information, whether the information was intended to be reused for machine learning, or whether making it computationally useful changes its sensitivity.

OpenAI has previously said that much of the publicly available material used for its model lineage came from sources such as Common Crawl and Wikipedia. Meta, Google, Alibaba and other developers also describe large public-web collections as major parts of their training pipelines. The exact dataset used by each current model differs, but the pathway is well established.

What mC4 shows

The multilingual C4 dataset is particularly useful because it can be inspected. Its Māori mi subset contains around 101,000 training records, plus a small validation set. The collection is not a curated Māori archive. It is what automated web collection and language classification produced.

Inside it are recognisable Te Ara pages in te reo Māori, including historical and cultural material. Dictionary of New Zealand Biography entries appear as well. There is material from Waatea, Toi Māori Aotearoa and other sources. A Ngāti Kahu website page is visible in the collection. There are religious texts, general websites and a mixture of other material.

There are also obvious errors. Some records classified as Māori are not Māori at all. Automated language detection struggles with short texts, mixed-language pages, names, copied templates and poor-quality web content. Some material is spam or machine-generated content.

That is why the right conclusion is not that AI developers possess a clean 101,000-document Māori library.

The stronger and more defensible conclusion is this:

A major Common-Crawl-derived AI corpus contains around 101,000 web records classified as Māori, and within those records are clearly identifiable Māori-language, iwi, Māori media and cultural sources.

That is direct evidence of the pathway from the Māori web into AI training data.

The collection can be inspected through the mC4 Māori dataset viewer.

Te Ara is substantial

Te Ara is one of the clearest examples of why public Māori material matters at scale. It contains a large body of Māori-focused history and cultural information, alongside the Dictionary of New Zealand Biography. Te Ara has described around 170 Māori-focused entries and 34 iwi entries, and its te reo Māori material together with translated biographies grew to more than one million words.

The Dictionary of New Zealand Biography contains thousands of biographies, including hundreds of Māori subjects with te reo Māori versions.

This is not obscure material hidden in an archival system. Much of it has been ordinary public HTML on one of New Zealand’s most prominent government history websites. Some of those pages are demonstrably present in mC4.

For AI training, public HTML is convenient. It is already machine-readable. It is linked. It can be crawled at scale. It does not require somebody to scan microfilm, perform OCR or negotiate archive access.

That convenience shapes what models learn.

Niupepa and historical text

The Niupepa collection is another important corpus. The University of Waikato’s historic Māori newspaper collection contains more than 17,000 pages from 34 periodicals published between 1842 and 1932. Around 70 percent of the collection is solely Māori, with much of the remainder bilingual.

This represents an extraordinary body of historic te reo Māori. Whether any particular current frontier model ingested the complete Niupepa corpus is not publicly established. We should not claim that it did without evidence.

But the existence of searchable digitised historical Māori text matters when thinking about what is technically available to dataset builders. Public digitisation changes the boundary between material that exists in an archive and material that can be gathered automatically into machine-learning pipelines.

The Niupepa collection can be explored through the Greenstone digital library.

Waitangi Tribunal material

The Waitangi Tribunal represents a different kind of scale. Its reports are publicly searchable and downloadable, often as large PDFs. Its inquiry document collections can include statements of claim, briefs of evidence, memoranda, submissions, appendices and other records.

The total volume is enormous and the content is often unusually rich in place-based information. Reports and evidence can include historical settlement patterns, ingoa wāhi, relationships between hapū and whenua, environmental observations, customary use, Crown actions, maps, legal descriptions and detailed locality evidence.

Much of this material is written in English.

That point is crucial because counting only te reo Māori badly underestimates Māori-related material available to AI training systems.

A language classifier may put a Waitangi Tribunal report into the English dataset even when the report contains hundreds of Māori place names, detailed whakapapa discussion and extensive material about iwi and hapū histories. The same is true of settlement deeds, environmental management plans, academic theses, council reports and many historical publications.

The Tribunal’s public reports and reports and documents search illustrate just how much material is available online.

Māori data is not Māori-language data

This may be the most important lesson from the research.

A model can learn a great deal about Māori from English-language documents. If we look only for mi language labels, we miss most of the material that describes Māori communities, history, land, law and geography in English.

An iwi management plan may be largely English while expressing iwi priorities, values and relationships with whenua. A Treaty settlement document may contain detailed descriptions of sites and historical events. An academic thesis may contain interviews or historical analysis. A council cultural-impact document may identify places and relationships. A nineteenth-century ethnographic text may contain material about tikanga and whakapapa, even if its interpretation is poor or colonial.

From an AI perspective, all of these can contribute patterns and associations during training.

So the visible Māori-language slice of a web corpus is probably only the tip of the Māori-related material within the wider English collection.

PDFs are becoming easier to ingest

There is also an important difference between older and newer model-training pipelines.

Early web corpora were strongest on ordinary HTML text. PDFs could be difficult. They might contain scanned pages, complex layouts, tables, maps or broken text extraction. That created some friction for large-scale use.

Modern model developers are increasingly solving that problem. Alibaba’s Qwen team explicitly says Qwen3’s data pipeline included PDF-like documents, with vision-language models used to extract text and additional processing to improve the extracted material. Qwen3’s disclosed training corpus was around 36 trillion tokens across 119 languages and dialects.

That does not prove that Qwen3 contains a particular Tribunal report. It does mean that “it was only available as a PDF” is no longer a strong reason to assume a modern training pipeline ignored it.

See the Qwen3 technical overview.

Māori Land Court records are different

Māori Land Court records illustrate why this research needs evidence levels.

Archives New Zealand describes minute books and related material extending back to the nineteenth century, including land information, succession, probate, adoption and whakapapa evidence. However, much of the historic corpus remains in physical, microfilm or controlled archival form rather than as an ordinary crawlable public website.

It would therefore be wrong to say that large language models have simply ingested the Māori Land Court minute books as a complete corpus.

Some indexes, excerpts, secondary sources and regional material may be online. Some academic or legal publications quote from the records. But the availability pathway is not the same as Te Ara or a public HTML iwi website.

See Archives New Zealand’s Māori Land Court research guide.

Public does not mean culturally neutral

A common response is that if a page was public, nothing improper could have happened when it entered AI training.

Legally and technically, public accessibility is highly relevant. Culturally, it does not answer every question.

Information can be published for a particular audience and purpose without its authors imagining that it will be copied into a global machine-learning corpus. A report can be public while containing sensitive or disputed material. A historic text can be public while representing a colonial interpretation that should not be treated as authoritative. A place may be named publicly while a new system that aggregates every reference to it changes practical accessibility.

AI training increases scale and changes reuse. Those changes are part of why Indigenous data-sovereignty frameworks ask questions beyond ordinary secrecy and privacy.

At the same time, it is important not to overstate what training does. Once a public page has influenced a model’s weights, the model is not necessarily carrying a retrievable copy of that page. Training is a statistical transformation. That does not remove concerns about authority or provenance, but it changes the technical problem.

Underrepresentation and extraction can coexist

There is an apparent contradiction in this subject.

Māori material may already have been harvested into global training datasets, yet models can still perform poorly with te reo Māori and Māori-specific context.

Both can be true because the global training corpus is so large. Even substantial Māori collections are small compared with the volume of English and Chinese material available online. Model quality is also affected by duplication, data quality, language balance, tokenizer behaviour, post-training and evaluation.

This creates an uncomfortable possibility: Māori material can be extracted without collective control while Māori remain underrepresented in the resulting technology.

The problem is therefore not solved merely by increasing the quantity of Māori training data. Authority, purpose, benefit, representation and who controls the resulting model still matter.

Which Māori worldview becomes machine-readable?

The deeper issue is not just quantity. It is selection.

What global models can easily acquire is what has been digitised and made accessible. For Māori history, that can heavily favour government reports, legislation, court decisions, public inquiry documents, colonial records, newspapers, academic writing and government-created datasets.

Locally held knowledge, oral history, restricted archives, whānau records and information deliberately kept away from public systems may be far less represented.

That means a model’s apparent “knowledge about Māori” may reflect what institutions wrote down and published rather than what Māori communities would consider the complete or most authoritative account.

For GIS this matters enormously. Official names, cadastral boundaries, published archaeological records and government maps are highly machine-readable. Relationships with whenua, local naming traditions, uncertainty, seasonal use and knowledge held through people may be much harder for a model to acquire.

A model can therefore sound knowledgeable while still representing a very particular documentary history.

Te Hiku Media shows another path

Te Hiku Media provides one of the strongest examples of Māori communities deliberately taking a different approach. Its Papa Reo work has developed Māori language technologies while retaining community control over important language data and establishing licensing conditions designed to protect kaitiakitanga over the data and derived technology.

That does not mean Māori material must never be used for AI. It demonstrates that AI development can be designed around authority rather than treating public availability or technical accessibility as sufficient permission.

See Papa Reo and Te Hiku Media’s work on Māori language technology.

What we can say with confidence

There are three useful evidence levels.

We can say with strong evidence that identifiable Māori material is present in some major public AI training corpora. mC4 proves that.

We can say with high likelihood that ordinary crawlable Māori-related public webpages have been available to other web-scale training pipelines, particularly where developers disclose heavy use of Common Crawl or comparable public-web data.

We should be much more careful with claims about specific PDFs, archive collections or named current models where the developer has not disclosed a source manifest.

That distinction between proven, likely and possible should be maintained throughout Māori AI discussion.

The practical lesson

The fact that public Māori material may already be present in model training is not a reason to abandon data sovereignty. It is a reason to become more precise about what sovereignty can still control.

An iwi cannot realistically remove every historical influence of a public webpage from every model already trained on web-scale data. But it can decide whether a new internal report is placed into a consumer chatbot. It can decide whether a document collection is indexed locally. It can decide which staff can search a sensitive archive. It can decide whether model outputs are allowed to become GIS features or public maps.

Historical model training and present-day data governance are related, but they are not the same problem.

The next article looks at that present-day question directly: What happens when you type into ChatGPT, Claude, Gemini, Copilot, DeepSeek or Qwen?.

Sources and further reading

Common Crawl: Common Crawl

mC4: Māori mi dataset viewer

Te Ara: A history of Te Ara

Niupepa: Historic Māori newspaper collection

Waitangi Tribunal: Tribunal reports

Archives New Zealand: Māori Land Court research guide

Qwen: Qwen3 technical overview

Papa Reo: Māori language technology under Māori control