What Māori material is already in AI training data?
When people ask whether ChatGPT, Claude, Gemini, Llama, Qwen or DeepSeek has been trained on Māori material, the honest answer is more complicated than yes or no. The developers of current frontier models do not publish complete lists of every webpage, PDF, book, image or other item used during training. That means we cannot point to a current model and prove that it learned from a particular Waitangi Tribunal report or iwi website unless the developer has disclosed it.
But that does not mean we know nothing. We can inspect some of the large public web corpora that have fed the modern AI ecosystem. When we do that, Māori material is plainly there.
The most useful evidence comes from mC4, the multilingual version of Google’s C4 corpus built from Common Crawl. The publicly accessible Māori-language subset contains roughly 101,000 training records labelled mi. Those records include genuine Māori-language and Māori-related pages from Te Ara, Dictionary of New Zealand Biography material, Māori media, arts organisations and at least one iwi website. They also include classification errors and rubbish, so the number must not be misrepresented as 101,000 clean Māori documents. What it does prove is that Māori online material has already been swept into web-scale corpora used by the AI research community.