What this means for Māori GIS: from protecting data to controlling the AI architecture
The previous articles in this series establish several uncomfortable but useful facts. Public Māori material has already entered web-scale AI training corpora. Māori-related information is likely far more extensive than the visible te reo Māori subset because large quantities of Māori history, planning, legal and environmental material are written in English. At the same time, models can remain weak in Māori language and context because the global training corpus is overwhelmingly larger than the Māori material within it.
We also know that entering a document into ChatGPT, Claude, Gemini, Copilot, DeepSeek or Qwen today is a separate issue from historic foundation-model training. Current prompts may be used only for inference, may be retained for a period, may be excluded from model training under enterprise terms, or may be eligible for later model improvement depending on the product and settings. Local AI can change the data path again by keeping current documents and inference on infrastructure controlled by the organisation.
For Māori GIS, this is not an abstract debate. Mapping is where information becomes spatially explicit. AI can make hidden relationships easier to discover, and GIS can then turn those findings into points, lines, polygons and maps that appear much more certain than the underlying evidence.
The practical challenge is therefore not to decide whether AI is good or bad for Māori mapping. It is to decide which parts of the AI architecture should be used for which kaupapa, with what information, and under whose authority.
Public training has already changed the starting point
The first implication is historical. We should assume that ordinary crawlable public Māori webpages may already have been captured into some international web-training datasets. The mC4 Māori subset provides direct evidence of this pathway, with identifiable Te Ara, Māori media, cultural and iwi material inside a major Common-Crawl-derived corpus.
That does not mean every current model used those exact records. It means the old idea that publicly available Māori information sat outside the global AI training ecosystem is no longer tenable.
This matters because some governance discussions begin as though the main decision is whether Māori information should be allowed to enter AI for the first time. For public material, that moment may have passed years ago.
But the conclusion should not be that sovereignty no longer matters. The more useful conclusion is that sovereignty now needs to distinguish between historic public exposure and new organisational information that remains controllable.
A public Te Ara article and a confidential iwi strategy document are not in the same position.
Do not surrender the controllable because the public web was scraped
There is a dangerous argument that can follow from this research: if global models already contain Māori information, perhaps there is little point protecting new material.
That would be a mistake.
Public historical material may already have influenced model weights, but private organisational information can still remain private. A field dataset can still be kept local. A sensitive spatial layer can still be excluded from cloud processing. An internal archive can still be indexed under role-based access. A new oral-history transcript can still remain outside foundation-model training. A whakapapa-linked collection can still be governed according to the authority that surrounds it.
The existence of historical extraction makes present-day architecture more important, not less.
Separate the map from the source material
GIS practitioners naturally think about the map as the visible output, but the sensitive part may sit much deeper in the workflow.
A final map may show only a general area. The project behind it may contain precise coordinates, field photographs, landowner notes, historical descriptions, cultural-impact reports, candidate wāhi tapu locations, interview material and a detailed source register.
If an AI assistant is given the complete project folder, the relevant data exposure may be far greater than anything visible on the published map.
This is why Māori GIS teams should treat AI data-flow design as part of GIS project design. Decide which source layers and documents the AI actually needs. In many cases the authoritative geometry can remain entirely inside QGIS, ArcGIS or a spatial database while AI works only with approved text or generalised attributes.
The question should not be “Can the AI read the GeoPackage?” It should be “Does the AI need the GeoPackage at all?”
AI can increase sensitivity without creating new source data
One of the most important capabilities of modern AI is aggregation.
A person might need days to search fifty environmental reports for references to an awa. A RAG system can retrieve those references in seconds. A model can group spelling variants, compare time periods and identify relationships across documents.
That is extremely useful.
It can also change the practical sensitivity of the collection. Information that was technically accessible but hard to discover becomes easy to query. Separate clues can be combined. Approximate descriptions can be ranked into likely locations. Photographs can reveal place through landscape, signage or background details.
GIS then adds another step by turning the result into geography.
This means access control should be designed around what AI makes possible, not merely around the permissions that existed before the AI was added.
Keep candidate information visibly provisional
AI is particularly dangerous when uncertain information becomes a precise GIS feature too early.
Suppose a model finds a historical name in a Tribunal report and suggests that it refers to a modern feature. If a practitioner immediately creates a point with six decimal places, the dataset now looks more certain than the evidence.
A better pattern is to retain fields such as place_as_written, source_document, source_page, candidate_match, geometry_source, review_status, confidence and reviewer_notes.
The AI finding should remain a candidate until the source and geography have been checked. If a local or customary name is not in the national Gazetteer, that does not make it wrong. It means the geographic evidence needs to be handled appropriately for that name and kaupapa.
The AI should help surface evidence, not silently settle it.
Historic training and current inference need different responses
The research suggests two distinct governance tracks.
The first concerns foundation-model provenance. What public Māori material may have entered global training? How were Indigenous languages treated? Were communities able to control reuse? Does the model reproduce colonial or institutional bias? How well is te reo Māori represented? What do we know about the developer’s training disclosures?
The second concerns current deployment. What happens to the documents we use now? Where does inference occur? Are prompts retained? Is model improvement enabled? Are embeddings stored externally? Do plugins call overseas services? Can we keep the workflow local?
The first problem may be impossible for an individual iwi to completely solve retrospectively.
The second is often highly controllable.
Governance should recognise that difference and put effort where architecture can materially improve outcomes.
Build a risk ladder rather than one rule
Not every AI task involving Māori GIS needs the same controls.
Asking a public model to explain the difference between NZTM2000 and WGS84 using public information is low exposure. Asking an enterprise model to summarise an ordinary internal project report may be acceptable under organisational controls. Uploading culturally sensitive coordinates to a consumer chatbot is a completely different proposition. Running a local model over an approved internal archive may reduce external exposure while still requiring strong internal access governance.
The consequence also matters. AI that helps draft metadata is different from AI that ranks households, influences enforcement or determines eligibility for services.
A useful risk approach therefore considers the information, the architecture, the authority and the consequence together.
Use public material to build capability
Māori organisations do not need to begin local AI experimentation with sensitive information.
There is enough public material to build substantial technical capability first. Public planning reports, published Tribunal material, open environmental datasets, public GIS metadata and ordinary government documents can be used to test local models, document extraction and RAG.
A team can learn how embeddings work. It can compare multilingual retrieval models. It can test te reo Māori and historical place names. It can deliberately ask questions that should return “not found”. It can inspect where indexes and logs are stored. It can disconnect the network and verify whether the system continues to work.
This creates technical confidence before the stakes increase.
The result is not merely better prompting. It is organisational knowledge about the architecture.
Treat local AI as an engineering option, not an ideological position
Local AI is sometimes presented as the morally superior alternative to cloud AI. That is not a useful framing.
Local AI has practical strengths. New documents can remain on hardware controlled by the organisation. Network exposure can be reduced. Storage and deletion are more inspectable. Open-weight models provide greater freedom to change providers or operate in disconnected environments.
Local AI also has costs. Hardware must be bought and maintained. Models may be weaker than frontier cloud systems. Security becomes the organisation’s responsibility. Software supply-chain risks remain. Access controls can be badly configured. A local employee can misuse data just as easily as a cloud user if governance is weak.
The purpose of local AI is not to prove ideological purity. It is to provide another architecture choice.
That choice is particularly valuable where external processing is the main unacceptable risk.
Enterprise cloud AI remains useful
The opposite mistake would be to treat all cloud AI as inappropriate.
Enterprise services can provide strong identity, access control, audit, retention and no-training commitments. For many ordinary organisational workloads those protections may be more robust than a poorly administered local server under someone’s desk.
The right question is whether the provider’s architecture and contract fit the information and kaupapa.
A Māori organisation may sensibly use Microsoft 365 Copilot for routine corporate work, an enterprise API for approved research tasks, and a local open-weight model for more sensitive document collections.
There is no technical requirement that one platform become the organisation’s entire AI strategy.
Avoid Copilot-shaped AI literacy
One of the larger capability risks is allowing whichever enterprise product has already been purchased to define the organisation’s understanding of AI.
Copilot may be useful. Gemini may be useful. ChatGPT Enterprise may be useful. Claude may be useful.
But the field also includes local inference, open-weight models, vector search, RAG, image classification, computer vision, agents, geospatial machine learning and dedicated AI infrastructure.
Māori AI capability should include people who understand those choices well enough to challenge vendors and design systems, not only people who know how to phrase prompts inside one product.
That is a sovereignty capability in its own right.
Develop model evaluation for Māori GIS
Model evaluation should move closer to the kaupapa.
A generic benchmark cannot tell an iwi GIS team whether a model reliably preserves macrons, distinguishes an iwi name from a place name, recognises historical spelling variants or knows when it has insufficient geographic evidence.
Teams can create their own small evaluation sets using authorised, non-sensitive examples. Ask the same questions across several models. Record whether the correct source was retrieved. Test false positives. Test OCR-damaged names. Test English and te reo Māori queries. Test whether the model invents coordinates when told not to.
This turns model selection from reputation into evidence.
The best model for general coding may not be the best model for local document retrieval. A smaller multilingual embedding model may improve Māori retrieval more than upgrading the language model itself.
Mapping should preserve whakapapa of evidence
GIS already has a concept that aligns strongly with good AI practice: provenance.
A derived layer should retain where it came from. Important transformations should be documented. The relationship between source, interpretation and output should remain visible.
AI makes that more important because its outputs can be fluent and plausible even when they are wrong.
For AI-assisted Māori GIS, provenance should include the original document or dataset, page or feature identifier, extraction method, model where relevant, date, prompt or task, review status, geographic source and any uncertainty.
The word whakapapa is often used in Māori data-governance discussion to describe relationships and provenance. In an AI-GIS workflow, preserving the whakapapa of a finding is not just a principle. It is technically practical and improves reproducibility.
Four layers of control
The series points toward four useful layers of AI sovereignty.
Data sovereignty concerns authority over the source information and its use.
Infrastructure sovereignty concerns who controls the computers, storage, networks and processing environment.
Model sovereignty concerns who controls the model weights, licence, updates and deployment choices, and what is known about model provenance.
Interpretive sovereignty concerns who has authority to decide what an output means and whether it should be acted upon or published.
Local AI can greatly strengthen infrastructure control. Open-weight models can improve model autonomy. Neither automatically grants authority over the source information or the cultural meaning of an answer.
The four layers should therefore be considered together.
Southland creates a new possibility
The Datagrid project in Southland adds another dimension. New Zealand may soon have vastly more domestic AI and high-performance computing capacity than it has had before.
That does not create Māori sovereignty merely because the servers are in Aotearoa. But it could create new technical options for locally hosted inference, shared GPU infrastructure and sovereign-cloud services.
The strategic question is whether Māori organisations can shape and access that infrastructure on terms that preserve meaningful control.
If Māori-controlled storage, Māori-governed compute, open-weight models and strong access controls can be combined, local AI stops being only a laptop experiment. It becomes an infrastructure proposition.
What organisations can do now
The practical path does not need to wait for national infrastructure.
Start by mapping the existing AI data flows in the organisation. Identify public consumer tools, enterprise tools, approved APIs and local experiments. Find out which settings enable model improvement and which products already provide no-training terms.
Choose a low-risk public document collection and build a small local RAG test. Verify that the model, embedding model and document index are local. Disconnect the network. Compare retrieval quality with a cloud model.
Identify one or two GIS workflows where AI can reduce repetitive effort without becoming the source of geographic truth. Metadata search, document discovery, QGIS expression assistance and source-linked place-name extraction are good examples.
Create a simple evidence standard for AI-derived GIS records. Require source, method, review status and geometry source.
Develop a small group of people who understand models, RAG, embeddings, APIs, local hardware and data flows. Give them enough time to experiment safely.
Then use the results to improve governance. The policy should describe systems the organisation actually understands, not generic imagined versions of AI.
A positive direction
The history of public-web AI training gives Māori legitimate reasons to question how knowledge was acquired and whose interests were considered. It also demonstrates that the global AI ecosystem is unlikely to wait for perfect governance before moving forward.
The practical response does not need to be withdrawal.
Māori organisations can become more technically capable. They can distinguish public information from protected collections. They can choose enterprise services deliberately. They can operate models locally. They can control document retrieval. They can build evaluation capability. They can keep geometry separate from unnecessary AI processing. They can insist on provenance. They can develop shared infrastructure.
Most importantly, they can move from asking only what AI companies are doing with Māori data to deciding what Māori organisations themselves want AI to do.
That is a more positive expression of rangatiratanga: not simply stronger barriers, but stronger ability to choose.
For a compact reference covering training, model providers, Māori source corpora, current consumer and enterprise data treatment, the Southland project and practical GIS checks, use AI training data, privacy and Māori GIS: practical reference.
This series
Start with How AI actually learns: training, inference, RAG and what happens to your data.
Then read What Māori material is already in AI training data?, What happens to your data in ChatGPT, Claude, Gemini, Copilot, DeepSeek and Qwen?, American and Chinese AI models, and Southland’s AI factory.