OPINION: Michael Swanson
I asked AI to model the 2026 General Election
Last Friday night, when I was sat on the couch watching Wheel of Time (because I am a cool guy), I was looking at some commentary around polls-of-polls and the models used to develop these, and it got me thinking…if I asked AI to set out a model for predicting the 2026 New Zealand General Election, what result would it give me? So I asked four different AI tools (Claude, Gemini, ChatGPT and Perplexity) to do this very thing, and the results were a mixture of interesting and bizarre.
I also want to say up front – this is not a deep academic dive into AI election modelling, nor is it a major critique on the use of AI in social sciences, rather it’s a bit of a “hmmm thats pretty bloody interesting” glance at how these models comprehend election modelling in New Zealand.

The four AI systems were asked to build an independent model forecasting the outcome of New Zealand’s 2026 general election. All four converge on the same mechanical understanding of MMP: a 120-seat base parliament, a 5% party-vote threshold (or one electorate win) to qualify for list seats, Sainte-Laguë proportional allocation, and the near-certainty of an overhang driven by Te Pāti Māori winning more Māori electorates than its party vote would otherwise earn. Where they diverge sharply is in the “current” polling data each claims to be working from, and that divergence, more than any modelling choice, explains why the four models reach different headline conclusions about who is likely to govern.
Methodological Similarities
Each model follows a recognisably similar three-stage structure: establish a polling baseline, apply MMP seat-allocation mechanics, then layer in qualitative judgement about factors polling cannot capture. All four explicitly flag the same handful of “known unknowns”: campaign-period gaffes and debate performance, differential turnout (especially on the Māori roll), late strategic or tactical voting near the 5% threshold, and the possibility that undecided voters break unevenly rather than mirroring decided voters.
All four also single out the same variable as the single biggest swing factor in the entire election: whether The Opportunity Party (TOP) crosses 5%. Because a party on 4.9% wins zero seats while one on 5.0% wins roughly five or six, this one threshold is treated by every model as capable of reshaping the governing arithmetic more than any other single input, a genuinely well-founded observation about how MMP concentrates risk at threshold boundaries.
There is also broad agreement on the overall character of the race: every model describes the election as extremely close, with margins of one or two seats separating a workable majority from a hung parliament. None of the four predicts a comfortable win for either bloc.
Methodological Differences
The models differ most visibly in how they communicate uncertainty. ChatGPT is the only one to run an explicit simulation (a stated approximately 20,000-iteration Monte Carlo), producing probabilistic outcomes: a 54% chance of a National-led government, 42% Labour-led, and 4% hung parliament. Perplexity takes a middle path, giving seat ranges (a left bloc of 60 to 66, a right bloc of 58 to 63) rather than a single number. Gemini and Claude, on the other hand, both present single-point seat tables with no stated confidence interval, despite each including prose caveats about how sensitive the outcome is to small shifts.
They also differ in how transparently they source their inputs. Claude cites and directly quotes a specific, dated Taxpayers’ Union-Curia poll (1 to 5 July 2026), correctly reporting that Labour sat at 31.5%, National at 30.5%, NZ First at 10.8%, and the Greens at 10.4%, figures that check out against the real published poll. Perplexity references a specific 1News poll from late June and links to the Electoral Commission’s own seat-allocation methodology. ChatGPT and Gemini, however, describe their polling inputs only in generic terms (”current evidence suggests”, “prominent mid-2026 public polls including Curia, Verian, Reid-Research, and Roy Morgan”) without citing a specific poll, date, or figure that can be independently checked.
Comparing the Results
The headline vote-share estimates vary more than ordinary poll-to-poll noise would predict. ChatGPT projects National on 31.0% and Labour on 29.5%, a National lead. Gemini projects Labour on 33.0% against National’s 30.0%. Claude projects an even larger Labour lead, 34.0% to 29.2%. Perplexity sits in between, at Labour 33% to National 31%. Checking these against the actual, verifiable July 2026 Curia poll (National 30.5%, Labour 31.5%) shows that ChatGPT’s relative positioning is closest to the real polling snapshot, while Gemini and Claude both significantly overstate Labour’s lead relative to any single published poll available at the time. This matters because it means at least some of the four systems are not working from a consistent, real polling record but are instead blending old and current polls, or generating plausible but unverified numbers, without flagging that they have done so.
This input disagreement flows straight through to the seat outcomes. Under ChatGPT’s numbers, the National-ACT-NZ First bloc falls short of a majority in its own seat table, yet its narrative conclusion, a 54% probability of a National-led government, sits somewhat awkwardly against that same table. Claude’s model is the only one of the four whose point estimate produces an actual bare majority for the incumbent coalition (62 of 123 seats). Gemini’s and Perplexity’s models both land on results where neither bloc clears the majority threshold on their central estimates, pointing instead toward a genuinely hung or fragmented parliament.
Issues With the Results
Several problems recur across the four models. First, there is a mismatch between the false precision of the outputs and the genuine uncertainty in the inputs. All four report party vote shares to one decimal place and seat totals as fixed numbers, despite every real poll carrying a margin of error of roughly plus or minus 3 percentage points per party, enough on its own to shift several seats in either direction. ChatGPT and Perplexity build that uncertainty into their final output structurally (via probabilities or ranges); Gemini and Claude present tables that look more confident than their own underlying assumptions justify.
Second, there are internal maths inconsistencies worth flagging. Gemini’s explanation of its Sainte-Laguë calculation renormalises vote share by the “effective” non-wasted vote rather than applying divisors to raw vote shares directly, and its rounding of 41.46 seats “up” to 42 does not follow ordinary rounding conventions; the method described does not fully match how Sainte-Laguë actually operates. Perplexity contains a plainer inconsistency: its own seat table implies a right-of-centre bloc total of 56 seats, yet its explanation separately claims a range of 58 to 63 for the same bloc, without reconciling the two figures anywhere in the text.
Third, the models make different, largely unstated assumptions about Te Pāti Māori’s electorate sweep. ChatGPT and Claude assume the party holds essentially all seven Māori electorates; Gemini assumes five of seven; Perplexity’s table shows six electorate wins alongside a party vote of just 2%, an aggressive assumption about roll-only voting patterns that Perplexity does not really justify or explain. All seem to struggle with Māori electorates and any sort of accurate predictions.
Finally, all four models treat the undecided or refused share of the electorate, typically 15 to 20% in real NZ polls, fairly loosely, either ignoring it or making a brief, unexamined assumption about how it will break. Given how tight every model’s headline result is, this unresolved bloc of voters is arguably a bigger source of uncertainty than the horse-race framing in any of the four models acknowledges.
Bottom Line
All four models agree on the mechanics of MMP and on the diagnosis that this will be an extraordinarily close election decided by small parties and threshold effects. Where they disagree is on the input data itself, and that disagreement, more than any difference in modelling sophistication, is what drives their differing conclusions about who is likely to govern.
Of the four, Claude is the most transparently sourced against a real, dated, checkable poll, though its resulting seat table still presents more certainty than its own caveats support. The more useful lesson from comparing these four outputs may not be which model’s seat maths is most elegant, but how easily an AI-generated forecast can look authoritative while quietly resting on unverified or inconsistent input assumptions.
Let’s be clear, this is not me trying to say “AI sucks” nor am I saying “just get AI to do it”. But, in a moment of downtime, I turned a “I wonder what would happen if” into and example of how AI might be able to support some of the work we do in projecting election results while highlighting its major pitfalls. Overall, I think the important point is that elections can be projected and predicted all we want, but the real thing can throw up all sorts of left-field results that models struggle to deal with, and AI tools still struggle to articulate some of its decision-making in how it deals with these issues.
If there is interest, in slower time I’ll publish what each model did, and its reasoning for “why” so people can see just how each AI tool went through the process.










1 Comment
Definitely more interesting that the “Wheel of Time.”
Glad you examined how the AI engines work (at a very high level) and indicated why they cannot be trusted. The simulation and modelling should be better in my view – but it isn’t.
I tend of think it should be called “Artificial Organisation” as “intelligence” is a bit of a stretch.