How to Verify AI Political Research
To verify AI-generated political research, do four things: click every citation and confirm it lands on the official record, read what the source actually says rather than what the summary says it says, treat a refusal to answer as a point in the tool's favour, and prefer tools that publish how they are measured. None of this requires technical knowledge. It requires about ninety seconds per claim.
Public affairs runs on being right. This guide sets out what checkable AI research looks like, using Emily as the worked example, and it applies to any tool you might put between the parliamentary record and your clients.
1. What a citation should let you do
A citation is a promise: click here and you will see the evidence for yourself. For UK political research the evidence has a home, and it is the official record. A claim about a debate should link to the Hansard page for that debate. A claim about a vote should link the division at votes.parliament.uk, where you can count the Ayes and Noes yourself. A claim about a written question should link the question's UIN at questions-statements.parliament.uk, where the minister's answer sits in full.
The test is one click to the record. A link that lands on a homepage, a search results page, or a news article about the thing rather than the thing itself has not kept the promise. Neither has a bare footnote with no link. If you cannot reach the primary source from the claim in one step, the claim is unverified, whatever wrote it.
When Emily cites, the link goes to the official page for the specific item: the debate, the division, the question. That is a design decision, and it is one you can hold any tool to.
2. Why abstention beats confident error
The costs of the two failure modes are wildly asymmetric. A gap in an answer costs you a follow-up search. A confident fabrication costs you the meeting where a client quotes it, the correction email afterwards, and a piece of your reputation. A made-up division result in a briefing is a professional injury; a sentence saying "I could not find a division on that" is a finding.
So when you evaluate a research tool, test the empty cases, because they are where the failure modes separate. Ask about something obscure enough that the honest answer is nothing: an issue no MP has raised, a bill that does not exist. A trustworthy tool says the record shows nothing and stops. An untrustworthy one fills the silence with plausible inventions, and it will do the same on the day it matters.
This is why Emily's public evaluation scores an abstention above a confident error, never below it. The stated rule in the methodology is that on a task designed to have no answer, saying so is the pass; on an answerable task, an abstention is reported as an abstention, never dressed up as anything else.
3. How to read a Sources block
Most AI research tools now end an answer with a list of sources. Read it actively. Three questions get you most of the value:
- Where do the links go? Count how many resolve to the primary record (parliament.uk domains, legislation.gov.uk, gov.uk) versus commentary about it. Both have uses, and they are different kinds of evidence.
- Does every load-bearing claim have one? Match the strongest claims in the answer to the list. A specific number, a named person's position, a date: each should trace to a listed source. A claim that traces to nothing is the model speaking on its own authority.
- What was searched and found nothing? The best tools disclose their misses. Emily records every search she ran for an answer, including the ones that returned zero results, and shows them with the reply. A zero-hit search is information: it tells you the silence was checked, and where.
4. How Emily's bench works, and how to check it yourself
Claims about accuracy are cheap, so Emily publishes a measurement instead. The Westminster Bench is a set of 19 research tasks a public affairs professional actually does: looking up written questions, checking division numbers, tracking a bill's stage, profiling a committee chair. Emily runs each task and is graded mechanically. Do the citations resolve to the official record? Do the claims survive checking against what was actually retrieved? Does the answer say the record shows nothing when it does?
Results are published per task, pass or fail, with no headline percentage. A single number invites gaming and hides exactly the failures you would want to know about, so the page shows you each task and what happened. On the latest published run, dated 25 August 2026, the outcomes were 16 passes, 1 failure and 2 abstentions across the 19 tasks, and all 50 citations Emily produced resolved to official-record domains. The failure is published alongside the passes, because a benchmark that only ever reports success is an advert.
To check it yourself rather than take this article's word:
- Read the methodology on the bench page, which states the grading rules, including the abstention rule quoted above.
- Open the raw results file at /bench/latest.json. It is the same data the page renders, machine-readable, with per-task citations, costs and model identifiers.
- Click the citations in the published answers and confirm they land where they claim to. The grading is automated; your click is the audit of the grader.
The task set and grading code are versioned with the product, so changing either to flatter a result would be a visible code change.
5. The honest limits
Verification also means knowing what a tool cannot promise. Two limits are worth understanding about Emily, and versions of both apply to any research tool.
The record and the commentary are different things. The official record tells you what was said and how members voted. It does not tell you what the lobby thinks it meant. Emily separates the two by source tier: claims grounded in the record carry record citations, while press commentary is labelled as press and can be excluded entirely if you want answers from the record alone. When you read any AI research output, ask which tier each claim rests on. A sentence about parliamentary arithmetic and a sentence about political mood need different kinds of evidence.
Quiet periods are real, and so is lag. Parliament rises for recess and the record largely stops moving; answers to written questions can publish days after they are given. An honest tool tells you when the quiet is real rather than inventing activity to fill a briefing, and it dates its claims against how complete its copy of the record is, never against the calendar. If a tool always has something confident to say about yesterday, including in mid-August, be sceptical.
6. A checklist for any tool
Five checks, none requiring more than a browser:
- Take one answer and click every citation. Count how many land on the official record page for the specific claim.
- Ask about something that does not exist. See whether you get a fabrication or an honest nothing.
- Ask the same factual question twice in a week. The record does not change; the answers should agree.
- Ask what was searched. A tool that cannot show its work is asking you to extend trust it has not earned.
- Ask the vendor how the product is measured, and whether you can see the measurement. If the answer is a percentage with no published method behind it, treat it as marketing.
The habit this builds is the point. AI research tools are becoming standard in public affairs, and the professionals who thrive with them will be the ones who verify cheaply and often, the same way a good researcher has always treated a second-hand quote.
See the measurement before the marketing.
The Westminster Bench publishes Emily's results task by task, failures included, with the raw data one click away.
Read the benchFrequently asked questions
How do I check whether an AI research citation is real?
Click it. A real citation to the UK parliamentary record lands on the official page for the specific thing claimed: a Hansard debate, a division at votes.parliament.uk, or a written question with its UIN at questions-statements.parliament.uk. A citation that lands on a homepage, a search page, or a page that does not contain the claim has verified nothing.
Why is it good when an AI tool says it could not find something?
Because the alternative is a confident guess. In political research a fabricated division number or an invented quote costs far more than a gap: it can reach a client, a board, or a minister before anyone checks it. A tool that reports "I searched and found nothing" is showing you its limits, which is exactly what makes its positive answers worth trusting.
What is the Westminster Bench?
The Westminster Bench is a public, task-level evaluation of Emily Politics on real Westminster research tasks: looking up written questions, checking division numbers, tracking bills, profiling members. Each task is graded mechanically on whether citations resolve to the official record and whether claims survive verification. Results are published per task at emilypolitics.com/bench, with no headline percentage.
Can AI political research be fully trusted without checking?
No, and a trustworthy tool will say so. The right standard is not blind trust but cheap verification: every claim should carry a link that lets you confirm it against the official record in one click, and the tool should tell you what it searched, what it found, and what it could not find.