The model's explanation had the right answer. Its verdict didn't.
A claim said that Charles Bukowski's father was born in 1895 and died in 1948. My fact-checker found a source that said "Heinrich (Henry) Bukowski (1895-1958)". The model read that source, put 1958 in its explanation, and returned supported.
So the explanation contained the correct death year, and the verdict was still wrong. The name matched and the birth year matched, but the death year did not, and the claim was marked supported anyway.
This happened again, in different forms. It became the main lesson I took from building Grounnel: an LLM's explanation and its verdict are two separate outputs, and nothing guarantees that they agree. I stopped treating the verdict as the final word and started checking the model's output against itself in code. That fixed many of these errors, but not all of them.
Then those checks started accusing true claims of being false, and I had to fix that too.
What the tool does
Grounnel is a fact-checker. You paste a text, and it:
- extracts the factual claims (up to 40 per run),
- excludes claims that are not suitable for external verification, such as opinions, predictions and personal statements,
- searches the web for sources,
- asks the model (Gemini 2.5 Flash Lite) to classify each claim against the retrieved sentences,
- runs the model's answer through a chain of deterministic checks,
- searches wider (3, then 5, then 8 sources) for claims that are still unresolved.
Each claim comes back as supported, partially_supported, contradicted, unsupported, unverifiable or excluded, with the source sentences behind it.

Where the model stops being trusted: everything below the line is code checking the model's output.
The model never writes a quote
One design choice comes first, because everything else depends on it.
Retrieved pages are split into sentences in code. Each sentence gets a number, and each source gets a letter. The model sees something like this:
{
"A": [
{ "n": 1, "text": "Heinrich (Henry) Bukowski (1895-1958) was ..." },
{ "n": 2, "text": "..." }
],
"B":
}
The model's answer can only point to sentences by letter and number:
{
"id": "c1",
"verdict": "contradicted",
"evidence_citations": [{ "source": "A", "n": 1 }],
"reason": "The source gives 1958 as the death year, not 1948.",
"confidence": 0.9
}
The text that the user sees is looked up afterwards, in code:
for (const { source, n } of citations) {
// `source` is model output, so it is treated as untrusted input
if (!Object.prototype.hasOwnProperty.call(sentencesBySource, source)) return { evidence: null, citations: [] };
const sentence = sentencesBySource[source]!.find((s) => s.n === n);
if (!sentence) return { evidence: null, citations: [] };
resolvedText.push(sentence.text);
resolvedCitations.push({ source, sentence: n, text: sentence.text });
}
If any citation points to a sentence that does not exist, all evidence for that claim is dropped. A schema rule makes sure that "no evidence" always means "no citations":
check: (c) => c.evidence !== null || c.citations.length === 0

The model points. Code retrieves. One invalid citation drops all evidence for the claim.
The model cannot invent a quote, because the output format has no place for one. This does not make the citation correct. It only means that every quote the user sees is a real sentence from a real page. The model can still point to the wrong sentence, but then the user sees that sentence next to the claim and can judge it.
This is not a new idea, and it does not solve the main problem. A model can cite the right sentence and still draw the wrong conclusion from it. That is what happened with Bukowski.
The explanation and the verdict disagree
After Bukowski, I collected more cases where the verdict was wrong but the explanation contained the information needed to reject it.
- The Wright brothers. A claim said the first flight on 17 December 1903 covered 852 feet. The first flight was 120 feet; 852 feet was the fourth and final flight. In one run, the model's explanation said "the airplane flew 852 ft on its fourth and final flight", and the verdict was still
supported. - Pluto. A claim said Pluto was reclassified in 2005. The sources said 2006. In both failed runs, the explanation named 2006. One run still returned
supported.
I don't have a theory of why this happens. What I observed is simple: the explanation and the verdict are generated as separate fields, and they can contradict each other.
First attempt: tell the model
My first fix was to add rules to the prompt. For dates, I added a section called DATE PRECISION, with the Bukowski case as its example. For the Wright brothers, I added a section called SEQUENCE POSITION.
I deployed SEQUENCE POSITION and ran the same article twice against the live system. It failed both times. In one run, the model's explanation was almost word for word the same as the example from my new prompt section, and the verdict was still wrong. I removed the section.
This does not mean prompt rules don't work. A rule about attribution strength ("led the team that built X" does not mean "personally designed X") worked on its first live test. But for this kind of failure, where the explanation already contains the right fact, adding instructions did not reliably fix the verdict.
Second attempt: check the explanation in code
So I wrote checks that read the model's explanation and compare it with its verdict.
What the year check detects. Take the Pluto case. The claim says Pluto was reclassified in 2005. The explanation says the reclassification happened in 2006. The verdict says supported. The check sees that the explanation states a different year for the same event and does not confirm the claim's year, so the verdict cannot be right.
The exact rules are narrow on purpose:
- It only runs if the claim contains exactly one year.
- It finds every year in the explanation and decides whether each one is stated or negated.
- It acts only if the claim's year is negated or missing, a different year is stated, and the sentence with that different year shares at least two key terms with the claim.
The last rule matters. Without it, an explanation like "the IAU decided this in 2006; the IAU was founded in 1919" could make 1919 look like a conflict.
What the check is allowed to do. It can flag an inconsistency. It cannot create evidence. When it fires, it marks the claim as contradicted, but with no evidence attached. The next gate in the chain does not allow contradicted without a cited sentence, so it downgrades the claim to unsupported and sends it back to the model once, with a note about the problem.
What happens next. The model verifies the claim again. If it now returns contradicted with a valid cited sentence, that verdict stands. If not, the claim stays unsupported. In the Pluto test, this is what happened: the check fired, the retry found the 2006 sentence, and the final verdict was contradicted with real evidence.

A check can only make a provisional verdict. A final contradicted needs a real cited sentence.
The whole chain
The year check is one of 12 deterministic gates that run after every model answer, always in the same order. 11 of them can change a verdict, and one only flags. Prompt rules are not counted as gates. Each gate exists because of a failure I saw in a real run.
The order matters as much as the gates. They fall into four groups:
- Check the explanation. The year check and its siblings: ordinal position ("first" vs "fourth"), negation, and general explanation/verdict consistency. One gate in this group reads the answer of a second, narrow model call that asks which event a source describes. These gates can make
contradictedprovisional; none of them can make it final. - An accusation needs evidence. A
contradictedverdict without a cited sentence is downgraded here. This group also catches an explanation that was written for a different claim in the same batch. - Compare values. Numbers and years in the claim are compared with the cited evidence. These are the only gates that can force
supported, not justcontradicted. The Bukowski claim is caught here: the cited source gives 1958 as the death year, and the claim says 1948. - Support needs evidence. This group runs last on purpose. A
supportedverdict, including one forced by the value comparison, must still carry a cited sentence.
If a gate reports an error, the claim goes back to the model once.

12 gates in four groups. The order is part of the design: gates that read only the explanation run before the evidence check, and the support check runs last.
Then my checks started accusing true claims
I added a few true negative claims to my test set, such as:
- "World War II did not end in 1943."
- "Buzz Aldrin was not the first man to walk on the Moon."
- "Microsoft did not create the iPhone."
I ran five such claims ten times each.
7 of 50 verdicts on true claims came back contradicted.
For a fact-checker, marking a true statement as false is the worst error it can make. Search was not the cause: in all seven cases, the correct evidence was retrieved. The model was not the cause either. Three of my explanation-reading checks caused all seven.
The mechanism is simple once you see it. The model correctly writes: "The war ended in 1945, not 1943." The year check sees that the claim's year (1943) is negated in the explanation and a different year (1945) is stated. That is exactly the pattern it was built to catch. But here the claim itself is negative, so the explanation agrees with the claim, and the check reads that agreement as a conflict.

Same explanation, two claims. The check cannot tell them apart.
The fix: the three checks now do nothing when the compared word in the claim sits inside a negation ("did not", "was not", and so on). They don't try to understand the negation. They just stop.
Before shipping, I replayed the change against all 109 past cases where these three checks had changed a verdict. Exactly 31 changed. In this replay set, all 31 were the known false accusations from the negative-claim tests, and the other 78 stayed the same. After deploying, I ran the five cases ten times each again: 50 of 50 were supported.
This was not the only time a check caused the error it was meant to prevent.
Apple earnings. The claim was about Apple's third fiscal quarter. The ordinal check finds the ordinal phrase ("third quarter") and compares the number attached to it with the claim's number. The explanation said "Apple reported $23.4 billion (or $23.43 billion) ... the third fiscal quarter". The check took the first number in that clause, 23.4, saw that it did not match the claim's 23.43, and turned a correct supported into contradicted. The fix was to take the number nearest to the ordinal phrase, not the first one.
A known risk that is live now. I added "last" and "final" to the ordinal check's word list. A code review then showed that "last" is usually not an ordinal at all ("last week", "last quarter"). In some cases, this can turn a true claim into contradicted, and contradictions from this check are protected from later correction. I decided to keep it and log every firing for manual review instead of guessing. I don't know yet how often it fires wrongly. A prompt change broke a code check
There is one more problem with checks that read model text: they depend on how the model words its explanation.
The Wright brothers check only works if the explanation names the flight by position ("the fourth flight"). If the model writes "the longest flight covered 852 feet", there is no ordinal word to compare, and the check does nothing.
I tried a new prompt section that described sequence positions in more detail. Tested alone, it looked good. In the full pipeline, it dropped detection of the false Wright brothers claim from 6 of 12 runs to 0 of 6. The new section did not change any verdict rule. It changed how the model phrased its explanations, and the check stopped matching. I reverted it.

The prompt did not change the check. It changed the text the check depended on.
Two lessons for me:
- A prompt change must be tested through the whole check chain, not on the model alone. A change made for an unrelated reason can silently switch a check off.
- This case is not solved. The false "first flight covered 852 feet" claim is caught in only about a quarter to a third of runs (16 of 51 in a pooled measurement; 1 of 5 in each of the last two). The true version has never been marked
contradicted, but the false one is still often markedsupported. I stopped working on it after many failed attempts, and my test now records the miss rate honestly instead of failing on it.
The general lesson: a deterministic check is not automatically safer than the model. It is predictable, testable and cheap. But it is written by a person looking at a few examples, and it can fail in ways no test covers until real text arrives.
The rule for uncertain cases
Not every gate exists to prevent false accusations. Bukowski and the Wright brothers are false confirmations. But when the system cannot be sure which way a claim goes, one rule decides: a false accusation is worse than a missed detection.
In practice:
- When unsure, the system says
unverifiable, notcontradicted. - If the model's self-reported confidence is below a threshold, the code downgrades the verdict to
unverifiable. The verdict is shown, not hidden. This confidence number is only a signal for being careful; it is not a probability that the verdict is correct. - A wider search that finds no evidence cannot overwrite an earlier verdict that had real evidence.
- A check that is not sure does nothing.
This rule costs detection. Some false claims come back as unsupported or unverifiable when a more aggressive system would mark them contradicted. I accept that cost.
What I don't know
I have not checked the contradicted verdicts from real use by hand, so I don't have an honest false-accusation rate, and the public stats page does not show one. The 7 of 50 above comes from my own test set.
These are verdict counts for all claims on the public site as of 6 October 2026 (302 runs since 7 August, including my own runs before the public launch). They are counts, not accuracy measurements, and they say nothing about how many claims in the world are true or false:
| Verdict | Claims |
|---|---|
supported | 2,886 |
unsupported | 550 |
contradicted | 277 |
excluded | 234 |
unverifiable | 220 |
partially_supported | 29 |
| no verdict | 43 |
Most of the traffic is me. In September, the system processed about 1,230 runs, and 73 of them came through the public site, on 50 different texts. The rest were my evaluation runs. Many of the 73 public runs use my own demo texts, and I cannot reliably separate my runs from visitors' runs, so I am not giving a user count.
Cost: my internal meter reported $7.67 for all 1,230 runs in September. The meter is known to read lower than Google's invoice, so I don't treat that number as exact. By the same meter, a typical run costs about 2 cents, and a full 40-claim run up to about 5 cents.
Limits
- One run is up to 40 claims, about 3,000 characters. The hosting plan (Vercel Hobby) stops any function after 300 seconds, and in this deployment a claim takes about 6.5 seconds on average, so 40 is a measured limit, not a chosen one.
- The checks cover the failures I have seen. New kinds of text will find new failures.
- The model is small and cheap. I compared it with stronger models on one narrow task: re-checking 9 real explanation/verdict pairs. Flash Lite got 8 of 9, Gemini 2.5 Flash got 7 of 9, and Gemini 3.1 Pro preview got 9 of 9. The test sent one pair at a time, while production sends batches, so it does not show that a stronger model would fix the pipeline. I have not tested a stronger model end to end.
Try it
grounnel.vercel.app. It is free, with no signup, and I build and run it alone. Paste any text. Every run gets a permanent link that anyone can open. If you find a wrong verdict, that link is the most useful thing you can send me.