Fact-checking in AI content is three problems, not one

Fact-checking in AI content is three problems, not one

Every tool in this category sells fact-checking as a solved feature. The published research says the hardest part of it is not solved anywhere, including in specialist academic systems. Here is what the evidence actually shows, and how to tell whether a page is verifiable.

For SEO and content teams · 7 min read · Every figure below links to a primary source

A note on how to read this. It is a post arguing that fact-checking claims are overstated, so a single soft statistic in it would refute the whole thing. Every number traces to a named primary source with a date. Where a source has a commercial interest, we say so. One figure we intended to use got cut because we could not verify it, and we tell you which one at the end.

The short version

  • Verifying that a source exists is close to solved. Verifying that it supports the claim citing it is not solved by anyone.
  • Eight AI search engines got citations wrong in over 60% of 1,600 queries. The best performer was wrong 37% of the time.
  • Deep research agents produce more citations than simpler tools and fabricate them at higher rates.
  • Half the facts worth publishing cannot be retrieved at all, because they were never published.

The word "fact-checking" hides three different jobs

When a tool says it fact-checks, it usually means one of three things, and they are not equally hard.

Diagram showing three levels of fact verification difficulty

Scores are macro-F1 on published benchmarks. Existence and metadata: CiteCheck, arXiv 2605.27700 (2026). Claim support: Liu, Stammbach and Henderson, Princeton, arXiv 2606.21155 (2026).

The research community is unusually blunt about where the wall is. CiteCheck, one of the stronger scientific citation verifiers, reports 88.7 macro-F1 on a 982-citation physics benchmark and then states its own scope in plain language: it focuses on citation existence and metadata fidelity, not on whether the cited paper supports a particular claim. The good news explicitly excludes the part that matters.

Law is where this gets tested with real consequences. A June 2026 Princeton benchmark found over 1,000 court filings containing fabricated citations, with the count rising year over year. When the authors tested automated checkers on 1,300 brief excerpts, the best configuration reached 82.8% recall and just 60.5% F1, averaging almost 17 reasoning steps per excerpt. Their conclusion is the whole story: the best agent reliably detects cases that do not exist, and struggles with misquotes and content misrepresentation.

Machines are getting good at catching citations that are fake. They remain bad at catching citations that are real and misused. The second kind is exactly what fluent AI writing produces.

The tools meant to be grounded in sources are not

In March 2025, the Tow Center for Digital Journalism at Columbia ran 1,600 queries across eight generative search tools, feeding them excerpts that a plain Google search returned in the top three results.

Bar chart of citation error rates by AI search engine

Jazwinska and Chandrasekar, "AI Search Has a Citation Problem", Columbia Journalism Review, 6 March 2025. Read the study.

Two details matter more than the headline. Grok 3 sent users to fabricated or broken URLs in 154 of 200 responses. And ChatGPT used hedging language in only 15 of its 134 wrong answers. The tools were not just wrong, they were confident, and the paid tiers were more confidently wrong than the free ones.

A 2026 study checked citation URLs across ten models on two large benchmarks, over 200,000 URLs in total. Between 3% and 13% were hallucinated, meaning no record in the Wayback Machine and almost certainly never real. Between 5% and 18% did not resolve at all. The finding worth pinning up: deep research agents generated substantially more citations per query than simpler search-augmented models, and hallucinated URLs at higher rates. More sources did not mean better sources.

The gap on the underlying skill is wide.

Bar chart comparing model and human accuracy at identifying cited papers

Press et al., "CiteME: Can Language Models Accurately Cite Scientific Claims?", NeurIPS 2024. CiteAgent is the paper's own tool-augmented system. Read the study.

Two kinds of fact, and only one is retrievable

This is the distinction that reorganised how we think about the problem.

Diagram contrasting published facts with experience facts

The second column is what Google calls "Experience", the first E in E-E-A-T. Google's guidance on creating helpful content.

A retrieval system can find the sentence "the average is X". It cannot find "that average is misleading, because in practice this ranges enormously depending on the situation", because nobody ever wrote that down. It lives in the head of someone who has done the work.

We learned this the expensive way. A generated article once carried a tax figure our retrieval had pulled from a page ranking in the top ten. The client, who works with drivers on ride-hailing platforms, corrected it: the real answer was not a single number at all, it was a wide range. The competitor's blog post had flattened a distribution into a tidy figure, and our pipeline had faithfully laundered that tidy figure into a new article.

The SERP is not a source. A competitor's blog post is not evidence. It is just the last place a number got repeated.

How a statistic survives with no source at all

That incident is one case of a general pattern. A number gets cited by a blog post, which is cited by a bigger blog post, which is cited by an ultimate guide, until repetition alone makes it feel authoritative. Trace it back and you find a vendor PDF that no longer exists, or nothing.

The closed-loop version has a name. Randall Munroe called it citogenesis in xkcd #978 in 2011, and Wikipedia now maintains a list of documented incidents: an unsourced claim appears somewhere, a journalist repeats it, then the original page cites the journalist. A fact bootstrapped into respectability with no ground truth underneath.

Language models are an accelerant here. They are trained on the repeated version, they reproduce the tidy number fluently, and as the URL data shows, they will happily attach a citation that never resolved. The tell is always the same: you can find the claim in fifty places and the measurement in none.

The one unforgivable failure

Being generic is a quality problem. Being confidently, specifically wrong is a liability problem.

In Mata v. Avianca (S.D.N.Y., June 2023), a judge sanctioned two attorneys and their firm $5,000 after they filed a brief citing six cases that ChatGPT had invented outright, complete with fabricated quotes attributed to real judges. That was 2023, and per the Princeton data above, better models did not fix it. The problem scaled.

Publishing carries the same exposure. CNET issued corrections on 41 of 77 AI-assisted finance articles in January 2023, some described by its own editor-in-chief as substantial. Sports Illustrated deleted product content published under fake author names with AI-generated headshots in November 2023. In March 2026, Hachette cancelled the novel Shy Girl after AI-detection analysis flagged the text. In each case the cost landed on the publisher, not the tool.

Almost nobody actually verifies

If you assume disciplined checking is the norm, one dataset should change your mind. A 2026 study surveyed 94 researchers about how they handle AI-generated citations in their own field, with the source one click away.

Chart showing gap between stated and actual citation verification behaviour

GhostCite, arXiv 2602.06718, 2026. The same paper found citation hallucination rates across benchmarked models ranging from 14.2% to 94.9%. Read the study.

These are scientists checking claims in their own discipline. There is no reason to assume marketing teams do better.

What to do about it

The test is simple: can a stranger check this claim in under two minutes? If not, it is not shippable. Four habits follow from that.

Classify before you defend. Every claim is a published fact, an experience fact, or neither. Published facts carry a primary source. Experience facts get attributed to the person who supplied them, on the record. Anything in the third bucket gets cut.

Ban the SERP as a source of truth. Trace claims to the study, the filing, the official statistic or the practitioner. If a number appears only in other content, treat it as unproven.

Scrutinise specifics hardest. A vague sentence is a quality problem you fix later. A fabricated statistic is a liability you ship once and pay for repeatedly.

Prefer the short, bulletproof article. Nine hundred words where every claim is sourced beats 2,500 words padded with soft ones, on trust and on defensibility.

What we promise, and what we do not

Be sceptical of any tool in this category, ours included, that markets fact-checking as solved. The research says the core of it is unsolved even in specialist systems, so a vendor claiming perfect fact-checking has either overclaimed or has not measured.

So here is our version, deliberately narrower than what the category promises. Two deterministic behaviours, visible in the product:

The narrow promise

1. Invented numbers get flagged, not shipped. When a finished article contains a figure that cannot be traced back to a retrieved source or an answer the client gave us, that number is flagged for a human. The article is not silently rewritten. A person is told.

2. Competitors are never cited as sources. A competitor's content is never used or cited as a source for a client's article.

That is the list. Not "we check every fact", because nobody can. The commitment is that every claim is either traced to a source or attributed on the record, and anything that is neither gets cut, even when that makes the article shorter. We think the narrowness is the point.

Sources, and one thing we cut

Disclosure: the two commitments above describe our own product. Treat them as claims from an interested party, verifiable by using it. We publish no accuracy percentage about our own output, because we have no proprietary measurement to publish, and inventing one would refute this post.

On epistemic status: Google's February 2023 guidance on AI content and its March 2024 scaled content abuse policy are confirmed primary documents. The contentEffort and OriginalContentScore attributes discussed elsewhere in SEO commentary come from the May 2024 Content Warehouse leak, and information gain comes from a patent. Those are inference about what Google values, not confirmed ranking mechanics. The 2026 papers cited here are arXiv preprints.

One claim we removed. We intended to cite a cross-database benchmark reporting that roughly two thirds of citations show metadata disagreement across CrossRef, OpenAlex and Semantic Scholar. We could not locate a genuine primary source for it, only the abstract text echoed inside unrelated pages, so we cut the figure rather than repeat it. That is the discipline this post argues for, applied to itself, and it is exactly the citogenesis pattern described above.

  1. Jazwinska, K. and Chandrasekar, A. "AI Search Has a Citation Problem." Columbia Journalism Review, 6 March 2025. cjr.org
  2. "Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents." arXiv 2604.03173, 2026. arxiv.org/abs/2604.03173
  3. Press, O. et al. "CiteME: Can Language Models Accurately Cite Scientific Claims?" NeurIPS 2024. arxiv.org/abs/2407.12861
  4. "CiteCheck: Retrieval-Grounded Detection of LLM Citation Hallucinations in Scientific Text." arXiv 2605.27700, 2026. arxiv.org/abs/2605.27700
  5. Liu, P., Stammbach, D. and Henderson, P. "Who Checks the Citations? Benchmarking Legal Hallucination Detection." Princeton, arXiv 2606.21155, 2026. arxiv.org/abs/2606.21155
  6. "GhostCite." arXiv 2602.06718, 2026. arxiv.org/abs/2602.06718
  7. Google Search Central. "Guidance about AI-generated content", 8 February 2023. developers.google.com
  8. Google Search Central. "March 2024 core update and new spam policies." developers.google.com
  9. Google Search Central. "Creating helpful, reliable, people-first content" (E-E-A-T). developers.google.com
  10. Munroe, R. "Citogenesis", xkcd #978, 2011. xkcd.com/978
  11. Mata v. Avianca, Inc., 678 F. Supp. 3d 443 (S.D.N.Y. 2023).
  12. CNET corrections: CNN Business, 25 January 2023. cnn.com
  13. Sports Illustrated: Futurism, 27 November 2023. futurism.com

A robot wrote this
article. You read
the whole thing.

That's Plume. SEO drafts that don't read like AI, plans from $26/mo for 10 articles.

Create your account