Loading...
Loading...
Loading...
Accurate enough to look convincing and not accurate enough to use unchecked. A 2025 evaluation of 400 references from eight chatbots found 26.5% fully correct and 39.8% fabricated or incorrect (Cabezas-Clavijo & Sidorenko-Bautista, arXiv:2505.18059). Retrieval-based tools fare better than pure generation but still return the wrong paper or blend metadata from several sources.
AI language models generate surprisingly convincing academic references—but how many of them are real? In a 2025 evaluation of 400 references generated by eight AI chatbots, only 26.5% were fully correct and 39.8% were fabricated or incorrect (Cabezas-Clavijo & Sidorenko-Bautista, arXiv:2505.18059).
Not the model's brand name. A 2026 study of 38 models—8,913 generated references across 24 topics—found that two variables predict whether a model recalls a real paper: the log of its parameter count, and the log of how often the topic appears in the literature. Together they explain about 60% of the variance, rising to 74–94% within a single model family (Smith et al., arXiv:2605.18732). Every reference in that study was checked with SourceVerify using the SVRIS comparison standard, which agreed with human reviewers 94.4% of the time (Cohen's κ = 0.887).
This is why a single "hallucination rate" per model misleads. Averaging across topics overstates reliability on the long tail— precisely where most real research questions sit. The same model can be dependable on a well-covered topic and unreliable on a niche one.
That study scores each generated reference on recall quality: is the reference real, and is it on topic. Three of its findings make the pattern concrete.
And the gap does not close with size. Raising the rarest topic tested to 0.90 quality by scale alone would need a model roughly 30× larger than the biggest one evaluated—around 50 trillion parameters. For the long tail, retrieval and verification are the only practical answers.
Estimate recall quality for your own topic and model size →
The curve that predicts recall is an S-shape, not a straight line. Adding parameters raises the ceiling on topics the model has already seen often enough to resolve, but it barely lifts the floor under rare ones, where the signal from a handful of training mentions stays buried under everything else compressed into the same space (Smith et al., arXiv:2605.18732). A bigger model is a model that reaches further down the tail—not one that stops inventing.
The practical consequence is arithmetic. At the 39.8% fabricated-or-incorrect rate measured by Cabezas-Clavijo & Sidorenko-Bautista (arXiv:2505.18059), a 40-reference bibliography carries roughly 16 entries that need correction before submission.
AI citation errors fall into several categories:
Tools that use retrieval-augmented generation (RAG) have lower hallucination rates because they search actual databases rather than generating from memory. Smith et al. reach the same conclusion from the other direction: below the frequency at which a topic can be recalled reliably, no amount of extra scale substitutes for looking the reference up (arXiv:2605.18732). But retrieval fails in its own ways. A RAG system can still:
Given that even the best AI models produce unreliable citations, the only safe approach is to verify every reference. This means checking against authoritative databases like OpenAlex, and Google Scholar.
SourceVerify automates this verification using the SVRIS standard—a transparent, auditable method that shows you exactly how each citation was verified. Unlike black-box AI checkers, SVRIS provides auditable results: you can see which fields matched (title, authors, year, venue) and which sources provided evidence.
AI-generated references are unreliable across every model tested, and how unreliable depends on the model's size and how well covered your topic is—not on which brand you used. In a 2026 study of 38 models, recall of real papers was predictable from those two variables alone (Smith et al., arXiv:2605.18732). No current model can be trusted to produce accurate citations without checking. For academic and professional work, use SVRIS-based verification to catch fabricated references before they damage your credibility.