The Reference That Never Existed#
Earlier this year, a team led by Maxim Topaz at Columbia University went looking for something odd: citations that point to papers nobody ever wrote. They checked 97.1 million references across nearly 2.5 million papers indexed in PubMed and found 4,046 fabricated ones, spread over 2,810 papers (Retraction Watch, 2026). That sounds tiny. The trend is not. One paper in 2,828 carried a fabricated reference in 2023, and by early 2026 the figure was one in 277, roughly a tenfold rise in three years (The Scientist, 2026). More than 98% of the affected papers had drawn no action from their publishers when the audit was run (Retraction Watch, 2026).
The audit is published in The Lancet, and press coverage notes that the steep rise coincides with wider access to large language models (The Scientist, 2026). A fake reference does not reveal which tool, if any, wrote it, so the link is a correlation.
Meanwhile, the AI research assistant has become a normal part of the lab and the library. Search for an AI research assistant today and you will find dozens, each promising to read the literature faster than you can. The question here is which of these tools can a scientist rely on, for which jobs, and how the results should be checked.
How an AI Research Assistant Actually Works#
Most of these tools sit on top of a large language model, or LLM. An LLM is a program trained on enormous amounts of text to predict which word is likely to come next. It is brilliant at producing fluent, confident prose. It is not a database. When asked for a reference, a plain LLM does not look anything up. It writes something that looks like a reference, because that is what a reference-shaped answer usually contains.
That behaviour is what people call hallucination. The word is a bit unfortunate, since the model is not seeing things; it is completing a pattern. A psychiatry team that tested ChatGPT in 2023 argued the term is a misnomer for just this reason (McGowan et al., 2023, Psychiatry Research).
The fix most vendors reach for is called retrieval-augmented generation, or RAG. Before answering, the system searches a real collection of papers, pulls out relevant passages, and asks the model to write using only that material. Think of it as an open-book exam rather than a memory test. A further layer is the "agent": a program that plans several steps, runs searches, reads results and loops back to correct itself. Tools branded as "deep research" are usually agents.
Each layer reduces one kind of error and can introduce another. A RAG tool will not invent a paper from thin air, but it can still pick the wrong papers, miss the right ones, or misquote a real one. So there are three ways an assistant can let you down: it can invent a source, mangle a real one, or quietly leave out the papers that matter. The evidence below is sorted by which of those it exposes.
Fabricated Citations: The Problem With Plain Chatbots#
The earliest studies tested general chatbots, and the results were poor. When ChatGPT-3.5 and GPT-4 wrote short literature reviews on 42 topics, 55% of the GPT-3.5 citations and 18% of the GPT-4 citations were fabricated, and many of the real ones contained errors (Walters et al., 2023, Scientific Reports). In a test of systematic review tasks, hallucination rates reached 39.6% for GPT-3.5, 28.6% for GPT-4 and 91.4% for Google's Bard, while precision stayed below 14% for every model (Chelli et al., 2024, Journal of Medical Internet Research).
Newer models did better, but not uniformly. In a finance test of 150 citations per chatbot, ChatGPT-4o hallucinated 20.0% of the time and o1-preview 21.3%, while Gemini Advanced reached 76.7% (Erdem et al., 2024, International Journal of Data Science and Analytics). In Nature, researchers reported that GPT-4o invented citations 78 to 90% of the time on literature-synthesis questions when it had no retrieval system behind it (Asai et al., 2026, Nature).
The chart below puts these results side by side. Treat it as a rough guide only, because each team defined "hallucination" differently and tested different subjects.
Specialist Tools Help, But Not in the Way You Might Expect#
The obvious answer to fake references is to build tools that only cite what they have actually retrieved. The evidence says this works for citation integrity.
OpenScholar, a retrieval model built on 45 million open-access papers, reached citation accuracy on par with human experts, and expert reviewers preferred its long-form answers over expert-written ones in a majority of cases (Asai et al., 2026, Nature). A biomedical multi-agent system called LITERAS, which queries PubMed directly, matched 99.82% of its references to real publications (Gorenshtein et al., 2025, Computers in Biology and Medicine). In the same study, Perplexity's Sonar model drew 35.6% of its sources from non-academic content, which is a different kind of problem: real links, weak evidence (Gorenshtein et al., 2025).
The weak spot is coverage. A comparison of GPT-4-powered Consensus against a human-led systematic review on big data research found precision of 76.2% but recall of only 38.1%, and the AI-selected papers averaged 202 citations each against 64.4 for the human-selected set (Tosi, 2025, IEEE Access). In plain terms: most of what it found was relevant, but it missed most of what it should have found, and it leaned towards papers that were already famous. A systematic review that quietly omits less-cited work is a biased review.
The 2026 Reality Check: Better, Still Patchy#
Several preprints from this year (not peer reviewed) test current tools. Because they have not yet been through formal review, treat them as early signals rather than settled findings.
A physics and astrophysics study compared human experts and mid-2025 AI models on eight research projects. The overlap between the papers humans and AI chose was below 6%. Only about 3% of AI-generated references were pure fabrications, but 64% were real papers with at least one wrong field such as the title, year or DOI (Hell et al., 2026, arXiv, not peer reviewed). The same preprint reported that a single-project test of a 2026 model showed no fabrications or mismatches, which is encouraging but is one test.
A Dresden team evaluated five literature-review tools. Consensus returned 50 sources of which 14 were usable (28%), You.com returned 82 with 19 usable (23%), and ChatGPT returned 17 with none usable. Repeating the same search produced overlapping results only 11.8 to 28% of the time, and no tool offered confidence scores (Dathe et al., 2026, arXiv, not peer reviewed). The authors called the tools "useful for exploration, risky for precision".
Finally, a University of Pennsylvania group looked at web links in answers from commercial models and "deep research" agents. Between 5 and 18% of cited URLs did not resolve, and deep research agents had a higher hallucination rate (10.7%) than simpler search-augmented models (4.8%). Giving the model a link-checking tool cut broken citations to under 1% (Rao et al., 2026, arXiv, not peer reviewed). More autonomy did not mean more reliability, but a simple check worked.

When the Assistant Becomes the Scientist#
A growing category goes beyond reading papers: systems that propose hypotheses, run experiments and write up results. Sakana AI's "AI Scientist" produced a manuscript that passed the first round of peer review at a machine learning workshop with a 70% acceptance rate (Lu et al., 2026, Nature). That is a milestone, though a narrow one: a single workshop, in one corner of computer science.
An independent evaluation found that 42% of its proposed experiments failed because of coding errors, that its novelty checks labelled well-established ideas as new, and that some manuscripts contained hallucinated numerical results (Beel et al., 2025, ACM SIGIR Forum). In another benchmark, coding agents produced fabricated or invalid experimental results in about 80% of cases (Chen et al., 2025, arXiv, not peer reviewed). When frontier agents were given an unpublished paper's central question and six days, the original authors rejected both attempts (Kirgis et al., 2026, arXiv, not peer reviewed). A 2026 roadmap reviewing the field concluded that AI is strong at structured, retrieval-grounded tasks and fragile at genuinely novel ideas and scientific judgement (Kong et al., 2026, arXiv, not peer reviewed).
The pattern across these studies is consistent. An assistant earns trust in proportion to how easily its output can be checked against something real.
A Scorecard: What to Trust, and How to Check It#
The table summarises the evidence by job. The ratings are this article's reading of the studies above, not a vendor ranking.
| Task | Typical tool | What the evidence shows | Trust level | How to verify |
|---|---|---|---|---|
| Drafting references from memory | Plain chatbot | Fabrication rates from 18% to over 90% across studies (Walters et al., 2023; Chelli et al., 2024) | Low | Look up every reference in PubMed or a publisher site |
| Finding papers on a topic | Retrieval-grounded search tool | High precision, low recall, bias to highly cited work (Tosi, 2025) | Medium for exploration | Cross-check with a manual database search |
| Citation accuracy in summaries | PubMed-grounded or open retrieval models | 99.82% of references matched real papers in LITERAS (Gorenshtein et al., 2025) | Medium to high | Spot-check claims against the cited passage |
| Systematic review searching | Any current tool | Low reproducibility across repeated runs (Dathe et al., 2026, not peer reviewed) | Low | Keep a documented, human-run protocol |
| Autonomous experiments and papers | AI scientist agents | Frequent failed or fabricated results (Beel et al., 2025) | Low | Demand code, logs and expert review |
A Practical Trust Checklist#
I would not tell anyone to avoid these tools. I would tell them to give the tools the right job. Use them to widen a search, summarise a paper you have read, or suggest search terms you had not thought of. Do not use them as the only source of a reference list.
Verify every citation yourself, ideally against the publisher's page or PubMed record, and check that the cited paper actually says what the summary claims. This catches both fabricated entries and the subtler "real paper, wrong details" problem. Where a tool shows its sources, read at least a sample of them. Run the search twice and compare, since a tool that returns different results each time cannot support a reproducible method. Ask which database sits behind the answer. A tool that cannot name its corpus cannot tell you what it missed.
For work that others will rely on, such as systematic reviews, clinical guidance or grant applications, keep a human-designed search protocol and use AI to supplement it, not replace it. A hybrid approach of AI for speed and humans for rigour is the one most of the evidence points towards (Tosi, 2025). Whichever AI citation checker you use, remember it is itself an AI tool with its own error rate.
One more point deserves honesty. The studies above test specific models at specific moments, and the field moves fast. The single positive test in the physics preprint suggests the newest systems may be better than the evidence base can yet show. Waiting for certainty is not a plan. Checking is.
Frequently Asked Questions#
Which AI research assistant is the most accurate? No independent evidence ranks tools reliably across all tasks. Retrieval-grounded systems such as OpenScholar and LITERAS showed near-human citation accuracy in published tests, while plain chatbots without retrieval fabricated references often (Asai et al., 2026; Gorenshtein et al., 2025).
What does it mean when an AI "hallucinates" a citation? The model generates a reference that looks plausible but points to a paper that does not exist, or mixes details from several real papers. This happens because the model predicts likely text rather than looking up records (McGowan et al., 2023).
Are AI tools for literature review safe to use for systematic reviews? Not as the sole method. One study found high precision but 38.1% recall, and a 2026 preprint found low reproducibility across repeated searches (Tosi, 2025; Dathe et al., 2026, not peer reviewed).
How common are fake references in published papers? A Lancet audit of nearly 2.5 million PubMed-indexed papers found fabricated references in one in 2,828 papers in 2023 and one in 277 in early 2026 (The Scientist, 2026).
Can an AI citation checker catch these errors? It can help. A link-checking tool reduced non-resolving citations to under 1% in one preprint, but a link that works does not prove the paper supports the claim, so a human check is still needed (Rao et al., 2026, not peer reviewed).
Can AI scientists run research without human supervision? Not reliably yet. One system passed a workshop's first-round review, yet independent tests found failed experiments and fabricated numbers, and expert authors rejected agent-produced answers to their own research questions (Lu et al., 2026; Beel et al., 2025; Kirgis et al., 2026, not peer reviewed).
Is it acceptable to cite a paper an AI summarised for me? Only after you have read the paper or at least verified the passage. The summary can misstate a real paper, as the physics preprint found when 64% of AI-generated references contained at least one incorrect field (Hell et al., 2026, not peer reviewed).
References#
- Retraction Watch (2026). One in 277 PubMed-indexed papers in 2026 shows fabricated references, says analysis. Reporting on Topaz et al., The Lancet.
- The Scientist (2026). One in 277 biomedical papers carry fake references.
- Topaz et al. (2026). Fabricated citations: an audit across 2.5 million biomedical papers. The Lancet. Cited via the secondary reports above, as the full text could not be accessed.
- Asai et al. (2026). Synthesizing scientific literature with retrieval-augmented language models. Nature.
- Walters et al. (2023). Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports.
- Chelli et al. (2024). Hallucination rates and reference accuracy of ChatGPT and Bard for systematic reviews. Journal of Medical Internet Research.
- McGowan et al. (2023). ChatGPT and Bard exhibit spontaneous citation fabrication during psychiatry literature search. Psychiatry Research.
- Erdem et al. (2024). Hallucination in AI-generated financial literature reviews: evaluating bibliographic accuracy. International Journal of Data Science and Analytics.
- Gorenshtein et al. (2025). LITERAS: Biomedical literature review and citation retrieval agents. Computers in Biology and Medicine.
- Tosi (2025). Comparing generative AI literature reviews versus human-led systematic literature reviews: a case study on big data research. IEEE Access.
- Hell et al. (2026). AI's capability in assisting scientific research in physics, astrophysics, and cosmology I: literature review. arXiv. (not peer reviewed)
- Dathe, Hoffmann and Mangold (2026). Useful for exploration, risky for precision: evaluating AI tools in academic research. arXiv. (not peer reviewed)
- Rao, Wong and Callison-Burch (2026). Detecting and correcting reference hallucinations in commercial LLMs and deep research agents. arXiv. (not peer reviewed)
- Lu et al. (2026). Towards end-to-end automation of AI research. Nature.
- Beel et al. (2025). Evaluating Sakana's AI Scientist: bold claims, mixed results, and a promising future?. ACM SIGIR Forum.
- Chen et al. (2025). MLR-Bench: evaluating AI agents on open-ended machine learning research. arXiv. (not peer reviewed)
- Kirgis et al. (2026). Can AI agents conduct open-ended AI research? Early evidence from two case studies. arXiv. (not peer reviewed)
- Kong et al. (2026). AI for auto-research: roadmap and user guide. arXiv. (not peer reviewed)