AI Ethics

Could AI Replace Human Peer Reviewers? Rethinking Scientific Peer Review in the Age of AI

QED Science's AI ranked 57,455 bioRxiv preprints and named a 'top 1%'. Blinded experts sided with the machine over journal prestige in 75% of cases. We examine the evidence, the backlash, and what AI triage means for pandemic-era science.

Every working scientist knows the quiet triage ritual. A new preprint appears on bioRxiv; you glance at the authors, the institution, perhaps the figures, and decide in seconds whether it deserves your attention. Yesterday, Nature reported that a machine has performed that ritual across a full year of life-science output, raising an uncomfortable question: what if our informal proxies for quality are worse than the algorithm's?[1]

The company behind the experiment, QED Science of Tel Aviv, used a multi-agent AI system to score 57,455 preprints posted on bioRxiv between May 2025 and April 2026 (a near-complete census of the server's output) and named the 574 papers it rates as the top one per cent for originality and validity.[1][2] Some researchers see a long-overdue correction to a prestige-driven publishing system; others see the birth of yet another metric to be gamed at scale.

For this blog's readers, the story matters beyond academic sociology. Preprint servers are the infrastructure through which virology moves at pandemic speed. Whatever one thinks of the QED Score, its question (how do we filter trustworthy signal from an unreviewed flood?) is exactly the one the next outbreak will ask of us.

Executive Summary#

QED Science, an AI company co-founded by Tel Aviv University molecular biologist Oded Rechavi and led by chief executive Niv Mastboim, has applied its "QED Score" (an AI-generated metric of manuscript originality and validity) to 57,455 bioRxiv preprints, releasing a ranked "top 1%" in June and defending the approach in a Nature News Q&A on 10 August 2026.[1][2] The system anonymises each manuscript, decomposes it into claims and supporting experiments, and scores it independently of author identity or prestige.[3] In the firm's validation studies, blinded experts preferred the AI-favoured paper over the journal-rank-favoured paper in 75 per cent of decisive comparisons.[1][3] Critics warn of single-metric reductionism, opacity and the risk of reinforcing the prestige culture the tool claims to dismantle.[1][4]

What Happened?#

The independently verifiable facts are these. QED Science offers a free AI review tool that assesses whether a manuscript's claims are supported by its data and identifies gaps in the work. Nature reports that it has been used by more than 10,000 laboratories across 1,500 institutions in over 70 countries.[1] In November 2025, openRxiv (the non-profit that operates bioRxiv and medRxiv) began piloting a pipeline letting authors send submissions directly to QED, receiving an automated review report, usually within about 30 minutes.[5]

Then, in June 2026, the company went considerably further, anonymously scoring all 57,455 bioRxiv preprints from a twelve-month window and naming "The 1%": 574 papers it judged the year's most original and valid.[2] Its white paper calls this the largest blind quality assessment of preprint science to date, estimating that equivalent human peer review would have taken 860,000 to 1,030,000 hours, more than 400 researcher-years.[3]

On 10 August, Nature published an interview with Mastboim defending the exercise. The 574 preprints, he said, "were selected solely according to QED's assessment of originality and validity, independently of author identity, institution or publication venue", partly to surface "amazing papers that have been missed by the current publishing system".[1]

Background: Why Preprints Became a Quality Problem#

To understand why this experiment lands with such force, recall how thoroughly preprints reshaped pandemic science. bioRxiv launched in 2013; medRxiv followed in 2019. When COVID-19 struck, both became the default channel for urgent findings: viral genomes, epidemiological models, vaccine and therapeutic results circulated months before journals could convene reviewers. The servers' founders have written that many scientists "found it hard to imagine how SARS-CoV-2 research could have moved so fast without preprints".[6]

But speed has a price. A preprint carries no journal rank, no citation history, no reviewer imprimatur: the traditional proxies for trust are, as QED's white paper puts it, "built for a different era of science and unavailable at the preprint stage".[3] The filtering problem is vast: more than 1.5 million life-science papers are published each year, perhaps 4.7 million manuscripts enter the submission pipeline, and peer review consumes an estimated 35-40 million expert hours in the life sciences alone.[3] Meanwhile, Nature reported last year that bioRxiv moderators are fighting a rising tide of AI-tainted content,[7] and only last week it described AI agents auditing the literature, in one analysis attempting to reproduce claims from ICML papers, with sobering results.[8]

Into this breach steps the QED Score. Each manuscript is anonymised, then decomposed into a meta-claim, main claims, related claims and supporting experiments. A panel of specialised AI agents probes figure inconsistencies, statistical soundness, contradictions with the literature, alternative hypotheses and reporting standards; a verification layer re-checks flagged issues, and an aggregator produces one calibrated score along two dimensions: originality and validity.[3] Mastboim told Nature the system was trained partly on what negative results look like, because "the published literature disproportionately represents successful and positive findings".[1]

Why This Matters#

For AI, this is a credibility test for a whole category of tooling. Automated manuscript assessment has mostly been benchmarked on computer-science conference submissions, where accept/reject decisions offer a convenient ground truth. QED chose a harder task: discriminating quality within already-published, peer-reviewed work.[3] If claim-level evaluation proves robust in the wet-lab life sciences, funders and journals will adopt it quickly.

For biology and medicine, the immediate consequence is a shift in what gets noticed. The company found that while the United States produced the most top-scoring preprints, Austria had the highest inclusion rate in the top one per cent;[4] its funder analysis claimed the small "G. Harold & Leila Y. Mathers Foundation" placed 9 of 32 supported papers in the top tier (28.1 per cent) against 2.0 per cent for the NIH.[2] Whether or not one trusts the metric, such tables will be read, quoted and acted upon.

For virology and pandemic preparedness, the stakes are more specific. The next fast-moving outbreak will again produce thousands of unreviewed manuscripts within months, some carrying the earliest signals on transmissibility, immune escape or candidate countermeasures. Health agencies, journalists and vaccine developers will need defensible ways to triage that flood. A validated, claim-level quality signal could help; an unvalidated one could amplify confident nonsense beneath a veneer of objectivity.

Critical Analysis#

The strongest evidence in QED's favour comes from its third validation study. The team identified 100 "contradiction pairs," manuscripts where the QED Score and the eventual journal's rank (SCImago Journal Rank, or SJR) disagreed, and asked blinded experts to judge which paper was stronger. Across 70 confident judgements, experts preferred the QED-favoured paper in 45 cases, the SJR-favoured paper in 15, with 10 ties; restricting to decisive judgements, that is 75 per cent siding with the AI (95% CI 63-84%, p < 0.001).[1][3] An internal benchmark on 925 expert-tiered papers found the score discriminated "Limited" from stronger work with an AUC of 0.867.[3] The company also reports that 12.9 per cent of high-scoring preprints ended up in lower-ranked journals: its so-called "hidden gems".[1][3]

These are respectable numbers. They are also, in the main, the company's own, reported in a self-published white paper rather than a peer-reviewed study. That is an irony not lost on critics. Independent researchers interviewed by The Scientist were measured. Pedro Beltrao of ETH Zürich called the score "an interesting proxy" but cautioned that scientific value "has many dimensions, and I think their method is probably scoring one dimension"; no scientist, he added, wants to be judged by a single metric.[4] Bluma Lesch, a Yale geneticist who joined the blinded assessment, found the alignment with expert ratings unsurprising, noting that "we all know that there are biases and imperfections in peer review".[4] Poonam Thakur of IISER Thiruvananthapuram anticipated community pushback and observed that QED sometimes proposes experiments that are simply not feasible, a reminder that human oversight remains essential.[4]

Further questions remain unresolved. Sixty decisive judgements is a thin foundation for a metric that could reorder careers. Anonymisation guards against author-identity bias but not against topical or methodological-fashion bias in training data. The company concedes the tool "assumes that the data and results are genuine": it evaluates reasoning, not fraud.[9] And Goodhart's law looms: the moment a QED Score influences hiring or funding, authors and AI-writing tools will optimise for it. Rechavi frames the tool as a complement, not a replacement: peer review "should be done. It's great. It's just not always available, and it fails often."[4]

A realistic timeline: expect pilot integrations like openRxiv's to expand over the next one to two years, at least one high-profile dispute over a wrongly ranked paper, and independent replication (the true test) within a similar window.

Expert Perspective#

Placed against previous milestones, The 1% is best understood as the third wave of AI's encroachment on scholarly infrastructure. The first was assistive: in 2023, bioRxiv trialled AI-generated summaries of preprints.[10] The second was adversarial: AI-tainted submissions forced servers to strengthen screening,[7] while AI fact-checkers began auditing legacy literature, in one case catching errors in a 75-year-old chemistry reference database.[8] The third wave, evaluative AI that ranks science itself, is qualitatively different because it confers or withholds attention, the scarcest resource in research.

Competing approaches exist. Established scientometric signals are slow and prestige-laden; newer entrants range from author-centred review assistants to agents that rerun analyses rather than grade prose.[8] What makes QED genuinely different is its combination of claim-level decomposition, enforced anonymisation, and validation against blinded expert judgement rather than accept/reject outcomes. What makes it genuinely risky is the same thing: it is confident enough to publish a league table.

Key Takeaways#

  1. The largest blinded AI assessment of preprints to date is public. QED Science scored 57,455 bioRxiv papers and named a top 1% of 574.[1][2]
  2. Early validation is encouraging but self-conducted. Blinded experts sided with the AI over journal rank in 75% of decisive disagreements, a company white-paper result awaiting independent replication.[3]
  3. The tool targets a real bottleneck. Peer review consumes tens of millions of expert hours a year, and preprints (the front line of outbreak science) carry no quality signal at all.[3]
  4. The criticisms are structural, not incidental. Single-metric reduction, training-data bias, opacity and Goodhart's-law gaming are all live concerns raised by independent scientists.[1][4]
  5. Pandemic preparedness is a silent stakeholder. How this debate resolves will shape how the next flood of outbreak preprints is filtered, trusted and acted upon.[6]

Frequently Asked Questions#

What exactly is the QED Score? An AI-generated metric of a manuscript's originality and validity, produced by a pipeline that anonymises the text, extracts claims and supporting experiments, and grades them against the literature.[3]

Is the QED Score peer reviewed? No. The validation studies appear in a company white paper. Nature covered the work as news, which is not an endorsement; independent replication has yet to appear.[1][3]

Does QED replace peer review? Its founders explicitly say no, positioning it as a complement (a fast, blind signal for work not yet formally assessed) and say they do not sell to journals or publishers.[1][4]

What is the connection to bioRxiv and medRxiv? openRxiv, the non-profit running both servers, began piloting QED as an opt-in author service in November 2025; the ranking was computed on bioRxiv content.[5][2]

Why does this matter for virology and pandemics? Preprints were the fastest channel for COVID-19 science and will be again in future outbreaks. Any system that ranks unreviewed papers at scale could influence which early findings (on variants, transmission or countermeasures) gain attention.[6]

What are the main criticisms? That a single number cannot capture scientific value; that the validation is small and self-published; that anonymisation does not remove topical bias; and that any influential score will be gamed.[1][4]

Can AI-generated papers fool AI reviewers? Possibly. QED assumes reported data are genuine and does not investigate fraud, a known limitation as AI-tainted submissions rise.[9][7]

References#

  1. Nature News Q&A — "This AI tool claims to pick the top 1% of preprints. Should researchers trust it?" (10 August 2026). https://www.nature.com/articles/d41586-026-02276-z
  2. QED Science press release — "QED Science Launches AI Infrastructure for Scientific Validation" (24 June 2026, GlobeNewswire). https://www.globenewswire.com/news-release/2026/06/24/3316709/0/en/qed-science-launches-ai-infrastructure-for-scientific-validation.html
  3. QED Science white paper — "QED Score: A Validated AI-Based Quality Metric" (company white paper, 2026; not peer reviewed). https://www.qedscience.com/white-papers/qed-score-a-validated-ai-based-quality-metric
  4. The Scientist — "Can AI Tools Spot Great Science Before Reviewers Do?" (25 June 2026). https://www.the-scientist.com/can-ai-tools-spot-great-science-before-reviewers-do-74677
  5. openRxiv — "Enabling options for review: from training and transparency to author-centered AI tools" (6 November 2025). https://openrxiv.org/enabling-review-options/
  6. Inglis, J. et al. — "How bioRxiv and medRxiv brought preprints to the life sciences" (mBio, 2026). https://repository.cshl.edu/id/eprint/42067/1/10.1128.mbio.02989-25.pdf
  7. Nature News — "AI content is tainting preprints: how moderators are fighting back" (12 August 2025). https://www.nature.com/articles/d41586-025-02469-y
  8. Nature News — "AI agents are checking the scientific literature — and spotting decades-old errors" (6 August 2026). https://www.nature.com/articles/d41586-026-02235-8
  9. Manusights — "q.e.d Science Review 2026: What to Watch" (quoting QED's own stated limitation that it assumes data are genuine). https://manusights.com/blog/qed-science-review-2026
  10. Nature News — "AI writes summaries of preprints in bioRxiv trial" (14 November 2023). https://www.nature.com/articles/d41586-023-03545-x

Related observations

Adjacent work from the same lines of enquiry.

Ebola's 100-Day Test: The Bundibugyo Vaccine Sprint

The Bundibugyo Ebola outbreak is now the largest in DRC history. With no licensed vaccine, three candidates have reached human trials in weeks — a real-world stress test of the 100 Days Mission, AI-assisted drug discovery and outbreak modelling.