Peer review · AI evidence review · 2026

AI in peer review: what the evidence shows

How much of peer review is already AI, whether AI reviews match human ones, how easily AI reviewers are fooled, how fast fabricated references are spreading, and what journals and funders allow. Every figure is quoted from its primary source and linked to PubMed.

Verified 20 September 2026 · 44 PubMed-indexed sources · preprints labelled
Key takeaways

Eight numbers that describe the state of play.

Each line is a complete, citable statement. The bracketed number is the source in the reference list.
  1. 011 in 277An audit of 2.5 million PubMed Central papers found 4,046 fabricated references in 2,810 papers; by the first seven weeks of 2026, one paper in 277 contained at least one fabricated reference, up from one in 2,828 in 2023. [1]
  2. 0220% / 12%About 20% of ICLR reviews and 12% of Nature Communications reviews in 2025 were classified as AI-generated by a detector trained on historical reviews (preprint). [2]
  3. 0322.5% non-complianceIn a randomized experiment at ICML 2026, the LLM-use policy assigned to reviewers had near-zero effects on decisions and scores, and 22.5% of reviewers under a prohibition still reported using an LLM (preprint). [3]
  4. 0460.0% vs 48.2%Rated by 45 domain scientists on 82 Nature-family papers, a GPT-5.2 reviewing agent scored above each paper's top-rated human reviewer (60.0% vs 48.2%), but AI reviewers overlapped with each other far more than humans did (21% vs 3%) (preprint). [4]
  5. 05κ ≤ 0.15Three frontier LLMs matched an orthopaedic journal's desk-review decisions 59–63% of the time with near-zero kappa, systematically over-rejected accepted manuscripts, and none identified plagiarism or dual submission. [5]
  6. 0684.4% flippedHidden nudges inserted into manuscripts flipped the recommendation of four commercial AI models 84.4% of the time, and warning the models about nudges barely helped (76.8%). [6]
  7. 0778 of 100Of the top 100 medical journals, 78 give guidance on AI in peer review; 59% of those prohibit it outright, 91% prohibit uploading manuscript content to AI, and 96% cite confidentiality as the reason. [7]
  8. 080.1% discloseAcross 5,114 journals and 5.2 million papers, 70% of journals had AI policies, AI-assisted writing rose regardless, and only about 0.1% of papers since 2023 disclosed AI use. [8]
Last verified 2026-09-20
Prevalence

How much AI is already in peer review?

Three independent measurements, one randomized experiment, and one disclosure count.

More than the policies suggest. The first large measurement found that between 6.5% and 16.9% of review text at four machine-learning conferences had been substantially modified by language models.[9] A later detector put the 2025 share at about a fifth of ICLR reviews and an eighth of Nature Communications reviews, with the sharpest rise in late 2024.[2] When ICML 2026 randomized reviewers to a prohibition or a permissive policy, the policy changed nothing about decisions or scores, and more than a fifth of the prohibited group reported using an LLM anyway.[3] On the author side, 70% of journals have AI policies and about 0.1% of papers disclose AI use.[8]

6.5–16.9%

Between 6.5% and 16.9% of the text submitted as peer reviews to four AI conferences could have been substantially modified by large language models. [9]

20% / 12%

About 20% of ICLR reviews and 12% of Nature Communications reviews in 2025 were classified as AI-generated by a detector trained on historical reviews (preprint). [2]

22.5% non-compliance

In a randomized experiment at ICML 2026, the LLM-use policy assigned to reviewers had near-zero effects on decisions and scores, and 22.5% of reviewers under a prohibition still reported using an LLM (preprint). [3]

0.1% disclose

Across 5,114 journals and 5.2 million papers, 70% of journals had AI policies, AI-assisted writing rose regardless, and only about 0.1% of papers since 2023 disclosed AI use. [8]

Uses

How is AI used in peer review today?

Two roles with different rules: journal-side screening, and author-side checking before submission.

A 2026 scoping review of 189 studies sorts current use into assistive roles, triage and reviewer support, and autonomous ones, review generation and prediction, and concludes that today’s systems lack the domain reasoning and ethical judgment for the autonomous kind.[10] Journal-side, the checkable jobs are moving first: publishers are building language-model tools to screen reference lists the way plagiarism software screens text[11], structured checklists let a model audit PRISMA adherence at about 79% accuracy[12], and a rule-constrained model picked the right statistical test for 99.3% of items while still anchoring on the authors’ own numbers.[13]

Author-side, the use is different: the manuscript is your own, and the question is how much of what reviewers will later say a tool can surface first. On ICLR 2026 papers an author-facing system covered 44.9% of the issues historical reviewers raised in one pass and up to 84.9% with deduplication and refill[14]; earlier, GPT-4 feedback overlapped with human reviewers about as much as two humans overlap with each other, and 57.4% of researchers who tried it found it helpful.[15]

189 studies

A scoping review of 189 studies found current AI systems lack the domain reasoning and ethical judgment for autonomous evaluation, with confidentiality breaches from third-party tools a key risk. [10]

reference screening

Publishers have begun building language-model tools to screen the references of submitted manuscripts, in the way plagiarism software screens text. [11]

79% vs 45%

Giving an LLM the structured PRISMA 2020 checklist raised adherence-checking accuracy to 78.7–79.7% from 45.2% with the manuscript alone; the best model reached 95.1% sensitivity but 49.3% specificity. [12]

99.3% method choice

Constrained by a rule-based prompt, an LLM chose the right statistical test for 99.3% of 148 items and reproduced chi-square values in 94.8% of tests, but anchored on authors' reported values. [13]

44.9% → 84.9%

An author-facing LLM system covered 44.9% of the issues historical reviewers had raised on ICLR 2026 papers with a single pass, and up to 84.9% with deduplication and refill at 5.2 times the tokens (preprint). [14]

30.85% vs 28.58%

GPT-4's feedback overlapped with human reviewers' points 30.85% of the time for Nature-family papers and 39.23% for ICLR, comparable to overlap between two human reviewers (28.58% and 35.25%). [15]

57.4%

57.4% of researchers who tried LLM-generated feedback found it helpful or very helpful, and 82.4% found it more beneficial than feedback from at least some human reviewers. [15]

Head to head

Can AI review a paper as well as a human reviewer?

Thirteen comparisons from 2024 to 2026, eight of them in clinical journals. Read the setting before the result; the designs differ.

The answer depends on what is being compared. Where expert scientists rated individual criticisms, an AI reviewing agent beat each paper’s best human reviewer and raised issues no human did, but the AI reviewers also agreed with each other far more than humans do and shared sixteen recurring weaknesses.[4] Where the test was agreement with the editor’s final decision, a cardiology journal found AI reviews non-inferior[16], a sports-medicine journal found a useful first-pass screen with a strong revision bias[17], and an orthopaedic journal found near-zero agreement, systematic over-rejection, and decision letters that scored higher than the editors’ own.[5] Where the test was the recommendation itself, models accepted almost everything: 2% rejections against 73% from human ophthalmology reviewers[19], and 95–98% acceptance of hand-surgery papers a stronger journal had rejected.[20]

AI-versus-human comparisons, 2024–2026. Figures as quoted in the ledger; preprints marked.
StudySettingFindingSource
Kim et al. 2026 (preprint)82 Nature-family papers; 2,960 criticisms rated by 45 domain scientistsGPT-5.2 agent scored 60.0% vs 48.2% for each paper's top human reviewer; AI reviewers overlapped 21% vs 3% for humans; 16 recurring AI weaknesses[4]
Eur Heart J Imaging Methods Pract 202640 cardiology manuscripts, blinded editorsConcordance with final decision 67.5% AI vs 71.9% human (P = 0.74); AI right on 60% of rejections, humans on 56%[16]
Am J Sports Med 202654 submissions, locally run 14B modelRecommended revision for 61.1%, of which 72.7% were rejected or cascaded; PPV 90.5% when it said reject or cascade[17]
Orthop Traumatol Surg Res 202632 submissions, ChatGPT, Gemini and ClaudeAccuracy 59–63%, kappa −0.05 to 0.15, over-rejection of accepted papers, no plagiarism detected; letters rated 4.46 vs 4.01 for editors[5]
Acad Radiol 202615 manuscripts × 8 models × 2 prompts × 3 runs0% reject across 720 reviews; prompt strictness shifted severity[18]
Medicine 2026300 ophthalmology manuscripts, 324 human reviewersHumans rejected 73.33%; ChatGPT and Gemini 2.00% each; no line numbers or references[19]
Hand Surg Rehab 202511 hand-surgery papers rejected by one journal, accepted by anotherChatGPT accepted 95–98%; concordance with the higher-impact journal 29–32%; its reviews scored 4.8–4.9 vs 2.8–3.2 for humans[20]
Postgrad Med J 2026398 BMJ Open reports on 119 articles vs ChatGPTHumans deeper and more diverse (interpretation, originality, applicability); AI structural, slightly better on format[21]
Adv Clin Exp Med 202613 orthopaedic RCTs, risk-of-bias checklistAI–AI agreement 91%; AI–expert 64–68%; 30–38.5% deviation on interpretive items, 100% on rule-based ones[22]
New Biotechnol 2026763 preprints, 12 grant proposalsAI more lenient than humans; grants scored 3.2–3.8 vs 2.5; AI detectors failed on reviews[23]
JMIR AI 2026200 transplantation papers, 5 open modelsBest exact quartile match 0.35; no affiliation bias; not reliable enough to replace reviewers[24]
Fichtl et al. 2026 (preprint)ICLR 2026 and Nature CommunicationsDetailed reviews with overly positive recommendations, generic criticism, uneven evidence grounding[25]
NEJM AI 20243,096 Nature-family and 1,709 ICLR papers; 308-researcher surveyGPT-4 overlap with human points 30.85% vs 28.58% between two humans; 57.4% found feedback helpful[15]
The split

What AI gets right, and what it misses.

The same pattern appears in every study that measured both.
Reliable, with a human checking
  • Rule-based checklist items: 100% agreement with the expert on risk-of-bias questions that have a rule.[22]
  • Reporting-guideline adherence, when the checklist is supplied: 78.7–79.7% accuracy against 45.2% without it.[12]
  • Statistical consistency: 99.3% correct test selection, 94.8% reproduced chi-square values.[13]
  • Coverage of the concerns reviewers will raise, if enough candidates are generated: up to 84.9%.[14]
  • Reference existence, once web retrieval is on: 98.6–100% citation integrity.[31]
Unreliable on its own
  • The accept-or-reject judgment: kappa near zero against editors, 0% rejections in 720 simulated reviews.[5][18]
  • Interpretive items such as allocation concealment and blinding: 30–38.5% deviation from the expert.[22]
  • Novelty, positioning and significance: less likely than humans to critique either.[23][21]
  • Specificity: no line numbers, no references, repeated stock phrases.[19]
  • Independence: AI reviewers agree with each other 21% of the time versus 3% for human pairs, and anchor on the authors’ own numbers.[4][13]
Manipulation

Can AI reviewers be manipulated?

Three 2026 experiments, one position paper, and the two countermeasures that worked.

Yes. Hidden nudges opposing a model’s initial verdict flipped it 84.4% of the time across four commercial models, and a prompt warning the models about nudges left the rate at 76.8%.[6] Invisible white text embedded in deliberately incoherent trial abstracts produced favourable decisions in 61% of queries against 0% in controls; ChatGPT and Gemini were fooled 93% of the time, DeepSeek 60%, and Claude never, detecting both the incoherence and the injection every time.[26] Injected text inflated older models’ quality scores by 0.82 points, while early-2026 frontier models penalised it and OCR preprocessing, which reads only what a human can see, removed the effect entirely.[27] Even without hidden text, having a language model rewrite a paper raises the scores AI reviewers give it.[28]

84.4% flipped

Hidden nudges inserted into manuscripts flipped the recommendation of four commercial AI models 84.4% of the time, and warning the models about nudges barely helped (76.8%). [6]

61% vs 0%

Invisible white text in incoherent trial abstracts produced favourable LLM decisions 61% of the time versus 0% in controls; ChatGPT and Gemini were fooled 93% of the time, Claude never. [26]

OCR removes it

Invisible text injection inflated older models' quality scores by 0.82 points, while early-2026 frontier models penalised it, and OCR preprocessing removed the effect entirely. [27]

gameable by rewriting

AI review scores can be raised simply by having a language model rewrite the paper, and AI reviewers show a 'hivemind' of excessive agreement (preprint). [28]

References

The reference problem: how fast are fabricated citations spreading?

One audit of 2.5 million papers, its preprint companion, a journal's self-audit, and the generation-side measurements.

The Lancet audit of PubMed Central found 4,046 fabricated references in 2,810 papers and a rate that rose more than twelve-fold in three years, to one paper in 277 by early 2026; 98.4% of the affected papers had seen no publisher action.[1] Preprints carried more than twice the rate, and twelve preprints reached journals with the fabrications intact.[29] On the generation side, 28.3% of references produced from memory by three 2026 models were completely fabricated[30], down from 55% for ChatGPT-3.5 in 2023[34]; switching on web retrieval took citation integrity to 98.6–100%.[31] The Lancet’s editors called a fabricated reference research misconduct without debate, and an ethics analysis sets out the three conditions under which US regulations would agree.[11][36]

1 in 277

An audit of 2.5 million PubMed Central papers found 4,046 fabricated references in 2,810 papers; by the first seven weeks of 2026, one paper in 277 contained at least one fabricated reference, up from one in 2,828 in 2023. [1]

12× in 3 years

The fabrication rate rose more than 12-fold, from about four per 10,000 papers in 2023 to 56.9 per 10,000 in early 2026, and review articles had a 57% higher rate than other paper types. [1]

98.4% no action

98.4% of the 2,810 papers found to contain fabricated references had received no publisher action at the time of the audit. [1]

2.27× in preprints

Fabricated references were more than twice as common in medRxiv preprints (37.2 per 10,000) as in peer-reviewed PubMed Central articles (16.4 per 10,000). [29]

12 carried into journals

Of 104 medRxiv preprints with fabricated references, 12 were later published in peer-reviewed journals with the fabricated references intact. [29]

28.3% fabricated

Under zero-shot, retrieval-disabled conditions, 55.0% of 300 neurocritical-care references generated by three LLMs contained an inaccuracy and 28.3% were completely fabricated; Grok-4 fabricated 50%, GPT-5.3 27% and DeepSeek-V3 8%. [30]

64.6% → 98.6%

With live web search switched on, citation integrity rose from 96.8% to 100% for GPT-5 Thinking and from 64.6% to 98.6% for Gemini 2.5 Pro on orthopaedic guideline questions. [31]

3,150 references

Across 3,150 references generated by three chatbots on rotator-cuff topics, ChatGPT had the lowest reference-hallucination score and Perplexity the highest. [32]

0 of 8,064

A nursing journal that screened all 8,064 verifiable references in its 2025–2026 articles found no fabricated citations, in contrast to the wider literature. [33]

55% / 18%

55% of the references ChatGPT-3.5 generated for literature reviews were fabricated, versus 18% for GPT-4; 43% and 24% of the real ones contained substantive errors. [34]

25.4%

A meta-analysis of 28 studies found 25.4% of quotations in medical journal articles were inaccurate: 11.9% major and 11.5% minor errors. [35]

misconduct

The Lancet's editors wrote that under the NIH definition of fabrication there is no debate that a fabricated reference represents research misconduct. [11]

three conditions

Hallucinated citations may qualify as research misconduct under US federal regulations when they are produced by a generative-AI tool, function as data, and were not checked for accuracy. [36]

reference screening

Publishers have begun building language-model tools to screen the references of submitted manuscripts, in the way plagiarism software screens text. [11]

Existence is not support

A reference can exist and still not say what the citing sentence claims. A meta-analysis of 28 studies found 25.4% of quotations in medical articles were inaccurate.[35] Checking that a reference resolves is the first test; checking that its abstract supports the claim is the second. Our reference checker does both.

Rules

What do journals and funders allow, and what do they require?

Policy surveys, the editorial bodies' own texts, one editors' consensus, and the funder ban.

For reviewers the rule is close to unanimous. Of the top 100 medical journals, 78 give guidance and 59% of those prohibit AI outright, with 91% prohibiting uploads of manuscript content and 96% citing confidentiality.[7] ICMJE requires reviewers to ask the journal’s permission before using AI and bars uploads where confidentiality cannot be assured[37]; COPE’s advice, updated after its March 2026 Forum, is that AI-generated reviewer comments should not be used[38]; NIH has banned generative AI from its grant review since 2023.[39] For authors the rule is disclosure and responsibility: ICMJE and COPE both say AI cannot be an author, that authors must say how and which tool was used, and that authors answer for every part of the manuscript.[37][44] The anaesthesia editors’ consensus adds the line that matters most for references: LLMs must not generate data, references, conclusions or manuscripts.[40]

Who says what, as captured 20 September 2026
BodyScopeRuleSource
ICMJEReviewersAsk the journal's permission first; do not upload where confidentiality cannot be assured; AI output can be incorrect, incomplete or biased[37]
ICMJEAuthorsDisclose AI-assisted technologies at submission; chatbots cannot be authors; humans are responsible and must review AI output[37]
COPE Forum advice, March 2026ReviewersAI-generated reviewer comments should not be used; no manuscript uploads; the reviewer is accountable for any AI suggestion included[38]
COPE position, 2023AuthorsAI cannot be an author; disclose how and which tool in Methods; authors fully responsible[44]
NIH, June 2023Grant reviewersGenerative AI prohibited for analysing or formulating critiques[39]
JAMA Network, 2023ReviewersAI policy extended to peer reviewers with reminders of the confidential nature of manuscripts[41]
53 anaesthesia and pain editors-in-chief, 2026AllLLMs must not generate data, references, conclusions or manuscripts, nor be used for editorial decisions or review reports[40]
Academic Medicine, 2026ReviewersAn AI tool cannot be assigned professional accountability; reviewers, like authors, must be human[42]
Top 100 medical journals, 2024 surveyReviewers78 give guidance; 59% prohibit; 91% forbid uploads; 96% cite confidentiality[7]
5,114 journals, 2026 analysisAuthors70% have AI policies; AI-assisted writing rose regardless; about 0.1% of papers disclose[8]
32 orthopaedic and biomedical journals, 2026Authors91% have a policy, 97% require disclosure; 45% address peer review, 48% AI-generated references[45]
For authors

What should an author do with all this?

The rules above, applied to the manuscript you are submitting.
  • Keep the work yours. AI cannot be an author and cannot carry responsibility; every claim, figure and reference is yours to defend.[37][44]
  • Never let a tool write references. The editors’ consensus bars it, the audit shows why, and a fabricated reference can meet the definition of misconduct.[40][1][11]
  • Use AI where it is reliable and verify where it is not. Checklists, reference existence and statistical consistency are the reliable jobs; novelty and significance are not.[22][12][23]
  • Disclose what the journal asks you to disclose. ICMJE puts writing assistance in the acknowledgments and data or figure work in the methods; COPE puts tool and use in the methods.[37][44]
  • Do not upload other people’s confidential manuscripts. That is the reviewer-side rule, and it is the one every body agrees on.[37][38][43]
Where PeerReviewAI sits in this

PeerReviewAI is built for authors, not reviewers: a pre-submission review of your own manuscript, with every reference checked against PubMed and Crossref with retraction screening, the reporting checklist audited item by item, and the target journal’s instructions checked, so the problems in the sections above are found before an editor sees them. It never issues an accept or reject verdict, and it is not offered as a tool for writing journal reviews. See the full sample review.

The benchmark

How good is the human review that AI is measured against?

Agreement with the editor and with each other, the errors reviewers miss, and what changes reviewer behaviour.

Every comparison above uses human reviewers or editors as the yardstick, so the yardstick’s own reliability matters. Across 48 studies, reviewer agreement is low, a mean kappa of 0.17[48]; at one journal reviewers disagreed with each other in the first round on three-quarters of the papers that were eventually published[49]; and BMJ reviewers found fewer than three of nine planted major errors.[50] First-round reports on oncology trials ran a median of 276 words and were constructive one time in five.[51] The full picture is on the statistics hub.

kappa 0.17

Across 48 studies and 19,443 manuscripts, agreement between reviewers is low: mean intraclass correlation 0.34 and mean Cohen's kappa 0.17. [48]

74.8% discordant

Across 16,006 manuscripts at one journal, reviewers disagreed with each other in the first round 74.8% of the time for papers later published and 85.8% for papers rejected; 6.3% of published papers had a majority recommendation to reject in the final round. [49]

2.58 of 9

BMJ reviewers found an average of 2.58 of nine major errors deliberately inserted into a test paper; training raised this only to about 3. [50]

276 words

First-round review reports on oncology trials had a median length of 276 words; 19% were constructive, 25% commented on the adequacy of conclusions, 9% on applicability and 7% on open research practices. [51]

23% → 65%

The same paper was rejected by 23% of reviewers when a Nobel laureate was the visible author, 48% when anonymized, and 65% when an unknown early-career author was shown. [52]

100M+ hours

Reviewers worldwide spent more than 100 million hours on peer review in 2020, the equivalent of more than 15,000 years. [53]

42.2% → 49.8%

Offering reviewers $250 raised the share of invitations that produced a completed review from 42.2% to 49.8% at Critical Care Medicine, cut review time by about a day, and did not change review quality. [54]

+0.5 RQI

An 18-week peer-review course for doctoral students raised the editor-judged quality of their reports by a median of 0.5 points on the Review Quality Instrument. [55]

Method

How this page was verified

Every statistic on this page is a row in our verification ledger: the claim as written here, the verbatim sentence from the source with its page number, the DOI and the PubMed ID, and the capture date. Sources were located through PubMed, Europe PMC, Crossref and arXiv, and the full texts are on file. Six sources are preprints and are labelled as such in the reference list; they are never combined with a peer-reviewed source in a single statement. Figures that could only be found in secondary roundups or news reports were left out.

Verified 20 September 2026. The evidence base is moving quickly; the capture date is stated wherever a figure is time-sensitive. Corrections: use the contact form and cite the statistic ID shown in the ledger.

Protected by Anthropic’s Zero Data Retention
Your manuscript is never stored, logged, or used for training.
FAQ

Questions, answered.

Don't see yours? Email us — we read every one.

Sources

References and sources

Every figure on this page is quoted from the primary source listed here. Verification dates are recorded per source.
  1. [1]Topaz M et al. Fabricated citations: an audit across 2·5 million biomedical papers. Lancet (London, England) 2026. doi:10.1016/s0140-6736(26)00603-3 PubMed 42107362.journal letter / editorial (full text on file; PubMed record verified) · verified 2026-09-20 · Topaz et al., Lancet correspondence; audit of PubMed Central OA subset 2023-01-01 to 2026-02-18
  2. [2]Shen S, Wang K. Detecting AI-generated content in academic peer reviews. arXiv:2602.00319 (v2, 30 January 2026). arxiv.org/abs/2602.00319preprint (arXiv; full text on file; not peer reviewed) · verified 2026-09-20
  3. [3]Kim SSY, Deng WH, Vaughan JW, et al. Use and effects of LLMs in peer review: a randomized experiment and survey at ICML 2026. arXiv:2609.19420 (16 September 2026). arxiv.org/abs/2609.19420preprint (arXiv; full text on file; not peer reviewed) · verified 2026-09-20
  4. [4]Kim S, Yoon D, Gashteovski K, et al. On the limits and opportunities of AI reviewers: reviewing the reviews of Nature-family papers with 45 expert scientists. arXiv:2605.20668 (20 May 2026). arxiv.org/abs/2605.20668preprint (arXiv; full text on file; not peer reviewed) · verified 2026-09-20
  5. [5]Lupon E et al. Can large language models provide high-quality desk review decisions in an orthopaedic surgery journal? A concordance study comparing three AI models to human editorial decisions. Orthopaedics & traumatology, surgery & research : OTSR 2026. doi:10.1016/j.otsr.2026.104801 PubMed 42498040.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20
  6. [6]Tomazini BM et al. Susceptibility of large language models to hidden nudge injection during simulated medical peer review: a quasi-experimental study. Research integrity and peer review 2026. doi:10.1186/s41073-026-00225-y PubMed 42237194.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20
  7. [7]Li ZQ et al. Use of Artificial Intelligence in Peer Review Among Top 100 Medical Journals. JAMA network open 2024. doi:10.1001/jamanetworkopen.2024.48609 PubMed 39625725.journal letter / editorial (full text on file; PubMed record verified) · verified 2026-09-20
  8. [8]He Y, Bu Y. Academic journals' AI policies fail to curb the surge in AI-assisted academic writing. Proceedings of the National Academy of Sciences of the United States of America 2026. doi:10.1073/pnas.2526734123 PubMed 41734077.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20
  9. [9]Liang W, Izzo Z, Zhang Y, et al. Monitoring AI-modified content at scale: a case study on the impact of ChatGPT on AI conference peer reviews. arXiv:2403.07183 (v4, 19 May 2026). arxiv.org/abs/2403.07183preprint (arXiv abstract captured) · verified 2026-09-12
  10. [10]Nabavi A et al. Artificial intelligence in scholarly peer review: a scoping review of applications, risks, and governance challenges. International journal of medical informatics 2026. doi:10.1016/j.ijmedinf.2026.106418 PubMed 41950628.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20
  11. [11]Bauchner H, Rivara FP. Fabricated references: a new threat to editorial integrity. Lancet (London, England) 2026. doi:10.1016/s0140-6736(26)00798-1 PubMed 42107358.journal letter / editorial (full text on file; PubMed record verified) · verified 2026-09-20 · Bauchner & Rivara comment on the audit
  12. [12]Kataoka Y et al. Large language models for automated PRISMA 2020 adherence checking. International journal of medical informatics 2026. doi:10.1016/j.ijmedinf.2026.106561 PubMed 42364467.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20
  13. [13]Dong M et al. Potential and limitations of a large language model in the statistical review of comparative categorical data: an exploratory study of structured prompt-guided approach. Research integrity and peer review 2026. doi:10.1186/s41073-026-00231-0 PubMed 42387591.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20
  14. [14]Parsa P, Rezaei A. More than mimicking reviewers: evaluating LLMs for pre-submission peer review. arXiv:2609.05788 (5 September 2026). arxiv.org/abs/2609.05788preprint (arXiv; full text on file; not peer reviewed) · verified 2026-09-20
  15. [15]Liang W, Zhang Y, Cao H, et al. Can large language models provide useful feedback on research papers? A large-scale empirical analysis. NEJM AI 2024. doi:10.1056/AIoa2400196 (arXiv:2310.01783). doi.org/10.1056/aioa2400196peer-reviewed article (Crossref record; arXiv abstract captured) · verified 2026-09-12
  16. [16]Zancanaro E et al. AI as a peer reviewer: a blinded comparative study of LLM-generated and human reviews in a cardiology journal. European heart journal. Imaging methods and practice 2026. doi:10.1093/ehjimp/qyag097 PubMed 42317395.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20
  17. [17]Patel R et al. Artificial Intelligence Cannot Replace Peer Reviewers but May Help Editors Triage: A Comparative Analysis of a Large Language Model and Human Reviewer Recommendations at the <i>American Journal of Sports Medicine</i>. The American journal of sports medicine 2026. doi:10.1177/03635465261463006 PubMed 42438216.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20
  18. [18]Erturk SM, Durmaz M. Large Language Models as Peer Reviewers: Prompt Sensitivity and Model-Dependent Reproducibility. Academic radiology 2026. doi:10.1016/j.acra.2026.04.046 PubMed 42624573.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20
  19. [19]Moshirfar M et al. Current capabilities of large language models as peer reviewers for manuscripts submitted to ophthalmology-related journals. Medicine 2026. doi:10.1097/md.0000000000048147 PubMed 41894292.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20
  20. [20]Marrella D et al. Comparing AI-generated and human peer reviews: A study on 11 articles. Hand surgery & rehabilitation 2025. doi:10.1016/j.hansur.2025.102225 PubMed 40691944.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20
  21. [21]Rajakumar HK et al. Peer review in the age of artificial intelligence: a comparative study of human and AI-generated review reports. Postgraduate medical journal 2026. doi:10.1093/postmj/qgag005 PubMed 41591890.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20
  22. [22]Lettner J et al. A new peer reviewer? Comparing AI with human performance in randomized controlled trial risk-of-bias assessment. Advances in clinical and experimental medicine : official organ Wroclaw Medical University 2026. doi:10.17219/acem/216070 PubMed 41953984.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20
  23. [23]Li W et al. Large language models for peer review in biotechnology. New biotechnology 2026. doi:10.1016/j.nbt.2026.03.007 PubMed 41903776.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20
  24. [24]Shen SM et al. Evaluation of Large Language Models for Peer Review in Transplantation Research: Algorithm Validation Study. JMIR AI 2026. doi:10.2196/84322 PubMed 41672474.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20
  25. [25]Fichtl AM, Ellinger L, Kelber J, Olík K, Groh G. AI-assisted peer review across research communities: from reviewer AI policies to LLM review quality. arXiv:2608.03581 (4 August 2026). arxiv.org/abs/2608.03581preprint (arXiv; full text on file; not peer reviewed) · verified 2026-09-20
  26. [26]De Cassai A et al. Prompt injection compromises large language model-based peer review: evidence from randomised controlled trial abstracts from anaesthesia journals. British journal of anaesthesia 2026. doi:10.1016/j.bja.2026.06.041 PubMed 42580931.journal letter / editorial (full text on file; PubMed record verified) · verified 2026-09-20 · pre-registered OSF rxzjv
  27. [27]Jarrett P et al. Frontier Language Models and Optical Character Recognition Preprocessing Against Invisible Text Injection in AI Peer Review. JAMA network open 2026. doi:10.1001/jamanetworkopen.2026.17356 PubMed 42258215.journal letter / editorial (full text on file; PubMed record verified) · verified 2026-09-20
  28. [28]Baumann J, Pei J, Koyejo S, Hovy D. Stop automating peer review without rigorous evaluation. arXiv:2605.03202 (v2, 4 May 2026). arxiv.org/abs/2605.03202preprint (arXiv; full text on file; not peer reviewed) · verified 2026-09-20
  29. [29]Topaz M et al. Fabricated references are twice as common in medRxiv preprints as in peer-reviewed articles. Journal of internal medicine 2026. doi:10.1111/joim.70139 PubMed 42444601.journal letter / editorial (full text on file; PubMed record verified) · verified 2026-09-20
  30. [30]Seifi A, Seyfi A. Hallucination Rate of Peer-Reviewed Citations Generated by Large Language Models in Neurocritical Care. Critical care explorations 2026. doi:10.1097/cce.0000000000001474 PubMed 42640622.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20
  31. [31]Şengül HB et al. Deep research capabilities in GPT-5 thinking and Gemini 2.5 Pro improve citation integrity and concordance with American Academy of Orthopaedic Surgeons anterior cruciate ligament and rotator cuff guidelines. Knee surgery, sports traumatology, arthroscopy : official journal of the ESSKA 2026. doi:10.1002/ksa.70315 PubMed 41649196.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20
  32. [32]Özbek İC, Bağcıer F. Reference Hallucination in AI-Assisted Academic Writing: A Comparative Analysis of ChatGPT, Gemini, and Perplexity in Rotator Cuff Literature. Indian journal of orthopaedics 2026. doi:10.1007/s43465-026-01807-0 PubMed 42558550.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20
  33. [33]Moons P. Fabricated references due to artificial intelligence hallucination in the European Journal of Cardiovascular Nursing: also here nurses are the most trustable profession. European journal of cardiovascular nursing 2026. doi:10.1093/eurjcn/zvag148 PubMed 42334371.journal letter / editorial (full text on file; PubMed record verified) · verified 2026-09-20
  34. [34]Walters WH, Wilder EI. Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific reports 2023. doi:10.1038/s41598-023-41032-5 PubMed 37679503.peer-reviewed article (Europe PMC / Crossref record) · verified 2026-09-12
  35. [35]Jergas H, Baethge C. Quotation accuracy in medical journal articles-a systematic review and meta-analysis. PeerJ 2015. doi:10.7717/peerj.1364 PubMed 26528420.peer-reviewed article (Europe PMC / Crossref record) · verified 2026-09-12
  36. [36]Resnik DB, Hosseini M. Hallucinated citations produced by generative artificial intelligence may constitute research misconduct when citations function as data in scholarly papers. Accountability in research 2026. doi:10.1080/08989621.2026.2645390 PubMed 41833014.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20
  37. [37]International Committee of Medical Journal Editors. Recommendations for the Conduct, Reporting, Editing, and Publication of Scholarly Work in Medical Journals: Defining the Role of Authors and Contributors; Responsibilities in the Submission and Peer-Review Process. Pages captured 20 September 2026. www.icmje.org/recommendations/browse/roles-and-responsibilities/responsibilities-in-the-submission-and-peer-peview-process.htmleditorial-body recommendations (web capture on file) · verified 2026-09-20
  38. [38]Committee on Publication Ethics. Case 26-02: Editors suspect that reviewers are using artificial intelligence. Advice from COPE Members, updated after the COPE Forum of March 2026. Captured 20 September 2026. publicationethics.org/guidance/case/editors-suspect-reviewers-are-using-artificial-intelligenceCOPE case advice (web capture on file; not formal COPE guidance) · verified 2026-09-20
  39. [39]National Institutes of Health. NOT-OD-23-149: The Use of Generative Artificial Intelligence Technologies is Prohibited for the NIH Peer Review Process. 23 June 2023. grants.nih.gov/grants/guide/notice-files/NOT-OD-23-149.htmlpolicy notice · verified 2026-09-12
  40. [40]De Cassai A et al. Responsible use of large language models in manuscript authorship, peer review, and editorial processes: a Delphi consensus among editors-in-chief of anaesthesia and pain medicine journals (RULE-AP). British journal of anaesthesia 2026. doi:10.1016/j.bja.2026.01.029 PubMed 41748337.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20
  41. [41]Flanagin A, Kendall-Taylor J, Bibbins-Domingo K. Guidance for Authors, Peer Reviewers, and Editors on Use of AI, Language Models, and Chatbots. JAMA 2023. doi:10.1001/jama.2023.12500 PubMed 37498593.journal letter / editorial (full text on file; PubMed record verified) · verified 2026-09-20
  42. [42]Fischer K et al. Artificial intelligence tools in scholarly publishing: guidance for peer reviewers. Academic medicine : journal of the Association of American Medical Colleges 2026. doi:10.1093/acamed/wvaf103 PubMed 41806860.journal letter / editorial (full text on file; PubMed record verified) · verified 2026-09-20
  43. [43]Matsubara S. Zero Risk Is Impossible: Confidentiality and Artificial Intelligence Use in Peer Review. JMA journal 2026. doi:10.31662/jmaj.2026-0060 PubMed 42577012.journal letter / editorial (full text on file; PubMed record verified) · verified 2026-09-20
  44. [44]COPE Council. COPE position: Authorship and AI tools. Last reviewed 13 February 2023. doi:10.24318/cCVRZBms. Captured 20 September 2026. doi.org/10.24318/cCVRZBmsCOPE position statement (web capture on file) · verified 2026-09-20
  45. [45]Kelly N, Neuner J, Sabharwal S. Variability in Journal and Publisher Policies on Artificial Intelligence Use in Manuscript Preparation: An Orthopaedic Perspective. JB & JS open access 2026. doi:10.2106/jbjs.oa.26.00221 PubMed 42698649.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20
  46. [46]Major J et al. Exploring the Endorsement and Implementation of Artificial Intelligence Guidelines in Leading Orthopaedic and Sports Medicine Journals: A Cross-Sectional Study. The Journal of bone and joint surgery. American volume 2026. doi:10.2106/jbjs.25.00373 PubMed 41706011.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20
  47. [47]Shen F et al. Prevalence and Characteristics of Generative Artificial Intelligence Policies in Vascular Surgery Journals: A Cross-Sectional Review. Annals of vascular surgery 2026. doi:10.1016/j.avsg.2025.09.056 PubMed 41067523.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20
  48. [48]Bornmann L, Mutz R, Daniel HD. A reliability-generalization study of journal peer reviews: a multilevel meta-analysis of inter-rater reliability and its determinants. PloS one 2010. doi:10.1371/journal.pone.0014331 PubMed 21179459.peer-reviewed article (Europe PMC / Crossref record) · verified 2026-09-12
  49. [49]De Las Cuevas C. Reviewer recommendations and editorial outcomes in peer review: a longitudinal analysis of agreement and disagreement across review rounds. Research integrity and peer review 2026. doi:10.1186/s41073-026-00200-7 PubMed 42260682.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20
  50. [50]Schroter S et al. What errors do peer reviewers detect, and does training improve their ability to detect them?. Journal of the Royal Society of Medicine 2008. doi:10.1258/jrsm.2008.080062 PubMed 18840867.peer-reviewed article (Europe PMC / Crossref record) · verified 2026-09-12
  51. [51]Logullo P et al. Peer review reports of randomized controlled trials in oncology can be short and superficial. Journal of clinical epidemiology 2025. doi:10.1016/j.jclinepi.2025.111893 PubMed 40602628.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20
  52. [52]Huber J et al. Nobel and novice: Author prominence affects peer review. Proceedings of the National Academy of Sciences of the United States of America 2022. doi:10.1073/pnas.2205779119 PubMed 36194633.peer-reviewed article (Europe PMC / Crossref record) · verified 2026-09-12
  53. [53]Aczel B, Szaszi B, Holcombe AO. A billion-dollar donation: estimating the cost of researchers' time spent on peer review. Research integrity and peer review 2021. doi:10.1186/s41073-021-00118-2 PubMed 34776003.peer-reviewed article (Europe PMC / Crossref record) · verified 2026-09-12
  54. [54]Cotton CS et al. Effect of Monetary Incentives on Peer Review Acceptance and Completion: A Quasi-Randomized Interventional Trial. Critical care medicine 2025. doi:10.1097/ccm.0000000000006637 PubMed 40047491.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20
  55. [55]Rohmann JL et al. Effectiveness of the researcher-led "Peerspectives" peer review training course on review quality, knowledge, and skills among doctoral students in the biomedical sciences: a pre-post study. Research integrity and peer review 2026. doi:10.1186/s41073-026-00220-3 PubMed 42210385.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20
  56. [56]Teixeira AL. AI in peer review: can artificial intelligence be an ally in reducing gender and geographical gaps in peer review? A randomized trial. Research integrity and peer review 2025. doi:10.1186/s41073-025-00182-y PubMed 41139793.peer-reviewed article (full text on file; PubMed record verified) · verified 2026-09-20

The checkable problems are findable before you submit.

A PeerReviewAI Peer Review checks your manuscript against the matching reporting checklist item by item, verifies every reference against PubMed and Crossref with retraction screening, audits it against your target journal's instructions, and applies the journal's formatting as tracked changes. $49 per manuscript, no subscription, no verdict.

For the manuscript you're submitting