Edition #020

|

13 Aug 2026

A journal sends a manuscript to a reviewer and expects that person to assess it. The reviewer is selected because of their knowledge, experience and ability to exercise scholarly judgement. The resulting report is assumed to represent the reviewer’s own evaluation of the work.

But can we still make that assumption?

Some reviewers are almost certainly created using generative AI. We cannot know how widespread the practice is because much of it will be undeclared. A reviewer may upload the complete manuscript and ask for a report, use AI to identify weaknesses, or generate a draft review that is edited before submission. An editor receiving a polished and detailed report has no reliable way of knowing how it was produced.

In my last newsletter edition, I considered whether reviewers should be prohibited from uploading manuscripts to AI tools when the authors may already have uploaded those same manuscripts themselves. That discussion focused largely on confidentiality, consent and transparency. But it leads to a more difficult question: if those concerns could be addressed, should reviewers be allowed to use AI to help evaluate papers?

The question is often framed as a choice between human peer review and AI peer review. That is too simplistic because AI can contribute to reviewing in many different ways.

What does using AI to peer review mean?

At one end of the spectrum, AI might carry out relatively mechanical checks. At the other end, it could replace the reviewer and perhaps even the editor. Between those extremes are several levels of involvement:

  1. Checking references, formatting and reporting requirements.
  2. Identifying suspicious phrases, citation problems or internal inconsistencies.
  3. Summarising the paper for the reviewer.
  4. Suggesting questions or weaknesses for the reviewer to investigate.
  5. Producing a draft review that the human reviewer edits and approves.
  6. Producing the final report and recommending accept, revise or reject.
  7. Making the editorial decision without meaningful human involvement.

Many people will probably be comfortable with the first few uses and increasingly uncomfortable as we move down the list. Yet the boundary is not as obvious as it might initially appear.

If AI identifies that a reference is not cited, that seems little different from an automated formatting check. If it points out that the sample size may be inadequate, it has begun to comment on the methodology. If it recommends rejection because the sample is too small, it is now exercising, or appearing to exercise, scholarly judgement.

Where should assistance end and evaluation begin?

The attraction is easy to understand

Peer review can be painfully slow. Authors may wait weeks or months (or even a year) while editors search for reviewers, invitations are declined and reports are repeatedly delayed.

A generative AI system can analyse a manuscript and produce comments within minutes. This does not mean that it would shorten the complete peer-review process by the same proportion. Journals would still need to appoint editors, find reviewers and manage revisions. But, it could greatly reduce the time a reviewer needs to produce an initial assessment.

AI could also make reviews more systematic. Even a conscientious human reviewer may overlook an uncited reference, a contradiction between two sections, a missing methodological detail or a claim that is stronger than the results justify. An AI tool could apply the same set of checks to every paper and flag issues for the reviewer to comment on.

Some tasks seem particularly well suited to automated assistance. AI might check whether every entry in the reference list appears in the text, identify potentially tortured phrases, locate unexplained abbreviations, compare the abstract with the findings, or detect phrases such as “As a large language model …” that suggest some of the text may have been AI-generated.

Used carefully, AI could enable reviewers to spend more time on the areas of the article that genuinely require expertise: whether the research question matters, whether the methods are appropriate, whether the results support the claims and whether the paper makes a meaningful contribution.

The danger of the 5,000-word review

The ability of AI to generate text quickly is both an advantage and a serious risk. A human reviewer who writes 500 words has probably made choices about which issues matter most. A generative AI system can produce 5,000 words of apparently detailed criticism at almost no additional cost to the reviewer.

But the cost has not disappeared. It has been transferred to the authors.

Every criticism may need to be investigated, answered and documented in a response letter. An AI-assisted report that takes a reviewer a few minutes to produce could require authors to spend days (weeks even) addressing repetitive, marginal or incorrect comments. The report may look thorough because it is long, while offering little sense of priority or scholarly judgement.

Longer reviews are not necessarily better reviews. AI may repeat the same concern in several forms, recommend unnecessary analyses, request references that do not exist or suggest changes that are incompatible with the research design. Authors may nevertheless feel obliged to respond to every point because the comments have arrived through the authority of the journal.

Editors have a responsibility that goes beyond collecting reports and forwarding them to authors. If AI-assisted reviewing is permitted, editors should identify duplicated comments, remove unreasonable requests and distinguish between changes that are necessary and those that are merely suggestions.

AI should reduce the burden of peer review, not move that burden from reviewers to authors.

Human responsibility must mean more than pressing “submit”

Most publisher guidance emphasises that the human reviewer remains responsible for the report. This is necessary, but it is not sufficient.

A reviewer cannot reasonably claim ownership of an assessment that they have barely read, checked or understood. Editing a few sentences in an AI-generated report does not turn it into an expert review. Nor should “the AI suggested it” become a defence when a report contains a fabricated citation, an unfounded allegation or a demand based on a misunderstanding of the method.

AI can produce criticisms that sound plausible even when they are wrong. It may favour familiar theories and conventional methods because these dominate the material on which it was trained. It may struggle to recognise a genuinely original contribution precisely because the work departs from established patterns. A system that is good at identifying what is typical may be less capable of judging what is novel.

Human oversight must therefore involve more than approving the final wording. The reviewer should verify every substantive criticism, check every suggested reference and be able to explain the reasoning behind the recommendation. If a reviewer would not be prepared to defend a comment without referring back to the AI, that comment should probably not be sent to the author.

Publisher policies already differ

The publishing sector has not reached a common position. Springer Nature tells reviewers not to upload manuscripts into generative AI tools, but asks them to disclose if an AI tool has supported any part of their evaluation. Reviewers remain accountable for the accuracy and views expressed in their reports.

Taylor & Francis states that reviewers must not use AI to generate review reports, although generative AI may be used to improve the language of a completed review. Elsevier takes a more restrictive position, telling reviewers not to upload manuscripts or related communications and not to use generative AI to assist with evaluating them.

These policies are understandable. Manuscripts are confidential documents, and public AI systems may create risks involving intellectual property, personal data and unpublished findings. AI can also produce inaccurate or biased conclusions.

However, prohibition does not necessarily prevent use. The Committee on Publication Ethics has considered cases in which editors suspected that reviewers were submitting AI-generated reports, including reports that appeared detailed but contained inaccuracies. The uncomfortable reality is that a policy may tell conscientious reviewers what they may do while having relatively little effect on those who are willing to conceal their use.

A rule that cannot be reliably monitored may encourage secrecy rather than responsible practice.

Publishers are already using AI

There is also an apparent imbalance between what publishers permit reviewers to do and what publishers do themselves. AI and automated tools are already used to support manuscript screening, reviewer identification, integrity checks, duplicate detection and other parts of the publication process.

In March 2026, Springer Nature reported that more than 1.5 million papers had benefited from nearly 60 AI tools during 2025. These tools supported activities including screening, editorial evaluation and research-integrity checks. Taylor & Francis uses Reviewer Locator, an AI-supported tool that matches submissions with potentially suitable reviewers..

There may be a distinction between a public generative AI service and a secure tool operated or approved by a publisher. A publisher-controlled system can be designed to protect confidentiality, restrict data retention and record how it has been used. Reviewers choosing whichever public tool they prefer cannot offer the same assurances.

Even so, authors should be told when AI is involved in processing or evaluating their work. They should know what the system does, what information it receives, whether its output can influence rejection, whether a human checks its conclusions and how an erroneous finding can be challenged.

Transparency should apply to publishers as well as to authors and reviewers.

From prohibition to regulated assistance

A blanket prohibition is attractive because it appears simple. However, it places checking a reference in the same general category as asking AI to write the entire review. It also risks becoming increasingly detached from actual practice as these tools become embedded in word processors, reference managers, journal platforms and research databases.

A more realistic policy would distinguish between levels of AI involvement. Journals might allow AI-assisted reviewing where:

  • The authors are informed that approved AI tools may be used.
  • The manuscript remains within a secure and confidential system.
  • The reviewer declares which tool was used and for what purpose.
  • Every substantive criticism and suggested reference is verified.
  • The reviewer personally determines the recommendation.
  • The editor removes repetitive, irrelevant or disproportionate demands.
  • Authors can challenge comments that appear inaccurate or machine-generated.
  • The journal regularly examines approved tools for errors and disciplinary bias.

This does not mean allowing AI to peer review papers independently. It would mean recognising that some forms of assistance may improve the process while establishing boundaries around the judgements that should remain human.

One boundary seems particularly important. AI may help a reviewer decide what to examine, but it should not decide whether a paper is accepted or rejected. The final recommendation should represent the informed judgement of the person invited to undertake the review, and the editorial decision should remain with the editor.

The choice we may actually face

AI will not solve every problem in peer review. It cannot guarantee that a paper is correct or replace deep subject knowledge. It may introduce new problems, including formulaic reports, fabricated criticisms and an escalation in the number of revisions demanded from authors.

But refusing to discuss legitimate uses will not keep AI out of peer review. Its use will remain hidden, inconsistent and unaccountable.

The better question may be not whether AI should be allowed to peer review a paper, but which tasks we are prepared to delegate and which judgements must remain human.

The choice may no longer be between human peer review and AI peer review. It may be between hidden, unregulated AI use and transparent, accountable AI assistance.


I publish two free newsletters each week

Governance & Research in HE
A weekly brief for higher education leaders and researchers examining governance, leadership and research in practice.
https://buff.ly/GhsZqVl

Publishing with Integrity
Examining scholarly publishing through an ethical lens, challenging assumptions and supporting academic career development.
https://buff.ly/WiG1yFb

If you subscribe, you will be alerted whenever I publish a new edition.


About the Author

Professor Graham Kendall is Vice-Chancellor of GlobalNxt University, Malaysia, and an Emeritus Professor at the University of Nottingham. Over the past 25 years, he has published more than 300 peer-reviewed papers and has served as Editor-in-Chief and Associate Editor of several international journals. Through Publishing with Integrity, he explores the ethical, governance and practical challenges facing scholarly publishing, encouraging greater transparency, integrity and informed debate across the global research community.

Originally published on LinkedIn

This edition was first published as part of my LinkedIn newsletter. If you use LinkedIn, I recommend reading it there, where you can also join the discussion. This version is provided particularly for readers who do not have a LinkedIn account.

Would you like to know when the next edition is published?

Receive an email whenever I publish a new edition of my newsletter.