When journals are slow to implement or enforce policies to manage sloppy, majority AI-authored papers, reviewers can and should push back

This is Jessica. If you’ve ever reviewed a heavily AI-generated paper, you know how unrewarding it can be. For example, last week I spent a morning reviewing a paper for a well-respected general science journal, which sounded interesting but was ultimately very marginal in terms of the delta over prior work. It had all the signs I’ve come to associate with heavy AI reliance–overloaded terms invented as unwarranted shorthand for basic statistical concepts; dense, fluent but vague writing, with lists of details and implied comparatives I couldn’t resolve even with repeated rereading; and a very incremental result—not a bad idea, just not a complete thought. It nonetheless took me over an hour and half to identify what they were trying to do and write up my thoughts in a review.

I’m now getting requests to review such papers a couple times a week–not from minor journals but from high impact, general science venues. The majority of the abstracts I’m sent these days get flagged as 100% AI generated on Pangram, which in my experience doesn’t happen if you are exerting even a small amount of effort to control what is said. I shouldn’t be surprised; I know of junior researchers submitting tens of papers a year, and some senior ones too. But I am surprised nonetheless, that I’m being asked so often by widely respected venues to give feedback to authors who have given me little reason to believe they even understand what they are submitting. Of course AI writing is not a perfect signal of low author involvement, and LLMs can lead to much better science and better writing when used well. But see enough of these mostly AI-written papers fall apart upon closer inspection and you start becoming skeptical of anything that’s flagged as mostly AI text.

The problem is the seemingly poor enforcement of policies that these journals have put in place to avoid low human involvement papers. From a quick check, many of the top general science journals have policies about human accountability (some of which are pretty complex, like Nature’s) and most require disclosure of AI use. In some cases, substantial AI generation is prohibited (like Royal Society). But my experience with review requests from such journals and what I’m hearing from others suggests these policies are not weeding out enough of the slop. I also know that as an AE for Science Advances, I have seen a fair amount of heavily AI-authored text but have yet to see an author mention AI use, though technically our policy requires this in the paper and cover letter. In trying not to overindex on writing, I’m pretty sure I’ve sent out at least a few sloppy papers for review. With their more polished surface and layers of technical detail, low quality AI-authored papers are harder to catch on a quick skim than human-generated ones. And so external reviewers are stuck doing much of the sorting that these policies are intended to take care of.

To better allocate my own attention as an external reviewer, I’ve started saying no to review requests when I’m pretty confident that I’m being presented with majority or fully AI-generated text. There are just too many papers to review for me to keep making time for those with little observable signal of human involvement.

And while I don’t think AI writing detection is a long term solution, it’s increasingly seeming that we may need to add some friction to the system to push venues to experiment with policies that could filter more effectively. So I wrote this reviewer statement with Auyon Siddiq, in the interest of normalizing reviewers’ exercising their autonomy and signaling to venues that better policy is needed.

The hope is that by encouraging researchers to assert their agency, it might help incentivize venues to move beyond the blanket “please disclose AI” statements that they don’t seem to be enforcing.

The main risk is that if many people start refusing to review on these grounds, it could make peer review much noisier temporarily. I personally think it’s worth the risk. If you use AI and detectors a lot, you get pretty good at telling when the writing involved little human effort. If you don’t use AI or detectors, you may be doing worse than chance. But if journals remain lax, we’ll likely end up with a lot of reviewers doing this kind of thing anyway. In the meantime, I recommend that reviewers take responsibility for becoming very familiar with the methods that exist, their limitations, and when they can be used without violating reviewer contracts.

There’s a much wider range of mechanisms journals could consider now: quotas on how many papers an author can submit in a given time frame, reciprocal reviewing requirements with desk rejection or future bans on submitting for non-compliance, higher standards for the clarity of writing, use of preliminary AI review to detect major issues, and so on. ML venues have been experimenting with most of these already. None of it is ideal, but when the old methods aren’t standing up to the scale of production that’s now possible, something needs to change. If we feel strongly that science requires human accountability, we should make it harder for authors to abuse the system.

P.S. Thanks to Aaron Clauset, Annie Liang, Eytan Adar, and Andrew for feedback on the statement.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论