Could AI Be Better at Peer Review Than Humans?

A couple of weeks ago, I wrote a Substack note about the power and potential of AI reviewers. An editor at the Chronicle of Higher Education liked it and invited me to submit an op-ed expanding on the idea. You can read it here (it’s free, but you have to create an account), or, if you are a paid subscriber, you can just read it below.

As you’ll see, I worry we'll lose something if we move entirely to AI reviewers. But the current system is terrible, and I take seriously the idea that we would be better off using LLMs, especially if we keep some human input into more ineffable factors like originality and interest. And I think the more modest idea of adding an AI reviewer to the mix is a no-brainer.

I suspect that not everyone will feel the same way.

One concern is confidentiality. This is the easiest to address; journals could use “closed” AIs, meaning anything put into the system isn’t accessible to anyone else. This is the sort of arrangement top law firms have, and they’re more concerned about confidentiality than professors are.

A second concern is moral. Now, AI reviewing doesn’t have the same problems here as in other domains. As I write,

This proposed usage of AI … wouldn’t replace scholarship, research, writing, teaching, or advising — core academic activities that many of us take great satisfaction in. It would replace peer reviewing, and while I enjoy the activity myself, I seem to be unusual in this regard. Most people hate it; they do it out of the goodness of their hearts and for the benefit of the field.

But still, there are other moral—or possibly aesthetic—objections. Maybe you believe that AI usage should be discouraged in general, regardless of how useful it might be. (Because of big tech, water usage, data centers, copyright infringement, a non-trivial p(doom), etc.) Or maybe you hold strongly to the view that, like apologizing, flirting, and reading a story to a child, reviewing journal articles is an intrinsically human activity.

A third concern is that AI reviews aren’t that good. This is the objection that I take most seriously and want to hear more about. But I have little patience for people who tried ChatGPT two years ago, saw that it hallucinated a lot, and refuse to check again. (My own “research” was done with Claude Opus 5.5 and GPT-5.6.) I’d urge people to try this experiment:

Take a paper of yours, accepted or rejected, and look again at the reviews you received. Then run your paper through a few high-level LLMs (not out-of-date versions — the current ones). Which reviews are better?

If you try the experiment, tell me in the comments how it worked out.

添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论