Should our journal publish AI-drafted manuscripts?

Forget both truth and beauty, I want to know about opportunity costs

Status: Rough conceptual model.This is a personal exploration of a live policy problem and definitely does not represent the opinion of the Alignment Journal itself.Given the context, I had best disclose my own AI usage in this article: transformative. Although the original model design was mine, it was made way better by iterative refinement and re-drafting by AI, and by no means would I have had time to write it purely by hand.

At the Alignment Journal we have been discussing whether to accept AI-drafted manuscripts for review. This is a relatively high-leverage question, as detecting AI-drafted prose is (currently, against authors who are not trying to hide it) surprisingly feasible, making it a cheap (albeit imperfect) proxy signal for us to use in desk review to filter out low-quality papers.[1]

The ideal policy would optimize for the overall quality of the journal’s output with regard to how that serves our readers. There are many components to that; quality, readability, professional and academic norms… I ignore most of those and bloody-mindedly focus on the economic effects of AI drafting.

On one hand, AI drafting lowers the cost of producing a manuscript, which may increase the flow of good work into our readers’ inboxes. On the other hand, it might instead increase the flow of low-quality AI slop into those inboxes. That is, we worry about the impact AI slop spam, as do many others (Gartenberg et al. 2026).

Insofar as the base costs of a high-quality paper exceed those of a slop manuscript, we might hope that permitting AI drafting provides a relatively larger benefit to slop authors than to substantive authors.

So, if we can detect AI drafting, should we do so? And: should we use it to desk-reject papers?

Here is a minimal economic model of that question, in the vein of Agrawal, Gans, and Goldfarb (2019), which models AI as a fall in the price of prediction. My assumptions are stylized and tractable rather than realistic or empirically calibrated. Where I need to choose some numbers, I have taken inspiration from the Alignment Journal, since that is closest to my heart.

Here, I assume AI drafting reduces the cost of producing prose: the writing-up stage of science gets cheap while the substance stage, for now, does not (as per Kwa et al. 2025). That is a shaky assumption even now (for example, both my drafting and substantive work in writing this very post were AI-expedited) and it will get shakier as AI progresses.

My goal here is not to persuade anyone of a particular optimal policy, but to lay out a concrete model of the first- and second-order economic effects of AI drafting so that we can debate these things better. Too often the discussion is conducted in purely moral terms (AI is “efficient” or “lazy”) or aesthetic ones (“AI prose is hard to read”); those lenses matter, but our duties in publishing surely also include the economic and strategic management of scientific knowledge production, and quantitative models of that seem under-supplied.

tl;drUnder my assumptions:
  1. AI drafting preferentially incentivises production of poor papers, which crowd out good ones: the time it saves good authors turns into extra manuscripts that are never read.
  2. A ban enforced by an AI-drafting detector can induce three different regimes, depending on detector quality:
  • At low true positive rate it changes nothing;
  • at intermediate levels it deters only the good authors, who give up AI before slop authors; then,
  • if the detector is good enough, it restores the no-AI outcome, forfeiting every hour AI would have saved.
  1. A flat submission fee lets everyone use AI while excluding spam, and beats the ban outright. The fee that does it is ballpark USD 2,500 to 7,500 per submission, which is outside the Overton window and moreover a business model indistinguishable from a predatory journal’s, so nobody, least of all me, is proposing it.
  2. A hybrid, where authors may pay to skip the detector, charges only AI drafters and publishes more good papers than the ban but fewer than the flat fee, because slop authors who draft by hand still clog the queue. The fee that does it is USD 4,000 to 12,000 per AI-drafted submission if the detector holds up against motivated rewriting..
  3. Not modelled: false positives, reputational effects, harmful slop, competing journals, and AI that does the science as well as the writing. And, clearly, I remain silent on whether AI can do good writing or not.
Putting that all together we can see that, regarded purely as an instrument for controlling slop rates, the detector ban is a clumsy approximation to a congestion charge, but it might be better than nothing and is more palateable than fees would be.

Setup

I assume there are two types of authors: high-quality authors (type ) and slop authors (type ). They differ only in the substance of their manuscripts and behave strategically about the presentation. The proportion of high-quality authors is , and the rest are slop authors. Articles have apparent substance , which is the quality the manuscript presents to a reviewer—the quality a reviewer can assess. I measure it in citoms, my idiosyncratic unit of “citability,” which happens to be what citations count. These are themselves an imperfect proxy for our true desired output, scientific progress , which is what a reader gets if the manuscript is published. I measure that in scientillas.[2]

A high-quality author spends substance time on each manuscript—the hours of research and thinking as opposed to writing up. The result is a manuscript that reads as good as it actually is, with as many citoms as scientillas:

A spread of author skill would enter here as a multiplier on ; I ignore it.

The rest of the authors produce slop, which in my model means manuscripts worth citoms at a glance but zero scientillas upon deep reading, , made with the minimum substance time . All of which is to say, the authors aren’t deceptive, whereas the authors produce low-quality outputs masquerading as adequate.

Every manuscript additionally requires drafting time : hand-drafting costs , AI drafting costs , where is the AI’s drafting efficiency. At AI saves nothing, so everyone might as well hand-draft; I call that no-AI state the hand-drafting equilibrium, and it is the baseline for all the other policies. An author has a time budget per period (year?), so an author who produces manuscripts at a cost of substance effort time satisfies . We write and for the substance time and manuscript count an -type author chooses, and for a slop author’s manuscript count; slop authors have no substance time to choose.

Reviewing a manuscript to the journal’s standard takes reviewer-hours, which we call the depth of review, and there are total reviewer-hours, so the journal can review at most manuscripts. Facing submissions, it reviews of them and rejects the rest unread, so a submission’s chance of being read at all is the coverage rate

A manuscript that receives a review is accepted with a probability depending on both its substance and the review time ,

i.e., deeper reviews detect the scientific substance more precisely. Here is the journal’s bar, the number of citoms at which a reviewed manuscript is accepted half the time. We write and for the acceptance probability of a reviewed high-quality manuscript and a reviewed slop manuscript, respectively. The journal holds the depth fixed with respect to queue size. Review slots are the scarce resource in everything that follows: every submission takes one, and a submission that goes unread costs its author nothing but costs the journal whatever it displaced.

Authors are assumed to be careerist, in that they care only about a payoff per accepted paper. The payoff is earned by clearing the bar in citoms; scientillas earn an author nothing beyond that. A high-quality author facing a drafting cost thus solves

maximizing accepted papers per unit of time. No author can move the coverage rate , and it scales the payoff of every candidate identically, so it drops out of the optimization. As such, congestion never influences the choice of how much effort to dedicate to substance. What does influence the choice is the drafting cost , which is by hand or with AI. Every hour spent on the substance of one manuscript is an hour not spent starting the next, and controls how much that next manuscript would have cost. The first-order condition is

As AI capability rises, falls, the next manuscript gets cheaper, and the optimum tilts toward more, thinner papers.

Now let us consider the target audience, the readers. Recall our benevolent planner goal to maximize the readers’ welfare. I operationalize that by the number of good papers published per period, the total amount of science we are pumping out to all readers. Of the high-quality manuscripts written, a fraction get reviewed and a fraction of those pass, so the flow of high-quality papers into print is the throughput

The obvious alternative is the average scientillas per published article, , roughly the amount of science a reader finds sampling the published literature at random; slop counts against , whereas merely ignores it. Every ranking of one policy against another below comes out the same way under both, so I track and mention only at the one point where they disagree, which is how high to set the fee in . Both count scientillas, which measure public-good outcomes. The authors’ incentives, by contrast, count citoms, career-wise goals measured in citations. Slop writing is the stuff that produces the latter without the former.

I make three other simplifying assumptions:

  1. AI drafting here introduces no errors and incurs no cognitive debt[3]
  2. Accepted slop delivers zero value but is not actively harmful.[4]
  3. The review budget is exogenous: is money divided by the going wage for qualified reviewer attention, and the model holds both fixed.

Policy A — Free-for-all

Under this policy anyone may use AI, and we triage submissions according to capacity (if there are too many papers we reject the excess). tl;dr everyone drafts their manuscripts with AI, as it is cheaper and goes unpunished, and the extra manuscripts mostly go unread. The equilibrium is easy to compute: the substance choice depends only on the drafting cost , submission counts follow from the time budget, and coverage follows from the counts.

Figure 2: The free-for-all as AI drafting capability grows: substance time per high-quality manuscript, the coverage rate , and the throughput . falls somewhat; the large damage comes from crowding.

Optimistically, we might hope that cheaper AI drafting frees up time compared to the manual alternative, allowing for more substance per paper. That does not happen under this model. The left edge of the capability sweep, , is the hand-drafting equilibrium. For the high-quality tier , the substance time per manuscript falls modestly as rises, from to about : cheaper drafting makes each manuscript cheaper, so authors choose quantity over quality. The fall is not too precipitous though: high-quality manuscripts stay above the bar ( drifts from down to scientillas, against a bar of citoms), and their authors write three of them where they used to write one. For the low-quality tier , the slop authors’ output grows much faster, because a slop manuscript is nearly all drafting time: rises from under to over .

The downside is visible through the effect on the queue. The submission pool ends up slop and coverage falls from roughly to , so the journal publishes about four times fewer high-quality papers, even though the high-quality authors are writing more of them than before. Most of their extra manuscripts are simply never read. In the units of the model that is falling from to good papers per author-period: a journal drawing on a hundred authors goes from printing 22 good papers a period to 6. The average quality of what it does print falls too, since slop passes a full review of the time and the pool is slop, so most of what the journal accepts is slop.

The collapse is smooth and the equilibrium is unique. The coverage rate cancels out of every author’s problem, so there is no feedback loop to amplify anything, and no tipping point. We could expand upon that but do not right now.

Policy B — Automated ban on AI-drafted manuscripts

A stylometric detector (e.g. Pangram) flags manuscripts it thinks are AI-drafted, and we desk-reject those.[5] This is the only desk-reject stage; anything it passes goes into the capacity triage of the review queue. I assume the detector has true-positive rate , which the machine-learning literature calls its recall: it correctly flags a fraction of AI-drafted manuscripts. For now, I assume it has a zero false-positive rate to keep it simple.[6] Desk-rejected papers consume no reviewer time. The rest get a normal review.

Before anyone changes behaviour, notice what the detector is to an author who drafts with AI: a fee. A fraction of their manuscripts is thrown away unread, so each one loses times its expected return, for a good manuscript and for slop. That is a proportional levy on expected return, and since it charges a good manuscript more than a slop one; a separating fee would do the opposite. We return to this with the . That by the way, is the sense in which the ban is an imperfect approximation to a fee.

Now, knowing this policy, everyone chooses how/if to write their papers. For low-quality authors, the trade-off is between the time saved by using AI and the risk of being desk-rejected. Their acceptance probability if they get past desk-rejection is the same however the manuscript was drafted, so it cancels from the comparison, and what remains is: papers per hour with AI, discounted by the survival rate , against papers per hour by hand. They keep AI drafting while , where

In this setting, the detector’s true-positive rate is pivotal, and the threshold it needs to clear depends on the time costs of producing slop. In particular rises with AI capability toward , the share of a hand-drafted slop manuscript’s time that goes on drafting— about at my parameters. As such, a detector with a true-positive rate below that threshold deters slop authors only while AI is weak enough that is under its rate; as capability grows past that point, the same detector stops working without any change in the policy or the detector. On unmodified LLM output the best commercial detector clears that comfortably: Pangram reports a false-negative rate of at a false-positive rate of (Glickenhaus et al. 2026), and independent tests put it near (Russell, Karpinska, and Iyyer 2025). The rate that matters, though, is the one sustained against slop authors who rewrite to evade detection once it cuts into their earnings, and there the numbers are grim: Pangram’s own figure on AI-edited human text is – (Glickenhaus et al. 2026), and on academic abstracts a rate falls to once rewrite prompts are searched against the detector (Ren, Raghavan, and Garg 2026) and below after a commercial humanizer (Karr et al. 2026). And because rises with capability, a ban that deters slop today can fail later without any policy change and without any decline in the detector. We ignore that twist for now.

The high-quality authors have a different behaviour threshold. Write for the best rate of expected acceptances per unit time that a high-quality author can achieve at drafting cost . High-quality authors abandon AI drafting once falls below , that is at

The two thresholds are strictly ordered: for every (at both are zero). To see this, let be the substance time a high-quality author chooses when drafting with AI. A hand-drafting author could choose that same ; it is not their optimum, so is at least what it yields, namely . The ratio increases in , and a slop author is at , so their ratio is smaller and their threshold is larger. In other words: the more substance time we invest in a manuscript, the smaller the share of its cost that AI drafting saves, so the less detection risk an author will accept to keep using it. A detector therefore stops high-quality authors from using AI strictly before it stops slop authors. Neither type is ever deterred from submitting: at worst, slop authors revert to hand-drafting, which restores the hand-drafting equilibrium’s spam rate, never less. Being flagged costs a high-quality author a manuscript full of work; it costs a slop author almost nothing.

Figure 3: The detector ban at a fixed (low) true-positive rate as AI capability grows, in the panels of Figure 2, with the free-for-all in grey for comparison. Vertical lines: where passes (left) and where does (right). Left of the first line the ban deters slop authors from AI drafting, so everyone hand-drafts and the hand-drafting equilibrium of persists whatever is; between the lines it deters only the high-quality authors, who hand-draft while slop authors keep using AI; right of the second it deters nobody and the free-for-all returns.

Three regimes appear as capability grows, with the detector's true-positive rate held fixed. While is still below the detector’s rate (left of the solid line), even slop authors dare not use AI, everyone hand-drafts, and the hand-drafting equilibrium persists whatever is: the ban works, at the price of forfeiting the hours that an AI would have saved. Once climbs past the detector (between the lines), the ban’s only behavioural effect is on the wrong people. High-quality authors still find AI not worth the risk and hand-draft, writing one manuscript where they could have drafted three; slop authors draft with AI, submit at full rate, and write off the fraction that is desk-rejected as a cost of doing business. The ban thins the queue here because chokes the inflow of spam, not rather than because anyone has stopped spamming; the good authors’ fewer, fuller papers raise the average quality of what is printed, not the count. Once climbs past the detector too (right of the dashed line), the ban deters nobody. Both author types draft with AI, the detector desk-rejects the same fraction of each type’s manuscripts, and so long as the queue is congested the review slots freed by desk rejection compensate for the good manuscripts lost, so the free-for-all equilibrium returns unchanged. Read against the detector instead, at fixed capability, the same three regimes run in reverse: nothing below , a tax on the compliant between and , and the hand-drafting equilibrium above . If we instead hold capability fixed but raise the true-positive rate , the same three regimes appear in the opposite order: below the ban changes nothing; between and it taxes only the compliant; above the hand-drafting equilibrium returns.

The standard objection to a ban is that it takes the time savings of AI drafting away from the compliant. My model agrees that the ban forfeits those savings, but so does the free-for-all, which converts them into extra manuscripts that mostly go unread. The savings are only worth having under a policy that can turn them into published papers, and that is what the submission fee of Policy C does.

Policy C — submission fees

Under this policy, we charge per submission, allow anything, triage and review whoever pays for it. To be clear: this is not on the table for the Alignment Journal, but is included as a baseline.

As you read this section you will note that we are talking about princely submission fees, relative to the academic norm. This is because of the bonkers economy of academic publishing which pours vast amounts of labour into producing papers, much of which is effectively un-budgeted.I could estimate the “value” of the paper in a few ways — the effective of a paper in getting tenure, the amount of submission fee that an author seems prepared pay, or the approximate value the papers have as represented by the dollar cost of the hours that go into a typical one.I chose the last here, since it seems the easiest to estimate (I’ve written papers), but in some ways this leads to surprisingly large valuations: it shows that papers are extremely expensive in time cost, even though academics tend to have smaller cash budgets to pay in submission fees than this time cost would indicate.Whether anyone would be prepared to pay fees obviously depends on not just their time cost, which is sunk, but the “value” of the venue (conference, journal) to which they submit, which is also nebulous. For reference, submitting to and attending a top tier conference can cost ~USD5000 in travel, registration, and accommodation fees.

An economist would say that submission fees are always charged, in waiting time if not in cash, and a journal chooses some mix of the two whether or not it admits to pricing (Cotton 2013). This does not change how it feels to be charged such a fee, nor the fact that many authors have no cash budget to pay it. Nonetheless…

Recall that a submission’s expected private return is times its acceptance probability: for slop, and roughly — much larger — for a high-quality paper. Any fee between those two numbers is separating: submitting slop now loses money, while submitting good work still pays. Nobody is forced back to hand-drafting, because the fee is the same however the paper was drafted — which is fine, since here at least we don’t care about provenance except as a proxy for spamminess.

If the submission fee is too low, slop authors keep entering until congestion drives their expected return down to the fee, . That pins the coverage rate at : slop manuscripts fill every review slot that high-quality manuscripts do not. Raising buys back coverage, but the accepted papers stay mostly slop until the fee exceeds what a slop author values each submission at.

Full exclusion of slop is sadly expensive. A slop author facing an uncongested queue () expects per submission, so the fee at which slop authors stop submitting entirely is

about an eighth of the private value of an acceptance (at my parameters). does not depend on : the review depth is fixed, so a slop manuscript’s acceptance odds are the same however large the flood, and the fee has to do all the deterring by itself, which gives this policy the nifty feature of excluding slop authors with a single fee setting even as capabilities advance, unlike the ban, which has to be re-checked as moves. At every slop author stops submitting, at least if they are rational expected-value maximizers.

Figure 4: The three policies as AI capability grows. Detector at true-positive rate and ; fee at the full-exclusion level . Only the fee regime turns more capable AI into more good papers published.

Interestingly, under the fee policy, more AI capability is simply good. High-quality authors keep all the time AI saves them and spend it writing more manuscripts. Huzzah! The journal, its queue protected, has the capacity to review and publish them; and the flow of good papers into print more than doubles across the plotted range. Under the other regimes, we burn that time surplus: the ban burns it by forcing high-quality researchers back to hand-drafting, and the free-for-all converts it into unread submissions. The fee curve in the figure is computed at , the price at which slop authors stop submitting, so no slop reaches the accepted papers at all. At any lower fee, slop authors keep submitting until the queue congestion lowers their expected return to the fee, so slop manuscripts fill every review slot the high-quality manuscripts leave free, and reviewers still accept some of those.

The fee’s lead has a second component that does not show: the share of what gets printed that is high-quality. We can understand the gap as a mismatch between the proxy (was AI used to write this?) and the true target (was this article high-quality?). The best the ban can do, at any true-positive rate, is restore the hand-drafting equilibrium, which at these parameters means about high-quality articles; a fee at removes all slop from the queue, so the share is . An entailment of that calibration is that the journal was hypothetically going to be about bullshit in the absence of AI drafting. That number follows from parameter choices I made for , and , which, as I said, are not crazy. Readers who believe their journal is purer than that should raise or and watch the fee’s lead on that share shrink accordingly: the fee’s advantage arises from the hand-drafted slop it eliminates, so the dirtier we think journals already are, the stronger the case for pricing submission over policing provenance, and vice versa.

Of course, many journals cannot charge submission fees, for reasons of custom and equity, and because no matter how eloquent the justification it will still look like predatory-journal behaviour; the Alignment Journal is not exempt. The deadweight losses of dealing with spam are still imposed on someone, and in fact authors bear them, paying in lost time queuing for a review slot: with cash fees near zero, first-response delay is the de facto submission price (Azar 2005). So Policy C is useful as a benchmark: it charges for review slots in cash rather than in waiting time, and the gap between it and the best policy we can feasibly adopt is the price of dwelling in the earthly realm of imperfect expected-utility maximizers and departmental budget constraints. A journal that did go this way would need to do it transparently, for instance by spending the fee revenue on reviewing, as a journal that pays its reviewers could; a fee refunded on acceptance charges bad papers more than good ones in expectation, at the cost of a moral hazard of its own. Elaborations are left as homework.

Policy D — authors pay to skip the detector

Authors could pay a fee per manuscript to have it reviewed without passing through the detector. Manuscripts that do not pay go through Policy B: the detector runs, and what it flags is desk-rejected. Hand-drafted manuscripts are never flagged, so their authors have no reason to pay. The fee is only ever paid by authors who drafted with AI and would rather not gamble on the detector. It imposes Policy C’s submission fee, but charges only AI drafters. Each manuscript now has three routes: draft by hand and pay nothing, draft with AI and risk desk rejection, or draft with AI and pay the fee.

Consider the slop authors first. A slop manuscript’s acceptance odds do not depend on how it was drafted, so if it reaches review it is worth to its author however it got there. An author who drafts with AI and does not pay has a fraction of their manuscripts desk-rejected, which costs them per manuscript in expectation. So paying the fee beats risking desk rejection only when the fee is less than . Paying the fee beats drafting by hand only when the fee is less than , by the same arithmetic that gave us under the ban. So a slop author pays the fee only when

Call the right-hand side the floor fee, : below that price every slop author pays to skip the detector, and above it none does. Any fee the journal wants slop authors not to pay must be at least that high. At a fee above the floor, slop authors ignore the exemption option and behave as under Policy B. Below they draft with AI and submit anyway, writing off the fraction of their manuscripts that is desk-rejected as a cost of doing business. Above they draft by hand, are never flagged, and submit at the hand-drafting equilibrium rate. The floor is lower than Policy C’s full-exclusion fee by the factor . Skipping the detector is worth exactly the fee it stood for, for slop, and in a congested queue a submission is reviewed only with probability , which scales its expected return and so the fee an author will pay ( was defined at ). That is the floor, not the fee the journal ends up charging. The optimal exemption fee comes out above at my parameters, as we see in a moment, because it is set by congestion among the high-quality authors rather than by what it takes to deter slop.

The high-quality authors need to estimate

which is a fancified version of the same calculation in Policy C, differing in a couple of elaborations.

First, the coverage rate no longer cancels. The fee is a fixed sum while the return scales with , so the amount an author writes now depends on how congested the queue is, and that congestion depends on how much everyone writes. The equilibrium is the point where the two agree. It is unique: a higher makes the fee smaller relative to the return, which makes authors write more, which lengthens the queue and lowers again. Second, they pay only while the exemption is worth buying. At a given substance choice, paying the fee beats risking desk rejection only when . Call this the cap, : the detector’s fee-equivalent for a good manuscript, and so the highest fee the high-quality authors will pay. If the detector rarely catches anyone, nobody will pay to skip it. For those who do pay, the fee acts like an extra drafting cost. It pushes toward fewer, more substantial manuscripts, the opposite of what cheap drafting did in the first-order condition above.

Which fee? Think of the review budget as a road with a fixed number of lanes. Every submission takes a review slot, and when the queue is congested () one more submission pushes some other manuscript out unread. If the pushed-out manuscript was good, the journal loses a good paper, and the author who did the pushing pays nothing for that loss. Economists call this a congestion externality, and the standard remedy is a congestion charge: we bill each author for the damage their submission does to everyone else’s. In this model the charge is easy to write down. In the congested case the throughput is , where is the number of accepted high-quality manuscripts per period and is the length of the queue. One more high-quality manuscript adds to and one slot to , so it changes by . The author counts only the first term, the paper they might get accepted. The second term accounts for the good papers they displace, which costs fall on the other high-quality authors. A fee equal to the second term, converted to money at per paper, makes the author’s calculation match the journal’s, and that is the optimal fee.

where is the high-quality share of the queue. In words: the return to a high-quality submission, discounted by the share of the queue that is high-quality. The fee must also be set above the floor , so that slop authors don’t pay it, and yet below the cap , so that high-quality authors do. That gives

Everything on the right depends on the equilibrium, and the equilibrium depends on the fee, so we solve this by numerically iterating to convergence. Is the formula right? Figure 5 checks it the slow way: hold and fixed, try every fee from zero upward, solve the equilibrium at each, and plot the throughput that results. The fee at which that curve peaks is the best fee the journal could charge, found by search rather than by formula; the solid vertical line marks where the formula says the best fee is, and the two agree. The folded cell after the figure repeats that comparison at four capability levels, from to , and six true-positive rates, from to , and reports the largest amount by which charging the formula fee falls short of the best fee found by search. It is zero at every point of the grid. When the congestion charge is below the cap, the best fee is itself. When would exceed the cap, the best the journal can do is charge the cap: any higher and the high-quality authors stop paying, take their chances with the detector, and falls off the cliff visible in Figure 5. At my parameters, the congestion charge is always above the floor, so the floor never matters. When the queue is uncongested, nobody displaces anybody, the charge is zero, and drops to the floor, which at is , a fraction of the of Policy C, because a slop author who cannot slip past the detector still has hand-drafting to fall back on.

Figure 5: The exemption fee at : throughput against the fee , at two true-positive rates. Solid verticals mark the formula ; dotted verticals mark the floor , below which slop authors pay the fee too. The cliff on the right of each curve is where high-quality authors stop paying and take their chances with the detector.

The detector’s true-positive rate does not appear in the congestion charge. In Policy D the detector does one thing: it gives AI-drafting authors a reason to pay, by desk-rejecting a fraction of AI-drafted, non-fee-paying manuscripts. The true positive rate enters only the floor and the cap, both of which rise with because skipping a better detector is worth more, so a better detector widens the range of fees that high-quality authors will pay and slop authors will not. At a low enough true-positive rate the cap falls below the congestion charge, and the journal can charge only what high-quality authors will still pay. Unlike , this fee changes as grows, because and do. This is the one place the average-quality metric would disagree with our overall value metric : $ is blind to volume so long as the average quality is high, so it would prefer the largest fee the high-quality authors will pay, right up to the cap. Near the fee pushes high-quality authors back to hand-drafting, because AI saves them almost nothing there, and Policy D degenerates into Policy B.

Figure 6: The four policies as AI capability grows. Detector and exemption fee at true-positive rate and ; submission fee at ; exemption fee at . The exemption fee is between the ban and the submission fee at every capability level .

The exemption fee lands between the ban and the submission fee at every capability level in Figure 6. Above the ban restores the hand-drafting equilibrium, good papers per author-period at my parameters, regardless of ; the exemption fee at the same true-positive rate publishes at (39 a period per hundred authors, against 22), because high-quality authors keep the time AI saves them and the congestion charge stops them from spending all of it on extra manuscripts. Below , at , the exemption fee roughly doubles the ban’s throughput, but both are poor, because the of AI-drafted slop manuscripts that the detector misses still fill the queue. What separates the exemption fee from the submission fee is hand-drafted slop: under Policy C slop authors pay whether or not they used AI, so they stop submitting; under Policy D a slop author who drafts by hand pays nothing, is never flagged, and stays in the queue. At every AI drafter pays, and Policy D is Policy C levied on AI drafters alone; the slop authors who draft by hand are the whole difference.

For the Alignment Journal the exemption fee faces the same objection as the submission fee, that we would be charging money from a cash-constrained author base. Only authors who choose to pay are charged. Paying is also a disclosure: a hand-drafted manuscript is never flagged, so the only reason to buy the exemption is that the manuscript was AI-drafted. Most journals now ask authors to declare AI use but can at best partially check the claim; here it arrives with a fee attached, so it is truthful by construction. A false positive still costs a hand-drafting author their manuscript, as under the ban. And the revenue arrives in proportion to the number of AI-drafted manuscripts the journal has to review, so a journal that pays its reviewers could route it straight back into .

Future work

On this analysis the detector ban is a clumsy approximation to a congestion charge, and Policy D makes the charge explicit: what we are charging for is constrained review slots. Whether the Alignment Journal, or any journal, has a reasonable economic basis to adopt the ban depends on two parameters I have only guessed here:

  • the true-positive rate a detector can sustain against motivated rewriting, compared against a we could estimate from the time costs of producing slop; and
  • how many good AI-drafted papers we would tolerate losing in the range of true-positive rates between the two thresholds.

Not modelled

  • Diluted review. Throughout, the journal holds review depth fixed and rations coverage by skipping papers, which is why every collapse above is smooth and every equilibrium unique. A journal that instead reviews everything less carefully puts congestion inside the authors’ incentives, and at some author mixes the decline becomes a fold with hysteresis: a cliff that does not reverse when capability is walked back. That variant has its own post, Diluted review; Bartolucci and Vivo (2026) works the same margin out properly in a queueing model.
  • More capable AI. Slop’s citoms are fixed here; the realistic case has them rising with , and as the range of separating fees closes, at which point the journal’s problem stops being mechanism design and becomes epistemology. More broadly, the premise that AI can write a paper but not do the science behind it is a statement about current capability, and the measured trend is that the boundary moves; none of these results survive the regime where it fails. Nor can high-quality authors oversell, since citoms and scientillas coincide for them by assumption; if AI drafting made real but modest work read as better than it is, they would gain an inflation margin of their own.
  • Harmful slop. Accepted slop is harmless here; if it instead poisons training corpora (Shumailov et al. 2023) or locks in error (Qiu et al. 2025), its value is negative and everything above understates the case for keeping it out.
  • Many journals. There is only one journal here; with many, slop authors send their manuscripts to whichever is least strict and there is a race to the bottom, which is a different and worse game.
  • False positives. The detector is assumed never to flag a hand-drafted manuscript, which is unlikely in practice; a false positive costs a compliant author a manuscript full of work, so even a small rate changes the outcomes near and erodes trust.
  • Reputations. If publishing slop were reputationally harmful, the return to slop would change and the amount of spam might also be moderated.

Ballpark real-world estimates

The model prices everything in units of , the value of one accepted paper to its author, so a fee in dollars needs a guess at in dollars. My best anchor is cost: a typical ICLR paper has historically taken about four researcher-months, and a careerist keeps spending that only if an acceptance is worth at least as much to them, so four researcher-months is a lower bound on . At USD 5,000 to USD 15,000 per researcher-month, from a PhD stipend to a fully loaded academic salary, that puts at USD 20,000 to USD 60,000. I will carry both ends through.

Under Policy C the fee at which slop authors stop submitting is , which at my parameters is about an eighth of : USD 2,500 to USD 7,500 per submission.

Under Policy D with a Pangram-grade detector, taking the true-positive rate from credible independent tests (Russell, Karpinska, and Iyyer 2025) as (i.e. ) with no false positives, that rate clears the large- limit, i.e. for every , so slop authors draft by hand, take their chances and never pay the fee. The fee is then a pure congestion charge. It rises slowly with , from at to at . Call it a fifth of : USD 4,000 to USD 12,000 per AI-drafted submission. The floor is –, well below that, so slop authors do not pay at any fee in this range. Those dollar figures assume the detector wins the arms race. At the that survives a rewrite attack, the exemption fee collapses to , a few hundred dollars, because skipping a detector that catches a quarter of manuscripts is worth little to anyone, and throughput barely improves on the free-for-all ( against good papers per author-period). A detector that slop authors can defeat prices nothing.

The ban has a price too, paid in time rather than money. Above every high-quality author drafts by hand, so each of their manuscripts costs more time than it would with AI, about time units at . On the same anchor, a paper takes about time units and four researcher-months, so the ban costs a good author about two researcher-months per manuscript: roughly half of , or USD 10,000 to USD 30,000. That is between two and three times the exemption fee, and the journal does not even collect it. A second way to price the ban gives its fee-equivalent directly: the exemption fee at which a high-quality author would be as well off as under the ban is at , USD 9,000 to USD 26,000, falling to at , USD 4,000 to USD 11,000, where AI saves little drafting time and the ban costs little. Either way the ban is a fee of the same order as the exemption fee, levied in time, on the good authors only, and collected by nobody.

The ideal exemption fee face value is higher than the submission fee. Under Policy C the fee only has to keep slop authors out, and once they are gone the queue is uncongested, so there is no congestion to charge for. Under Policy D the slop authors who draft by hand stay, the queue stays congested, and the fee is a congestion charge on the high-quality authors.

Both fees are large next to what journals that do charge actually ask; economics journals charge around USD 100–300 per submission, I believe. Either the going rate is low, or is smaller than I have guessed for most authors, or my is too high. A journal whose full review lets through fewer than of slop manuscripts needs a proportionately smaller fee, since scales with .

Incoming

References

Agrawal, Gans, and Goldfarb. 2019. The Economics of Artificial Intelligence: An Agenda.

Akram. 2023. “An Empirical Study of AI Generated Text Detection Tools.”

Azar. 2005. “The Review Process in Economics: Is It Too Fast?” Southern Economic Journal.

Bartolucci, and Vivo. 2026. “Queue & AI: When Faster Tasks Slow Down the Workflow.”

Bauer. 2026. “The Last Costly Signal: How Generative AI Collapses Competence Signaling and Why Liability Sustains Markets for Expert Services.”

Cotton. 2013. “Submission Fees and Response Times in Academic Publishing.” American Economic Review.

Cui, Dias, and Ye. 2025. “Signaling in the Age of AI: Evidence from Cover Letters.”

Emi, and Spero. 2024. “Technical Report on the Pangram AI-Generated Text Classifier.”

Fourie. 2026. “Messy Research, Certification and the Monetization of Science.”

Gartenberg, Hasan, Murray, et al. 2026. “More Versus Better: Artificial Intelligence, Incentives, and the Emerging Crisis in Peer Review.” Organization Science.

Glickenhaus, Thai, Russell, et al. 2026. “Pangram 4 Technical Report.”

Karr, Khvatskii, Hua, et al. 2026. “Why AI Detection Fails for Academic Integrity.”

Kosmyna, Hauptmann, Yuan, et al. 2025. “Your Brain on ChatGPT: Accumulation of Cognitive Debt When Using an AI Assistant for Essay Writing Task.”

Kwa, West, Becker, et al. 2025. “Measuring AI Ability to Complete Long Tasks.”

Lee. 2025. “The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers.”

Qiu, He, Chugh, et al. 2025. “The Lock-in Hypothesis: Stagnation by Algorithm.” In.

Ren, Raghavan, and Garg. 2026. “Hitting a Moving Target: Test-Time Adaptation for AI Text Detection Under Continual Distribution Shift.”

Russell, Karpinska, and Iyyer. 2025. “People Who Frequently Use ChatGPT for Writing Tasks Are Accurate and Robust Detectors of AI-Generated Text.”

Shumailov, Shumaylov, Zhao, et al. 2023. “The Curse of Recursion: Training on Generated Data Makes Models Forget.”

Footnotes

  1. Aside: the quality bar for academic prose is sadly low.
  2. The obvious alternative, the millikuhn, one thousandth of a paradigm shift, is too coarse for most things a journal publishes.
  3. pace Kosmyna et al. (2025) and Lee (2025).
  4. We could be cute here. The model treats slop as zero scientillas rather than hypothesizing slopons, the antiparticle of the scientilla, .
  5. The detector reads style, not substance, and the two are coming apart: in the neighbouring market for cover letters, an AI writing tool cut the correlation between text quality and callbacks by half (Cui, Dias, and Ye 2025).
  6. In the more realistic case that it has a nonzero false-positive rate, .
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论