When they can perform a task, AIs are much cheaper than humans

AI systems are increasingly capable of substantial work. I'm old enough to remember 2025, when METR's time-horizon graph climbed from seven-minute tasks at 80% reliability at the start of the year to tasks taking more than an hour by the end. The time horizons for Astra and Fable 5.1 are now so long that METR's current task suite cannot reliably estimate them. But the recent Hugging Face attack and a slew of mathematics results show that frontier systems are now capable of some tasks that would take months or years of human effort.

Sure, AIs have now solved a Millennium Prize problem, but at what cost? Doing a task is not the same as doing it cheaply. AI systems will only replace human workers if running them costs less than paying the workers. Perhaps these feats are expensive, so that even once AI can do the work, compute limits how many workers it replaces. But if AI systems are already cheap and AGI really is a few years away, we could soon be living in a world with vast numbers of digital workers, which could threaten not only our jobs but also our continued control over the future.

The data

To investigate, I used GPT-6 Astra and Fable 5.1 to create a database of AI tasks, comparing the compute an AI system needs to perform a task with the time a human would take to perform the same task. The estimates of both are necessarily crude. Frontier AI companies don't publish how much compute their latest models use, how many parameters they have, or what their architecture is, so our compute figures are usually guesses. Nor is it always easy to know how long a task would take a human, which in any case can be quite variable. Nevertheless, for a rough order-of-magnitude picture of compute costs, the estimates are good enough.

The graph below plots the AI compute required against the human active time for the tasks in the dataset. Where several AI systems have done a task, only the cheapest that reached at least human performance is shown. You can explore the individual points, including the more expensive runs, in the interactive version.

Human active time against AI compute, one point per task

As can be seen in the graph, the compute required to perform a task is roughly linear in the time a human would take, with FLOP required per second of human time. There is substantial scatter, but aside from a cluster of very cheap tasks there is no obvious trend in the FLOPs per second required to complete a task.

Distribution of compute per human-second, stacked by task category

The median task takes FLOP per second of human time, and 90% of tasks lie between and FLOP per second of human time. The most expensive tasks, at around FLOP per second, are the Navier–Stokes proof, a Pokémon Crystal play-through, and an ARC-AGI-3 environment.

Costing FLOPs

A FLOP is a single floating-point operation, one addition or multiplication of two numbers. For a modern H100:

H100 Source
Peak rate, 8-bit (FLOP/s) NVIDIA
Rental price (USD per GPU-hour) 3 shattered.io (2026)
Typical utilization (%) 10 DeepSeek (2025)
Cost per delivered FLOP (USD) Derived from the rows above

So for the median task, at FLOP per second of human time, the GPU cost of replicating an hour of human work is about 4 cents. Even for expensive tasks, at FLOP per second, it is about $15 an hour. These costs are considerably lower than the US median wage of $25 per hour, and far lower than the fully loaded $100 to $300 per hour cost of skilled white-collar workers, like research scientists or software engineers.

While obviously there is a lot of uncertainty in these numbers, the gap is so large for nearly every task that we can be confident in the conclusion: when AI systems can do a human task, they almost always are far cheaper than humans. Put the other way round, at hardware prices a worker's wage buys around FLOP per second, so an AI system would have to spend more than that per second of human time before it cost more than the human.

Current prices for language-model inference are considerably higher than this. On the tasks in the dataset where we know what the AI run cost, the price per FLOP comes out more than ten times the hardware cost in the table. Frontier companies have to cover their training and research costs, and they also take generous margins. In the long run, however, the marginal cost of serving a model is set by the cost per FLOP, as training and research costs are amortized over ever more use.

Implications for AGI costs

I think we may well develop AGI—AI systems that can fully automate white-collar work—within the next few years. We are clearly not there yet, despite the periodic announcements to the contrary, but the tasks frontier systems can now do are broad and complex enough that development within the next few decades seems almost inevitable.

What would such systems cost to run? Current systems give a guide. Typical tasks in the dataset are performed at FLOP per second of human time. Real jobs, however, are a mixture of tasks, and what matters is the average over the mixture, weighted by the time spent on each. That average is set by the expensive tasks even when they are rare: a job that spends 1% of its time on tasks at FLOP per second costs at least FLOP per second overall.

The mean over the dataset is FLOP per second of human time, but the dataset is not weighted the way real human work is. A more accurate assessment, which I am not attempting here, would look at what jobs actually consist of and ask how cheaply each part can already be matched. The dataset does illustrate how much the expensive tasks matter, though: its mean is essentially driven by three tasks at FLOP per second; without these, the mean is FLOP per second. If I had to guess, I would put costs at between and FLOP per second of human time within a year of AGI.

This is, of course, speculative, because we do not really understand why current systems fall short of AGI. It is conceivable that what they lack is very compute-intensive, and that we are greatly underestimating the cost. For example, perhaps the expensive tasks are not yet in the dataset because AI cannot do them, or because they would be too expensive to run with current methods.

I think this is unlikely. AIs already perform a very broad range of tasks, and most jobs can be seen pretty straightforwardly as a sequence of such tasks, so the missing "secret sauce" would have to be far more compute-expensive than anything we have seen. Even the most ambitious tasks, like the Navier–Stokes proof, cost only around FLOP per second of human time, and that task is neither typical of white-collar work nor, I suspect, particularly well optimized.

If I had to make the case for AGI costs being considerably greater than current model costs, it would look something like continual learning. Current models have frozen weights, and methods that update them online could cost substantially more compute. Online learning would also cause problems for parallelization, since GPU economics rest on serving many users at once from a single copy of the weights. Nevertheless, I don't think this is that likely.

How many AGI workers?

If AGI runs at to FLOP per second of human time, AI workers are far cheaper than human workers and could be deployed in great number. The table below gives the cost of an hour of human work at present hardware costs, the number of H100s needed to replace a full-time worker, and how many workers the 27.6 million H100-equivalents (H100e) that Epoch estimates had been sold by June 2026 could run:

AGI compute rate (FLOP/s) Cost of human work (USD/hr) H100s per full-time worker Workers the June 2026 stock could run (millions)
0.014 0.0011 24,000
0.14 0.011 2,400
1.4 0.11 240
14 1.1 24

In Part 1 of my series on the AI industrial explosion I priced AI labor using the FLOP per second brain estimate, which is high compared with what the data here suggests. Even at that rate I found that automating everyone was cheap, and that the cost of the compute barely mattered for growth, since compute is a small part of the physical capital stock next to factories, mines, and utilities.

These are present-day costs and compute stocks. Setting AGI aside for the moment, we would expect the price of delivered compute to keep falling at roughly 25 to 30% a year, as it has since the mid-2000s, and the supply of compute to keep growing. The table below gives the history of both since 2020 and a naive extrapolation out to 2030, assuming that cost and installed stock continue their historical trends. (I suspect these trends are somewhat optimistic.)

Year (end) Basis Cost per delivered FLOP (USD) Installed stock (millions of H100e) At FLOP/s: Cost of human work (USD/hr) At FLOP/s: Workers the stock could run (millions)
2020 actual — 9.3 —
2022 actual 0.3 5.0 2.7
2024 actual 6.7 2.7 59
2025 actual 20 2.0 180
2026 projected 46 1.4 410
2028 projected 140 0.81 1,200
2030 projected 290 0.46 2,500

The arrival of AGI would change this picture radically. In the short run, demand for compute would leap as soon as systems could do a worker's job. Whether that produces a compute crunch depends on when AGI arrives and how much compute it needs.

If AGI running at FLOP per second arrived at the end of this year, the world's chips could run about 40 million workers, far fewer than the hundreds of millions of remote-capable jobs, and the price of compute would rise from today's few dollars an hour towards the wages of the workers it was replacing. If instead AGI arrives around 2030 at or FLOP per second, the stock by then could run more workers than there are white-collar jobs in the world, and we would already be past the crunch. Even then the value of compute would almost certainly sit well above the naive extrapolation, as there would be a great deal of unmet demand for labor at these prices.

In the long run, the compute stock will depend on the demand for compute, and it can grow as fast as fabs, packaging, power, and ultimately the whole economy can be built out. Much of the demand for AI labor will in turn require building robots and other physical capital for the AI to work with. I discuss this at length in my series on the AI industrial explosion.

How much cheaper could AI get?

The price of getting a fixed task done by AI has fallen greatly over the past few years. Gundlach et al. (2025) priced running whole benchmarks at a fixed score, using the tokens each model actually consumed, and found the cost falling 5 to 10 times a year for frontier models over 2024 and 2025. For open-weight models the fall is about 3 times a year, almost all of it from better models rather than cheaper hardware.

I don't see a comparable decline in my dataset, though this may just reflect a lack of data. Most of the long-running agentic tasks only became possible in the last year or two, so few models have been compared on each, and the estimates are noisy.

Some of the dataset's points show that, for certain tasks, the computational floor can be far below what humans require. Chess moves, image classification, and simple arithmetic can be done at around or below FLOP per second of human time. These are specialized systems, though, and a general model doing the same task costs far more: a chess move from GPT-4.1 comes to around FLOP per second of human time. So even if most short human tasks can be done very cheaply by dedicated systems, that need not mean a low floor for work as a whole, whose cost is dominated by the most expensive tasks.

The other evidence people have pointed to is the human brain. Carlsmith (2020) counts the brain's spikes and neuron updates, prices each as a floating-point operation, and gets to FLOP per second for the whole brain. When Carlsmith wrote in 2020, systems that could do the tasks in the dataset did not exist. Now that AI matches human performance at and below the lower end of his range across a broad range of tasks, the dataset is stronger evidence about how cheaply AGI will run than any analogy from the brain.

The brain also provides some weak evidence about the floor. Evolution presumably economizes on operations, so one might expect the brain to be reasonably close to the limits. But the median task in the dataset already runs at FLOP per second of human time, below Carlsmith's whole range, so the brain is evidently not the floor. Evolution works under strong constraints from an inherited architecture, and it would not shock me if the brain were several orders of magnitude less efficient than what deliberate engineering can achieve.

Training compute

I have also collected cases where an AI system has been trained to learn a task, and compared that against the time humans take to learn the same task. These comparisons are obviously even more difficult to make, because AI training and human learning look very different, and it is difficult to make consistent judgments about start and end points. Nevertheless, I think it is still helpful to get an order-of-magnitude sense of how costly it is for AI systems to learn things that take humans a lot of time.

AI training compute against the time a human takes to learn the same task

The dataset here is obviously small, and very much driven by the cases where at least some vague comparison was possible. Nevertheless, I think it is interesting that on these tasks AI systems use on the order of FLOP per second of human learning time to mimic human learning of various skills.

Compared to inference, this is a few hundred times less efficient, which suggests that relative to a human baseline we are substantially worse at training than at inference. Measured against the brain rather than against inference, though, training is not wildly expensive. Carlsmith's estimate for the human brain is around FLOP per second, so a model learning a skill uses about as much compute as the brain does over the time the human takes to learn it.

Whole training runs give another comparison. Qwen3-235B-A22B, an open model trained in 2025, used about FLOP to train. In Appendix B I attempt to estimate what it learned in terms of human learning time. Matching its benchmark scores to human curricula, subject by subject and language by language, gives on the order of 100,000 hours of study for a human. If we instead consider the Wikipedia facts it can recall, on the order of 60 million, these would take a person on the order of 3 million hours to memorize. These suggest the model learns facts and skills at a rate of around FLOP per second of human learning time. The comparison is tricky: the model has many abilities that are hard to elicit or measure, and unlike a human it is purely text-based. Its knowledge is far beyond any human's, despite its otherwise spiky and mixed abilities.

On these metrics, then, training an AI costs about as much as the equivalent human wages, unlike inference, which is far cheaper. The largest frontier runs are on the order of FLOP, at least twenty times Qwen3's, and hence cost on the order of thirty careers' worth of wages, but those costs are amortized over the model's existence and so are quickly becoming insignificant.

Despite this, there are some important disanalogies. Data efficiency is much poorer than a human's, and this is a practical barrier to deploying current systems. Parallelization also makes it more expensive in practice to serve models with specialized weights. Humans are an existence proof that far more data-efficient learning is possible, so better algorithms are likely to come. But feeding a model enormous amounts of data is easy, so massive models trained on everything may remain the cheapest path regardless.

Appendix A: Methodology

To build the dataset I had GPT-6 Astra and Fable 5.1 search for cases where both quantities could reasonably be estimated: benchmarks and studies that record how long humans took and how many tokens the AI used, and individual feats such as the Navier–Stokes proof. Each row has a research note giving the derivation, and I set the rules and ruled on the cases the agents raised. The largest sources are Terminal-Bench and Terminal-Bench-Science, METR's time-horizon suite, SWE-Marathon, Epoch AI's benchmark runs, and GDPval. In all there are 190 sources, most contributing a single task.

The dataset only provides rough order-of-magnitude estimates, but since the quantities of interest span many orders of magnitude the results are nevertheless informative about AI costs. If someone wants to create a better dataset or collect their own data, I'd love to see it! I'm particularly interested in better estimates for compute requirements, since these are rarely published despite being more fundamental than either token usage or price.

Compute. Almost no source reports the compute required to achieve a result. What benchmarks and papers report, at best, is how many tokens the model used or what the run cost, and often neither. So compute is reconstructed: tokens, taken from the source, worked back from its cost at list price, or estimated from the task, times FLOP per token, taken as twice the model's active parameter count, plus attention over the context. Frontier companies do not disclose parameter counts, so for their models the count is a guess from pricing, speed, and model family, with a range of around a factor of three either way. The guess is shared by every row of that model, so the error is systematic per model. Open models and systems that are not language models publish their architectures, so their figures are firmer.

Training compute is six FLOP per parameter per training token from published sizes, or the paper's own figure where one was reported.

Of the 1,684 compute figures, 12 rest on measured operation counts, 54 on documented inputs, and 1,618 on assumed inputs from reported costs, token usage, or both.

Human time. Where a study timed people on the task, I use those timings. Otherwise the figure is an estimate, either the source's own or one made by the LLM agents, from related data where there was any and by judgment where there was not.

How the human time was obtained Rows
Timings recorded on the task 473
The source's own estimate 744
LLM estimate from related data 132
LLM judgment 232
Fixed by the task's definition 103

Performance. Each run is labeled below, matching, or above the human baseline. Matching means the reported performance is broadly comparable, or that the human time was estimated for reproducing the AI's output at its quality. Below means the AI basically did the job, but noticeably worse than the human. Runs where the AI fell well short of that are left out of the dataset altogether, since a catastrophic failure says little about how hard a task is. The dividing line was a judgment, guided by whether the AI reached about half the human's score on the benchmark's own metric after subtracting the chance floor.

Many tasks appear under several models at very different costs. Throughout this post I use the cheapest run that reached or beat the human baseline, which the dataset flags on each task.

The dataset is larger than the analysis uses, because most tasks carry runs from several models:

What is counted Count
Data points 1,684
Of these, at least matching human performance 1,111
Unique tasks 484
Of these, with at least one matching point 395
Of these, training runs 45
Unique models 275

Appendix B: how much learning does a trained model embody?

It is hard to get a sense of how much a modern language model knows, but it is vast compared to a person. As a simple estimate, consider Qwen3-235B-A22B. It has 235 billion parameters, 22 billion of them active per token, and was trained on 36 trillion tokens covering 119 languages. The technical report gives no overall FLOP estimate for training. Using Epoch's 6 FLOP per active parameter per token gives a total of FLOP for pretraining. Post-training and reinforcement learning are small costs by comparison; including these I would estimate a total spend of FLOP.

Like other modern language models, Qwen3 has knowledge across a very broad range of areas. Matching its benchmark scores to human curricula, I estimate this represents on the order of 100,000 hours of study. Comparing instead to the facts it recalls from Wikipedia gives even more, something like 2 to 6 million hours of human time.

While these estimates are extremely rough, they suggest the model used on the order of FLOP per second of human learning time to acquire its knowledge. Compared to Carlsmith's estimate for the brain, current training methods do not look broadly dissimilar in efficiency to human learning; language models get their breadth from being trained for far longer than a human lives.

Domains of knowledge

One good approach is to look at all of the academic disciplines and languages that Qwen3 is able to perform well at on tests, and compare that to how long it would take a human to acquire a similar degree of proficiency in the subject. Scores are Qwen3's thinking-mode results, and the hours for degrees come from the UK credit framework:

Ability Human level matched Evidence Hours Basis for the hours
English Native, educated adult — 20,000 Pre-school immersion
Schooling to end of high school Completed MMLU-Pro, 86% 10,600 OECD instruction time
53 languages Graduate-level academic work in the language MMLU-ProX, 94% of English 94,800 FSI class hours
23 further languages Partial, reading ahead of writing Belebele, 83 to 91% 16,000 400 class hours each
Physics, chemistry, and biology Doctorate GPQA, 62.3% 27,000 Degree plus doctoral work in three fields
Medicine Qualified physician MMLU-Pro, 88% 8,000 A medical degree
Mathematics Strong undergraduate, not research MATH-500, 98% 3,600 A three-year degree
11 programming languages beyond Python Working competence Multi-LCB, 62.9% 4,400 400 hours each
Eight further subjects Bachelor's degree MMLU-Pro 28,800 3,600 hours a field
Total 213,200

The total, 213,200 hours, is 24.3 years around the clock, or about 107 working years at standard full-time work. Dividing FLOP by it gives FLOP per second of human learning time.

Obviously, this approach should be taken with substantial salt. Benchmark scores may overestimate how much Qwen3 knows of the relevant subjects:

  • Qwen3 may match human performance on these benchmarks while only having acquired a subset of the abilities and skills a human would have. For example, while it can match a physics PhD on GPQA questions, it does not necessarily have the skills to run experiments or perform novel research.
  • In particular, domain knowledge includes visual and auditory knowledge that Qwen3 completely lacks.
  • Learning related languages (e.g., Spanish and Italian) and overlapping disciplines (physics and mathematics) takes less time than studying each independently, because knowledge transfers between subjects.
  • It's possible Qwen3 was trained on some or all of these benchmarks, which could lead to a substantial overstatement of abilities.

Taking these issues into account, I would guess a more conservative estimate is something like 80,000 hours of human effort to match Qwen3 across these domains, or FLOP per second of human learning time.

Facts

Benchmarks miss most of what a model knows: the miscellaneous facts that no curriculum covers. To estimate that knowledge I treated English Wikipedia as the universe of facts and measured what fraction of it Qwen3 can recall closed-book.

A fact is one standalone question with a single answer, written from one sentence or table row of an article. I sampled articles in ten strata, the five levels of Wikipedia's Vital Articles lists and the remaining seven million articles split into five bands by length, and from each article drew four to six blocks, a block being one paragraph or table row. A language model turned each block into questions, a second pass filtered out questions that were definitional, inferable from their own wording, open-ended, or time-relative, and I read every surviving question by hand. Qwen3 answered them closed-book with reasoning off, a model graded the answers against the stored reference, and I answered the same questions myself.

Results were as follows:

Stratum Articles Sampled Blocks Blocks drawn Questions Qwen3 correct Me correct
Vital level 1 10 10 876 40 69 41 13
Vital level 2 90 11 926 44 46 34 13
Vital level 3 900 11 1,173 44 55 30 8.7
Vital level 4 9,965 11 1,133 44 117 57 2.5
Vital level 5 34,227 11 582 44 63 20 1
Other, under 2 KB 1,419,119 15 149 50 60 12 0
Other, 2 to 5 KB 2,309,149 20 272 79 110 23 1
Other, 5 to 15 KB 2,435,261 25 695 100 155 36 0
Other, 15 to 50 KB 852,341 25 1,592 148 268 39 0
Other, over 50 KB 179,745 15 2,805 86 165 45 1
Total 7,240,807 154 10,203 679 1,108 337 40

Given this, we can estimate the total number of facts in Wikipedia, the facts known by Qwen3, and the facts known by me. For each sampled article, facts equal its block count times the surviving questions per drawn block, and facts held replace surviving questions with correct answers. Each stratum's mean over its sampled articles, times its article count, gives the stratum total. The table shows of each total with one standard error from the sampling of articles. My recall on the random articles is too small to estimate band by band, two correct in 758, so for me the five bands are pooled.

Stratum Facts in stratum () Facts Qwen3 knows () Facts I know ()
Vital level 1 3.21 ± 0.10 2.98 ± 0.14 2.44 ± 0.17
Vital level 2 3.89 ± 0.09 3.75 ± 0.12 3.27 ± 0.19
Vital level 3 5.16 ± 0.22 4.93 ± 0.28 4.25 ± 0.23
Vital level 4 6.40 ± 0.15 6.09 ± 0.17 4.42 ± 0.26
Vital level 5 6.37 ± 0.16 5.88 ± 0.13 4.68 ± 0.43
Other, under 2 KB 7.20 ± 0.19 6.57 ± 0.31 —
Other, 2 to 5 KB 7.58 ± 0.09 6.75 ± 0.18 —
Other, 5 to 15 KB 8.00 ± 0.06 7.36 ± 0.10 —
Other, 15 to 50 KB 7.93 ± 0.07 7.03 ± 0.08 —
Other, over 50 KB 7.74 ± 0.15 7.27 ± 0.24 —
Other, all five bands 8.47 ± 0.04 7.79 ± 0.08 5.89 ± 0.31
Total 8.48 ± 0.04 7.80 ± 0.08 5.94 ± 0.28

It looks like English Wikipedia holds about 300 million facts by this standard, of which Qwen3 recalls about 60 million, or one fifth. By contrast, I know about 900,000 of these facts, so Qwen3 knows about 70 times as many facts as I do.

(Some facts are duplicated across articles. Having looked at the questions, my sense is that this is not very common, because most facts are very obscure. However, the most common facts in the most popular articles are probably the most likely to be repeated, and these are also the ones I am most likely to know. This suggests the measure probably understates how much more Qwen3 knows than me. I would guess this effect is not more than a factor of two overall, but I have not tried to correct for it.)

We can ballpark how much human learning this represents in two ways. The first is to compare to myself. I just turned 29, have spent most of my life in school or working as a researcher, and spend a lot of my free time on nonfiction of some form or other. If we estimate that I have spent about a quarter of my waking life in somewhat active study, that equates to around 40,000 hours, which implies I've acquired around 20 facts per hour of study. That implies Qwen3's 60 million facts are equivalent to about 3 million human hours, for an estimate of FLOP per second of human time.

Another option is to compare to how many facts humans learn per hour of study. No study reports this rate directly, since the memory literature reports retention rather than throughput, but it can be worked out from the studies that tested recall long after learning. Cepeda et al. (2008) taught 1,354 people obscure trivia facts and tested them closed-book up to a year later; counting all study time, about 13 facts an hour were still recalled. Bahrick et al. (1993) taught foreign vocabulary over several years and found 7 to 10 words an hour still recalled one to five years on. Spaced-repetition users do better: Woźniak (1990) reports about 33 items an hour kept recallable indefinitely, counting every review, and Gwern's review of Anki data gives 12 to 51 an hour. Taking 10 to 30 facts an hour, Qwen3's 60 million facts are 2 to 6 million hours of human study, comparable to my own estimated rate of 20 an hour. This gives an estimate on the order of to FLOP per second of human learning time.

  1. About 11 hours a day for five years.
  2. The high-school-labeled questions across all 14 fields, from HELM.
  3. MMLU-ProX rates 24 of the 53 individually, at 86 to 100% of its English score. The Qwen3 report covers all of them, together with English and a second Chinese script, as a pooled average over 55 languages on a translated AIME paper, at 94% of its English score, so at most five of them can be doing badly.
  4. To professional working proficiency. The 12 languages FSI does not schedule are taken at the average of the 41 it does, plus FSI's self-study allowance.
  5. Reading comprehension across the ten language families, reported in the Qwen3 report only for Qwen3-32B, a smaller sibling trained on the same data, which the 235B beats wherever both are measured, so this is a floor.
  6. About a third of the way to professional proficiency.
  7. Against 65% for people holding or pursuing a PhD in the matching field, and 34% for skilled non-experts with unrestricted web access, from HELM.
  8. The professional medicine questions, drawn from licensing material.
  9. Where a three-time IMO gold medalist scored 90%. Also 54.8% on Omni-MATH olympiad problems, and 8.5% on FrontierMath tiers 1 to 3, advanced undergraduate to early-career research, against 0% on its research tier.
  10. Over twelve languages, from 77.5% in Python down to 40.0% in Scala; the eleven non-Python languages at about 80% of its Python level.
  11. College- and professional-labeled questions at close to its rate on the high-school ones, eight of eight on college computer science and four of four on machine learning. The subjects are computer science, business, economics, engineering, psychology, philosophy, history, and the benchmark's residual category.
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论