A Four-Axis Bayesian Epoch Capabilities Index with Human Baselines

This is a crosspost from the General-Purpose AI Policy Lab research blog.

The Epoch Capabilities Index compresses many benchmark scores into one for each model, following the framework of the Rosetta Stone paper. In a previous post, we added human baselines to the same scale to see how the models compare to humans. But one issue is that most humans score near-perfectly on abstract reasoning benchmarks like ARC-AGI or VPCT and sit near chance on GPQA-type benchmarks, while many models show the opposite pattern. One index cannot produce both of these orderings, so in this post we move to a Bayesian setup with four skill axes instead of one, and proper uncertainty estimation.

TL;DR

  • We extended the Epoch Capabilities Index to four skill axes, in a Bayesian setup where every ability comes with its uncertainty and the human tiers are fitted inside the model as test-takers.
  • Sparse data (6% of the test-taker-by-benchmark matrix) means the model does not settle on one answer by itself. We need two ordering priors to help identify it, a hard ordering on the human tiers, and a soft expected improvement along releases of the same model family.
  • Main Results:
    • The four axes that come out of the model are Fluid Intelligence, Scientific Knowledge and Reasoning, Agentic Capabilities, and Legacy QA.
    • Models have now passed every human tier on Scientific Knowledge and Reasoning, though this arguably says more about breadth of recall than about doing actual science.
    • On Agentic Capabilities, the frontier is reaching the Skilled Generalist baseline around now (summer 2026).
    • On Fluid Intelligence, AI models have likely passed Average Humans, but human experts still lead. The trend, if it stays linear, will cross top tiers in 2027 and 2028.
    • Four of the ten runs converge on a second mode that places human tiers slightly differently and move these dates by a few months to years.

One-Dimensional Setup

Following Alexander Barry's Bayesian version of the Epoch Capabilities Index (ECI), we rebuilt the model in Python using PyMC. In this Bayesian setup, instead of finding one single best value for each parameter, we sample a whole distribution of plausible ones, so every ability comes with uncertainty. The full setup is in the model section below. We also fit the nine human tiers as test-takers next to the models, from Average Human up to Committees of Domain Experts, plus two high-school tiers. The baselines and the benchmark table have been updated and expanded since the previous post (the full tables with sources are in the appendix), and we added a partial ordering prior on human tiers (see following sections). For the 1D fit only, we excluded the human-easy benchmarks like ARC-AGI or VPCT (in a similar fashion to the previous post; the list is in the Appendix).

Here's our rebuilt index (called ECI-H with H for the human baselines) matched with the ECI scores for the state-of-the-art models:

Our values differ from Epoch's for three main reasons. First, we do not fit the same table, since our benchmark set is larger and drops eight human-easy benchmarks. Second, Epoch publishes one value per model, taking its best score on each benchmark, while we fit every thinking-effort variant as its own test-taker with its own scores. Third, we include human baselines.

Each benchmark also gets a difficulty on the same scale, so we can plot models, benchmarks and the human tiers together over time:

However, putting AIs and humans on a single axis is arguably quite objectionable.

Multidimensional extension

As we mentioned earlier and discussed in the previous post, some benchmarks are trivial for humans and hard for AI models, which breaks the single difficulty axis. Epoch's Benchmark Scores = General Capability + Claudiness also points to scores carrying more than one dimension (and that's between models alone). We test this by extending the model to four skill axes using an MIRT (Multidimensional Item Response Theory) model, commonly used in psychometrics, while keeping all the benchmarks and human baselines.

The intuition behind the model is that each test-taker has four abilities , that form its skill profile, the way a student can be strong in algebra and weak in essay writing. Each benchmark weighs those skills through its four positive loadings , one per skill, which say how much each skill counts for that benchmark. The loadings also set how a benchmark separates its test-takers, what psychometrics literature calls discrimination. So a benchmark with a large loading splits weak models from strong ones clearly, while one with small loadings doesn't react to skill. The difficulty is the bar the weighted skills must clear to get more than the midpoint on the benchmark, and the S-curve turns the result into a score between 0 and 1. We also need to take into account the random-guessing for each benchmark so we fix each benchmark's guessing floor in advance and start the curve there instead of at 0, so scores on a four-option exam bottom out at 0.25 rather than 0.

As such the expected score for test-taker and benchmark is

and the observed score scatters around it with Beta noise (following Barry's post),

Each benchmark gets its own noise level .

The form we use above to define what's inside the sigmoid belongs to one of three families common in the IRT literature and it is called compensatory because a strong skill can make up for a weak one inside the sum. In the non-compensatory family a benchmark needs all its skills at once, and the sum becomes a product of per-axis curves, . The semi-compensatory family sits in between and adds interaction terms to the compensatory sum. We tried both alternatives and the non-compensatory fit did not converge, while the semi-compensatory one converged only under heavy constraints and made worse predictions.

Prior Assumptions

We first tried to fit this model with no other assumptions than the ones explained above, but the model did not settle on one answer. This is due to our data being very sparse (the test-taker by benchmark matrix is filled only at 6%, and the average test-taker has about six scores) and the fact that many arrangements of abilities and loadings explain the scores equally well, so repeated runs fall on different solutions. Given this, we needed to put more prior information into the model to help it converge to one answer.

Human Ordering (hard prior)

In the data, non-skilled humans are mostly tested on human-easy benchmarks and experts are tested mostly on hard benchmarks. Yet, we know that average humans would do worse than experts on the hard benchmarks, and that experts would do at least as well as average humans on easy benchmarks. So we gave the model a prior ordering where a Domain Expert is at least as good as a Skilled Generalist, and a committee is at least as good as one of its members, on every skill. The ordering says nothing about the size of the gap between the tiers (where no ranking is obvious, like between a Top Performer and a committee of experts, we don't impose any ordering). The two high-school tiers join the ordering by a Domain Expert being at least as good as a High School Qualifier and a Top Performer at least as good as a High School Top Performer.

Model families (soft prior)

Recent models sometimes lack data to estimate their ability scores, but within one release chain, like the GPT flagships and the Claude Opus line, we can expect each new release to improve on the one before it. A release can regress if the data says so; we only nudge it towards improving. We also use time between releases for the difference in abilities, so the expected gain grows with the gap between releases, and a lab shipping many small updates is not expected to gain more than one shipping a single big release over the same year. Regarding the thinking-effort variants, they are tied to their base release, and we do not order them among themselves, as a higher effort can potentially overthink.


To see how the priors help, we can compare the runs. First without any prior, they split into two different sets of axes. After adding the human ordering alone we still get two solutions that disagree about the axes. Adding the model family assumption finally makes all the runs agree on one set of axes. We also tried three axes with all the assumptions on, and the runs still split in two. These assumptions also help improve the model's predictive ability since, on left-out scores (leave-one-out cross-validation) the final model beats the no-prior version by about 107 ± 18 and the 1D index by about 1,000 ± 33, comparing on the rows where the comparison is reliable.

Results

With all of the assumptions above in place, the model settles on a single answer for most of the runs. Let's look at what it found.

The Axes

We name each axis after the benchmarks whose loading vectors are most collinear with it, i.e. the benchmarks that draw on that skill and almost nothing else. We get:

  • Axis 1: We call it Fluid Intelligence, defined by ARC-AGI-2, ARC-AGI and VPCT, abstract puzzle benchmarks.
  • Axis 2: We call it Scientific Knowledge and Reasoning, defined by WMDP Chemistry and Biology, the GPQA science subsets and FrontierMath.
  • Axis 3: We call it Agentic Capabilities, defined by GBAEval, the Remote Labor Index and SWE-Bench Pro, benchmarks where a model works through long tasks rather than answering questions.
  • Axis 4: We call it Legacy QA, consisting mainly of older question-answering benchmarks, largely saturated, OpenBookQA, ARC (AI2), BoolQ and similar benchmarks.

And here's how the top models compare on each axis to humans:

The frontier models sit above every human tier on Scientific Knowledge and Reasoning. On Fluid Intelligence the opposite holds as every tier (except Average Human) sits above the best models. On Legacy QA the humans also sit on top, but this is more of a data artifact since the eight benchmarks that define this axis most purely (axis share above one half) were last run on models from mid-2024 or earlier, with most of them already scoring around 0.9 there, and no frontier model was ever measured on them. So the human lead on this axis is a comparison against a frozen pool of older models, and remains untested against the actual frontier.

Forecasting

For these forecasts, only models whose ability on the axis is well estimated enter the pool (posterior SD below 0.33, plus flagged frontier releases). We take the record-setting frontier models for each plausible set of abilities the model produced, then fit a per-draw record envelope (running max of the frontier in each posterior draw), and extend it at its rate over a 1.5-year window to see when it reaches each human tier . And as mentioned above, given how the recent models' positions on the Legacy QA axis come from the prior rather than actual ability, we leave this axis out of the forecasts. The headline numbers below are for the majority mode (6 of the 10 runs, see the Appendix for the minority mode).

On the Fluid Intelligence axis, the frontier likely sits above Average Humans, at the Skilled Generalist level, but still below expert baselines. It is on track to reach the Domain Expert by spring 2027 and pass the Top Performer by mid-2028.

On the Agentic Capabilities axis, models have already passed some of the lower tiers, with the frontier reaching the Skilled Generalist baseline right about now (we're writing this during the last days of August 2026), and projected to pass the Committee of Domain Experts by mid-2027.

On the Scientific Knowledge and Reasoning axis, models have passed every human tier already, with probabilities of 0.85 to 0.95 (since on WMDP Chemistry the Domain Expert baseline is 0.433 against a best model score of 0.809, on WMDP Biology 0.605 against 0.875, and on GPQA Diamond 0.812 against 0.946). In fact every model on this axis since 2023 sits above the human baselines, which says more about what these benchmarks reward, that is breadth of recall across a whole field, than about doing science and research (the Skilled Generalist is below chance on GPQA Diamond, 0.22, and even the in-domain PhD gets 0.43 on WMDP Chemistry).

The main assumption here is that the forecasted trend based on the recent rate keeps its slope, which seems reasonable since within our window the frontier shows no sign of decelerating, and on the Agentic Capabilities axis it even seems to accelerate with the latest releases. These dates should be read with the uncertainty the model gives them (at 95% the crossing windows stretch by several more years, and far longer on the Agentic Capabilities axis), and the best way to tighten them would be better human baselines, especially on the Agentic Capabilities axis, where the human tiers rest on a handful of measurements.

Limitations

The first limitation is the data itself. As explained in the previous post, we need more of it and data of better quality, especially for the human baselines, but also for the models, since we fill only 6% of the test-taker by benchmark matrix and this sparsity forced us to add more assumptions to the model.

The second one, which we share with the Rosetta Stone paper and the ECI in general, is that we fit benchmark-level scores instead of item-level answers, so we do not fully respect the assumptions of the MIRT setup.

The last one is the calibration, since our predictive intervals are wider than the data requires, making our model more conservative than it should be.

Further Work

Beyond collecting more data, a few directions are worth exploring:

  • Running the legacy benchmarks on current frontier models. This would help settle the Legacy QA axis with data instead of it being an artifact, and it fills more of the grid at the same time.
  • Ceilings on saturating benchmarks. We already fix a guessing floor per benchmark, so a ceiling either inferred from the data or fixed seems reasonable, and would help saturation be read as such and not as extreme difficulty.
  • Testing the forecasts in both directions. Fitting the model on an older snapshot of the data and checking its forecasts against the releases that came out since, and going forward, keeping track of the predictions we made here and see how they pan out.

Code

The code is available at github.com/General-Purpose-AI-Policy-Lab/Multiaxis_ECI/tree/blogpost-frozen.

Appendix

Appendix A

Below is the full graphical specification of the model:

Appendix B

This appendix records the previous fits and how each assumption affects the model. For the 1D Setup, we use all the benchmarks to compare it with our final fit.

Fit

Runs x Draws

Divergences

Posterior modes (axis systems)

elpd (LOO) ± se

Δ vs final

1D index

10 × 10,000

0

1

6,249.6 ± 112.1

−999.6 ± 33.2

4 axes, no priors

12 × 3,000

778

2

7,583.9 ± 75.9

−107.2 ± 18.3

4 axes, human ordering only

12 × 3,000

80

2

7,583.3 ± 73.1

−109.2 ± 17.2

4 axes, both priors

10 × 12,000

37

1

7,710.4 ± 76.3

0

3 axes, both priors

12 × 3,000

12

2

7,447.8 ± 77.4

−191.1 ± 14.3

Appendix C

Four of the ten runs place the human tiers differently. This section records what moves and how the forecasts change.

  • Minority minus majority, averaged over the nine tiers: −2.71 on Agentic, +1.58 on Legacy QA, +0.71 on Fluid Intelligence, −0.04 on Scientific Knowledge and Reasoning.
  • Both groups share one axis system and describe the scores equally well.

Here are the forecasts for the minority chains:

On these minority forecasts, for the Agentic axis, we will soon have surpassed all of the human tiers. For Fluid Intelligence, frontier models are still below the average human contrary to what the majority chains predict.

Appendix D

Here below are the crossover dates with the 95% intervals for both the majority and minority chains:

Majority chains: for all axes, the 95% timespans are much wider, but still close by 2035 at most.

Minority chains: the Fluid Intelligence axis keeps its 95% intervals relatively tight, but otherwise uncertainties are huge.

Appendix E

For each observed score we compute where it lands inside the model's predictive distribution (called the probability integral transform, PIT). A perfectly calibrated model would spread these values evenly but our histogram bulges in the middle instead meaning that the observed scores land near the center of the predictive intervals more often than they should, so the intervals are wider than the data requires and make our model pretty conservative.

Appendix F

Down below are the human scores we used as well as the exhaustive list of benchmarks included in the setup.

Human Baselines

Benchmark

Human Group

Score

Source

Information

ARC-AGI

Average Human

0.77

ARC Prize

MTurk

ARC-AGI-2

Average Human

0.6

ARC Prize

average test-taker

BIG-Bench Hard (BBH)

Average Human

0.677

BIG-Bench Hard

average human raters

BoolQ

Average Human

0.9

BoolQ

Human annotators

CSQA2

Average Human

0.903

CSQA2

average accuracy of humans

MMLU

Average Human

0.345

MMLU

MTurk

OpenBookQA

Average Human

0.92

OpenBookQA

probability from random human subjects

ScienceQA

Average Human

0.884

ScienceQA

MTurk workers with a high school degree or higher who passed the qualification examples

SimpleBench

Average Human

0.837

SimpleBench

nine non-specialized humans

SuperGLUE

Average Human

0.898

SuperGLUE

human performance estimates after training phase

TriviaQA

Average Human

0.797

TriviaQA

human performance level

VPCT

Average Human

0.999

Epoch AI

three volunteers

ARC-AGI

Committee of Average Humans

0.98

ARC Prize

Human panel (at least two participants solved one or more sub-pairs within their first two attempts)

ARC-AGI-2

Committee of Average Humans

0.999

ARC Prize

Human panel

CSQA2

Committee of Average Humans

0.941

CSQA2

majority vote

HellaSwag

Committee of Average Humans

0.956

HellaSwag

majority vote of 5 crowd workers (MTurk)

WinoGrande

Committee of Average Humans

0.94

WinoGrande

majority vote of crowd workers (MTurk)

ARC-AGI

Skilled Generalist

0.98

ARC Prize

STEM Graduates

GPQA Diamond

Skilled Generalist

0.219

GPQA Diamond

highly skilled and incentivized non-experts who have or are pursuing PhDs in other domains

GPQA Diamond Biology

Skilled Generalist

0.22

GPQA Diamond Biology

not in-domain PhD

GPQA Diamond Chemistry

Skilled Generalist

0.22

GPQA Diamond Chemistry

not in-domain PhD

GPQA Main Biology

Skilled Generalist

0.43

GPQA Main Biology

not in-domain PhD

GPQA Main Chemistry

Skilled Generalist

0.31

GPQA Main Chemistry

not in-domain PhD

GSM8K

Skilled Generalist

0.9677

GSM8K

qualified human annotators who have passed a qualification exam with at least a bachelor's degree

MATH Level 5

Skilled Generalist

0.4

MATH Level 5

a computer science PhD student who does not especially like mathematics

OS World (Screenshot)

Skilled Generalist

0.724

OS World

individuals not familiar with the software

SimpleQA Verified

Skilled Generalist

0.944

SimpleQA

human annotator going through the test

Visual Task Assessment (VISTA)

Skilled Generalist

0.554

Scale SEAL

16 full-time employees

PIQA

Committee of Skilled Generalists

0.949

PIQA

majority vote of top annotators

OTIS Mock AIME 2024-2025

High School Qualifier

0.53

OTIS Mock AIME

average score by high school students from the OTIS program (percentage from number of questions answered)

OTIS Mock AIME 2024-2025

High School Top Performer

0.93

OTIS Mock AIME

top scorer from the OTIS program (percentage from number of questions answered)

BioLP-bench

Domain Expert

0.384

RAND report

Bachelor's w/ lab experience

GPQA Diamond

Domain Expert

0.812

GPQA Diamond

in-domain PhD validators, GPQA paper Table 2

GPQA Diamond Biology

Domain Expert

0.831

RAND report

In-domain PhD

GPQA Diamond Chemistry

Domain Expert

0.831

RAND report

In-domain PhD

GPQA Main Biology

Domain Expert

0.667

RAND report

In-domain PhD

GPQA Main Chemistry

Domain Expert

0.72

RAND report

In-domain PhD

LAB-Bench Cloning

Domain Expert

0.6

RAND report

In-domain PhD

LAB-Bench LitQA2

Domain Expert

0.7

RAND report

In-domain PhD

LAB-Bench Protocol

Domain Expert

0.79

RAND report

In-domain PhD

LAB-Bench SeqQA

Domain Expert

0.78

RAND report

In-domain PhD

MATH Level 5

Domain Expert

0.9

MATH Level 5

a three-time IMO gold medalist university student

MMLU

High School Top Performer

0.898

MMLU

estimation from the authors based the 95th percentile of student results

MMLU Biology

High School Top Performer

0.9

RAND report

idem

MMLU Chemistry

High School Top Performer

0.9

RAND report

idem

WMDP Biology

Domain Expert

0.605

RAND report

In-domain PhD (RAND)

WMDP Chemistry

Domain Expert

0.433

RAND report

In-domain PhD (RAND)

FrontierMath

Committee of Domain Experts

0.35

Epoch AI

solved collectively across all teams (40 exceptional math undergraduates and subject-matter experts) in four and a half hours and with internet access

PRBench Finance

Committee of Domain Experts

0.796

PRBench Finance

agreement between human experts

PRBench Legal

Committee of Domain Experts

0.796

PRBench Legal

agreement between human experts

BIG-Bench Hard (BBH)

Top Performer

0.944

BIG-Bench Hard

max human raters

GeoBench

Top Performer

0.9

GeoBench

top player

List of benchmarks

  • APEX Agents
  • ARC (AI2)
  • ARC-AGI
  • ARC-AGI-2
  • Adversarial NLI
  • Aider Polyglot
  • AudioMultiChallenge
  • BALROG
  • BIG-Bench Hard (BBH)
  • BioLP-bench
  • BlueprintBench 2
  • BoolQ
  • CAD-Eval
  • CL-Bench
  • CL-Bench Life
  • CSQA2
  • Chess Puzzles
  • CritPt
  • CursorBench
  • Cybench
  • DeepResearchBench
  • DeepSWE
  • EBR-bench
  • EnigmaEval
  • ExploitBench
  • Fiction.LiveBench
  • ForecastBench
  • FrontierCode
  • FrontierMath
  • FrontierMath Tier 4
  • GBAEval
  • GDP.pdf
  • GDPval
  • GPQA Diamond
  • GPQA Diamond Biology
  • GPQA Diamond Chemistry
  • GPQA Main Biology
  • GPQA Main Chemistry
  • GSM8K
  • GSO-Bench
  • GeoBench
  • HellaSwag
  • Humanity's Last Exam
  • LAB-Bench Cloning
  • LAB-Bench LitQA2
  • LAB-Bench Protocol
  • LAB-Bench SeqQA
  • LAMBADA
  • Lech Mazur Writing
  • LiveBench
  • MATH Level 5
  • MCP Atlas
  • METR Time Horizons
  • MMLU
  • MMLU Biology
  • MMLU Chemistry
  • MMLU Pro Biology
  • MMLU Pro Chemistry
  • MMLU-Pro
  • MindCube
  • MultiChallenge
  • MultiNRC
  • Mystery Game Puzzles
  • OS World (Screenshot)
  • OS World 2
  • OTIS Mock AIME 2024-2025
  • OpenBookQA
  • PIQA
  • PRBench Finance
  • PRBench Legal
  • PostTrainBench
  • ProofBench
  • Remote Labor Index
  • SEAL Instruction Following
  • SEAL Tool Use (Enterprise)
  • SWE-Bench Pro
  • SWE-Bench Pro (Private)
  • SWE-Bench Verified
  • SciCode
  • ScienceQA
  • SimpleBench
  • SimpleQA Verified
  • SpatialViz-Bench
  • SuperGLUE
  • Surface Evolver Bench
  • TerminalBench
  • The Agent Company
  • TriviaQA
  • TutorBench
  • VPCT
  • Video-MME
  • Visual Task Assessment (VISTA)
  • VisualToolBench
  • WMDP Biology
  • WMDP Chemistry
  • WeirdML
  • WinoGrande

Benchmarks excluded from the 1D fit

  • ARC-AGI
  • ARC-AGI-2
  • HellaSwag
  • OpenBookQA
  • PIQA
  • SimpleBench
  • VPCT
  • WinoGrande
  1. The full mathematical specification is in the appendix.
  2. Details in the appendix.
  3. Details about the previous fits and the comparisons between them are in the appendix
  4. Four of the ten runs place the human tiers and some older models differently on the Agentic Capabilities axis, the frontier results are unchanged. Details in the appendix.
  5. The four runs that place the human tiers differently move these dates by a few months to years. Details in the appendix.
  6. The 95% crossover dates are in the appendix
  7. We discussed the quality problems of human baselines at length in the previous post, and they still affect the results here.
  8. More details about calibration in the appendix
  9. Paired LOO deltas use only the rows with Pareto-k below 0.7 in both fits.
  10. Divergences are small and only affect one parameter associated to the GSM8K benchmark.
  11. The final model's two groups share one axis system and differ only on the human tiers and 18 older or small models on the Agentic Capabilities axis (Appendix C), contrary to the other fits which don't agree on the axes.
  12. K=3 with all assumptions splits ten runs against two, trading two of the axes between the solutions.
添加评论
点赞收藏
点踩分享查看原文
评论
?
参与讨论