Are Manifold prediction markets well-calibrated?
I wanted to do a study of whether prediction markets are well-calibrated aka when the market predicts a P% chance of some event, do such events really happen P% of the time?
I asked Opus 5 to use Manifold's API to scrape all binary resolved markets (excluding N/A), probabilities over time and total trading volume in "mana". I used 3 different snapshots of probabilities: 3 days after the market opened, 50% midpoint between creation date and resolution date, and 3 days before the resolution. Two filters were used to filter out extremely low-quality data:
1) Removed markets with <7 days lifespan (resolution date - creation date)
2) Removed markets with <3 traders in total
This filtered out around a third of all markets though. And these aren't even particularly strict filters.
Below are the calibration graphs for all three probability snapshots (3 days after creation, midpoint, 3 days before resolution):
Early on, the probability mass is clustered around 50% rather than at the edges, and as the market gets closer to resolution, the distribution becomes U-shaped with probabilities concentrated near 0% and 100%; Thomas Bayes would love to see it. Interestingly though, when the market is close to resolution, more values fall within the [0, 0.1] range than the [0.9, 1.0] range. I guess it's because YES base rate across the whole Manifold is ~36%, not 50%.
Conclusion 1: early on, markets overestimate the probability.
On the first graph, the blue curve is entirely under the orange line, biased in one direction. So instead of Nothing Ever Happens, it's... Everything Always Happens? This bias goes away over time, but even at the midpoint markets are still leaning more towards Everything Always Happens. I thought everyone was in Nothing Ever Happens mode these days. Maybe this is a quirk of Manifold specifically, but I don't know what could be causing it. The average magnitude of the bias (predicted YES - actual YES) is around 8% early on, around 6% at the midpoint and shrinks to 1% as resolution approaches.
Practical takeaway: if you see a market where the resolution is far away, the probability is usually over- rather than underestimated.
Now let's look at how Brier score depends on the total trading volume:
Weirdly, for markets with the largest volume, accuracy of predictions is worse (early, midpoint) or at least not better (late) compared to markets with volume of ~10k mana. Opus 5 hypothesizes that these are elections, stuff like FIFA world cup, etc. where the irreducible uncertainty is high (spoiler: topic-level data doesn't support this). Or maybe those high volume markets have the most normies who aren't good at forecasting, but I doubt normies know about Manifold in the first place.
Opus 5 also suggested the following "Brier skill score" (which is used in weather forecasting):
1 − Brier/(base_rate * (1 − base_rate))
1 = perfect score, 0 = no better than predicting this bin/group's base rate, <0 = worse than just predicting the base rate. This allows one to disambiguate between "low raw Brier because the markets are super good at forecasting" and "low raw Brier because the base rate is close to 1 or 0".
The "above 10k predictions are worse or at least not better" conclusion didn't change. This score also allows us to better interpret how markets perform at low volume. They are close to 0, meaning that at low volume relying on prediction markets is only mildly better than just using the base rate.
Conclusion 2: higher trading volume means better predictions (except when it's too high somehow).
Practical takeaway: for markets with trading volume between 100 mana and 10,000 mana, predictions improve as trading volume increases. But beyond 10k, predictions don't improve or even get less accurate. Also, the difference between 1k and 10k is not very large. The first 1k of trading volume buys most of the accuracy.
Now let's look at the part most relevant for forecasting AI capabilities years ahead - how well do prediction markets do at varying market lifespans?
Raw Brier:
Brier skill score:
Conclusion 3: a few days after creation, a market resolving a year out carries essentially no information beyond its base rate.
The blue early Brier skill curve goes down as market lifespan increases, meaning that events that are far away are inherently difficult to predict in advance. Raw Brier stays roughly the same: long-lived markets have a different base rate, which confounds it.
Practical takeaway: if you are looking at a recently made prediction market with a resolution date many years into the future, don't trust the probability too much (at least not more than you would trust the base rate).
Conclusion 4: long-term events aren't intrinsically unpredictable, they're unpredictable early.
The red late Brier skill curve goes up as market lifespan increases (raw Brier decreases), meaning that as time goes on...uh, idk, enough information is accumulated for the prediction to be nearly settled? I'm not sure how to explain this, maybe there is some confounder that I can't think of. Btw, trading volume is not correlated with a market's lifespan, in case you are wondering about that.
Practical takeaway: if you are looking at an old market where resolution is coming soon, you can trust the probability.
I'm really curious why as lifespan increases, early and late Brier skill curves diverge.
Ok, how about we make a grid and see how the Brier skill score depends on both the trading volume and lifespan?
Conclusion 5: markets (mostly) perform better than just using the base rate.
Markets (with at least 3 traders and lifespan of at least 7 days) aren't so dumb as to underperform the base rate. For most cells in the three grids above, Brier skill score is >0. Except in the early snapshot, for markets whose resolution is many months away and whose trading volume is low. Markets with short lifespans (7-20 days) and low volume (<=175 mana) are the most low-skill (aka unpredictable) ones even near their resolution. Something about markets with short lifespans makes them "coin-flippy" (relative to markets with long lifespans) even at high trading volume.
Practical takeaway: if you want to be really sure you're getting useful information (beyond the base rate), look for markets with resolution date no less than 20 days from the creation date and with total trading volume of at least 350 mana.
So far we haven't looked at the content of markets. Let's calculate base rates and Brier skill scores for 100 Manifold topics (essentially tags) with the most markets. The data below is sorted by the early Brier skill score (3 days after creation).
topic n base rate skill:early mid late
-------------------------------------------------------------------------------
death-markets 433 0.552 0.519 0.642 0.743
elon-musk-9SnzlR8p2A 812 0.353 0.359 0.503 0.700
chatgpt 333 0.366 0.305 0.403 0.671
predictions-on-predictions 306 0.422 0.289 0.433 0.608
twitter 915 0.372 0.285 0.422 0.637
uk-politics 558 0.403 0.282 0.405 0.693
basketball 437 0.364 0.274 0.412 0.611
uk 370 0.384 0.270 0.371 0.648
europe 436 0.339 0.270 0.408 0.667
donald-trump 2,691 0.381 0.265 0.418 0.709
united-states 634 0.390 0.262 0.332 0.617
us-politics 6,146 0.356 0.253 0.382 0.671
road-bicycle-racing 296 0.338 0.250 0.419 0.651
2024-us-presidential-election 2,712 0.354 0.243 0.357 0.670
magaland 1,535 0.350 0.242 0.412 0.688
elon-musk-14d9d9498c7e 1,451 0.323 0.241 0.427 0.695
metaculus 548 0.288 0.235 0.502 0.792
trumps-second-term 681 0.335 0.234 0.421 0.709
elections-world 388 0.466 0.232 0.370 0.601
openai 1,314 0.336 0.230 0.402 0.683
politics-default 6,584 0.349 0.227 0.375 0.670
the-life-of-biden 620 0.324 0.227 0.346 0.709
nba 695 0.456 0.212 0.297 0.464
elections 512 0.430 0.210 0.328 0.591
metaforecasting 417 0.410 0.208 0.318 0.507
science-default 1,291 0.292 0.205 0.366 0.641
world-default 2,001 0.300 0.204 0.371 0.669
russia 1,000 0.252 0.201 0.488 0.788
climate 399 0.424 0.199 0.410 0.713
news 663 0.314 0.199 0.329 0.637
space 682 0.343 0.197 0.393 0.726
business 1,056 0.354 0.178 0.321 0.648
spacex 394 0.398 0.178 0.382 0.754
sports-betting 411 0.436 0.177 0.279 0.473
soccer 974 0.437 0.177 0.263 0.468
apple 529 0.338 0.176 0.406 0.728
finance 1,367 0.434 0.175 0.297 0.505
please-resolve 299 0.368 0.173 0.324 0.661
democratic-party 311 0.277 0.173 0.274 0.573
health 357 0.339 0.169 0.263 0.560
sports-default 3,412 0.413 0.169 0.287 0.517
ai 4,215 0.318 0.165 0.312 0.637
stocks 1,247 0.482 0.159 0.270 0.550
ukrainerussia-war 1,428 0.234 0.159 0.479 0.771
manifold-users 595 0.422 0.157 0.344 0.580
china 683 0.264 0.156 0.301 0.659
technical-ai-timelines 962 0.368 0.156 0.295 0.631
movies 1,186 0.431 0.156 0.340 0.685
internet 746 0.338 0.156 0.344 0.666
118th-congress 298 0.309 0.153 0.243 0.586
technology-default 5,394 0.323 0.152 0.318 0.653
manifold-6748e065087e 2,733 0.369 0.151 0.308 0.597
fun 1,439 0.434 0.151 0.234 0.398
economics-default 3,099 0.403 0.151 0.302 0.590
2024-us-election 374 0.417 0.146 0.265 0.658
football 1,031 0.428 0.145 0.257 0.476
bitcoin 1,108 0.442 0.145 0.357 0.689
boxoffice 333 0.453 0.143 0.354 0.786
manifold-leagues 671 0.441 0.140 0.308 0.552
global-macro 644 0.275 0.139 0.345 0.655
entertainment 1,799 0.397 0.136 0.302 0.633
arabisraeli-conflict 761 0.304 0.136 0.338 0.641
prices 398 0.357 0.135 0.313 0.659
culture-default 1,523 0.344 0.133 0.313 0.634
weather 449 0.359 0.132 0.292 0.641
formula-1 858 0.404 0.131 0.235 0.319
celebrities 356 0.343 0.130 0.335 0.593
crime 318 0.305 0.130 0.356 0.607
youtube 740 0.405 0.128 0.280 0.704
ai-impacts 393 0.305 0.127 0.308 0.587
bitcoin-maxi 719 0.448 0.126 0.344 0.685
crypto-speculation 1,700 0.396 0.123 0.324 0.668
law-order 630 0.314 0.119 0.320 0.609
middle-east 490 0.253 0.118 0.318 0.615
metamarkets 575 0.397 0.117 0.234 0.437
lesswrong-annual-review 366 0.180 0.117 0.212 0.440
iran 469 0.222 0.114 0.345 0.565
llms 316 0.335 0.113 0.323 0.667
nfl 1,417 0.471 0.109 0.178 0.344
ai-safety 464 0.244 0.109 0.275 0.656
nonpredictive 1,213 0.465 0.105 0.201 0.378
israel 1,138 0.286 0.104 0.344 0.636
israelhamas-conflict-2023 893 0.300 0.104 0.312 0.614
ukraine 728 0.224 0.103 0.434 0.736
crypto-prices 795 0.475 0.100 0.263 0.507
personal 980 0.460 0.095 0.199 0.422
wars 1,191 0.212 0.092 0.342 0.697
geopolitics 780 0.212 0.092 0.255 0.627
music-f213cbf1eab5 523 0.371 0.091 0.213 0.571
competition-math 359 0.440 0.089 0.185 0.412
manifold-business-future 602 0.294 0.082 0.311 0.671
tesla 432 0.319 0.073 0.355 0.684
personal-goals 1,719 0.456 0.055 0.169 0.433
manifold-user-retention 286 0.322 0.053 0.346 0.596
mathematics 400 0.347 0.050 0.092 0.394
destinygg 2,279 0.315 0.043 0.295 0.642
manifold-features-25bad7c7792e 864 0.315 0.041 0.230 0.601
gaming 1,522 0.365 0.023 0.211 0.599
sccsq4 440 0.564 -0.007 0.011 0.057
new-years-resolutions-2024 6,082 0.163 -0.277 0.144 0.746
And here's a scatter plot of early Brier skill score and base rate:
This goes against something I said above: that elections-related markets would be among the least predictable ones. They're actually quite predictable. 'football' and 'soccer' aren't at the bottom of the table either.
'mathematics' is near the bottom, to my surprise. Naively, I thought it would be the topic with the least irreducible uncertainty.
'chatgpt' being near the top of the table while 'ai-safety' is near the bottom is another mystery.
'new-years-resolutions-2024' having significantly negative Brier skill score early on is interesting. Btw, with 6082 markets (after filtering) and YES base rate of 16.3%, this topic probably accounts for a good chunk of the overestimation bias I mentioned earlier. It's the only popular topic where simply predicting the base rate will get you ahead of everyone else (though as resolution approaches, Brier skill score rises to ~0.75).
'death-markets' being the most predictable one is unexpected. I guess actuarial tables are excellent priors?
'sccsq4' is the least predictable topic overall if we look at all 3 snapshots. It was a tag for some kind of user-run competitive leaderboard, mostly featuring stocks and crypto-related markets. It's odd that it has such a low late Brier skill score, given that 'stocks', 'crypto-speculation' and 'crypto-prices' all have higher Brier skill scores.
If anyone wants to play around with scraping and plotting, here: https://github.com/Expertium/manifold-calibration