Polymarket Calibration Analysis Across 4 Domains

Analysis evaluating Polymarket’s forecasting performance across NBA games, film awards, U.S. elections and economic releases. Aims to test both predictive accuracy and whether market prices behave as well-calibrated probabilities.
Polymarket Calibration Analysis Across 4 Domains
Like

Share this post

Choose a social network to share with, or copy the URL to share elsewhere

This is a representation of how your post may appear on social media. The actual post will vary between social networks

Introduction

Prediction markets have emerged as one of most significant innovations in modern forecasting. By allowing participants to buy and sell contracts linked to future events, platforms such as Polymarket convert dispersed information and collective intelligence into market-implied probabilities. Unlike conventional models that rely primarily on historical data, prediction market prices continuously update as new information is released, producing live forecasts. However, the question remains how accurately does Polymarket predict real-world outcomes, and to what extent do its prices represent well-calibrated probabilities?

Methodology

Polymarket's forecasting performance was evaluated across four domains that differ in nature of their uncertainty and the relevant information accessible about them.

NBA games: frequent, clear outcomes, lots of historical game data, real-time information such as injuries, line-up changes and fatigue.

Film awards: not purely statistical events but they also involve taste, reputation, campaigning and media narratives. 

US Elections:  incorporate news, scandals, debate performances, polling data, and shifts in public sentiment.

Macroeconomic forecasts: shaped by signals in financial markets, policy decisions, commodity prices as well as central bank reports. 

Comparing these domains therefore makes it possible to examine whether prediction-market quality depends on the type of event being forecast, rather than treating market performance as a single universal characteristic. 

Data Collection

The datasets were constructed using Polymarket's Gamma API to identify markets and the CLOB price-history API to retrieve time-stamped historical contract prices. For NBA games, the 2025–26 season (1,230 games) was examined. Because fewer historical markets were available in the remaining domains, resolved events from film-award, election and economic markets were collected from 2023 onward to obtain sufficiently large but still recent samples.

Not every event contained price history far in advance of the resolution. A fixed cohort was therefore established for each domain at the earliest horizon that preserved a sufficiently large sample while still allowing meaningful price development to be studied. The same events were then followed through all subsequent horizons, ensuring that movements in the results reflect changing contract prices rather than changing sample composition.

  • NBA games: 1,019 games from the 2025–26 season, observed over the final 5 days before the game.
  • Film awards: 100 resolved events with price data available from 39 days before resolution.
  • U.S. elections: 95 resolved events, observed from 5 months before resolution.
  • U.S. economic indicators: 139 releases covering CPI, Federal Reserve decisions, GDP and unemployment, with price data available from 7 days before release.

Forecast Evaluation

Forecasting performance was evaluated using accuracy, Brier score and calibration.

Accuracy - records how often the highest-priced contract correctly identified the eventual outcome. 

Brier score - measures the accuracy of probabilistic predictions by computing the average squared difference between predicted probabilities. 

The accuracy graph reveals substantial differences in predictability across the domains. US election markets remain the most accurate throughout, generally between 94% and 99%. US economic indicators hold steady around 88%. Film-award markets climbed the most, from roughly 62% at 39 days to about 80% at the final hour. By contrast, NBA accuracy increases only modestly from about 66% five days before tip-off to 68% one hour beforehand, suggesting that much of the remaining error reflects irreducible uncertainty in games rather than a lack of available information. Most series also stabilise before the final day, indicating that last-minute trading rarely changes the market’s predicted winner. 

Accuracy only changes when the predicted winner changes and therefore misses improvements in the probability assigned to that outcome. The Brier-score, on the other hand, provides a more sensitive measure of forecasting quality whereby a lower score represents smaller probability errors.

The Brier score improves in every domain as resolution approaches. Film awards decline the most from approximately 0.164 to 0.130, NBA games from 0.216 to 0.199, economic indicators from 0.087 to 0.080, and election markets settle near 0.02. The relatively high NBA Brier score should not be interpreted as poor forecasting since individual games are inherently uncertain and produce many probabilities around 30–70%. Consequently, even well-calibrated NBA forecasts naturally generate larger squared errors. Film awards display the biggest improvements in Brier score, giving evidence that new information improves probability forecasts over time. This is supported by the prediction study by Pardoe & Simonton (2008) which revealed that awards such as the Oscars rely heavily on results from precursor awards. By one week before the award this improvement has already occurred, suggesting that this additional information for improving predictions eventually reaches a limit.   

Contract-Level Calibration

While accuracy and Brier scores measure overall forecasting performance, they do not reveal whether individual contract prices represent reliable probabilities. Contract-level calibration assesses whether contracts trading at a given probability win at approximately the same frequency. For example, contracts priced at 70% should win about 70% of the time. To evaluate this, contracts were grouped into 10-percentage-point price bands and the mean price in each band was compared with the corresponding actual win rate. Calibration was examined at the earliest consistent observation point, an intermediate horizon, and one hour before resolution.

NBA markets show the strongest calibration. Observed win rates remain very close to the 45-degree perfect-calibration line at all three horizons, with ECE remaining very low. This suggests that much of the relevant information is already incorporated several days before games begin. Professional basketball provides extensive public information, including historical performance, injuries, line-ups and external betting odds, which gives traders strong benchmarks against which Polymarket prices can be evaluated. Thus the high Brier score seen earlier reflects the underlying event uncertainty rather than poor probability calibration.

Election markets present the opposite combination. They have extremely low Brier scores but larger calibration errors, with ECE falling from approximately 0.068 five months out to 0.033 one day before resolution. Many election contracts are concentrated close to zero or one, allowing the market to achieve a very low average squared error by correctly identifying near-certain winners and losers. However it seems to be less precisely calibrated in genuinely competitive contests. Nevertheless, the decline in ECE shows that polling and other incoming information does improve calibration as elections approach.

Film-award markets display the strongest learning effect. At 39 days, ECE is approximately 0.069, with many low- and mid-priced nominees overvalued and stronger favourites winning more frequently than their prices imply. By seven days and especially one hour before the outcome, both ECE and the Brier score fall substantially. This convergence is consistent with the changing information environment during awards season. 

The U.S. economic-indicator markets, by contrast, show more stable calibration with the ECE changing minimally from 0.023 seven days before the outcome to 0.024 one hour before release. This implies that, unlike election campaigns or award races, the principal information needed to forecast macroeconomic releases is already available well before the final day.

Logistic Recalibration

The binned calibration graphs must be interpreted cautiously because several intermediate probability bands in the film-award, election and economic datasets contain fewer than 20 contracts, with some economic bins containing only three or four. Apparent deviations from perfect calibration may therefore reflect sampling variation rather than genuine mispricing.

To provide a more robust assessment, logistic recalibration uses the full distribution of contract prices:

Perfect logistic calibration corresponds to α=0 and b=1. The slope, b, measures probability dispersion where b >1 indicates underconfidence (probabilities are too compressed towards 50%), whereas b<1 indicates overconfidence (probabilities are too extreme).

Statistical uncertainty was assessed using event-level bootstrapping. At every horizon, whole events were resampled with replacement 1,000 times and the logistic model was refitted to each sample to construct 95% confidence intervals. Resampling at the event rather than contract level preserves dependence between multiple contracts belonging to the same underlying event.

The recalibration results reinforce the domain differences. NBA and economic-indicator markets remain close to b=1, with bootstrap confidence intervals containing one throughout their observed horizons. There is therefore little evidence of over- or underconfidence. Film awards show moderate underconfidence. Polymarket generally identifies favourites and longshots correctly but does not separate their probabilities strongly enough, consistent with a favourite-longshot bias.

The strongest deviation occurs in U.S. elections. Despite their very low Brier scores and high accuracy, election probabilities are systematically too underconfident. This may partly reflect the large number of mutually exclusive low-probability contracts. Early slope estimates are highly unstable because the sample contains only 95 events and many probabilities cluster near 0 or 1, producing wide bootstrap intervals. More importantly, once the slope stabilises near resolution, it remains clearly above one, indicating persistent underconfidence even as more information becomes available. 

Conclusion

Overall, these results show that Polymarket's forecasting quality cannot be reduced to a single measure of accuracy. NBA markets are almost perfectly calibrated yet produce high Brier scores simply because games are inherently uncertain, while election markets combine near-perfect accuracy with the strongest evidence of underconfidence. Film awards demonstrate the clearest learning effect, with both calibration and Brier score improving substantially as new information arrives. This suggests that prediction-market quality is shaped less by the amount of available information than by the structure of uncertainty within each domain. 

Please sign in

If you are a registered user on Laidlaw Scholars Network, please sign in