The Temporal Structure of Forecast Error on Polymarket: A Decomposition of Brier Loss over the 30 Days before Operational Market Closure
Author: Seiryu Ando (Graduate School of Economics, Kyoto University)
JEL Classification: G14, D83, C53, C58
Keywords: Prediction Markets, Polymarket, Brier Loss, Forecast Error, Murphy Decomposition, Brier Skill Score
What this study measures
Accuracy at one point in time does not show when a prediction market's forecast error declined. This study aligns Polymarket prices at six checkpoints from 30 days to 12 hours before operational market closure and measures how their error relative to the final outcome changes over time.
The outcome measure is Brier loss, the squared difference between a binary final outcome and the market price. A decline means that the price became more accurate relative to the realised outcome. It does not directly measure objective uncertainty, new information, or market participants' subjective uncertainty.
Measurement design and data
The study covers October 25, 2020 through July 3, 2026. From a candidate population of 548,839 markets and 87,691 events, I drew a stratified random sample of 30,000 markets across 468 strata defined by domain, closure period, volume and liquidity, listing duration, and event structure.
Gamma closedTime defines operational market closure. The checkpoints are 30 days, 14 days, 7 days, 3 days, 1 day, and 12 hours before closure. At each checkpoint, the analysis uses the last Yes-token API price observed at or before the target, requires it to be no more than 12 hours old, and does not interpolate prices. Outcomes come from official CLOB winner records.
Markets are first averaged within events, after which each event receives equal weight. This prevents events containing many related markets from dominating the mean. Confidence intervals use an event-cluster bootstrap.
| Sample stage | Markets | Events |
|---|---|---|
| Candidate population | 548,839 | 87,691 |
| Stratified random sample | 30,000 | 21,349 |
| Structurally eligible at 30 days | 1,912 | 1,432 |
| Fixed cohort observed at all six checkpoints | 1,818 | 1,393 |
The primary analysis follows these same 1,818 markets at every checkpoint, avoiding changes caused solely by turnover in the set of observed markets.
Main result
Mean Brier loss declines from 0.088253 at 30 days to 0.034962 at 12 hours. The absolute decline is 0.053291 (95% confidence interval: 0.045783 to 0.061038), and the relative decline from the 30-day value is 60.384% (95% CI: 54.248% to 66.110%).
| Time to operational closure | Mean Brier loss | 95% confidence interval |
|---|---|---|
| 30 days | 0.088253 | 0.080205 to 0.096359 |
| 14 days | 0.064492 | 0.057513 to 0.071742 |
| 7 days | 0.054552 | 0.047597 to 0.061720 |
| 3 days | 0.044141 | 0.037857 to 0.050781 |
| 1 day | 0.036863 | 0.031391 to 0.042677 |
| 12 hours | 0.034962 | 0.029473 to 0.040768 |
Comparing rates of decline
Because the checkpoints are unevenly spaced, the exploratory analysis compares mean decline per day. The estimate is 0.001485 per day from 30 to 14 days and 0.002187 per day over the combined period from 14 days to 12 hours. The difference is 0.000702 (95% CI: 0.000138 to 0.001275).
This comparison is exploratory. The 14-day boundary was selected after examining the data, and the interval does not account for boundary selection or multiple comparisons. Estimates across the five adjacent intervals are not monotonic, and the final interval from one day to 12 hours includes zero in its confidence interval. The result therefore does not establish continuous acceleration.
Murphy decomposition and Brier Skill Score
In the weighted Murphy decomposition, uncertainty remains 0.139128 because the outcome composition of the fixed cohort does not change. Resolution—the extent to which different probability forecasts separate different realised outcome frequencies—increases from 0.058806 at 30 days to 0.107274 at 12 hours. This increase is the dominant accounting component of the decline in Brier loss.
The Brier Skill Score relative to a base-rate forecast rises from 0.365671 at 30 days to 0.748707 at 12 hours. The latter means that Brier loss is about 74.9% lower than under the base-rate forecast for the same cohort. Resolution and BSS describe forecast accuracy; neither directly measures information acquisition or the rate at which uncertainty is resolved.
Robustness and the price-measurement audit
The relative decline from 30 days to 12 hours is 60.384% under equal-event weighting, 61.214% under equal-market weighting, 61.190% under Sampling-IPW, 59.809% under Observation-IPW, and 60.640% when both IPW schemes are combined. The magnitude of the main result is broadly maintained across these alternatives.
The API-versus-raw-trade concordance audit uses 91 markets. The mean Brier difference is not statistically distinguishable from zero, but the median timestamp gap is 95.1 minutes and the cohort is small and unrepresentative. This does not establish equivalence between API and raw prices or the absence of measurement error.
Limits of interpretation
The main estimates are conditional on Polymarket's API price-history series. The aggregation rule within the API fidelity setting is not established, and the data do not directly measure the historical last trade, midpoint, or order-book consensus.
Gamma closedTime marks operational closure, not when the outcome became known or the exact settlement time. The fixed cohort is also restricted to long-lived markets that existed at 30 days and have prices at all six checkpoints. The findings do not generalise directly to short-duration markets or to all Polymarket markets. The bootstrap intervals are conditional on the sampled 30,000 markets and exclude first-stage sampling uncertainty.
What the results mean
Prediction-market “accuracy” is not a single number. A reproducible estimate requires choices about the event-time anchor, lead time, fixed versus changing cohorts, market versus event weighting, and allowable price staleness.
The practical contribution of the paper is not to rank Polymarket against other forecasting methods. It is to specify the design conditions needed to measure prediction-market accuracy transparently and reproducibly.
Paper
The Temporal Structure of Forecast Error on Polymarket: A Decomposition of Brier Loss over the 30 Days before Operational Market Closure
- Posted on SSRN: July 24, 2026
- Last revised: August 15, 2026
Read the revised paper on SSRN ↗
Related Book Project
I am also writing a Japanese-language book that extends the paper's questions beyond forecasting to hedging and market design: Prediction Markets: Know the Future, Prepare for the Future.
