Overview
Introduction
Prediction-market accuracy is the degree to which market-implied probabilities match outcomes across a defined set of contracts at a stated forecast horizon, under a disclosed scoring and sampling method. Prediction markets can produce useful and sometimes well-calibrated forecasts, but accuracy is conditional and cannot be reduced to one all-purpose percentage.
A market may aggregate information quickly, yet its result still depends on the questions selected, when prices are sampled, which quote represents the forecast, and what benchmark is used. A correct favorite is not enough to prove that the probabilities were good. Likewise, one unexpected result does not make a 30% forecast wrong. A credible test needs repeated forecasts, resolved outcomes, proper probability scores, fixed timing, and transparent exclusions. A proper scoring rule is designed so a forecaster's expected score is optimized when the report matches the probability they believe, with no reward for shading it toward a preferred outcome.
Key takeaways
How Accurate Are Prediction Markets?
Prediction markets have produced useful forecasts, but no result applies to every venue, category, or horizon. The Commodity Futures Trading Commission says they can sometimes forecast outcomes better than polling or other methods. Regulation, activity, or a familiar name does not certify a displayed probability.
Accuracy belongs to a predeclared forecast sample. Specify which contracts qualify, when prices are observed, how nonstandard endings are handled, and which score is calculated. Readers who need the lifecycle first can review how prediction markets work .
Evidence carries more weight when comparable contracts resolve cleanly and prices were recorded before the outcomes became knowable.
Three questions need separating. Probability quality asks whether quoted chances match outcomes. A winner call asks only whether the favorite won. Profitability depends on price, spread, fees, fills, and size. Calibration need not offer a profitable trade, while a lucky profit need not validate a method. A 2004 research review found supportive evidence alongside favorite-longshot effects and small-market delays. Favorite-longshot behavior means markets tend to overprice low-probability outcomes and underprice high-probability ones.
What Accuracy Means for a Probability Forecast
A binary price commonly serves as an implied market probability, not verified truth. Risk preferences, capital constraints, market design, stale quotes, and trader beliefs can create a wedge between price and objective probability.
The word “accurate” combines five tests. Each answers a useful question, but none provides the whole verdict.
| Test | What it answers and misses |
|---|---|
| Favorite-won hit rate | Shows how often the side above 50% won, but gives a 51% and a 99% favorite equal credit. |
| Calibration | Tests whether events quoted near a probability occur at about that frequency, but can reward an uninformative base-rate forecast. |
| Brier or log score | Measures probability error across outcomes, but changes with the event set, horizon, weighting, and score convention. |
| Discrimination (Brier resolution) | Shows whether forecasts separate events with different outcome rates, but does not prove that probability levels are calibrated. |
| Benchmark skill | Shows whether a market adds value over a predeclared alternative, but only when the target, sample, and timing match. |
A hit rate is secondary evidence. Obvious favorites can produce an excellent classification result while saying little about difficult events. Proper rules grade the probability itself. Gneiting and Raftery's scoring-rule framework explains why quadratic and logarithmic scores suit that job.
The event sample sets the difficulty. Report its base rate, probability bands, category mix, horizon, contract count, and event count beside any aggregate score.
How Brier Scores and Log Loss Test Market Probabilities
Proper scores grade the full probability. Lower loss is better. The formula, timestamp, quote type, event universe, and weighting rule must accompany the result.
Brier Score
For a binary event using the one-term convention:
Brier score = mean((p - y)^2)
Here, p is the forecast probability from 0 to 1 and y equals 1 for Yes and 0 for No. This version ranges from 0 to 1. A 70% Yes forecast loses (0.70 - 1)^2 = 0.09 when Yes occurs and (0.70 - 0)^2 = 0.49 when No occurs. An always-50% forecast scores 0.25 under this one-term convention, but that is not a universal “chance” benchmark.
Some research sums the Yes and No terms as a two-outcome vector. It ranges from 0 to 2 and doubles the one-term value, so an always-50% forecast scores 0.50. A naive benchmark should reflect the event sample's base rate.
A stale last trade, the midpoint, and the current best bid price can each imply a different probability for the same contract at the same timestamp. The current best ask price may represent an executable purchase, subject to available size. Disclose one primary quote rule and test reasonable alternatives.
Log Loss
For a binary event:
log loss = -(y × ln(p) + (1 - y) × ln(1 - p))
Log loss penalizes confident misses more sharply than the Brier score. It becomes unbounded when an event assigned zero probability occurs. Predeclare whether exact 0 and 1 values are allowed or clipped to fixed bounds. Never change those bounds after outcomes are known.
| Metric choice | What it emphasizes |
|---|---|
| One-term Brier score | Squared probability error on a 0-to-1 scale |
| Two-outcome Brier score | The same squared error, doubled onto a 0-to-2 scale |
| Log loss | Steeper penalty for confident tail misses |
| Brier skill score | Gain over a stated baseline, where higher is better and zero means no gain. |
No good Brier threshold applies without the event set, horizon, convention, and baseline. Easier contracts can produce a lower score without a better method.
Why Calibration Alone Is Not Enough
Calibration, also called reliability, compares quoted probabilities with observed frequencies. If many comparable contracts priced near 70% resolve Yes about seven times in ten, that bin is well calibrated. A reliability diagram should show the mean quoted probability, observed frequency, number of observations, and uncertainty intervals for each bin.
Reliability and Resolution
Perfect aggregate calibration can still carry no discriminatory information. Suppose half the events in a sample resolve Yes and a forecaster assigns 50% to every one. The forecast is calibrated in aggregate, but it never distinguishes a likely event from an unlikely one. In the Brier decomposition, that lack of discrimination means it has no resolution.
Allan Murphy's Brier-score decomposition expresses the relationship in loss notation:
Brier score = reliability - resolution + Brier uncertainty
The reliability term measures how far quoted probabilities sat from observed frequencies, so a smaller reliability term means better calibration. Resolution rewards separation into groups with different outcome rates. Brier uncertainty reflects the base rate of the evaluated outcomes. Sharpness describes how far forecasts move from the base rate, but it is useful only when paired with calibration.
Expected calibration error changes with bin boundaries. A venue-wide estimate can hide category-level errors that run in opposite directions and cancel in the average. Stratify by probability band, horizon, and category, with bin counts and uncertainty intervals reported for every stratum. Page and Clemen's 2013 calibration study shows why horizon matters.
What Research Finds About Prediction-Market Accuracy
The research supports conditional usefulness, not a category-wide accuracy rate. Evidence status is as important as the headline result because peer-reviewed designs, operator analyses, and current working papers carry different levels of independence and finality.
| Finding | Scope and limit |
|---|---|
| NBER working paper: markets were generally fairly accurate and often beat moderately sophisticated benchmarks. | Wolfers and Zitzewitz reviewed older settings in 2004, while also documenting bias and small-market limits. |
| Peer-reviewed election study: Iowa Electronic Markets (IEM) vote-share forecasts were closer than contemporaneous raw polls in 74% of 964 comparisons. | Berg, Nelson, and Rietz matched five US elections from 1988 through 2004. The figure is not a hit rate or category-wide accuracy percentage. |
| Peer-reviewed calibration study: longer-horizon prices displayed favorite-longshot behavior. | Page and Clemen analyzed 1,787 markets grouped into 597 competitions. Their screened, older sample is not a current venue score. |
| Peer-reviewed randomized comparison: aggregated beliefs and prices carried complementary information. | Dana and coauthors studied 535 forecasters and 113 geopolitical questions using a 0-to-2 Brier convention. |
| Peer-reviewed combination study: combinations beat component methods on average. | Graefe and coauthors combined polls, experts, models, and IEM forecasts across six US elections. |
| Operator-hosted working paper: a Kalshi analysis reports large cohort, tail, and clock effects. | Kagan and Baiocchi report 2,243,741 resolved markets through mid-2026. The paper is not independent peer-reviewed proof. |
| Operator self-study: Hypermind reported that, after adjusting for category mix, its Brier scores were on par with Polymarket and Kalshi and its calibration was similar to both. | Hypermind’s analysis covers 1,141 resolved markets. It is first-party work, not independent validation or a venue ranking. |
| Working paper under review: price discovery in a Polymarket dataset was attributed mostly to about 3% of accounts. | Gomez-Cram and coauthors revised the work in June 2026. The estimate is dataset-specific. |
The IEM election study also reported average absolute error of 1.20 percentage points for market forecasts versus 1.62 for contemporaneous polls over the final five days. In the randomized matched-information study, Dana's pre-specified belief aggregate scored 0.210 versus 0.227 for prices under the two-outcome convention, but that difference was not statistically significant. A hybrid significantly beat prices alone.
Page and Clemen reported greater favorite-longshot distortion in political markets than in the other categories in their screened historical sample. Their results do not apply to every current contract.
Under the one-term convention, the Kalshi working paper reports that its changing cohort, which admits short-duration markets as the horizon narrows, scored 0.126 at the one-day horizon after roughly 492,000 short-duration markets entered. Its fixed cohort of 35,562 contracts, which follows the same contracts at every horizon, scored 0.039 at the same one-day horizon. The gap between those scores shows cohort sensitivity. The changing cohort scored worse rather than better, so a short duration does not by itself make a contract easy. The direction of a cohort effect has to be measured for the sample in hand, not assumed from the horizon. Separate analyses in the paper show that clock choice and tail exclusions also change results, so the findings do not establish a venue ranking.
Prediction Markets vs Polls, Experts, and Models
A poll, model, expert aggregate, and market price need not forecast the same quantity. Vote intention is not a winner probability. A poll model transforms survey data through turnout and uncertainty assumptions. Comparing raw percentages across these targets is a category error.
| Method | What it actually estimates |
|---|---|
| Raw poll | Responses, opinions, expectations, or vote intention in a sampled population |
| Poll-derived probability model | Outcome probability after modeling sampling error and other assumptions |
| Expert aggregate | A combined judgment from selected forecasters under an aggregation rule |
| Market price | A tradable estimate of what the contract pays, which is an implied probability for a binary contract and an implied vote share for a share-indexed one |
| Combined forecast | A rule-based blend of several component estimates |
Markets have beaten useful benchmarks, but matched belief aggregates and combinations can equal or outperform them. Dana and coauthors asked participants for beliefs before they traded, aligning the information set. Graefe and coauthors' forecast-combination study found average combinations more accurate than components in its six-election vote-share sample, with error reductions of 16% to 59% against average component errors.
A fair contest needs the same target, event universe, timestamp, information cutoff, weighting, and proper score. Compare methods event by event with paired uncertainty intervals. If one forecast is revised later, information sets no longer match. “Markets beat polls” requires a study design, not a product claim.
How Time, Liquidity, Bias, and Manipulation Change Accuracy
Scores can improve as information arrives, but an easier late sample can mimic learning.
Forecast Horizon and the Event Clock
Report the fixed and the changing cohort together to separate learning from composition.
If the outcome becomes knowable before the platform's administrative close or the sampled closing quote, scoring that later quote measures post-event information instead of forecast quality. Useful checkpoints include 90 days, 30 days, seven days, and 24 hours before the event under one eligibility rule.
Liquidity and Participant Mix
Liquidity in event markets includes more than volume. Trader count, spread, depth, quote age, and independent information are different signals. Thin books can remain stale, while high volume can be concentrated.
Page and Clemen found no significant volume effect within their screened sample, while Kalshi's working paper reports correlations with volume and trader count. Gomez-Cram and coauthors attribute much price discovery to a small minority of accounts in their dataset. None supplies a causal cutoff that applies across venues.
Bias and Manipulation
Page and Clemen measured the pattern directly: prices near 0.20 resolved Yes 15.3% of the time, while prices near 0.80 resolved Yes 87.4% of the time in their screened sample, where the gap widened with time to expiration.
Attempted distortions may attract arbitrage across market prices, but correction is not guaranteed. Hanson, Oprea, and Porter found correction in a laboratory study. A contrary experiment by Deck, Lin, and Porter found destructive manipulation under different incentives and funding.
A temporary quote move affects the score only if present at the observation timestamp, whereas a persistent distortion survives whenever the quote is sampled. Outcome interference changes the real-world result rather than the price, while settlement or resolution interference targets the source or rule that determines how the contract pays. A distortion that moves the price also moves measured volatility, but measuring market volatility does not itself measure forecast error. Rules and enforcement belong in the framework used to evaluate prediction market integrity .
| Factor | What the evidence supports |
|---|---|
| Horizon | More information can improve forecasts, but fixed and changing cohorts must be separated. |
| Liquidity | Activity may support correction, but volume and headcount do not guarantee calibration. |
| Bias | Favorite-longshot effects are conditional by horizon, category, and design. |
| Manipulation | Persistence depends on capital, incentives, depth, counter-trading, and the mechanism involved. |
How to Audit a Prediction Market Accuracy Claim
A reproducible audit fixes its rules before outcomes are inspected. It records the disposition of every eligible contract and ties all horizons to the real event clock.
Freeze the Sample and Clock
- Define the venue, category, date range, metric, and claim being tested.
- Freeze the eligible contract universe and every exclusion rule before scoring outcomes, so famous winners cannot be selected later.
- Group duplicate listings, mutually exclusive candidates, and threshold ladders by underlying event instead of treating them as independent evidence.
- Record unresolved, canceled, voided, split, corrected, and disputed dispositions before identifying the subset that can be scored.
- Anchor the event clock to when the result became knowable, not the platform's later administrative close or settlement.
- Sample a fixed cohort at predeclared horizons, or show fixed and changing cohorts together when market duration differs.
Choose the Quote and Score
- Predeclare last trade, midpoint, bid, ask, or another quote rule, then report staleness, spread, and missing observations.
- Give each underlying event equal primary weight. Show contract-weighted, trade-weighted, or volume-weighted alternatives only as labeled sensitivity analyses because one active question can otherwise dominate.
- Calculate at least one proper score and print its formula, scale, and treatment of exact zero and one probabilities.
- Report calibration, bin counts, uncertainty intervals, and Brier resolution, so an uninformative base-rate forecast is not mistaken for a useful one.
Compare and Report
- Select a base-rate, model, expert, poll, or venue benchmark before examining results, using the identical event set and timestamps.
- Stratify results by category, horizon, probability band, duration, and liquidity conditions before making broad claims.
- Compare paired events with uncertainty intervals that account for dependent contracts, using clustering or block resampling where appropriate.
- Publish provenance, dataset versions, research status, conflicts, unavailable records, and exclusions needed for replication.
Historical collection should use stable identifiers, native timestamps, raw rules, and resolution states. The process to access prediction market data supports a reproducible sample without changing the evaluation design.
| Audit check | Required disclosure |
|---|---|
| Sample | Eligible universe, exclusions, dispositions, event grouping |
| Time | Underlying cutoff, fixed horizons, quote timestamp |
| Price | Quote type, staleness, spread, weighting |
| Score | Formula, convention, calibration, Brier resolution, uncertainty intervals |
| Comparison | Identical events, times, target, information cutoff, benchmark |
| Provenance | Sources, versions, research status, conflicts, missing data |
How Much Weight Should You Give a Market Price?
A current quote can be useful evidence without being a verified probability or profitable opportunity. Calibration belongs to a series of comparable forecasts, not to one quote. Before relying on the quote, ask seven questions:
- Does the contract define the outcome, source, cutoff, and cancellation treatment clearly?
- Is the timestamp from before the outcome became knowable?
- Is the displayed quote recent enough to reflect current information?
- Do the spread and available depth support the interpretation being made?
- Does historical evidence cover this category and forecast horizon?
- Is the cited score based on a complete, fixed, comparable sample?
- How does the same forecast compare with an external benchmark at the same time?
The mechanics used to interpret prediction market odds explain price, payout, fees, and break-even calculations. Reading how a crypto order book works shows why bid, ask, spread, and depth each place a different constraint on what a displayed probability supports.
When conditions are unclear, reduce the evidential weight assigned to the quote. That is an interpretation control, not a trading signal or confidence threshold.
Check Current Markets Without Treating Price as Proof
Readers can inspect current prediction market platforms after understanding the sampling and scoring controls. A displayed contract price remains one forecast observation, not evidence that the venue or category is historically accurate.
CryptoSlate also publishes live crypto asset prices. A spot price shows an asset's current trading level, useful background for an event contract linked to that asset. A crypto price move is not a prediction-market accuracy score.
Apply the same caution to leaderboards. Without matched questions, times, quote rules, and dispositions, a venue-wide score can reflect category mix rather than forecast skill. Current screens support observation and discovery, not standalone validation.
Frequently Asked Questions
Can a well-calibrated market still get an event wrong?
Yes. Across comparable cases quoted near 70%, about seven in ten should resolve Yes and about three should resolve No. Calibration is a repeated-sample property, so one surprise does not show that a forecast was defective. Evaluation needs enough cases in that band to separate calibration from luck. Report the bin count alongside the observed frequency.
What is a good Brier score for a prediction market?
There is no universally good Brier threshold. A Brier score needs a stated baseline and must specify whether it uses the 0-to-1 binary form or the 0-to-2 two-outcome form. A score on near-certain contracts cannot fairly be compared with one on difficult questions without matching or stratification, and duration does not reliably indicate which is which.
Compare the score with a predeclared base-rate or competing forecast on identical events and timestamps, and report paired uncertainty intervals instead of applying an unsupported quality label.
Are prediction markets more accurate than polls?
Markets have beaten raw polls in some matched studies, including the five-election IEM comparison. Aggregated beliefs and combined forecasts can match or outperform markets in other designs. A fair comparison uses the same target and matched pre-event observations.
A raw vote-intention percentage and a win probability are different targets, so convert both to the same question before comparing performance.
Does more liquidity make prediction markets more accurate?
Liquidity can make correction easier, but it does not provide a causal accuracy threshold. Treat volume, spreads, depth, quote age, and trader count as separate diagnostics. The Kalshi operator-hosted working paper reports correlations with activity, not independent peer-reviewed proof that more liquidity caused better forecasts.
Can one trader manipulate a prediction market forecast?
One trader can move a quoted probability, and laboratory experiments have produced both correction and lasting distortion. For accuracy, what matters is whether the move survives until the predeclared observation used for scoring. A thin-book move is not proof of persistent distortion when the quote reverts before that timestamp.
Rules analysis separately asks whether the conduct violated the contract or venue policy.
Which prediction market is most accurate?
No venue is a timeless winner across all event sets and horizons. A fair comparison requires a matched sample and one predeclared scoring method, with uncertainty intervals reported. Venue-wide scores built from unmatched categories cannot establish which market is generally most accurate.
Recheck when category mix, design, or data access changes because a reproducible matched sample can support a narrow finding, not a permanent ranking.