The record

The record

When a model disagrees with the market, who is right?

We measured it, on every settled selection we have, for two different forecasts. The answer is uncomfortable and it is the most useful thing on this site: the further a forecast strays from the closing price, the worse it does — and by the scoring rule that punishes confident mistakes, neither forecast beats the price in any band at all.

Which forecast

Two forecasts, measured the same way against the same closing prices. They do not have the same record and they do not have the same threshold — which is the point of showing both.

01 Use it

Check a disagreement

Put in what a forecast says and what the market says. The record answers with what happened the last time a gap that size was settled — and refuses to answer when the gap is too small to mean anything.

Your model, a tipster, your own read — any probability that did not come from the price.

The de-vigged price. Take the sharpest book you can see and strip the margin.

Disagreement 8.0pp

Reading the record…

Measured over settled selections, a Dixon-Coles goals model against the Shin-de-vigged closing price. It describes THAT model. A different forecast has a different record — and its own bar: bands under pp are withheld here, because below that the “right” count stops tracking which forecast is actually better.

02 The finding

Every band, with its sample

Each bar is how often the market finished closer to the result than the forecast did. The hairline is the coin flip. Read them downward: a forecast holds its own while it agrees, and loses ground as it argues.

How often the market was closer
 

Each row is a band of disagreement between the forecast and the sharp closing price, and the figure is how often the market turned out to be closer to the result. When the two roughly agree it is a coin flip. The further the forecast strays, the more reliably it is the one that was wrong. That is why a model holds a veto here and never a price.

Below 6pp this forecast is not quoted beside a selection · in-sample

03 The catch

Two scorekeepers, and they disagree

There are two ways to ask who was right, and they do not always answer the same. One counts how often each side finished closer to the result. The other scores every forecast by how much probability it put on what actually happened — a Brier score, which punishes being confidently wrong rather than merely being on the wrong side.

Where they split, the count is the one doing the flattering. It rewards a forecast for winning near-ties by a hair and never notices it losing the occasional band by a mile. A model with crude inputs hedges toward the middle, wins more hairs, and looks better than it is.

This is why the threshold is a property of the forecast and not of the site. Each one is withheld below the gap where its own two scorekeepers stop contradicting each other. Setting a single number for everything would have meant mislabelling whichever model it was not written for.

04 What it is not

The limits, stated

It describes these forecasts. A Dixon-Coles goals sheet fitted per season, and an Elo ladder built from results alone. A different forecast — yours, a tipster’s, a neural net’s — has its own record, and neither of these transfers to it. What almost certainly does transfer is the shape: a forecast fighting a liquid closing price is fighting the aggregate of everyone who has already done the work.

Under its own bar each forecast says nothing, deliberately. In the narrow bands a “model right 51%” headline drawn from a tie would state, in the only way a reader parses it, that the model beats the market. It does not — and where the closer-count says otherwise, the Brier score says it plainly.

It is about accuracy, not profit. Being closer to the truth and being a winning bet are different things — a price can be wrong in your favour on a match your model has no read on at all. Our own settled record, including the parts that are not yet significant enough to publish, is on the performance page, and the method behind all of it is on how it works.