SETTLED.

Prediction-market claims, checked against resolution data.

Numeric forecasts vs temperature-bucket prices

No LLM this time: pure meteorology against market asks, at seven different lead times, with three forecast models and per-station bias correction. The asks won every single round.

The hypothesisDaily highest-temperature markets are priced by amateurs, and pricing should be worst far from resolution. A gaussian model over professional weather forecasts — entered at the right lead time — will beat the bucket asks. If raw forecasts fail, per-station bias correction across multiple models will find the residual edge.

The test

An hourly daemon snapshotted every open city-temperature market: the multi-model forecast for the resolution station, our model's probability for each temperature bucket, and the bucket's live executable ask. Paper entries fired at seven lead-time checkpoints so "bet early vs bet late" became measurable rather than debatable. When raw forecasts failed, we trained per-station bias corrections (48 stations, three forecast models against observed settlement truth) and re-ran the entire entry gate out-of-sample — the strongest version of the thesis we could construct.

The result

Every checkpoint lost, from −40% to −64% ROI at executable asks over dozens-to-hundreds of settled entries each. Later entries lost less — the market converges — but no lead time was positive. The bias-correction rescue also failed: despite the blend being measurably more accurate than any single model (blend error 1.61° vs 2.37° for the worst source), the corrected gate scored −80.1% out-of-sample, statistically identical to the uncorrected −79.1%. The asks were efficient against every forecast configuration we could build.

Conclusion

Whoever prices these buckets already uses forecasts at least as good as the public models — improving the forecast doesn't create an edge when the price already contains it. This null pairs with our audit of actual weather-market winners: the one genuinely profitable trader we verified earns single-digit ROI from station-specific settlement quirks, not from out-forecasting anyone. The data platform survives as intelligence for manual trading; the automated entry thesis is dead at every lead time.