Model Accuracy
PRO BETAPer-model forecast accuracy & bias tracking across stations and time windows
Select a station to view model accuracy data.
The pipeline
Model forecasts are captured continuously as they publish — every run of every model, hour by hour — and archived unmodified. Once a day (overnight), each archived forecast is compared against what actually happened, and the scores on this page are recomputed. Nothing is graded until the truth for a day is in; nothing is regraded to make history look better. Forecasts are raw model output: no bias correction of any kind is applied before scoring.
What counts as "the day"
Every score uses the official climate day: midnight-to-midnight in the station's local standard time, year-round — the same window the NWS CLI report summarizes and the same window daily weather markets settle on. During daylight-saving time this runs 1:00am to 1:00am on the local clock; that is not a bug, it is exactly how the settlement products work. Forecast data is sliced to this same window before totaling, so a model is never credited or penalized for weather that fell outside the settlement day.
Truth
Temperature: the verified daily high/low from the NWS CLI (or DSM) climate report when available, otherwise the day's most extreme METAR observation. At METAR-only international stations, values are rounded from whole-Celsius readings and can differ ±1°F from the true extreme.
Precipitation: the NWS CLI daily precipitation total, preferring the next-morning final report over same-evening preliminaries. A day is "wet" when measurable rain (> 0.005") fell; a trace counts as dry — matching how rain markets settle. Historical truth was backfilled from the official NWS text archive through the identical parser used for live capture.
Forecast leads (precipitation)
Each day is scored at three leads: the most recent model run issued at least 0h / 12h / 24h before the market day began (local-standard midnight). 0h grades the freshest possible run; 24h grades what the model said a full day ahead — the number that matters if you position the day before. A run only qualifies if its forecast hours cover at least 20 of the day's 24 hours; short-range models (HRRR/RRFS) can only qualify via their extended runs, which is why their history is shorter.
How a model's daily rain call is derived
Accumulation models (GFS, ECMWF, ICON, …): hourly precipitation is de-duplicated to one value per hour and summed over the day; the call is "rain" when the total exceeds 0.005". Amount errors (MAE/bias/RMSE, in inches) are also scored for these models.
Rate-based models (HRRR, RRFS): scored wet/dry only — "rain" when any in-window hour shows measurable precipitation. No amount stats.
NWS PoP (NBM-PoP): the National Blend probability-of-precipitation guidance — the statistical basis of the official NWS forecast PoP. The call is "rain" when any 6-hour period overlapping the day carries PoP ≥ 50% (the "more likely than not" reading). A day of persistent 20% chances is therefore graded as "says dry" — at 20% the forecast's own claim is probably not. This threshold makes PoP a deliberately conservative caller: expect a high "Rain Calls Verified" percentage with a lower share of rainy days caught. Scored wet/dry only, and excluded from the consensus vote below.
Comparing across call styles — read this before comparing POD/FAR between models
The three derivation rules above set very different bars for "saying rain." Accumulation models trigger a rain call on any measurable day total — one damp hour counts — so they call rain often and carry more false alarms. NBM-PoP needs 50% confidence, so it calls rain rarely and verifies at a high rate. Neither is cheating; they are different kinds of statements graded on their own terms. The practical rule: use CSI to rank models against each other (it is robust to how chatty a caller is), and read Rain Calls Verified / POD as a description of each model's personality — trustworthy-but-quiet versus alert-but-noisy — rather than as a head-to-head score.
The metrics
Every settled day lands in one of four buckets: hit (called rain, rained), miss (called dry, rained), false alarm (called rain, stayed dry), correct negative (called dry, stayed dry). From those counts:
- Rain Calls Verified (success ratio) = hits ÷ (hits + false alarms) — of the days the model said rain, the share where it rained. The "can I trust a rain call?" number. Gameable by rarely calling rain, so read it with POD.
- POD (probability of detection) = hits ÷ (hits + misses) — of the days it actually rained, the share the model called. The "will it warn me?" number.
- FAR (false alarm ratio) = false alarms ÷ (hits + false alarms) — the complement of Rain Calls Verified.
- CSI (critical success index) = hits ÷ (hits + misses + false alarms) — the ranking metric. It ignores easy correct-dry days entirely, so it can't be inflated in dry climates, and it punishes both crying wolf and sleeping through rain.
- Accuracy = (hits + correct negatives) ÷ all days — intuitive but flattered by dry climates (always saying "dry" in Las Vegas scores ~97%); provided for context only.
- Wet days (base rate) = the share of days in the window with measurable rain — the difficulty of the field the models are playing on.
Consensus & the Brier score
The consensus treats the share of accumulation/rate models calling "wet" (≥3 votes required) as a probability and grades it with the Brier score — the mean squared gap between stated probability and outcome (0 = perfect, lower is better). Its benchmark is the climatology Brier, base_rate × (1 − base_rate): what you'd score by always forecasting the local wet-day frequency. The consensus only has real skill where it beats that number. This is also the fair way to think about probabilities in general: a single 70% forecast that busts is not an error — 70% forecasts should bust 30% of the time; only miscalibration across many days is.
The all-stations aggregate
Per-station samples accumulate slowly (a 30-day window is at most 30 samples), so the ALL view pools every scored station. Contingency counts are summed and the ratios recomputed — exact, not averaged — while amount errors and Brier scores are sample-weighted means. Its unit is station-days (one station scored for one day). Caveat: the pool blends very different climates (Miami's ~40% wet rate with Las Vegas's ~2%), so it answers "is this model good at rain generally" — use the per-station view for "should I trust it in this market."
Temperature specifics
Temperature uses the day's verified high or low. Bias is forecast − actual (positive = model runs warm); Correction is the same error with the sign flipped — the amount to add to a forecast to fix it on average. Colors never flip between the two views: red always means a warm-running model, blue cold-running. Run-time and lead-time tables break the same errors out by model cycle (00Z/06Z/12Z/18Z) and by hours-before-the-extreme. A run is only scored when the actual extreme occurred inside its forecast coverage.
Known limitations
- Precipitation forecast history begins late April 2026 for the global fleet; HRRR-EXT/RRFS precip capture began mid-August 2026 — their samples are still small, and small samples produce flashy but unstable scores. The Days / Stn-Days column is the tell.
- Some stations lack some models entirely (e.g. Central Park has no GFS/NAM/RAP point forecasts), so per-model sample sizes differ within one station.
- Precipitation truth requires an NWS CLI product, so precipitation scoring covers US stations only.
- Preliminary CLI values are replaced automatically when the final report publishes; days marked prelim in the daily grid may still adjust.