WEATHERMAPS.AI · THE AI MODEL DESK OF THE FORECAST NETWORK FORECAST EAGLE FORECAST EURO WEATHER.CITY WEATHER.FOOTBALL
weathermaps.ai
Skill · weathermaps.ai/skill-v1 · last 30 days

Which model was closest to what happened?

Every model's 2 m temperature forecast for 497 major cities, checked against the temperature reported by the nearest airport (within 30 km) at 00, 06, 12 and 18Z. Every model is scored on the same cases. Lower is better.

Day 1Day 2Day 3Day 5Day 7Day 10
AIFS 1.18n 1342 0 so far 0 so far 0 so far 0 so far 0 so far
HRES 1.24n 1342 0 so far 0 so far 0 so far 0 so far 0 so far
WN2 1.22n 1342 0 so far 0 so far 0 so far 0 so far 0 so far
ENS mean 1.25n 1342 0 so far 0 so far 0 so far 0 so far 0 so far
Error without station biasThe RMSE after removing each model's average error at each airport. A grid box that sits higher than the runway is steadily cold there; this score does not count that against the model.
RMSERoot-mean-square error in °C, bias included. Favours the finer grid (HRES, 0.1°), which is closer to each airport's real height.
BiasAverage forecast minus observation. Negative means the model ran cold. The outlined cell is the value closest to zero.

How much has been scored

Day N needs runs that are at least N days old, so the right-hand columns fill in over the first two weeks. Observations have been archived since 16 Sep 06Z.

Observation hours400, 06, 12, 18Z · METAR
Runs archived200 and 12Z, every model
Cities with an airport497 / 602nearest METAR within 30 km
Minimum per cell30fewer cases shows a dash

What this can and cannot tell you

An airport is not a city, and a grid box is not an airport. The comparison is fair between models because every model is scored on exactly the same cases, but no single number here is "the error of the forecast for that city".

The forecasts are rounded to whole degrees before they reach this site, which adds about 0.3 °C to every RMSE equally. Thirty days is a short sample: a difference of a few hundredths between two models is noise.

The ensemble mean usually scores best at longer lead times. That is expected, not a finding: averaging many forecasts removes the detail that cannot be predicted, which lowers squared error. Its spread is what the Disagreement Index is measured against.