Which model was closest to what happened?
Every model's 2 m temperature forecast for 497 major cities, checked against the temperature reported by the nearest airport (within 30 km) at 00, 06, 12 and 18Z. Every model is scored on the same cases. Lower is better.
| Day 1 | Day 2 | Day 3 | Day 5 | Day 7 | Day 10 | |
|---|---|---|---|---|---|---|
| AIFS | 1.18n 1342 | —0 so far | —0 so far | —0 so far | —0 so far | —0 so far |
| HRES | 1.24n 1342 | —0 so far | —0 so far | —0 so far | —0 so far | —0 so far |
| WN2 | 1.22n 1342 | —0 so far | —0 so far | —0 so far | —0 so far | —0 so far |
| ENS mean | 1.25n 1342 | —0 so far | —0 so far | —0 so far | —0 so far | —0 so far |
| Day 1 | Day 2 | Day 3 | Day 5 | Day 7 | Day 10 | |
|---|---|---|---|---|---|---|
| AIFS | 1.78n 1342 | —0 so far | —0 so far | —0 so far | —0 so far | —0 so far |
| HRES | 1.91n 1342 | —0 so far | —0 so far | —0 so far | —0 so far | —0 so far |
| WN2 | 1.92n 1342 | —0 so far | —0 so far | —0 so far | —0 so far | —0 so far |
| ENS mean | 1.92n 1342 | —0 so far | —0 so far | —0 so far | —0 so far | —0 so far |
| Day 1 | Day 2 | Day 3 | Day 5 | Day 7 | Day 10 | |
|---|---|---|---|---|---|---|
| AIFS | -0.52n 1342 | —0 so far | —0 so far | —0 so far | —0 so far | —0 so far |
| HRES | -0.57n 1342 | —0 so far | —0 so far | —0 so far | —0 so far | —0 so far |
| WN2 | -0.77n 1342 | —0 so far | —0 so far | —0 so far | —0 so far | —0 so far |
| ENS mean | -0.72n 1342 | —0 so far | —0 so far | —0 so far | —0 so far | —0 so far |
How much has been scored
Day N needs runs that are at least N days old, so the right-hand columns fill in over the first two weeks. Observations have been archived since 16 Sep 06Z.
What this can and cannot tell you
An airport is not a city, and a grid box is not an airport. The comparison is fair between models because every model is scored on exactly the same cases, but no single number here is "the error of the forecast for that city".
The forecasts are rounded to whole degrees before they reach this site, which adds about 0.3 °C to every RMSE equally. Thirty days is a short sample: a difference of a few hundredths between two models is noise.
The ensemble mean usually scores best at longer lead times. That is expected, not a finding: averaging many forecasts removes the detail that cannot be predicted, which lowers squared error. Its spread is what the Disagreement Index is measured against.