ArXiv

Model selection with proper scoring rules on data sets of time series

Authors
Giorgio Corani, Stefano Damato, Dario Azzimonti...
Categories
stat.ML, cs.LG
arXiv
https://arxiv.org/abs/2606.24715v1
PDF
https://arxiv.org/pdf/2606.24715v1

Brief

Model selection for probabilistic time-series models is analyzed by comparing three aggregations of per-series proper-scoring-rule values (mean, median, mean rank). The authors trace conflicting selection outcomes to skewness in score distributions, prove that longer per-series test sets drive the criteria to agreement, and report that for short tests the mean score best recovers the true model; experiments use intermittent series including the M5 dataset.

Why it matters

Giorgio Corani, Stefano Damato, Dario Azzimonti, and Lorenzo Zambon (arXiv:2606.24715v1, published 2026-06-23) show that aggregating per-series proper-scoring-rule values by mean, median, or mean rank can yield conflicting model-selection decisions due to skewness in the distribution of per-series scores.

Key details

  • As per-series test length increases the three aggregation criteria (mean score, median score, mean rank) progressively converge; for short test sets only the mean score reliably identifies the true model as best. Experiments on intermittent time series, including the M5 competition dataset, also find mean-rank selection invariant to scaling factors.
Source evidence

Abstract

We consider the problem of model selection between probabilistic models on data sets of time series. Chosen a proper scoring rule, we denote by the term \textit{score} the average value of the scoring rule on the test of an individual time series. For model selection, we need aggregating the values of the scores across multiple time series. Three summary statistics are commonly used for model selection: mean score, median score, and mean rank. Results in previous papers show that these statistics can yield conflicting decisions; we show how the conflicting conclusions are due to the skewness of the distribution of scores. We also show that as the test set of each time series of the data set increases, the different model selection criteria progressively converge to the same conclusion. However, for short tests sets, only the mean score identifies the true model as the best. We illustrate these phenomena with an analysis on intermittent time series, including the data set of the M5 competition, where we underline the importance of having a large test set. In such experiments, we further notice that model selection based on mean ranks remains unchanged using different scaling factors.