Evaluation of Forecasting Models for District Heating

Comparing a storage-aware surrogate with operational replay on identical forecast outputs to determine whether the two methods select the same model.

Forecasting model families
6
Random seeds
30
Shared evaluation weeks
104

Data and methods

  • The models were trained and tuned on data from 2014–2022, then tested on data from 2023–2024.
  • The comparison covers six forecasting model families, 30 random seeds, and 104 common evaluation weeks.
  • Each set of forecasts was assessed using both a storage-aware surrogate and an operational replay that accounts for unit commitment and state transitions.

Results

  • HGB-r3 ranked sixth under surrogate evaluation but first under operational replay.
  • When model selection was based on the first 52 weeks, the two methods selected different models for every one of the 30 random seeds.
  • Over the subsequent 52 weeks, the selected models differed by an average of 5.414 weighted cost units per time step.

Status

The manuscript, “Forecast Model Selection for District Heating: Storage-Aware Surrogate Evaluation versus Operational Replay,” is currently under review.