4. Results
The first step of the analysis was the visual assessment of the forecasting error series. The visual analysis of the forecasting error time series consists in plotting those errors on the timeline. Such diagrams may reveal the existing patterns, including cyclicality, seasonality or the trend, which were not visible in the analysis of the forecast value time series. For example, if the series of forecasting errors display regular fluctuations in specific periods of time, this may suggest that the forecasting model has difficulties predicting certain seasonal patterns. The behavior of the analyzed time series is presented in
Figure 3.
The visual analysis of the forecasting error times series is an important stage of the forecasting model analysis. By means of the reliable understanding of error series patterns and properties, the researchers and analysts may identify significant relationships and aspects which are worth analyzing further and in more detail. Such an approach allows to understand the forecasting error dynamics and potential model-related problems better. Following the visual analysis, it is possible to carry out a more advanced statistical analysis. Calculations of the basic parameters of the forecasting error distribution, including the mean value, standard deviation or skewness, may provide information on the error characteristics and asymmetry. Moreover, the STL decomposition (Seasonal and Trend decomposition using Loess) allows to determine the components of the trend, seasonality and the remainder which may help to identify the major sources of errors in the forecasts. Statistical hypothesis testing plays an important role in the analysis. Determination of the p-value for the tests with the hypothesis concerning the absence of any trend or seasonality allows to find out whether there are any statistically significant deviations from those assumptions. The basic numerical characteristics of the analyzed error time series are presented in
Table 4.
The analysis of the forecasting error time series for different channels revealed diversified error patterns and characteristics in those channels. Some channels tend to overestimate, whether others to underestimate the forecast values. The differences of the standard deviation, coefficient of variation, skewness and kurtosis indicate diverse error variation. For every channel, the analysis of those parameters may offer valuable guidelines for further optimization and improvement of forecasting models. For Channel_01, the mean value of error is 172, whereas the median is 176, which suggests that most errors are below the mean value. However, the asymmetry coefficient value indicates that there is a poor asymmetry of error distribution. However, high standard deviation (1,095) and high value of the coefficient of variation (CV = 6.212) point to high error variation. For Channel_02, the mean error is 245, whereas the median is 507, suggesting that the models tend to underestimate the predicted values. High standard deviation (2,268) and kurtosis (11.280) indicate significant variation of the error distribution. Analyzing Channel_03, a conclusion can be drawn that the mean error is close to zero, but low median (15.5) and high standard deviation (174) indicate diverse error characteristics. Skewness is close to zero and kurtosis (4.387) proves higher value concentration than in the normal distribution (kurtosis is 0). For Channel_04, the mean error is 574, whereas the median is 297, suggesting the underestimation of the predicted values. High standard deviation (5,665) and kurtosis (2.270) indicate significant error variation and a certain degree of the analyzed values dispersion. The distribution is right-skewed. The mean error in Channel_05 is 126 and the median is 149, suggesting small value undervaluation. High standard deviation (1,017) and the coefficient of variation (8.040) indicate significant variation. The distribution is left-skewed. For Channel_06, the mean error is 25, whereas the median is 29.5, suggesting small value underestimation. Low standard deviation (52) and kurtosis (0.567) indicate relatively low variation and the distribution close to normal. The distribution is left-skewed. The mean error for Channel_07 (-1,387) and the median (-687) are negative, suggesting the tendency to overestimate the predicted values. High standard deviation (2,822) and kurtosis (-0.487) indicate significant error variation and platykurtic distribution. The distribution is left-skewed. Channel_08 is characterized by the mean error of -706 and the median -175, suggesting overestimation of the predicted values. High standard deviation (1,602) and kurtosis (-0.774) indicate certain error variation and platykurtic distribution. The distribution is left-skewed. For Channel_09, the mean error (-1,583) and the median are negative (-732), suggesting overestimation of the predicted values. High standard deviation (2,969) and kurtosis (-0.726) indicate significant error variation and platykurtic distribution. The distribution is left-skewed. For Channel_10, the mean error is -228, whereas the median is -147, suggesting value overestimation. High standard deviation (420) and kurtosis (-0.354) indicate error variation. The distribution is left-skewed.
Generally speaking, the value of the coefficient of variation (CV = Std.Dev / Mean) indicates high variation in the analyzed error distributions.
The subsequent analytical step was to analyze the forecasting error randomness. The results are presented in
Table 5
The analysis of the forecasting error randomness indicates that each analyzed series can be considered random (in the sense of one of the tests used and alpha = 0.05). Moreover, low p-values for Channel_02, Channel_07, Channel_09 and Channel_10 in some tests may suggest the presence of certain irregularities in the error behavior.
The analysis of the stationarity of the analyzed error series using ADF (Augmented Dickey–Fuller test) indicates that the series may be considered stationary (p-value <= 0.01 for every series). The results of the series autocorrelation analysis are not homogeneous and may point to irregularities. The detailed values of coefficients and critical significances (p-values) for the first seven delays are presented in
Table 6. For the test, the values of ACF coefficients and Ljung-Box test were used.
According to the initial analyses, the regularities may refer to each analyzed series. In every analyzed case, the autocorrelation is present for the first seven rows.
The results of the analysis of the forecasting error time series are presented in
Table 7,
Table 8 and
Table 9.
Columns in
Table 7 contain the following information:
“Trend_stl” – the value determined using the equation (3), informing about the strength of the trend component in the STL decomposition (the closer it is to 1, the higher the significance of the trend in the error is), “Season_stl” – the value determined using the equation (4), informing about the strength of the seasonality component in the STL decomposition (similar to the preceding value, the closer it is to 1, the higher the significance of the component in the error is), “MAE_error” – value of the MAE error (1) for the product, “MAPE_error” – value of the MAPE error (2) for the product, “Remainder_MAE_stl” – “non-systematic” error understood as the MAE value for the error series, calculated for the remainder component in the STL decomposition (the mean of the absolute values of the remainder component in the error series), informing about MAE error excluding the systematic components of the error series, “Quotient_stl” – relative “non-systematic” error, understood as the quotient of “Remainder_MAE_stl” and “MAE_error”, informing what part of the general MAE error is taken by MAE, calculated solely based on the remainder component of STL decomposition.
The data in the table were ordered based on the non-decreasing values of the measure (4) determining the strength of the seasonality component in the error series. In STL decomposition, the frequency of 7 was assumed for every analyzed series as the operator works 7 days a week and the data refers to daily values. The results in
Table 7 do not show direct, strong and unambiguous relationships between the values. Solely, (Pearson’s) correlations between the following values can be considered significant (alpha = 0.05):
Between the strength of the trend component (Trend_stl) and the strength of the seasonal component (Season_stl), r = 0.59 (t = 2.426, p = 0.034). The more significant the trend component is, the higher the significance of the seasonal component.
Between the strength of the trend component (Trend_stl) and the relative “non-systematic” error (Quotient_stl), r = -0.69 (t = -3.163, p = 0.009). The more significant the trend component in the errors is, the smaller the error relating to the exclusion of that component.
Between the strength of the seasonal component (Season_stl) and the relative “non-systematic” error (Quotient_stl), r = -0.70 (t = -3.251, p = 0.007). The more significant the seasonal component, the smaller the “non-systematic” error.
Between the “non-systematic” error (Remainder_MAE_stl) and MAE error (MAE_error), r = 0.88 (t = 6.185, p < 0.001). The higher the absolute error, the higher the absolute “non-systematic” error. Generally speaking, this relation can be deemed obvious.
Referring to section one, attention should be paid to the fact that the maximum value of the indicator (3) in the analyzed series is 0.158 and, generally speaking, proves small strength of the trend component in the analyzed error series. In just two cases, the strength of the trend component is higher than the strength of the seasonal component (Channel_02, Channel_09). In the analyzed problem, the seasonal component of the error series is more significant.
The numerical aspects relating to the method of identifying systematic components using STL method should be emphasized. The general determined trend is not linear and decomposition parameter changes can be used to control trend variation. Moreover, this is closely connected with the seasonal component with a simultaneous absence of any impact on the remainder component. From this perspective, systematic components should be analyzed jointly. For the pre-determined decomposition parameters, the systematic components are naturally correlated. This means that the correlations in sections two and three should be considered natural.
Despite a general low strength of the trend component, the results of the trend presence analysis using Student’s t-test, Mann–Kendall test, WAVK test (Lyubchich V. et al. 2023) point to an important trend presence in most analyzed series. The detailed results are presented in
Table 8.
The results presented in
Table 8 point to the trend presence for the forecasting errors in channel_02, channel_07, channel_08 and channel_09. However, the visual assessment of the phenomenon in the function of time does not confirm any clear trend.
The following tests were used to analyze a significant seasonal component in the analyzed time series : combined.kwr - Ollech and Webel’s combined seasonality test (Ollech, D., Webel, K., 2020), test QS (qs.p), Friedman Rank test (fried.p), Kruskall Wallis test (kw.p), F-Test on seasonal dummies (seasdum.p) and Welch seasonality test (welch.p).
The results of the tests carried out point to clear seasonality in error series referring to channel_10, channel_07 and channel_08. For channel_09, low p-value suggest possible presence of significant seasonality as well. The results are consistent with those from the analysis of the strength of seasonality (4).
Figure 4 and
Figure 5 present the visualized decompositions performed for two extreme examples.
Figure 4 presents error decomposition for channel_03 characterized by the lower share of systematic components in the overall error.
Figure 5 depicts error decomposition for channel_10 characterized by the highest share of systematic components.
The major difference of the systematic components’ strength lies in the error scale. For channel_03, the trend ranges from ca. -30 to ca. 10, seasonality from ca. -56 to ca. 23, whereas the overall error from -786 to 740. For channel_10, the trend ranges from ca. -340 to ca. -23, seasonality from ca. -457 to ca. 253, whereas the overall error from -1,388 to 681. This means that the error decomposition visualization can also be used to assess the strength and significance of the systematic error components. It should also be stressed that a key indicator here can be the range of individual component changes.