The results of the econometric analysis for the period 2000–2024 indicate an increasing role for investment in artificial intelligence and digitalization in supporting economic growth in Saudi Arabia in both the short and long run. These findings are consistent with the strategic objectives of Saudi Vision 2030, which aims to diversify the economic base, enhance total factor productivity, and strengthen the digital economy.
The estimated models reveal that variables associated with digital transformation contribute positively to explaining growth dynamics, reflecting the effectiveness of government digital policies and technological transformation programs in improving economic performance.
Moreover, the forecasting results suggest that strengthening investment in artificial intelligence and digital infrastructure constitutes an important pathway for sustaining economic growth and improving the efficiency of resource allocation and public expenditure. These findings align with the vision’s objective of building a competitive knowledge-based economy driven by innovation and advanced technologies. Accordingly, the results of this study provide empirical support for policymakers to continue expanding digital transformation policies as a key driver of long-term economic growth in Saudi Arabia.
The dataset used in this study covers the period 2000–2024 and includes GDP growth, internet usage, government expenditure, inflation, and patent applications. The full dataset is provided in Appendix A.
Missing values in the patent series for the years 2008, 2009, 2012, 2022, 2023, and 2024 were treated using the linear interpolation method, which is a commonly applied technique in time-series analysis when limited and intermittent gaps exist in the data. This procedure aims to preserve the continuity of the time series and avoid excluding years from the sample, thereby reducing potential estimation bias and improving the efficiency of econometric results when applying regression models or subsequent time-series models.
The study relies on annual data obtained from official national and international sources to ensure reliability and reproducibility. Due to the absence of direct data on artificial intelligence investment, the analysis employs proxy indicators that reflect innovation activity and digital transformation. This approach is widely adopted in contemporary economic literature. Furthermore, the temporal units of the data were standardized, and variables were converted into real values or growth rates when necessary to ensure the validity of econometric estimation and its alignment with the objectives of the study.
15.3. Statistical Analysis
The empirical analysis was conducted using Statistical Package for the Social Sciences (SPSS) software to estimate the econometric models and perform the necessary statistical and diagnostic tests.
Preliminary analysis indicates that all variables become stationary after first differencing, confirming their integration of order one (I(1)) and supporting their suitability for inclusion in dynamic econometric models.
The model fit results indicate a moderate explanatory capacity, as summarized in
Table 2, which reports key goodness-of-fit and forecasting indicators.
The model fit results indicate a moderate explanatory capacity, with the coefficient of determination reaching 0.542, suggesting that the model explains approximately 54.2% of the variation in the dependent variable after accounting for the dynamic characteristics of the time series. The Stationary R-squared, which also equals 0.542, further confirms the model’s ability to explain variations in the data after removing non-stationarity.
Regarding forecasting accuracy, the model reports moderate prediction errors, as reflected by the values of RMSE and MAE, indicating that the model’s forecasts deviate from the actual values within a reasonable range on average. In contrast, relatively high values of MAPE are observed, suggesting that percentage-based error measures are sensitive to the presence of small actual values or sharp fluctuations during certain periods, which reduces the reliability of MAPE as a standalone indicator of forecasting performance.
The Normalized Bayesian Information Criterion (BIC), with a value of 3.770, indicates an acceptable balance between model fit and model complexity, supporting the use of the model in comparative analysis with alternative econometric specifications.
Although the coefficient of determination indicates a moderate explanatory power, the primary objective of this study is not solely to maximize R2, but rather to capture the dynamic temporal relationships and both short-run and long-run interactions among the variables. Moreover, the relatively high values of the percentage error metric (MAPE) can be attributed to the presence of small actual values and sharp fluctuations in certain periods. Therefore, absolute error measures and information criteria—such as RMSE, MAE, and BIC—provide more reliable indicators for evaluating model performance in this context.
The diagnostic statistics of the model are reported in
Table 3, including the Ljung–Box Q test and residual diagnostics.
The model results indicate a moderate explanatory capacity, with the stationary coefficient of determination reported as Stationary . This implies that the model explains approximately 54.2% of the variation in the first difference of the dependent variable using four explanatory variables. In addition, the Normalized Bayesian Information Criterion (BIC = 3.770) reflects an acceptable balance between model fit and model complexity.
With regard to residual diagnostics, the Ljung–Box Q test at lag 18 reports a statistically significant result , indicating the presence of some remaining autocorrelation in the residuals at certain lag lengths. However, no outliers were detected .
Overall, these findings suggest that the model captures a substantial portion of the underlying temporal dynamics, although minor improvements in the model specification may still be possible. Such improvements could involve introducing additional lag terms or adjusting the model order to further reduce residual autocorrelation and enhance the model’s dynamic representation.
The estimated parameters of the ARIMA model are presented in
Table 4, highlighting the autoregressive and moving average components.
The estimation results of the ARIMA model for the first difference of the dependent variable indicate that the intercept term is statistically insignificant . This suggests the absence of a constant mean in the differenced series, a pattern consistent with the statistical properties of stationary time series after first differencing.
Regarding the dynamic components, both the first-order autoregressive coefficient (AR(1)) and the moving average coefficient (MA(1)) are statistically insignificant, with probability values of and , respectively. These findings suggest a limited degree of serial dependence in the current changes of the series after incorporating the explanatory variables, implying that the internal time-series structure is relatively weak compared with the influence of external factors.
With respect to the explanatory variables, the results indicate that most variables are not statistically significant at conventional significance levels. The only variable approaching statistical significance is , with a probability value of and a negative coefficient, suggesting a possible short-run effect of this variable on changes in the dependent variable, although it does not reach the conventional 5% significance level. In contrast, , , and do not exhibit statistically significant short-run effects.
Overall, these findings highlight several important insights:
The short-run dynamics appear relatively weak when using an ARIMA specification with exogenous variables expressed only in differenced form.
The results suggest that economic relationships may not be efficiently captured by the ARIMA framework alone, particularly when multiple explanatory variables are involved.
This observation strengthens the methodological motivation for employing more flexible econometric models, such as the ARDL framework, which is capable of capturing both:
- ○
short-run and long-run effects, and
- ○
long-run equilibrium relationships (cointegration) among variables.
The residual behavior of the model is examined through autocorrelation analysis, as illustrated in
Figure 1.
The Residual ACF and Residual PACF plots indicate that most autocorrelation coefficients of the residuals fall within the confidence bounds, with only a few isolated values exceeding the limits at certain lag lengths. These exceedances do not follow a systematic pattern nor exhibit a gradual decay structure. This behavior suggests the absence of a systematic autocorrelation structure in the residuals and supports the assumption that the residuals behave approximately as white noise.
This finding is consistent with the previously reported diagnostic tables, where both the Residual ACF Summary and Residual ACF Table show that autocorrelation coefficients fluctuate around zero without persistence across lag periods. This pattern confirms that the observed exceedances are sporadic rather than systematic, and therefore do not indicate a general misspecification of the model.
Although the Ljung–Box Q(18) test indicates statistical significance at the 5% level, this result can be interpreted as evidence of limited residual autocorrelation that does not substantially undermine the adequacy of the model, particularly given the moderate length of the time-series sample.
Combined with the Model Fit indicators, including Stationary and Normalized , The results suggest that the model achieves an acceptable level of explanatory fit and successfully captures the main temporal dynamics of the series. Any remaining minor autocorrelation effects are unlikely to materially affect the overall conclusions of the study.
A comparison between observed and fitted values is presented in
Figure 2, demonstrating the model’s ability to capture the overall trend.
The figure presents a comparison between the observed values and the model’s fitted values. The fitted line generally follows the overall pattern of fluctuations over time with a reasonable degree of accuracy, indicating that the model is capable of capturing the main movements of the series after first differencing.
However, some deviations are observed at sharp peaks and troughs, particularly during periods of economic shocks. These discrepancies reflect the model’s limited ability to fully capture extreme short-run fluctuations in the data.
These observations are consistent with the previous residual diagnostic results (ACF/PACF and the Ljung–Box test), which suggested the presence of minor and irregular residual autocorrelation. Overall, the figure supports the adequacy of the model for general explanatory purposes and methodological comparison, while acknowledging that more flexible econometric models may further improve forecasting accuracy, particularly in periods characterized by extreme values or sudden shocks.
The descriptive statistics of the transformed variables are reported in
Table 5, providing insights into their central tendency and dispersion.
The descriptive statistics of the variables after applying the first-difference transformation indicate that the mean values of all series—except for DIFF(X4,1)—are close to zero. This pattern reflects the absence of a systematic trend in the annual changes and is consistent with the statistical properties of a stationary time series.
The standard deviation values reveal noticeable variation in the volatility of the variables. In particular, DIFF(Y,1) and DIFF(X1,1) exhibit relatively higher levels of fluctuation compared with the other variables, indicating greater short-run variability in these series.
In contrast, the variable DIFF(X4,1) records a relatively higher positive mean along with a larger standard deviation, suggesting sharper changes and wider fluctuations over time. This behavior may reflect differences in the measurement scale or the intrinsic characteristics of this variable compared with the other series.
Finally, the sample size remains constant across all variables (n = 24), ensuring consistency in statistical comparisons and enhancing the reliability of the descriptive analysis across the dataset.
The pairwise relationships among the variables are examined using Pearson correlation coefficients, as shown in
Table 6.
The Pearson correlation coefficients among the variables after applying the first-difference transformation indicate varying degrees of short-run linear relationships. A moderate positive correlation is observed between DIFF(Y,1) and DIFF(X1,1) , suggesting that both variables tend to move in the same direction during short-run fluctuations.
In contrast, DIFF(Y,1) exhibits a moderate negative correlation with DIFF(X2,1) , indicating a clear inverse relationship in contemporaneous changes.
The variables DIFF(X3,1) and DIFF(X4,1) do not display statistically significant correlations with the dependent variable, as their correlation coefficients are weak and statistically insignificant. This suggests the absence of a direct short-run linear relationship with the dependent variable.
Regarding the relationships among the explanatory variables, statistically significant negative correlations are observed between DIFF(X1,1) and DIFF(X2,1) , as well as between DIFF(X1,1) and DIFF(X4,1) . Nevertheless, the magnitude of these correlations remains below the commonly accepted thresholds associated with severe multicollinearity.
It should be noted that the reported correlations reflect short-run linear associations only and do not imply causality. Consequently, concerns related to multicollinearity remain limited and are further addressed within the dynamic modeling framework adopted in the study.
The overall regression performance is summarized in
Table 7, which reports the model’s explanatory power and goodness-of-fit measures.
The model summary results indicate a moderate overall correlation between the explanatory variables and the dependent variable, with a reported correlation coefficient of . The coefficient of determination suggests that the model explains approximately 36.3% of the variation in the dependent variable. After adjusting for the number of explanatory variables and the sample size, the Adjusted decreases to 0.229, indicating a moderate explanatory capacity once model complexity is taken into account.
The change statistics further show that the inclusion of the explanatory variables improves the overall model fit. However, the F-change test is only marginally significant, with . This result indicates statistical significance at the 10% level, but not at the conventional 5% significance level. Such findings suggest that the joint effect of the explanatory variables becomes more evident in the short run, while the overall explanatory strength remains somewhat limited given the available sample size.
Overall, the model exhibits moderate explanatory power, with joint effects that are marginally significant, reflecting both the constraints of the sample size and the short-run dynamics characterizing the dataset.
The statistical significance of the regression model is evaluated through the ANOVA results presented in
Table 8.
The ANOVA results indicate that the overall explanatory model accounts for a portion of the variation in the dependent variable. The regression sum of squares (SSR = 231.247) is observed relative to the error sum of squares (SSE = 406.342), reflecting the share of variance explained by the model compared with the unexplained residual variation.
The F-statistic with a p-value of 0.062 suggests that the model exhibits marginal statistical significance at the 10% level, although it does not reach the conventional 5% significance threshold. This result indicates that the explanatory variables collectively contribute to explaining the variation in the dependent variable, albeit with moderate explanatory strength.
This outcome can partly be attributed to the limited sample size (n = 24) and the characteristics of time-series data, which may reduce the statistical power of the test. Accordingly, the ANOVA results support the use of the model for interpretative analysis and methodological comparison, while suggesting caution when relying on the model for precise forecasting purposes.
The estimated regression coefficients and their statistical significance are reported in
Table 9.
The estimated Ordinary Least Squares (OLS) regression model can be expressed as follows:
Based on the estimated coefficients, the empirical regression equation is:
where:
represents the economic growth rate,
denotes the percentage of internet users, reflecting digital infrastructure and technology diffusion,
represents government expenditure as a percentage of GDP,
denotes the inflation rate,
represents resident patent applications, reflecting technological innovation,
represents the error term.
The estimated coefficients capture the marginal effect of each explanatory variable on economic growth, holding the other variables constant.
The estimated regression coefficients indicate that the intercept term is not statistically significant , suggesting the absence of a constant mean in the changes observed after the first-difference transformation. This result is consistent with the statistical properties of stationary time series.
Regarding the explanatory variables, the results show that DIFF(X1,1), DIFF(X3,1), and DIFF(X4,1) do not exhibit statistically significant effects on the dependent variable, as their corresponding p-values exceed the conventional significance thresholds. This finding suggests that these variables do not exert a meaningful short-run impact on the dependent variable within the estimated regression framework.
In contrast, the variable DIFF(X2,1) exhibits a negative effect with marginal statistical significance . This result suggests that an increase in the short-run changes of this variable is associated with a decline in the corresponding changes of the dependent variable.
Furthermore, the standardized coefficient indicates that DIFF(X2,1) represents the most influential explanatory variable within the model in relative terms, despite its statistical significance lying just above the conventional 5% threshold. This finding implies that the variable may exert a notable short-run influence on the dependent variable, although the strength of this relationship should be interpreted with caution, given its borderline statistical significance.
The coefficient estimates reflect short-run effects only and should be interpreted with caution, particularly given the small sample size and the marginal significance of some parameters.
Multicollinearity diagnostics, including VIF and tolerance values, are presented in
Table 10.
The partial and semi-partial (part) correlation coefficients indicate that the variable DIFF(X2,1) exhibits the strongest partial association with the dependent variable after controlling for the influence of the remaining explanatory variables. This suggests that DIFF(X2,1) contributes relatively more to explaining the variance in the dependent variable, although the magnitude of this effect remains moderate.
In contrast, the variables DIFF(X1,1), DIFF(X3,1), and DIFF(X4,1) display relatively weak partial correlations, indicating limited independent short-run effects within the model.
With respect to multicollinearity diagnostics, the Tolerance values remain within acceptable levels (greater than 0.1), while the Variance Inflation Factor (VIF) values remain below the commonly accepted critical threshold (VIF < 5). These results indicate the absence of severe multicollinearity among the explanatory variables.
Therefore, the limited statistical significance of some coefficients cannot be attributed to multicollinearity, but rather to the limited short-run explanatory effects of certain variables or the constraints imposed by the sample size.
The correlations among the estimated regression coefficients are shown in
Table 11, providing insights into parameter stability.
The Coefficient Correlations table presents the correlations among the estimated regression coefficients, rather than the correlations among the original variables. This diagnostic is useful for evaluating the stability and reliability of the estimated parameters within the regression model.
The results indicate moderate correlations between some coefficients, most notably the positive correlation between the coefficients of DIFF(X1,1) and DIFF(X2,1) , as well as between DIFF(X1,1) and DIFF(X4,1) . These relationships suggest a partial overlap in the information captured by these variables within the model specification.
In contrast, relatively weak or negative correlations are observed among several other coefficients, particularly between DIFF(X3,1) and the remaining explanatory variables, indicating a relatively independent contribution of this variable within the regression framework. In addition, the covariance matrix shows relatively small covariance values, supporting the stability of the estimated parameters and indicating that the standard errors of the coefficients are not excessively inflated.
Taken together with the previously reported VIF and Tolerance diagnostics, these results suggest that the correlations among the coefficients do not reach critical levels that would threaten the stability of the estimation or bias the results. Rather, the limited statistical significance observed for some coefficients is more likely attributable to the short-run nature of the relationships and the relatively small sample size, rather than to severe multicollinearity among the explanatory variables.
Additional collinearity diagnostics based on eigenvalues and condition indices are reported in
Table 12.
The multicollinearity diagnostic results based on Eigenvalues and the Condition Index indicate that no severe multicollinearity problem exists in the model. All Condition Index values remain well below the commonly accepted critical thresholds (typically 20 or 30), with the highest reported value reaching CI = 2.86. Such a low value suggests a high degree of stability in the estimated regression coefficients.
With respect to Variance Proportions, no dimension exhibits simultaneously high variance proportions (greater than 0.50) for more than one explanatory variable in conjunction with a high Condition Index. This condition is generally considered the primary indicator of severe multicollinearity. Although some variables—such as DIFF(X1,1), DIFF(X2,1), and DIFF(X4,1)—display relatively higher variance proportions in the fifth dimension, the accompanying low Condition Index value confirms that this overlap is not statistically problematic.
Overall, these findings suggest that the model coefficients are stable and that the limited statistical significance observed for some variables cannot be attributed to multicollinearity. Rather, it is more likely associated with the short-run nature of the relationships or the constraints imposed by the relatively small sample size.
The distribution of residuals and prediction errors is summarized in
Table 13.
In addition, the standardized values of both the predicted values and the residuals exhibit relatively stable distributions, which supports the adequacy of the model with respect to the underlying statistical assumptions.