4. Testing Results
This section empirically estimates the national tax revenue model and forecasts projected estimates empirically. In the empirical analysis, effective methodologies will be used to improve the performance of the regression analysis model and enhance the accuracy of numerical predictions despite insufficient data.
First, during the data processing stage, outliers were removed, missing values were handled by replacing them with the mean, and scale differences of input variables were reduced to improve the model’s performance, stability, and convergence speed. Second, in the variable selection stage, correlation coefficients and Variance Inflation Factors (VIF) were checked to remove unnecessary variables with low correlations or high multicollinearity, thereby improving the model’s performance. New variables were created through combinations of existing variables to enhance predictive power. Additionally, categorical variables were converted into dummy variables and interaction terms were added by multiplying variables to further increase prediction accuracy. Third, during the model estimation stage, stable equilibrium relationships were utilized by verifying unit roots and cointegration according to time-series econometric methods to identify optimal parameters. Furthermore, results of estimating static and dynamic cointegrated regression models were compared to select the optimal estimate. Fourth, in the prediction stage, dynamic predictions were used to supplement static forecasts made by experts. Data were divided into in-sample and out-of-sample segments. Both in-sample and out-of-sample predictions were conducted simultaneously to improve the prediction performance.
A time-series stability test was conducted to assess the stability of variables needed for the empirical analysis of the tax revenue model. The stability of the time-series variables used in the analysis was checked through unit root tests.
Table 2 shows results of the (Augmented) Dickey-Fuller unit root test. As a result of the Wald test for the unit root model, for most variables, we could not reject the null hypothesis that there is a unit root at the 5% significance level. It was found that unstable level variables became stationary after first differencing. Therefore, it is necessary to test whether a cointegrating relationship exists among level variables.
To confirm the long-term equilibrium relationship between dependent and explanatory variables in individual tax models, a cointegration test was conducted for the regression model below. The equilibrium relationship between time series variables to be used in the regression model was verified by cointegration tests.
Table 3 presents results of cointegration tests proposed by [
13,
14]. Results of testing the linear trend model with structural breaks, as proposed by [
14], using the appropriate lag based on AIC, showed that for all individual tax models, the null hypothesis of the presence of a unit root in the residuals (indicating no cointegration) was rejected at the 5% significance level. Therefore, cointegration regression analysis was performed using level variables of tax models.
As it was confirmed earlier that individual variables in tax models exhibited unit roots without showing a cointegration relationship, a Granger causality test was first conducted to estimate the cointegration regression model. Based on a four-lag specification of explanatory variables, results of the test rejected the null hypothesis of no causality for most individual tax models, assuming the existence of causal relationships.
Now that a stable cointegration relationship was established, static and dynamic OLS estimators for regression models were used to estimate tax functions and discuss results. Estimated regression equations, cointegration regression model estimation results, and projections on tax revenue were discussed by separating them into individual tax categories.
From now on, we will specify the cointegration tax revenue regression model, assess its robustness, estimate it, and make predictions. We performed optimal estimation on cointegration regression models for each tax category and then made forecasts for tax revenue.
Analysis procedures were as follows. First, when determining the tax revenue estimation model, if there are many independent variables, overfitting will reduce errors but increase the variance of parameters, leading to lower prediction reliability. Conversely, if there are fewer independent variables, underfitting will increase errors (bias), although the standard deviation of parameters is lower. Therefore, endogenous variables were adjusted to reduce errors and increase reliability based on the economic theory.
Second, since time series data changed over time, the method to reduce prediction errors based on past information (learning) used the prediction that minimized expected value of the loss function. For optimal prediction, the tax revenue estimation error (= actual settlement value − budget forecast value) was used as a control variable for estimation.
Third, the method of adjusting the tax revenue loss function aimed to improve the model’s prediction performance by processing and utilizing the data according to the analysis objective and characteristics of the time series data, considering trend, seasonality, and outlier removal. Additionally, regularization techniques were used to prevent excessive weights in the linear regression model’s cost function. Regularization could control the model’s complexity, reduce data dimensions, and mitigate the risk of overfitting.
Fourth, the ideal tax revenue prediction strategy differed by country but depended on the availability of data information. It is important to predict each individual tax, as each tax category contributes differently to the forecast based on its taxable base. Since Korea’s estimation errors arose from asymmetric errors across tax categories, utilizing information from tax categories with significant weight and errors could improve prediction accuracy. However, predicting tax revenue is challenging due to complex interactions between economic fluctuations and government policies. Therefore, optimal forecasting values estimated were compared with those from a random simulation considering shocks as a forecasting method for single time series.
4.1. Income Tax
The estimation model for income tax revenue was divided into labor income tax, comprehensive income tax, capital gains tax, and interest and dividend income tax as shown below. Through principal component analysis and factor (covariance-correlation matrix) analysis, variables having low correlations with tax revenue were eliminated to reduce dimensionality of data. Multicollinearity checks were then performed to exclude explanatory variables having high correlations from the analysis. The estimated equation for level variables with realized prediction error (= actual value − realized forecast value) as a control variable is presented along with estimation results.
4.1.1. Labor Income Tax
Estimation results for level variable equation of labor income tax after adjustment considering key factors such as multicollinearity were obtained from y = f(X
1, X
2), where y was the labor income tax, X
1 was the wage, X
2 was the forecasting error of labor income tax, and f(.) was the functional form. Based on estimated values of the labor income tax cointegration model, in-sample labor income tax revenue predictions reflecting basic forecast values of labor income tax explanatory variables are shown in
Table 4.
The comparison of model forecasting accuracy using RMSE for in-sample labor income tax revenue forecasts showed that, compared to the government’s forecast of 1,573,702 million KRW for the labor income tax, OLS had a forecast of 13,056,065 million KRW, DOLS had a forecast of 12,079,926 million KRW, and FMLS had a forecast of 12,634,118 million KRW. These results revealed that the government’s prediction method had the highest forecasting accuracy, followed by the DOLS method.
For the next year’s tax revenue, excluding the OLS method due to endogeneity issues, results comparing the DOLS method (which had the second highest forecasting accuracy after the government’s forecast) with other estimation techniques are shown in
Table 5.
The forecast for labor income tax revenue in 2025 was estimated at 74,227,332 million KRW using the DOLS method and 68,123,375 million KRW using the dynamic model DOLS method. The actual time series forecast (random simulation) of 74,432,574 million KRW served as a reference estimate. However, the forecast from the labor income tax cointegration regression model could be sensitive to assumptions regarding predicted values of the explanatory variable such as salary, the chosen estimation model, the chosen estimation method, and shocks related to labor and labor income.
4.1.2. Comprehensive Income Tax
Estimation results for the level variable equation of comprehensive income tax after adjustment considering key factors such as multicollinearity were obtained from y = f(X
1, X
2), where y was the comprehensive income tax, X
1 was the business income and real estate rental income, X
2 was the forecasting error of comprehensive income tax, and f(.) was the functional form. Based on estimated values of the labor income tax cointegration model, in-sample comprehensive income tax revenue predictions reflecting the basic forecast values of the comprehensive income tax explanatory variables are shown in
Table 6.
The comparison of model forecasting accuracy using RMSE for in-sample comprehensive income tax revenue forecasts showed that, compared to the government’s forecast of 2,870,935 million KRW, OLS had a forecast of 4,378,596 million KRW, DOLS had a forecast of 16,559,414 million KRW and FMLS had a forecast of 10,932,554 million KRW. These results showed that the government’s prediction method had the highest forecasting accuracy, followed by FMLS and then DOLS.
To estimate the tax revenue for the next year, the FMLS method’s forecast was then compared with forecasts of other estimation techniques. Results are shown in
Table 7.
The forecast for comprehensive income tax revenue in 2025 was estimated at 17,173,790 million KRW using the FMLS method. The actual time series forecast (random simulation) of 25,177,055 million KRW served as a reference estimate. However, the forecast from the comprehensive income tax cointegration regression model could be sensitive to assumptions regarding predicted values of explanatory variables such as business income and real estate rental income, the chosen estimation model, the chosen estimation method, and shocks related to business income and real estate rental income.
4.1.3. Capital Gains Tax
Estimation results for the level variable equation of comprehensive income tax after adjustment considering key factors such as multicollinearity were obtained from y = f(X
1, X
2, X
3), where y was the capital gains tax, X
1 was the transfer price, X
2 was the capital gain, X
3 was the forecasting error of capital gains tax, and f(.) was the functional form. Based on estimated values of the capital gains tax cointegration model, in-sample capital gains tax revenue predictions reflecting basic forecast values of capital gains tax explanatory variables are shown in
Table 8.
Comparison of model forecasting accuracy using RMSE for in-sample capital gains tax revenue forecasts showed that, compared to the government’s forecast of 8,715,320 million KRW, OLS had a forecast of 6,428,878 million KRW and FMLS had a forecast of 6,067,648 million KRW. These results revealed that the FMLS method had the highest forecasting accuracy, followed by the OLS method, while the government’s forecast had the lowest accuracy. This indicates that the forecasting error for capital gains tax has increased recently, leading to a larger shortfall in capital gains tax revenue.
To estimate tax revenue for the next year, the FMLS method’s forecast was then compared to forecasts with other estimation forecasts. Results are shown in
Table 9.
The forecast for capital gains tax revenue in 2025 was 19,081,552 million KRW using the FMLS method, as it had the highest forecasting accuracy. The actual time series forecast (random simulation) of 20,551,933 million KRW served as a reference estimate. However, the forecast from the capital gains tax cointegration regression model could be sensitive to assumptions regarding predicted values of explanatory variables such as asset transfer price and capital gain, the chosen estimation model, the chosen estimation method, and shocks related to asset transfer price and capital gain.
4.1.4. Interest and Dividend Income Tax
Estimation results for the level variable equation of interest and dividend income tax after adjustment considering key factors such as multicollinearity were obtained from y = f(X
1, X
2, X
3), where y was the interest and dividend income tax, X
1 was the interest and dividend income, X
2 was the balance of deposits and bonds, X
3 was the forecasting error of interest and dividend income tax, and f(.) was the functional form. Based on estimated values of the interest and dividend income tax cointegration model, in-sample interest and dividend income tax revenue predictions reflecting basic forecast values of the interest and dividend income tax explanatory variables are shown in
Table 10.
The comparison of model forecasting accuracy using RMSE for in-sample interest and dividend income tax revenue forecasts showed that, compared to the government’s forecast of 1,141,991 million KRW, OLS had a forecast of 2,172,656 million KRW and FMLS had a forecast of 1,302,785 million KRW. These results revealed that the government’s prediction method had the highest forecasting accuracy, followed by the FMLS method, whereas the OLS method had the lowest accuracy.
To estimate the tax revenue for the next year, the FMLS method’s forecast was compared with other estimation forecasts. Results are shown in
Table 11.
The forecast for interest and dividend income tax revenue in 2025 was estimated at 8,016,698 million KRW using the FMLS method. The actual time series forecast (random simulation) of 11,040,117 million KRW served as a reference estimate. However, the forecast from the interest and dividend income tax cointegration regression model could be sensitive to assumptions regarding predicted values of explanatory variables such as interest and dividend income, the chosen estimation model, the chosen estimation method, shocks related to interest and dividend income, and balances of deposits and bonds.
4.2. Corporate Tax
Estimation results for the level variable equation of corporate tax after adjustment considering key factors such as multicollinearity were obtained from y = f(X
1, X
2, X
3), where y was the corporate tax, X
1 was the corporate net income, X
2 was the capital investment, X
3 was the corporate tax forecasting error, and f(.) was the functional form. Based on estimated values of the corporate tax cointegration model, in-sample corporate tax revenue predictions reflecting basic forecast values of corporate tax explanatory variables are shown in
Table 12.
The forecasting accuracy of corporate tax revenue can be assessed using prediction error loss. The comparison of model prediction performance using the RMSE for in-sample corporate tax revenue predictions showed that the government’s forecast was 17,382,371 million KRW, whereas OLS’ forecast was 23,140,958 million KRW, DOLS’s forecast was 12,308,331 million KRW, and FMLS’s forecast was 20,905,217 million KRW. These results revealed that the DOLS method had higher forecasting accuracy than the government’s prediction method, suggesting that the recent increase in corporate tax forecasting errors might have led to a larger corporate tax deficit.
As the DOLS prediction showed a higher forecasting accuracy, the predicted corporate tax revenue for the following year compared with other time series forecasts is shown in
Table 13.
The forecast for corporate tax revenue in 2025 was estimated at 81,549,804 million KRW using the DOLS method. The actual time series forecast for reference was estimated at 81,974,971 million KRW. However, the forecast from the corporate tax cointegration regression model could be sensitive to assumptions regarding predicted values of explanatory variables such as net income and capital investment, the chosen estimation model, the chosen estimation method, and shocks related to corporate sales or profits.
4.3. Inheritance and Gift Tax
Estimation results for the level variable equation of inheritance and gift tax after adjustment considering key factors such as multicollinearity were obtained from y = f(X
1, X
2, X
3, X
4), where y was the inheritance and gift tax, X
1 was the value of inherited & donated property, X
2 was the number of heirs, X
3 was the forecasting error, X
4 was the number of gift determiners, and f(.) was the functional form. Based on estimated values of the inheritance and gift tax cointegration model, the in-sample tax revenue predictions reflecting basic forecast values of explanatory variables are shown in
Table 14.
To evaluate the forecasting accuracy of inheritance and gift tax, the RMSE for in-sample inheritance and gift tax revenue forecasts was used to compare different models’ predictive performances. Results showed that the government’s forecast was 1,991,453 million KRW, while OLS had a forecast of 2,327,545 million KRW, DOLS had a forecast of 1,341,686, million KRW, and FMLS had a forecast of 3,927,365. Therefore, the DOLS method had higher forecasting accuracy than the government’s method.
As the DOLS forecast showed higher accuracy, the next year’s tax revenue was compared with tax revenue of other time series forecasting methods. Results are shown in
Table 15.
The forecast for inheritance and gift tax revenue in 2025 was estimated at 15,212,612 million KRW using the DOLS method. The actual time series forecast of 17,897,784 million KRW served as a reference estimate. However, the forecast from the inheritance and gift tax cointegration regression model could be sensitive to assumptions regarding predicted values of explanatory variables such as inherited and gifted property values, number of heirs, number of gift decisions, the chosen estimation model, the chosen estimation method, and shocks related to inheritance and gifts.
4.4. Securities Transaction Tax
Estimation results for the level variable equation of securities transaction tax after adjustment considering key factors such as multicollinearity were obtained from y = f(X
1, X
2, X
3), where y was the securities transaction tax, X
1 was the stock, X
2 was the number of stock transaction, X
3 was the forecasting error, and f(.) was the functional form. Based on estimated values of the tax cointegration model, in-sample tax revenue predictions reflecting the basic forecast values of the explanatory variables are shown in
Table 16.
To evaluate the forecasting accuracy for securities transaction tax, the RMSE for in-sample securities transaction tax revenue forecasts was used to compare models’ predictive performances. Results showed that the government’s forecast was 1,172,516 million KRW, while OLS had a forecast of 1,823,285 million KRW, DOLS had a forecast of 889,078 million KRW and FMLS had a forecast of 1,166,516 million KRW. Therefore, the DOLS method had the highest forecasting accuracy, followed by FMLS, the government’s forecast, and OLS.
Since the DOLS forecast demonstrated a higher accuracy than others, the next year’s tax revenue was compared with other time series forecasts. Results are shown in
Table 17.
The securities transaction tax revenue in 2025 was forecasted to be 6,346,607 million KRW based on the DOLS method. The reference estimate from the FMLS method was 6,522,208 million KRW. However, the forecast from the securities transaction tax cointegration regression model could be sensitive to assumptions regarding predicted values of explanatory variables such as KOSPI, KOSDAQ, other securities, the number of securities transactions, the chosen estimation model and method, and shocks related to securities and their transaction volumes.
4.5. Stamp Duty Tax
Estimation results for the level variable equation of stamp duty tax after adjustment considering key factors such as multicollinearity were obtained from y = f(X
1, X
2), where y was the stamp duty tax, X
1 was the number of tax stamp documents, X
2 was the forecasting error, and f(.) was the functional form. Based on estimated values of the tax cointegration model, in-sample tax revenue predictions reflecting basic forecast values of explanatory variables are shown in
Table 18.
To assess the predictive power of the stamp duty, the prediction performance of the model was compared using RMSE based on in-sample stamp duty forecast values. As a result, compared to the government’s prediction of 106,358 million KRW, the OLS method predicted 50,672 million KRW, DOLS predicted 147,634 million KRW, and FMLS predicted 14,524 million KRW. Therefore, the FMLS method showed the highest predictive power, followed by the OLS method, the government’s prediction, and the DOLS method.
Since the FMLS prediction demonstrated the highest predictive power, its forecast for the following year’s tax revenue was compared with other time series forecasts. Results are shown in
Table 19.
The 2025 stamp duty revenue forecast was estimated to be 793,055 million KRW based on the FMLS method. The government’s forecast of 629,011 million KRW was provided as a reference estimate. However, the forecast from the stamp duty cointegration regression model could be sensitive to assumptions regarding forecasted values of taxed documents, the chosen estimation model and method, and shocks related to the stamp duty tax.
4.6. Comprehensive Real Estate Tax
Estimation results for the level variable equation of comprehensive real estate tax after adjustment considering key factors such as multicollinearity were obtained from y = f(X
1, X
2, X
3, X
4), where y was the comprehensive real estate tax, X
1 was the national asset, X
2 was the real estate income, X
3 was the number of real estate tax payers, X
4 was the forecasting error, and f(.) was the functional form. Based on estimated values of the tax cointegration model, the in-sample tax revenue predictions reflecting basic forecast values of explanatory variables are shown in
Table 20.
To assess the predictive accuracy of the Comprehensive Real Estate Tax, the Root Mean Square Error (RMSE) for predicted values of the tax in the sample was used to compare performances of different models. As a result, the RMSE was 1,510,862 million KRW for government’s prediction method, 1,289,866 million KRW for the OLS method, and 1,980,777 million KRW for the FMLS method. Therefore, the OLS prediction method demonstrated the highest predictive accuracy, followed by the government prediction method and then the FMLS method.
Tax estimates for the following year were predicted using the FMLS method. Results are shown in
Table 21.
The 2025 estimate for the Comprehensive Real Estate Tax was projected to be 4,336,215 million KRW based on the FMLS method. The estimate from the single time-series forecasting method was 3,648,200 million KRW, which could be referred to as a benchmark estimate. However, the forecast from the cointegration regression model for the Comprehensive Real Estate Tax might be sensitive to changes in assumptions regarding explanatory variables such as national assets and business real estate income, the selected estimation model and method, shocks related to national assets, business real estate income, and the number of taxpayers for the Comprehensive Real Estate Tax.
4.7. Value-Added Tax
Estimation results for the level variable equation of value-added tax after adjustment considering key factors such as multicollinearity were obtained from y = f(X
1, X
2, X
3), where y was the value-added tax, X
1 was the revenue, X
2 was the import, X
3 was the forecasting error, and f(.) was the functional form. Based on estimated values of the tax cointegration model, in-sample tax revenue predictions reflecting basic forecast values of explanatory variables are shown in
Table 22.
To evaluate the forecasting accuracy of VAT (Value Added Tax), we compared the model’s forecasting performance using the RMSE of VAT predictions within the sample. Results showed that, compared to the government forecast of 6,863,105 million KRW, the RMSE was 7,505,311 million KRW for OLS, 5,600,372 million KRW for DOLS, and 8,099,686 million KRW for FMLS. Therefore, the DOLS method exhibited the highest forecasting accuracy, followed by the government method, the OLS method, and the FMLS method.
Forecast results for the following year’s tax revenue using the DOLS method are presented in
Table 23.
The 2025 forecast for Value Added Tax (VAT) revenue was estimated to be 80,494,634 million KRW based on the DOLS method. The dynamic model DOLS method estimated it at 79,326,416 million KRW and the single time series forecasting method estimated it at 76,822,040 million KRW. However, the forecast from the VAT cointegration regression model might be sensitive to assumptions regarding predicted values of explanatory variables such as sales and income, the chosen estimation model and method, and shocks related to business revenues and overseas income.
4.8. Liquor Tax
Estimation results for the level variable equation of liquor tax after adjustment considering key factors such as multicollinearity were obtained from y = f(X
1, X
2), where y was the liquor tax, X
1 was the liquor revenue, X
2 was the forecasting error, and f(.) was the functional form. Based on estimated values of the tax cointegration model, in-sample tax revenue predictions reflecting basic forecast values of explanatory variables are shown in
Table 24.
To assess the predictive power of the liquor tax, we used the RMSE (Root Mean Square Error) of predicted values for liquor tax to compare performances of models. In comparison, the government’s forecast method had an RMSE of 20,607 million KRW, while the RMSE was 310,292 million KRW for the OLS method, 699,372 million KRW for the DOLS method, and 303,585 million KRW for the FMLS method. Therefore, the government’s forecast method had the highest predictive power, followed by the FMLS method, the OLS method, and the DOLS method.
The forecasted liquor tax revenue for the next year was predicted using the FMLS method. Results are shown in
Table 25.
The 2025 liquor tax revenue forecast was estimated to be 3,599,901 million KRW using the FMLS method. The reference estimate using the single time series forecasting method was 3,594,276 million KRW. However, the forecast from the liquor tax cointegration regression model was sensitive to assumptions regarding predicted values of alcoholic beverage sales and import values, the chosen estimation model and method, and shocks related to domestic sales and imports of alcoholic beverages.
4.9. Transportation, Energy, and Environmental Tax
Estimation results for the level variable equation of transportation, energy, environmental tax after adjustment considering key factors such as multicollinearity were obtained from y = f(X
1, X
2, X
3), where y was the transportation, energy, environmental tax, X
1 was gasoline consumption, X
2 was diesel consumption, X
3 was the forecasting error, and f(.) was the functional form. Based on estimated values of the tax cointegration model, in-sample tax revenue predictions reflecting basic forecast values of explanatory variables are shown in
Table 26.
To evaluate the predictive power of transportation, energy, and environmental tax models, we used the RMSE of predicted values within the sample to compare models’ performances. Results showed that the RMSE was 262,668 million KRW for the government’s forecast method, 3,749,250 million KRW for OLS, 209,574 million KRW for DOLS, and 3,574,776 million KRW for FMLS. Therefore, the DOLS method demonstrated the highest predictive accuracy, followed by the government’s forecast method, the FMLS method, and the OLS method.
Results of predicting the next year’s tax revenue using the DOLS method are shown in
Table 27.
The 2025 forecast for the transportation, energy, and environmental tax was estimated at 10,908,554 million KRW using the DOLS method. The single time series forecasting method provided a reference estimate of 12,792,677 million KRW. However, the forecast from the transportation, energy, and environmental tax cointegration regression model might be sensitive to assumptions regarding predicted values of gasoline and diesel consumption, the chosen estimation model and method, and related shocks to transportation, energy, and environmental factors.
4.10. Excise Tax
Estimation results for the level variable equation of excise tax adjusted considering key factors such as multicollinearity were obtained from y = f(X
1, X
2, X
3), where y was the excise tax, X
1 was the final consumption, X
2 was the taxable goods and place of sale, X
3 was the forecasting error, and f(.) was the functional form. Based on estimated values of the tax cointegration model, in-sample tax revenue predictions reflecting basic forecast values of the explanatory variables are shown in
Table 28.
To assess the predictive accuracy of the individual consumption tax, the model’s predictive performance was compared using the RMSE (Root Mean Square Error) based on forecasted values of the individual consumption tax in the sample. Results are shown as follows. The RMSE was 1,132,344 million KRW for the government’s prediction method, 1,106,527 million KRW for the OLS (Ordinary Least Squares), 2,423,149 million KRW for the DOLS (Dynamic Ordinary Least Squares), and 1,063,139 million KRW for the FMLS (Fully Modified Least Squares). Therefore, the FMLS method demonstrated the highest predictive accuracy, followed by the OLS method, the government’s forecast method, and the DOLS method.
The next year’s tax revenue was forecasted using the FMLS method. Results are shown in
Table 29 below.
The 2025 individual consumption tax forecast was estimated to be 11,006,891 million KRW based on the FMLS method. The reference estimate using the single time series forecasting method was 8,805,095 million KRW. However, the forecast from the individual consumption tax cointegration regression model was sensitive to assumptions regarding predicted values of final consumption and sales of taxable goods and locations, as well as the adopted estimation model and techniques. It may change due to shocks related to final consumption in stores and sales of luxury taxable goods and locations.
4.11. Tariff
Estimation results for the level variable equation of tariff, adjusted considering key factors such as multicollinearity, were obtained from y = f(X
1, X
2, X
3), where y was the tariff, X
1 was the imports, X
2 was the effective tariff rate, X
3 was the term of trade, and f(.) was the functional form. Based on estimated values of the tax cointegration model, in-sample tax revenue predictions reflecting basic forecast values of explanatory variables are shown in
Table 30.
To assess the predictive accuracy of tariffs, the RMSE (Root Mean Squared Error) was used to compare performances of various models for tariff predictions within the sample. Results show that, compared to the government’s RMSE of 2,436,291 million KRW, the OLS (Ordinary Least Squares) method had an RMSE of 1,315,168 million KRW and the FMLS (Fully Modified Least Squares) method had an RMSE of 1,388,382 million KRW. Therefore, the FMLS method showed the highest predictive accuracy, followed by the OLS method and then the government’s forecast method.
The tariff revenue for the following year was predicted using the FMLS method. Results are shown in
Table 31.
The 2025 tariff revenue was estimated to be 8,776,880 million KRW using the FMLS method. The reference estimate from the univariate time series forecasting method was 5,946,109 million KRW. However, the forecast from the tariff cointegrated regression model might vary significantly based on assumptions about predicted values for customs import amounts, effective tariff rates, trade conditions, and chosen estimation models and techniques. Additionally, external import shocks, effective tariff rates, and trade condition changes can impact the forecast.
4.12. Projections for National Tax Revenue
Finally, based on estimated values from the cointegrated revenue model for each tax type, forecasted national tax revenues were compared with the government’s forecasts and actual revenues. Additionally, future tax revenue projections are provided. For convenience, projections for the tax revenue, excluding the education tax, agricultural and fishery special tax, and other income taxes, are presented in
Table 32.
The government’s tax revenue forecasts for 2023 and 2024 were approximately 374,763,900 million KRW and 378,420,800 million KRW, respectively. However, the actual settled tax revenue was more than 50 trillion KRW lower. In comparison, the cointegration predicted tax revenue was much closer to the actual settled tax revenue. As a result, the forecast for 2025 was 341,524,525 million KRW.
Recent forecasting errors have been attributed more to long-term asymmetric errors by tax type than to temporary economic shocks. From this perspective, utilizing information about the weight and error in taxes with significant discrepancies can improve forecasting accuracy. Thus, national tax revenue was estimated by summing optimal DOLS or FMLS estimates for each tax type using the cointegration model.
For example, looking at large tax revenue shortfalls in income tax, corporate tax, and value-added tax, the corporate tax forecast using the cointegration DOLS method showed a high accuracy. The value-added tax forecast also performed well using the cointegration DOLS method. Among income taxes, the government’s forecast method performed the best for labor income tax, comprehensive income tax, and interest and dividend income taxes, with the cointegration method being the second best. However, for capital gains tax, the cointegration FMLS method had the highest prediction performance. These results suggest that optimizing tax forecasts by combining different estimation techniques for each tax type, especially considering asymmetric errors, can lead to more accurate tax revenue predictions.