To evaluate the performance of the proposed finite distribution function estimators,
,
, and
, we conducted simulation studies based on two populations: A synthetic finite population from Chen et al. [
16], and the 2023 Korean Survey of Household Finances and Living Conditions.
Performance was then assessed over
simulation replications using percentage relative bias (%RB) and relative root mean squared error (RRMSE), where
with
denoting the estimate from replication
r and
the target parameter. For the finite distribution function, the bootstrap variance, and quantiles, the corresponding choices were
The performance of the Woodruff confidence interval was assessed by its coverage probability (
), lower error rate (%L), and upper error rate (%U):
where
and
denote, respectively, the lower and upper Woodruff CI bounds for the
qth quantile in replication
r.
5.1. Study 1
Following the simulation design of Chen et al. [
16], we generated a finite population of size
. The study variable
y and auxiliary variables
are generated from
where
follow the design in Chen et al. [
16], and the error terms
. The parameter
is chosen such that the correlation coefficient
between
y and the linear predictor
equals 0.5.
We consider four model specification scenarios
TT: Both and M are correctly specified.
TF: is correctly specified, but M is misspecified, with omitted from the model.
FT: M is correctly specified, but is misspecified, with omitted from the model.
FF: Both models are misspecified, with omitted in each model.
The analysis uses a nonprobability sample
of size
and a probability sample
of size
.
Table 1 reports %RB and RRMSE for the proposed finite distribution function estimators. Under TT, all estimators exhibit low bias and error, indicating stable performance. Under TF and FT, the DR estimator attains lower bias and error than the alternatives, highlighting the advantages of the doubly robust property. By contrast, under FF, performance deteriorates substantially for all estimators.
Table 2 compares the bootstrap variance estimators in terms of %RB and
. Under TT, all variance estimators perform satisfactorily. Under TF and FT, despite model misspecification, the variance estimator associated with the DR method retains low bias and an
close to 95%, indicating stable reliability and accuracy. Conversely, under FF, coverage performance deteriorates markedly across all methods.
Table 3 summarizes the results for the quantile estimators. Mirroring the findings for the finite distribution function estimators, all methods perform well under the TT scenario. Under TF and FT, the DR-based quantiles remain stable, confirming the robustness of the doubly robust approach. By contrast, under FF, overall estimation accuracy deteriorates.
Table 4 reports the Woodruff CI results for the quantile estimators, including
, %L, and %U.Consistent with previous findings, all methods perform well under the TT scenario. Under TF and FT, the DR-based intervals maintain
close to the nominal 95% with balanced tail errors, indicating high reliability. By contrast, under FF, coverage deteriorates substantially across methods.
falls below the nominal level and both tail error rates increase, signaling degraded interval performance.
5.2. Study 2
In the second simulation study, we treat the 2023 Korean Survey of Household Finances and Living Conditions (SHFLC;
) as the finite population and repeatedly draw subsamples from it.
Table 5 summarizes the key variables used in the experiment and their definitions.
The nonprobability sample
was generated to mimic structures commonly observed in practice. The propensity score model was specified as a logistic regression,
where
was chosen so that
. Under this design, households with higher educational attainment of the household head, non-single households, apartment residents, and households without debt were more likely to be included in
. The nonprobability sample
was then selected by Poisson sampling with inclusion probabilities
.
The probability sample was stratified into nine strata defined by GEO, HOME, and SIZE. A mixed allocation scheme—combining Neyman and proportional allocation—was used to determine stratum specific sample sizes, followed by simple random sampling without replacement within each stratum. The sample sizes were set to and .
The study variable of interest was current income (INCOME). Because the true outcome model was unknown, we included EXP1 and EXP2- the covariates with comparatively strong explanatory power- as regressors in the working model. This setup allows us to assess the impact of model misspecification on estimation performance and to isolate efficiency gains attributable to the DR estimator. We consider two scenarios regarding the propensity score model:
Table 6 reports the results for the distribution–function estimators. Overall, the REG estimator performs reasonably well, although its bias and error are somewhat larger at lower quantiles than at middle and upper quantiles, likely reflecting the limited explanatory power of the auxiliary variables and the possible over-representation of high-income households. Under Scenario A, the IPW estimator and the DR estimator both exhibit low bias and error, confirming the effectiveness of propensity score adjustment. Under Scenario B, REG estimator is the most stable, while the DR estimator inherits some bias from the misspecified IPW component and thus loses efficiency. In summary, when the propensity score model is correctly specified, the IPW estimator, the REG estimator, and the DR estimator all yield stable results. However, when the propensity-score model is misspecified, only the REG estimator and the DR estimator perform well, with the REG estimator performing best. These findings highlight that the choice of estimator may critically depend on the availability and explanatory power of the auxiliary variables.
Table 7 compares the bootstrap variance estimators in terms of %RB and
. Consistent with the findings for the finite distribution function estimators, the REG estimator shows degraded variance performance at lower quantiles. The IPW estimator maintains coverage close to 95%
under Scenario A, but its
declined markedly under Scenario B. The DR estimator achieves both low bias and stable
across scenarios, indicating reliable variance estimation.
Table 8 compares the quantile estimators in terms of %RB and RRMSE. The REG estimator shows substantial bias at lower quantiles, whereas the IPW estimator performs well under Scenario A but deteriorates under Scenario B. The DR estimator maintains moderate bias and error across both scenarios, yielding comparatively stable performance overall.
Table 9 reports results for the Woodruff confidence intervals of the quantile estimators-
, %L, and %U. The IPW estimator attains
close to the nominal 95% under Scenario A, but coverage drops sharply under Scenario B, accompanied by an upward bias in %U, indicating sensitivity to propensity score misspecification. The REG estimator performs well at the middle and upper quantiles, but shows increased %L at lower quantiles. The DR estimator maintains stable
across scenarios, with only a slight upward bias in %U under Scenario B.