Preprint
Article

This version is not peer-reviewed.

City-Level Business-Registration Centers and Entity-Class Location Patterns in China: An MSP-MBKMeans Analysis of 185 Million Records

Submitted:

09 August 2026

Posted:

14 August 2026

You are already at the latest version

Abstract
Business-registration coordinates describe legal addresses, not verified workplaces, and shared addresses can create artificial density peaks. This study identifies registration-center structure and compares the center--periphery positions of four historical entity classes in China. The January 2021 archive contains 185,033,392 deduplicated entities in 364 city-level source units; 345 have observations. The proposed Multi-Scale Persistent-Peak-Constrained MiniBatch K-Means algorithm (MSP-MBKMeans) combines peaks repeated across two H3 scales with density-valley, minimum-share, adaptive-distance, address-weight, and spatial-resampling checks. In a prespecified holdout of 200 synthetic cities, it recovered the substantive center count in 99.0%; HDBSCAN and OPTICS reached 100.0%, so no universal superiority is claimed. The fixed 8% primary definition identified 1,410 centers, with a city median of four; 25 cities remained unresolved. At 6/8/10%, primary-center position was stable in 71.3% of cities, whereas center count was often definition-dependent. An independent comparison in 323 units placed the primary registration center a median 3.32 km from the GHSL 2020 population node, with 82.7% within 10 km. After class-wise false-discovery-rate control, foreign-identified and Hong Kong/Macao/Taiwan-identified entities were more central than the private-dominated residual class in 79 and 82 cities; their city-equal median advantages remained 3.74 and 4.57 km. State-identified entities had no common national direction. Complete center-share sensitivity analyses at 6/8/10% preserved 90.7% of the 1,035 city-class labels, and leave-class-out centers preserved 99.3%. A supplementary catalogue reports corrected conclusions and separate structural and class-location sensitivity warnings for every unit. Conclusions concern registration locations, not production, employment, or causal ownership effects.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  

1. Introduction

Cities concentrate organizations, infrastructure, information, and administrative services, yet the word “center” can refer to very different objects. An employment center is identified from jobs, a commercial center from customer-facing activity, a built-up center from settlement morphology, and a business-registration center from legal addresses. These objects can overlap without being identical. A national archive of geocoded registration records offers a rare opportunity to examine where organizations are legally anchored, but it also creates a strong temptation to call those coordinates workplaces or economic activity points. This paper starts by refusing that shortcut. Its subject is registration geography: the spatial structure of recorded legal locations and the way historically identified entity classes are positioned within that structure.
The distinction is not merely terminological. A registration address may be a genuine office, a headquarters, a shared service address, an incubator, or an address maintained for regulatory purposes. Chinese company practice explicitly recognizes possible inconsistency between a registered and an actual business address [1]; research in other settings likewise shows that an address can act as an organizational signal rather than a direct measure of activity [2]. If a building hosts thousands of legal registrations, a conventional density surface can place a prominent peak there even when the corresponding entities operate elsewhere. Conversely, erasing multiplicity and counting every address once may discard real organizational concentration. A scientifically useful analysis must expose this uncertainty instead of burying it in preprocessing.
Registration data nonetheless have distinctive value. Administrative business records can cover organizations that are absent from surveys, consumer platforms, remote-sensing products, and employment datasets. Large Chinese enterprise-registration data have already been used to support spatiotemporal industrial analysis and to address missing registration attributes [3]. More generally, emerging urban data sources allow city patterns to be studied at scales that would be impractical with field enumeration alone [4]. The present snapshot contains more than 185 million deduplicated entities and reaches hundreds of city-level units. Its scale makes national comparison possible, but scale alone is not a contribution. The underlying business-data project raised a practical question: how can an enormous coordinate archive be turned into understandable city findings when shared addresses, uncertain center counts, and missing cities cannot be ignored? Without clear definitions and complete city-by-city evidence, more records can simply make a biased result look more precise.
Three linked questions motivate the study. First, how many repeatable business-registration centers can be identified in each observed city without fixing the number in advance? Second, once those centers are defined independently of entity class, do state-identified, foreign-identified, and Hong Kong/Macao/Taiwan-identified entities occupy systematically different center–periphery positions from the private-dominated residual class? Third, can the densest part of each center catchment be delineated as a compact registration zone without joining separate centers across empty space? These questions form one research line. Center discovery supplies the reference geometry; class comparison gives that geometry a social-economic interpretation; and zone delineation translates point concentrations into reviewable spatial outputs.
Existing tools solve parts of the problem but not the whole research problem. K-means minimizes within-cluster squared distance [5], while MiniBatch K-Means makes related optimization practical for very large datasets [6]. Neither algorithm decides, by itself, how many substantively meaningful urban centers exist. Density-peak methods identify locally prominent modes [7], but a peak seen at one spatial resolution may be a grid artifact, an address spike, or a minor satellite. Standard cluster validation, including silhouette reasoning [8], can favor geometrically useful partitions that do not correspond to interpretable urban centers. Stability assessment can also be misleading when analysts fix the selected K and repeatedly refit centers under that fixed K: conditional stability is then mistaken for selection stability.
We address these problems with Multi-Scale Persistent-Peak-Constrained MiniBatch K-Means, abbreviated MSP-MBKMeans. The procedure discovers peaks at two nested H3 scales, retains only cross-scale correspondences, applies minimum-share, adaptive-separation, and intervening-density-valley constraints, and uses the surviving peaks as fixed starting centers for MiniBatch K-Means. Most importantly, each H3-block resample repeats peak discovery and selection of K. A city can therefore be reported as exactly stable, core-stable with optional centers, usable with uncertainty, or unresolved. This graduated outcome is more informative than either forcing one solution for every city or discarding all cities that are not perfectly stable.
Concentrated addresses are handled through three prespecified weighting schemes. Raw entity multiplicity (A) preserves the archive literally. Exact-address equality (B) gives the strongest correction. City-specific 99th-percentile truncation (C) retains most count information while bounding the largest shared addresses and serves as the primary estimate. The same A/B/C logic is used both to test center robustness and to evaluate R-relative class differences. A directional national claim is allowed only when all three city-equal medians agree and their province-cluster intervals exclude zero. This rule makes robustness part of the estimand rather than an appendix added after seeing the preferred result.
The empirical findings are concise even though the calculations are extensive. Across 300 synthetic cities, MSP-MBKMeans recovered the correct center count in 100.0% of six calibration geometries and 87.8% of nine stress geometries; HDBSCAN and OPTICS reached 88.9% under stress. In 345 units with source samples, the median selected count was four. Center uncertainty was not negligible: 69 cities had exactly stable counts, 88 stable cores, 163 usable structures with uncertainty, and 25 unresolved structures. An independent GHSL comparison was possible for 323 units with accepted auxiliary boundaries; primary registration centers had a 3.32 km median distance from the population-weighted settlement reference, and 82.7% were within 10 km.
The central social-economic result is not an abstract statement that cities are “heterogeneous.” Foreign-identified and Hong Kong/Macao/Taiwan-identified entities are closer to the primary business-registration center than the private-dominated R class in the city-equal national distribution. Under the primary C specification, the median advantages are 3.74 and 4.57 km; both remain positive in the strict subset that requires an exactly stable center count and an address-robust primary center. State-identified entities have no nationally common center–periphery direction. The latter result is substantive rather than a failed test: state-related registration locations can reflect central administrative institutions, peripheral land-intensive units, local state-enterprise histories, and shared addresses in different combinations across cities.
The paper makes four contributions. First, it combines multi-scale peak persistence with constrained MiniBatch K-Means to identify interpretable registration centers and tests the design against five competing algorithms and five substantive ablations. Second, it reselects K under spatial perturbation instead of assuming that the full-sample center count is certain. Third, it tests whether concentrated shared addresses or the mixture of city-level administrative units changes the conclusions. Fourth, it provides a complete national application: 1,410 center catchments, 1,359 zones reaching their density-coverage target, findings for 345 observed units, and explicit no-sample statements for 19 units. All 364 units appear in the city conclusion catalogue.
The contribution is deliberately bounded. MSP-MBKMeans is a domain-specific combination of established methods and new rules, not a new theory of clustering. The results do not identify a true central business district, production location, workplace, firm migration, or an ownership effect. The four labels do not restore the lost 429 source strings. K is not a development score. The analysis is descriptive, cross-sectional with respect to the archive snapshot, and tied to administrative city units. These limits narrow the claims, but they also make the results usable: readers know exactly what was measured, which cities support formal inference, and where the method declined to decide.
The remainder of the paper reviews the relevant urban-center and spatial-clustering research, explains the data and method, presents algorithmic and substantive results, interprets eight cities, reports all 364 units, and discusses what registration centrality can and cannot tell us about urban organization.

3. Materials and Methods

3.1. Study Object, Snapshot, and Inferential Unit

The object of this study is the business-registration location: the coordinate attached to a legal registration record. It is not assumed to be a shop front, factory, office, headquarters, workplace, or place at which employees and customers are physically present. This distinction governs the title, variables, maps, and interpretation. The expression “registration center” therefore denotes a concentration center estimated from registration coordinates. Similarly, a “registration agglomeration zone” denotes a high-density portion of those coordinates; it is not called an industrial cluster.
The source snapshot was created on 11–12 January 2021. It is a historical cumulative collection containing active, cancelled, revoked, and other registration states. Besides companies, it contains a small number of public institutions, foundations, schools, law firms, and other registered social entities. Because complete establishment, closure, and status histories are unavailable, the snapshot cannot support a cohort analysis, a 2021 entry analysis, or a claim about all firms operating on one date. Our estimand is the spatial organization visible in this cumulative registration archive.
The source data were lawfully purchased by the institution represented by the author for research use. Entity-level records remain subject to third-party copyright and institutional agreement and therefore cannot be redistributed.
The source inventory began with 185,107,617 valid physical records in four source partitions. Cross-file identity checks, duplicate resolution, damaged-record handling, coordinate validity checks, and city reconciliation produced 185,033,392 deduplicated entities in the formal analysis layer. All 364 source city or regional units remain in the final city inventory. Of these, 345 contain usable source observations and 19 contain none. A no-sample unit is not silently dropped or assigned a zero-density estimate: it receives a separate “no source sample; not estimable” output.
The city is the inferential unit for national comparisons. This choice prevents a few very large registration archives from determining a national median. Province is the resampling cluster because cities in the same province can share administrative practice, data lineage, and regional context. Entity counts remain important within a city for fitting density and distance distributions, but national summaries assign equal weight to each eligible city.

3.2. Historical Entity Labels and Their Limits

The retained field contains four historical output labels, denoted R, S, F, and H (Table 1). S, F, and H resulted from matching a now-lost organization-type string to 56 state-related, 193 foreign-related, and 180 Hong Kong/Macao/Taiwan-related source mappings. The 429 mapping strings are not 429 observed categories in the retained data. Because the original organization-type field is unavailable, no entity can be reassigned to one of those strings and no classifier is trained to reconstruct it.
R is the residual class not matched by S, F, or H. It is overwhelmingly private-enterprise dominated but also includes the small non-enterprise portion of the archive. We therefore use the operational name “private-dominated other registration entities,” not the stronger claim “all private firms.” The other classes are likewise called “state-identified,” “foreign-identified,” and “Hong Kong/Macao/Taiwan-identified.” These names describe historical identification rules, not a verification of beneficial ownership in 2021.
The four counts sum to the complete 185,033,392-entity analysis layer. S, F, and H together account for 0.9717% of the formal archive. A class-city comparison is reportable only when the class contains at least 100 raw entities and spans at least eight H3 resolution-6 spatial blocks. This rule gives rare classes a meaningful spatial footprint and avoids treating a large R denominator as evidence that a geographically concentrated rare class is precisely estimated.

3.3. Coordinate Reference System, City Identity, and Grid Representation

The source longitude–latitude pairs were confirmed to be BD-09. They were converted to WGS84 (EPSG:4326) using the same specified transformation throughout the study. Across province samples, the median displacement induced by conversion ranged from 854.91 to 1,382.87 m and the 99th percentile ranged from 941.46 to 1,477.76 m. Retaining the unconverted coordinates would therefore introduce a systematic displacement of the same order as some local center-location errors. Raw coordinates were preserved in the coordinate checks, while all modeling uses WGS84.
City identity follows the source-file inventory and a 2021 administrative-code crosswalk. An older local boundary layer is used for cartographic context and anomaly review only; it neither deletes nor reallocates an entity. This protects the analysis from silently changing the source population when a coordinate lies outside an imperfect polygon. It also means that city units are not geometrically comparable in a strict morphological sense: municipalities, prefecture-level cities, autonomous prefectures, and directly administered county-level units can cover very different extents.
To test whether that mixture determines the national result, the complete analysis is summarized for three nested inventories fixed before comparison: all 364 source units; 330 prefecture-level units after excluding province-level municipalities and special county-level units; and 291 standard prefecture-level cities whose official names end in “city.” Center reliability and the C-specification S/F/H contrasts are recomputed within each inventory. This sensitivity analysis tests the composition of the inventory, not hypothetical changes to boundaries within any one unit.
Exact WGS84 sites are aggregated to the H3 discrete global grid. Resolution 9 is the fitting input; resolutions 8 and 7 support multi-scale peak detection, and resolution 7 also defines stability blocks. Hexagonal aggregation provides compact adjacency and a reproducible hierarchy, while true cell area is used whenever counts become densities. The method does not require a planar national projection. Great-circle distances are computed in kilometers, and local weighted fitting coordinates are handled consistently within each city. H3 is a spatial representation used in the analysis, not evidence that the underlying process is naturally hexagonal.

3.4. Three Address-Weighting Specifications

Shared registration agents, incubators, office buildings, and administrative addresses can hold unusually many entities. Treating every record literally may turn one legal address into an apparent urban center; treating every address equally may remove genuine differences between large and small sites. We therefore analyze three prespecified weights for exact site i, where n i is its raw entity count and q . 99 , c is the city-specific 99th percentile of site counts:
w i A = n i , w i B = 1 , w i C = min ( n i , q . 99 , c ) .
Specification A retains all entity multiplicity. B assigns equal weight to each exact address and is the strongest correction for mass registration. C winsorizes only the extreme upper tail and is the primary specification. A conclusion is called address-robust only when its direction survives all three specifications under the stated inferential rule. The three specifications are interpreted as sensitivity bounds, not as competing estimates from which the most favorable result may be selected.

3.5. MSP-MBKMeans Center Discovery

Figure 1 summarizes the method. Center discovery uses total registration weight only. R/S/F/H labels are withheld until centers have been estimated, preventing the outcome comparison from influencing the spatial reference system.
For each city, resolution-9 weights are aggregated to resolutions 8 and 7. At a cell g, the peak score is the total registration weight in its occupied second-order hexagonal neighborhood, Q g = h N 2 ( g ) W h . A cell is a local peak if no neighboring score is larger; tied plateaus retain only the lowest H3 identifier. This deterministic tie rule prevents one flat high area from becoming several peaks. Resolution-8 and resolution-7 peaks are found separately. A resolution-8 peak is persistent only when its resolution-7 parent or one of that parent’s first-order neighbors is a resolution-7 peak. Persistent candidates are ranked by Q g and limited to 96. The final rule has no relative peak-height cutoff.
Candidates are considered in that order. A candidate a must be at least d min = max ( 2 km , 0.08 r 90 ) from every retained peak, where r 90 is the weighted 90th-percentile city radius. For candidate peaks a and b, their valley ratio is
V a b = min g P a b { a , b } Q g min ( Q a , Q b ) ,
where P a b is the shortest H3 path between them. A separate center is permitted only when V a b 0.50 ; adjacent peaks have ratio one, and an unavailable path is assigned ratio zero. The tentative nearest-seed allocation must give every seed at least 8% of fitting weight, and at most eight seeds are retained. Thus a small satellite is not promoted merely because it is a local maximum.
The retained peaks determine both starting positions and candidate K for weighted MiniBatch K-Means. Fitting uses a city-local equal-distance projection and the radial range containing 99.5% of cumulative weight, preventing isolated remote records from pulling the fitted centers. MiniBatch K-Means uses batch size 4,096, at most 300 iterations, one supplied initialization, fixed seed 20260808, and no random reassignment. If a fitted center carries less than 8% of fitting weight, it is removed and the fit is repeated. Centers are ordered by fitted share, with rank 1 defined as the primary registration center. Table 2 records the complete fixed rule. This design keeps the large-sample advantage of MiniBatch K-Means [6] while making center number and starting positions depend on visible spatial structure.

3.6. Spatial-Block Stability with Repeated Selection of K

A fixed K can produce apparently stable centers even when the data do not clearly support that center count. We therefore multiply each H3 resolution-7 block by an independent mean-one Gamma ( 16 , 1 / 16 ) weight and repeat the complete peak selection, candidate filtering, K selection, and MiniBatch fit 20 times. Repeated centers are paired with full-sample centers by the minimum-total WGS84 geodesic distance assignment. The consensus and four reliability statuses are defined in Table 2; unresolved cities do not enter formal center–periphery comparison, although exploratory maps and numerical diagnostics remain available.
Twenty repetitions make appearance rates observable in five-percentage-point increments. They provide a reproducible grading diagnostic rather than a precise probability estimate. A separate repetition-count sensitivity analysis compares 20, 50, and 100 repetitions in a structured sample of 12 cities selected to span all four reliability levels and the observed range of H3 input sizes; its results are reported in Section 4. The fixed random seed makes the stated analysis reproducible. The spatial perturbations assess whether the center structure persists; they do not enumerate every possible random ordering of MiniBatch updates.

3.7. Computational Scale

The 185,033,392 entities occupy 19,175,721 exact coordinate sites and 2,905,086 H3 resolution-9 cells in the 345 observed units. Replacing repeated coordinates by one coordinate and its multiplicity preserves their exact K-means contribution, while H3 resolution 9 supplies the finer spatial representation used for center fitting. A city has a median 6,494 occupied fitting cells and a maximum of 51,414, so the national archive is not held in memory as one point matrix.
Let G c denote occupied fitting cells, P c 96 retained candidates, K c 8 centers, q 4 , 096 batch size, T MiniBatch updates, and B spatial repetitions in city c. Nested H3 aggregation and bounded-neighborhood peak scoring grow approximately with G c ; the capped candidate comparisons do not require all cell pairs. MiniBatch fitting requires approximately O ( T q K c ) calculations, and spatial repetition scales this term approximately by B. Exact-site class summaries grow with the city’s site count, weighted quantiles add a sorting step, and density-priority zone construction grows approximately as O ( G 8 c log G 8 c ) . Cities are independent after their source assignment is fixed and can therefore be evaluated concurrently. The reported analysis was feasible on an Apple M4 workstation with 10 processor cores and 16 GB memory, with at most four cities evaluated concurrently and numerical calculation for each city restricted to one processor core; method-comparison times for identical synthetic samples are reported in Table 4.

3.8. Center–Periphery Outcomes and Inference

For every eligible city, class, and address specification, we calculate weighted mean and quartile distances to the primary and nearest centers; shares within 5, 10, and 20 km of the primary center; shares within 5 and 10 km of any center; core, intermediate, and peripheral shares defined by city-wide distance quartiles; and allocation across center catchments. The principal distance contrast for class m { S , F , H } is
G m , c = E ( D R , c ) E ( D m , c ) .
A positive value means that class m is closer to the registration center than R in city c. We also report the probability-scale index
C m , c = 2 Pr ( D m , c < D R , c ) 1 ,
which ranges from 1 to 1 and equals zero when neither distribution has a centrality advantage.
City-level uncertainty is estimated by deleting H3 resolution-6 blocks in turn. If G ^ ( b ) is the gap after deleting block b, the jackknife standard error is [ ( B 1 ) B 1 b ( G ^ ( b ) G ¯ ( ) ) 2 ] 1 / 2 and the interval is G ^ ± 1.96 S E . Formal assessment requires at least 100 raw class entities, eight occupied class blocks, and eight usable delete-block estimates. A two-sided p value is reconstructed from the symmetric normal interval. Benjamini–Hochberg control at q = 0.05 is then applied separately to the 345 S, F, and H city families. A city is formally called more central or more peripheral only when its A contrast survives this correction and the A/B/C point estimates all have that direction. A corrected A contrast followed by an A/B/C reversal is called address-sensitive; all other supported cases are no clear difference. Unresolved center structure overrides the distance result. The catalogue therefore reports corrected city statements, not unadjusted screening labels.
The interval procedure was checked before use with fixed-total random labels and a known spatial shift. Under the null, 1,000 simulations produced a 1.0% rejection rate and 99.0% coverage; under the alternative, 500 simulations produced 99.8% detection and 100.0% correct direction. A separate null experiment with 2,000 simulations for the model-gain block procedure gave 4.7% rejection and 95.3% coverage. These simulations check calibration under their stated designs; they do not make the observed archive a random population sample.
National estimates are medians of city effects and shares of cities with positive effects. Their 95% intervals use 5,000 province-cluster bootstrap repetitions. A national directional conclusion requires all A/B/C median effects to share a sign and all corresponding cluster intervals to exclude zero. Structure, size, and province summaries are descriptive stratifications; they are not causal moderators.

3.9. Registration Agglomeration Zones

Occupied resolution-8 cells are allocated to the nearest fitted center, creating non-overlapping Voronoi-like catchments. Within each catchment, C-weight density is computed per square kilometer. A density-priority frontier grows outward from the highest-density occupied cell near the center until it carries 50% of catchment C weight. One empty cell may be bridged for morphological closure only if it belongs to the same nearest-center catchment. When the anchored patch is insufficient, a secondary patch is admitted only if it carries at least 5% of catchment weight. If qualifying patches cannot reach 50%, the zone is reported as incomplete rather than expanded across large empty areas.
For each zone we report area, raw entities, C weight, density, 90th-percentile radius, city share, and R/S/F/H composition. Class enrichment uses the location quotient L Q m = p m , z / p m , c . A class is called prominently enriched only when L Q m 1.20 and exceeds the upper 95% bound from the hypergeometric random-label distribution conditional on city class totals and zone capacity. This dual threshold prevents enormous sample size from turning negligible compositional differences into substantive claims.

3.10. Validation Design

The first validation set contains six synthetic geometries repeated 20 times each: compact one-center, elongated one-center, balanced two-center, imbalanced two-center, balanced four-center, and a main center with a 4% satellite. A second stress set contains nine additional geometries covering overlapping centers, unequal scales, three and six centers, a curved non-convex field, uniform noise, a concentrated address spike, and satellites immediately below and above the substantive share threshold. Each stress geometry is also repeated 20 times, giving 300 synthetic cities. Expected K and generating centers are known. MSP-MBKMeans is compared on identical coordinates with the legacy one-standard-error MiniBatch K-Means selector, full-covariance Gaussian mixtures selected by BIC, HDBSCAN, mean shift, and OPTICS. Baseline settings are fixed by method rather than selected separately for each scenario: the mixture searches one to eight components; HDBSCAN uses an 8% minimum cluster size and 10-point core rule; mean shift estimates one common 20th-percentile bandwidth; and OPTICS uses a 1% neighborhood minimum, an 8% minimum cluster size, and ξ = 0.05 . Accuracy of K, matched center-position error, and computation time are recorded. Five ablations remove cross-scale persistence, the density-valley rule, the minimum-share rule, adaptive separation, or the 99.5% fitting range one at a time.
Because the expanded stress set informed the final parameter selection, it is treated as model-selection evidence rather than an independent evaluation. A separate synthetic evaluation was specified after the algorithm and thresholds had been finalized. It contains ten new geometries—rotated anisotropy, unequal three-center and irregular five-center mixtures, 7% and 9% satellites, uniform contamination, an off-center address spike, a two-node coastal arc, four unequal scales, and a skewed one-center corridor—with 20 repetitions each. Its generating scenarios, expected substantive K, random seed, comparison methods, and parameter settings were determined before evaluation and remained unchanged throughout the analysis. All 1,200 method evaluations are reported, including failures.
Parameter sensitivity is evaluated on all 345 observed units. For the substantively important 6/8/10% minimum-share choice, the complete analysis is repeated: 6% and 10% each use 20 spatial-block repetitions with full K reselection, every A/B/C class profile is recalculated, and the same class-wise false-discovery-rate rule is applied. Center count is share-stable only when all three selected K values agree. A primary center is share-stable when both alternative centers lie within max ( 2 km , 0.10 r 90 ) of the 8% center; a city-class conclusion is stable only when its corrected label agrees at 6%, 8%, and 10%. These three statements are intentionally separate. Valley ratios of 0.40/0.50/0.60 and adaptive distance fractions of 0.06/0.08/0.10 retain the lighter one-factor comparison because they scarcely changed K in the initial all-city grid.
Center fitting uses all classes, so the statement that labels are not supplied as model features does not by itself rule out reference dependence: a concentrated rare class could still help create the center against which it is compared. At the fixed 8% definition, a leave-class-out diagnostic therefore subtracts the C-weight contribution of S, F, or H, refits the corresponding center system, and recalculates the class contrast. Full-data and leave-class-out fits are both deterministic in this diagnostic; resampling uncertainty remains represented by the primary model. We compare K, primary-center displacement, and the corrected city conclusion. This is a sensitivity check, not a replacement center definition.
Real-data external agreement was assessed against an independently constructed GHSL 2020 settlement reference [21]. For each of the 323 observed units with an accepted auxiliary boundary, the boundary was applied to the complete GHSL grid rather than to registration-occupied cells. The most restrictive SMOD tier among class 30, class 23 or above, and class 21 or above was retained when it contained at least three grid cells and 10% of boundary population. Eight-neighbor components carrying at least 5% of population in that tier were kept, up to 12 per unit; component rank and centroid used GHSL population alone. Registration coordinates, counts, classes, and fitted centers played no part in constructing this reference. Twenty-one units lacked a matched boundary. Tieling was also excluded because its licensed auxiliary polygon contains a disconnected southern component that captures unrelated settlements; no registration coordinate was used to trim or repair that boundary. These 22 units were marked unavailable rather than assigned a registration-derived envelope.
The first comparison pairs the primary registration and GHSL nodes. A second uses minimum-cost one-to-one matching between all centers in the same unit. Because this produces only min ( K reg , K GHSL ) pairs, unmatched nodes and distance shares using both matched pairs and all evaluated registration centers are reported separately. GHSL reflects population and built settlement rather than legal registration, so these distances test cross-source positional agreement, not the true registration-center count. The auxiliary boundaries are older than the registration snapshot and are used only for this external check, never to reassign source records. Address sensitivity and spatial-resampling status provide two further checks: the first asks whether shared registration addresses change the result, and the second asks whether the same center structure survives neighborhood-level perturbation.
The GHSL construction was developed during manuscript review and is therefore a post-review external check, not a preregistered validation. Its one-factor sensitivity varies the tier population requirement from 10% to 5/15%, component share from 5% to 3/8%, and node cap from 12 to 8. Every setting rebuilds the reference from GHSL and boundaries alone. We compare selected tier, GHSL node count, primary-reference displacement, primary registration distance, and all-node matching, thereby separating primary-location robustness from secondary-node sensitivity.

4. Results

4.1. Data Conservation and National Coverage

The formal analysis layer contains 185,033,392 deduplicated registration entities. The same total is conserved when entities are aggregated to exact sites, H3 fitting cells, city profiles, center catchments, and registration agglomeration zones. The final city inventory contains all 364 source city-level units: 345 have usable records and 19 are retained as explicitly non-estimable because no source sample was available. This distinction matters. The national findings describe the 345 observed units and must not be interpreted as evidence that the 19 remaining units lacked registered entities in reality.
All city-center models use the total entity distribution without reference to R/S/F/H. The rare-class eligibility rule later yielded 307 formal city comparisons for S, 263 for F, and 220 for H. H spans 30 province-level clusters, whereas S and F span 31. Variation in eligible counts reflects the minimum requirement of 100 raw class entities across eight spatial blocks, not discretionary removal after inspecting effects.

4.2. Synthetic Recovery, Competing Methods, and Ablations

Table 3 shows that MSP-MBKMeans recovered the prespecified number of centers in all 120 synthetic evaluations (100.0%). The legacy one-standard-error MiniBatch K-Means selector was correct in 20 of 120 evaluations (16.7%). Its characteristic error was over-selection: clear one-, two-, and four-center patterns were frequently split into six to eight clusters. The comparison used exactly the same simulated coordinates and weights, so the difference is attributable to the selection framework rather than easier scenarios for MSP-MBKMeans.
The expanded benchmark changes the strength, but not the direction, of the algorithmic claim (Table 4). MSP-MBKMeans, HDBSCAN, and OPTICS all recovered K in 100.0% of the six original scenarios. In the nine stress scenarios, HDBSCAN and OPTICS reached 88.9%, while MSP-MBKMeans reached 87.8%. MSP-MBKMeans failed deliberately on the curved non-convex field by returning several persistent modes and missed one of 20 overlapping and one of 20 unequal-scale pairs. HDBSCAN and OPTICS instead merged every overlapping pair. The methods therefore have different failure boundaries; the experiment does not support universal superiority.
The ablations identify which constraints carried the strongest observed protection. Removing the 8% minimum-share rule reduced original-scenario accuracy from 100.0% to 82.5% and stress accuracy from 87.8% to 46.1%. Removing cross-scale persistence reduced the two rates to 83.3% and 78.3%; removing the density-valley rule reduced them to 97.5% and 81.1%. Adaptive separation had a small aggregate effect in this benchmark (87.2% under stress), while removing the 99.5% fitting range changed no synthetic K. The latter controls remain relevant to geographically extensive real units and remote outliers, but the synthetic evidence should not be overstated.
The prespecified synthetic holdout provides the stronger generalization check. MSP-MBKMeans recovered the substantive center count in 198 of 200 new cities (99.0%), with a 149.5 m median matched-center error when K was correct. Its two misses occurred in the unequal-scale four-center geometry. HDBSCAN and OPTICS each reached 100.0%, while GMM-BIC, mean shift, and the legacy selector reached 60.0%, 58.5%, and 20.0%, respectively (Table 5). The result supports strong recovery in the prespecified holdout while preserving the narrower claim that MSP-MBKMeans is not the universally most accurate clustering method. Because the author designed these scenarios after examining the method, this remains an internal holdout rather than an externally preregistered benchmark.
The real-city threshold analysis separates center-count sensitivity, primary-position sensitivity, and class-location stability. The initial one-factor grid showed that changing the minimum share from 8% to 6% or 10% changed raw K much more often than changing the valley or distance settings (Table 6). The share rule was therefore subjected to the complete confirmatory analysis rather than left as a local check. After 20 spatial repetitions, the consensus-K medians were five, four, and four at 6%, 8%, and 10%; agreement with the fixed 8% K was 39.7% and 47.2% at the two alternatives. Only 246 cities (71.3%) kept the primary center within the adaptive displacement limit at both alternatives. Corrected city-class conclusions agreed at all three shares in 310 S comparisons (89.9%), 313 F comparisons (90.7%), and 316 H comparisons (91.6%), or 939 of 1,035 comparisons (90.7%) overall. Table 7 reports the full corrected directional counts, and the catalogue reports the three K values and primary-position status for each city. Thus, the 1,410 centers and the city forms are results of the fixed 8% definition; a stable class statement must not be read as proof that the number or exact position of every center is definition-independent.
In the 345 observed cities, MSP-MBKMeans selected a median of four centers versus six under the legacy method; the legacy selector chose more centers in 76.8% of cities. Figure 2 summarizes this difference with the benchmarks, ablations, and GHSL distances. Four is not a universal urban optimum; it is the median under the stated substantive constraints.

4.3. Observed Center Counts and Reliability

Under the fixed 8% C-weight primary definition, 1,410 centers were identified across 345 cities. The city distribution was K = 1 for 17 cities, K = 2 for 45, K = 3 for 65, K = 4 for 81, K = 5 for 64, K = 6 for 50, K = 7 for 20, and K = 8 for 3. The mode and median were both four. This distribution should not be converted into a development ranking: a geographically large prefecture that contains several county nodes can have a larger K than a compact, economically intensive municipality. It is also conditional on the 8% minimum-share definition, as the separate 6/8/10% catalogue fields make explicit.
Table 8 shows how spatial resampling separated structural information from uncertainty. Exactly stable center counts were obtained for 69 cities. Another 88 retained stable core centers but had optional peripheral centers. For 163 cities, the principal structure remained usable only with an explicit center-count uncertainty label. Twenty-five cities were unresolved and were excluded from formal center–periphery inference. Put differently, 157 cities had either exactly stable counts or stable cores; 320 had at least a qualified usable structure; and the method openly declined to make a formal structural claim for the remaining 25. The lower exact-stability count than in the initial model is a direct consequence of admitting unequal-scale candidate peaks and repeating K selection 20 times under every address specification.
The three address specifications provide a different diagnostic. Fifty-seven cities had identical center counts under A/B/C, 67 passed the full-structure comparison that allows at most one optional center, and 209 had a primary center stable under all three specifications. These numbers explain why center reliability and address robustness are reported separately: a city can have a stable C-weight bootstrap solution yet move its primary center when thousands of entities sharing one exact address are changed from raw multiplicity to address-equal weight.
The repetition-count diagnostic clarifies the resolution of these labels without changing the prespecified 20-repetition primary definition. The structured sample contains the smallest, median-sized, and largest H3 input within each of the four primary reliability levels. Compared with 100 repetitions, the 20-repetition result retained the same K and reliability level in 10 of 12 cities; 50 repetitions retained the same K in 10 and the same reliability level in 11 (Table 9). All three cities sampled from the exactly stable group remained exactly stable, and all three sampled unresolved cities remained unresolved. Changes concerned optional centers or the boundary between core-stable and usable-with-uncertainty. Median primary-center displacement was effectively zero because most primary centers were unchanged, but the largest displacement reached 3.54 km at 20 repetitions and 2.78 km at 50. These results support the reliability grades as coarse diagnostics while confirming that their numerical appearance rates are not precise probabilities.

4.4. External Positional Agreement with GHSL

The independent comparison covers 323 observed units with an accepted auxiliary boundary; 22 observed units are unavailable. The MSP-MBKMeans primary registration center was a median 3.32 km from the population-weighted GHSL 2020 primary node, with a 7.31 km 75th percentile. Overall, 64.4% of evaluated cities were within 5 km, 82.7% within 10 km, and 90.1% within 20 km. Median and 75th-percentile distances were 2.11 and 3.62 km for exactly stable cities, 5.11 and 10.97 km for stable-core cities, 3.09 and 6.23 km for usable-with-uncertainty cities, and 4.36 and 8.01 km for unresolved cities.
The all-center comparison evaluates 1,349 registration centers and 1,337 independently constructed GHSL nodes in those 323 units. Minimum-cost matching produced 1,130 pairs and left 219 registration centers and 207 GHSL nodes unmatched. Among matched pairs, median separation was 6.66 km; 40.3% were within 5 km, 63.4% within 10 km, and 81.2% within 20 km. Using all 1,349 evaluated registration centers as the denominator, the corresponding confirmed-coverage shares were 33.7%, 53.1%, and 68.1%. Of 1,026 evaluated secondary registration centers, 810 were matched and 216 were unmatched; matched secondary nodes had an 8.69 km median separation, with 54.7% within 10 km and 76.4% within 20 km. These results support broad positional agreement but do not justify claiming that every fitted secondary node has an external counterpart or that GHSL supplies the true registration K.
Large deviations were not automatically classified as failures. The ten largest accepted comparisons involved Aba Tibetan and Qiang Autonomous Prefecture (227.1 km), Qamdo (225.9 km), Heihe (194.0 km), Yichun in Heilongjiang (139.5 km), Altay Prefecture (124.2 km), Suihua (108.5 km), Huanggang (95.8 km), Shangluo (83.2 km), Gannan Tibetan Autonomous Prefecture (68.4 km), and Dingxi (66.1 km). Many are extensive administrative territories with separated settlements. Tieling’s former 221.8 km value was removed after review showed that a disconnected auxiliary-boundary component captured unrelated southern settlements. GHSL and MSP-MBKMeans can otherwise emphasize different nodes because one represents population-weighted settlement morphology and the other cumulative registration density.
The post-review GHSL sensitivity separates a stable primary comparison from definition-dependent secondary counts (Table 10). Moving the tier population requirement to 5% or 15% left the primary-distance median at 3.32 and 3.29 km and 82.7% and 82.4% within 10 km. Changing component share from 5% to 3% or 8% did not move any primary node, but changed the reference count from 1,337 to 1,690 or 918 and changed all-node 10 km agreement from 63.4% to 64.5% or 70.6%. The apparent improvement under 8% follows from retaining fewer secondary nodes; it is not evidence that 8% is truer. Capping at eight nodes changed K in only 1.9% of units. Primary positional agreement is robust in this grid, whereas secondary-node enumeration remains conditional on the component definition.

4.5. What Spatial Patterns Do the Four Classes Actually Show?

The class labels do not enter center fitting. More central means a shorter mean distance than R with a direction supported across A/B/C; more peripheral means longer. No clear difference means no repeatable ordering, not identical distributions, and address-sensitive means that shared-address treatment changes the direction.
The city counts in Table 11 make the national pattern concrete after multiplicity control. F is more central in 79 cities and more peripheral in 8; H is more central in 82 and more peripheral in 3. Both F and H are more central in 41 cities. Within those 41, S shows no clear difference from R in 27, is more peripheral in 6, and is also more central in 8. S alone is more central in 24 cities and more peripheral in 25, so neither “state entities are always central” nor “state entities are always peripheral” describes the country. In 82 cities, S, F, and H all show no clear difference from R. This does not prove identical distributions; it says that the available evidence does not support a repeatable ordering after address checks and false-discovery-rate control.
The correction deliberately changed the evidential status of many screening results. Before correction, the 1,035 city-class comparisons contained 234 more-central, 60 more-peripheral, and 18 address-sensitive screening labels. The corresponding formal counts are 185, 36, and 7. These are not lost observations: most moved to “no clear difference” because their individual intervals were not strong enough relative to the number of city comparisons. Among formally eligible comparisons, usable delete-block counts were never close to the minimum: minima for S/F/H were 17/55/55 and medians were 459/476/466.
Table 12 reports the central national result. F had positive A/B/C primary gaps of 3.25, 2.47, and 3.74 km; H had 4.39, 3.28, and 4.57 km, with every interval excluding zero. Address equality (B) attenuated but did not erase either pattern. The probability indices show modest distributional shifts, not percentages located downtown. S had gaps of 0.04, 1.01 , and 0.11 km, so its signs and intervals did not support a national direction.
Under C, nearest-center gaps were 1.06 km for F and 1.37 km for H, smaller than their primary-center gaps. Figure 3 therefore supports orientation toward the dominant registration node, not a causal location preference.

4.6. Strict Reliability and Address-Robustness Subsets

The main F/H result survived progressively narrower quality subsets. Among cities with exactly stable counts or stable cores, the C-specification median gap was 3.20 km for F ( n = 129 ) and 4.00 km for H ( n = 107 ); S remained near zero at 0.24 km ( n = 150 ). Among cities whose primary center was address-robust, medians were 3.99 km for F ( n = 169 ), 4.76 km for H ( n = 136 ), and 0.03 km for S ( n = 190 ).
The most restrictive subset required both an exactly stable center count and an A/B/C-robust primary center. F retained a 2.49 km median advantage across 43 cities, with a 95% interval of [1.57, 6.75] km. H retained a 3.46 km advantage across 35 cities, with [1.14, 8.59] km. S was 0.21 km across 49 cities, with [ 2.19 , 1.79] km. Thus, the F/H finding was not driven by unresolved center structures or by cities whose primary center depended on concentrated addresses. Conversely, the absence of a uniform S direction was reproduced under the strictest quality conditions.
At the city-class level, the corrected 1,035 comparisons comprise 185 more-central, 36 more-peripheral, 7 address-sensitive, 562 no-clear, and 245 not-estimable results. The 41-city joint F/H pattern and the 82 cities with no clear S/F/H difference are evidence categories, not city rankings. The complete threshold propagation further shows that 939 of 1,035 labels (90.7%) are identical at 6%, 8%, and 10%. Results that change are explicitly marked in the city catalogue rather than removed.
The leave-class-out check shows that the main reference is rarely created by the class being compared (Table 13). Removing S, F, or H preserved deterministic K in 314, 332, and 332 cities and preserved the corrected city conclusion in 342, 343, and 343 cities. Across all 1,035 comparisons, 1,028 conclusions (99.3%) agree. The seven changes occurred in Taizhou (S), Huainan (S), Hebi (H), Zhongshan (S and H), Qiannan (F), and Wenshan (F). They remain flagged in the catalogue; the high overall agreement does not license hiding those exceptions.
The conclusion was also stable to the mixture of administrative units in the source inventory. Under C, the full 364-unit inventory yielded median gaps of 0.11 , 3.74, and 4.57 km for S, F, and H. Restricting the analysis to 330 prefecture-level units yielded 0.24 , 3.74, and 4.66 km; restricting it further to 291 standard prefecture-level cities yielded 0.15 , 3.73, and 4.66 km. In all three inventories the median selected center count was four. Table 14 therefore shows that F/H centrality and the absence of a common S direction are not created by municipalities, autonomous prefectures, regions, or directly administered county-level units.

4.7. Variation by Center Structure, Archive Size, and Province

Figure 4 shows positive F/H medians in all reportable center structures and weaker gaps in the lowest size quartile, but no monotonic size relationship. Province medians were positive in 22/23 eligible F and 15/18 eligible H province units; S split 15 positive versus 12 negative. These are descriptive distributions, not effects of structure, size, or province.

4.8. Registration Agglomeration Zones

The 1,410 centers produced 1,410 catchments covering 1,571,280 occupied cells and conserving all entities. A total of 1,359 zones reached the 50% C-weight target; 51 remained incomplete. All 64,975 one-cell bridges stayed within their originating catchment, so none joined different center systems.
Combined zones carried a median 48.7% of raw city entities (IQR 46.2–51.2%) at a median density of 482 entities/km2 (Figure 5). Ninety-nine zones belonged to unresolved center structures and remained exploratory; 45 zones attached to otherwise reliable centers were incomplete.
Prominent enrichment required both L Q 1.20 and the conditional random-label upper bound; it describes unusual within-city composition, not an industrial cluster or causal formation process.

4.9. Eight Illustrative Cities and Complete City Outputs

Figure 6 and Figure 7 show Beijing, Shanghai, Wuhan, Chengdu, Xi’an, Chongqing, Guangzhou, and Shenzhen using the same visual grammar. The maps display estimated centers, nearest-center catchments, and high-density registration zones. They were chosen to illustrate different administrative extents and center structures, not because they had the strongest effects. Guangzhou is particularly useful as a caution: its unresolved status means that mapped centers remain exploratory rather than a formal city structure conclusion.
Supplementary Table S1 gives the corresponding conclusion for every source unit. Each entry states the center form, the S/F/H pattern relative to R, the concentration carried by registration zones, and the reliability limit. This is the main city-level product of the study: it distinguishes a substantive finding from an unavailable or unresolved comparison without turning the cities into a ranking. Keeping the catalogue as a separately paginated, searchable part of the article package preserves all 364 results without interrupting the national argument.

5. Eight Cities in Plain Language

The eight cities in Figure 6 and Figure 7 show how the same measures lead to different, readable conclusions. They are examples rather than a ranking. Table 15 reports the evidence used in the descriptions. A positive class gap means that the class is closer to the primary registration center than R.
Beijing. One strong primary center carries 70.0% of fitted weight and is accompanied by two secondary centers whose status is less certain. S, F, and H are all closer to the primary center than R, by 5.37, 9.23, and 7.53 km. Three registration zones carry 48.4% of entities. The clear finding is four-class centrality; the exact secondary-center list is not fixed.
Shanghai. Four centers share the registration pattern more evenly: the primary center carries 51.3%. S, F, and H are all more central than R, with F showing the largest gap at 9.42 km. The primary center is 7.36 km from the independent GHSL reference, so registration geography and settlement geography are close but not identical.
Wuhan. All C-weight resamples support two centers, led by a primary center carrying 65.6% of fitted weight. S, F, and H are all more central than R, by 3.77, 3.02, and 4.62 km. The two registration zones carry 45.5% of entities at high density. The center count is stable even though address-equal weighting can add one further minor center.
Chengdu. A strong primary center carries 76.4%, with two smaller centers and some uncertainty about the exact count. F and H are clearly closer to the primary center than R; S shows no clear difference. This is the common national pattern in a concrete city: international classes are central, while S does not follow a fixed direction.
Xi’an. This is the cleanest two-center case: all three address specifications select two centers and the count is stable. F and H are closer to the primary center than R, but S is 2.19 km farther away. Xi’an shows why the national S result is not a weak version of F/H; its direction can genuinely differ by city.
Chongqing. Two C-weight centers organize an exceptionally large municipal area, while the A/B specifications select three and four. F and H are about 60 km closer to the primary center than R, while S has no clear difference. These large values reflect Chongqing’s spatial extent and should not be compared directly with compact cities. Its zones cover a large area but have low density.
Guangzhou. Three density peaks are visible and the primary peak is close to GHSL, but the center system changes too much under spatial perturbation. The apparent class gaps are therefore not formal findings. Guangzhou is included to show that a plausible map does not override unstable center evidence.
Shenzhen. Four relatively balanced centers accompany the highest zone density among the eight cities. Yet S, F, and H show no clear distance difference from R. High registration density therefore does not automatically produce strong separation among historical entity classes.
Together, the cases show that a city conclusion needs three pieces of information: center form, class position, and evidence strength. Any one of these alone can mislead. The complete catalogue applies this same plain-language structure to all 364 source units.

6. Discussion

6.1. What the National Pattern Means

The main empirical result is straightforward. Legal addresses identified as foreign (F) or as Hong Kong, Macao, or Taiwan (H) more often lie near the principal concentration of all registrations. After class-wise false-discovery-rate control, F is more central in 79 cities and more peripheral in 8, while H is more central in 82 and more peripheral in 3. The city-equal typical differences are 3.74 and 4.57 km and keep the same national direction under three treatments of shared addresses. State-identified registrations (S) have no common national direction, being more central in 24 cities and more peripheral in 25. These are shifts in address distributions, not claims that every F or H entity is central, that R avoids centers, or that a central address improves performance.
Three recurring city patterns make the national result concrete. In 8 cities, S, F, and H are all closer to the primary center than R. In 27, F and H are closer while S has no clear difference; in 6, F and H are closer but S is farther out. In 82 cities, none of S/F/H has a clear difference from R after correction. Other cities show only one class difference, address or center-share sensitivity, insufficient observations, or an unresolved center system. This classification identifies intelligible regularities without inventing an economic story for every city.
Several mechanisms could produce F/H centrality: access to professional or administrative services, finance, transport, bilingual networks, established foreign-business districts, recognizable central addresses, or centrally located registration agents. The data cannot distinguish them. Likewise, the mixed S pattern may combine central administrative, financial, and headquarters registrations with peripheral utilities, industry, transport, resources, and development-zone entities. The surviving S label is too coarse to separate these forms. Both interpretations are hypotheses for linked industry, administrative, and longitudinal data, not mechanisms established here.
F/H differences are larger relative to the primary center than to the nearest center. An entity in a polycentric city may be close to a secondary node without being close to the dominant citywide node. Reporting both distances therefore preserves center hierarchy while recognizing meaningful secondary concentrations.

6.2. Why Concentrated-Address Sensitivity Matters

Mass registration at one legal address can alter the density surface, the primary center, and class distances. If a service address hosts many F or H records, raw counts may help create both the reference center and the apparent class advantage, even though class labels do not enter center fitting.
The A/B/C design makes this risk visible. A counts every entity, B gives each exact location equal weight, and C bounds only extreme multiplicities. B is not automatically more truthful: a genuine complex containing many independent organizations need not count the same as an isolated address. C is the prespecified compromise, while agreement across all three specifications supports the strongest interpretation. F/H effects shrink under B but remain positive with intervals excluding zero; under C they are similar to or somewhat larger than under A. Thus some concentrated addresses may amplify raw patterns, while bounding extreme sites can also prevent unrelated mass registrations from shifting the center.
Only 7 of 1,035 city–class comparisons retain an address-sensitive label after multiplicity control, down from 18 screening labels, but this does not prove that address bias is absent: 245 comparisons are unavailable or unresolved, and many others show no clear difference. A/B/C controls observed coordinate multiplicity; it cannot determine whether any unique address is an operating location. The leave-class-out check addresses a related concern: 1,028 of 1,035 conclusions remain unchanged after the compared class is removed from center construction. The seven exceptions remain named in the catalogue.

6.3. Methodological Contribution and Its Boundary

The study does not claim to invent K-means, MiniBatch K-Means, density peaks, hierarchical grids, bootstrap, Voronoi allocation, location quotients, or hypergeometric inference. MSP-MBKMeans contributes a task-specific connection among them. A candidate center must persist across spatial resolutions before clustering and must then satisfy separation, an intervening density valley, and a post-fit share requirement. The expanded benchmark led to removal of an early relative peak-score cutoff, retaining original recoveries while reducing missed unequal-scale centers. The final method rejects elongated single-center splitting and a 4% satellite that generic held-out quantization error over-partitions.
Uncertainty is part of the output. Re-selecting K in every spatial-block bootstrap replicate evaluates whether the reported center system itself is repeatable. A separate A/B/C comparison evaluates sensitivity to shared-address weighting. Their combination distinguishes exactly stable systems, stable main centers with optional minor nodes, broader count uncertainty, and unresolved cities; it also provides a strict subset for testing the F/H result. The 25 unresolved cities remain visible rather than being removed.
The method then converts centers into catchments, density-priority zones, and conditional class enrichment for every source unit. This end-to-end structure is useful because each city result can be traced to explicit spatial and reliability conditions. It is nevertheless a domain-designed assembly of established methods, not a universal clustering theorem. MSP-MBKMeans tied HDBSCAN and OPTICS in the original six scenarios but was one percentage point lower in the expanded stress set. Its advantage is an explicit, substantively constrained K, scalable MiniBatch K-Means fitting, and repeatable reliability reporting. Curved non-convex fields, other registration systems, or other grid hierarchies may require different thresholds or a non-K-means fitting stage.

6.4. Center Counts and Independent Positional Evidence

The median of four centers is not an urban ranking. Administrative units differ in area, settlement hierarchy, terrain, and included counties; a large autonomous prefecture may contain several legitimate county-seat nodes, while a compact municipality may have one dominant registration concentration. K means only the number of registration nodes supported within the source unit under the fixed 8% rule. Reliability is therefore essential: among 345 observed cities, 69 have exactly stable counts, 88 have optional peripheral centers, 163 have broader count uncertainty, and 25 are unresolved. The 6/8/10% sensitivity analysis adds a different uncertainty: only 39.7% and 47.2% retain the 8% count at 6% and 10%, respectively. The catalogue reports this definition sensitivity separately from fixed-rule resampling stability.
The independent GHSL comparison supports primary-center position without defining registration-center truth. Across 323 boundary-supported units, the primary-node distance has a 3.32 km median and 82.7% are within 10 km. Among all nodes, 63.4% of 1,130 matched pairs are within 10 km, while 219 of 1,349 evaluated registration centers remain unmatched. Of the secondary centers, 810 are matched and 216 are unmatched; 54.7% of matched secondary centers lie within 10 km. Reference-parameter changes leave primary agreement nearly fixed but alter secondary-node counts, so the evidence is stronger for primary position than for secondary count. Tieling is excluded because its auxiliary boundary contains an anomalously disconnected component. GHSL measures built settlement, not legal registration; employment, activity, or verified local landmarks would test different meanings.

6.5. What Registration Zones Add

Centroids do not show how far a concentration extends. The catchment-relative zones add a footprint to each node and carry a median 48.7% of city registration weight, close to the intended half-weight target. The 51 incomplete zones identify catchments that cannot reach the target without crossing excessive empty space or accepting tiny fragments; forcing completion would create smoother but less defensible maps. No morphological bridge crosses a center boundary.
Conditional LQ measures composition relative to the same city. A zone with many F entities need not be enriched if F is already common citywide, while a smaller count can be enriched where F is rare. Requiring both a 20% proportional elevation and a random-label exceedance improves interpretability, but still measures composition rather than supplier links, knowledge exchange, or production. Used under the correct label, zones can guide registry-quality review, comparisons with employment or POI centers, and targeted field validation; they cannot by themselves establish transport or infrastructure demand.

6.6. National Inference, Boundaries, and Limitations

Equal city weighting estimates the typical eligible city contrast, not the experience of a randomly selected entity. Province-cluster intervals acknowledge regional dependence but cannot remove the modifiable areal unit problem [58]. Source units mix municipalities, prefecture-level cities, autonomous prefectures, and special county-level units. Because local polygons were outdated, they were not used to reassign entities; this preserves the source inventory but retains administrative heterogeneity. Restricting the analysis to prefecture-level units and then standard prefecture-level cities leaves C-specification F/H gaps near 3.7/4.7 km and S near zero. Administrative composition therefore does not explain the main direction, although boundaries within units can still matter. Structure, size, and province splits remain descriptive rather than causal.
The largest limitation is semantic: registration location is not verified activity location. Address weighting cannot support claims about employment, output, commuting, emissions, land use, or infrastructure. The snapshot is cumulative, contains multiple registration states, and lacks a complete event history, so it cannot reveal migration, center formation, survival, or policy effects. Class labels are also coarse: the lost original organization-type field prevents recovery of 429 mapping strings; R is private-dominated rather than purely private; and S/F/H are historical identified classes rather than current ownership verification.
Geocoding quality cannot be reconstructed without original addresses. BD-09 to WGS84 conversion removes a systematic offset, and H3 aggregation plus the 99.5% fitting range reduces isolated anomalies, but neither removes city-specific error. Thresholds also embody domain judgment. The complete 6/8/10% sensitivity analysis preserves 90.7% of city-class labels, but center count agrees with the 8% definition in only 39.7% and 47.2% of cities at 6% and 10%, and only 71.3% retain the primary center under the joint displacement rule. The selected 8% is an explicit substantive definition, not a unique optimum. Likewise, the 12-city repetition-count diagnostic preserves the reliability level in 10 cities at 20 versus 100 repetitions and in 11 at 50 versus 100; the four levels are therefore practical evidence grades, not precise stability probabilities. The catalogue marks each kind of threshold sensitivity separately. The GHSL check covers 323 of 345 observed units, was constructed during review rather than preregistered, uses older auxiliary boundaries, and compares settlement with registration. Twenty-one units lack a matched boundary and Tieling is excluded. Finally, the entire study is descriptive: sector, history, capital, policy, rent, transport, services, and data practice can all correlate with center distance, so F/H differences are not ownership effects or causal preferences.

6.7. Future Research

The most useful extension is targeted linkage rather than another nationwide raster. Establishment and status dates could create cohorts; industry codes could test composition; verified operating addresses could estimate registration-to-activity displacement; and rent, transport, development-zone, and administrative-service data could test mechanisms. Method development can compare alternative grids, estimate center count and location jointly, and validate against employment, POI, transit, nighttime light, and local records. The complete 364-unit catalogue makes a balanced validation sample possible: cities can be selected across reliability, GHSL distance, structure, size, and address robustness instead of checking only famous cases.

7. Conclusions

This study turns 185 million cumulative historical registration coordinates into three readable outputs: each city’s number of registration centers and its reliability, the position of S/F/H relative to the private-dominated R class, and the high-density registration zones around those centers. Under the fixed 8% primary definition, 69 of 345 observed units have an exactly stable center count, 88 have stable main centers with optional small nodes, 163 have a readable main pattern with count uncertainty, and 25 remain unresolved. The 1,410 fitted centers support 1,410 zones, of which 1,359 reach the prespecified coverage target without crossing center boundaries. The 6/8/10% sensitivity analysis shows why this definition must remain visible: center counts agree with the 8% result in only 39.7% and 47.2% of cities at 6% and 10%, while primary-center position remains stable at both alternatives in 246 cities (71.3%).
The four-class result is equally concrete after correction for 1,035 city comparisons. F is more central in 79 cities and more peripheral in 8; H is more central in 82 and more peripheral in 3. Their city-equal typical advantages over R are 3.74 and 4.57 km. Both are more central in 41 cities: S has no clear difference in 27, is more peripheral in 6, and is also more central in 8. Across all cities, S is more central in 24 and more peripheral in 25, while 82 cities show no clear S/F/H difference from R. The main national regularity is therefore F/H orientation toward the primary registration node; S varies by city, and many cities have no supported class ordering. Restricting the inventory to prefecture-level or standard prefecture-level cities leaves the national directions nearly unchanged.
This finding has social-economic value without requiring a causal story. Historically identified internationally connected registrations are more closely associated with the leading legal-address concentration, potentially through professional and administrative services, established foreign-business districts, recognizable addresses, or registration intermediaries. Central administrative and headquarters-type S registrations may coexist with peripheral utilities, industry, transport, resources, and development-zone entities. Industry, organization subtype, and longitudinal data are needed to distinguish these explanations.
The city catalogue is a core output, not an appendix of raw indicators. All 364 source units receive an explicit conclusion: 345 rows describe spatial form, class position, registration concentration, and confidence in complete sentences; 19 report that no usable coordinates are available. Beijing, for example, has a dominant center with secondary nodes and S/F/H all closer than R; Wuhan has a stable two-center structure with the same class ordering. Urumqi is a stable single-center city where F is closer, S has no clear difference, and H is address-sensitive. Guiyang and Guangzhou have visible density peaks but unresolved center systems, so class gaps are not promoted to formal findings. No-clear, unresolved, and unavailable outcomes are scientific results rather than missing success cases.
MSP-MBKMeans retains MiniBatch K-Means for scalable fitting but constrains how centers enter the fit and how uncertainty is reported. Candidates must persist across two spatial scales, carry a meaningful share, remain separated, and have a density valley between them; K is re-selected in every spatial-block resample. The prespecified holdout achieves 99.0% correct K across 200 new synthetic cities, compared with 100.0% for HDBSCAN and OPTICS. Although the share threshold changes many center counts, 939 of 1,035 corrected city-class labels are identical at 6%, 8%, and 10%, and 1,028 remain identical when the compared class is removed before refitting the center. The catalogue therefore gives each city its three center counts, primary-position status, and class-label sensitivity separately. Independent GHSL evidence supports primary-center position but retains 219 unmatched registration centers. The method is therefore well supported for this registration-center task, not universally superior and not proof of every city’s physical center count.
For future registration-data projects, four principles follow. First, use one documented coordinate system and reconcile city totals before interpretation. Second, compare raw counts, address equality, and bounded extreme-address weights. Third, fit centers from the full registration field without supplying class labels as features, then verify with leave-class-out centers that a rare class did not create its own reference; report structure, class position, zones, and reliability together. Fourth, retain unresolved and no-sample cities, and never interpret a larger K or denser zone as higher economic development. Combining uncertainty, address sensitivity, center-definition sensitivity, and GHSL discrepancy can guide balanced address verification and later linkage to industry, transport, rent, policy, and operating locations.
The conclusion is deliberately narrow. China’s cumulative business-registration archive contains repeatable urban registration centers, although confidence in their number varies by city. F and H registrations are generally more oriented toward the primary registration center than R, whereas S has no national direction. Registration zones locate these concentrations, and the city catalogue states where each conclusion is supported, uncertain, or unavailable. These findings concern legal registration locations, not production, employment, office activity, performance, or causal ownership effects.

Supplementary Materials

The following supporting information can be downloaded at the website of this paper posted on Preprints.org.

Author Contributions

Conceptualization, H.W.; methodology, H.W.; validation, H.W.; formal analysis, H.W.; investigation, H.W.; resources, H.W.; data curation, H.W.; writing—original draft preparation, H.W.; writing—review and editing, H.W.; visualization, H.W.; project administration, H.W. The author has read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable. The study analyzes organizational registration records and contains no experiment involving human participants or animals.

Data Availability Statement

The underlying entity-level records originated from Tianyancha and were lawfully acquired for institutional research use. They cannot be redistributed because of third-party copyright and institutional contractual restrictions. For confidential assessment, the analysis code, synthetic-data generator and specifications, non-identifying city-level aggregate results, and machine-readable bilingual city catalogue will be provided to journal editors and reviewers upon request. Readers may request the same non-identifying research materials from the corresponding author at loveweihaitong@foxmail.com, subject to the original rights and agreements. Enterprise-level records, coordinates, and identities cannot be provided.

Acknowledgments

The author thanks the institutions listed in the affiliations for the professional setting in which the research problem was developed. Language and formatting assistance was used during manuscript preparation; all methodological choices, analyses, interpretations, and final wording were reviewed and approved by the author.

Conflicts of Interest

The author declares no conflict of interest.

References

  1. Wang, Y. Handling of Inconsistencies Between the Registered Address and the Actual Business Address. In Essential Knowledge and Legal Practices for Establishing and Operating Companies in China; Springer, 2022; pp. 163–169. [CrossRef]
  2. Tezer, A.; Bodur, H.O.; Grohmann, B. Location, Location ... Mailing Location? The Impact of Address as a Signal. Journal of Business Research 2021, 130, 683–698. [CrossRef]
  3. Zhao, M.; Liu, X.; Derudder, B.; Zhong, Y.; Shen, W. Big Enterprise Registration Data Imputation: Supporting Spatiotemporal Analysis of Industries in China. Computers, Environment and Urban Systems 2018, 70, 9–23. [CrossRef]
  4. Arribas-Bel, D. Accidental, Open and Everywhere: Emerging Data Sources for the Understanding of Cities. Applied Geography 2014, 49, 45–53. [CrossRef]
  5. Lloyd, S.P. Least Squares Quantization in PCM. IEEE Transactions on Information Theory 1982, 28, 129–137. [CrossRef]
  6. Sculley, D. Web-Scale K-Means Clustering. In Proceedings of the Proceedings of the 19th International Conference on World Wide Web, 2010, pp. 1177–1178. [CrossRef]
  7. Rodriguez, A.; Laio, A. Clustering by Fast Search and Find of Density Peaks. Science 2014, 344, 1492–1496. [CrossRef]
  8. Rousseeuw, P.J. Silhouettes: A Graphical Aid to the Interpretation and Validation of Cluster Analysis. Journal of Computational and Applied Mathematics 1987, 20, 53–65. [CrossRef]
  9. McDonald, J.F. The Identification of Urban Employment Subcenters. Journal of Urban Economics 1987, 21, 242–258. [CrossRef]
  10. Giuliano, G.; Small, K.A. Subcenters in the Los Angeles Region. Regional Science and Urban Economics 1991, 21, 163–182. [CrossRef]
  11. McMillen, D.P. Nonparametric Employment Subcenter Identification. Journal of Urban Economics 2001, 50, 448–473. [CrossRef]
  12. Redfearn, C.L. Persistence in Urban Form: The Long-Run Durability of Employment Centers in Metropolitan Areas. Regional Science and Urban Economics 2009, 39, 224–232. [CrossRef]
  13. Wei, H. Exploring and Practicing the Quantification of Interior Design Colors from an IKEA Design Perspective. Journal of Sensor Networks and Data Communications 2024, 4, 1–12. [CrossRef]
  14. Kloosterman, R.C.; Musterd, S. The Polycentric Urban Region: Towards a Research Agenda. Urban Studies 2001, 38, 623–633. [CrossRef]
  15. Meijers, E. Measuring Polycentricity and Its Promises. European Planning Studies 2008, 16, 1313–1323. [CrossRef]
  16. Burger, M.J.; Meijers, E.J. Form Follows Function? Linking Morphological and Functional Polycentricity. Urban Studies 2012, 49, 1127–1149. [CrossRef]
  17. Wei, H. DecorPGNet: Functional Area Division and Layout Algorithm Model in Living Rooms of Chinese Apartment-Style Family Homes. Civil Engineering Research Journal 2024, 15, 555902. [CrossRef]
  18. Wang, Y.; Gu, Y.; Dou, M.; Qiao, M. Rethinking the Identification of Urban Centers from the Perspective of Function Distribution: A Framework Based on Point-of-Interest Data. Sustainability 2020, 12, 1543. [CrossRef]
  19. Lou, G.; Chen, Q.; He, K.; Zhou, Y.; Shi, Z. Using Nighttime Light Data and POI Big Data to Detect the Urban Centers of Hangzhou. Remote Sensing 2019, 11, 1821. [CrossRef]
  20. Schneider, A.; Friedl, M.A.; Potere, D. A New Map of Global Urban Extent from MODIS Satellite Data. Environmental Research Letters 2009, 4, 044003. [CrossRef]
  21. Melchiorri, M.; Florczyk, A.J.; Freire, S.; Schiavina, M.; Pesaresi, M.; Kemper, T. The Global Human Settlement Layer Sets a New Standard for Global Urban Data Reporting with the Urban Centre Database. Frontiers in Environmental Science 2022, 10, 1003862. [CrossRef]
  22. Dijkstra, L.; Florczyk, A.J.; Freire, S.; Kemper, T.; Melchiorri, M.; Pesaresi, M.; Schiavina, M. Applying the Degree of Urbanisation to the Globe: A New Harmonised Definition Reveals a Different Picture of Global Urbanisation. Journal of Urban Economics 2021, 125, 103312. [CrossRef]
  23. Cayo, M.R.; Talbot, T.O. Positional Error in Automated Geocoding of Residential Addresses. International Journal of Health Geographics 2003, 2, 10. [CrossRef]
  24. Zandbergen, P.A. Geocoding Quality and Implications for Spatial Analysis. Geography Compass 2009, 3, 647–680. [CrossRef]
  25. Wei, H. DynRFM-Hazard: Repurchase Probability Estimation Using Dynamic RFM and Discrete-Interval Hazards. Preprints 2026. Preprint, version 1. [CrossRef]
  26. Krugman, P. Increasing Returns and Economic Geography. Journal of Political Economy 1991, 99, 483–499. [CrossRef]
  27. Fujita, M.; Krugman, P.; Venables, A.J. The Spatial Economy: Cities, Regions, and International Trade; MIT Press: Cambridge, MA, 1999. [CrossRef]
  28. Rosenthal, S.S.; Strange, W.C. Evidence on the Nature and Sources of Agglomeration Economies. In Handbook of Regional and Urban Economics: Cities and Geography; Elsevier, 2004; Vol. 4, pp. 2119–2171. [CrossRef]
  29. Ellison, G.; Glaeser, E.L. Geographic Concentration in U.S. Manufacturing Industries: A Dartboard Approach. Journal of Political Economy 1997, 105, 889–927. [CrossRef]
  30. Duranton, G.; Overman, H.G. Testing for Localization Using Micro-Geographic Data. The Review of Economic Studies 2005, 72, 1077–1106. [CrossRef]
  31. Duranton, G.; Overman, H.G. Exploring the Detailed Location Patterns of U.K. Manufacturing Industries Using Microgeographic Data. Journal of Regional Science 2008, 48, 213–243. [CrossRef]
  32. Head, K.; Ries, J.; Swenson, D. Agglomeration Benefits and Location Choice: Evidence from Japanese Manufacturing Investments in the United States. Journal of International Economics 1995, 38, 223–247. [CrossRef]
  33. Guimaraes, P.; Figueiredo, O.; Woodward, D. Agglomeration and the Location of Foreign Direct Investment in Portugal. Journal of Urban Economics 2000, 47, 115–135. [CrossRef]
  34. Crozet, M.; Mayer, T.; Mucchielli, J.L. How Do Firms Agglomerate? A Study of FDI in France. Regional Science and Urban Economics 2004, 34, 27–54. [CrossRef]
  35. Basile, R.; Castellani, D.; Zanfei, A. Location Choices of Multinational Firms in Europe: The Role of EU Cohesion Policy. Journal of International Economics 2008, 74, 328–340. [CrossRef]
  36. Cheng, L.K.; Kwan, Y.K. What Are the Determinants of the Location of Foreign Direct Investment? The Chinese Experience. Journal of International Economics 2000, 51, 379–400. [CrossRef]
  37. Huang, Y.; Shirai, S. Intrametropolitan FDI Firm Location in Guangzhou, China. The Annals of Regional Science 2000, 34, 535–555. [CrossRef]
  38. Fan, C.C.; Scott, A.J. Industrial Agglomeration and Development: A Survey of Spatial Economic Issues in East Asia and a Statistical Analysis of Chinese Regions. Economic Geography 2003, 79, 295–319. [CrossRef]
  39. Lu, J.; Tao, Z. Trends and Determinants of China’s Industrial Agglomeration. Journal of Urban Economics 2009, 65, 167–180. [CrossRef]
  40. Long, C.; Zhang, X. Patterns of China’s Industrialization: Concentration, Specialization, and Clustering. China Economic Review 2012, 23, 593–612. [CrossRef]
  41. Jain, A.K. Data Clustering: 50 Years beyond K-Means. Pattern Recognition Letters 2010, 31, 651–666. [CrossRef]
  42. Tibshirani, R.; Walther, G.; Hastie, T. Estimating the Number of Clusters in a Data Set via the Gap Statistic. Journal of the Royal Statistical Society Series B 2001, 63, 411–423. [CrossRef]
  43. Comaniciu, D.; Meer, P. Mean Shift: A Robust Approach toward Feature Space Analysis. IEEE Transactions on Pattern Analysis and Machine Intelligence 2002, 24, 603–619. [CrossRef]
  44. Campello, R.J.G.B.; Moulavi, D.; Sander, J. Density-Based Clustering Based on Hierarchical Density Estimates. In Proceedings of the Advances in Knowledge Discovery and Data Mining. Springer, 2013, pp. 160–172. [CrossRef]
  45. von Luxburg, U. A Tutorial on Spectral Clustering. Statistics and Computing 2007, 17, 395–416. [CrossRef]
  46. Wei, H.; Wang, X. Financial Risk Management Early-Warning Model for Chinese Enterprises. Journal of Risk and Financial Management 2024, 17, 255. [CrossRef]
  47. Wei, H. AI-Driven Employee Turnover Warning for Talent-Retention Decision Support in Finance and Taxation Education: Model Development and Comparison. Preprints 2026. Preprint, version 1. [CrossRef]
  48. Sahr, K. Central Place Indexing: Hierarchical Linear Indexing Systems for Mixed-Aperture Hexagonal Discrete Global Grid Systems. Cartographica: The International Journal for Geographic Information and Geovisualization 2019, 54, 16–29. [CrossRef]
  49. Getis, A.; Ord, J.K. The Analysis of Spatial Association by Use of Distance Statistics. Geographical Analysis 1992, 24, 189–206. [CrossRef]
  50. Ord, J.K.; Getis, A. Local Spatial Autocorrelation Statistics: Distributional Issues and an Application. Geographical Analysis 1995, 27, 286–306. [CrossRef]
  51. Baddeley, A.; Turner, R.; Moller, J.; Hazelton, M. Residual Analysis for Spatial Point Processes (with Discussion). Journal of the Royal Statistical Society Series B 2005, 67, 617–666. [CrossRef]
  52. Kulldorff, M. A Spatial Scan Statistic. Communications in Statistics—Theory and Methods 1997, 26, 1481–1496. [CrossRef]
  53. Efron, B. Bootstrap Methods: Another Look at the Jackknife. The Annals of Statistics 1979, 7, 1–26. [CrossRef]
  54. Politis, D.N.; Romano, J.P. Large Sample Confidence Regions Based on Subsamples under Minimal Assumptions. The Annals of Statistics 1994, 22, 2031–2050. [CrossRef]
  55. Lahiri, S.N. Theoretical Comparisons of Block Bootstrap Methods. The Annals of Statistics 1999, 27, 386–404. [CrossRef]
  56. Wei, H. Negative Indicators and Ordering Stability in Exploratory Factor Analysis: A Sign-Orientation Theory with Reproducible Simulation Evidence. Preprints 2026. Preprint, version 1. [CrossRef]
  57. Wei, H. From Program Design to Continuous Improvement: An AI-Supported Decision Framework for Vocational Training Operations. Preprints 2026. Concept paper; DOI supplied by the author and not yet registered in Crossref on 8 August 2026. [CrossRef]
  58. Fotheringham, A.S.; Wong, D.W.S. The Modifiable Areal Unit Problem in Multivariate Statistical Analysis. Environment and Planning A 1991, 23, 1025–1044. [CrossRef]
Figure 1. Analysis sequence. Entity labels do not participate in center selection.
Figure 1. Analysis sequence. Entity labels do not participate in center selection.
Preprints 227566 g001
Figure 2. Algorithm benchmark, component ablation, observed center counts, and external positional agreement. The GHSL curve covers the 323 units with accepted auxiliary boundaries and is displayed to 25 km so that the central part remains legible; larger deviations are retained in all summaries.
Figure 2. Algorithm benchmark, component ablation, observed center counts, and external positional agreement. The GHSL curve covers the 323 units with accepted auxiliary boundaries and is displayed to 25 km so that the central part remains legible; larger deviations are retained in all summaries.
Preprints 227566 g002
Figure 3. Center reliability and city-equal national primary-center contrasts. Error bars are province-cluster 95% intervals.
Figure 3. Center reliability and city-equal national primary-center contrasts. Error bars are province-cluster 95% intervals.
Preprints 227566 g003
Figure 4. Descriptive primary-center gaps by center structure and registration-count quartile. Points are group medians under the C address specification; they are not causal effects of structure or size.
Figure 4. Descriptive primary-center gaps by center structure and registration-count quartile. Points are group medians under the C address specification; they are not causal effects of structure or size.
Preprints 227566 g004
Figure 5. City distributions of combined registration-agglomeration-zone coverage, area, and raw density. Black segments show interquartile ranges and white circles show medians; area and density use logarithmic scales.
Figure 5. City distributions of combined registration-agglomeration-zone coverage, area, and raw density. Black segments show interquartile ranges and white circles show medians; area and density use logarithmic scales.
Preprints 227566 g005
Figure 6. Registration-center maps for Beijing, Shanghai, Wuhan, and Chengdu. Pale cells show nearest-center catchments, saturated cells show dense registration zones, and numbered circles show MSP-MBKMeans centers.
Figure 6. Registration-center maps for Beijing, Shanghai, Wuhan, and Chengdu. Pale cells show nearest-center catchments, saturated cells show dense registration zones, and numbered circles show MSP-MBKMeans centers.
Preprints 227566 g006
Figure 7. Registration-center maps for Xi’an, Chongqing, Guangzhou, and Shenzhen. Guangzhou is unresolved, so its mapped centers are exploratory.
Figure 7. Registration-center maps for Xi’an, Chongqing, Guangzhou, and Shenzhen. Guangzhou is unresolved, so its mapped centers are exploratory.
Preprints 227566 g007
Table 1. Operational entity classes used in the analysis.
Table 1. Operational entity classes used in the analysis.
Label Formal name Entities Share (%) Historical mappings Interpretation boundary
R Private-dominated other registration entities 183,235,463 99.0283 Residual Dominated by private entities; includes some other registered subjects
S State-identified registration entities 734,375 0.3969 56 Historical string match; not a complete verification of contemporary ownership
F Foreign-identified registration entities 534,945 0.2891 193 Historical foreign-related string match
H Hong Kong/Macao/Taiwan-identified entities 528,609 0.2857 180 Historical Hong Kong, Macao, or Taiwan-related string match
Table 2. Fixed MSP-MBKMeans rule used for the 8% primary definition.
Table 2. Fixed MSP-MBKMeans rule used for the 8% primary definition.
Component Fixed rule
Input and peaks Weighted H3 resolution-9 cells; resolution-8 and -7 peak scores sum occupied cells in a second-order neighborhood; one lowest-ID peak per tied plateau; r8 peak must persist at r7 parent or first-order neighbor.
Candidate screen At most 96 ranked persistent peaks; no relative-height cutoff; separation at least max ( 2 km , 0.08 r 90 ) ; valley ratio at most 0.50; tentative and fitted share at least 8%; at most 8 centers.
MiniBatch fit City-local equal-distance coordinates; central 99.5% weighted radial range; batch 4,096; maximum 300 iterations; one peak-based start; seed 20260808; reassignment ratio 0.
Spatial repetition 20 complete reselections of K after independent Gamma ( 16 , 1 / 16 ) multipliers on H3 resolution-7 blocks. Centers are paired by minimum-total WGS84 geodesic distance.
Consensus and reliability Peaks appearing in at least 50% of repetitions enter the consensus fit, with the highest peak always retained. The displacement limit is max ( 1.5 km , 0.10 r 90 ) . Exact stability requires K agreement 0.80 , every retained peak appearance 0.80 , 90th-percentile paired displacement within the limit, and 90th-percentile share difference 0.10 . Core stability drops only the K-agreement condition; usable-with-uncertainty requires appearance 0.50 , displacement within twice the limit, and share difference 0.20 . Other cases are unresolved.
Table 3. Synthetic validation of MSP-MBKMeans.
Table 3. Synthetic validation of MSP-MBKMeans.
Scenario Repetitions Correct K (%) Median matched error (m)
Compact one center 20 100.0 47.5
Elongated one center 20 100.0 135.4
Balanced two centers 20 100.0 93.6
Imbalanced two centers 20 100.0 175.8
Balanced four centers 20 100.0 87.3
Subthreshold 4% satellite 20 100.0 Not applicable
All MSP-MBKMeans evaluations 120 100.0
Legacy one-standard-error selector 120 16.7
Table 4. Center-count recovery in the expanded synthetic benchmark.
Table 4. Center-count recovery in the expanded synthetic benchmark.
Method Original scenarios (%) Stress scenarios (%) Stress median time (s)
MSP-MBKMeans 100.0 87.8 0.044
HDBSCAN 100.0 88.9 0.095
OPTICS 100.0 88.9 2.124
Gaussian mixture, BIC 83.3 55.6 0.271
Mean shift 70.0 33.9 0.131
Legacy one-SE MiniBatch K-Means 16.7 27.8 0.266
Each column contains 20 repetitions per geometry: 120 evaluations in the six original scenarios and 180 evaluations in the nine stress scenarios. Computation times describe these synthetic samples on the study computer and are not a general hardware benchmark.
Table 5. Prespecified synthetic holdout evaluation.
Table 5. Prespecified synthetic holdout evaluation.
Method Evaluations Correct K (%) Matched error (m)
MSP-MBKMeans 200 99.0 149.5
HDBSCAN 200 100.0 121.0
OPTICS 200 100.0 87.1
Gaussian mixture, BIC 200 60.0 92.4
Mean shift 200 58.5 107.0
Legacy one-SE MiniBatch K-Means 200 20.0
The generating scenarios, expected center counts, random seed, comparison methods, and parameter settings were specified before evaluation. Errors are medians among evaluations with correct K. Parameters remained unchanged throughout the evaluation.
Table 6. Real-city sensitivity to key MSP-MBKMeans thresholds.
Table 6. Real-city sensitivity to key MSP-MBKMeans thresholds.
Setting Median raw K Same K as base (%) S gap (km) F gap (km) H gap (km)
Minimum share 6% 6 38.3 0.00 3.72 4.90
Base: 8%, valley 0.50, distance 0.08 5 100.0 0.19 3.72 4.61
Minimum share 10% 4 44.6 0.27 3.58 4.49
Valley ratio 0.40 5 99.7 0.19 3.72 4.61
Valley ratio 0.60 5 99.7 0.19 3.72 4.61
Distance fraction 0.06 5 99.7 0.16 3.72 4.61
Distance fraction 0.10 5 99.7 0.19 3.72 4.58
Raw K is the deterministic full-sample candidate count; the final spatial-resampling consensus has median four. Gaps use fixed eligible sets of 307/263/220 cities for S/F/H, so changes isolate center-threshold sensitivity. Positive values indicate greater centrality than R.
Table 7. Complete propagation of the 6/8/10% minimum-center-share definition.
Table 7. Complete propagation of the 6/8/10% minimum-center-share definition.
Share Median consensus K Unresolved cities S central/peripheral F central/peripheral H central/peripheral S stable F stable H stable
6% 5 15 24/25 87/8 83/3 Compared across all three rows
8% 4 25 24/25 79/8 82/3 310 313 316
10% 4 24 24/23 80/7 80/2 Compared across all three rows
Central/peripheral counts are class-wise Benjamini–Hochberg-corrected city conclusions. Stable is the number, out of 345, whose complete corrected label is identical at 6%, 8%, and 10%. The 8% row remains the primary specification; alternatives are not searched for favorable results.
Table 8. Observed center structure, reliability, and address sensitivity.
Table 8. Observed center structure, reliability, and address sensitivity.
Diagnostic Cities Interpretation
Exactly stable center count 69 Count and center configuration are repeatable
Stable core; optional peripheral centers 88 Core is repeatable, minor centers may appear or disappear
Usable with center-count uncertainty 163 Main pattern can be described with an uncertainty warning
Unresolved 25 No formal center–periphery conclusion
Identical K under A/B/C 57 Center count is insensitive to address weighting
Full center structure matched under A/B/C 67 Main configuration agrees while allowing at most one optional center
Primary center stable under A/B/C 209 Main reference center survives all address weights
Table 9. Sensitivity to the number of spatial repetitions in 12 structured-sample cities.
Table 9. Sensitivity to the number of spatial repetitions in 12 structured-sample cities.
Comparison with 100 repetitions Same K Same reliability level Median primary shift (km) Maximum primary shift (km)
20 repetitions 10/12 10/12 0.00 3.54
50 repetitions 10/12 11/12 0.00 2.78
The sample takes the minimum, median, and maximum resolution-9 cell count within each primary reliability level. It spans size and reliability for diagnosis and is not a random national sample. The primary analysis retains its prespecified 20 repetitions.
Table 10. Sensitivity of the independent GHSL settlement reference.
Table 10. Sensitivity of the independent GHSL settlement reference.
Setting GHSL nodes Same K (%) Primary median (km) Primary within 10 km (%) All-node within 10 km (%)
Base: tier 10%, component 5%, cap 12 1,337 100.0 3.32 82.7 63.4
Tier population 5% 1,299 97.2 3.32 82.7 64.2
Tier population 15% 1,352 95.4 3.29 82.4 63.4
Component population 3% 1,690 53.9 3.32 82.7 64.5
Component population 8% 918 44.3 3.32 82.7 70.6
Maximum eight nodes 1,326 98.1 3.32 82.7 63.1
Each row rebuilds the GHSL reference without registration data. Same K compares GHSL node count with the base reference, not with MSP-MBKMeans. All-node percentages use minimum-cost matched pairs and must be read with the changing node count.
Table 11. City-level positional conclusions for the four classes.
Table 11. City-level positional conclusions for the four classes.
Class More central More peripheral No clear Sensitive Not estimable
S 24 25 256 2 38
F 79 8 175 1 82
H 82 3 131 4 125
Notes: S, F, and H denote state-, foreign-, and Hong Kong/Macao/Taiwan-identified entities. Each row totals 345 observed cities. Directional and address-sensitive labels pass Benjamini–Hochberg control at q = 0.05 within their class family, use at least eight delete-block estimates, and follow the A/B/C rule. Supplementary Table S1 gives every corrected city conclusion.
Table 12. City-equal national center–periphery contrasts relative to R.
Table 12. City-equal national center–periphery contrasts relative to R.
Class Weight Cities Primary gap, km (95% CI) Probability index Nearest gap, km
S A 307 0.04 [ 0.81 , 0.84] 0.010 0.33
S B 307 1.01 [ 2.43 , 0.25 ] 0.027 0.63
S C 307 0.11 [ 0.99 , 0.71] 0.012 0.36
F A 263 3.25 [2.25, 4.45] 0.079 1.02
F B 263 2.47 [1.70, 3.02] 0.039 0.35
F C 263 3.74 [2.63, 4.79] 0.084 1.06
H A 220 4.39 [2.27, 6.43] 0.115 1.34
H B 220 3.28 [2.00, 5.00] 0.077 0.56
H C 220 4.57 [2.94, 6.96] 0.122 1.37
Notes: Gap = E ( D R ) E ( D m ) ; positive values indicate that class m is closer to a center. Medians are city-equal. Confidence intervals use 5,000 province-cluster bootstrap repetitions. C is the primary specification.
Table 13. Sensitivity to excluding the compared class from center construction.
Table 13. Sensitivity to excluding the compared class from center construction.
Removed Same K Median shift (km) P90 shift (km) Same FDR label
S 314 (91.0%) 0.16 1.52 342 (99.1%)
F 332 (96.2%) 0.10 0.57 343 (99.4%)
H 332 (96.2%) 0.08 0.50 343 (99.4%)
Full and leave-class-out centers use deterministic full-sample fits at the fixed 8% definition so that this diagnostic isolates reference dependence. Primary-model resampling uncertainty is reported separately.
Table 14. Sensitivity to the composition of city-level administrative units under the primary C specification.
Table 14. Sensitivity to the composition of city-level administrative units under the primary C specification.
Inventory Observed units S gap, km (95% CI) F gap, km (95% CI) H gap, km (95% CI)
All source units 345 0.11 [ 0.99 , 0.71] 3.74 [2.63, 4.79] 4.57 [2.94, 6.96]
Prefecture-level units 316 0.24 [ 1.29 , 0.70] 3.74 [2.63, 4.96] 4.66 [2.92, 7.00]
Standard prefecture-level cities 285 0.15 [ 1.06 , 0.80] 3.73 [2.63, 4.90] 4.66 [2.92, 6.96]
The class-specific eligible counts are 307/263/220, 292/253/211, and 263/240/207 for S/F/H, respectively. All estimates are city-equal; intervals use province-cluster resampling.
Table 15. Evidence summary for the eight illustrative cities under the primary C specification.
Table 15. Evidence summary for the eight illustrative cities under the primary C specification.
City Entities K A/B K Primary share Reliability GHSL km S gap F gap H gap Zone share Density
Beijing 4,815,679 3 4/2 0.700 Usable–uncertain 1.39 5.37 9.23 7.53 0.484 2,581
Shanghai 5,210,935 4 3/4 0.513 Usable–uncertain 7.36 6.90 9.42 7.26 0.478 2,402
Wuhan 2,262,224 2 2/3 0.656 Exactly stable 7.13 3.77 3.02 4.62 0.455 6,410
Chengdu 3,806,872 3 2/3 0.764 Usable–uncertain 5.81 1.06 7.70 6.61 0.489 2,707
Xi’an 2,614,817 2 2/2 0.899 Exactly stable 2.03 2.19 4.21 3.75 0.496 4,152
Chongqing 3,209,412 2 3/4 0.718 Usable–uncertain 6.83 15.87 63.41 60.67 0.508 265
Guangzhou 3,842,124 3 2/3 0.706 Unresolved 4.03 5.35 5.62 1.93 0.487 3,106
Shenzhen 5,190,364 4 5/5 0.474 Usable–uncertain 7.76 0.95 2.11 1.13 0.501 10,576
A/B K gives center counts under raw-entity and address-equal weights; the main K uses C. Gaps are E(DR) − E(Dm) in km. Density is entities/km2. Guangzhou’s gaps are exploratory because its center system is unresolved.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.