Submitted:
09 August 2026
Posted:
14 August 2026
You are already at the latest version
Abstract
Business-registration coordinates describe legal addresses, not verified workplaces, and shared addresses can create artificial density peaks. This study identifies registration-center structure and compares the center--periphery positions of four historical entity classes in China. The January 2021 archive contains 185,033,392 deduplicated entities in 364 city-level source units; 345 have observations. The proposed Multi-Scale Persistent-Peak-Constrained MiniBatch K-Means algorithm (MSP-MBKMeans) combines peaks repeated across two H3 scales with density-valley, minimum-share, adaptive-distance, address-weight, and spatial-resampling checks. In a prespecified holdout of 200 synthetic cities, it recovered the substantive center count in 99.0%; HDBSCAN and OPTICS reached 100.0%, so no universal superiority is claimed. The fixed 8% primary definition identified 1,410 centers, with a city median of four; 25 cities remained unresolved. At 6/8/10%, primary-center position was stable in 71.3% of cities, whereas center count was often definition-dependent. An independent comparison in 323 units placed the primary registration center a median 3.32 km from the GHSL 2020 population node, with 82.7% within 10 km. After class-wise false-discovery-rate control, foreign-identified and Hong Kong/Macao/Taiwan-identified entities were more central than the private-dominated residual class in 79 and 82 cities; their city-equal median advantages remained 3.74 and 4.57 km. State-identified entities had no common national direction. Complete center-share sensitivity analyses at 6/8/10% preserved 90.7% of the 1,035 city-class labels, and leave-class-out centers preserved 99.3%. A supplementary catalogue reports corrected conclusions and separate structural and class-location sensitivity warnings for every unit. Conclusions concern registration locations, not production, employment, or causal ownership effects.
Keywords:
business-registration location
; urban centers
; MiniBatch K-Means
; density peaks
; spatial resampling
; H3
; China
; foreign-identified entities
; registration agglomeration zones
1. Introduction
Cities concentrate organizations, infrastructure, information, and administrative services, yet the word “center” can refer to very different objects. An employment center is identified from jobs, a commercial center from customer-facing activity, a built-up center from settlement morphology, and a business-registration center from legal addresses. These objects can overlap without being identical. A national archive of geocoded registration records offers a rare opportunity to examine where organizations are legally anchored, but it also creates a strong temptation to call those coordinates workplaces or economic activity points. This paper starts by refusing that shortcut. Its subject is registration geography: the spatial structure of recorded legal locations and the way historically identified entity classes are positioned within that structure.
The distinction is not merely terminological. A registration address may be a genuine office, a headquarters, a shared service address, an incubator, or an address maintained for regulatory purposes. Chinese company practice explicitly recognizes possible inconsistency between a registered and an actual business address [1]; research in other settings likewise shows that an address can act as an organizational signal rather than a direct measure of activity [2]. If a building hosts thousands of legal registrations, a conventional density surface can place a prominent peak there even when the corresponding entities operate elsewhere. Conversely, erasing multiplicity and counting every address once may discard real organizational concentration. A scientifically useful analysis must expose this uncertainty instead of burying it in preprocessing.
Registration data nonetheless have distinctive value. Administrative business records can cover organizations that are absent from surveys, consumer platforms, remote-sensing products, and employment datasets. Large Chinese enterprise-registration data have already been used to support spatiotemporal industrial analysis and to address missing registration attributes [3]. More generally, emerging urban data sources allow city patterns to be studied at scales that would be impractical with field enumeration alone [4]. The present snapshot contains more than 185 million deduplicated entities and reaches hundreds of city-level units. Its scale makes national comparison possible, but scale alone is not a contribution. The underlying business-data project raised a practical question: how can an enormous coordinate archive be turned into understandable city findings when shared addresses, uncertain center counts, and missing cities cannot be ignored? Without clear definitions and complete city-by-city evidence, more records can simply make a biased result look more precise.
Three linked questions motivate the study. First, how many repeatable business-registration centers can be identified in each observed city without fixing the number in advance? Second, once those centers are defined independently of entity class, do state-identified, foreign-identified, and Hong Kong/Macao/Taiwan-identified entities occupy systematically different center–periphery positions from the private-dominated residual class? Third, can the densest part of each center catchment be delineated as a compact registration zone without joining separate centers across empty space? These questions form one research line. Center discovery supplies the reference geometry; class comparison gives that geometry a social-economic interpretation; and zone delineation translates point concentrations into reviewable spatial outputs.
Existing tools solve parts of the problem but not the whole research problem. K-means minimizes within-cluster squared distance [5], while MiniBatch K-Means makes related optimization practical for very large datasets [6]. Neither algorithm decides, by itself, how many substantively meaningful urban centers exist. Density-peak methods identify locally prominent modes [7], but a peak seen at one spatial resolution may be a grid artifact, an address spike, or a minor satellite. Standard cluster validation, including silhouette reasoning [8], can favor geometrically useful partitions that do not correspond to interpretable urban centers. Stability assessment can also be misleading when analysts fix the selected K and repeatedly refit centers under that fixed K: conditional stability is then mistaken for selection stability.
We address these problems with Multi-Scale Persistent-Peak-Constrained MiniBatch K-Means, abbreviated MSP-MBKMeans. The procedure discovers peaks at two nested H3 scales, retains only cross-scale correspondences, applies minimum-share, adaptive-separation, and intervening-density-valley constraints, and uses the surviving peaks as fixed starting centers for MiniBatch K-Means. Most importantly, each H3-block resample repeats peak discovery and selection of K. A city can therefore be reported as exactly stable, core-stable with optional centers, usable with uncertainty, or unresolved. This graduated outcome is more informative than either forcing one solution for every city or discarding all cities that are not perfectly stable.
Concentrated addresses are handled through three prespecified weighting schemes. Raw entity multiplicity (A) preserves the archive literally. Exact-address equality (B) gives the strongest correction. City-specific 99th-percentile truncation (C) retains most count information while bounding the largest shared addresses and serves as the primary estimate. The same A/B/C logic is used both to test center robustness and to evaluate R-relative class differences. A directional national claim is allowed only when all three city-equal medians agree and their province-cluster intervals exclude zero. This rule makes robustness part of the estimand rather than an appendix added after seeing the preferred result.
The empirical findings are concise even though the calculations are extensive. Across 300 synthetic cities, MSP-MBKMeans recovered the correct center count in 100.0% of six calibration geometries and 87.8% of nine stress geometries; HDBSCAN and OPTICS reached 88.9% under stress. In 345 units with source samples, the median selected count was four. Center uncertainty was not negligible: 69 cities had exactly stable counts, 88 stable cores, 163 usable structures with uncertainty, and 25 unresolved structures. An independent GHSL comparison was possible for 323 units with accepted auxiliary boundaries; primary registration centers had a 3.32 km median distance from the population-weighted settlement reference, and 82.7% were within 10 km.
The central social-economic result is not an abstract statement that cities are “heterogeneous.” Foreign-identified and Hong Kong/Macao/Taiwan-identified entities are closer to the primary business-registration center than the private-dominated R class in the city-equal national distribution. Under the primary C specification, the median advantages are 3.74 and 4.57 km; both remain positive in the strict subset that requires an exactly stable center count and an address-robust primary center. State-identified entities have no nationally common center–periphery direction. The latter result is substantive rather than a failed test: state-related registration locations can reflect central administrative institutions, peripheral land-intensive units, local state-enterprise histories, and shared addresses in different combinations across cities.
The paper makes four contributions. First, it combines multi-scale peak persistence with constrained MiniBatch K-Means to identify interpretable registration centers and tests the design against five competing algorithms and five substantive ablations. Second, it reselects K under spatial perturbation instead of assuming that the full-sample center count is certain. Third, it tests whether concentrated shared addresses or the mixture of city-level administrative units changes the conclusions. Fourth, it provides a complete national application: 1,410 center catchments, 1,359 zones reaching their density-coverage target, findings for 345 observed units, and explicit no-sample statements for 19 units. All 364 units appear in the city conclusion catalogue.
The contribution is deliberately bounded. MSP-MBKMeans is a domain-specific combination of established methods and new rules, not a new theory of clustering. The results do not identify a true central business district, production location, workplace, firm migration, or an ownership effect. The four labels do not restore the lost 429 source strings. K is not a development score. The analysis is descriptive, cross-sectional with respect to the archive snapshot, and tied to administrative city units. These limits narrow the claims, but they also make the results usable: readers know exactly what was measured, which cities support formal inference, and where the method declined to decide.
The remainder of the paper reviews the relevant urban-center and spatial-clustering research, explains the data and method, presents algorithmic and substantive results, interprets eight cities, reports all 364 units, and discusses what registration centrality can and cannot tell us about urban organization.
2. Related Work and Methodological Positioning
2.1. What Kind of Urban Center Is Being Measured?
Urban-center research has never had a single universally correct dependent variable. Classical work on metropolitan subcenters often begins with employment density and thresholds that distinguish large local concentrations from surrounding areas. McDonald [9] formalized early employment-subcenter identification; the Los Angeles study by Giuliano and Small [10] converted the intuitive idea into replicable density and employment criteria; and McMillen [11] later developed nonparametric identification that reduced dependence on arbitrary candidate locations. Redfearn’s long-run analysis further showed that employment centers can persist even as metropolitan form changes [12]. These approaches concern employment landscapes. Their importance here lies in the principles that centers should be locally prominent, spatially separated, large enough to matter, and evaluated for persistence—not in an assumption that registration counts are jobs.
The broader measurement problem is how to turn a qualitative construct into numerical rules without silently changing its meaning. Applied quantitative-design research provides one example of defining observable components before comparison [13]. Here, the same measurement principle represents a registration center through local prominence, separation, minimum share, and persistence rather than through informal visual judgment.
Polycentricity also has morphological and functional meanings. The polycentric urban-region literature treats multiple nodes as a research agenda rather than a single self-evident statistic [14]. Meijers [15] shows that alternative measures can yield different polycentricity assessments, while Burger and Meijers [16] distinguish the size and spatial arrangement of nodes from flows and relations among them. MSP-MBKMeans estimates a morphological pattern in one particular administrative point field. It does not observe commuting, supply chains, inter-office communication, or customer flows. Accordingly, the paper uses “monocentric,” “dominant-core polycentric,” and “balanced polycentric” only as explicit summaries of fitted registration-center shares. These labels do not establish a functionally integrated polycentric urban region.
Spatial-layout research likewise distinguishes geometric partitioning from the functional meaning assigned to a partition [17]. The distinction is directly relevant here: nearest-center catchments are geometric allocations, whereas a functional business district would require activity and interaction data that are unavailable.
New urban data sources have expanded center identification beyond employment. Point-of-interest distributions can characterize functional mix and activity concentrations; Wang et al. [18], for example, identify centers through distributions of urban functions, while Lou et al. combine nighttime lights and POIs to detect Hangzhou’s urban centers [19]. Remote sensing provides globally consistent but definition-dependent urban extents [20]. The GHSL Urban Centre Database and the harmonized Degree of Urbanisation provide a population-and-built-up-area reference designed for cross-country comparison [21,22]. Registration records add a different layer: they locate legal organizational presence, including many entities invisible in imagery or consumer POI platforms. The correct validation question is therefore whether registration centers are plausibly aligned with independent settlement centers, not whether the two coordinate sets must coincide.
This distinction explains the study’s validation design. A small GHSL distance supports positional plausibility. A large distance may reveal an MSP-MBKMeans error, a registration-specific node, a difference between administrative and morphological city definitions, or a prefecture containing several distant settlements. Treating GHSL as ground truth would erase these possibilities. The reliability hierarchy, A/B/C address comparison, and city map are therefore read alongside external distance. No single number certifies a city center.
2.2. Administrative Registration Records as Spatial Evidence
Administrative business data combine broad coverage with measurement ambiguity. Zhao et al. [3] demonstrate the analytical potential of large Chinese enterprise-registration records and the practical importance of imputing incomplete attributes for spatiotemporal industry research. Their work supports the proposition that registration archives can illuminate economic geography, but it also highlights the need to document field completeness and transformation. The present study takes a complementary path: rather than reconstructing the lost organization-type detail or building a time series, it freezes the four surviving historical labels and concentrates on spatial center structure.
The legal and operational literature provides a direct warning about address meaning. A company’s registered address and actual business address can be inconsistent in Chinese practice [1]. Address itself may communicate status or legitimacy to stakeholders [2], creating incentives that are not reducible to physical production requirements. More generally, geocoding studies show that address matching and coordinate assignment introduce positional uncertainty that can propagate into spatial statistics [23,24]. Shared registration services and development zones can add a different problem: many legally distinct entities can concentrate at one valid coordinate. Consequently, a registration point may represent a real facility, an administrative anchor, or both. The archive cannot discriminate among these cases at national scale, so coordinate error and address multiplicity require separate diagnostics.
Three implications follow. First, the analysis should not use remote-sensing land cover to “correct” registration points: a legal address in a mixed or residential building can still be valid. Second, exact-address multiplicity must remain visible because it is both a possible bias and a feature of registration organization. Third, the language of conclusions must stay at the registration level. These principles motivated the A/B/C weights and the decision to omit WorldCover and DEM from the core algorithm. Environmental covariates could explain where physical facilities are feasible, but they cannot validate the meaning of a legal address.
The archive is also a historical cumulative collection rather than a clean stock of active firms. This feature changes the interpretation of density. A center can represent persistent administrative attractiveness accumulated across multiple registration states. It cannot, without event dates and status histories, be interpreted as the current mass of operating establishments. The snapshot is still meaningful: legal registrations are institutional traces of organizational formation. The resulting map is an accumulated registration landscape.
Time-indexed event models require explicit event intervals rather than an undated cumulative history [25]. The same data requirement applies here: without complete establishment and status intervals, this study can estimate accumulated spatial form but cannot estimate center emergence, movement, or current operating risk.
2.3. Firm Location, Agglomeration, and Ownership-Related Classes
Economic geography explains why organizations may concentrate. Increasing returns, market access, input sharing, labor pooling, and knowledge spillovers can support spatial agglomeration [26,27,28]. Empirical concentration measures also warn that visible clustering must be evaluated relative to industrial composition and an appropriate counterfactual [29]. Micro-geographic tests demonstrate that localization can occur at scales hidden by broad administrative summaries [30,31]. These foundations justify studying fine-scale registration concentration, but they do not turn a concentration of legal addresses into evidence of production linkages.
Studies of mobile international investment repeatedly find that existing concentrations and policy environments shape location choice, including evidence from Japanese investment in the United States and from Portugal, France, and the European Union [32,33,34,35]. Research on China is more directly relevant. Cheng and Kwan [36] examined determinants of provincial FDI location and emphasized market size, infrastructure, policy, and labor-cost conditions; intrametropolitan work in Guangzhou examined foreign-invested firms within a Chinese city [37]. Chinese industrial concentration also changes across periods, regions, and territorial scales [38,39,40]. A central registration address may offer proximity to government services, professional intermediaries, finance, transport, information, or prestigious business districts, but it may also be the address of an agent. The national F/H gaps are therefore consistent with center-oriented registration, not proof of agglomeration economies or any single channel.
The state-identified class has reasons to be less uniform. Central administrative and financial organizations may register in core areas, whereas utilities, resource firms, industrial units, transport operators, and land-intensive state entities may be peripheral. Local histories of state-sector restructuring and development-zone registration can shift the balance. Because the data contain neither industry detail nor a recovered organization-type string, these channels cannot be separated. The absence of a national S direction is therefore best treated as a substantive boundary on generalization. It warns against turning one-city observations into a nationwide ownership narrative.
The four surviving labels are historical operational categories, not clean ownership treatments. R includes a small number of non-enterprise entities; S/F/H are string-match outputs; and beneficial ownership may have changed after registration. The study consequently avoids propensity-score language, treatment effects, or counterfactual claims. Its estimands are distributional differences in registration distance, conditional only on the city-specific center geometry and the eligibility rules.
2.4. K-Means Scalability and the Unresolved Problem of K
Lloyd’s algorithm [5] remains a foundational solution for partitioning observations by squared Euclidean distance, but the broader clustering literature stresses that no single algorithm or validity criterion is uniformly appropriate [41]. For massive point sets, repeatedly assigning every point to every candidate center can be prohibitively expensive. MiniBatch K-Means reduces computation by updating centers with small batches and has demonstrated web-scale utility [6]. This computational property is essential for a national archive, even after exact sites have been aggregated to H3 cells.
Scalability does not solve model selection. K-means will return K clusters for any requested K, and a reduction in quantization error is guaranteed as K increases. Generic validity indices such as the silhouette [8] and the gap statistic [42] assess different aspects of partition quality, but neither defines when a cluster is a substantively meaningful registration center. An elongated city, a dispersed prefecture, or a shared-address spike can make a geometrically useful partition look like a center system. A one-standard-error rule applied to held-out quantization error can still choose an unnecessarily large K when the error curve is shallow and correlated folds are treated optimistically. This is precisely what the synthetic comparison uncovered.
Density-peak clustering begins from a different intuition: cluster centers have high local density and are relatively distant from points of higher density [7]. This is closer to the urban idea of a prominent node, yet direct application remains insufficient. Grid resolution changes neighborhood relationships; a narrow address spike may outrank a broad center; several peaks may sit on the same density plateau; and a small but isolated satellite can appear prominent. MSP-MBKMeans therefore uses density peaks as constrained candidates and initializers, not as final centers. The MiniBatch optimization re-estimates their weighted positions, and post-fit share pruning enforces substantive scale.
Alternative families clarify the design choice. Mean shift searches for density modes but remains bandwidth dependent [43]; hierarchical density clustering represents variable-density structure and noise without a fixed K [44]; and spectral clustering can recover non-convex graph structure at higher graph-construction cost [45]. The evaluation therefore compares MSP-MBKMeans with mean shift, HDBSCAN, OPTICS, and a BIC-selected Gaussian mixture, as well as the legacy MiniBatch K-Means selector. MSP-MBKMeans does not assume that these methods are inferior. Its combination is intentionally modular: multi-scale persistence addresses resolution-specific noise, the valley rule checks separation, adaptive distance reflects city extent, minimum share rejects negligible satellites, and MiniBatch fitting preserves scalability. The methodological claim is that this explicit combination addresses the operational registration-center definition and documented legacy failure, not that it dominates every clustering algorithm on every data type.
Applied model comparisons in enterprise risk and employee-turnover studies illustrate the narrower methodological principle that alternatives should be judged on a defined decision task rather than one favorable score [46,47]. Here, that task is recovery of substantively sized spatial centers, so correct K, center displacement, computation time, and failure cases are reported together.
2.5. Discrete Grids, Spatial Dependence, and Honest Uncertainty
Hierarchical hexagonal grids support reproducible aggregation across scales and compact adjacency operations. Work on discrete global grid indexing formalizes how hierarchical cells can be addressed and related [48]. In this study, H3 resolutions do four jobs: reduce an enormous exact-site problem to weighted cells, provide nested scales for peak persistence, define contiguous neighborhoods for zones, and supply spatial blocks for uncertainty. The grid hierarchy is therefore necessary for applying the same rules across all cities.
Spatial observations are not independent. Nearby cells share urban context and may be affected by the same concentrated address or local node. Classical local-association statistics, point-process residual diagnostics, and spatial scan methods formalize different manifestations of this dependence [49,50,51,52]. The ordinary bootstrap provides the general resampling foundation [53], but record-level iid resampling would fragment spatial features and produce implausibly narrow uncertainty here. Subsampling theory provides a route to inference under weak assumptions [54], while block-bootstrap theory explains why block construction matters for spatially dependent data [55]. The Gamma-multiplier procedure therefore perturbs whole H3 resolution-7 blocks rather than copying individual entities.
The study extends this logic from parameter uncertainty to model-selection uncertainty. Every bootstrap replicate repeats peak detection and K selection. This distinction is central. If the full-sample K were held fixed, the resulting interval would answer only, “Where are the centers, assuming this center count is unquestionably correct?” The actual question is, “Would a spatially perturbed version of the city support the same centers?” The four reliability grades communicate the answer without reducing a complex stability distribution to a binary pass/fail label.
Sign orientation is part of that communication. Work on negative indicators and ordering stability shows that reversing an indicator without fixing its interpretation can reverse rankings while leaving the underlying information unchanged [56]. The present study therefore defines once, states that positive means more central, and applies the same orientation to every city, address treatment, table, and figure.
National uncertainty introduces another dependence level. Cities within a province can share policies, registry practices, historical development, and preprocessing lineage. Province-cluster bootstrap intervals therefore complement city-level block uncertainty. The two procedures answer different questions: block deletion and multiplier resampling assess a city’s spatial structure; province-cluster resampling assesses how sensitive the national city-equal summary is to regional composition.
2.6. Registration Zones, Location Quotients, and Interpretation Discipline
After centers are fitted, a nearest-center allocation provides an exhaustive, non-overlapping partition of occupied cells. A simple global density threshold would favor large-city cores and make small-city zones disappear. A catchment-relative 50% target instead asks where half of each center’s capped registration weight is most densely packed. Density-priority growth preserves connection to the center, while the one-cell bridge and 5% secondary-patch rule balance fragmentation against artificial contiguity.
Location quotients compare a zone’s class composition with its own city baseline. They are descriptive ratios, not causal statistics. In an archive of this size, a tiny deviation can exceed a random-label confidence bound. Requiring both and an upper-bound exceedance separates statistical detectability from practical prominence. The conditional hypergeometric calculation also avoids comparing a rare class with an unrestricted national null when its opportunity set is city specific.
The term “agglomeration zone” is retained because the algorithm identifies spatial concentration, but it is always modified by “registration.” An industrial cluster normally implies related firms, specialization, production linkages, knowledge spillovers, or labor-market interaction. None of these is observed. A registration zone can be a valuable planning or data-review object without being an industrial cluster.
Decision-support research separates an evidence-based starting point from later review and revision [57]. Registration zones have this limited role: they identify where address verification or linkage can begin, while later field evidence may revise the thresholds. They do not, by themselves, authorize an urban-development decision.
3. Materials and Methods
3.1. Study Object, Snapshot, and Inferential Unit
The object of this study is the business-registration location: the coordinate attached to a legal registration record. It is not assumed to be a shop front, factory, office, headquarters, workplace, or place at which employees and customers are physically present. This distinction governs the title, variables, maps, and interpretation. The expression “registration center” therefore denotes a concentration center estimated from registration coordinates. Similarly, a “registration agglomeration zone” denotes a high-density portion of those coordinates; it is not called an industrial cluster.
The source snapshot was created on 11–12 January 2021. It is a historical cumulative collection containing active, cancelled, revoked, and other registration states. Besides companies, it contains a small number of public institutions, foundations, schools, law firms, and other registered social entities. Because complete establishment, closure, and status histories are unavailable, the snapshot cannot support a cohort analysis, a 2021 entry analysis, or a claim about all firms operating on one date. Our estimand is the spatial organization visible in this cumulative registration archive.
The source data were lawfully purchased by the institution represented by the author for research use. Entity-level records remain subject to third-party copyright and institutional agreement and therefore cannot be redistributed.
The source inventory began with 185,107,617 valid physical records in four source partitions. Cross-file identity checks, duplicate resolution, damaged-record handling, coordinate validity checks, and city reconciliation produced 185,033,392 deduplicated entities in the formal analysis layer. All 364 source city or regional units remain in the final city inventory. Of these, 345 contain usable source observations and 19 contain none. A no-sample unit is not silently dropped or assigned a zero-density estimate: it receives a separate “no source sample; not estimable” output.
The city is the inferential unit for national comparisons. This choice prevents a few very large registration archives from determining a national median. Province is the resampling cluster because cities in the same province can share administrative practice, data lineage, and regional context. Entity counts remain important within a city for fitting density and distance distributions, but national summaries assign equal weight to each eligible city.
3.2. Historical Entity Labels and Their Limits
The retained field contains four historical output labels, denoted R, S, F, and H (Table 1). S, F, and H resulted from matching a now-lost organization-type string to 56 state-related, 193 foreign-related, and 180 Hong Kong/Macao/Taiwan-related source mappings. The 429 mapping strings are not 429 observed categories in the retained data. Because the original organization-type field is unavailable, no entity can be reassigned to one of those strings and no classifier is trained to reconstruct it.
R is the residual class not matched by S, F, or H. It is overwhelmingly private-enterprise dominated but also includes the small non-enterprise portion of the archive. We therefore use the operational name “private-dominated other registration entities,” not the stronger claim “all private firms.” The other classes are likewise called “state-identified,” “foreign-identified,” and “Hong Kong/Macao/Taiwan-identified.” These names describe historical identification rules, not a verification of beneficial ownership in 2021.
The four counts sum to the complete 185,033,392-entity analysis layer. S, F, and H together account for 0.9717% of the formal archive. A class-city comparison is reportable only when the class contains at least 100 raw entities and spans at least eight H3 resolution-6 spatial blocks. This rule gives rare classes a meaningful spatial footprint and avoids treating a large R denominator as evidence that a geographically concentrated rare class is precisely estimated.
3.3. Coordinate Reference System, City Identity, and Grid Representation
The source longitude–latitude pairs were confirmed to be BD-09. They were converted to WGS84 (EPSG:4326) using the same specified transformation throughout the study. Across province samples, the median displacement induced by conversion ranged from 854.91 to 1,382.87 m and the 99th percentile ranged from 941.46 to 1,477.76 m. Retaining the unconverted coordinates would therefore introduce a systematic displacement of the same order as some local center-location errors. Raw coordinates were preserved in the coordinate checks, while all modeling uses WGS84.
City identity follows the source-file inventory and a 2021 administrative-code crosswalk. An older local boundary layer is used for cartographic context and anomaly review only; it neither deletes nor reallocates an entity. This protects the analysis from silently changing the source population when a coordinate lies outside an imperfect polygon. It also means that city units are not geometrically comparable in a strict morphological sense: municipalities, prefecture-level cities, autonomous prefectures, and directly administered county-level units can cover very different extents.
To test whether that mixture determines the national result, the complete analysis is summarized for three nested inventories fixed before comparison: all 364 source units; 330 prefecture-level units after excluding province-level municipalities and special county-level units; and 291 standard prefecture-level cities whose official names end in “city.” Center reliability and the C-specification S/F/H contrasts are recomputed within each inventory. This sensitivity analysis tests the composition of the inventory, not hypothetical changes to boundaries within any one unit.
Exact WGS84 sites are aggregated to the H3 discrete global grid. Resolution 9 is the fitting input; resolutions 8 and 7 support multi-scale peak detection, and resolution 7 also defines stability blocks. Hexagonal aggregation provides compact adjacency and a reproducible hierarchy, while true cell area is used whenever counts become densities. The method does not require a planar national projection. Great-circle distances are computed in kilometers, and local weighted fitting coordinates are handled consistently within each city. H3 is a spatial representation used in the analysis, not evidence that the underlying process is naturally hexagonal.
3.4. Three Address-Weighting Specifications
Shared registration agents, incubators, office buildings, and administrative addresses can hold unusually many entities. Treating every record literally may turn one legal address into an apparent urban center; treating every address equally may remove genuine differences between large and small sites. We therefore analyze three prespecified weights for exact site i, where is its raw entity count and is the city-specific 99th percentile of site counts:
Specification A retains all entity multiplicity. B assigns equal weight to each exact address and is the strongest correction for mass registration. C winsorizes only the extreme upper tail and is the primary specification. A conclusion is called address-robust only when its direction survives all three specifications under the stated inferential rule. The three specifications are interpreted as sensitivity bounds, not as competing estimates from which the most favorable result may be selected.
3.5. MSP-MBKMeans Center Discovery
Figure 1 summarizes the method. Center discovery uses total registration weight only. R/S/F/H labels are withheld until centers have been estimated, preventing the outcome comparison from influencing the spatial reference system.
For each city, resolution-9 weights are aggregated to resolutions 8 and 7. At a cell g, the peak score is the total registration weight in its occupied second-order hexagonal neighborhood, . A cell is a local peak if no neighboring score is larger; tied plateaus retain only the lowest H3 identifier. This deterministic tie rule prevents one flat high area from becoming several peaks. Resolution-8 and resolution-7 peaks are found separately. A resolution-8 peak is persistent only when its resolution-7 parent or one of that parent’s first-order neighbors is a resolution-7 peak. Persistent candidates are ranked by and limited to 96. The final rule has no relative peak-height cutoff.
Candidates are considered in that order. A candidate a must be at least from every retained peak, where is the weighted 90th-percentile city radius. For candidate peaks a and b, their valley ratio is
where is the shortest H3 path between them. A separate center is permitted only when ; adjacent peaks have ratio one, and an unavailable path is assigned ratio zero. The tentative nearest-seed allocation must give every seed at least 8% of fitting weight, and at most eight seeds are retained. Thus a small satellite is not promoted merely because it is a local maximum.
The retained peaks determine both starting positions and candidate K for weighted MiniBatch K-Means. Fitting uses a city-local equal-distance projection and the radial range containing 99.5% of cumulative weight, preventing isolated remote records from pulling the fitted centers. MiniBatch K-Means uses batch size 4,096, at most 300 iterations, one supplied initialization, fixed seed 20260808, and no random reassignment. If a fitted center carries less than 8% of fitting weight, it is removed and the fit is repeated. Centers are ordered by fitted share, with rank 1 defined as the primary registration center. Table 2 records the complete fixed rule. This design keeps the large-sample advantage of MiniBatch K-Means [6] while making center number and starting positions depend on visible spatial structure.
3.6. Spatial-Block Stability with Repeated Selection of K
A fixed K can produce apparently stable centers even when the data do not clearly support that center count. We therefore multiply each H3 resolution-7 block by an independent mean-one Gamma weight and repeat the complete peak selection, candidate filtering, K selection, and MiniBatch fit 20 times. Repeated centers are paired with full-sample centers by the minimum-total WGS84 geodesic distance assignment. The consensus and four reliability statuses are defined in Table 2; unresolved cities do not enter formal center–periphery comparison, although exploratory maps and numerical diagnostics remain available.
Twenty repetitions make appearance rates observable in five-percentage-point increments. They provide a reproducible grading diagnostic rather than a precise probability estimate. A separate repetition-count sensitivity analysis compares 20, 50, and 100 repetitions in a structured sample of 12 cities selected to span all four reliability levels and the observed range of H3 input sizes; its results are reported in Section 4. The fixed random seed makes the stated analysis reproducible. The spatial perturbations assess whether the center structure persists; they do not enumerate every possible random ordering of MiniBatch updates.
3.7. Computational Scale
The 185,033,392 entities occupy 19,175,721 exact coordinate sites and 2,905,086 H3 resolution-9 cells in the 345 observed units. Replacing repeated coordinates by one coordinate and its multiplicity preserves their exact K-means contribution, while H3 resolution 9 supplies the finer spatial representation used for center fitting. A city has a median 6,494 occupied fitting cells and a maximum of 51,414, so the national archive is not held in memory as one point matrix.
Let denote occupied fitting cells, retained candidates, centers, batch size, T MiniBatch updates, and B spatial repetitions in city c. Nested H3 aggregation and bounded-neighborhood peak scoring grow approximately with ; the capped candidate comparisons do not require all cell pairs. MiniBatch fitting requires approximately calculations, and spatial repetition scales this term approximately by B. Exact-site class summaries grow with the city’s site count, weighted quantiles add a sorting step, and density-priority zone construction grows approximately as . Cities are independent after their source assignment is fixed and can therefore be evaluated concurrently. The reported analysis was feasible on an Apple M4 workstation with 10 processor cores and 16 GB memory, with at most four cities evaluated concurrently and numerical calculation for each city restricted to one processor core; method-comparison times for identical synthetic samples are reported in Table 4.
3.8. Center–Periphery Outcomes and Inference
For every eligible city, class, and address specification, we calculate weighted mean and quartile distances to the primary and nearest centers; shares within 5, 10, and 20 km of the primary center; shares within 5 and 10 km of any center; core, intermediate, and peripheral shares defined by city-wide distance quartiles; and allocation across center catchments. The principal distance contrast for class is
A positive value means that class m is closer to the registration center than R in city c. We also report the probability-scale index
which ranges from to 1 and equals zero when neither distribution has a centrality advantage.
City-level uncertainty is estimated by deleting H3 resolution-6 blocks in turn. If is the gap after deleting block b, the jackknife standard error is and the interval is . Formal assessment requires at least 100 raw class entities, eight occupied class blocks, and eight usable delete-block estimates. A two-sided p value is reconstructed from the symmetric normal interval. Benjamini–Hochberg control at is then applied separately to the 345 S, F, and H city families. A city is formally called more central or more peripheral only when its A contrast survives this correction and the A/B/C point estimates all have that direction. A corrected A contrast followed by an A/B/C reversal is called address-sensitive; all other supported cases are no clear difference. Unresolved center structure overrides the distance result. The catalogue therefore reports corrected city statements, not unadjusted screening labels.
The interval procedure was checked before use with fixed-total random labels and a known spatial shift. Under the null, 1,000 simulations produced a 1.0% rejection rate and 99.0% coverage; under the alternative, 500 simulations produced 99.8% detection and 100.0% correct direction. A separate null experiment with 2,000 simulations for the model-gain block procedure gave 4.7% rejection and 95.3% coverage. These simulations check calibration under their stated designs; they do not make the observed archive a random population sample.
National estimates are medians of city effects and shares of cities with positive effects. Their 95% intervals use 5,000 province-cluster bootstrap repetitions. A national directional conclusion requires all A/B/C median effects to share a sign and all corresponding cluster intervals to exclude zero. Structure, size, and province summaries are descriptive stratifications; they are not causal moderators.
3.9. Registration Agglomeration Zones
Occupied resolution-8 cells are allocated to the nearest fitted center, creating non-overlapping Voronoi-like catchments. Within each catchment, C-weight density is computed per square kilometer. A density-priority frontier grows outward from the highest-density occupied cell near the center until it carries 50% of catchment C weight. One empty cell may be bridged for morphological closure only if it belongs to the same nearest-center catchment. When the anchored patch is insufficient, a secondary patch is admitted only if it carries at least 5% of catchment weight. If qualifying patches cannot reach 50%, the zone is reported as incomplete rather than expanded across large empty areas.
For each zone we report area, raw entities, C weight, density, 90th-percentile radius, city share, and R/S/F/H composition. Class enrichment uses the location quotient . A class is called prominently enriched only when and exceeds the upper 95% bound from the hypergeometric random-label distribution conditional on city class totals and zone capacity. This dual threshold prevents enormous sample size from turning negligible compositional differences into substantive claims.
3.10. Validation Design
The first validation set contains six synthetic geometries repeated 20 times each: compact one-center, elongated one-center, balanced two-center, imbalanced two-center, balanced four-center, and a main center with a 4% satellite. A second stress set contains nine additional geometries covering overlapping centers, unequal scales, three and six centers, a curved non-convex field, uniform noise, a concentrated address spike, and satellites immediately below and above the substantive share threshold. Each stress geometry is also repeated 20 times, giving 300 synthetic cities. Expected K and generating centers are known. MSP-MBKMeans is compared on identical coordinates with the legacy one-standard-error MiniBatch K-Means selector, full-covariance Gaussian mixtures selected by BIC, HDBSCAN, mean shift, and OPTICS. Baseline settings are fixed by method rather than selected separately for each scenario: the mixture searches one to eight components; HDBSCAN uses an 8% minimum cluster size and 10-point core rule; mean shift estimates one common 20th-percentile bandwidth; and OPTICS uses a 1% neighborhood minimum, an 8% minimum cluster size, and . Accuracy of K, matched center-position error, and computation time are recorded. Five ablations remove cross-scale persistence, the density-valley rule, the minimum-share rule, adaptive separation, or the 99.5% fitting range one at a time.
Because the expanded stress set informed the final parameter selection, it is treated as model-selection evidence rather than an independent evaluation. A separate synthetic evaluation was specified after the algorithm and thresholds had been finalized. It contains ten new geometries—rotated anisotropy, unequal three-center and irregular five-center mixtures, 7% and 9% satellites, uniform contamination, an off-center address spike, a two-node coastal arc, four unequal scales, and a skewed one-center corridor—with 20 repetitions each. Its generating scenarios, expected substantive K, random seed, comparison methods, and parameter settings were determined before evaluation and remained unchanged throughout the analysis. All 1,200 method evaluations are reported, including failures.
Parameter sensitivity is evaluated on all 345 observed units. For the substantively important 6/8/10% minimum-share choice, the complete analysis is repeated: 6% and 10% each use 20 spatial-block repetitions with full K reselection, every A/B/C class profile is recalculated, and the same class-wise false-discovery-rate rule is applied. Center count is share-stable only when all three selected K values agree. A primary center is share-stable when both alternative centers lie within of the 8% center; a city-class conclusion is stable only when its corrected label agrees at 6%, 8%, and 10%. These three statements are intentionally separate. Valley ratios of 0.40/0.50/0.60 and adaptive distance fractions of 0.06/0.08/0.10 retain the lighter one-factor comparison because they scarcely changed K in the initial all-city grid.
Center fitting uses all classes, so the statement that labels are not supplied as model features does not by itself rule out reference dependence: a concentrated rare class could still help create the center against which it is compared. At the fixed 8% definition, a leave-class-out diagnostic therefore subtracts the C-weight contribution of S, F, or H, refits the corresponding center system, and recalculates the class contrast. Full-data and leave-class-out fits are both deterministic in this diagnostic; resampling uncertainty remains represented by the primary model. We compare K, primary-center displacement, and the corrected city conclusion. This is a sensitivity check, not a replacement center definition.
Real-data external agreement was assessed against an independently constructed GHSL 2020 settlement reference [21]. For each of the 323 observed units with an accepted auxiliary boundary, the boundary was applied to the complete GHSL grid rather than to registration-occupied cells. The most restrictive SMOD tier among class 30, class 23 or above, and class 21 or above was retained when it contained at least three grid cells and 10% of boundary population. Eight-neighbor components carrying at least 5% of population in that tier were kept, up to 12 per unit; component rank and centroid used GHSL population alone. Registration coordinates, counts, classes, and fitted centers played no part in constructing this reference. Twenty-one units lacked a matched boundary. Tieling was also excluded because its licensed auxiliary polygon contains a disconnected southern component that captures unrelated settlements; no registration coordinate was used to trim or repair that boundary. These 22 units were marked unavailable rather than assigned a registration-derived envelope.
The first comparison pairs the primary registration and GHSL nodes. A second uses minimum-cost one-to-one matching between all centers in the same unit. Because this produces only pairs, unmatched nodes and distance shares using both matched pairs and all evaluated registration centers are reported separately. GHSL reflects population and built settlement rather than legal registration, so these distances test cross-source positional agreement, not the true registration-center count. The auxiliary boundaries are older than the registration snapshot and are used only for this external check, never to reassign source records. Address sensitivity and spatial-resampling status provide two further checks: the first asks whether shared registration addresses change the result, and the second asks whether the same center structure survives neighborhood-level perturbation.
The GHSL construction was developed during manuscript review and is therefore a post-review external check, not a preregistered validation. Its one-factor sensitivity varies the tier population requirement from 10% to 5/15%, component share from 5% to 3/8%, and node cap from 12 to 8. Every setting rebuilds the reference from GHSL and boundaries alone. We compare selected tier, GHSL node count, primary-reference displacement, primary registration distance, and all-node matching, thereby separating primary-location robustness from secondary-node sensitivity.
4. Results
4.1. Data Conservation and National Coverage
The formal analysis layer contains 185,033,392 deduplicated registration entities. The same total is conserved when entities are aggregated to exact sites, H3 fitting cells, city profiles, center catchments, and registration agglomeration zones. The final city inventory contains all 364 source city-level units: 345 have usable records and 19 are retained as explicitly non-estimable because no source sample was available. This distinction matters. The national findings describe the 345 observed units and must not be interpreted as evidence that the 19 remaining units lacked registered entities in reality.
All city-center models use the total entity distribution without reference to R/S/F/H. The rare-class eligibility rule later yielded 307 formal city comparisons for S, 263 for F, and 220 for H. H spans 30 province-level clusters, whereas S and F span 31. Variation in eligible counts reflects the minimum requirement of 100 raw class entities across eight spatial blocks, not discretionary removal after inspecting effects.
4.2. Synthetic Recovery, Competing Methods, and Ablations
Table 3 shows that MSP-MBKMeans recovered the prespecified number of centers in all 120 synthetic evaluations (100.0%). The legacy one-standard-error MiniBatch K-Means selector was correct in 20 of 120 evaluations (16.7%). Its characteristic error was over-selection: clear one-, two-, and four-center patterns were frequently split into six to eight clusters. The comparison used exactly the same simulated coordinates and weights, so the difference is attributable to the selection framework rather than easier scenarios for MSP-MBKMeans.
The expanded benchmark changes the strength, but not the direction, of the algorithmic claim (Table 4). MSP-MBKMeans, HDBSCAN, and OPTICS all recovered K in 100.0% of the six original scenarios. In the nine stress scenarios, HDBSCAN and OPTICS reached 88.9%, while MSP-MBKMeans reached 87.8%. MSP-MBKMeans failed deliberately on the curved non-convex field by returning several persistent modes and missed one of 20 overlapping and one of 20 unequal-scale pairs. HDBSCAN and OPTICS instead merged every overlapping pair. The methods therefore have different failure boundaries; the experiment does not support universal superiority.
The ablations identify which constraints carried the strongest observed protection. Removing the 8% minimum-share rule reduced original-scenario accuracy from 100.0% to 82.5% and stress accuracy from 87.8% to 46.1%. Removing cross-scale persistence reduced the two rates to 83.3% and 78.3%; removing the density-valley rule reduced them to 97.5% and 81.1%. Adaptive separation had a small aggregate effect in this benchmark (87.2% under stress), while removing the 99.5% fitting range changed no synthetic K. The latter controls remain relevant to geographically extensive real units and remote outliers, but the synthetic evidence should not be overstated.
The prespecified synthetic holdout provides the stronger generalization check. MSP-MBKMeans recovered the substantive center count in 198 of 200 new cities (99.0%), with a 149.5 m median matched-center error when K was correct. Its two misses occurred in the unequal-scale four-center geometry. HDBSCAN and OPTICS each reached 100.0%, while GMM-BIC, mean shift, and the legacy selector reached 60.0%, 58.5%, and 20.0%, respectively (Table 5). The result supports strong recovery in the prespecified holdout while preserving the narrower claim that MSP-MBKMeans is not the universally most accurate clustering method. Because the author designed these scenarios after examining the method, this remains an internal holdout rather than an externally preregistered benchmark.
The real-city threshold analysis separates center-count sensitivity, primary-position sensitivity, and class-location stability. The initial one-factor grid showed that changing the minimum share from 8% to 6% or 10% changed raw K much more often than changing the valley or distance settings (Table 6). The share rule was therefore subjected to the complete confirmatory analysis rather than left as a local check. After 20 spatial repetitions, the consensus-K medians were five, four, and four at 6%, 8%, and 10%; agreement with the fixed 8% K was 39.7% and 47.2% at the two alternatives. Only 246 cities (71.3%) kept the primary center within the adaptive displacement limit at both alternatives. Corrected city-class conclusions agreed at all three shares in 310 S comparisons (89.9%), 313 F comparisons (90.7%), and 316 H comparisons (91.6%), or 939 of 1,035 comparisons (90.7%) overall. Table 7 reports the full corrected directional counts, and the catalogue reports the three K values and primary-position status for each city. Thus, the 1,410 centers and the city forms are results of the fixed 8% definition; a stable class statement must not be read as proof that the number or exact position of every center is definition-independent.
In the 345 observed cities, MSP-MBKMeans selected a median of four centers versus six under the legacy method; the legacy selector chose more centers in 76.8% of cities. Figure 2 summarizes this difference with the benchmarks, ablations, and GHSL distances. Four is not a universal urban optimum; it is the median under the stated substantive constraints.
4.3. Observed Center Counts and Reliability
Under the fixed 8% C-weight primary definition, 1,410 centers were identified across 345 cities. The city distribution was for 17 cities, for 45, for 65, for 81, for 64, for 50, for 20, and for 3. The mode and median were both four. This distribution should not be converted into a development ranking: a geographically large prefecture that contains several county nodes can have a larger K than a compact, economically intensive municipality. It is also conditional on the 8% minimum-share definition, as the separate 6/8/10% catalogue fields make explicit.
Table 8 shows how spatial resampling separated structural information from uncertainty. Exactly stable center counts were obtained for 69 cities. Another 88 retained stable core centers but had optional peripheral centers. For 163 cities, the principal structure remained usable only with an explicit center-count uncertainty label. Twenty-five cities were unresolved and were excluded from formal center–periphery inference. Put differently, 157 cities had either exactly stable counts or stable cores; 320 had at least a qualified usable structure; and the method openly declined to make a formal structural claim for the remaining 25. The lower exact-stability count than in the initial model is a direct consequence of admitting unequal-scale candidate peaks and repeating K selection 20 times under every address specification.
The three address specifications provide a different diagnostic. Fifty-seven cities had identical center counts under A/B/C, 67 passed the full-structure comparison that allows at most one optional center, and 209 had a primary center stable under all three specifications. These numbers explain why center reliability and address robustness are reported separately: a city can have a stable C-weight bootstrap solution yet move its primary center when thousands of entities sharing one exact address are changed from raw multiplicity to address-equal weight.
The repetition-count diagnostic clarifies the resolution of these labels without changing the prespecified 20-repetition primary definition. The structured sample contains the smallest, median-sized, and largest H3 input within each of the four primary reliability levels. Compared with 100 repetitions, the 20-repetition result retained the same K and reliability level in 10 of 12 cities; 50 repetitions retained the same K in 10 and the same reliability level in 11 (Table 9). All three cities sampled from the exactly stable group remained exactly stable, and all three sampled unresolved cities remained unresolved. Changes concerned optional centers or the boundary between core-stable and usable-with-uncertainty. Median primary-center displacement was effectively zero because most primary centers were unchanged, but the largest displacement reached 3.54 km at 20 repetitions and 2.78 km at 50. These results support the reliability grades as coarse diagnostics while confirming that their numerical appearance rates are not precise probabilities.
4.4. External Positional Agreement with GHSL
The independent comparison covers 323 observed units with an accepted auxiliary boundary; 22 observed units are unavailable. The MSP-MBKMeans primary registration center was a median 3.32 km from the population-weighted GHSL 2020 primary node, with a 7.31 km 75th percentile. Overall, 64.4% of evaluated cities were within 5 km, 82.7% within 10 km, and 90.1% within 20 km. Median and 75th-percentile distances were 2.11 and 3.62 km for exactly stable cities, 5.11 and 10.97 km for stable-core cities, 3.09 and 6.23 km for usable-with-uncertainty cities, and 4.36 and 8.01 km for unresolved cities.
The all-center comparison evaluates 1,349 registration centers and 1,337 independently constructed GHSL nodes in those 323 units. Minimum-cost matching produced 1,130 pairs and left 219 registration centers and 207 GHSL nodes unmatched. Among matched pairs, median separation was 6.66 km; 40.3% were within 5 km, 63.4% within 10 km, and 81.2% within 20 km. Using all 1,349 evaluated registration centers as the denominator, the corresponding confirmed-coverage shares were 33.7%, 53.1%, and 68.1%. Of 1,026 evaluated secondary registration centers, 810 were matched and 216 were unmatched; matched secondary nodes had an 8.69 km median separation, with 54.7% within 10 km and 76.4% within 20 km. These results support broad positional agreement but do not justify claiming that every fitted secondary node has an external counterpart or that GHSL supplies the true registration K.
Large deviations were not automatically classified as failures. The ten largest accepted comparisons involved Aba Tibetan and Qiang Autonomous Prefecture (227.1 km), Qamdo (225.9 km), Heihe (194.0 km), Yichun in Heilongjiang (139.5 km), Altay Prefecture (124.2 km), Suihua (108.5 km), Huanggang (95.8 km), Shangluo (83.2 km), Gannan Tibetan Autonomous Prefecture (68.4 km), and Dingxi (66.1 km). Many are extensive administrative territories with separated settlements. Tieling’s former 221.8 km value was removed after review showed that a disconnected auxiliary-boundary component captured unrelated southern settlements. GHSL and MSP-MBKMeans can otherwise emphasize different nodes because one represents population-weighted settlement morphology and the other cumulative registration density.
The post-review GHSL sensitivity separates a stable primary comparison from definition-dependent secondary counts (Table 10). Moving the tier population requirement to 5% or 15% left the primary-distance median at 3.32 and 3.29 km and 82.7% and 82.4% within 10 km. Changing component share from 5% to 3% or 8% did not move any primary node, but changed the reference count from 1,337 to 1,690 or 918 and changed all-node 10 km agreement from 63.4% to 64.5% or 70.6%. The apparent improvement under 8% follows from retaining fewer secondary nodes; it is not evidence that 8% is truer. Capping at eight nodes changed K in only 1.9% of units. Primary positional agreement is robust in this grid, whereas secondary-node enumeration remains conditional on the component definition.
4.5. What Spatial Patterns Do the Four Classes Actually Show?
The class labels do not enter center fitting. More central means a shorter mean distance than R with a direction supported across A/B/C; more peripheral means longer. No clear difference means no repeatable ordering, not identical distributions, and address-sensitive means that shared-address treatment changes the direction.
The city counts in Table 11 make the national pattern concrete after multiplicity control. F is more central in 79 cities and more peripheral in 8; H is more central in 82 and more peripheral in 3. Both F and H are more central in 41 cities. Within those 41, S shows no clear difference from R in 27, is more peripheral in 6, and is also more central in 8. S alone is more central in 24 cities and more peripheral in 25, so neither “state entities are always central” nor “state entities are always peripheral” describes the country. In 82 cities, S, F, and H all show no clear difference from R. This does not prove identical distributions; it says that the available evidence does not support a repeatable ordering after address checks and false-discovery-rate control.
The correction deliberately changed the evidential status of many screening results. Before correction, the 1,035 city-class comparisons contained 234 more-central, 60 more-peripheral, and 18 address-sensitive screening labels. The corresponding formal counts are 185, 36, and 7. These are not lost observations: most moved to “no clear difference” because their individual intervals were not strong enough relative to the number of city comparisons. Among formally eligible comparisons, usable delete-block counts were never close to the minimum: minima for S/F/H were 17/55/55 and medians were 459/476/466.
Table 12 reports the central national result. F had positive A/B/C primary gaps of 3.25, 2.47, and 3.74 km; H had 4.39, 3.28, and 4.57 km, with every interval excluding zero. Address equality (B) attenuated but did not erase either pattern. The probability indices show modest distributional shifts, not percentages located downtown. S had gaps of 0.04, , and km, so its signs and intervals did not support a national direction.
Under C, nearest-center gaps were 1.06 km for F and 1.37 km for H, smaller than their primary-center gaps. Figure 3 therefore supports orientation toward the dominant registration node, not a causal location preference.
4.6. Strict Reliability and Address-Robustness Subsets
The main F/H result survived progressively narrower quality subsets. Among cities with exactly stable counts or stable cores, the C-specification median gap was 3.20 km for F () and 4.00 km for H (); S remained near zero at km (). Among cities whose primary center was address-robust, medians were 3.99 km for F (), 4.76 km for H (), and 0.03 km for S ().
The most restrictive subset required both an exactly stable center count and an A/B/C-robust primary center. F retained a 2.49 km median advantage across 43 cities, with a 95% interval of [1.57, 6.75] km. H retained a 3.46 km advantage across 35 cities, with [1.14, 8.59] km. S was 0.21 km across 49 cities, with [, 1.79] km. Thus, the F/H finding was not driven by unresolved center structures or by cities whose primary center depended on concentrated addresses. Conversely, the absence of a uniform S direction was reproduced under the strictest quality conditions.
At the city-class level, the corrected 1,035 comparisons comprise 185 more-central, 36 more-peripheral, 7 address-sensitive, 562 no-clear, and 245 not-estimable results. The 41-city joint F/H pattern and the 82 cities with no clear S/F/H difference are evidence categories, not city rankings. The complete threshold propagation further shows that 939 of 1,035 labels (90.7%) are identical at 6%, 8%, and 10%. Results that change are explicitly marked in the city catalogue rather than removed.
The leave-class-out check shows that the main reference is rarely created by the class being compared (Table 13). Removing S, F, or H preserved deterministic K in 314, 332, and 332 cities and preserved the corrected city conclusion in 342, 343, and 343 cities. Across all 1,035 comparisons, 1,028 conclusions (99.3%) agree. The seven changes occurred in Taizhou (S), Huainan (S), Hebi (H), Zhongshan (S and H), Qiannan (F), and Wenshan (F). They remain flagged in the catalogue; the high overall agreement does not license hiding those exceptions.
The conclusion was also stable to the mixture of administrative units in the source inventory. Under C, the full 364-unit inventory yielded median gaps of , 3.74, and 4.57 km for S, F, and H. Restricting the analysis to 330 prefecture-level units yielded , 3.74, and 4.66 km; restricting it further to 291 standard prefecture-level cities yielded , 3.73, and 4.66 km. In all three inventories the median selected center count was four. Table 14 therefore shows that F/H centrality and the absence of a common S direction are not created by municipalities, autonomous prefectures, regions, or directly administered county-level units.
4.7. Variation by Center Structure, Archive Size, and Province
Figure 4 shows positive F/H medians in all reportable center structures and weaker gaps in the lowest size quartile, but no monotonic size relationship. Province medians were positive in 22/23 eligible F and 15/18 eligible H province units; S split 15 positive versus 12 negative. These are descriptive distributions, not effects of structure, size, or province.
4.8. Registration Agglomeration Zones
The 1,410 centers produced 1,410 catchments covering 1,571,280 occupied cells and conserving all entities. A total of 1,359 zones reached the 50% C-weight target; 51 remained incomplete. All 64,975 one-cell bridges stayed within their originating catchment, so none joined different center systems.
Combined zones carried a median 48.7% of raw city entities (IQR 46.2–51.2%) at a median density of 482 entities/km2 (Figure 5). Ninety-nine zones belonged to unresolved center structures and remained exploratory; 45 zones attached to otherwise reliable centers were incomplete.
Prominent enrichment required both and the conditional random-label upper bound; it describes unusual within-city composition, not an industrial cluster or causal formation process.
4.9. Eight Illustrative Cities and Complete City Outputs
Figure 6 and Figure 7 show Beijing, Shanghai, Wuhan, Chengdu, Xi’an, Chongqing, Guangzhou, and Shenzhen using the same visual grammar. The maps display estimated centers, nearest-center catchments, and high-density registration zones. They were chosen to illustrate different administrative extents and center structures, not because they had the strongest effects. Guangzhou is particularly useful as a caution: its unresolved status means that mapped centers remain exploratory rather than a formal city structure conclusion.
Supplementary Table S1 gives the corresponding conclusion for every source unit. Each entry states the center form, the S/F/H pattern relative to R, the concentration carried by registration zones, and the reliability limit. This is the main city-level product of the study: it distinguishes a substantive finding from an unavailable or unresolved comparison without turning the cities into a ranking. Keeping the catalogue as a separately paginated, searchable part of the article package preserves all 364 results without interrupting the national argument.
5. Eight Cities in Plain Language
The eight cities in Figure 6 and Figure 7 show how the same measures lead to different, readable conclusions. They are examples rather than a ranking. Table 15 reports the evidence used in the descriptions. A positive class gap means that the class is closer to the primary registration center than R.
Beijing. One strong primary center carries 70.0% of fitted weight and is accompanied by two secondary centers whose status is less certain. S, F, and H are all closer to the primary center than R, by 5.37, 9.23, and 7.53 km. Three registration zones carry 48.4% of entities. The clear finding is four-class centrality; the exact secondary-center list is not fixed.
Shanghai. Four centers share the registration pattern more evenly: the primary center carries 51.3%. S, F, and H are all more central than R, with F showing the largest gap at 9.42 km. The primary center is 7.36 km from the independent GHSL reference, so registration geography and settlement geography are close but not identical.
Wuhan. All C-weight resamples support two centers, led by a primary center carrying 65.6% of fitted weight. S, F, and H are all more central than R, by 3.77, 3.02, and 4.62 km. The two registration zones carry 45.5% of entities at high density. The center count is stable even though address-equal weighting can add one further minor center.
Chengdu. A strong primary center carries 76.4%, with two smaller centers and some uncertainty about the exact count. F and H are clearly closer to the primary center than R; S shows no clear difference. This is the common national pattern in a concrete city: international classes are central, while S does not follow a fixed direction.
Xi’an. This is the cleanest two-center case: all three address specifications select two centers and the count is stable. F and H are closer to the primary center than R, but S is 2.19 km farther away. Xi’an shows why the national S result is not a weak version of F/H; its direction can genuinely differ by city.
Chongqing. Two C-weight centers organize an exceptionally large municipal area, while the A/B specifications select three and four. F and H are about 60 km closer to the primary center than R, while S has no clear difference. These large values reflect Chongqing’s spatial extent and should not be compared directly with compact cities. Its zones cover a large area but have low density.
Guangzhou. Three density peaks are visible and the primary peak is close to GHSL, but the center system changes too much under spatial perturbation. The apparent class gaps are therefore not formal findings. Guangzhou is included to show that a plausible map does not override unstable center evidence.
Shenzhen. Four relatively balanced centers accompany the highest zone density among the eight cities. Yet S, F, and H show no clear distance difference from R. High registration density therefore does not automatically produce strong separation among historical entity classes.
Together, the cases show that a city conclusion needs three pieces of information: center form, class position, and evidence strength. Any one of these alone can mislead. The complete catalogue applies this same plain-language structure to all 364 source units.
6. Discussion
6.1. What the National Pattern Means
The main empirical result is straightforward. Legal addresses identified as foreign (F) or as Hong Kong, Macao, or Taiwan (H) more often lie near the principal concentration of all registrations. After class-wise false-discovery-rate control, F is more central in 79 cities and more peripheral in 8, while H is more central in 82 and more peripheral in 3. The city-equal typical differences are 3.74 and 4.57 km and keep the same national direction under three treatments of shared addresses. State-identified registrations (S) have no common national direction, being more central in 24 cities and more peripheral in 25. These are shifts in address distributions, not claims that every F or H entity is central, that R avoids centers, or that a central address improves performance.
Three recurring city patterns make the national result concrete. In 8 cities, S, F, and H are all closer to the primary center than R. In 27, F and H are closer while S has no clear difference; in 6, F and H are closer but S is farther out. In 82 cities, none of S/F/H has a clear difference from R after correction. Other cities show only one class difference, address or center-share sensitivity, insufficient observations, or an unresolved center system. This classification identifies intelligible regularities without inventing an economic story for every city.
Several mechanisms could produce F/H centrality: access to professional or administrative services, finance, transport, bilingual networks, established foreign-business districts, recognizable central addresses, or centrally located registration agents. The data cannot distinguish them. Likewise, the mixed S pattern may combine central administrative, financial, and headquarters registrations with peripheral utilities, industry, transport, resources, and development-zone entities. The surviving S label is too coarse to separate these forms. Both interpretations are hypotheses for linked industry, administrative, and longitudinal data, not mechanisms established here.
F/H differences are larger relative to the primary center than to the nearest center. An entity in a polycentric city may be close to a secondary node without being close to the dominant citywide node. Reporting both distances therefore preserves center hierarchy while recognizing meaningful secondary concentrations.
6.2. Why Concentrated-Address Sensitivity Matters
Mass registration at one legal address can alter the density surface, the primary center, and class distances. If a service address hosts many F or H records, raw counts may help create both the reference center and the apparent class advantage, even though class labels do not enter center fitting.
The A/B/C design makes this risk visible. A counts every entity, B gives each exact location equal weight, and C bounds only extreme multiplicities. B is not automatically more truthful: a genuine complex containing many independent organizations need not count the same as an isolated address. C is the prespecified compromise, while agreement across all three specifications supports the strongest interpretation. F/H effects shrink under B but remain positive with intervals excluding zero; under C they are similar to or somewhat larger than under A. Thus some concentrated addresses may amplify raw patterns, while bounding extreme sites can also prevent unrelated mass registrations from shifting the center.
Only 7 of 1,035 city–class comparisons retain an address-sensitive label after multiplicity control, down from 18 screening labels, but this does not prove that address bias is absent: 245 comparisons are unavailable or unresolved, and many others show no clear difference. A/B/C controls observed coordinate multiplicity; it cannot determine whether any unique address is an operating location. The leave-class-out check addresses a related concern: 1,028 of 1,035 conclusions remain unchanged after the compared class is removed from center construction. The seven exceptions remain named in the catalogue.
6.3. Methodological Contribution and Its Boundary
The study does not claim to invent K-means, MiniBatch K-Means, density peaks, hierarchical grids, bootstrap, Voronoi allocation, location quotients, or hypergeometric inference. MSP-MBKMeans contributes a task-specific connection among them. A candidate center must persist across spatial resolutions before clustering and must then satisfy separation, an intervening density valley, and a post-fit share requirement. The expanded benchmark led to removal of an early relative peak-score cutoff, retaining original recoveries while reducing missed unequal-scale centers. The final method rejects elongated single-center splitting and a 4% satellite that generic held-out quantization error over-partitions.
Uncertainty is part of the output. Re-selecting K in every spatial-block bootstrap replicate evaluates whether the reported center system itself is repeatable. A separate A/B/C comparison evaluates sensitivity to shared-address weighting. Their combination distinguishes exactly stable systems, stable main centers with optional minor nodes, broader count uncertainty, and unresolved cities; it also provides a strict subset for testing the F/H result. The 25 unresolved cities remain visible rather than being removed.
The method then converts centers into catchments, density-priority zones, and conditional class enrichment for every source unit. This end-to-end structure is useful because each city result can be traced to explicit spatial and reliability conditions. It is nevertheless a domain-designed assembly of established methods, not a universal clustering theorem. MSP-MBKMeans tied HDBSCAN and OPTICS in the original six scenarios but was one percentage point lower in the expanded stress set. Its advantage is an explicit, substantively constrained K, scalable MiniBatch K-Means fitting, and repeatable reliability reporting. Curved non-convex fields, other registration systems, or other grid hierarchies may require different thresholds or a non-K-means fitting stage.
6.4. Center Counts and Independent Positional Evidence
The median of four centers is not an urban ranking. Administrative units differ in area, settlement hierarchy, terrain, and included counties; a large autonomous prefecture may contain several legitimate county-seat nodes, while a compact municipality may have one dominant registration concentration. K means only the number of registration nodes supported within the source unit under the fixed 8% rule. Reliability is therefore essential: among 345 observed cities, 69 have exactly stable counts, 88 have optional peripheral centers, 163 have broader count uncertainty, and 25 are unresolved. The 6/8/10% sensitivity analysis adds a different uncertainty: only 39.7% and 47.2% retain the 8% count at 6% and 10%, respectively. The catalogue reports this definition sensitivity separately from fixed-rule resampling stability.
The independent GHSL comparison supports primary-center position without defining registration-center truth. Across 323 boundary-supported units, the primary-node distance has a 3.32 km median and 82.7% are within 10 km. Among all nodes, 63.4% of 1,130 matched pairs are within 10 km, while 219 of 1,349 evaluated registration centers remain unmatched. Of the secondary centers, 810 are matched and 216 are unmatched; 54.7% of matched secondary centers lie within 10 km. Reference-parameter changes leave primary agreement nearly fixed but alter secondary-node counts, so the evidence is stronger for primary position than for secondary count. Tieling is excluded because its auxiliary boundary contains an anomalously disconnected component. GHSL measures built settlement, not legal registration; employment, activity, or verified local landmarks would test different meanings.
6.5. What Registration Zones Add
Centroids do not show how far a concentration extends. The catchment-relative zones add a footprint to each node and carry a median 48.7% of city registration weight, close to the intended half-weight target. The 51 incomplete zones identify catchments that cannot reach the target without crossing excessive empty space or accepting tiny fragments; forcing completion would create smoother but less defensible maps. No morphological bridge crosses a center boundary.
Conditional LQ measures composition relative to the same city. A zone with many F entities need not be enriched if F is already common citywide, while a smaller count can be enriched where F is rare. Requiring both a 20% proportional elevation and a random-label exceedance improves interpretability, but still measures composition rather than supplier links, knowledge exchange, or production. Used under the correct label, zones can guide registry-quality review, comparisons with employment or POI centers, and targeted field validation; they cannot by themselves establish transport or infrastructure demand.
6.6. National Inference, Boundaries, and Limitations
Equal city weighting estimates the typical eligible city contrast, not the experience of a randomly selected entity. Province-cluster intervals acknowledge regional dependence but cannot remove the modifiable areal unit problem [58]. Source units mix municipalities, prefecture-level cities, autonomous prefectures, and special county-level units. Because local polygons were outdated, they were not used to reassign entities; this preserves the source inventory but retains administrative heterogeneity. Restricting the analysis to prefecture-level units and then standard prefecture-level cities leaves C-specification F/H gaps near 3.7/4.7 km and S near zero. Administrative composition therefore does not explain the main direction, although boundaries within units can still matter. Structure, size, and province splits remain descriptive rather than causal.
The largest limitation is semantic: registration location is not verified activity location. Address weighting cannot support claims about employment, output, commuting, emissions, land use, or infrastructure. The snapshot is cumulative, contains multiple registration states, and lacks a complete event history, so it cannot reveal migration, center formation, survival, or policy effects. Class labels are also coarse: the lost original organization-type field prevents recovery of 429 mapping strings; R is private-dominated rather than purely private; and S/F/H are historical identified classes rather than current ownership verification.
Geocoding quality cannot be reconstructed without original addresses. BD-09 to WGS84 conversion removes a systematic offset, and H3 aggregation plus the 99.5% fitting range reduces isolated anomalies, but neither removes city-specific error. Thresholds also embody domain judgment. The complete 6/8/10% sensitivity analysis preserves 90.7% of city-class labels, but center count agrees with the 8% definition in only 39.7% and 47.2% of cities at 6% and 10%, and only 71.3% retain the primary center under the joint displacement rule. The selected 8% is an explicit substantive definition, not a unique optimum. Likewise, the 12-city repetition-count diagnostic preserves the reliability level in 10 cities at 20 versus 100 repetitions and in 11 at 50 versus 100; the four levels are therefore practical evidence grades, not precise stability probabilities. The catalogue marks each kind of threshold sensitivity separately. The GHSL check covers 323 of 345 observed units, was constructed during review rather than preregistered, uses older auxiliary boundaries, and compares settlement with registration. Twenty-one units lack a matched boundary and Tieling is excluded. Finally, the entire study is descriptive: sector, history, capital, policy, rent, transport, services, and data practice can all correlate with center distance, so F/H differences are not ownership effects or causal preferences.
6.7. Future Research
The most useful extension is targeted linkage rather than another nationwide raster. Establishment and status dates could create cohorts; industry codes could test composition; verified operating addresses could estimate registration-to-activity displacement; and rent, transport, development-zone, and administrative-service data could test mechanisms. Method development can compare alternative grids, estimate center count and location jointly, and validate against employment, POI, transit, nighttime light, and local records. The complete 364-unit catalogue makes a balanced validation sample possible: cities can be selected across reliability, GHSL distance, structure, size, and address robustness instead of checking only famous cases.
7. Conclusions
This study turns 185 million cumulative historical registration coordinates into three readable outputs: each city’s number of registration centers and its reliability, the position of S/F/H relative to the private-dominated R class, and the high-density registration zones around those centers. Under the fixed 8% primary definition, 69 of 345 observed units have an exactly stable center count, 88 have stable main centers with optional small nodes, 163 have a readable main pattern with count uncertainty, and 25 remain unresolved. The 1,410 fitted centers support 1,410 zones, of which 1,359 reach the prespecified coverage target without crossing center boundaries. The 6/8/10% sensitivity analysis shows why this definition must remain visible: center counts agree with the 8% result in only 39.7% and 47.2% of cities at 6% and 10%, while primary-center position remains stable at both alternatives in 246 cities (71.3%).
The four-class result is equally concrete after correction for 1,035 city comparisons. F is more central in 79 cities and more peripheral in 8; H is more central in 82 and more peripheral in 3. Their city-equal typical advantages over R are 3.74 and 4.57 km. Both are more central in 41 cities: S has no clear difference in 27, is more peripheral in 6, and is also more central in 8. Across all cities, S is more central in 24 and more peripheral in 25, while 82 cities show no clear S/F/H difference from R. The main national regularity is therefore F/H orientation toward the primary registration node; S varies by city, and many cities have no supported class ordering. Restricting the inventory to prefecture-level or standard prefecture-level cities leaves the national directions nearly unchanged.
This finding has social-economic value without requiring a causal story. Historically identified internationally connected registrations are more closely associated with the leading legal-address concentration, potentially through professional and administrative services, established foreign-business districts, recognizable addresses, or registration intermediaries. Central administrative and headquarters-type S registrations may coexist with peripheral utilities, industry, transport, resources, and development-zone entities. Industry, organization subtype, and longitudinal data are needed to distinguish these explanations.
The city catalogue is a core output, not an appendix of raw indicators. All 364 source units receive an explicit conclusion: 345 rows describe spatial form, class position, registration concentration, and confidence in complete sentences; 19 report that no usable coordinates are available. Beijing, for example, has a dominant center with secondary nodes and S/F/H all closer than R; Wuhan has a stable two-center structure with the same class ordering. Urumqi is a stable single-center city where F is closer, S has no clear difference, and H is address-sensitive. Guiyang and Guangzhou have visible density peaks but unresolved center systems, so class gaps are not promoted to formal findings. No-clear, unresolved, and unavailable outcomes are scientific results rather than missing success cases.
MSP-MBKMeans retains MiniBatch K-Means for scalable fitting but constrains how centers enter the fit and how uncertainty is reported. Candidates must persist across two spatial scales, carry a meaningful share, remain separated, and have a density valley between them; K is re-selected in every spatial-block resample. The prespecified holdout achieves 99.0% correct K across 200 new synthetic cities, compared with 100.0% for HDBSCAN and OPTICS. Although the share threshold changes many center counts, 939 of 1,035 corrected city-class labels are identical at 6%, 8%, and 10%, and 1,028 remain identical when the compared class is removed before refitting the center. The catalogue therefore gives each city its three center counts, primary-position status, and class-label sensitivity separately. Independent GHSL evidence supports primary-center position but retains 219 unmatched registration centers. The method is therefore well supported for this registration-center task, not universally superior and not proof of every city’s physical center count.
For future registration-data projects, four principles follow. First, use one documented coordinate system and reconcile city totals before interpretation. Second, compare raw counts, address equality, and bounded extreme-address weights. Third, fit centers from the full registration field without supplying class labels as features, then verify with leave-class-out centers that a rare class did not create its own reference; report structure, class position, zones, and reliability together. Fourth, retain unresolved and no-sample cities, and never interpret a larger K or denser zone as higher economic development. Combining uncertainty, address sensitivity, center-definition sensitivity, and GHSL discrepancy can guide balanced address verification and later linkage to industry, transport, rent, policy, and operating locations.
The conclusion is deliberately narrow. China’s cumulative business-registration archive contains repeatable urban registration centers, although confidence in their number varies by city. F and H registrations are generally more oriented toward the primary registration center than R, whereas S has no national direction. Registration zones locate these concentrations, and the city catalogue states where each conclusion is supported, uncertain, or unavailable. These findings concern legal registration locations, not production, employment, office activity, performance, or causal ownership effects.
Supplementary Materials
The following supporting information can be downloaded at the website of this paper posted on Preprints.org.
Author Contributions
Conceptualization, H.W.; methodology, H.W.; validation, H.W.; formal analysis, H.W.; investigation, H.W.; resources, H.W.; data curation, H.W.; writing—original draft preparation, H.W.; writing—review and editing, H.W.; visualization, H.W.; project administration, H.W. The author has read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Institutional Review Board Statement
Not applicable. The study analyzes organizational registration records and contains no experiment involving human participants or animals.
Informed Consent Statement
Not applicable.
Data Availability Statement
The underlying entity-level records originated from Tianyancha and were lawfully acquired for institutional research use. They cannot be redistributed because of third-party copyright and institutional contractual restrictions. For confidential assessment, the analysis code, synthetic-data generator and specifications, non-identifying city-level aggregate results, and machine-readable bilingual city catalogue will be provided to journal editors and reviewers upon request. Readers may request the same non-identifying research materials from the corresponding author at loveweihaitong@foxmail.com, subject to the original rights and agreements. Enterprise-level records, coordinates, and identities cannot be provided.
Acknowledgments
The author thanks the institutions listed in the affiliations for the professional setting in which the research problem was developed. Language and formatting assistance was used during manuscript preparation; all methodological choices, analyses, interpretations, and final wording were reviewed and approved by the author.
Conflicts of Interest
The author declares no conflict of interest.
References
- Wang, Y. Handling of Inconsistencies Between the Registered Address and the Actual Business Address. In Essential Knowledge and Legal Practices for Establishing and Operating Companies in China; Springer, 2022; pp. 163–169. [CrossRef]
- Tezer, A.; Bodur, H.O.; Grohmann, B. Location, Location ... Mailing Location? The Impact of Address as a Signal. Journal of Business Research 2021, 130, 683–698. [CrossRef]
- Zhao, M.; Liu, X.; Derudder, B.; Zhong, Y.; Shen, W. Big Enterprise Registration Data Imputation: Supporting Spatiotemporal Analysis of Industries in China. Computers, Environment and Urban Systems 2018, 70, 9–23. [CrossRef]
- Arribas-Bel, D. Accidental, Open and Everywhere: Emerging Data Sources for the Understanding of Cities. Applied Geography 2014, 49, 45–53. [CrossRef]
- Lloyd, S.P. Least Squares Quantization in PCM. IEEE Transactions on Information Theory 1982, 28, 129–137. [CrossRef]
- Sculley, D. Web-Scale K-Means Clustering. In Proceedings of the Proceedings of the 19th International Conference on World Wide Web, 2010, pp. 1177–1178. [CrossRef]
- Rodriguez, A.; Laio, A. Clustering by Fast Search and Find of Density Peaks. Science 2014, 344, 1492–1496. [CrossRef]
- Rousseeuw, P.J. Silhouettes: A Graphical Aid to the Interpretation and Validation of Cluster Analysis. Journal of Computational and Applied Mathematics 1987, 20, 53–65. [CrossRef]
- McDonald, J.F. The Identification of Urban Employment Subcenters. Journal of Urban Economics 1987, 21, 242–258. [CrossRef]
- Giuliano, G.; Small, K.A. Subcenters in the Los Angeles Region. Regional Science and Urban Economics 1991, 21, 163–182. [CrossRef]
- McMillen, D.P. Nonparametric Employment Subcenter Identification. Journal of Urban Economics 2001, 50, 448–473. [CrossRef]
- Redfearn, C.L. Persistence in Urban Form: The Long-Run Durability of Employment Centers in Metropolitan Areas. Regional Science and Urban Economics 2009, 39, 224–232. [CrossRef]
- Wei, H. Exploring and Practicing the Quantification of Interior Design Colors from an IKEA Design Perspective. Journal of Sensor Networks and Data Communications 2024, 4, 1–12. [CrossRef]
- Kloosterman, R.C.; Musterd, S. The Polycentric Urban Region: Towards a Research Agenda. Urban Studies 2001, 38, 623–633. [CrossRef]
- Meijers, E. Measuring Polycentricity and Its Promises. European Planning Studies 2008, 16, 1313–1323. [CrossRef]
- Burger, M.J.; Meijers, E.J. Form Follows Function? Linking Morphological and Functional Polycentricity. Urban Studies 2012, 49, 1127–1149. [CrossRef]
- Wei, H. DecorPGNet: Functional Area Division and Layout Algorithm Model in Living Rooms of Chinese Apartment-Style Family Homes. Civil Engineering Research Journal 2024, 15, 555902. [CrossRef]
- Wang, Y.; Gu, Y.; Dou, M.; Qiao, M. Rethinking the Identification of Urban Centers from the Perspective of Function Distribution: A Framework Based on Point-of-Interest Data. Sustainability 2020, 12, 1543. [CrossRef]
- Lou, G.; Chen, Q.; He, K.; Zhou, Y.; Shi, Z. Using Nighttime Light Data and POI Big Data to Detect the Urban Centers of Hangzhou. Remote Sensing 2019, 11, 1821. [CrossRef]
- Schneider, A.; Friedl, M.A.; Potere, D. A New Map of Global Urban Extent from MODIS Satellite Data. Environmental Research Letters 2009, 4, 044003. [CrossRef]
- Melchiorri, M.; Florczyk, A.J.; Freire, S.; Schiavina, M.; Pesaresi, M.; Kemper, T. The Global Human Settlement Layer Sets a New Standard for Global Urban Data Reporting with the Urban Centre Database. Frontiers in Environmental Science 2022, 10, 1003862. [CrossRef]
- Dijkstra, L.; Florczyk, A.J.; Freire, S.; Kemper, T.; Melchiorri, M.; Pesaresi, M.; Schiavina, M. Applying the Degree of Urbanisation to the Globe: A New Harmonised Definition Reveals a Different Picture of Global Urbanisation. Journal of Urban Economics 2021, 125, 103312. [CrossRef]
- Cayo, M.R.; Talbot, T.O. Positional Error in Automated Geocoding of Residential Addresses. International Journal of Health Geographics 2003, 2, 10. [CrossRef]
- Zandbergen, P.A. Geocoding Quality and Implications for Spatial Analysis. Geography Compass 2009, 3, 647–680. [CrossRef]
- Wei, H. DynRFM-Hazard: Repurchase Probability Estimation Using Dynamic RFM and Discrete-Interval Hazards. Preprints 2026. Preprint, version 1. [CrossRef]
- Krugman, P. Increasing Returns and Economic Geography. Journal of Political Economy 1991, 99, 483–499. [CrossRef]
- Fujita, M.; Krugman, P.; Venables, A.J. The Spatial Economy: Cities, Regions, and International Trade; MIT Press: Cambridge, MA, 1999. [CrossRef]
- Rosenthal, S.S.; Strange, W.C. Evidence on the Nature and Sources of Agglomeration Economies. In Handbook of Regional and Urban Economics: Cities and Geography; Elsevier, 2004; Vol. 4, pp. 2119–2171. [CrossRef]
- Ellison, G.; Glaeser, E.L. Geographic Concentration in U.S. Manufacturing Industries: A Dartboard Approach. Journal of Political Economy 1997, 105, 889–927. [CrossRef]
- Duranton, G.; Overman, H.G. Testing for Localization Using Micro-Geographic Data. The Review of Economic Studies 2005, 72, 1077–1106. [CrossRef]
- Duranton, G.; Overman, H.G. Exploring the Detailed Location Patterns of U.K. Manufacturing Industries Using Microgeographic Data. Journal of Regional Science 2008, 48, 213–243. [CrossRef]
- Head, K.; Ries, J.; Swenson, D. Agglomeration Benefits and Location Choice: Evidence from Japanese Manufacturing Investments in the United States. Journal of International Economics 1995, 38, 223–247. [CrossRef]
- Guimaraes, P.; Figueiredo, O.; Woodward, D. Agglomeration and the Location of Foreign Direct Investment in Portugal. Journal of Urban Economics 2000, 47, 115–135. [CrossRef]
- Crozet, M.; Mayer, T.; Mucchielli, J.L. How Do Firms Agglomerate? A Study of FDI in France. Regional Science and Urban Economics 2004, 34, 27–54. [CrossRef]
- Basile, R.; Castellani, D.; Zanfei, A. Location Choices of Multinational Firms in Europe: The Role of EU Cohesion Policy. Journal of International Economics 2008, 74, 328–340. [CrossRef]
- Cheng, L.K.; Kwan, Y.K. What Are the Determinants of the Location of Foreign Direct Investment? The Chinese Experience. Journal of International Economics 2000, 51, 379–400. [CrossRef]
- Huang, Y.; Shirai, S. Intrametropolitan FDI Firm Location in Guangzhou, China. The Annals of Regional Science 2000, 34, 535–555. [CrossRef]
- Fan, C.C.; Scott, A.J. Industrial Agglomeration and Development: A Survey of Spatial Economic Issues in East Asia and a Statistical Analysis of Chinese Regions. Economic Geography 2003, 79, 295–319. [CrossRef]
- Lu, J.; Tao, Z. Trends and Determinants of China’s Industrial Agglomeration. Journal of Urban Economics 2009, 65, 167–180. [CrossRef]
- Long, C.; Zhang, X. Patterns of China’s Industrialization: Concentration, Specialization, and Clustering. China Economic Review 2012, 23, 593–612. [CrossRef]
- Jain, A.K. Data Clustering: 50 Years beyond K-Means. Pattern Recognition Letters 2010, 31, 651–666. [CrossRef]
- Tibshirani, R.; Walther, G.; Hastie, T. Estimating the Number of Clusters in a Data Set via the Gap Statistic. Journal of the Royal Statistical Society Series B 2001, 63, 411–423. [CrossRef]
- Comaniciu, D.; Meer, P. Mean Shift: A Robust Approach toward Feature Space Analysis. IEEE Transactions on Pattern Analysis and Machine Intelligence 2002, 24, 603–619. [CrossRef]
- Campello, R.J.G.B.; Moulavi, D.; Sander, J. Density-Based Clustering Based on Hierarchical Density Estimates. In Proceedings of the Advances in Knowledge Discovery and Data Mining. Springer, 2013, pp. 160–172. [CrossRef]
- von Luxburg, U. A Tutorial on Spectral Clustering. Statistics and Computing 2007, 17, 395–416. [CrossRef]
- Wei, H.; Wang, X. Financial Risk Management Early-Warning Model for Chinese Enterprises. Journal of Risk and Financial Management 2024, 17, 255. [CrossRef]
- Wei, H. AI-Driven Employee Turnover Warning for Talent-Retention Decision Support in Finance and Taxation Education: Model Development and Comparison. Preprints 2026. Preprint, version 1. [CrossRef]
- Sahr, K. Central Place Indexing: Hierarchical Linear Indexing Systems for Mixed-Aperture Hexagonal Discrete Global Grid Systems. Cartographica: The International Journal for Geographic Information and Geovisualization 2019, 54, 16–29. [CrossRef]
- Getis, A.; Ord, J.K. The Analysis of Spatial Association by Use of Distance Statistics. Geographical Analysis 1992, 24, 189–206. [CrossRef]
- Ord, J.K.; Getis, A. Local Spatial Autocorrelation Statistics: Distributional Issues and an Application. Geographical Analysis 1995, 27, 286–306. [CrossRef]
- Baddeley, A.; Turner, R.; Moller, J.; Hazelton, M. Residual Analysis for Spatial Point Processes (with Discussion). Journal of the Royal Statistical Society Series B 2005, 67, 617–666. [CrossRef]
- Kulldorff, M. A Spatial Scan Statistic. Communications in Statistics—Theory and Methods 1997, 26, 1481–1496. [CrossRef]
- Efron, B. Bootstrap Methods: Another Look at the Jackknife. The Annals of Statistics 1979, 7, 1–26. [CrossRef]
- Politis, D.N.; Romano, J.P. Large Sample Confidence Regions Based on Subsamples under Minimal Assumptions. The Annals of Statistics 1994, 22, 2031–2050. [CrossRef]
- Lahiri, S.N. Theoretical Comparisons of Block Bootstrap Methods. The Annals of Statistics 1999, 27, 386–404. [CrossRef]
- Wei, H. Negative Indicators and Ordering Stability in Exploratory Factor Analysis: A Sign-Orientation Theory with Reproducible Simulation Evidence. Preprints 2026. Preprint, version 1. [CrossRef]
- Wei, H. From Program Design to Continuous Improvement: An AI-Supported Decision Framework for Vocational Training Operations. Preprints 2026. Concept paper; DOI supplied by the author and not yet registered in Crossref on 8 August 2026. [CrossRef]
- Fotheringham, A.S.; Wong, D.W.S. The Modifiable Areal Unit Problem in Multivariate Statistical Analysis. Environment and Planning A 1991, 23, 1025–1044. [CrossRef]
Figure 1.
Analysis sequence. Entity labels do not participate in center selection.

Figure 2.
Algorithm benchmark, component ablation, observed center counts, and external positional agreement. The GHSL curve covers the 323 units with accepted auxiliary boundaries and is displayed to 25 km so that the central part remains legible; larger deviations are retained in all summaries.
Figure 2.
Algorithm benchmark, component ablation, observed center counts, and external positional agreement. The GHSL curve covers the 323 units with accepted auxiliary boundaries and is displayed to 25 km so that the central part remains legible; larger deviations are retained in all summaries.

Figure 3.
Center reliability and city-equal national primary-center contrasts. Error bars are province-cluster 95% intervals.
Figure 3.
Center reliability and city-equal national primary-center contrasts. Error bars are province-cluster 95% intervals.

Figure 4.
Descriptive primary-center gaps by center structure and registration-count quartile. Points are group medians under the C address specification; they are not causal effects of structure or size.
Figure 4.
Descriptive primary-center gaps by center structure and registration-count quartile. Points are group medians under the C address specification; they are not causal effects of structure or size.

Figure 5.
City distributions of combined registration-agglomeration-zone coverage, area, and raw density. Black segments show interquartile ranges and white circles show medians; area and density use logarithmic scales.
Figure 5.
City distributions of combined registration-agglomeration-zone coverage, area, and raw density. Black segments show interquartile ranges and white circles show medians; area and density use logarithmic scales.

Figure 6.
Registration-center maps for Beijing, Shanghai, Wuhan, and Chengdu. Pale cells show nearest-center catchments, saturated cells show dense registration zones, and numbered circles show MSP-MBKMeans centers.
Figure 6.
Registration-center maps for Beijing, Shanghai, Wuhan, and Chengdu. Pale cells show nearest-center catchments, saturated cells show dense registration zones, and numbered circles show MSP-MBKMeans centers.

Figure 7.
Registration-center maps for Xi’an, Chongqing, Guangzhou, and Shenzhen. Guangzhou is unresolved, so its mapped centers are exploratory.
Figure 7.
Registration-center maps for Xi’an, Chongqing, Guangzhou, and Shenzhen. Guangzhou is unresolved, so its mapped centers are exploratory.

Table 1.
Operational entity classes used in the analysis.
| Label | Formal name | Entities | Share (%) | Historical mappings | Interpretation boundary |
|---|---|---|---|---|---|
| R | Private-dominated other registration entities | 183,235,463 | 99.0283 | Residual | Dominated by private entities; includes some other registered subjects |
| S | State-identified registration entities | 734,375 | 0.3969 | 56 | Historical string match; not a complete verification of contemporary ownership |
| F | Foreign-identified registration entities | 534,945 | 0.2891 | 193 | Historical foreign-related string match |
| H | Hong Kong/Macao/Taiwan-identified entities | 528,609 | 0.2857 | 180 | Historical Hong Kong, Macao, or Taiwan-related string match |
Table 2.
Fixed MSP-MBKMeans rule used for the 8% primary definition.
| Component | Fixed rule |
|---|---|
| Input and peaks | Weighted H3 resolution-9 cells; resolution-8 and -7 peak scores sum occupied cells in a second-order neighborhood; one lowest-ID peak per tied plateau; r8 peak must persist at r7 parent or first-order neighbor. |
| Candidate screen | At most 96 ranked persistent peaks; no relative-height cutoff; separation at least ; valley ratio at most 0.50; tentative and fitted share at least 8%; at most 8 centers. |
| MiniBatch fit | City-local equal-distance coordinates; central 99.5% weighted radial range; batch 4,096; maximum 300 iterations; one peak-based start; seed 20260808; reassignment ratio 0. |
| Spatial repetition | 20 complete reselections of K after independent Gamma multipliers on H3 resolution-7 blocks. Centers are paired by minimum-total WGS84 geodesic distance. |
| Consensus and reliability | Peaks appearing in at least 50% of repetitions enter the consensus fit, with the highest peak always retained. The displacement limit is . Exact stability requires K agreement , every retained peak appearance , 90th-percentile paired displacement within the limit, and 90th-percentile share difference . Core stability drops only the K-agreement condition; usable-with-uncertainty requires appearance , displacement within twice the limit, and share difference . Other cases are unresolved. |
Table 3.
Synthetic validation of MSP-MBKMeans.
| Scenario | Repetitions | Correct K (%) | Median matched error (m) |
|---|---|---|---|
| Compact one center | 20 | 100.0 | 47.5 |
| Elongated one center | 20 | 100.0 | 135.4 |
| Balanced two centers | 20 | 100.0 | 93.6 |
| Imbalanced two centers | 20 | 100.0 | 175.8 |
| Balanced four centers | 20 | 100.0 | 87.3 |
| Subthreshold 4% satellite | 20 | 100.0 | Not applicable |
| All MSP-MBKMeans evaluations | 120 | 100.0 | — |
| Legacy one-standard-error selector | 120 | 16.7 | — |
Table 4.
Center-count recovery in the expanded synthetic benchmark.
| Method | Original scenarios (%) | Stress scenarios (%) | Stress median time (s) |
|---|---|---|---|
| MSP-MBKMeans | 100.0 | 87.8 | 0.044 |
| HDBSCAN | 100.0 | 88.9 | 0.095 |
| OPTICS | 100.0 | 88.9 | 2.124 |
| Gaussian mixture, BIC | 83.3 | 55.6 | 0.271 |
| Mean shift | 70.0 | 33.9 | 0.131 |
| Legacy one-SE MiniBatch K-Means | 16.7 | 27.8 | 0.266 |
Each column contains 20 repetitions per geometry: 120 evaluations in the six original scenarios and 180 evaluations in the nine
stress scenarios. Computation times describe these synthetic samples on the study computer and are not a general hardware
benchmark.
Table 5.
Prespecified synthetic holdout evaluation.
| Method | Evaluations | Correct K (%) | Matched error (m) |
|---|---|---|---|
| MSP-MBKMeans | 200 | 99.0 | 149.5 |
| HDBSCAN | 200 | 100.0 | 121.0 |
| OPTICS | 200 | 100.0 | 87.1 |
| Gaussian mixture, BIC | 200 | 60.0 | 92.4 |
| Mean shift | 200 | 58.5 | 107.0 |
| Legacy one-SE MiniBatch K-Means | 200 | 20.0 | — |
The generating scenarios, expected center counts, random seed, comparison methods, and parameter settings were specified before evaluation. Errors are medians among evaluations with correct K. Parameters remained unchanged throughout the evaluation.
Table 6.
Real-city sensitivity to key MSP-MBKMeans thresholds.
| Setting | Median raw K | Same K as base (%) | S gap (km) | F gap (km) | H gap (km) |
|---|---|---|---|---|---|
| Minimum share 6% | 6 | 38.3 | 0.00 | 3.72 | 4.90 |
| Base: 8%, valley 0.50, distance 0.08 | 5 | 100.0 | 3.72 | 4.61 | |
| Minimum share 10% | 4 | 44.6 | 3.58 | 4.49 | |
| Valley ratio 0.40 | 5 | 99.7 | 3.72 | 4.61 | |
| Valley ratio 0.60 | 5 | 99.7 | 3.72 | 4.61 | |
| Distance fraction 0.06 | 5 | 99.7 | 3.72 | 4.61 | |
| Distance fraction 0.10 | 5 | 99.7 | 3.72 | 4.58 |
Raw K is the deterministic full-sample candidate count; the final spatial-resampling consensus has median four. Gaps use fixed
eligible sets of 307/263/220 cities for S/F/H, so changes isolate center-threshold sensitivity. Positive values indicate greater
centrality than R.
Table 7.
Complete propagation of the 6/8/10% minimum-center-share definition.
| Share | Median consensus K | Unresolved cities | S central/peripheral | F central/peripheral | H central/peripheral | S stable | F stable | H stable |
|---|---|---|---|---|---|---|---|---|
| 6% | 5 | 15 | 24/25 | 87/8 | 83/3 | Compared across all three rows | ||
| 8% | 4 | 25 | 24/25 | 79/8 | 82/3 | 310 | 313 | 316 |
| 10% | 4 | 24 | 24/23 | 80/7 | 80/2 | Compared across all three rows | ||
Central/peripheral counts are class-wise Benjamini–Hochberg-corrected city conclusions. Stable is the number, out of 345, whose complete corrected label is identical at 6%, 8%, and 10%. The 8% row remains the primary specification; alternatives are not searched for favorable results.
Table 8.
Observed center structure, reliability, and address sensitivity.
| Diagnostic | Cities | Interpretation |
|---|---|---|
| Exactly stable center count | 69 | Count and center configuration are repeatable |
| Stable core; optional peripheral centers | 88 | Core is repeatable, minor centers may appear or disappear |
| Usable with center-count uncertainty | 163 | Main pattern can be described with an uncertainty warning |
| Unresolved | 25 | No formal center–periphery conclusion |
| Identical K under A/B/C | 57 | Center count is insensitive to address weighting |
| Full center structure matched under A/B/C | 67 | Main configuration agrees while allowing at most one optional center |
| Primary center stable under A/B/C | 209 | Main reference center survives all address weights |
Table 9.
Sensitivity to the number of spatial repetitions in 12 structured-sample cities.
| Comparison with 100 repetitions | Same K | Same reliability level | Median primary shift (km) | Maximum primary shift (km) |
|---|---|---|---|---|
| 20 repetitions | 10/12 | 10/12 | 0.00 | 3.54 |
| 50 repetitions | 10/12 | 11/12 | 0.00 | 2.78 |
The sample takes the minimum, median, and maximum resolution-9 cell count within each primary reliability level. It spans size
and reliability for diagnosis and is not a random national sample. The primary analysis retains its prespecified 20 repetitions.
Table 10.
Sensitivity of the independent GHSL settlement reference.
| Setting | GHSL nodes | Same K (%) | Primary median (km) | Primary within 10 km (%) | All-node within 10 km (%) |
|---|---|---|---|---|---|
| Base: tier 10%, component 5%, cap 12 | 1,337 | 100.0 | 3.32 | 82.7 | 63.4 |
| Tier population 5% | 1,299 | 97.2 | 3.32 | 82.7 | 64.2 |
| Tier population 15% | 1,352 | 95.4 | 3.29 | 82.4 | 63.4 |
| Component population 3% | 1,690 | 53.9 | 3.32 | 82.7 | 64.5 |
| Component population 8% | 918 | 44.3 | 3.32 | 82.7 | 70.6 |
| Maximum eight nodes | 1,326 | 98.1 | 3.32 | 82.7 | 63.1 |
Each row rebuilds the GHSL reference without registration data. Same K compares GHSL node count with the base reference,
not with MSP-MBKMeans. All-node percentages use minimum-cost matched pairs and must be read with the changing node
count.
Table 11.
City-level positional conclusions for the four classes.
| Class | More central | More peripheral | No clear | Sensitive | Not estimable |
|---|---|---|---|---|---|
| S | 24 | 25 | 256 | 2 | 38 |
| F | 79 | 8 | 175 | 1 | 82 |
| H | 82 | 3 | 131 | 4 | 125 |
Notes: S, F, and H denote state-, foreign-, and Hong Kong/Macao/Taiwan-identified entities. Each row totals 345 observed cities. Directional and address-sensitive labels pass Benjamini–Hochberg control at within their class family, use at least eight delete-block estimates, and follow the A/B/C rule. Supplementary Table S1 gives every corrected city conclusion.
Table 12.
City-equal national center–periphery contrasts relative to R.
| Class | Weight | Cities | Primary gap, km (95% CI) | Probability index | Nearest gap, km |
|---|---|---|---|---|---|
| S | A | 307 | 0.04 [, 0.84] | 0.010 | |
| S | B | 307 | [, ] | ||
| S | C | 307 | [, 0.71] | 0.012 | |
| F | A | 263 | 3.25 [2.25, 4.45] | 0.079 | 1.02 |
| F | B | 263 | 2.47 [1.70, 3.02] | 0.039 | 0.35 |
| F | C | 263 | 3.74 [2.63, 4.79] | 0.084 | 1.06 |
| H | A | 220 | 4.39 [2.27, 6.43] | 0.115 | 1.34 |
| H | B | 220 | 3.28 [2.00, 5.00] | 0.077 | 0.56 |
| H | C | 220 | 4.57 [2.94, 6.96] | 0.122 | 1.37 |
Notes: Gap ; positive values indicate that class m is closer to a center. Medians are city-equal. Confidence intervals use 5,000 province-cluster bootstrap repetitions. C is the primary specification.
Table 13.
Sensitivity to excluding the compared class from center construction.
| Removed | Same K | Median shift (km) | P90 shift (km) | Same FDR label |
|---|---|---|---|---|
| S | 314 (91.0%) | 0.16 | 1.52 | 342 (99.1%) |
| F | 332 (96.2%) | 0.10 | 0.57 | 343 (99.4%) |
| H | 332 (96.2%) | 0.08 | 0.50 | 343 (99.4%) |
Full and leave-class-out centers use deterministic full-sample fits at the fixed 8% definition so that this diagnostic isolates reference dependence. Primary-model resampling uncertainty is reported separately.
Table 14.
Sensitivity to the composition of city-level administrative units under the primary C specification.
Table 14.
Sensitivity to the composition of city-level administrative units under the primary C specification.
| Inventory | Observed units | S gap, km (95% CI) | F gap, km (95% CI) | H gap, km (95% CI) |
|---|---|---|---|---|
| All source units | 345 | [, 0.71] | 3.74 [2.63, 4.79] | 4.57 [2.94, 6.96] |
| Prefecture-level units | 316 | [, 0.70] | 3.74 [2.63, 4.96] | 4.66 [2.92, 7.00] |
| Standard prefecture-level cities | 285 | [, 0.80] | 3.73 [2.63, 4.90] | 4.66 [2.92, 6.96] |
The class-specific eligible counts are 307/263/220, 292/253/211, and 263/240/207 for S/F/H, respectively. All estimates are
city-equal; intervals use province-cluster resampling.
Table 15.
Evidence summary for the eight illustrative cities under the primary C specification.
| City | Entities | K | A/B K | Primary share | Reliability | GHSL km | S gap | F gap | H gap | Zone share | Density |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Beijing | 4,815,679 | 3 | 4/2 | 0.700 | Usable–uncertain | 1.39 | 5.37 | 9.23 | 7.53 | 0.484 | 2,581 |
| Shanghai | 5,210,935 | 4 | 3/4 | 0.513 | Usable–uncertain | 7.36 | 6.90 | 9.42 | 7.26 | 0.478 | 2,402 |
| Wuhan | 2,262,224 | 2 | 2/3 | 0.656 | Exactly stable | 7.13 | 3.77 | 3.02 | 4.62 | 0.455 | 6,410 |
| Chengdu | 3,806,872 | 3 | 2/3 | 0.764 | Usable–uncertain | 5.81 | 7.70 | 6.61 | 0.489 | 2,707 | |
| Xi’an | 2,614,817 | 2 | 2/2 | 0.899 | Exactly stable | 2.03 | 4.21 | 3.75 | 0.496 | 4,152 | |
| Chongqing | 3,209,412 | 2 | 3/4 | 0.718 | Usable–uncertain | 6.83 | 15.87 | 63.41 | 60.67 | 0.508 | 265 |
| Guangzhou | 3,842,124 | 3 | 2/3 | 0.706 | Unresolved | 4.03 | 5.35 | 5.62 | 0.487 | 3,106 | |
| Shenzhen | 5,190,364 | 4 | 5/5 | 0.474 | Usable–uncertain | 7.76 | 0.95 | 2.11 | 1.13 | 0.501 | 10,576 |
A/B K gives center counts under raw-entity and address-equal weights; the main K uses C. Gaps are E(DR) − E(Dm) in km.
Density is entities/km2. Guangzhou’s gaps are exploratory because its center system is unresolved.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.