8. Experimental Results
8.1. Diagnostic Trajectories Across the Hard-Regime Sweep
The primary empirical result of this paper is the behavior of the learned metric models across the Cefalú hard-regime sweep. Rather than evaluating model families only at isolated hard points, we track how the principal geometry-sensitive diagnostics evolve across the selected benchmark cases. This makes it possible to study not only relative model quality at fixed parameter values, but also the structure of learned geometric degradation under controlled hardening of the underlying quartic geometry. The trajectories are not uniformly monotone in ; in particular, the local-model negative-frequency channel is non-monotone across the sweep. This already indicates that hard-regime fidelity is structured rather than captured by a single scalar hardness coordinate.
Across the sweep, the principal quantities of interest are negative-eigenvalue frequency and projective-invariance drift, together with supporting lower-tail stability summaries such as the minimum-eigenvalue mean and spectral_tail_mean. The resulting trajectories show how different fidelity channels evolve, whether lower-tail instability appears before broader breakdown, and whether invariance-sensitive and positivity-sensitive diagnostics remain aligned across model families or separate as the regime changes.
These trajectories also serve an interpretive role beyond raw comparison. If distinct diagnostics deteriorate at different rates, then learned geometric breakdown is structured rather than uniform. In that case, the sweep becomes a means of identifying which aspects of learned fidelity fail first and whether those failure patterns differ systematically between local-input and globally invariant models. For this reason, the trajectory analysis is not merely a plotting convenience; it is one of the main devices through which this paper studies regime-dependent geometric fidelity.
Figure 9 summarizes these diagnostic trajectories across the Cefalú sweep. The figure is intended to show both the absolute behavior of the primary diagnostics and the structured separation of the competing model families across different fidelity channels.
8.2. Global Versus Local Degradation Profiles
A central question is how the globally invariant and local-input models separate across the Cefalú sweep. The answer depends on which aspect of learned geometric fidelity is being measured. The strongest and cleanest sweep-wide signal is the projective-invariance advantage of the globally invariant model: across the full hard-regime sweep, the global model consistently achieves substantially lower projective-invariance drift than the local baseline. This confirms that the invariant architecture robustly preserves the global projective structure that motivates it, not only at isolated hard points but across the broader Cefalú family.
At the same time, the positivity-oriented and lower-tail stability diagnostics tell a different story. Across the same sweep, the local baseline consistently outperforms the globally invariant model on negative-eigenvalue frequency, minimum-eigenvalue mean, and especially spectral_tail_mean. The regime-dependent comparison therefore does not support a single scalar ordering of the model families. Instead, it reveals a structured diagnostic tradeoff: the globally invariant architecture is stronger on invariance-sensitive fidelity, while the local baseline is stronger on positivity-tail stability.
This makes the degradation-profile analysis more scientifically informative than a simple extension of the earlier work benchmark. The main question is no longer whether the global model “wins” everywhere, but how different notions of learned geometric fidelity separate across the hard regime. In this sense, the Cefalú sweep does more than rank models; it shows that learned geometric breakdown is multi-axis, with different diagnostics selecting different strengths and different hardness patterns across the family.
8.3. Effect of Geometry-Aware Training
A second question is whether geometry-aware regularization improves the behavior of the globally invariant model in the hardest regimes. The architecture result of earlier work established that globally invariant structure provides a strong starting point for invariance-sensitive fidelity. The objective-ablation study asks whether modest but explicit geometric constraints in the training objective can further improve particular hard-regime fidelity channels.
The ablation results show that the effect of regularization is real but selective. Projective-consistency regularization produces the clearest gains on projective-invariance drift in the key overlap hard cases, while negativity-oriented regularization improves lower-tail stability as measured by spectral_tail_mean. The combined objective improves some metrics simultaneously in selected cases, but does not produce a universal monotone improvement across the full ablation set. In particular, the current ablation study does not materially change negative-eigenvalue frequency in the frozen outputs, and the combined objective can expose tradeoffs rather than eliminate them.
Thus, geometry-aware regularization does not simply “improve the model” in one scalar sense. Different objective terms act on different geometric channels. Projective-oriented regularization improves invariance behavior, while negativity-oriented regularization improves lower-tail stability. The ablation study therefore strengthens the central theme: learned geometric fidelity in hard quartic regimes is structured, and targeted objective design can reshape that structure without collapsing it into a single dominant notion of quality.
8.4. Degeneration-Sensitive Class Profiles Across the Cefalú Sweep
Beyond the learned-fidelity diagnostics emphasized in the main sweep analysis, we attach a first degeneration-sensitive geometric profile to the Cefalú hard-regime family. The purpose of this extension is not to claim a full singularity calculation, but to begin organizing the benchmark in a way that is compatible with later degeneration- and node-aware studies. In particular, we seek a computable family-level proxy for geometric fragility together with global class-facing summaries that can be tracked across the sweep.
To this end, we introduce a degeneration-sensitive fragility score derived from the sampled quartic geometry. In the present benchmark, this score is intended only as a computable proxy for near-degenerate behavior rather than as a singularity invariant in the strict algebro-geometric sense. Concretely, for each case in the sweep we evaluate a geometry-side fragility statistic , together with low-quantile and thresholded summaries that record how strongly fragile geometric sectors are represented at that value of .
Alongside this degeneration-sensitive profile, we track global characteristic-class-facing summaries across the same family. At minimum, these include the Euler proxy together with determinant- and log-determinant-facing grouped summaries. The purpose of these global quantities is not to certify topology in the present paper, but to provide a first computational shadow of how class-like geometric information varies across a controlled hardening of the family.
The resulting sweep-level analysis serves two purposes. First, it gives the hard-regime family a geometry-side stress coordinate rather than treating only as a benchmark label. Second, it begins to connect learned metric behavior to global class-sensitive geometry, which is important for later degeneration, singularity, and finite-node conifold studies. In this sense, the present extension should be read as a first degeneration-aware layer rather than as a final characteristic-class computation.
8.5. Localized Class and Instability Decomposition Near Fragile Sectors
To move beyond sweep-level averages, we supplement the hard-regime analysis with a support decomposition based on the geometry-side fragility proxy introduced above. Using the GeoCY fragility export, each sampled point is assigned to either a fragile sector or a complementary regular sector, and the resulting partition is joined directly to the pointwise learned-metric diagnostics through the common key. This yields a pointwise support decomposition over the full hard-regime sweep without requiring additional clustering assumptions.
The first result of this localized analysis is that instability does not distribute uniformly across the geometry. Fragile sectors carry higher negative-eigenvalue frequency in four of the five sweep cases for both the local and globally invariant baselines, and they carry worse lower-tail behavior in all five local cases and in four of the five global cases. Thus, the lower-tail degradation identified earlier in the paper is not only a sweep-level effect: it localizes preferentially in degeneration-sensitive geometry.
Alongside these instability summaries, we compute modest class-shadow-style grouped quantities from the same pointwise exports, including local moments and weighted means of euler_density_ proxy, determinant_g, logdet_g, and logdet_residual_proxy. These are not singular characteristic classes in the strict sense, but they provide a first computational shadow of how class-facing geometric content is distributed across fragile and regular sectors. In the present results, the class-shadow concentration is mixed rather than uniformly fragile-enhanced. This is itself informative: unlike lower-tail instability, class-facing concentration does not yet collapse to a single fragile-sector narrative.
The support-conditioned layer is therefore useful in two ways. First, it shows that degeneration-sensitive localization is already visible in the learned metric data. Second, it distinguishes channels that localize strongly in fragile geometry from channels whose support is more distributed.
Figure 1 and
Figure 2 summarize the strongest support-conditioned signals for instability and lower-tail behavior,
Figure 4 records the corresponding equation-facing residual comparison, and
Figure 3 shows the class-shadow-style comparison.
Figure 1.
Fragile-versus-regular comparison for negative-eigenvalue frequency across the Cefalú hard-regime sweep. Fragile sectors carry higher negative-frequency in most sweep cases for both local and globally invariant models, indicating that instability localizes preferentially in degeneration-sensitive geometry.
Figure 1.
Fragile-versus-regular comparison for negative-eigenvalue frequency across the Cefalú hard-regime sweep. Fragile sectors carry higher negative-frequency in most sweep cases for both local and globally invariant models, indicating that instability localizes preferentially in degeneration-sensitive geometry.
Figure 2.
Fragile-versus-regular comparison for lower-tail stability. The lower metric-eigenvalue tail degrades more strongly in fragile sectors, especially for the local baseline, showing that lower-tail failure is not uniformly distributed across the geometry.
Figure 2.
Fragile-versus-regular comparison for lower-tail stability. The lower metric-eigenvalue tail degrades more strongly in fragile sectors, especially for the local baseline, showing that lower-tail failure is not uniformly distributed across the geometry.
Figure 3.
Fragile-versus-regular comparison for the euler_density_proxy. Unlike lower-tail instability, the class-shadow concentration is mixed rather than uniformly fragile-enhanced, suggesting that class-facing localization is more structured and less one-dimensional than instability localization alone.
Figure 3.
Fragile-versus-regular comparison for the euler_density_proxy. Unlike lower-tail instability, the class-shadow concentration is mixed rather than uniformly fragile-enhanced, suggesting that class-facing localization is more structured and less one-dimensional than instability localization alone.
Figure 4.
Fragile-versus-regular comparison for the equation-facing residual proxy logdet_residual_proxy. The fragile-sector separation is stronger and more consistent for the local baseline than for the globally invariant model, indicating channel-dependent localization of equation-facing error.
Figure 4.
Fragile-versus-regular comparison for the equation-facing residual proxy logdet_residual_proxy. The fragile-sector separation is stronger and more consistent for the local baseline than for the globally invariant model, indicating channel-dependent localization of equation-facing error.
8.6. Equation-Facing Qualification of Fragile-Sector Localization
The support decomposition becomes more meaningful if the same fragile sectors are also visible in an equation-facing quantity rather than only in benchmark-side instability summaries. For this reason, we augment the pointwise export layer with the residual proxy
which provides a lightweight volume-form / Monge–Ampère-type measure of deviation at the sampled-point level. The role of this proxy is not to certify Ricci-flatness, but to test whether degeneration-sensitive localization also appears in a quantity closer to the defining geometric equation.
The resulting comparison shows that residual concentration is present, but not uniformly across model classes. In the present sweep, the fragile-versus-regular separation in logdet_residual_proxy is stronger and more consistent for the local baseline than for the globally invariant model. This is scientifically important because it shows that the localization pattern is itself channel-sensitive: lower-tail instability localizes clearly for both model families, whereas equation-facing residual concentration is a sharper fragile-sector separator for the local model than for the global one.
The same structured picture appears in the hardest-case ablation subset at . There, negativity-oriented and combined objective variants slightly soften fragile-sector lower-tail suppression, but they do not clearly remove fragile-sector negative-frequency concentration. This means that the objective interventions are not simply erasing degeneration-sensitive localization. Rather, they modify specific channels while leaving the broader fragile-sector structure visible.
Taken together, these results justify a modest but meaningful conclusion. The present benchmark supports a first support-conditioned shadow in which instability, lower-tail behavior, and equation-facing residuals can all be resolved relative to fragile geometry. The alignment is not uniform across channels, and the class-shadow layer remains mixed, but the extension already shows that degeneration-sensitive sectors are computationally visible and scientifically interpretable within the current framework.
Figure 4 summarizes the fragile-versus-regular residual comparison, and
Table 5 records the compact support-conditioned summary used in the text.
8.7. Low-Dimensional Collective Response Structure
The degeneration-aware support decomposition makes it possible to ask a stronger question than whether fragile sectors merely exist. Namely: do the resulting support-conditioned response objects behave as if they were freely independent, or do they instead organize into a lower-dimensional collective structure? In the present study, we examine this question using the grouped fragile-versus-regular response objects constructed from instability, class-shadow-style, and residual-facing summaries. The goal is not to claim a theorem-level gluing law, but to test whether the support-conditioned response is already constrained in a numerically collective way.
To do this, we assemble response vectors from the grouped response channels
We then analyze their covariance, correlation, and singular-value structure. Across this 13-channel response space, the resulting data are clearly low-dimensional: both fragile-sector and regular-sector response objects occupy a substantially compressed subspace relative to the ambient channel count. At the same time, the first-pass collective signal is more subtle than a simple “fragile sectors are more collective” slogan. In the present results, the collective compression is stronger for the regular-sector response objects than for the fragile ones. Concretely, the fragile effective-rank ratio is approximately
, whereas the regular effective-rank ratio is approximately
; similarly, the first principal component explains about
of the fragile-sector variance but about
of the regular-sector variance. Thus, the evidence supports a real low-dimensional collective response structure, but it does not support the stronger claim that fragile sectors are uniquely or maximally collective in this first pass.
The dominant collective modes are also scientifically informative. The leading fragile-sector mode is not a pure instability axis, nor a pure residual axis, nor a pure class-shadow axis. Rather, it is a mixed mode whose strongest loadings include euler_density_proxy_mean, logdet_g_mean, euler_density_proxy_weighted_mean, min_eigenvalue_q10, and logdet_residual_mean. This indicates that the support-conditioned response is organized collectively across several channels at once. In that limited but meaningful sense, the grouped response is not well described as a free sum of independent channel contributions.
The ablation subset shows a similarly cautious pattern. There too, the response objects are strongly compressed, and the fragile-only ablation structure is even more low-dimensional, with an effective-rank ratio of approximately and the first principal component explaining about of the variance. However, the variant-to-variant shifts are small, so the current objective interventions should not be interpreted as materially reorganizing the collective-mode structure. The right conclusion is therefore modest: the support-conditioned response exhibits clear low-dimensional collective structure, but the strongest collective compression in this first pass is not uniquely fragile-sector-specific. These low-dimensionality summaries are extracted from a relatively small support-conditioned response sample and should therefore be interpreted as exploratory collective diagnostics rather than as stable asymptotic statistics.
Figure 5,
Figure 6 and
Figure 7 summarize the principal low-dimensionality and coupling signals, while
Table 6 records the compact collective-response statistics used in the text.
Table 6.
Compact collective-response summary for the support-conditioned response objects. Lower effective-rank ratios and higher variance capture in the leading principal modes indicate stronger collective compression relative to the full 13-channel response space.
Table 6.
Compact collective-response summary for the support-conditioned response objects. Lower effective-rank ratios and higher variance capture in the leading principal modes indicate stronger collective compression relative to the full 13-channel response space.
| Subset |
Effective-rank ratio |
PC1 variance explained |
Interpretation |
| Fragile sectors |
0.249 |
0.630 |
low-dimensional, but not maximally compressed |
| Regular sectors |
0.171 |
0.793 |
stronger first-pass collective compression |
|
ablation fragile subset |
0.154 |
0.795 |
highly compressed, small variant-to-variant reorganization |
Figure 5.
Singular-value decay for the support-conditioned response objects. The rapid decay relative to the full 13-channel response space indicates that both fragile-sector and regular-sector summaries occupy a substantially lower-dimensional collective structure.
Figure 5.
Singular-value decay for the support-conditioned response objects. The rapid decay relative to the full 13-channel response space indicates that both fragile-sector and regular-sector summaries occupy a substantially lower-dimensional collective structure.
Figure 6.
Explained-variance profile for the leading principal components of the support-conditioned response objects. The first principal mode captures a large fraction of the variance in both fragile and regular sectors, confirming that the grouped response is not freely distributed across all channels.
Figure 6.
Explained-variance profile for the leading principal components of the support-conditioned response objects. The first principal mode captures a large fraction of the variance in both fragile and regular sectors, confirming that the grouped response is not freely distributed across all channels.
Figure 7.
Correlation structure of the support-conditioned response channels. The mixed loading pattern across instability, class-shadow, and residual-facing quantities shows that the leading collective organization is not confined to a single diagnostic family.
Figure 7.
Correlation structure of the support-conditioned response channels. The mixed loading pattern across instability, class-shadow, and residual-facing quantities shows that the leading collective organization is not confined to a single diagnostic family.
Figure 8.
Collective-mode structure in the ablation subset. The response remains strongly compressed, but the objective variants do not materially reorganize the dominant collective modes in this first pass.
Figure 8.
Collective-mode structure in the ablation subset. The response remains strongly compressed, but the objective variants do not materially reorganize the dominant collective modes in this first pass.
8.8. Hardest-Case Analysis
The hardest-case analysis provides the most concentrated view of learned geometric failure in the benchmark. The regime sweep shows that the notion of hardness is not diagnostic-independent: cases that are hardest for projective-invariance drift are not identical to the cases that are hardest for lower-tail stability or training loss. In the present results, the drift-sensitive notion of hardness points toward the regime, while positivity-tail and loss-sensitive criteria identify different parts of the sweep as hardest for the globally invariant model. This split is itself a significant result, because it shows that the hard regime is not one-dimensional.
The purpose of the hardest-case analysis is therefore not simply to isolate one difficult benchmark row, but to study how different notions of learned failure compete in the most demanding settings. In particular, this analysis asks whether the globally invariant model preserves its invariance advantage in the drift-hardest regime, whether the lower-tail weakness seen in the sweep remains visible under a more detailed comparison, and whether geometry-aware regularization can improve one fidelity channel without destabilizing another. In this setting, one objective variant may improve the lower metric-eigenvalue tail while only weakly changing drift, whereas another may improve drift while leaving the hardest positivity failures largely unresolved.
For that reason, the hardest-case analysis serves as the most detailed lens in the paper. It is where the benchmark moves beyond aggregate comparison and begins to study the structure of failure itself: which metrics deteriorate most severely, which interventions improve them, and which tradeoffs emerge once the regime enters its most demanding range.
Figure 12 summarizes the detailed comparison in the hardest overlap case used by the objective-ablation study, while
Table 8 records the corresponding objective-level summary for the globally invariant model.
Table 7.
Core results across the full hard-regime sweep.
Table 7.
Core results across the full hard-regime sweep.
| Cefalú Case () |
Model |
Negative-eigenvalue frequency |
Projective-invariance drift |
spectral_ tail_mean |
| 0.50 |
Local |
0.015000 |
3.733362e-08 |
0.026163 |
| 0.50 |
Global |
0.046667 |
1.564001e-08 |
0.007294 |
| 0.75 |
Local |
0.025000 |
3.886719e-08 |
0.023318 |
| 0.75 |
Global |
0.045000 |
1.728535e-08 |
0.007518 |
| 0.90 |
Local |
0.011667 |
3.917764e-08 |
0.038782 |
| 0.90 |
Global |
0.046667 |
1.593338e-08 |
0.013410 |
| 1.00 |
Local |
0.006667 |
3.617257e-08 |
0.048261 |
| 1.00 |
Global |
0.046667 |
1.783793e-08 |
0.012506 |
| 1.10 |
Local |
0.005000 |
3.794829e-08 |
0.058322 |
| 1.10 |
Global |
0.038333 |
1.740176e-08 |
0.017665 |
Table 8.
Objective-ablation summary for the globally invariant model at the hardest overlap ablation case, Cefalú .
Table 8.
Objective-ablation summary for the globally invariant model at the hardest overlap ablation case, Cefalú .
| Variant |
Projective-invariance drift |
spectral_tail_mean |
Interpretation |
| Baseline |
1.643319e-08 |
0.009408 |
reference objective |
| Negativity-regularized |
1.654960e-08 |
0.012471 |
strongest lower-tail stability among variants (tied) |
| Projective-regularized |
1.575332e-08 |
0.009408 |
strongest drift among variants |
| Combined |
1.625779e-08 |
0.012471 |
mixed tradeoff; near-best drift with strongest lower-tail stability (tied) |
8.9. Diagnostic Disagreement and Fidelity Breakdown
A further result is that learned geometric breakdown does not occur uniformly across diagnostics. Instead, different diagnostics fail at different rates and identify different regions of the sweep as hardest. This means that learned metric failure is not adequately described by a single score or threshold. Rather, it has internal structure: some aspects of fidelity remain comparatively stable while others deteriorate, and those deterioration patterns differ across model families and objective variants.
This is especially important because the benchmark combines several distinct notions of geometric quality. Negative-eigenvalue frequency, minimum-eigenvalue mean, and spectral_tail_mean probe different aspects of lower-tail stability, while projective-invariance drift probes consistency with the global projective structure of the quartic geometry. Determinant-based summaries, the Euler proxy, and training loss contribute additional context. The sweep results show that these diagnostics do not fail together and do not produce a single consistent hardness ordering. In particular, the globally invariant model is sweep-wide superior on projective-invariance drift, while the local baseline is sweep-wide superior on positivity-tail stability.
The diagnostic-disagreement analysis is therefore one of the most scientifically informative aspects of the study. It shows that learned geometric fidelity in hard quartic regimes is multi-axis rather than one-dimensional. The benchmark does not merely compare models; it reveals a hierarchy and tradeoff structure among different notions of geometric fidelity. This structured view of failure is stronger and more informative than a simple benchmark ranking, because it identifies which aspects of learned geometry are improved by global invariant structure, which remain weaker, and which can be partially reshaped by geometry-aware objective design.
Figure 9.
Diagnostic trajectories across the Cefalú hard-regime sweep. The figure emphasizes that different diagnostics evolve differently across the sweep: projective-invariance drift consistently favors the globally invariant model, while positivity-oriented and lower-tail diagnostics favor the local baseline.
Figure 9.
Diagnostic trajectories across the Cefalú hard-regime sweep. The figure emphasizes that different diagnostics evolve differently across the sweep: projective-invariance drift consistently favors the globally invariant model, while positivity-oriented and lower-tail diagnostics favor the local baseline.
Figure 10.
Comparison of degradation profiles for local and globally invariant models. The main scientific signal is not a uniform winner, but a structured separation between invariance-sensitive and positivity-tail-sensitive notions of learned geometric fidelity.
Figure 10.
Comparison of degradation profiles for local and globally invariant models. The main scientific signal is not a uniform winner, but a structured separation between invariance-sensitive and positivity-tail-sensitive notions of learned geometric fidelity.
Figure 11.
Effect of geometry-aware objective variants on hardest-case fidelity. Projective-oriented regularization improves invariance drift on the key overlap hard cases, while negativity-oriented regularization improves lower-tail stability, showing that different objective terms act on different fidelity channels.
Figure 11.
Effect of geometry-aware objective variants on hardest-case fidelity. Projective-oriented regularization improves invariance drift on the key overlap hard cases, while negativity-oriented regularization improves lower-tail stability, showing that different objective terms act on different fidelity channels.
Figure 12.
Detailed comparison at the hardest overlap Cefalú regime used in the objective-ablation study. The figure is intended to highlight how invariance-sensitive and lower-tail-sensitive diagnostics respond differently to model choice and geometry-aware regularization.
Figure 12.
Detailed comparison at the hardest overlap Cefalú regime used in the objective-ablation study. The figure is intended to highlight how invariance-sensitive and lower-tail-sensitive diagnostics respond differently to model choice and geometry-aware regularization.