Submitted:
25 August 2026
Posted:
27 August 2026
You are already at the latest version
Abstract
The reproducibility of empirical machine learning is often compromised by heterogeneous preprocessing pipelines, ad-hoc model selection, and insufficient statistical rigor. This paper provides the complete methodological foundations of the adaptive framework \(\mathcal{F} := (\Phi, \Pi, \Xi)\), designed to overcome these issues through a unified representation space, a standardized operational protocol, and a rigorous evaluation metric. We formalize two core components: the transformation operator \(\Phi\), which maps inputs from arbitrary domains into a shared space \(\mathcal{Z}\) with a provable bound on inter-domain distance (Proposition~1), and the operational protocol \(\Pi\), which guarantees invariance of parameter estimator variance across domains (Proposition~2). We provide algorithmic instantiations, complete theoretical proofs, and complexity analysis. Empirical validation on ten diverse real-world datasets (regression and classification) confirms that \(\Phi\) reduces average inter-domain MMD distance by over 53.6\% without degrading predictive signal, and that \(\Pi\) yields parameter estimates with a coefficient of variation whose standard deviation across domains is only 0.007, thereby meeting the invariance criterion. The framework reveals statistically significant improvements on five of the ten datasets—most notably on Wine (\(\Delta\eta = +0.933\))—while transparently flagging non-significant or negative results (e.g., AutoMPG). All code, pre-registered protocols, and data are released for full reproducibility, establishing \(\mathcal{F}\) as a robust methodological benchmark for honest model comparison.
Keywords:
reproducible research
; model comparison
; domain generalization
; machine learning
; cross-validation
; representation learning
; statistical significance
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.