Submitted:
02 October 2026
Posted:
06 October 2026
You are already at the latest version
Abstract
A probe R², a judge-human agreement rate or a cross-validated correlation is often read as evidence of fine-grained structure. When the target is grouped, a predictor that knows only the group can attain such a score. Building on conditional probing (Hewitt et al., 2021), we compare selected published headlines H with training-only group baselines B in the paper's own metric and, for nested comparisons, quantify the conditional predictive gain (CPG): the out-of-sample reduction in squared error from adding the representation to the group predictor, split exactly into group-mean and within-group parts. The aim is to quantify what a published score allows one to infer about the resolution of prediction. In our worked example, the space and time probes of Gurnee & Tegmark, country, state or borough means match or exceed the published space scores, and resolution oracles built from the target exceed the time scores. Beyond country and state means, residual readouts from four 6-7B models remove 20-34% of the remaining coordinate error (an increase in total R² of .008-.013); for one of them the CPG point estimates met thresholds frozen before evaluation, although with whole countries resampled the World interval reaches zero. On the eligible USA subset with counties observed in training, the saved readouts' net gain over state means comes from between-county differences, and a cross-fitted linear recalibration within counties yields gains near zero. Below the century, decade or year the gains are small. The same diagnostic applies to a language-model probe of political ideology, a Twitter prediction of county heart-disease mortality and LLM judges: on retained non-tie votes, a model-identity baseline calibrated from other human votes in the same collection attains much of the judges' agreement with humans, while the judges reduce its disagreement rate by 12-28%. We recommend reporting both numbers, at more than one grouping level, whenever a salient grouping could plausibly attain much of the reported performance.

Keywords:
probing
; evaluation
; group baseline
; conditional predictive gain
; language-model representations
; LLM-as-a-judge
; static word embeddings
; reproducibility
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.