Preprint
Article

This version is not peer-reviewed.

CoT-Core: Accelerating LLM Evaluation via CoT-Aware Coreset Selection

Submitted:

17 July 2026

Posted:

20 July 2026

You are already at the latest version

Abstract
Evaluating Large Language Models (LLMs) incurs prohibitive computational overhead during continuous development processes. While coreset selection accelerates evaluation, existing methods either suffer from a severe "cold start'' bottleneck requiring massive historical logs (e.g., Item Response Theory) or exhibit a surface lexical bias that misses the underlying reasoning manifold of tasks. We propose CoT-Core, a novel training-free core question selection framework. Recognizing that lexically disparate questions can share equivalent underlying logic, CoT-Core prompts LLMs to unroll zero-shot Chain-of-Thought (CoT) reasoning trajectories. Projecting these paths into a latent space effectively clusters questions by intrinsic logical equivalence rather than superficial text similarity. Extensive experiments on GSM8K, MMLU, MMLU-Pro, and GPQA demonstrate that CoT-Core drastically reduces evaluation costs while maintaining high-fidelity score estimation, and delineate the boundary conditions of reasoning-aware pruning, revealing that its efficacy is intrinsically gated by task complexity.
Keywords: 
;  ;  ;  ;  

1. Introduction

The rapid advancement of Large Language Models (LLMs) [2,20] has necessitated the development of comprehensive evaluation frameworks and benchmarks, such as MMLU [8], GSM8K [6], and HELM [10]. However, evaluating modern LLMs across tens of thousands of test cases introduces prohibitive computational and financial overhead [17]. Furthermore, evaluation is rarely a one-time event; during the continuous pre-training, alignment, and fine-tuning phases, frequent and repeated evaluations across numerous model checkpoints are strictly required to monitor progress and steer model development [4,7], severely exacerbating the consumption of computational resources.
Recent literature highlights that this exhaustive evaluation paradigm is highly inefficient due to inherent dataset redundancy [21,26]. This creates an urgent need for coresets—minimal subsets of data that accurately reconstruct full-dataset evaluation scores. While various sampling heuristics exist, the most prominent and systematic approaches for constructing evaluation coresets generally align with two main paradigms. Psychometric methods, which score questions based on historical model correctness using Item Response Theory (IRT) [12], are fundamentally bottlenecked by their heavy reliance on massive, pre-existing evaluation logs. Consequently, they face a severe “cold start” problem: for newly constructed benchmarks or private evaluation datasets, the extensive historical response data required to calibrate item parameters is simply unavailable. This strict prerequisite renders history-dependent methods completely inapplicable to emerging evaluation scenarios where ground-truth correctness records across a wide range of models have not yet been accumulated. On the other hand, geometric methods, such as the widely adopted k-center greedy algorithm [14], typically operate on shallow textual embeddings of the questions. While computationally efficient, these dense encoders exhibit a strong bias toward surface lexical overlap [16,19], fundamentally failing to capture the latent reasoning structures underlying the tasks.
Figure 1. The motivation and overview of CoT-Core.
Figure 1. The motivation and overview of CoT-Core.
Preprints 223734 g001
To overcome the limitations of history-dependent metrics and shallow lexical features, we propose CoT-Core, a novel training-free core question selection framework. Our approach is motivated by a critical observation: lexically disparate questions—which appear entirely unrelated in standard text embedding spaces—can share equivalent underlying logic or reasoning structures. To capture this underlying reasoning manifold, CoT-Core first prompts LLMs to unroll explicit Chain-of-Thought (CoT) reasoning trajectories. By directly projecting these unrolled reasoning paths into a dense embedding space, CoT-Core effectively overcomes lexical divergence, achieving highly precise and cognitive-aligned question clustering. Building upon this logically structured latent space, we apply standard clustering algorithms to extract a minimal, highly representative core subset. This selected coreset serves as a high-fidelity proxy for the full benchmark, drastically mitigating the computational and financial overhead of continuous model evaluation while maintaining accurate performance estimation.
Our main contributions are summarized as follows: 1) We propose CoT-Core, a novel training-free coreset selection framework that circumvents surface-level lexical traps. By projecting unrolled reasoning trajectories into a latent space, it clusters questions based on their underlying reasoning manifold rather than superficial text features. 2) Methodologically, we decouple spatial coreset selection from downstream proxy estimation. CoT-Core employs centroid-based geometric sampling to extract representative instances, naturally supporting Standard Aggregation (SA) for a strictly training-free pipeline, while remaining seamlessly compatible with advanced probabilistic estimators (e.g., GP-IRT). 3) Extensive experiments demonstrate that CoT-Core drastically reduces evaluation overhead while maintaining high-fidelity score estimation. Furthermore, we delineate the boundary conditions of reasoning-aware pruning, revealing that its efficacy is intrinsically tied to task complexity.

3. Methodology

In this section, we present the technical details of CoT-Core. We first formally define the evaluation coreset problem, and then detail the three primary stages of our pipeline: CoT Trajectory Generation, Contextualized Trajectory Embedding, and Isomorphic Clustering and Selection.

3.1. Problem Formulation

Let D = { ( q i , a i ) } i = 1 N denote a comprehensive LLM evaluation benchmark comprising N test instances, where q i is the question prompt and a i is the reference answer. For a given large language model M , the standard evaluation process computes its true overall performance by averaging the exact match scores across the entire dataset:
S D ( M ) = 1 N i = 1 N m ( M ( q i ) , a i )
where M ( q i ) represents the model’s generated output and m ( · , · ) is a specific evaluation metric function (e.g., exact match accuracy).
The overarching objective of coreset evaluation intrinsically consists of two orthogonal stages: Subset Selection and Performance Estimation.
First, the selection stage identifies a minimal representative subset S D , strictly bounded by a budget constraint | S | = B N . Second, the estimation stage employs a proxy scoring function, denoted as f e s t , to predict the model’s full-dataset capability based solely on its responses to the constrained subset S .
The ultimate goal is to formulate a subset S and an estimator f e s t such that the predicted proxy score closely approximates the true full-dataset performance for any unseen model M :
S ^ D ( M ) = f e s t ( M , S ) S D ( M )
Unlike history-dependent baselines that tightly couple their selection algorithms with specific parameterized estimators, CoT-Core operates fundamentally as an estimator-agnostic framework. Our primary objective is to deliver a highly informative subset S in a strictly training-free setting. The resulting coreset is inherently flexible: it naturally supports Standard Aggregation (SA) as f e s t for a deterministic baseline, while remaining fully compatible with advanced probabilistic estimators if historical data is accessible.

3.2. CoT-Core Framework Overview

The severe absence of historical data in real-world scenarios forces coreset construction into a strict training-free setting. Under this severe constraint, conventional text embeddings fundamentally fail to capture the underlying logic. Compelled by this limitation, CoT-Core leverages unrolled reasoning trajectories as a necessary proxy to measure true semantic distance.
As illustrated in Figure 2, our framework systematically maps the reasoning manifold of evaluation datasets through a straightforward, three-stage pipeline: (1) eliciting zero-shot reasoning trajectories via a generator LLM (Section 3.3), (2) projecting these contextualized paths into a dense latent space (Section 3.4), and (3) applying training-free geometric clustering to extract the final representative coreset (Section 3.5).

3.3. Stage 1: CoT Trajectory Generation

As highlighted in our motivation, embedding the raw question text q i captures only superficial lexical features (e.g., vocabulary overlap), fundamentally failing to capture its underlying logic. To address this, our first stage shifts the representation space from the problem definition to the problem resolution path.
For each instance in the dataset D , we utilize an instruction-tuned LLM as the generator, denoted as M g e n . We append a standard Chain-of-Thought trigger (e.g., “Let’s think step by step”) to the question q i , compelling the model to articulate its complete reasoning process. The generated trajectory is defined as:
t i = M g e n ( q i p c o t )
where ⊕ denotes string concatenation and p c o t is the CoT prompt. Crucially, we retain this trajectory t i regardless of whether the final derived answer is correct. The unrolled path itself—even if ultimately flawed—explicitly exposes the underlying logic and step-by-step reasoning structure required to navigate the problem space.

3.4. Stage 2: Contextualized Trajectory Embedding

Instead of relying on history-dependent performance metrics, we directly leverage the structural density of the generated reasoning paths. Crucially, to ensure that the unrolled logic remains grounded in its original premise, we do not embed the trajectory in isolation. For each instance, we concatenate the raw question q i and the generated trajectory t i into a structured input sequence:
s i = Question : q i Reason : t i
where ⊕ denotes string concatenation. This contextualized sequence s i is then mapped into a continuous, high-dimensional space using a pre-trained dense retrieval model E , yielding the reasoning feature matrix E R N × d :
e i = E ( s i )
By embedding the logically unrolled CoT explicitly grounded in its problem context, e i naturally captures the underlying reasoning manifold of the task. This effectively overcomes lexical divergence between lexically disparate but logically isomorphic problems.

3.5. Stage 3: Isomorphic Clustering and Selection

Because the embedding space now inherently encodes the underlying logic and structural equivalence, we apply standard k-means clustering to partition the N trajectory embeddings into B clusters, aligning with the coreset budget constraint. For each cluster, we select the problem instance whose reasoning embedding is closest to the geometric centroid. This training-free sampling strategy maximizes coverage of the reasoning manifold and systematically eliminates dataset redundancy without requiring prior evaluation data or trainable parameters.
Crucially, CoT-Core decouples the coreset selection S from the downstream proxy scoring function f e s t . By default, it naturally accommodates Standard Aggregation (SA)—a deterministic baseline that applies selection-derived instance weights or reduces to an unweighted empirical mean. Furthermore, this formulation remains seamlessly compatible with advanced probabilistic models (e.g., GP-IRT) for parameterized estimation when historical response data is available.

4. Experiments

4.1. Experimental Setup

Datasets & Implementation Details. We evaluate CoT-Core on four widely adopted reasoning benchmarks: GSM8K [6], GPQA (Diamond split) [13], MMLU [8], and MMLU-Pro [22]. For trajectory generation ( M g e n ), we employ the open-weight Phi-4 model [1] via the OpenCompass framework [11] to ensure standardized reproducibility. The contextualized trajectories are then mapped into a continuous feature space using the BGE-M3 encoder [5] ( E ).
Baselines & Data Splits. We compare CoT-Core against two paradigms of coreset selection. Training-free baselines include Random sampling, Question-Emb K-Means (clustering raw problem text embeddings), and the k-Center Greedy algorithm [14]. History-dependent baselines include Correctness K-Means (clustering historical binary response matrices) and IRT[12]. Evaluation logs are sourced from TinyBenchmarks [12] and the Open LLM Leaderboard [3]. For a rigorous evaluation of generalization, we hold out 100 disjoint models as the test set. For history-dependent methods, we fit latent parameters using a strictly disjoint training set of N = 25 models to simulate practical low-resource settings (results for larger N are deferred to Appendix C).
Evaluation Protocol. Coresets are extracted at strict budgets of 1 % , 5 % , 10 % , 15 % , and 20 % of the full dataset size. To isolate selection quality from downstream aggregation, all subsets are evaluated using two decoupled estimators: Standard Aggregation (SA)—a deterministic baseline that applies selection-derived instance weights or reduces to an unweighted empirical mean—and GP-IRT[12], a parameterized historical estimator. The discrepancy between proxy estimations and true full-dataset scores is quantified via Mean Absolute Error (MAE) (↓) and Ranking Similarity (Sim) (↑). To rigorously mitigate initialization variance, we report the mean performance across 5 independent random seeds for all algorithmic methods, and 20 independent trials for the Random baseline.

4.2. Main Results

Delineating Boundary Conditions via Task Complexity. A cross-sectional analysis of Table 1 and Table 2 delineates our approach’s boundary conditions: CoT-Core’s efficacy is intrinsically gated by benchmark difficulty. On straightforward tasks like GSM8K, the method hits a lower boundary (simplicity floor), yielding noticeable but bounded improvements. However, on expert-level benchmarks (MMLU-Pro, GPQA) where raw text acts as an information bottleneck, entering this operative domain via CoT-Core provides massive informational lift. Remarkably, our training-free approach bridges the gap with history-dependent methods. Under the advanced GP-IRT estimator, CoT-Core frequently matches or surpasses IRT and Correctness K-Means at larger budgets (e.g., 10 % on MMLU/MMLU-Pro), proving that capturing latent logical structure can serve as a highly effective alternative to compiling expensive historical evaluation logs.
Robustness Across Estimators. Under the strictly training-free SA estimator (Table 1), CoT-Core’s CoT-aware coresets demonstrate robust representativeness compared to Question-Emb K-Means and Random sampling across the majority of selection budgets. When transitioning to the advanced GP-IRT estimator (Table 2), the absolute predictive errors (MAE) naturally compress across all methods due to the estimator’s parameterized advantage. Despite this tighter global margin, CoT-Core maintains highly competitive—and often leading—performance among training-free methods. Most notably, our approach showcases exceptional resilience in extreme high-compression regimes (e.g., 5 % to 10 % ), where it effectively preserves the relative performance ranking of evaluated models regardless of the scoring function. For instance, with a mere 5 % budget under GP-IRT, CoT-Core achieves a remarkable Ranking Similarity (Sim) of 0.5123 on GPQA, nearly doubling the correlation preserved by traditional text clustering ( 0.2759 ), underscoring the structural stability of our reasoning-aware selection.

4.3. Analysis

The Simplicity Floor and Informational Boundaries.

CoT-Core’s performance gains are bounded by task complexity, directly reflecting the variance in discriminative information gain introduced by the CoT trajectories. In straightforward benchmarks like GSM8K (our lower boundary), reasoning paths predominantly involve homogeneous basic arithmetic. Consequently, unrolled trajectories offer limited structural variance, and CoT verbosity occasionally introduces linguistic noise rather than useful signals. Conversely, reasoning-aware pruning finds its robust operative domain in complex benchmarks demanding diverse underlying logical structures. Invoking advanced theorems for STEM questions, or explicitly evaluating distractors for non-STEM questions, provides powerful spatial discriminators. Crucially, this informational advantage is highly resilient: even when the generator ( M g e n ) hallucinates on excessively difficult questions, the specific nature of its attempted retrieval or logical confusion provides a far more informative clustering signal than compressed raw question text.

Exposing Cognitive Isomorphism via Logical Trajectories

To understand how CoT-Core externalizes implicit reasoning paths into explicit structural features within the representation space, we qualitatively analyze the reasoning manifold of the embeddings. We focus on GPQA, a highly complex benchmark spanning Physics, Chemistry, and Biology, as it provides an ideal testbed for observing cross-disciplinary semantic shifts.
Macro-Level Manifold: Breaking Lexical Traps.Figure 3 visualizes the embedding spaces before and after introducing Chain-of-Thought (CoT) trajectories. In the Question-Only space (left), the dense retriever naturally separates the three sciences into distinct clusters. However, this apparent separability is largely driven by lexical traps. Because domain-specific vocabularies (e.g., "cell" in Biology vs. "velocity" in Physics) are highly disjoint, the encoder creates artificial boundaries based on surface-level entities rather than intrinsic logical structures.
Conversely, the Question+CoT space (right) exhibits a fascinating phenomenon: while maintaining global structure, we observe deliberate inter-cluster merging. Specifically, the originally monolithic domain clusters (e.g., Biology, in green) begin to fragment, with individual instances clearly breaking away from their main group. Rather than being noise, this structural dispersion indicates that CoT-Core forces the embedding space to align with the underlying reasoning structures. Questions within the same domain that require fundamentally different logical trajectories (e.g., a biology question requiring mathematical calculation versus one requiring factual recall) are pulled apart, disrupting the artificial, lexicon-driven boundaries.
Micro-Level Case Study: Correcting the Logical Manifold. To validate that spatial reorganization captures underlying reasoning rather than surface semantics, we examine two contrasting GPQA scenarios (detailed in Appendix B).
Case 1: Uncovering Logical Isomorphism (False-Negative Correction). Two physics problems—one regarding a CERN Bubble Chamber and another on a spherical detector—exhibit a high Question-Only cosine distance of 0.5461 due to disjoint vocabularies. However, CoT unrolling reveals an identical mathematical skeleton (relativistic kinematics and time dilation). By encoding this shared logic, CoT-Core collapses their distance by 61% to 0.2128, ensuring clustering targets reasoning skills rather than experimental dressing.
Case 2: Escaping Lexical Traps (False-Positive Correction). Conversely, traditional encoders frequently conflate distinct problems sharing deceptive surface vocabulary. Two astrophysics questions about a 6000 K spotted star (quantum state population vs. macroscopic photometric transit) are dangerously entangled with a distance of 0.1957. CoT-Core exposes their divergent physical laws (microscopic Boltzmann distribution vs. macroscopic Stefan-Boltzmann law), shattering the lexical illusion and nearly doubling their distance to 0.3890 to prevent pruning non-redundant knowledge.

4.4. Ablation Studies

Impact of the CoT Generator: The Verbosity Penalty. We investigate the impact of the LLM used to externalize reasoning trajectories on MMLU-Pro, varying the generator from lightweight (1B) to frontier models (72B) while fixing the encoder to BGE-M3.
Table 3. Ablation on generator scale (MMLU-Pro). Lightweight models perform remarkably competitively, particularly at extreme compression ratios.
Table 3. Ablation on generator scale (MMLU-Pro). Lightweight models perform remarkably competitively, particularly at extreme compression ratios.
Generator 1% Budget 10% Budget 20% Budget
LLaMA-3-1B 0.0280 ± 0.0204 0.0109 ± 0.0082 0.0057 ± 0.0043
Mistral-8B 0.0316 ± 0.0240 0.0081 ± 0.0058 0.0066 ± 0.0051
Phi-4 0.0295 ± 0.0241 0.0082 ± 0.0061 0 . 0054 ± 0 . 0043
Qwen-72B 0.0305 ± 0.0236 0 . 0077 ± 0 . 0060 0.0065 ± 0.0048
A counter-intuitive observation is that scaling to a massive 72B generator yields marginal, often statistically insignificant, clustering improvements over lightweight alternatives. This highlights a core premise of our framework: trajectory correctness is secondary to logical structure. Even when 1B models derive incorrect final answers, they successfully externalize the requisite reasoning steps, fully exposing the problem’s underlying reasoning manifold. Furthermore, frontier models suffer from a Verbosity Penalty. Extensive instruction tuning causes 72B models to generate excessive conversational boilerplate. Because dense encoders are sensitive to lexical frequency [15], this universal linguistic noise artificially inflates similarities between fundamentally disparate problems. Consequently, the compact, raw logical trajectories of smaller models prove highly robust for geometric clustering, confirming CoT-Core’s cost-efficiency.
Impact of Hierarchical Cleaning: The Optimal Abstraction Horizon. To isolate CoT-Core’s mechanisms, we evaluate three abstraction levels on GPQA using the Phi-4 generator (prompts in Appendix A.2): Raw CoT (unaltered), Cleaned CoT (conversational boilerplate removed, entities retained), and Abstract CoT (extreme compression retaining only pure logical structures).
As Table 4 shows, the intermediate Cleaned CoT consistently minimizes predictive error. Stripping universally shared boilerplate drastically reduces linguistic noise, sharpening manifold boundaries. However, extreme compression (Abstract CoT) severely degrades performance, exposing a critical failure mode: Semantic Starvation. Unlike elementary math, expert domains require domain-specific semantic anchors to contextualize the underlying logic. Stripping these disciplinary groundings deprives the encoder of essential contextual dimensionality. Thus, CoT-Core operates optimally by pruning generic noise while preserving the semantic anchors required to contextualize the latent logical structure.

5. Conclusions

We presented CoT-Core, a training-free coreset framework that circumvents surface lexical traps by clustering the underlying reasoning manifold of unrolled Chain-of-Thought trajectories, successfully aligning lexically disparate problems via their intrinsic logical isomorphism. Extensive evaluations confirm CoT-Core consistently yields high-fidelity proxy estimations for both standard and advanced probabilistic estimators. Crucially, we delineate the boundary conditions of reasoning-aware pruning: its efficacy intrinsically scales with task complexity, bounded by a simplicity floor on basic arithmetic while delivering massive informational gains on expert-level benchmarks. Furthermore, achieving optimal representation requires balancing model-induced conversational noise (Verbosity Penalty) against the severe loss of disciplinary context (Semantic Starvation). Ultimately, CoT-Core provides a scalable, data-free alternative to compiling expensive historical evaluation logs, proving that capturing latent logical structures is sufficient for robust benchmark compression and accelerating continuous LLM development.

Appendix A. Implementation Details

This section provides comprehensive details regarding our experimental setup to facilitate reproducibility.

Appendix A.1. Model Configurations

We employ the pre-trained dense encoder (e.g., BGE-M3) for contextualized trajectory embedding. For generating the Unrolled Reasoning Trajectories, we utilize the generator M g e n with the following generation hyperparameters: temperature T = 0.7 , and maximum token length set to 1024.

Appendix A.2. Prompt Templates for Hierarchical Cleaning

To elicit implicit reasoning from the generator and subsequently perform our hierarchical cleaning ablations (as discussed in Section 4.4), we utilize specific system prompts. To strictly control for confounding variables, both the initial reasoning generation and the subsequent hierarchical cleaning steps employ the identical Phi-4 model.
For the initial generation, we append a standard CoT trigger p c o t (e.g., “Let’s think step by step”) to the original question q i . The subsequent abstraction processes utilize the exact prompt templates presented below.
Preprints 223734 i001Preprints 223734 i002

Appendix B. Case Study: Full Problem Descriptions and Reasoning Trajectories

This appendix provides the complete original text for the four GPQA questions analyzed in our micro-level case study (Section 4.2), alongside their unrolled Chain-of-Thought (CoT) trajectories generated by the M g e n (Phi-4) model.

Appendix B.1. Case 1: Uncovering Logical Isomorphism (False-Negative Correction)

These two problems are lexically disjoint but share an identical underlying mathematical skeleton rooted in relativistic kinematics and time dilation.

Question A (Experimental Resolution):

In the CERN Bubble Chamber a decay occurs, X 0 Y + Z in τ 0 = 8 × 10 16 s, i.e. the proper lifetime of X 0 . What minimum resolution is needed to observe at least 30% of the decays? Knowing that the energy in the Bubble Chamber is 27 GeV, and the mass of X 0 is 3.41 GeV.
Preprints 223734 i003Preprints 223734 i004

Question B (Kinematic State):

Particles are collided at the center of a spherical detector producing new type of particles that travel uninterrupted at ultra-relativistic velocities highly centered around Lorentz factor of 20 . On average, one third of these fast-decaying particles reaches the detector inner walls. The radius of the detector is 30 meters. What Lorentz factor is needed in order to have about two thirds of these particles reaching the detector inner walls?
Preprints 223734 i005Preprints 223734 i006

Analysis of the Isomorphism:

This case illustrates the False-Negative problem inherent in standard text-embedding approaches and how unrolled Chain-of-Thought (CoT) trajectories mitigate it.
If we rely solely on the problem formulations, Question A and Question B appear semantically disjoint. Question A is framed as an experimental high-energy physics problem (featuring entities like “CERN Bubble Chamber,” “27 GeV,” and “resolution”), whereas Question B is framed as a macroscopic kinematic problem (featuring “spherical detector,” “30 meters,” and “inner walls”). A naive embedding model would map these two questions to distant regions in the latent space due to their lexical divergence.
However, the unrolled CoTs reveal a shared, identical mathematical skeleton. Both reasoning trajectories sequentially invoke the same underlying physical principles:
1.
Relativistic Kinematics: Calculating or utilizing the Lorentz factor ( γ ).
2.
Time Dilation: Linking proper lifetime to dilated lifetime ( τ = γ τ 0 ).
3.
Mean Decay Length: Formulating the distance traveled before decay ( d c γ τ 0 ).
4.
Threshold Probability: Correlating the traveled distance with the survival/decay ratio.
By exposing this implicit reasoning structure, the CoT acts as a semantic equalizer. The generation of shared specialized tokens (e.g., γ , τ 0 , v c ) and structurally identical mathematical equations forces the embedding space to recognize their deep logical isomorphism. Consequently, these two instances, which would be falsely separated (False-Negative) in a purely lexical feature space, are successfully clustered together based on their true cognitive and reasoning demands.1

Appendix B.2. Case 2: Escaping Lexical Traps (False-Positive Correction)

These two problems share extensive surface-level vocabulary regarding spotted stars and effective temperatures, but require fundamentally orthogonal physical laws (quantum statistical mechanics vs. macroscopic photometric geometry).

Question C (Quantum States):

Astronomers are studying a star with a 1.5 solar radius and 1.1 solar masses. When the star’s surface is not covered by dark spots, its T eff is 6000K. However, when 40% of its surface is covered by spots, the overall photospheric effective temperature decreases to 5500K. In the stellar photosphere, when examining the ratio of the number of neutral atoms of Ti in two energetic levels (level 1 and level 2), astronomers have observed that this ratio decreases when the star has spots. What is the factor by which this ratio changes when the star does not have spots compared to when it has spots? Note that the transition between the energy levels under consideration corresponds to a wavelength of approximately 1448 Å. Assume that the stellar photosphere is in LTE.
Preprints 223734 i007Preprints 223734 i008

Question D (Photometric Flux):

Astronomers are currently observing a star with a radius equal to that of the Sun. One hemisphere of the star is covered in dark spots with a filling factor of 20%. The star has an effective temperature ( T eff ) of 6000K, and the spots exhibit a temperature difference of 1000K. As only one hemisphere is spotty, photometric time-series observations will reveal periodic variations in brightness due to rotational modulation. Interestingly, this situation can closely resemble the presence of an exoplanet. To produce the same amplitude signal in the star’s light curve (if the star was not covered by spots!), what should be the radius of a hypothetical exoplanet relative to the radius of the host star (i.e. R pl / R star )?
Preprints 223734 i009Preprints 223734 i010

Analysis of the Orthogonality:

This case demonstrates the “Lexical Trap” (False-Positive) problem in benchmark evaluation and how unrolled CoT prevents erroneous clustering.
At the surface level, Question C and Question D are lexically entangled. Both problem descriptions share a high density of domain-specific keywords and identical numerical setups: “star,” “dark spots,” “effective temperature,” and “6000K.” A naive text-embedding model, which heavily weights term frequencies and surface-level semantic proximity, would inevitably map these two questions into the same cluster, classifying them as highly similar stellar astrophysics problems.
However, the unrolled CoT trajectories expose fundamentally orthogonal underlying mathematical and physical engines:
  • Question C (Quantum Statistical Mechanics): The reasoning trajectory is entirely driven by atomic physics. It invokes the Boltzmann distribution to calculate microscopic population states, generating specialized tokens like energy differences ( Δ E ), Planck’s constant (h), and exponential decay functions dependent on temperature ( exp ( Δ E / k T ) ).
  • Question D (Macroscopic Photometry): The reasoning trajectory operates on geometric and thermodynamic scales. It relies on the Stefan-Boltzmann law (where flux is proportional to T 4 ) and geometrical area ratios (projected disk areas and filling factors) to calculate macroscopic light curve variations.
By expanding the implicit reasoning paths, the CoT introduces dimensionally distinct vocabularies (e.g., h c / λ k vs. T 4 ratios). This forced articulation of the underlying physical laws acts as a semantic repeller in the latent space, effectively pushing these two lexically similar but logically distinct problems apart, thereby resolving the False-Positive collision.2

Appendix C. Extended Main Results

In the main text (Section 4.2), we reported the performance under the low-resource setting ( N = 25 ). Here, we present the full evaluation results across larger historical budgets N { 50 , 100 , 200 } .
Table A1. Performance comparison at N=50 using SA proxy. Best performance among Training-Free methods is highlighted in bold. Trainable methods are shown for reference only.
Table A1. Performance comparison at N=50 using SA proxy. Best performance among Training-Free methods is highlighted in bold. Trainable methods are shown for reference only.
Benchmark Method Sampling Ratio
1%
MAE (↓) / Sim (↑)
5%
MAE (↓) / Sim (↑)
10%
MAE (↓) / Sim (↑)
15%
MAE (↓) / Sim (↑)
20%
MAE (↓) / Sim (↑)
Training-Free Methods (Zero-Shot)
GSM8K Random 0.0890 / 0.9044 0.0382 / 0.9747 0.0246 / 0.9851 0.0187 / 0.9897 0.0142 / 0.9920
Question-Emb K-Means 0.0875 / 0.8951 0.0321 / 0.9743 0.0241 / 0.9797 0.0175 / 0.9871 0.0133 / 0.9931
k-Center Greedy 0.0683 / 0.9118 0.0349 / 0.9738 0.0283 / 0.9847 0.0218 / 0.9897 0.0175 / 0.9925
CoT-Feature (Ours) 0.0755 / 0.8884 0.0331 / 0.9696 0.0209 / 0.9853 0.0197 / 0.9884 0.0154 / 0.9918
Trainable Methods (Require Historical Data, N=50)
Correctness K-Means 0.0623 / 0.9130 0.0339 / 0.9748 0.0256 / 0.9816 0.0230 / 0.9848 0.0195 / 0.9887
IRT 0.0670 / 0.9053 0.0329 / 0.9754 0.0241 / 0.9833 0.0169 / 0.9908 0.0143 / 0.9907
Training-Free Methods (Zero-Shot)
MMLU Random 0.0330 / 0.9489 0.0149 / 0.9880 0.0102 / 0.9935 0.0079 / 0.9957 0.0063 / 0.9969
Question-Emb K-Means 0.0298 / 0.9484 0.0215 / 0.9877 0.0154 / 0.9923 0.0203 / 0.9943 0.0252 / 0.9958
k-Center Greedy 0.0477 / 0.9461 0.0530 / 0.9867 0.0393 / 0.9909 0.0360 / 0.9942 0.0317 / 0.9952
CoT-Feature (Ours) 0.0324 / 0.9353 0.0118 / 0.9881 0.0081 / 0.9941 0.0081 / 0.9964 0.0060 / 0.9971
Trainable Methods (Require Historical Data, N=50)
Correctness K-Means 0.0272 / 0.9638 0.0191 / 0.9789 0.0161 / 0.9825 0.0152 / 0.9838 0.0144 / 0.9851
IRT 0.0225 / 0.9588 0.0113 / 0.9896 0.0083 / 0.9938 0.0069 / 0.9960 0.0065 / 0.9968
Training-Free Methods (Zero-Shot)
MMLU-Pro Random 0.0311 / 0.9556 0.0132 / 0.9899 0.0093 / 0.9941 0.0070 / 0.9956 0.0058 / 0.9967
Question-Emb K-Means 0.0300 / 0.9579 0.0133 / 0.9890 0.0111 / 0.9947 0.0116 / 0.9957 0.0100 / 0.9969
k-Center Greedy 0.0369 / 0.9524 0.0432 / 0.9857 0.0334 / 0.9923 0.0306 / 0.9936 0.0272 / 0.9957
CoT-Feature (Ours) 0.0296 / 0.9682 0.0124 / 0.9893 0.0082 / 0.9941 0.0072 / 0.9965 0.0055 / 0.9969
Trainable Methods (Require Historical Data, N=50)
Correctness K-Means 0.0283 / 0.9769 0.0172 / 0.9903 0.0152 / 0.9933 0.0131 / 0.9942 0.0113 / 0.9947
IRT 0.0412 / 0.9642 0.0208 / 0.9902 0.0151 / 0.9938 0.0122 / 0.9955 0.0102 / 0.9963
Training-Free Methods (Zero-Shot)
GPQA Random - / - 0.1252 / 0.2851 0.0797 / 0.4480 0.0635 / 0.5579 0.0536 / 0.6022
Question-Emb K-Means - / - 0.1098 / 0.2856 0.0859 / 0.3656 0.0615 / 0.5338 0.0522 / 0.6263
k-Center Greedy - / - 0.1232 / 0.2197 0.0821 / 0.4036 0.0608 / 0.5355 0.0565 / 0.6151
CoT-Feature (Ours) - / - 0.1044 / 0.4334 0.0731 / 0.5839 0.0601 / 0.5646 0.0516 / 0.6013
Trainable Methods (Require Historical Data, N=50)
Correctness K-Means - / - 0.1269 / 0.2524 0.0942 / 0.3297 0.0771 / 0.4398 0.0666 / 0.4601
IRT - / - 0.1047 / 0.3878 0.0780 / 0.4071 0.0656 / 0.5104 0.0540 / 0.5748
Table A2. Performance comparison at N=50 using GPIRT proxy. Best performance among Training-Free methods is highlighted in bold. Trainable methods are shown for reference only.
Table A2. Performance comparison at N=50 using GPIRT proxy. Best performance among Training-Free methods is highlighted in bold. Trainable methods are shown for reference only.
Benchmark Method Sampling Ratio
1%
MAE (↓) / Sim (↑)
5%
MAE (↓) / Sim (↑)
10%
MAE (↓) / Sim (↑)
15%
MAE (↓) / Sim (↑)
20%
MAE (↓) / Sim (↑)
Training-Free Methods (Zero-Shot)
GSM8K Random 0.0933 / 0.8809 0.0325 / 0.9670 0.0203 / 0.9839 0.0157 / 0.9894 0.0129 / 0.9917
Question-Emb K-Means 0.0781 / 0.8796 0.0310 / 0.9679 0.0204 / 0.9815 0.0161 / 0.9871 0.0123 / 0.9930
k-Center Greedy 0.0972 / 0.8676 0.0333 / 0.9658 0.0214 / 0.9833 0.0167 / 0.9899 0.0143 / 0.9929
CoT-Feature (Ours) 0.0849 / 0.8713 0.0297 / 0.9592 0.0192 / 0.9856 0.0163 / 0.9893 0.0133 / 0.9918
Trainable Methods (Require Historical Data, N=50)
Correctness K-Means 0.0698 / 0.9103 0.0292 / 0.9718 0.0183 / 0.9855 0.0152 / 0.9887 0.0140 / 0.9926
IRT 0.0805 / 0.8858 0.0278 / 0.9642 0.0191 / 0.9844 0.0149 / 0.9910 0.0128 / 0.9924
Training-Free Methods (Zero-Shot)
MMLU Random 0.0207 / 0.9658 0.0106 / 0.9908 0.0082 / 0.9945 0.0067 / 0.9961 0.0056 / 0.9971
Question-Emb K-Means 0.0220 / 0.9597 0.0123 / 0.9899 0.0108 / 0.9934 0.0146 / 0.9948 0.0192 / 0.9961
k-Center Greedy 0.0208 / 0.9651 0.0239 / 0.9892 0.0244 / 0.9925 0.0257 / 0.9945 0.0244 / 0.9955
CoT-Feature (Ours) 0.0233 / 0.9535 0.0100 / 0.9900 0.0068 / 0.9949 0.0066 / 0.9965 0.0056 / 0.9971
Trainable Methods (Require Historical Data, N=50)
Correctness K-Means 0.0207 / 0.9723 0.0141 / 0.9868 0.0129 / 0.9887 0.0125 / 0.9892 0.0121 / 0.9894
IRT 0.0213 / 0.9655 0.0105 / 0.9911 0.0077 / 0.9949 0.0066 / 0.9964 0.0065 / 0.9971
Training-Free Methods (Zero-Shot)
MMLU-Pro Random 0.0239 / 0.9624 0.0097 / 0.9906 0.0072 / 0.9943 0.0058 / 0.9957 0.0049 / 0.9967
Question-Emb K-Means 0.0238 / 0.9593 0.0098 / 0.9899 0.0075 / 0.9944 0.0073 / 0.9957 0.0069 / 0.9968
k-Center Greedy 0.0249 / 0.9582 0.0146 / 0.9895 0.0149 / 0.9936 0.0163 / 0.9947 0.0163 / 0.9956
CoT-Feature (Ours) 0.0226 / 0.9691 0.0105 / 0.9899 0.0070 / 0.9939 0.0059 / 0.9959 0.0047 / 0.9966
Trainable Methods (Require Historical Data, N=50)
Correctness K-Means 0.0226 / 0.9668 0.0102 / 0.9909 0.0103 / 0.9946 0.0101 / 0.9951 0.0093 / 0.9954
IRT 0.0278 / 0.9518 0.0100 / 0.9914 0.0083 / 0.9947 0.0079 / 0.9958 0.0076 / 0.9966
Training-Free Methods (Zero-Shot)
GPQA Random - / - 0.1281 / 0.2834 0.0596 / 0.4492 0.0414 / 0.6098 0.0352 / 0.6768
Question-Emb K-Means - / - 0.1370 / 0.2460 0.0601 / 0.4870 0.0379 / 0.6004 0.0362 / 0.6651
k-Center Greedy - / - 0.1172 / 0.3289 0.0612 / 0.4532 0.0434 / 0.5656 0.0364 / 0.6267
CoT-Feature (Ours) - / - 0.1242 / 0.3571 0.0612 / 0.5275 0.0412 / 0.5904 0.0362 / 0.6616
Trainable Methods (Require Historical Data, N=50)
Correctness K-Means - / - 0.1028 / 0.3905 0.0491 / 0.4433 0.0382 / 0.5559 0.0373 / 0.6144
IRT - / - 0.1117 / 0.3240 0.0505 / 0.4854 0.0411 / 0.6455 0.0352 / 0.6663
Table A3. Performance comparison at N=100 using SA proxy. Best performance among Training-Free methods is highlighted in bold. Trainable methods are shown for reference only.
Table A3. Performance comparison at N=100 using SA proxy. Best performance among Training-Free methods is highlighted in bold. Trainable methods are shown for reference only.
Benchmark Method Sampling Ratio
1%
MAE (↓) / Sim (↑)
5%
MAE (↓) / Sim (↑)
10%
MAE (↓) / Sim (↑)
15%
MAE (↓) / Sim (↑)
20%
MAE (↓) / Sim (↑)
Training-Free Methods (Zero-Shot)
GSM8K Random 0.0890 / 0.9044 0.0382 / 0.9747 0.0246 / 0.9851 0.0187 / 0.9897 0.0142 / 0.9920
Question-Emb K-Means 0.0875 / 0.8951 0.0321 / 0.9743 0.0241 / 0.9797 0.0175 / 0.9871 0.0133 / 0.9931
k-Center Greedy 0.0683 / 0.9118 0.0349 / 0.9738 0.0283 / 0.9847 0.0218 / 0.9897 0.0175 / 0.9925
CoT-Feature (Ours) 0.0755 / 0.8884 0.0331 / 0.9696 0.0209 / 0.9853 0.0197 / 0.9884 0.0154 / 0.9918
Trainable Methods (Require Historical Data, N=100)
Correctness K-Means 0.0522 / 0.9226 0.0344 / 0.9699 0.0299 / 0.9843 0.0266 / 0.9865 0.0233 / 0.9903
IRT 0.0650 / 0.9049 0.0295 / 0.9734 0.0216 / 0.9797 0.0168 / 0.9880 0.0141 / 0.9893
Training-Free Methods (Zero-Shot)
MMLU Random 0.0330 / 0.9489 0.0149 / 0.9880 0.0102 / 0.9935 0.0079 / 0.9957 0.0063 / 0.9969
Question-Emb K-Means 0.0298 / 0.9484 0.0215 / 0.9877 0.0154 / 0.9923 0.0203 / 0.9943 0.0252 / 0.9958
k-Center Greedy 0.0477 / 0.9461 0.0530 / 0.9867 0.0393 / 0.9909 0.0360 / 0.9942 0.0317 / 0.9952
CoT-Feature (Ours) 0.0324 / 0.9353 0.0118 / 0.9881 0.0081 / 0.9941 0.0081 / 0.9964 0.0060 / 0.9971
Trainable Methods (Require Historical Data, N=100)
Correctness K-Means 0.0184 / 0.9762 0.0111 / 0.9906 0.0100 / 0.9917 0.0091 / 0.9927 0.0082 / 0.9936
IRT 0.0222 / 0.9565 0.0105 / 0.9895 0.0089 / 0.9932 0.0071 / 0.9950 0.0063 / 0.9962
Training-Free Methods (Zero-Shot)
MMLU-Pro Random 0.0311 / 0.9556 0.0132 / 0.9899 0.0093 / 0.9941 0.0070 / 0.9956 0.0058 / 0.9967
Question-Emb K-Means 0.0300 / 0.9579 0.0133 / 0.9890 0.0111 / 0.9947 0.0116 / 0.9957 0.0100 / 0.9969
k-Center Greedy 0.0369 / 0.9524 0.0432 / 0.9857 0.0334 / 0.9923 0.0306 / 0.9936 0.0272 / 0.9957
CoT-Feature (Ours) 0.0296 / 0.9682 0.0124 / 0.9893 0.0082 / 0.9941 0.0072 / 0.9965 0.0055 / 0.9969
Trainable Methods (Require Historical Data, N=100)
Correctness K-Means 0.0289 / 0.9815 0.0219 / 0.9919 0.0192 / 0.9937 0.0175 / 0.9940 0.0157 / 0.9948
IRT 0.0274 / 0.9561 0.0111 / 0.9923 0.0085 / 0.9938 0.0064 / 0.9953 0.0058 / 0.9961
Training-Free Methods (Zero-Shot)
GPQA Random - / - 0.1252 / 0.2851 0.0797 / 0.4480 0.0635 / 0.5579 0.0536 / 0.6022
Question-Emb K-Means - / - 0.1098 / 0.2856 0.0859 / 0.3656 0.0615 / 0.5338 0.0522 / 0.6263
k-Center Greedy - / - 0.1232 / 0.2197 0.0821 / 0.4036 0.0608 / 0.5355 0.0565 / 0.6151
CoT-Feature (Ours) - / - 0.1044 / 0.4334 0.0731 / 0.5839 0.0601 / 0.5646 0.0516 / 0.6013
Trainable Methods (Require Historical Data, N=100)
Correctness K-Means - / - 0.1124 / 0.4882 0.0925 / 0.4670 0.0792 / 0.5425 0.0742 / 0.5639
IRT - / - 0.1125 / 0.3425 0.0713 / 0.4955 0.0511 / 0.6022 0.0466 / 0.6193
Table A4. Performance comparison at N=100 using GPIRT proxy. Best performance among Training-Free methods is highlighted in bold. Trainable methods are shown for reference only.
Table A4. Performance comparison at N=100 using GPIRT proxy. Best performance among Training-Free methods is highlighted in bold. Trainable methods are shown for reference only.
Benchmark Method Sampling Ratio
1%
MAE (↓) / Sim (↑)
5%
MAE (↓) / Sim (↑)
10%
MAE (↓) / Sim (↑)
15%
MAE (↓) / Sim (↑)
20%
MAE (↓) / Sim (↑)
Training-Free Methods (Zero-Shot)
GSM8K Random 0.0881 / 0.8824 0.0292 / 0.9732 0.0200 / 0.9856 0.0161 / 0.9899 0.0130 / 0.9921
Question-Emb K-Means 0.0899 / 0.8723 0.0276 / 0.9723 0.0210 / 0.9804 0.0158 / 0.9877 0.0125 / 0.9931
k-Center Greedy 0.0831 / 0.8703 0.0279 / 0.9725 0.0216 / 0.9832 0.0177 / 0.9897 0.0149 / 0.9925
CoT-Feature (Ours) 0.0818 / 0.8751 0.0285 / 0.9664 0.0188 / 0.9860 0.0162 / 0.9889 0.0135 / 0.9917
Trainable Methods (Require Historical Data, N=100)
Correctness K-Means 0.0741 / 0.9099 0.0263 / 0.9727 0.0205 / 0.9893 0.0195 / 0.9910 0.0177 / 0.9934
IRT 0.0651 / 0.8984 0.0260 / 0.9748 0.0195 / 0.9809 0.0152 / 0.9894 0.0128 / 0.9902
Training-Free Methods (Zero-Shot)
MMLU Random 0.0199 / 0.9662 0.0098 / 0.9909 0.0076 / 0.9944 0.0062 / 0.9960 0.0053 / 0.9970
Question-Emb K-Means 0.0206 / 0.9626 0.0107 / 0.9895 0.0086 / 0.9940 0.0103 / 0.9949 0.0132 / 0.9962
k-Center Greedy 0.0209 / 0.9684 0.0129 / 0.9902 0.0141 / 0.9935 0.0161 / 0.9953 0.0164 / 0.9959
CoT-Feature (Ours) 0.0218 / 0.9631 0.0096 / 0.9911 0.0069 / 0.9952 0.0061 / 0.9964 0.0054 / 0.9971
Trainable Methods (Require Historical Data, N=100)
Correctness K-Means 0.0190 / 0.9715 0.0100 / 0.9906 0.0083 / 0.9932 0.0075 / 0.9941 0.0066 / 0.9948
IRT 0.0188 / 0.9684 0.0099 / 0.9913 0.0076 / 0.9945 0.0064 / 0.9957 0.0057 / 0.9966
Training-Free Methods (Zero-Shot)
MMLU-Pro Random 0.0229 / 0.9629 0.0094 / 0.9907 0.0066 / 0.9947 0.0053 / 0.9960 0.0045 / 0.9969
Question-Emb K-Means 0.0227 / 0.9581 0.0093 / 0.9899 0.0069 / 0.9946 0.0061 / 0.9959 0.0054 / 0.9970
k-Center Greedy 0.0229 / 0.9640 0.0102 / 0.9900 0.0086 / 0.9941 0.0096 / 0.9949 0.0099 / 0.9958
CoT-Feature (Ours) 0.0209 / 0.9706 0.0099 / 0.9903 0.0065 / 0.9944 0.0051 / 0.9965 0.0043 / 0.9972
Trainable Methods (Require Historical Data, N=100)
Correctness K-Means 0.0295 / 0.9550 0.0103 / 0.9927 0.0091 / 0.9956 0.0085 / 0.9970 0.0086 / 0.9970
IRT 0.0255 / 0.9520 0.0088 / 0.9929 0.0066 / 0.9948 0.0051 / 0.9962 0.0046 / 0.9971
Training-Free Methods (Zero-Shot)
GPQA Random - / - 0.1022 / 0.2447 0.0549 / 0.4685 0.0415 / 0.5990 0.0377 / 0.6439
Question-Emb K-Means - / - 0.0990 / 0.2167 0.0557 / 0.3998 0.0412 / 0.5598 0.0376 / 0.6586
k-Center Greedy - / - 0.1040 / 0.1707 0.0565 / 0.3994 0.0446 / 0.5275 0.0411 / 0.5964
CoT-Feature (Ours) - / - 0.0836 / 0.3995 0.0482 / 0.6417 0.0380 / 0.6166 0.0369 / 0.6505
Trainable Methods (Require Historical Data, N=100)
Correctness K-Means - / - 0.0653 / 0.4379 0.0443 / 0.5324 0.0415 / 0.5921 0.0431 / 0.6339
IRT - / - 0.0944 / 0.2877 0.0490 / 0.5302 0.0360 / 0.6203 0.0330 / 0.6608
Table A5. Performance comparison at N=200 using SA proxy. Best performance among Training-Free methods is highlighted in bold. Trainable methods are shown for reference only.
Table A5. Performance comparison at N=200 using SA proxy. Best performance among Training-Free methods is highlighted in bold. Trainable methods are shown for reference only.
Benchmark Method Sampling Ratio
1%
MAE (↓) / Sim (↑)
5%
MAE (↓) / Sim (↑)
10%
MAE (↓) / Sim (↑)
15%
MAE (↓) / Sim (↑)
20%
MAE (↓) / Sim (↑)
Training-Free Methods (Zero-Shot)
GSM8K Random 0.0890 / 0.9044 0.0382 / 0.9747 0.0246 / 0.9851 0.0187 / 0.9897 0.0142 / 0.9920
Question-Emb K-Means 0.0875 / 0.8951 0.0321 / 0.9743 0.0241 / 0.9797 0.0175 / 0.9871 0.0133 / 0.9931
k-Center Greedy 0.0683 / 0.9118 0.0349 / 0.9738 0.0283 / 0.9847 0.0218 / 0.9897 0.0175 / 0.9925
CoT-Feature (Ours) 0.0755 / 0.8884 0.0331 / 0.9696 0.0209 / 0.9853 0.0197 / 0.9884 0.0154 / 0.9918
Trainable Methods (Require Historical Data, N=200)
Correctness K-Means 0.0546 / 0.9239 0.0348 / 0.9694 0.0303 / 0.9800 0.0277 / 0.9817 0.0259 / 0.9870
IRT 0.0723 / 0.9048 0.0283 / 0.9757 0.0205 / 0.9841 0.0172 / 0.9866 0.0148 / 0.9895
Training-Free Methods (Zero-Shot)
MMLU Random 0.0330 / 0.9489 0.0149 / 0.9880 0.0102 / 0.9935 0.0079 / 0.9957 0.0063 / 0.9969
Question-Emb K-Means 0.0298 / 0.9484 0.0215 / 0.9877 0.0154 / 0.9923 0.0203 / 0.9943 0.0252 / 0.9958
k-Center Greedy 0.0477 / 0.9461 0.0530 / 0.9867 0.0393 / 0.9909 0.0360 / 0.9942 0.0317 / 0.9952
CoT-Feature (Ours) 0.0324 / 0.9353 0.0118 / 0.9881 0.0081 / 0.9941 0.0081 / 0.9964 0.0060 / 0.9971
Trainable Methods (Require Historical Data, N=200)
Correctness K-Means 0.0211 / 0.9699 0.0139 / 0.9885 0.0112 / 0.9898 0.0099 / 0.9911 0.0085 / 0.9935
IRT 0.0223 / 0.9558 0.0097 / 0.9901 0.0072 / 0.9928 0.0060 / 0.9940 0.0055 / 0.9956
Training-Free Methods (Zero-Shot)
MMLU-Pro Random 0.0311 / 0.9556 0.0132 / 0.9899 0.0093 / 0.9941 0.0070 / 0.9956 0.0058 / 0.9967
Question-Emb K-Means 0.0300 / 0.9579 0.0133 / 0.9890 0.0111 / 0.9947 0.0116 / 0.9957 0.0100 / 0.9969
k-Center Greedy 0.0369 / 0.9524 0.0432 / 0.9857 0.0334 / 0.9923 0.0306 / 0.9936 0.0272 / 0.9957
CoT-Feature (Ours) 0.0296 / 0.9682 0.0124 / 0.9893 0.0082 / 0.9941 0.0072 / 0.9965 0.0055 / 0.9969
Trainable Methods (Require Historical Data, N=200)
Correctness K-Means 0.0332 / 0.9847 0.0244 / 0.9918 0.0222 / 0.9927 0.0194 / 0.9937 0.0174 / 0.9940
IRT 0.0233 / 0.9582 0.0123 / 0.9901 0.0080 / 0.9939 0.0061 / 0.9957 0.0054 / 0.9963
Training-Free Methods (Zero-Shot)
GPQA Random - / - 0.1252 / 0.2851 0.0797 / 0.4480 0.0635 / 0.5579 0.0536 / 0.6022
Question-Emb K-Means - / - 0.1098 / 0.2856 0.0859 / 0.3656 0.0615 / 0.5338 0.0522 / 0.6263
k-Center Greedy - / - 0.1232 / 0.2197 0.0821 / 0.4036 0.0608 / 0.5355 0.0565 / 0.6151
CoT-Feature (Ours) - / - 0.1044 / 0.4334 0.0731 / 0.5839 0.0601 / 0.5646 0.0516 / 0.6013
Trainable Methods (Require Historical Data, N=200)
Correctness K-Means - / - 0.1261 / 0.2994 0.1072 / 0.3593 0.0915 / 0.4311 0.0839 / 0.4810
IRT - / - 0.1055 / 0.2956 0.0776 / 0.4353 0.0610 / 0.5367 0.0529 / 0.5793
Table A6. Performance comparison at N=200 using GPIRT proxy. Best performance among Training-Free methods is highlighted in bold. Trainable methods are shown for reference only.
Table A6. Performance comparison at N=200 using GPIRT proxy. Best performance among Training-Free methods is highlighted in bold. Trainable methods are shown for reference only.
Benchmark Method Sampling Ratio
1%
MAE (↓) / Sim (↑)
5%
MAE (↓) / Sim (↑)
10%
MAE (↓) / Sim (↑)
15%
MAE (↓) / Sim (↑)
20%
MAE (↓) / Sim (↑)
Training-Free Methods (Zero-Shot)
GSM8K Random 0.0881 / 0.8761 0.0287 / 0.9744 0.0192 / 0.9859 0.0152 / 0.9898 0.0125 / 0.9921
Question-Emb K-Means 0.0869 / 0.8594 0.0279 / 0.9729 0.0203 / 0.9811 0.0153 / 0.9874 0.0121 / 0.9931
k-Center Greedy 0.0879 / 0.8655 0.0265 / 0.9731 0.0192 / 0.9817 0.0158 / 0.9891 0.0135 / 0.9923
CoT-Feature (Ours) 0.0738 / 0.8793 0.0278 / 0.9693 0.0187 / 0.9860 0.0157 / 0.9886 0.0131 / 0.9914
Trainable Methods (Require Historical Data, N=200)
Correctness K-Means 0.0705 / 0.9242 0.0260 / 0.9762 0.0198 / 0.9875 0.0175 / 0.9890 0.0163 / 0.9918
IRT 0.0706 / 0.9023 0.0252 / 0.9758 0.0181 / 0.9857 0.0142 / 0.9892 0.0130 / 0.9916
Training-Free Methods (Zero-Shot)
MMLU Random 0.0186 / 0.9710 0.0093 / 0.9919 0.0074 / 0.9948 0.0062 / 0.9963 0.0053 / 0.9972
Question-Emb K-Means 0.0201 / 0.9657 0.0104 / 0.9906 0.0090 / 0.9942 0.0112 / 0.9953 0.0150 / 0.9964
k-Center Greedy 0.0188 / 0.9720 0.0160 / 0.9914 0.0173 / 0.9937 0.0192 / 0.9955 0.0190 / 0.9959
CoT-Feature (Ours) 0.0216 / 0.9617 0.0085 / 0.9919 0.0063 / 0.9956 0.0059 / 0.9966 0.0052 / 0.9974
Trainable Methods (Require Historical Data, N=200)
Correctness K-Means 0.0186 / 0.9722 0.0094 / 0.9925 0.0080 / 0.9938 0.0073 / 0.9944 0.0064 / 0.9956
IRT 0.0184 / 0.9714 0.0086 / 0.9921 0.0067 / 0.9938 0.0055 / 0.9954 0.0049 / 0.9964
Training-Free Methods (Zero-Shot)
MMLU-Pro Random 0.0225 / 0.9646 0.0094 / 0.9908 0.0069 / 0.9942 0.0055 / 0.9955 0.0048 / 0.9963
Question-Emb K-Means 0.0228 / 0.9611 0.0094 / 0.9904 0.0074 / 0.9943 0.0069 / 0.9952 0.0062 / 0.9967
k-Center Greedy 0.0228 / 0.9673 0.0117 / 0.9901 0.0111 / 0.9940 0.0127 / 0.9949 0.0127 / 0.9958
CoT-Feature (Ours) 0.0216 / 0.9747 0.0098 / 0.9907 0.0069 / 0.9937 0.0056 / 0.9954 0.0047 / 0.9962
Trainable Methods (Require Historical Data, N=200)
Correctness K-Means 0.0234 / 0.9685 0.0097 / 0.9933 0.0092 / 0.9954 0.0087 / 0.9966 0.0090 / 0.9968
IRT 0.0221 / 0.9619 0.0095 / 0.9908 0.0072 / 0.9940 0.0052 / 0.9958 0.0046 / 0.9963
Training-Free Methods (Zero-Shot)
GPQA Random - / - 0.1108 / 0.2931 0.0875 / 0.4409 0.0578 / 0.5951 0.0423 / 0.6625
Question-Emb K-Means - / - 0.1111 / 0.2699 0.0826 / 0.4144 0.0521 / 0.6110 0.0448 / 0.6466
k-Center Greedy - / - 0.1170 / 0.2745 0.0863 / 0.4452 0.0623 / 0.5391 0.0487 / 0.5753
CoT-Feature (Ours) - / - 0.1031 / 0.3926 0.0737 / 0.5764 0.0492 / 0.6092 0.0423 / 0.6502
Trainable Methods (Require Historical Data, N=200)
Correctness K-Means - / - 0.0805 / 0.3920 0.0605 / 0.4511 0.0449 / 0.4943 0.0406 / 0.5771
IRT - / - 0.0837 / 0.3378 0.0586 / 0.4497 0.0484 / 0.5412 0.0374 / 0.6340

References

  1. Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J Hewett, Mojan Javaheripi, Piero Kauffmann, et al. Phi-4 technical report. arXiv preprint arXiv:2412.08905, 2024.
  2. Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
  3. Edward Beeching, Clémentine Fourrier, Nathan Habib, Sheon Han, Nathan Lambert, Nazneen Rajani, Omar Sanseviero, Lewis Tunstall, and Thomas Wolf. Open llm leaderboard. https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard, 2023.
  4. Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. Pythia: A suite for analyzing large language models across training and scaling. In International conference on machine learning, pp. 2397–2430. PMLR, 2023.
  5. Jianlv Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, and Zheng Liu. Bge m3-embedding: Multi-lingual, multi-functionality, multi-granularity text embeddings through self-knowledge distillation. arXiv preprint arXiv:2402.03216, 4(5), 2024.
  6. Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
  7. Dirk Groeneveld, Iz Beltagy, Evan Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, et al. Olmo: Accelerating the science of language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 15789–15809, 2024. [CrossRef]
  8. Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020.
  9. Peiyu Li, Xiuxiu Tang, Si Chen, Ying Cheng, Ronald Metoyer, Ting Hua, and Nitesh V Chawla. Adaptive testing for llm evaluation: A psychometric alternative to static benchmarks. arXiv preprint arXiv:2511.04689, 2025.
  10. Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110, 2022.
  11. OpenCompass Contributors. Opencompass: A universal evaluation platform for foundation models. https://github.com/open-compass/opencompass, 2023.
  12. Felipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun, Gongjun Xu, and Mikhail Yurochkin. tinybenchmarks: evaluating llms with fewer examples. arXiv preprint arXiv:2402.14992, 2024.
  13. David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman. Gpqa: A graduate-level google-proof q&a benchmark. In First conference on language modeling, 2024.
  14. Ozan Sener and Silvio Savarese. Active learning for convolutional neural networks: A core-set approach. arXiv preprint arXiv:1708.00489, 2017.
  15. Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H Chi, Nathanael Schärli, and Denny Zhou. Large language models can be easily distracted by irrelevant context. In International Conference on Machine Learning, pp. 31210–31227. PMLR, 2023.
  16. Koustuv Sinha, Prasanna Parthasarathi, Joelle Pineau, and Adina Williams. Unnatural language inference. In Proceedings of the 59th annual meeting of the Association for Computational Linguistics and the 11th international joint conference on natural language processing (Volume 1: Long Papers), pp. 7329–7346, 2021. [CrossRef]
  17. Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on machine learning research, 2023.
  18. Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A Smith, and Yejin Choi. Dataset cartography: Mapping and diagnosing datasets with training dynamics. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 9275–9293, 2020.
  19. Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. Beir: A heterogenous benchmark for zero-shot evaluation of information retrieval models. arXiv preprint arXiv:2104.08663, 2021.
  20. Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
  21. Shaobo Wang, Cong Wang, Wenjie Fu, Yue Min, Mingquan Feng, Isabel Guan, Xuming Hu, Conghui He, Cunxiang Wang, Kexin Yang, et al. Rethinking llm evaluation: Can we evaluate llms with 200x less data? arXiv preprint arXiv:2510.10457, 2025.
  22. Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, et al. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. Advances in Neural Information Processing Systems, 37:95266–95290, 2024. [CrossRef]
  23. Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837, 2022. [CrossRef]
  24. Sheldon Yu, Yuxin Xiong, Junda Wu, Xintong Li, Tong Yu, Xiang Chen, Ritwik Sinha, Jingbo Shang, and Julian McAuley. Explainable chain-of-thought reasoning: An empirical analysis on state-aware reasoning dynamics. arXiv preprint arXiv:2509.00190, 2025.
  25. Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. Automatic chain of thought prompting in large language models. arXiv preprint arXiv:2210.03493, 2022.
  26. Zicheng Zhang, Xiangyu Zhao, Xinyu Fang, Chunyi Li, Xiaohong Liu, Xiongkuo Min, Haodong Duan, Kai Chen, and Guangtao Zhai. Redundancy principles for mllms benchmarks. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 12492–12504, 2025. [CrossRef]
1
It is worth noting that the LLM exhibits a minor reasoning hallucination in Question B (applying a linear proportion to exponential decay). However, this precisely highlights the robustness of our approach: the embedding alignment relies on the structural and conceptual vocabulary (Lorentz transformations, time dilation) extracted by the CoT, rather than the strict arithmetic correctness of the final step.
2
It should be noted that in Question D, the LLM utilizes a simplified geometric assumption regarding the spot area on a sphere versus a projected cross-sectional disk (ignoring limb darkening effects). Nevertheless, this pedagogical simplification does not hinder the embedding process. The core contribution of the CoT remains intact: successfully shifting the semantic representation from the misleading lexical surface to the correct physical domain (macroscopic geometric flux rather than quantum states).
Figure 2. Overview of CoT-Core. Our framework unrolls LLM reasoning trajectories, embeds them, and applies isomorphic clustering to select coresets. In contrast to raw text embeddings which fall into lexical traps (falsely grouping Q1 & Q2), CoT-Core effectively aligns lexically disparate problems that share an identical latent logical structure (Q1 & Q3).
Figure 2. Overview of CoT-Core. Our framework unrolls LLM reasoning trajectories, embeds them, and applies isomorphic clustering to select coresets. In contrast to raw text embeddings which fall into lexical traps (falsely grouping Q1 & Q2), CoT-Core effectively aligns lexically disparate problems that share an identical latent logical structure (Q1 & Q3).
Preprints 223734 g002
Figure 3. t-SNE of GPQA embeddings. (Left) Question-only embeddings are hindered by surface lexical traps and domain-specific vocabulary. (Right) CoT-Core reorganizes the latent space via logical isomorphisms, capturing underlying reasoning structures that transcend superficial text similarities.
Figure 3. t-SNE of GPQA embeddings. (Left) Question-only embeddings are hindered by surface lexical traps and domain-specific vocabulary. (Right) CoT-Core reorganizes the latent space via logical isomorphisms, capturing underlying reasoning structures that transcend superficial text similarities.
Preprints 223734 g003
Table 1. Performance comparison at N = 25 using the SA estimator. Best performance among training-free methods is highlighted in bold. History-dependent methods are shown for reference only. Note that results at the 1% budget for GPQA are omitted, as this strictly equates to a single instance, rendering coreset evaluation statistically meaningless.
Table 1. Performance comparison at N = 25 using the SA estimator. Best performance among training-free methods is highlighted in bold. History-dependent methods are shown for reference only. Note that results at the 1% budget for GPQA are omitted, as this strictly equates to a single instance, rendering coreset evaluation statistically meaningless.
Benchmark Method Sampling Ratio
1%
MAE (↓) / Sim (↑)
5%
MAE (↓) / Sim (↑)
10%
MAE (↓) / Sim (↑)
15%
MAE (↓) / Sim (↑)
20%
MAE (↓) / Sim (↑)
Training-Free Methods (Zero-Shot)
GSM8K Random 0.0890 / 0.9044 0.0382 / 0.9747 0.0246 / 0.9851 0.0187 / 0.9897 0.0142 / 0.9920
Question-Emb K-Means 0.0875 / 0.8951 0.0321 / 0.9743 0.0241 / 0.9797 0.0175 / 0.9871 0.0133 / 0.9931
k-Center Greedy 0.0683 / 0.9118 0.0349 / 0.9738 0.0283 / 0.9847 0.0218 / 0.9897 0.0175 / 0.9925
CoT-Feature (Ours) 0.0755 / 0.8884 0.0331 / 0.9696 0.0209 / 0.9853 0.0197 / 0.9884 0.0154 / 0.9918
Trainable Methods (Require Historical Data, N=25)
Correctness K-Means 0.0596 / 0.9129 0.0347 / 0.9700 0.0259 / 0.9770 0.0231 / 0.9838 0.0198 / 0.9886
IRT 0.0664 / 0.9172 0.0308 / 0.9649 0.0196 / 0.9817 0.0175 / 0.9852 0.0145 / 0.9879
Training-Free Methods (Zero-Shot)
MMLU Random 0.0330 / 0.9489 0.0149 / 0.9880 0.0102 / 0.9935 0.0079 / 0.9957 0.0063 / 0.9969
Question-Emb K-Means 0.0298 / 0.9484 0.0215 / 0.9877 0.0154 / 0.9923 0.0203 / 0.9943 0.0252 / 0.9958
k-Center Greedy 0.0477 / 0.9461 0.0530 / 0.9867 0.0393 / 0.9909 0.0360 / 0.9942 0.0317 / 0.9952
CoT-Feature (Ours) 0.0324 / 0.9353 0.0118 / 0.9881 0.0081 / 0.9941 0.0081 / 0.9964 0.0060 / 0.9971
Trainable Methods (Require Historical Data, N=25)
Correctness K-Means 0.0304 / 0.9616 0.0203 / 0.9754 0.0181 / 0.9773 0.0169 / 0.9781 0.0160 / 0.9785
IRT 0.0253 / 0.9589 0.0112 / 0.9895 0.0081 / 0.9938 0.0068 / 0.9966 0.0056 / 0.9969
Training-Free Methods (Zero-Shot)
MMLU-Pro Random 0.0311 / 0.9556 0.0132 / 0.9899 0.0093 / 0.9941 0.0070 / 0.9956 0.0058 / 0.9967
Question-Emb K-Means 0.0300 / 0.9579 0.0133 / 0.9890 0.0111 / 0.9947 0.0116 / 0.9957 0.0100 / 0.9969
k-Center Greedy 0.0369 / 0.9524 0.0432 / 0.9857 0.0334 / 0.9923 0.0306 / 0.9936 0.0272 / 0.9957
CoT-Feature (Ours) 0.0296 / 0.9682 0.0124 / 0.9893 0.0082 / 0.9941 0.0072 / 0.9965 0.0055 / 0.9969
Trainable Methods (Require Historical Data, N=25)
Correctness K-Means 0.0257 / 0.9789 0.0144 / 0.9902 0.0123 / 0.9929 0.0107 / 0.9936 0.0095 / 0.9941
IRT 0.0379 / 0.9540 0.0151 / 0.9891 0.0111 / 0.9930 0.0092 / 0.9943 0.0073 / 0.9962
Training-Free Methods (Zero-Shot)
GPQA Random - / - 0.1252 / 0.2851 0.0797 / 0.4480 0.0635 / 0.5579 0.0536 / 0.6022
Question-Emb K-Means - / - 0.1098 / 0.2856 0.0859 / 0.3656 0.0615 / 0.5338 0.0522 / 0.6263
k-Center Greedy - / - 0.1232 / 0.2197 0.0821 / 0.4036 0.0608 / 0.5355 0.0565 / 0.6151
CoT-Feature (Ours) - / - 0.1044 / 0.4334 0.0731 / 0.5839 0.0601 / 0.5646 0.0516 / 0.6013
Trainable Methods (Require Historical Data, N=25)
Correctness K-Means - / - 0.0936 / 0.3125 0.0743 / 0.3435 0.0660 / 0.4498 0.0570 / 0.4770
IRT - / - 0.1054 / 0.2576 0.0723 / 0.3806 0.0552 / 0.6081 0.0513 / 0.6176
Table 2. Performance comparison at N = 25 using the GP-IRT estimator. Best performance among training-free methods is highlighted in bold. History-dependent methods are shown for reference only. Note that results at the 1% budget for GPQA are omitted, as this strictly equates to a single instance, rendering coreset evaluation statistically meaningless.
Table 2. Performance comparison at N = 25 using the GP-IRT estimator. Best performance among training-free methods is highlighted in bold. History-dependent methods are shown for reference only. Note that results at the 1% budget for GPQA are omitted, as this strictly equates to a single instance, rendering coreset evaluation statistically meaningless.
Benchmark Method Sampling Ratio
1%
MAE (↓) / Sim (↑)
5%
MAE (↓) / Sim (↑)
10%
MAE (↓) / Sim (↑)
15%
MAE (↓) / Sim (↑)
20%
MAE (↓) / Sim (↑)
Training-Free Methods (Zero-Shot)
GSM8K Random 0.0827 / 0.8846 0.0297 / 0.9701 0.0202 / 0.9843 0.0158 / 0.9893 0.0131 / 0.9919
Question-Emb K-Means 0.0975 / 0.8443 0.0296 / 0.9715 0.0206 / 0.9776 0.0158 / 0.9866 0.0123 / 0.9928
k-Center Greedy 0.1016 / 0.8597 0.0289 / 0.9707 0.0216 / 0.9830 0.0166 / 0.9895 0.0136 / 0.9929
CoT-Feature (Ours) 0.0800 / 0.8700 0.0290 / 0.9642 0.0197 / 0.9844 0.0166 / 0.9882 0.0135 / 0.9914
Trainable Methods (Require Historical Data, N=25)
Correctness K-Means 0.0830 / 0.8608 0.0281 / 0.9691 0.0183 / 0.9830 0.0156 / 0.9886 0.0121 / 0.9940
IRT 0.0789 / 0.8957 0.0308 / 0.9659 0.0177 / 0.9838 0.0158 / 0.9877 0.0134 / 0.9909
Training-Free Methods (Zero-Shot)
MMLU Random 0.0212 / 0.9639 0.0116 / 0.9897 0.0089 / 0.9938 0.0072 / 0.9957 0.0060 / 0.9968
Question-Emb K-Means 0.0225 / 0.9606 0.0120 / 0.9891 0.0105 / 0.9933 0.0130 / 0.9947 0.0165 / 0.9961
k-Center Greedy 0.0198 / 0.9695 0.0190 / 0.9893 0.0191 / 0.9929 0.0209 / 0.9951 0.0205 / 0.9957
CoT-Feature (Ours) 0.0223 / 0.9577 0.0111 / 0.9901 0.0076 / 0.9948 0.0068 / 0.9962 0.0059 / 0.9969
Trainable Methods (Require Historical Data, N=25)
Correctness K-Means 0.0222 / 0.9648 0.0142 / 0.9856 0.0121 / 0.9864 0.0113 / 0.9850 0.0110 / 0.9836
IRT 0.0209 / 0.9651 0.0131 / 0.9883 0.0095 / 0.9932 0.0073 / 0.9961 0.0066 / 0.9968
Training-Free Methods (Zero-Shot)
MMLU-Pro Random 0.0250 / 0.9598 0.0104 / 0.9904 0.0074 / 0.9944 0.0060 / 0.9955 0.0051 / 0.9966
Question-Emb K-Means 0.0244 / 0.9547 0.0103 / 0.9895 0.0081 / 0.9949 0.0083 / 0.9955 0.0074 / 0.9968
k-Center Greedy 0.0240 / 0.9670 0.0146 / 0.9896 0.0149 / 0.9935 0.0162 / 0.9943 0.0162 / 0.9955
CoT-Feature (Ours) 0.0240 / 0.9724 0.0102 / 0.9901 0.0069 / 0.9943 0.0058 / 0.9963 0.0047 / 0.9968
Trainable Methods (Require Historical Data, N=25)
Correctness K-Means 0.0234 / 0.9747 0.0091 / 0.9925 0.0081 / 0.9948 0.0075 / 0.9957 0.0070 / 0.9958
IRT 0.0257 / 0.9549 0.0109 / 0.9903 0.0095 / 0.9943 0.0085 / 0.9960 0.0075 / 0.9971
Training-Free Methods (Zero-Shot)
GPQA Random - / - 0.1019 / 0.2807 0.0679 / 0.4782 0.0568 / 0.5842 0.0497 / 0.6248
Question-Emb K-Means - / - 0.0887 / 0.2759 0.0703 / 0.3900 0.0552 / 0.5704 0.0482 / 0.6388
k-Center Greedy - / - 0.0973 / 0.2351 0.0701 / 0.4334 0.0551 / 0.5547 0.0519 / 0.6335
CoT-Feature (Ours) - / - 0.0922 / 0.5123 0.0649 / 0.6206 0.0545 / 0.5812 0.0482 / 0.6318
Trainable Methods (Require Historical Data, N=25)
Correctness K-Means - / - 0.0728 / 0.3970 0.0597 / 0.4130 0.0559 / 0.5022 0.0503 / 0.5293
IRT - / - 0.0887 / 0.2545 0.0622 / 0.4248 0.0500 / 0.6402 0.0473 / 0.6475
Table 4. Ablation on Hierarchical Cleaning (GPQA). Cleaned CoT achieves the optimal balance.
Table 4. Ablation on Hierarchical Cleaning (GPQA). Cleaned CoT achieves the optimal balance.
Abstraction Level GPQA Budget (MAE ↓ / Sim ↑)
5% 10% 20%
Raw CoT 0.1044 / 0.4334 0.0731 / 0.5839 0.0516 / 0.6013
Cleaned CoT 0.1007 / 0.5041 0.0666 / 0.4970 0.0442 / 0.6438
Abstract CoT 0.1206 / 0.3561 0.0831 / 0.4506 0.0518 / 0.5714
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings