Submitted:
17 July 2026
Posted:
20 July 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction

2. Related Work
LLM Benchmark Evaluation and Redundancy
Geometric and Training-free Methods.
Performance-Based and Psychometric Methods.
CoT-Aware Representations for Coreset Selection
3. Methodology
3.1. Problem Formulation
3.2. CoT-Core Framework Overview
3.3. Stage 1: CoT Trajectory Generation
3.4. Stage 2: Contextualized Trajectory Embedding
3.5. Stage 3: Isomorphic Clustering and Selection
4. Experiments
4.1. Experimental Setup
4.2. Main Results
4.3. Analysis
The Simplicity Floor and Informational Boundaries.
Exposing Cognitive Isomorphism via Logical Trajectories
4.4. Ablation Studies
| Generator | 1% Budget | 10% Budget | 20% Budget |
|---|---|---|---|
| LLaMA-3-1B | |||
| Mistral-8B | |||
| Phi-4 | |||
| Qwen-72B |
5. Conclusions
Appendix A. Implementation Details
Appendix A.1. Model Configurations
Appendix A.2. Prompt Templates for Hierarchical Cleaning


Appendix B. Case Study: Full Problem Descriptions and Reasoning Trajectories
Appendix B.1. Case 1: Uncovering Logical Isomorphism (False-Negative Correction)
Question A (Experimental Resolution):


Question B (Kinematic State):


Analysis of the Isomorphism:
- 1.
- Relativistic Kinematics: Calculating or utilizing the Lorentz factor ().
- 2.
- Time Dilation: Linking proper lifetime to dilated lifetime ().
- 3.
- Mean Decay Length: Formulating the distance traveled before decay ().
- 4.
- Threshold Probability: Correlating the traveled distance with the survival/decay ratio.
Appendix B.2. Case 2: Escaping Lexical Traps (False-Positive Correction)
Question C (Quantum States):


Question D (Photometric Flux):


Analysis of the Orthogonality:
- Question C (Quantum Statistical Mechanics): The reasoning trajectory is entirely driven by atomic physics. It invokes the Boltzmann distribution to calculate microscopic population states, generating specialized tokens like energy differences (), Planck’s constant (h), and exponential decay functions dependent on temperature ().
- Question D (Macroscopic Photometry): The reasoning trajectory operates on geometric and thermodynamic scales. It relies on the Stefan-Boltzmann law (where flux is proportional to ) and geometrical area ratios (projected disk areas and filling factors) to calculate macroscopic light curve variations.
Appendix C. Extended Main Results
| Benchmark | Method | Sampling Ratio | ||||
|---|---|---|---|---|---|---|
| 1% MAE (↓) / Sim (↑) |
5% MAE (↓) / Sim (↑) |
10% MAE (↓) / Sim (↑) |
15% MAE (↓) / Sim (↑) |
20% MAE (↓) / Sim (↑) |
||
| Training-Free Methods (Zero-Shot) | ||||||
| GSM8K | Random | 0.0890 / 0.9044 | 0.0382 / 0.9747 | 0.0246 / 0.9851 | 0.0187 / 0.9897 | 0.0142 / 0.9920 |
| Question-Emb K-Means | 0.0875 / 0.8951 | 0.0321 / 0.9743 | 0.0241 / 0.9797 | 0.0175 / 0.9871 | 0.0133 / 0.9931 | |
| k-Center Greedy | 0.0683 / 0.9118 | 0.0349 / 0.9738 | 0.0283 / 0.9847 | 0.0218 / 0.9897 | 0.0175 / 0.9925 | |
| CoT-Feature (Ours) | 0.0755 / 0.8884 | 0.0331 / 0.9696 | 0.0209 / 0.9853 | 0.0197 / 0.9884 | 0.0154 / 0.9918 | |
| Trainable Methods (Require Historical Data, N=50) | ||||||
| Correctness K-Means | 0.0623 / 0.9130 | 0.0339 / 0.9748 | 0.0256 / 0.9816 | 0.0230 / 0.9848 | 0.0195 / 0.9887 | |
| IRT | 0.0670 / 0.9053 | 0.0329 / 0.9754 | 0.0241 / 0.9833 | 0.0169 / 0.9908 | 0.0143 / 0.9907 | |
| Training-Free Methods (Zero-Shot) | ||||||
| MMLU | Random | 0.0330 / 0.9489 | 0.0149 / 0.9880 | 0.0102 / 0.9935 | 0.0079 / 0.9957 | 0.0063 / 0.9969 |
| Question-Emb K-Means | 0.0298 / 0.9484 | 0.0215 / 0.9877 | 0.0154 / 0.9923 | 0.0203 / 0.9943 | 0.0252 / 0.9958 | |
| k-Center Greedy | 0.0477 / 0.9461 | 0.0530 / 0.9867 | 0.0393 / 0.9909 | 0.0360 / 0.9942 | 0.0317 / 0.9952 | |
| CoT-Feature (Ours) | 0.0324 / 0.9353 | 0.0118 / 0.9881 | 0.0081 / 0.9941 | 0.0081 / 0.9964 | 0.0060 / 0.9971 | |
| Trainable Methods (Require Historical Data, N=50) | ||||||
| Correctness K-Means | 0.0272 / 0.9638 | 0.0191 / 0.9789 | 0.0161 / 0.9825 | 0.0152 / 0.9838 | 0.0144 / 0.9851 | |
| IRT | 0.0225 / 0.9588 | 0.0113 / 0.9896 | 0.0083 / 0.9938 | 0.0069 / 0.9960 | 0.0065 / 0.9968 | |
| Training-Free Methods (Zero-Shot) | ||||||
| MMLU-Pro | Random | 0.0311 / 0.9556 | 0.0132 / 0.9899 | 0.0093 / 0.9941 | 0.0070 / 0.9956 | 0.0058 / 0.9967 |
| Question-Emb K-Means | 0.0300 / 0.9579 | 0.0133 / 0.9890 | 0.0111 / 0.9947 | 0.0116 / 0.9957 | 0.0100 / 0.9969 | |
| k-Center Greedy | 0.0369 / 0.9524 | 0.0432 / 0.9857 | 0.0334 / 0.9923 | 0.0306 / 0.9936 | 0.0272 / 0.9957 | |
| CoT-Feature (Ours) | 0.0296 / 0.9682 | 0.0124 / 0.9893 | 0.0082 / 0.9941 | 0.0072 / 0.9965 | 0.0055 / 0.9969 | |
| Trainable Methods (Require Historical Data, N=50) | ||||||
| Correctness K-Means | 0.0283 / 0.9769 | 0.0172 / 0.9903 | 0.0152 / 0.9933 | 0.0131 / 0.9942 | 0.0113 / 0.9947 | |
| IRT | 0.0412 / 0.9642 | 0.0208 / 0.9902 | 0.0151 / 0.9938 | 0.0122 / 0.9955 | 0.0102 / 0.9963 | |
| Training-Free Methods (Zero-Shot) | ||||||
| GPQA | Random | - / - | 0.1252 / 0.2851 | 0.0797 / 0.4480 | 0.0635 / 0.5579 | 0.0536 / 0.6022 |
| Question-Emb K-Means | - / - | 0.1098 / 0.2856 | 0.0859 / 0.3656 | 0.0615 / 0.5338 | 0.0522 / 0.6263 | |
| k-Center Greedy | - / - | 0.1232 / 0.2197 | 0.0821 / 0.4036 | 0.0608 / 0.5355 | 0.0565 / 0.6151 | |
| CoT-Feature (Ours) | - / - | 0.1044 / 0.4334 | 0.0731 / 0.5839 | 0.0601 / 0.5646 | 0.0516 / 0.6013 | |
| Trainable Methods (Require Historical Data, N=50) | ||||||
| Correctness K-Means | - / - | 0.1269 / 0.2524 | 0.0942 / 0.3297 | 0.0771 / 0.4398 | 0.0666 / 0.4601 | |
| IRT | - / - | 0.1047 / 0.3878 | 0.0780 / 0.4071 | 0.0656 / 0.5104 | 0.0540 / 0.5748 | |
| Benchmark | Method | Sampling Ratio | ||||
|---|---|---|---|---|---|---|
| 1% MAE (↓) / Sim (↑) |
5% MAE (↓) / Sim (↑) |
10% MAE (↓) / Sim (↑) |
15% MAE (↓) / Sim (↑) |
20% MAE (↓) / Sim (↑) |
||
| Training-Free Methods (Zero-Shot) | ||||||
| GSM8K | Random | 0.0933 / 0.8809 | 0.0325 / 0.9670 | 0.0203 / 0.9839 | 0.0157 / 0.9894 | 0.0129 / 0.9917 |
| Question-Emb K-Means | 0.0781 / 0.8796 | 0.0310 / 0.9679 | 0.0204 / 0.9815 | 0.0161 / 0.9871 | 0.0123 / 0.9930 | |
| k-Center Greedy | 0.0972 / 0.8676 | 0.0333 / 0.9658 | 0.0214 / 0.9833 | 0.0167 / 0.9899 | 0.0143 / 0.9929 | |
| CoT-Feature (Ours) | 0.0849 / 0.8713 | 0.0297 / 0.9592 | 0.0192 / 0.9856 | 0.0163 / 0.9893 | 0.0133 / 0.9918 | |
| Trainable Methods (Require Historical Data, N=50) | ||||||
| Correctness K-Means | 0.0698 / 0.9103 | 0.0292 / 0.9718 | 0.0183 / 0.9855 | 0.0152 / 0.9887 | 0.0140 / 0.9926 | |
| IRT | 0.0805 / 0.8858 | 0.0278 / 0.9642 | 0.0191 / 0.9844 | 0.0149 / 0.9910 | 0.0128 / 0.9924 | |
| Training-Free Methods (Zero-Shot) | ||||||
| MMLU | Random | 0.0207 / 0.9658 | 0.0106 / 0.9908 | 0.0082 / 0.9945 | 0.0067 / 0.9961 | 0.0056 / 0.9971 |
| Question-Emb K-Means | 0.0220 / 0.9597 | 0.0123 / 0.9899 | 0.0108 / 0.9934 | 0.0146 / 0.9948 | 0.0192 / 0.9961 | |
| k-Center Greedy | 0.0208 / 0.9651 | 0.0239 / 0.9892 | 0.0244 / 0.9925 | 0.0257 / 0.9945 | 0.0244 / 0.9955 | |
| CoT-Feature (Ours) | 0.0233 / 0.9535 | 0.0100 / 0.9900 | 0.0068 / 0.9949 | 0.0066 / 0.9965 | 0.0056 / 0.9971 | |
| Trainable Methods (Require Historical Data, N=50) | ||||||
| Correctness K-Means | 0.0207 / 0.9723 | 0.0141 / 0.9868 | 0.0129 / 0.9887 | 0.0125 / 0.9892 | 0.0121 / 0.9894 | |
| IRT | 0.0213 / 0.9655 | 0.0105 / 0.9911 | 0.0077 / 0.9949 | 0.0066 / 0.9964 | 0.0065 / 0.9971 | |
| Training-Free Methods (Zero-Shot) | ||||||
| MMLU-Pro | Random | 0.0239 / 0.9624 | 0.0097 / 0.9906 | 0.0072 / 0.9943 | 0.0058 / 0.9957 | 0.0049 / 0.9967 |
| Question-Emb K-Means | 0.0238 / 0.9593 | 0.0098 / 0.9899 | 0.0075 / 0.9944 | 0.0073 / 0.9957 | 0.0069 / 0.9968 | |
| k-Center Greedy | 0.0249 / 0.9582 | 0.0146 / 0.9895 | 0.0149 / 0.9936 | 0.0163 / 0.9947 | 0.0163 / 0.9956 | |
| CoT-Feature (Ours) | 0.0226 / 0.9691 | 0.0105 / 0.9899 | 0.0070 / 0.9939 | 0.0059 / 0.9959 | 0.0047 / 0.9966 | |
| Trainable Methods (Require Historical Data, N=50) | ||||||
| Correctness K-Means | 0.0226 / 0.9668 | 0.0102 / 0.9909 | 0.0103 / 0.9946 | 0.0101 / 0.9951 | 0.0093 / 0.9954 | |
| IRT | 0.0278 / 0.9518 | 0.0100 / 0.9914 | 0.0083 / 0.9947 | 0.0079 / 0.9958 | 0.0076 / 0.9966 | |
| Training-Free Methods (Zero-Shot) | ||||||
| GPQA | Random | - / - | 0.1281 / 0.2834 | 0.0596 / 0.4492 | 0.0414 / 0.6098 | 0.0352 / 0.6768 |
| Question-Emb K-Means | - / - | 0.1370 / 0.2460 | 0.0601 / 0.4870 | 0.0379 / 0.6004 | 0.0362 / 0.6651 | |
| k-Center Greedy | - / - | 0.1172 / 0.3289 | 0.0612 / 0.4532 | 0.0434 / 0.5656 | 0.0364 / 0.6267 | |
| CoT-Feature (Ours) | - / - | 0.1242 / 0.3571 | 0.0612 / 0.5275 | 0.0412 / 0.5904 | 0.0362 / 0.6616 | |
| Trainable Methods (Require Historical Data, N=50) | ||||||
| Correctness K-Means | - / - | 0.1028 / 0.3905 | 0.0491 / 0.4433 | 0.0382 / 0.5559 | 0.0373 / 0.6144 | |
| IRT | - / - | 0.1117 / 0.3240 | 0.0505 / 0.4854 | 0.0411 / 0.6455 | 0.0352 / 0.6663 | |
| Benchmark | Method | Sampling Ratio | ||||
|---|---|---|---|---|---|---|
| 1% MAE (↓) / Sim (↑) |
5% MAE (↓) / Sim (↑) |
10% MAE (↓) / Sim (↑) |
15% MAE (↓) / Sim (↑) |
20% MAE (↓) / Sim (↑) |
||
| Training-Free Methods (Zero-Shot) | ||||||
| GSM8K | Random | 0.0890 / 0.9044 | 0.0382 / 0.9747 | 0.0246 / 0.9851 | 0.0187 / 0.9897 | 0.0142 / 0.9920 |
| Question-Emb K-Means | 0.0875 / 0.8951 | 0.0321 / 0.9743 | 0.0241 / 0.9797 | 0.0175 / 0.9871 | 0.0133 / 0.9931 | |
| k-Center Greedy | 0.0683 / 0.9118 | 0.0349 / 0.9738 | 0.0283 / 0.9847 | 0.0218 / 0.9897 | 0.0175 / 0.9925 | |
| CoT-Feature (Ours) | 0.0755 / 0.8884 | 0.0331 / 0.9696 | 0.0209 / 0.9853 | 0.0197 / 0.9884 | 0.0154 / 0.9918 | |
| Trainable Methods (Require Historical Data, N=100) | ||||||
| Correctness K-Means | 0.0522 / 0.9226 | 0.0344 / 0.9699 | 0.0299 / 0.9843 | 0.0266 / 0.9865 | 0.0233 / 0.9903 | |
| IRT | 0.0650 / 0.9049 | 0.0295 / 0.9734 | 0.0216 / 0.9797 | 0.0168 / 0.9880 | 0.0141 / 0.9893 | |
| Training-Free Methods (Zero-Shot) | ||||||
| MMLU | Random | 0.0330 / 0.9489 | 0.0149 / 0.9880 | 0.0102 / 0.9935 | 0.0079 / 0.9957 | 0.0063 / 0.9969 |
| Question-Emb K-Means | 0.0298 / 0.9484 | 0.0215 / 0.9877 | 0.0154 / 0.9923 | 0.0203 / 0.9943 | 0.0252 / 0.9958 | |
| k-Center Greedy | 0.0477 / 0.9461 | 0.0530 / 0.9867 | 0.0393 / 0.9909 | 0.0360 / 0.9942 | 0.0317 / 0.9952 | |
| CoT-Feature (Ours) | 0.0324 / 0.9353 | 0.0118 / 0.9881 | 0.0081 / 0.9941 | 0.0081 / 0.9964 | 0.0060 / 0.9971 | |
| Trainable Methods (Require Historical Data, N=100) | ||||||
| Correctness K-Means | 0.0184 / 0.9762 | 0.0111 / 0.9906 | 0.0100 / 0.9917 | 0.0091 / 0.9927 | 0.0082 / 0.9936 | |
| IRT | 0.0222 / 0.9565 | 0.0105 / 0.9895 | 0.0089 / 0.9932 | 0.0071 / 0.9950 | 0.0063 / 0.9962 | |
| Training-Free Methods (Zero-Shot) | ||||||
| MMLU-Pro | Random | 0.0311 / 0.9556 | 0.0132 / 0.9899 | 0.0093 / 0.9941 | 0.0070 / 0.9956 | 0.0058 / 0.9967 |
| Question-Emb K-Means | 0.0300 / 0.9579 | 0.0133 / 0.9890 | 0.0111 / 0.9947 | 0.0116 / 0.9957 | 0.0100 / 0.9969 | |
| k-Center Greedy | 0.0369 / 0.9524 | 0.0432 / 0.9857 | 0.0334 / 0.9923 | 0.0306 / 0.9936 | 0.0272 / 0.9957 | |
| CoT-Feature (Ours) | 0.0296 / 0.9682 | 0.0124 / 0.9893 | 0.0082 / 0.9941 | 0.0072 / 0.9965 | 0.0055 / 0.9969 | |
| Trainable Methods (Require Historical Data, N=100) | ||||||
| Correctness K-Means | 0.0289 / 0.9815 | 0.0219 / 0.9919 | 0.0192 / 0.9937 | 0.0175 / 0.9940 | 0.0157 / 0.9948 | |
| IRT | 0.0274 / 0.9561 | 0.0111 / 0.9923 | 0.0085 / 0.9938 | 0.0064 / 0.9953 | 0.0058 / 0.9961 | |
| Training-Free Methods (Zero-Shot) | ||||||
| GPQA | Random | - / - | 0.1252 / 0.2851 | 0.0797 / 0.4480 | 0.0635 / 0.5579 | 0.0536 / 0.6022 |
| Question-Emb K-Means | - / - | 0.1098 / 0.2856 | 0.0859 / 0.3656 | 0.0615 / 0.5338 | 0.0522 / 0.6263 | |
| k-Center Greedy | - / - | 0.1232 / 0.2197 | 0.0821 / 0.4036 | 0.0608 / 0.5355 | 0.0565 / 0.6151 | |
| CoT-Feature (Ours) | - / - | 0.1044 / 0.4334 | 0.0731 / 0.5839 | 0.0601 / 0.5646 | 0.0516 / 0.6013 | |
| Trainable Methods (Require Historical Data, N=100) | ||||||
| Correctness K-Means | - / - | 0.1124 / 0.4882 | 0.0925 / 0.4670 | 0.0792 / 0.5425 | 0.0742 / 0.5639 | |
| IRT | - / - | 0.1125 / 0.3425 | 0.0713 / 0.4955 | 0.0511 / 0.6022 | 0.0466 / 0.6193 | |
| Benchmark | Method | Sampling Ratio | ||||
|---|---|---|---|---|---|---|
| 1% MAE (↓) / Sim (↑) |
5% MAE (↓) / Sim (↑) |
10% MAE (↓) / Sim (↑) |
15% MAE (↓) / Sim (↑) |
20% MAE (↓) / Sim (↑) |
||
| Training-Free Methods (Zero-Shot) | ||||||
| GSM8K | Random | 0.0881 / 0.8824 | 0.0292 / 0.9732 | 0.0200 / 0.9856 | 0.0161 / 0.9899 | 0.0130 / 0.9921 |
| Question-Emb K-Means | 0.0899 / 0.8723 | 0.0276 / 0.9723 | 0.0210 / 0.9804 | 0.0158 / 0.9877 | 0.0125 / 0.9931 | |
| k-Center Greedy | 0.0831 / 0.8703 | 0.0279 / 0.9725 | 0.0216 / 0.9832 | 0.0177 / 0.9897 | 0.0149 / 0.9925 | |
| CoT-Feature (Ours) | 0.0818 / 0.8751 | 0.0285 / 0.9664 | 0.0188 / 0.9860 | 0.0162 / 0.9889 | 0.0135 / 0.9917 | |
| Trainable Methods (Require Historical Data, N=100) | ||||||
| Correctness K-Means | 0.0741 / 0.9099 | 0.0263 / 0.9727 | 0.0205 / 0.9893 | 0.0195 / 0.9910 | 0.0177 / 0.9934 | |
| IRT | 0.0651 / 0.8984 | 0.0260 / 0.9748 | 0.0195 / 0.9809 | 0.0152 / 0.9894 | 0.0128 / 0.9902 | |
| Training-Free Methods (Zero-Shot) | ||||||
| MMLU | Random | 0.0199 / 0.9662 | 0.0098 / 0.9909 | 0.0076 / 0.9944 | 0.0062 / 0.9960 | 0.0053 / 0.9970 |
| Question-Emb K-Means | 0.0206 / 0.9626 | 0.0107 / 0.9895 | 0.0086 / 0.9940 | 0.0103 / 0.9949 | 0.0132 / 0.9962 | |
| k-Center Greedy | 0.0209 / 0.9684 | 0.0129 / 0.9902 | 0.0141 / 0.9935 | 0.0161 / 0.9953 | 0.0164 / 0.9959 | |
| CoT-Feature (Ours) | 0.0218 / 0.9631 | 0.0096 / 0.9911 | 0.0069 / 0.9952 | 0.0061 / 0.9964 | 0.0054 / 0.9971 | |
| Trainable Methods (Require Historical Data, N=100) | ||||||
| Correctness K-Means | 0.0190 / 0.9715 | 0.0100 / 0.9906 | 0.0083 / 0.9932 | 0.0075 / 0.9941 | 0.0066 / 0.9948 | |
| IRT | 0.0188 / 0.9684 | 0.0099 / 0.9913 | 0.0076 / 0.9945 | 0.0064 / 0.9957 | 0.0057 / 0.9966 | |
| Training-Free Methods (Zero-Shot) | ||||||
| MMLU-Pro | Random | 0.0229 / 0.9629 | 0.0094 / 0.9907 | 0.0066 / 0.9947 | 0.0053 / 0.9960 | 0.0045 / 0.9969 |
| Question-Emb K-Means | 0.0227 / 0.9581 | 0.0093 / 0.9899 | 0.0069 / 0.9946 | 0.0061 / 0.9959 | 0.0054 / 0.9970 | |
| k-Center Greedy | 0.0229 / 0.9640 | 0.0102 / 0.9900 | 0.0086 / 0.9941 | 0.0096 / 0.9949 | 0.0099 / 0.9958 | |
| CoT-Feature (Ours) | 0.0209 / 0.9706 | 0.0099 / 0.9903 | 0.0065 / 0.9944 | 0.0051 / 0.9965 | 0.0043 / 0.9972 | |
| Trainable Methods (Require Historical Data, N=100) | ||||||
| Correctness K-Means | 0.0295 / 0.9550 | 0.0103 / 0.9927 | 0.0091 / 0.9956 | 0.0085 / 0.9970 | 0.0086 / 0.9970 | |
| IRT | 0.0255 / 0.9520 | 0.0088 / 0.9929 | 0.0066 / 0.9948 | 0.0051 / 0.9962 | 0.0046 / 0.9971 | |
| Training-Free Methods (Zero-Shot) | ||||||
| GPQA | Random | - / - | 0.1022 / 0.2447 | 0.0549 / 0.4685 | 0.0415 / 0.5990 | 0.0377 / 0.6439 |
| Question-Emb K-Means | - / - | 0.0990 / 0.2167 | 0.0557 / 0.3998 | 0.0412 / 0.5598 | 0.0376 / 0.6586 | |
| k-Center Greedy | - / - | 0.1040 / 0.1707 | 0.0565 / 0.3994 | 0.0446 / 0.5275 | 0.0411 / 0.5964 | |
| CoT-Feature (Ours) | - / - | 0.0836 / 0.3995 | 0.0482 / 0.6417 | 0.0380 / 0.6166 | 0.0369 / 0.6505 | |
| Trainable Methods (Require Historical Data, N=100) | ||||||
| Correctness K-Means | - / - | 0.0653 / 0.4379 | 0.0443 / 0.5324 | 0.0415 / 0.5921 | 0.0431 / 0.6339 | |
| IRT | - / - | 0.0944 / 0.2877 | 0.0490 / 0.5302 | 0.0360 / 0.6203 | 0.0330 / 0.6608 | |
| Benchmark | Method | Sampling Ratio | ||||
|---|---|---|---|---|---|---|
| 1% MAE (↓) / Sim (↑) |
5% MAE (↓) / Sim (↑) |
10% MAE (↓) / Sim (↑) |
15% MAE (↓) / Sim (↑) |
20% MAE (↓) / Sim (↑) |
||
| Training-Free Methods (Zero-Shot) | ||||||
| GSM8K | Random | 0.0890 / 0.9044 | 0.0382 / 0.9747 | 0.0246 / 0.9851 | 0.0187 / 0.9897 | 0.0142 / 0.9920 |
| Question-Emb K-Means | 0.0875 / 0.8951 | 0.0321 / 0.9743 | 0.0241 / 0.9797 | 0.0175 / 0.9871 | 0.0133 / 0.9931 | |
| k-Center Greedy | 0.0683 / 0.9118 | 0.0349 / 0.9738 | 0.0283 / 0.9847 | 0.0218 / 0.9897 | 0.0175 / 0.9925 | |
| CoT-Feature (Ours) | 0.0755 / 0.8884 | 0.0331 / 0.9696 | 0.0209 / 0.9853 | 0.0197 / 0.9884 | 0.0154 / 0.9918 | |
| Trainable Methods (Require Historical Data, N=200) | ||||||
| Correctness K-Means | 0.0546 / 0.9239 | 0.0348 / 0.9694 | 0.0303 / 0.9800 | 0.0277 / 0.9817 | 0.0259 / 0.9870 | |
| IRT | 0.0723 / 0.9048 | 0.0283 / 0.9757 | 0.0205 / 0.9841 | 0.0172 / 0.9866 | 0.0148 / 0.9895 | |
| Training-Free Methods (Zero-Shot) | ||||||
| MMLU | Random | 0.0330 / 0.9489 | 0.0149 / 0.9880 | 0.0102 / 0.9935 | 0.0079 / 0.9957 | 0.0063 / 0.9969 |
| Question-Emb K-Means | 0.0298 / 0.9484 | 0.0215 / 0.9877 | 0.0154 / 0.9923 | 0.0203 / 0.9943 | 0.0252 / 0.9958 | |
| k-Center Greedy | 0.0477 / 0.9461 | 0.0530 / 0.9867 | 0.0393 / 0.9909 | 0.0360 / 0.9942 | 0.0317 / 0.9952 | |
| CoT-Feature (Ours) | 0.0324 / 0.9353 | 0.0118 / 0.9881 | 0.0081 / 0.9941 | 0.0081 / 0.9964 | 0.0060 / 0.9971 | |
| Trainable Methods (Require Historical Data, N=200) | ||||||
| Correctness K-Means | 0.0211 / 0.9699 | 0.0139 / 0.9885 | 0.0112 / 0.9898 | 0.0099 / 0.9911 | 0.0085 / 0.9935 | |
| IRT | 0.0223 / 0.9558 | 0.0097 / 0.9901 | 0.0072 / 0.9928 | 0.0060 / 0.9940 | 0.0055 / 0.9956 | |
| Training-Free Methods (Zero-Shot) | ||||||
| MMLU-Pro | Random | 0.0311 / 0.9556 | 0.0132 / 0.9899 | 0.0093 / 0.9941 | 0.0070 / 0.9956 | 0.0058 / 0.9967 |
| Question-Emb K-Means | 0.0300 / 0.9579 | 0.0133 / 0.9890 | 0.0111 / 0.9947 | 0.0116 / 0.9957 | 0.0100 / 0.9969 | |
| k-Center Greedy | 0.0369 / 0.9524 | 0.0432 / 0.9857 | 0.0334 / 0.9923 | 0.0306 / 0.9936 | 0.0272 / 0.9957 | |
| CoT-Feature (Ours) | 0.0296 / 0.9682 | 0.0124 / 0.9893 | 0.0082 / 0.9941 | 0.0072 / 0.9965 | 0.0055 / 0.9969 | |
| Trainable Methods (Require Historical Data, N=200) | ||||||
| Correctness K-Means | 0.0332 / 0.9847 | 0.0244 / 0.9918 | 0.0222 / 0.9927 | 0.0194 / 0.9937 | 0.0174 / 0.9940 | |
| IRT | 0.0233 / 0.9582 | 0.0123 / 0.9901 | 0.0080 / 0.9939 | 0.0061 / 0.9957 | 0.0054 / 0.9963 | |
| Training-Free Methods (Zero-Shot) | ||||||
| GPQA | Random | - / - | 0.1252 / 0.2851 | 0.0797 / 0.4480 | 0.0635 / 0.5579 | 0.0536 / 0.6022 |
| Question-Emb K-Means | - / - | 0.1098 / 0.2856 | 0.0859 / 0.3656 | 0.0615 / 0.5338 | 0.0522 / 0.6263 | |
| k-Center Greedy | - / - | 0.1232 / 0.2197 | 0.0821 / 0.4036 | 0.0608 / 0.5355 | 0.0565 / 0.6151 | |
| CoT-Feature (Ours) | - / - | 0.1044 / 0.4334 | 0.0731 / 0.5839 | 0.0601 / 0.5646 | 0.0516 / 0.6013 | |
| Trainable Methods (Require Historical Data, N=200) | ||||||
| Correctness K-Means | - / - | 0.1261 / 0.2994 | 0.1072 / 0.3593 | 0.0915 / 0.4311 | 0.0839 / 0.4810 | |
| IRT | - / - | 0.1055 / 0.2956 | 0.0776 / 0.4353 | 0.0610 / 0.5367 | 0.0529 / 0.5793 | |
| Benchmark | Method | Sampling Ratio | ||||
|---|---|---|---|---|---|---|
| 1% MAE (↓) / Sim (↑) |
5% MAE (↓) / Sim (↑) |
10% MAE (↓) / Sim (↑) |
15% MAE (↓) / Sim (↑) |
20% MAE (↓) / Sim (↑) |
||
| Training-Free Methods (Zero-Shot) | ||||||
| GSM8K | Random | 0.0881 / 0.8761 | 0.0287 / 0.9744 | 0.0192 / 0.9859 | 0.0152 / 0.9898 | 0.0125 / 0.9921 |
| Question-Emb K-Means | 0.0869 / 0.8594 | 0.0279 / 0.9729 | 0.0203 / 0.9811 | 0.0153 / 0.9874 | 0.0121 / 0.9931 | |
| k-Center Greedy | 0.0879 / 0.8655 | 0.0265 / 0.9731 | 0.0192 / 0.9817 | 0.0158 / 0.9891 | 0.0135 / 0.9923 | |
| CoT-Feature (Ours) | 0.0738 / 0.8793 | 0.0278 / 0.9693 | 0.0187 / 0.9860 | 0.0157 / 0.9886 | 0.0131 / 0.9914 | |
| Trainable Methods (Require Historical Data, N=200) | ||||||
| Correctness K-Means | 0.0705 / 0.9242 | 0.0260 / 0.9762 | 0.0198 / 0.9875 | 0.0175 / 0.9890 | 0.0163 / 0.9918 | |
| IRT | 0.0706 / 0.9023 | 0.0252 / 0.9758 | 0.0181 / 0.9857 | 0.0142 / 0.9892 | 0.0130 / 0.9916 | |
| Training-Free Methods (Zero-Shot) | ||||||
| MMLU | Random | 0.0186 / 0.9710 | 0.0093 / 0.9919 | 0.0074 / 0.9948 | 0.0062 / 0.9963 | 0.0053 / 0.9972 |
| Question-Emb K-Means | 0.0201 / 0.9657 | 0.0104 / 0.9906 | 0.0090 / 0.9942 | 0.0112 / 0.9953 | 0.0150 / 0.9964 | |
| k-Center Greedy | 0.0188 / 0.9720 | 0.0160 / 0.9914 | 0.0173 / 0.9937 | 0.0192 / 0.9955 | 0.0190 / 0.9959 | |
| CoT-Feature (Ours) | 0.0216 / 0.9617 | 0.0085 / 0.9919 | 0.0063 / 0.9956 | 0.0059 / 0.9966 | 0.0052 / 0.9974 | |
| Trainable Methods (Require Historical Data, N=200) | ||||||
| Correctness K-Means | 0.0186 / 0.9722 | 0.0094 / 0.9925 | 0.0080 / 0.9938 | 0.0073 / 0.9944 | 0.0064 / 0.9956 | |
| IRT | 0.0184 / 0.9714 | 0.0086 / 0.9921 | 0.0067 / 0.9938 | 0.0055 / 0.9954 | 0.0049 / 0.9964 | |
| Training-Free Methods (Zero-Shot) | ||||||
| MMLU-Pro | Random | 0.0225 / 0.9646 | 0.0094 / 0.9908 | 0.0069 / 0.9942 | 0.0055 / 0.9955 | 0.0048 / 0.9963 |
| Question-Emb K-Means | 0.0228 / 0.9611 | 0.0094 / 0.9904 | 0.0074 / 0.9943 | 0.0069 / 0.9952 | 0.0062 / 0.9967 | |
| k-Center Greedy | 0.0228 / 0.9673 | 0.0117 / 0.9901 | 0.0111 / 0.9940 | 0.0127 / 0.9949 | 0.0127 / 0.9958 | |
| CoT-Feature (Ours) | 0.0216 / 0.9747 | 0.0098 / 0.9907 | 0.0069 / 0.9937 | 0.0056 / 0.9954 | 0.0047 / 0.9962 | |
| Trainable Methods (Require Historical Data, N=200) | ||||||
| Correctness K-Means | 0.0234 / 0.9685 | 0.0097 / 0.9933 | 0.0092 / 0.9954 | 0.0087 / 0.9966 | 0.0090 / 0.9968 | |
| IRT | 0.0221 / 0.9619 | 0.0095 / 0.9908 | 0.0072 / 0.9940 | 0.0052 / 0.9958 | 0.0046 / 0.9963 | |
| Training-Free Methods (Zero-Shot) | ||||||
| GPQA | Random | - / - | 0.1108 / 0.2931 | 0.0875 / 0.4409 | 0.0578 / 0.5951 | 0.0423 / 0.6625 |
| Question-Emb K-Means | - / - | 0.1111 / 0.2699 | 0.0826 / 0.4144 | 0.0521 / 0.6110 | 0.0448 / 0.6466 | |
| k-Center Greedy | - / - | 0.1170 / 0.2745 | 0.0863 / 0.4452 | 0.0623 / 0.5391 | 0.0487 / 0.5753 | |
| CoT-Feature (Ours) | - / - | 0.1031 / 0.3926 | 0.0737 / 0.5764 | 0.0492 / 0.6092 | 0.0423 / 0.6502 | |
| Trainable Methods (Require Historical Data, N=200) | ||||||
| Correctness K-Means | - / - | 0.0805 / 0.3920 | 0.0605 / 0.4511 | 0.0449 / 0.4943 | 0.0406 / 0.5771 | |
| IRT | - / - | 0.0837 / 0.3378 | 0.0586 / 0.4497 | 0.0484 / 0.5412 | 0.0374 / 0.6340 | |
References
- Marah Abdin, Jyoti Aneja, Harkirat Behl, Sébastien Bubeck, Ronen Eldan, Suriya Gunasekar, Michael Harrison, Russell J Hewett, Mojan Javaheripi, Piero Kauffmann, et al. Phi-4 technical report. arXiv preprint arXiv:2412.08905, 2024.
- Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, Florencia Leoni Aleman, Diogo Almeida, Janko Altenschmidt, Sam Altman, Shyamal Anadkat, et al. Gpt-4 technical report. arXiv preprint arXiv:2303.08774, 2023.
- Edward Beeching, Clémentine Fourrier, Nathan Habib, Sheon Han, Nathan Lambert, Nazneen Rajani, Omar Sanseviero, Lewis Tunstall, and Thomas Wolf. Open llm leaderboard. https://huggingface.co/spaces/open-llm-leaderboard/open_llm_leaderboard, 2023.
- Stella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley, Kyle O’Brien, Eric Hallahan, Mohammad Aflah Khan, Shivanshu Purohit, USVSN Sai Prashanth, Edward Raff, et al. Pythia: A suite for analyzing large language models across training and scaling. In International conference on machine learning, pp. 2397–2430. PMLR, 2023.
- Jianlv Chen, Shitao Xiao, Peitian Zhang, Kun Luo, Defu Lian, and Zheng Liu. Bge m3-embedding: Multi-lingual, multi-functionality, multi-granularity text embeddings through self-knowledge distillation. arXiv preprint arXiv:2402.03216, 4(5), 2024.
- Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021.
- Dirk Groeneveld, Iz Beltagy, Evan Walsh, Akshita Bhagia, Rodney Kinney, Oyvind Tafjord, Ananya Jha, Hamish Ivison, Ian Magnusson, Yizhong Wang, et al. Olmo: Accelerating the science of language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 15789–15809, 2024. [CrossRef]
- Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding. arXiv preprint arXiv:2009.03300, 2020.
- Peiyu Li, Xiuxiu Tang, Si Chen, Ying Cheng, Ronald Metoyer, Ting Hua, and Nitesh V Chawla. Adaptive testing for llm evaluation: A psychometric alternative to static benchmarks. arXiv preprint arXiv:2511.04689, 2025.
- Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, et al. Holistic evaluation of language models. arXiv preprint arXiv:2211.09110, 2022.
- OpenCompass Contributors. Opencompass: A universal evaluation platform for foundation models. https://github.com/open-compass/opencompass, 2023.
- Felipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun, Gongjun Xu, and Mikhail Yurochkin. tinybenchmarks: evaluating llms with fewer examples. arXiv preprint arXiv:2402.14992, 2024.
- David Rein, Betty Li Hou, Asa Cooper Stickland, Jackson Petty, Richard Yuanzhe Pang, Julien Dirani, Julian Michael, and Samuel R Bowman. Gpqa: A graduate-level google-proof q&a benchmark. In First conference on language modeling, 2024.
- Ozan Sener and Silvio Savarese. Active learning for convolutional neural networks: A core-set approach. arXiv preprint arXiv:1708.00489, 2017.
- Freda Shi, Xinyun Chen, Kanishka Misra, Nathan Scales, David Dohan, Ed H Chi, Nathanael Schärli, and Denny Zhou. Large language models can be easily distracted by irrelevant context. In International Conference on Machine Learning, pp. 31210–31227. PMLR, 2023.
- Koustuv Sinha, Prasanna Parthasarathi, Joelle Pineau, and Adina Williams. Unnatural language inference. In Proceedings of the 59th annual meeting of the Association for Computational Linguistics and the 11th international joint conference on natural language processing (Volume 1: Long Papers), pp. 7329–7346, 2021. [CrossRef]
- Aarohi Srivastava, Abhinav Rastogi, Abhishek Rao, Abu Awal Md Shoeb, Abubakar Abid, Adam Fisch, Adam R Brown, Adam Santoro, Aditya Gupta, Adrià Garriga-Alonso, et al. Beyond the imitation game: Quantifying and extrapolating the capabilities of language models. Transactions on machine learning research, 2023.
- Swabha Swayamdipta, Roy Schwartz, Nicholas Lourie, Yizhong Wang, Hannaneh Hajishirzi, Noah A Smith, and Yejin Choi. Dataset cartography: Mapping and diagnosing datasets with training dynamics. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 9275–9293, 2020.
- Nandan Thakur, Nils Reimers, Andreas Rücklé, Abhishek Srivastava, and Iryna Gurevych. Beir: A heterogenous benchmark for zero-shot evaluation of information retrieval models. arXiv preprint arXiv:2104.08663, 2021.
- Hugo Touvron, Louis Martin, Kevin Stone, Peter Albert, Amjad Almahairi, Yasmine Babaei, Nikolay Bashlykov, Soumya Batra, Prajjwal Bhargava, Shruti Bhosale, et al. Llama 2: Open foundation and fine-tuned chat models. arXiv preprint arXiv:2307.09288, 2023.
- Shaobo Wang, Cong Wang, Wenjie Fu, Yue Min, Mingquan Feng, Isabel Guan, Xuming Hu, Conghui He, Cunxiang Wang, Kexin Yang, et al. Rethinking llm evaluation: Can we evaluate llms with 200x less data? arXiv preprint arXiv:2510.10457, 2025.
- Yubo Wang, Xueguang Ma, Ge Zhang, Yuansheng Ni, Abhranil Chandra, Shiguang Guo, Weiming Ren, Aaran Arulraj, Xuan He, Ziyan Jiang, et al. Mmlu-pro: A more robust and challenging multi-task language understanding benchmark. Advances in Neural Information Processing Systems, 37:95266–95290, 2024. [CrossRef]
- Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in neural information processing systems, 35:24824–24837, 2022. [CrossRef]
- Sheldon Yu, Yuxin Xiong, Junda Wu, Xintong Li, Tong Yu, Xiang Chen, Ritwik Sinha, Jingbo Shang, and Julian McAuley. Explainable chain-of-thought reasoning: An empirical analysis on state-aware reasoning dynamics. arXiv preprint arXiv:2509.00190, 2025.
- Zhuosheng Zhang, Aston Zhang, Mu Li, and Alex Smola. Automatic chain of thought prompting in large language models. arXiv preprint arXiv:2210.03493, 2022.
- Zicheng Zhang, Xiangyu Zhao, Xinyu Fang, Chunyi Li, Xiaohong Liu, Xiongkuo Min, Haodong Duan, Kai Chen, and Guangtao Zhai. Redundancy principles for mllms benchmarks. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 12492–12504, 2025. [CrossRef]
| 1 | It is worth noting that the LLM exhibits a minor reasoning hallucination in Question B (applying a linear proportion to exponential decay). However, this precisely highlights the robustness of our approach: the embedding alignment relies on the structural and conceptual vocabulary (Lorentz transformations, time dilation) extracted by the CoT, rather than the strict arithmetic correctness of the final step. |
| 2 | It should be noted that in Question D, the LLM utilizes a simplified geometric assumption regarding the spot area on a sphere versus a projected cross-sectional disk (ignoring limb darkening effects). Nevertheless, this pedagogical simplification does not hinder the embedding process. The core contribution of the CoT remains intact: successfully shifting the semantic representation from the misleading lexical surface to the correct physical domain (macroscopic geometric flux rather than quantum states). |


| Benchmark | Method | Sampling Ratio | ||||
|---|---|---|---|---|---|---|
| 1% MAE (↓) / Sim (↑) |
5% MAE (↓) / Sim (↑) |
10% MAE (↓) / Sim (↑) |
15% MAE (↓) / Sim (↑) |
20% MAE (↓) / Sim (↑) |
||
| Training-Free Methods (Zero-Shot) | ||||||
| GSM8K | Random | 0.0890 / 0.9044 | 0.0382 / 0.9747 | 0.0246 / 0.9851 | 0.0187 / 0.9897 | 0.0142 / 0.9920 |
| Question-Emb K-Means | 0.0875 / 0.8951 | 0.0321 / 0.9743 | 0.0241 / 0.9797 | 0.0175 / 0.9871 | 0.0133 / 0.9931 | |
| k-Center Greedy | 0.0683 / 0.9118 | 0.0349 / 0.9738 | 0.0283 / 0.9847 | 0.0218 / 0.9897 | 0.0175 / 0.9925 | |
| CoT-Feature (Ours) | 0.0755 / 0.8884 | 0.0331 / 0.9696 | 0.0209 / 0.9853 | 0.0197 / 0.9884 | 0.0154 / 0.9918 | |
| Trainable Methods (Require Historical Data, N=25) | ||||||
| Correctness K-Means | 0.0596 / 0.9129 | 0.0347 / 0.9700 | 0.0259 / 0.9770 | 0.0231 / 0.9838 | 0.0198 / 0.9886 | |
| IRT | 0.0664 / 0.9172 | 0.0308 / 0.9649 | 0.0196 / 0.9817 | 0.0175 / 0.9852 | 0.0145 / 0.9879 | |
| Training-Free Methods (Zero-Shot) | ||||||
| MMLU | Random | 0.0330 / 0.9489 | 0.0149 / 0.9880 | 0.0102 / 0.9935 | 0.0079 / 0.9957 | 0.0063 / 0.9969 |
| Question-Emb K-Means | 0.0298 / 0.9484 | 0.0215 / 0.9877 | 0.0154 / 0.9923 | 0.0203 / 0.9943 | 0.0252 / 0.9958 | |
| k-Center Greedy | 0.0477 / 0.9461 | 0.0530 / 0.9867 | 0.0393 / 0.9909 | 0.0360 / 0.9942 | 0.0317 / 0.9952 | |
| CoT-Feature (Ours) | 0.0324 / 0.9353 | 0.0118 / 0.9881 | 0.0081 / 0.9941 | 0.0081 / 0.9964 | 0.0060 / 0.9971 | |
| Trainable Methods (Require Historical Data, N=25) | ||||||
| Correctness K-Means | 0.0304 / 0.9616 | 0.0203 / 0.9754 | 0.0181 / 0.9773 | 0.0169 / 0.9781 | 0.0160 / 0.9785 | |
| IRT | 0.0253 / 0.9589 | 0.0112 / 0.9895 | 0.0081 / 0.9938 | 0.0068 / 0.9966 | 0.0056 / 0.9969 | |
| Training-Free Methods (Zero-Shot) | ||||||
| MMLU-Pro | Random | 0.0311 / 0.9556 | 0.0132 / 0.9899 | 0.0093 / 0.9941 | 0.0070 / 0.9956 | 0.0058 / 0.9967 |
| Question-Emb K-Means | 0.0300 / 0.9579 | 0.0133 / 0.9890 | 0.0111 / 0.9947 | 0.0116 / 0.9957 | 0.0100 / 0.9969 | |
| k-Center Greedy | 0.0369 / 0.9524 | 0.0432 / 0.9857 | 0.0334 / 0.9923 | 0.0306 / 0.9936 | 0.0272 / 0.9957 | |
| CoT-Feature (Ours) | 0.0296 / 0.9682 | 0.0124 / 0.9893 | 0.0082 / 0.9941 | 0.0072 / 0.9965 | 0.0055 / 0.9969 | |
| Trainable Methods (Require Historical Data, N=25) | ||||||
| Correctness K-Means | 0.0257 / 0.9789 | 0.0144 / 0.9902 | 0.0123 / 0.9929 | 0.0107 / 0.9936 | 0.0095 / 0.9941 | |
| IRT | 0.0379 / 0.9540 | 0.0151 / 0.9891 | 0.0111 / 0.9930 | 0.0092 / 0.9943 | 0.0073 / 0.9962 | |
| Training-Free Methods (Zero-Shot) | ||||||
| GPQA | Random | - / - | 0.1252 / 0.2851 | 0.0797 / 0.4480 | 0.0635 / 0.5579 | 0.0536 / 0.6022 |
| Question-Emb K-Means | - / - | 0.1098 / 0.2856 | 0.0859 / 0.3656 | 0.0615 / 0.5338 | 0.0522 / 0.6263 | |
| k-Center Greedy | - / - | 0.1232 / 0.2197 | 0.0821 / 0.4036 | 0.0608 / 0.5355 | 0.0565 / 0.6151 | |
| CoT-Feature (Ours) | - / - | 0.1044 / 0.4334 | 0.0731 / 0.5839 | 0.0601 / 0.5646 | 0.0516 / 0.6013 | |
| Trainable Methods (Require Historical Data, N=25) | ||||||
| Correctness K-Means | - / - | 0.0936 / 0.3125 | 0.0743 / 0.3435 | 0.0660 / 0.4498 | 0.0570 / 0.4770 | |
| IRT | - / - | 0.1054 / 0.2576 | 0.0723 / 0.3806 | 0.0552 / 0.6081 | 0.0513 / 0.6176 | |
| Benchmark | Method | Sampling Ratio | ||||
|---|---|---|---|---|---|---|
| 1% MAE (↓) / Sim (↑) |
5% MAE (↓) / Sim (↑) |
10% MAE (↓) / Sim (↑) |
15% MAE (↓) / Sim (↑) |
20% MAE (↓) / Sim (↑) |
||
| Training-Free Methods (Zero-Shot) | ||||||
| GSM8K | Random | 0.0827 / 0.8846 | 0.0297 / 0.9701 | 0.0202 / 0.9843 | 0.0158 / 0.9893 | 0.0131 / 0.9919 |
| Question-Emb K-Means | 0.0975 / 0.8443 | 0.0296 / 0.9715 | 0.0206 / 0.9776 | 0.0158 / 0.9866 | 0.0123 / 0.9928 | |
| k-Center Greedy | 0.1016 / 0.8597 | 0.0289 / 0.9707 | 0.0216 / 0.9830 | 0.0166 / 0.9895 | 0.0136 / 0.9929 | |
| CoT-Feature (Ours) | 0.0800 / 0.8700 | 0.0290 / 0.9642 | 0.0197 / 0.9844 | 0.0166 / 0.9882 | 0.0135 / 0.9914 | |
| Trainable Methods (Require Historical Data, N=25) | ||||||
| Correctness K-Means | 0.0830 / 0.8608 | 0.0281 / 0.9691 | 0.0183 / 0.9830 | 0.0156 / 0.9886 | 0.0121 / 0.9940 | |
| IRT | 0.0789 / 0.8957 | 0.0308 / 0.9659 | 0.0177 / 0.9838 | 0.0158 / 0.9877 | 0.0134 / 0.9909 | |
| Training-Free Methods (Zero-Shot) | ||||||
| MMLU | Random | 0.0212 / 0.9639 | 0.0116 / 0.9897 | 0.0089 / 0.9938 | 0.0072 / 0.9957 | 0.0060 / 0.9968 |
| Question-Emb K-Means | 0.0225 / 0.9606 | 0.0120 / 0.9891 | 0.0105 / 0.9933 | 0.0130 / 0.9947 | 0.0165 / 0.9961 | |
| k-Center Greedy | 0.0198 / 0.9695 | 0.0190 / 0.9893 | 0.0191 / 0.9929 | 0.0209 / 0.9951 | 0.0205 / 0.9957 | |
| CoT-Feature (Ours) | 0.0223 / 0.9577 | 0.0111 / 0.9901 | 0.0076 / 0.9948 | 0.0068 / 0.9962 | 0.0059 / 0.9969 | |
| Trainable Methods (Require Historical Data, N=25) | ||||||
| Correctness K-Means | 0.0222 / 0.9648 | 0.0142 / 0.9856 | 0.0121 / 0.9864 | 0.0113 / 0.9850 | 0.0110 / 0.9836 | |
| IRT | 0.0209 / 0.9651 | 0.0131 / 0.9883 | 0.0095 / 0.9932 | 0.0073 / 0.9961 | 0.0066 / 0.9968 | |
| Training-Free Methods (Zero-Shot) | ||||||
| MMLU-Pro | Random | 0.0250 / 0.9598 | 0.0104 / 0.9904 | 0.0074 / 0.9944 | 0.0060 / 0.9955 | 0.0051 / 0.9966 |
| Question-Emb K-Means | 0.0244 / 0.9547 | 0.0103 / 0.9895 | 0.0081 / 0.9949 | 0.0083 / 0.9955 | 0.0074 / 0.9968 | |
| k-Center Greedy | 0.0240 / 0.9670 | 0.0146 / 0.9896 | 0.0149 / 0.9935 | 0.0162 / 0.9943 | 0.0162 / 0.9955 | |
| CoT-Feature (Ours) | 0.0240 / 0.9724 | 0.0102 / 0.9901 | 0.0069 / 0.9943 | 0.0058 / 0.9963 | 0.0047 / 0.9968 | |
| Trainable Methods (Require Historical Data, N=25) | ||||||
| Correctness K-Means | 0.0234 / 0.9747 | 0.0091 / 0.9925 | 0.0081 / 0.9948 | 0.0075 / 0.9957 | 0.0070 / 0.9958 | |
| IRT | 0.0257 / 0.9549 | 0.0109 / 0.9903 | 0.0095 / 0.9943 | 0.0085 / 0.9960 | 0.0075 / 0.9971 | |
| Training-Free Methods (Zero-Shot) | ||||||
| GPQA | Random | - / - | 0.1019 / 0.2807 | 0.0679 / 0.4782 | 0.0568 / 0.5842 | 0.0497 / 0.6248 |
| Question-Emb K-Means | - / - | 0.0887 / 0.2759 | 0.0703 / 0.3900 | 0.0552 / 0.5704 | 0.0482 / 0.6388 | |
| k-Center Greedy | - / - | 0.0973 / 0.2351 | 0.0701 / 0.4334 | 0.0551 / 0.5547 | 0.0519 / 0.6335 | |
| CoT-Feature (Ours) | - / - | 0.0922 / 0.5123 | 0.0649 / 0.6206 | 0.0545 / 0.5812 | 0.0482 / 0.6318 | |
| Trainable Methods (Require Historical Data, N=25) | ||||||
| Correctness K-Means | - / - | 0.0728 / 0.3970 | 0.0597 / 0.4130 | 0.0559 / 0.5022 | 0.0503 / 0.5293 | |
| IRT | - / - | 0.0887 / 0.2545 | 0.0622 / 0.4248 | 0.0500 / 0.6402 | 0.0473 / 0.6475 | |
| Abstraction Level | GPQA Budget (MAE ↓ / Sim ↑) | ||
|---|---|---|---|
| 5% | 10% | 20% | |
| Raw CoT | 0.1044 / 0.4334 | 0.0731 / 0.5839 | 0.0516 / 0.6013 |
| Cleaned CoT | 0.1007 / 0.5041 | 0.0666 / 0.4970 | 0.0442 / 0.6438 |
| Abstract CoT | 0.1206 / 0.3561 | 0.0831 / 0.4506 | 0.0518 / 0.5714 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).