Submitted:
12 July 2026
Posted:
13 July 2026
You are already at the latest version
Abstract
Keywords:
1. Introduction
2. Background and Related Work
2.1. Global Workspace Theory and the J-Space Result
2.2. Activation Steering
2.3. Lenses and Readback
2.4. Active Inference for Experiment Selection
2.5. Pratyabhij ñ ā
3. The Doctrine as Engineering Specification
3.1. What: Verbalizable Concept Codes
3.2. Where: The Band, in Band Coordinates
3.3. When: Uncommitted Moments
3.4. How to Verify: Band-Targeted Readback
3.5. Budget: Freedom and Its Cost
3.6. Scope
4. System and Research-Loop Architecture
4.1. Control Loop
4.2. Dual-Closure Gates
4.3. Expected-Free-Energy Experiment Selection
4.4. Adversarial Review
4.5. Compute Discipline
5. Experimental Setup
5.1. Models
5.2. Corpora and Concepts
5.3. Seeding and Independence
5.4. Tiers, Gates, and Statistics
5.5. Reproducibility
6. Results
6.1. F1: Loadability Replicates and Scales; Pruning Collapses It
6.2. F2: Band Content Is Legible Only to a Band-Targeted Lens, and the Gradient Is Model-Intrinsic
6.3. F3: Uncapped Writing Steers but Destroys Freedom; the Budget Binds
6.4. F4: Timing Is the Mechanism; the Flash Variant Failed and Was Revised
6.5. F5: The Core Claim at Six Independent-Stream Seeds
6.6. F6a: Dose Response by Timing Arm, with a Documented Supersession
6.7. F6b: The Calibration Recipe Transfers Across Plants; Geometry Does Not
6.8. F6c: Readback Tracks Behaviour at Screen Tier; Pooled Generalisation Fails
6.9. F6d: Corpus Robustness Is an Amplitude Question
6.10. F7: An Untrained Associative-Memory Write Path Matches the Analytic Path
6.11. F8: Gated Steering Is a Legibility Mechanism, Not a Throughput Maximiser
6.12. F9: Naive Contrastive Directions Do Not Improve Alignment on an Aligned Model
6.13. F10: Band-Targeted Reading Is Twice as Sensitive as the Final-Target Baseline at Subtle Doses
6.14. F11: Gated Placement Is 2.3 Times More Write-Efficient than Continuous Flooding
6.15. F12: Recognition-Gated Hardening Is a Deployable Moat on Some Models, Predicted Before Deployment
7. Core Innovations and Defensibility
7.1. a Doctrine, Not a Vector
7.2. Reading in the Band’s Own Basis
7.3. Recognition as a Deployable Defence
7.4. a Reproducible Research Instrument
7.5. an Integrated, Released Product
8. Discussion
8.1. What the Doctrine Buys
8.2. Transparency Rather than Throughput
8.3. Loadability as a Model Property
8.4. the Research Loop as a Result
8.5. Relation to the Source Philosophy
8.6. Limitations
9. Released Artifacts and Reproducibility
10. Conclusions
Acknowledgments
Appendix A. Gate Record Schema
Appendix B. Complete Gate Ledger
| Loop | Headline gate(s) | Tier | Verdict | Principal outcome |
|---|---|---|---|---|
| L0 | gate_L0 | smoke | pass | Scaffold, contracts, TRIZ contradictions C1–C3 resolved. |
| L1 | gate_L1 | screen | fail | Qwen3-4B loadability HR@5 0.10 (null 0.0104, ); band contrast 0.306; reportability sub-hypothesis fails. |
| L1b | gate_L1b | screen | pass | Qwen3.6-27B loadability HR@5 0.55 (22/40, null 0.068); loadability scales with size. |
| L2 | gate_L2 | screen | fail | Nemotron twin loadability 0.00 (= null = control); articulation gradient (). |
| L2b | gate_L2b | screen | registered outcome | Band-exit lens reads 0.20 across 7 concepts; final-target lens reads 0.00. |
| L3 | gate_L3 | screen | fail | Uncapped writes: lift 0.40 at ; freedom cost binding (20/40 rejections). |
| L4 | gate_L4, gate_L4b | screen | fail, then pass | Flash-gating fails (0.175); corrected entropy gating passes: lift 0.40 at , ≈9.85 writes/gen vs. prefill 0.20. |
| L5 | gate_L5_tau (6 runs) | screen | pass | robustness 6/6 across P40/P60/P80 percentiles × 2 seeds; lift 0.35–0.40 in budget. |
| L6 | gate_L6_align, _confirm | confirm | pass / split | Gated 0.40 (7.15 writes) > rate-matched 0.225 (5.0) > prefill 0.20; 3-seed budget claim 3/3, margin claim 1/3. |
| L7 | gate_L7_articulation_null | screen | pass | Articulation gradient model-intrinsic: logit-lens null vs. Jacobian , both . |
| L8 | gate_L8_dose | screen | pass (superseded) | Dose grid; gated levels later shown ≈0.1 high (pre-stream-fix); superseded by L18. |
| L9 | gate_L9_probe, _alignconf, _flash, _writecost | screen | pass / fail / fail / pass | Stream fix (per-generation seeds); clean 3-seed gated 0.30–0.35, margin //; flash adds nothing; throughput cost of writes nil (24.8–25.0 vs. 19.7 tok/s). |
| L10 | gate_L10_cross | screen | fail | Borrowed geometry on Qwen3-4B: gated 0.05; target Jacobians ≈10× weaker; generality boundary. |
| L11 | gate_L11_rep | confirm | pass | Core claim at 6 seeds: gated 0.30–0.35 in budget; advantage sign-consistent 6/6, . |
| L12 | gate_L12_flash (3 seeds) | screen | pass | Uncommitted-moment gating beats commitment-flash 3/3 ( to ). |
| L13 | gate_L13_recipe, probes | screen | pass | Calibrated amplitude (3×) transfers: Qwen3-4B gated 0.40 vs. prefill 0.175; site probes eliminate layer choice. |
| L14 | gate_L14_amp, _multiseed, _readback | confirm / screen | pass | Monotone dose 0.05/0.20/0.40/0.775; recipe 4/4 seeds (0.325–0.475); readback BA 0.684 (36-setting sweep disclosed). |
| L15 | gate_L15_amp_joint, _readback | confirm / screen | pass | Per-seed monotone 3/3; two-plant ordering (screen); held-out readback BA 0.637 (margin inside CI; downgraded by review #12). |
| L16 | gate_L16_corpus, _fine | screen | pass | Corpus robustness 2/3; pooled readback BA 0.590 (negative exploratory); donor fine grid saturates near . |
| L17 | gate_L17_xdose, _cvar | screen | pass / fail | Arm-set robustness holds; corpus-A seed-fragile (2/4 registered cells), self-audit negative. |
| L18 | gate_L18_l8redo, _npretry | screen | pass | L8 supersession measured ( at all three amplitudes); fragility repaired by amplitude (3/3 at ). |
| L19 | gate_L19_cax, _l8ms | confirm / screen | fail-on-margin / pass | Both corpora double under amplitude; one per-seed margin fails (0.025 < 0.05); supersession offset confirmed at , all 6 offsets negative. |
| L20 | gate_L20_confirm | confirm | pass (equivalence fail-on-margin) | Cold associative store steers 3/3 (0.4445/0.5556/0.4445); equivalence to analytic 2/3, third seed one concept-hit apart. |
| L21 | gate_L21_baselines (3 seeds), _jailbreak, _truthful | screen | fail (all registered) | Continuous CSR 0.96 vs. gated 0.16 at matched entropy cost; AdvBench ASR 0.25 unchanged; truthfulness criterion fails. |
| L22 | gate_L22_benchmark, _lens_headtohead, _floor_a* | confirm | pass (head-to-head fail-on-margin) | Band vs. final-target lens: 0.475 vs. 0.2375 at (), converge at saturation (1.00 vs. 0.95); gated lift-per-write 2.32× continuous, 6/6 cells. |
| L23 | gate_L23_harden | screen | fail | Fused prayoga-prabodha harden loop: gated preserves freedom (11.8 vs. 50 writes) but does not resist an activation-level attacker; motivates recognition gating. |
| char | gate_char_qwen3-4b, _nemotron-mini-4b | screen | pass | Both prompt-robust (attack 0.00) but activation-vulnerable (0.90 / 1.00); locates the harmful signature in the residual stream. |
| L24 | gate_L24_innovation | screen | fail | Weight/prompt-space hardening mechanisms weak (best restore-prefill 0.1 leaves ASR 0.70); pushes toward recognition gating. |
| L25 | gate_L25_promptspace | screen | fail | Prompt-space defences uninformative on this corpus (baseline wrapped ASR already 0.00). |
| L26 | gate_L26_moat_proof, _llama, _qwen, _smol | screen (exploratory) | pass / pass / honest neg. / honest neg. | Recognition-gated moat: Gemma-2-2B 0.50→0.25 and Llama-3.2-1B 0.25→0.083 at 0 over-refusal; Qwen2.5-1.5B ineffective; SmolLM2-1.7B backfires 0.583→0.917; clean-gap predicts 4/4. |
Appendix C. Expected-Free-Energy Selector
Appendix D. Adversarial Review Catalogue
Appendix E. Glossary of Sanskrit Terms
Appendix F. Module and Configuration Map
Appendix G. Figure Provenance
References
- Bricken, T.; Templeton, A.; Batson, J.; Chen, B.; Jermyn, A.; Conerly, T.; Turner, N.; Anil, C.; Denison, C.; Askell, A.; et al. Towards Monosemanticity: Decomposing Language Models with Dictionary Learning; Transformer Circuits series, 2023. [Google Scholar]
- Cunningham, H.; Ewart, A.; Riggs, L.; Huben, R.; Sharkey, L. Sparse Autoencoders Find Highly Interpretable Features in Language Models. arXiv 2023, arXiv:2309.08600. [Google Scholar]
- Turner, A.M.; Thiergart, L.; Leech, G.; Udell, D.; Vazquez, J.J.; Mini, U.; MacDiarmid, M. Activation Addition: Steering Language Models Without Optimization. arXiv 2023, arXiv:2308.10248. [Google Scholar]
- Rimsky, N.; Gabrieli, N.; Schulz, J.; Tong, M.; Hubinger, E.; Turner, A. Steering Llama 2 via Contrastive Activation Addition. Proceedings of the Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics 2024, Volume 1, 15504–15522. [Google Scholar] [CrossRef]
- Zou, A.; Phan, L.; Chen, S.; Campbell, J.; Guo, P.; Ren, R.; Pan, A.; Yin, X.; Mazeika, M.; Dombrowski, A.K.; et al. Representation Engineering: A Top-Down Approach to AI Transparency. arXiv 2023, arXiv:2310.01405. [Google Scholar]
- nostalgebraist. Interpreting GPT: The Logit Lens. LessWrong Blog Post. 2020. [Google Scholar]
- Belrose, N.; Furman, Z.; Smith, L.; Halawi, D.; Ostrovsky, I.; McKinney, L.; Biderman, S.; Steinhardt, J. Eliciting Latent Predictions from Transformers with the Tuned Lens. arXiv 2023, arXiv:2303.08112. [Google Scholar]
- Anthropic Interpretability Team. Verbalizable Representations Form a Global Workspace in Language Models; Transformer Circuits series, 2026. [Google Scholar]
- Anthropic Interpretability Team. jlens: Jacobian Lens Reference Implementation. GitHub repository, Apache-2.0 licence. Companion code for the workspace article; vendored unmodified in the prabodha repository. 2026. [Google Scholar] [CrossRef] [PubMed]
- Torella, R. The Īśvarapratyabhijñākārikā of Utpaladeva with the Author’s Vṛtti: Critical Edition and Annotated Translation, corrected ed.; Motilal Banarsidass: Delhi, 2002. [Google Scholar]
- Friston, K.; Rigoli, F.; Ognibene, D.; Mathys, C.; Fitzgerald, T.; Pezzulo, G. Active Inference and Epistemic Value. Cogn. Neurosci. 2015, 6, 187–214. [Google Scholar] [CrossRef] [PubMed]
- Da Costa, L.; Parr, T.; Sajid, N.; Veselic, S.; Neacsu, V.; Friston, K. Active Inference on Discrete State-Spaces: A Synthesis. J. Math. Psychol. 2020, 99, 102447. [Google Scholar] [CrossRef] [PubMed]
- Parr, T.; Pezzulo, G.; Friston, K.J. Active Inference: The Free Energy Principle in Mind, Brain, and Behavior; MIT Press: Cambridge, MA, 2022. [Google Scholar]
- Sathish, S.; Ahsan, M.; Latifi, M. Active Circuit Discovery: A Multi-Action POMDP Agent for Causal Feature Identification in Transformer Attribution Graphs. Symmetry 2026, 18, 1043. [Google Scholar] [CrossRef]
- Sathish, S. Refusal as a Broken Symmetry: Mechanistic Interpretability of Output-Policy Capture Across Jailbreak, Hypnosis, and Vaśīkaraṇa. Preprints 2026, Article 2026070139. [Google Scholar] [CrossRef]
- Arditi, A.; Obeso, O.; Syed, A.; Paleka, D.; Panickssery, N.; Gurnee, W.; Nanda, N. Refusal in Language Models Is Mediated by a Single Direction. In Proceedings of the Advances in Neural Information Processing Systems, 2024; 37. [Google Scholar]
- Baars, B.J. A Cognitive Theory of Consciousness; Cambridge University Press: Cambridge, 1988. [Google Scholar]
- Dehaene, S.; Kerszberg, M.; Changeux, J.P. A Neuronal Model of a Global Workspace in Effortful Cognitive Tasks. Proc. Natl. Acad. Sci. 1998, 95, 14529–14534. [Google Scholar] [CrossRef] [PubMed]
- Mashour, G.A.; Roelfsema, P.; Changeux, J.P.; Dehaene, S. Conscious Processing and the Global Neuronal Workspace Hypothesis. Neuron 2020, 105, 776–798. [Google Scholar] [CrossRef] [PubMed]
- Dehaene, S. Consciousness and the Brain: Deciphering How the Brain Codes Our Thoughts; Viking: New York, 2014. [Google Scholar]
- Butlin, P.; Long, R.; Elmoznino, E.; Bengio, Y.; Birch, J.; Constant, A.; Deane, G.; Fleming, S.M.; Frith, C.; Ji, X.; et al. Consciousness in Artificial Intelligence: Insights from the Science of Consciousness. arXiv 2023, arXiv:2308.08708. [Google Scholar]
- Todd, E.; Li, M.L.; Sharma, A.S.; Mueller, A.; Wallace, B.C.; Bau, D. Function Vectors in Large Language Models. In Proceedings of the The Twelfth International Conference on Learning Representations, 2024. [Google Scholar]
- Templeton, A.; Conerly, T.; Marcus, J.; Lindsey, J.; Bricken, T.; Chen, B.; Pearce, A.; Citro, C.; Ameisen, E.; Jones, A.; et al. Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet; Transformer Circuits series, 2024. [Google Scholar]
- Friston, K. The Free-Energy Principle: A Unified Brain Theory? Nat. Rev. Neurosci. 2010, 11, 127–138. [Google Scholar] [CrossRef] [PubMed]
- Dyczkowski, M.S.G. The Stanzas on Vibration: The Spandakārikā with Four Commentaries; State University of New York Press: Albany, 1992. [Google Scholar]
- Altshuller, G.S. Creativity as an Exact Science: The Theory of the Solution of Inventive Problems; Gordon and Breach: New York, 1984. [Google Scholar]
- Yang, A.; et al. Qwen3 Technical Report. arXiv 2025, arXiv:2505.09388. [Google Scholar]
- Sreenivas, S.T.; Muralidharan, S.; Joshi, R.; Chochowski, M.; Patwary, M.; Shoeybi, M.; Catanzaro, B.; Kautz, J.; Molchanov, P. LLM Pruning and Distillation in Practice: The Minitron Approach. arXiv 2024, arXiv:2408.11796. [Google Scholar]
- Gemma Team. Gemma 2: Improving Open Language Models at a Practical Size. arXiv 2024, arXiv:2408.00118. [Google Scholar]
- Llama Team; @ Meta, A.I. The Llama 3 Herd of Models. arXiv 2024, arXiv:2407.21783. [Google Scholar]
- Yang, A.; et al. Qwen2.5 Technical Report. arXiv 2024, arXiv:2412.15115. [Google Scholar]
- Allal, L.B.; Lozhkov, A.; Bakouch, E.; Blázquez, G.M.; Penedo, G.; Tunstall, L.; Marafioti, A.; Kydlíček, H.; et al. SmolLM2: When Smol Goes Big — Data-Centric Training of a Small Language Model. arXiv 2025, arXiv:2502.02737. [Google Scholar]
- Ramsauer, H.; Schäfl, B.; Lehner, J.; Seidl, P.; Widrich, M.; Adler, T.; Gruber, L.; Holzleitner, M.; Pavlović, M.; Sandve, G.K.; et al. Hopfield Networks Is All You Need. In Proceedings of the International Conference on Learning Representations, 2021. [Google Scholar]
- Zou, A.; Wang, Z.; Carlini, N.; Nasr, M.; Kolter, J.Z.; Fredrikson, M. Universal and Transferable Adversarial Attacks on Aligned Language Models. arXiv 2023, arXiv:2307.15043. [Google Scholar]
- Lin, S.; Hilton, J.; Evans, O. TruthfulQA: Measuring How Models Mimic Human Falsehoods. Proceedings of the Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics 2022, Volume 1, 3214–3252. [Google Scholar] [CrossRef]
- Sathish, S. prabodha: Recognition-Gated Workspace Steering for Language Models. GitHub repository, Apache-2.0 licence. Gate ledger, configurations, and research journal under gates/ and research/. 2026. [Google Scholar] [CrossRef] [PubMed]

















| Family | Checkpoint | Size | Role |
|---|---|---|---|
| Qwen3 | Qwen3-4B | 4B | primary plant, L1–L22 |
| Qwen3 | Qwen3.6-27B | 27B | loadability scaling, L1b |
| Nemotron | Nemotron-Mini-4B | 4B | doctrine development, twin |
| Gemma-2 | Gemma-2-2B | 2B | moat, works (L26) |
| Llama-3.2 | Llama-3.2-1B | 1B | moat, works (L26) |
| Qwen2.5 | Qwen2.5-1.5B | 1.5B | moat, fails (L26) |
| SmolLM2 | SmolLM2-1.7B | 1.7B | moat, backfires (L26) |
| Seed | Gated lift | Advantage | (nats) |
|---|---|---|---|
| 42 | 0.30 | ||
| 123 | 0.35 | ||
| 777 | 0.35 | ||
| 2024 | 0.35 | ||
| 31415 | 0.35 | ||
| 999 | 0.35 | ||
| mean | 0.342 |
| Arm | CSR | ASR | Refusal | |
|---|---|---|---|---|
| baseline | 0.000 | 0.800 | 0.200 | |
| prefill | 0.160 | 1.000 | 0.000 | |
| entropy-gated | 0.160 | 1.000 | 0.000 | |
| logit bias | 0.400 | 0.800 | 0.200 | |
| continuous | 0.960 | 1.000 | 0.000 |
| Model | Clean | ASR | Recognition-gated | Moat | |
|---|---|---|---|---|---|
| gap | none | ASR | OR | ||
| Gemma-2-2B | yes | 0.500 | 0.250 | 0.00 | works |
| Llama-3.2-1B | yes | 0.250 | 0.083 | 0.00 | works |
| Qwen2.5-1.5B | no | 0.333 | 0.333 | 0.20 | fails |
| SmolLM2-1.7B | no | 0.583 | 0.917 | 0.00 | backfires |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).