Submitted:
18 July 2026
Posted:
20 July 2026
You are already at the latest version
Abstract
Keywords:
MSC: https://github.com/xufengduan/Awesome-Language-MI-Survey
1. Introduction
Scope and Audience

2. Interpretability Methodology
2.1. Vocabulary Projection
Linguistic Leverage
2.2. Activation Patching
Linguistic Leverage
2.3. Sparse Autoencoders
Linguistic Leverage
2.4. Circuit Discovery
Linguistic Leverage
2.5. Activation Steering
2.5.0.6. Linguistic Leverage
3. MI on Linguistic Competence
3.1. Dynamic View: How It Emerges
What We Learn
3.2. Mechanistic View: Specialized vs. Shared Circuits
What We Learn
3.3. Cognitive View: Partial Alignment with Human Language Processing
What We Learn
4. Multilingualism
4.1. Language-Selective Components
What We Learn
4.2. Cross-Lingual Transfer
What We Learn
4.3. Mechanisms of Semantic Alignment
What We Learn
5. Challenges and Open Questions
5.1. Scaling vs. Interpretability Trade-Off
5.2. Evaluation Gaps
5.3. Causality and Model Editing
5.4. Cognitive Plausibility and Human-Like Generalization
5.5. Ethical and Societal Dimensions
5.6. Implications of Synthetic Data and Model Collapse
6. Conclusion
Limitations
Appendix A
Appendix A.1. Related Works
Appendix A.1.1. Foundational Interpretability Studies
Appendix A.1.2. Benchmarks and Linguistic Probing
Appendix A.1.3. Shift from Model-Agnostic to Model-Specific Interpretability
Appendix A.1.4. Running Example: Subject–Verb Agreement

| Dataset | Language | Size | Paradigms |
|---|---|---|---|
| BLiMP [164] | English | 67k | 67 |
| CLiMP [172] | Chinese | 16k | 16 |
| SLING [147] | Chinese | 38k | 38 |
| ZhoBLiMP [103] | Chinese | 35k | 118 |
| JBLiMP [145] | Japanese | 331 | 39 |
| RuBLiMP [152] | Russian | 45k | 45 |
| BLiMP-NL [150] | Dutch | 8.4k | 84 |
| QFrBLiMP [12] | Quebec-French | 1761 | 20 |
| TurBLiMP [11] | Turkish | 16k | 16 |
| UrBLiMP [1] | Urdu | 5696 | 19 |
| Irish-BLiMP [108] | Irish | 1020 | 102 |
| Arabic MPs [4] | Arabic | 3000 | 9 |
| BLiMP-IT [9] | Italian | 2899 | 78 |
| CLAMS [118] | English, French, German, Hebrew and Russian | 229.9k | 7 |
| MultiBLiMP 1.0 [81] | 101 languages | 128321 | 2 |
| BHS [90] | Basque, Hindi, Swahili | 300 | 3 |
| Methodology | Representative Works | Key Idea | Strengths and Limitations |
| Methodology | Representative Works | Key Idea | Strengths and Limitations |
| Vocabulary Projection | Logit Lens [50,124]; Tuned Lens [14]; Future Lens [128]. | Decodes intermediate hidden states into vocabulary tokens to reveal the model’s evolving predictions layer-by-layer. | Strengths: Maps word vectors across languages rapidly; no model retraining needed. Limitations: Ineffective for distant languages; performance depends on initial dictionary quality; struggles with new words or complex tasks. |
| Patching | Activation Patching/Attribution Patching [7,43,110,121,151,177] | Identifies causally decisive components by swapping activations between inputs or using gradient-based approximations. | Strengths: Pinpoints layers/heads encoding key information; enables transplant experiments; shows causal roles of activations. Limitations: Sensitive to superposition; computationally costly. |
| Sparse Autoencoders | SAE, Transcoder, Cross-layer Transcoder, Crosscoder and Binary Autoencoder [16,26,39,76] | Decomposes polysemantic neurons into sparse, interpretable feature directions, resolving superposition. | Strengths: Disentangles polysemantic neurons; reveals interpretable subspaces; quantifies feature entanglement. Limitations: Requires training or specialized modeling; imperfect disentanglement. |
| Automated Circuit Discovery | Circuit Tracing [5,32,107] | Automatically finds minimal subgraphs (component- or feature-level) that implement specific tasks. | Strengths: Traces information flow via minimal subgraphs; reveals emergent “circuit reuse”; mechanistically grounded. Limitations: Hard to automate fully; scalability remains challenging. |
| Activation Steering | ActAdd, CAA, STA [42,110,111,134,157,159,163] | Manipulates model behavior via inference-time activation injection (Steering) or permanent weight modification (Editing). | Strengths: Directly alters internal behavior; useful for debiasing and control; connects interpretability with intervention. Limitations: Can disrupt unrelated behaviors; hard to find minimal safe edits. |
| Paper Title | Type | Venue |
| Benchmarks | ||
| MIB: A Mechanistic Interpretability Benchmark [117] | Method | ICML |
| AXBENCH: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders [169] | Method | ICML |
| SAEBench: A Comprehensive Benchmark for Sparse Autoencoders [84] | Method | ICML |
| InterpBench: Semi-Synthetic Transformers for Evaluating MI Techniques [56] | Method | NIPS |
| Thank You, Stingray: Multilingual Large Language Models Can Not (Yet) Disambiguate Cross-Lingual Word Senses [18] | Method | NAACL |
| CausalGym: Benchmarking causal interpretability methods on linguistic tasks [7] | Method | ACL |
| Holmes: A Benchmark to Assess the Linguistic Competence of Language Models [160] | Method | TACL |
| SyntaxGym: An Online Platform for Targeted Evaluation of Language Models [49] | Method | ACL |
| Probing | ||
| Lexical Popularity: Quantifying the Impact of Pre-training for LLM Performance [136] | Method | Arxiv |
| Steering Embedding Models with Geometric Rotation: Mapping Semantic Relationships Across Languages and Models [46] | Method | Arxiv |
| Emergence of Phonemic, Syntactic, and Semantic Representations in Artificial Neural Networks [126] | Application | Arxiv |
| ShifCon: Enhancing Non-Dominant Language Capabilities with a Shift-based Multilingual Contrastive Framework [178] | Method | ACL |
| Lost in Multilinguality: Dissecting Cross-lingual Factual Inconsistency in Transformer Language Models [162] | Application | ACL |
| Probing Syntax in Large Language Models: Successes and Remaining Challenges [34] | Application | COLM |
| Probing Internal Representations of Multi-Word Verbs in Large Language Models [87] | Application | MWE |
| Can Cross-Lingual Transferability of Multilingual Transformers Be Activated Without End-Task Data? [24] | Method | ACL |
| Finding Neurons in a Haystack: Case Studies with Sparse Probing [61] | Application | Arxiv |
| Probing Classifiers: Promises, Shortcomings, and Advances [13] | Method | CL |
| First Align, then Predict: Understanding the Cross-Lingual Ability of Multilingual BERT [120] | Application | EACL |
| Finding Universal Grammatical Relations in Multilingual BERT [23] | Application | ACL |
| On the Language Neutrality of Pre-trained Multilingual Representations [95] | Application | EMNLP |
| Emergent Linguistic Structure in Artificial Neural Networks Trained by Self-supervision [106] | Application | PNAS |
| A Structural Probe for Finding Syntax in Word Representations [69] | Application | NAACL |
| Under the Hood: Using Diagnostic Classifiers to Investigate and Improve how Language Models Track Agreement Information [51] | Application | BlackboxNLP |
| Vocabulary Projection (Logit Lens) | ||
| Eliciting Latent Predictions from Transformers with the Tuned Lens [14] | Method | Arxiv |
| Future Lens: Anticipating Subsequent Tokens from a Single Hidden State [128] | Method | CoNLL |
| Interpreting GPT: the logit lens [124] | Method | Blog |
| Unraveling Syntax: How Language Models Learn Context-Free Grammars [138] | Application | Arxiv |
| The Semantic Hub Hypothesis: Language Models Share Semantic Representations [171] | Application | ICLR |
| Do Multilingual LLMs Think in English? [139] | Application | Workshop |
| Do Llamas Work in English? On the Latent Language of Multilingual Transformers [166] | Application | ACL |
| DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models [31] | Application | ICLR |
| Sparse AutoEncoders (SAE) | ||
| Are Sparse Autoencoders Useful? A Case Study in Sparse Probing [82] | Method | ICML |
| Binary Autoencoder for Mechanistic Interpretability of Large Language Models [26] | Method | Arxiv |
| LinguaLens: Towards Interpreting Linguistic Mechanisms via Sparse Auto-Encoder [79] | Method | EMNLP |
| SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models [67] | Method | Arxiv |
| A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious Minima [155] | Method | Arxiv |
| The Grounding Gap: How LLMs Anchor the Meaning of Abstract Concepts Differently from Humans [25] | Application | Arxiv |
| From Syntax to Emotion: A Mechanistic Analysis of Emotion Inference in LLMs [142] | Application | Arxiv |
| Cross-Architecture Model Diffing with Crosscoders: Unsupervised Discovery of Differences Between LLMs [80] | Application | Arxiv |
| What’s the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering [104] | Application | ICLR |
| Crosscoding Through Time: Tracking Emergence of Linguistic Representations [10] | Application | ICML |
| Large Language Models Share Representations of Latent Grammatical Concepts [17] | Application | NAACL |
| Incremental Sentence Processing Mechanisms in Autoregressive Transformer Language Models [62] | Application | NAACL |
| Sparse Autoencoders Can Capture Language-Specific Concepts Across Diverse Languages [6] | Application | Arxiv |
| Unveiling Language-Specific Features via Sparse Autoencoders [33] | Application | ACL |
| Analyzing Multilingualism in Large Language Models with Sparse Autoencoders [27] | Application | COLM |
| Sparse Autoencoders Can Capture Language-Specific Concepts [6] | Application | Arxiv |
| Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders [64] | Application | Arxiv |
| Causal Language Control in Multilingual Transformers via Sparse Feature Steering [30] | Application | Arxiv |
| Semantic Convergence: Investigating Shared Representations Across Scaled LLMs [135] | Application | SRW |
| Extended Abstract for “Linguistic Universals”: Emergent Shared Features in Independent Monolingual Language Models via Sparse Autoencoders [187] | Application | MRL |
| Sparse Autoencoders Find Highly Interpretable Features in Language Models [76] | Method | ICLR |
| Transcoders Find Interpretable LLM Feature Circuits [38] | Method | NIPS |
| Activation Patching | ||
| Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context [57] | Method | Arxiv |
| How to use and interpret activation patching [68] | Method | Arxiv |
| Is This the Subspace You Are Looking for? An Interpretability Illusion [105] | Method | ICLR |
| Towards Best Practices of Activation Patching in Language Models [177] | Method | ICLR |
| Function Vectors in Large Language Models [157] | Method | ICLR |
| Attribution Patching Outperforms Automated Circuit Discovery [151] | Method | Blackbox |
| Attribution Patching: Activation Patching At Industrial Scale [121] | Method | BLOG |
| Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models [45] | Method | IJCNLP |
| The Dual-Route Model of Induction [44] | Application | COLM |
| Neuron Analysis | ||
| From Directions to Regions: Decomposing Activations in Language Models via Local Geometry[140] | Method | Arxiv |
| LANDeRMT: Dectecting and Routing Language-Aware Neurons for Selectively Finetuning LLMs to Machine Translation [189] | Method | ACL |
| Task-Specific Skill Localization in Fine-tuned Language Models [129] | Method | ICML |
| Language Arithmetics: Towards Systematic Language Neuron Identification and Manipulation [58] | Application | AACL |
| Cross-Lingual Generalization and Compression [133] | Application | ACL |
| How Syntax Specialization Emerges in Language Models [36] | Application | Arxiv |
| Language Lives in Sparse Dimensions: Interpretable Multilingual Control [186] | Application | Arxiv |
| Inducing Dyslexia in Vision Language Models [70] | Application | Arxiv |
| Different types of syntactic agreement recruit the same units [89] | Application | Arxiv |
| Sparse Subnetwork Enhancement for Underrepresented Languages in Large Language Models [59] | Application | Arxiv |
| Multilingual Knowledge Editing with Language-Agnostic Factual Neurons [181] | Application | COLING |
| From Language to Cognition: How LLMs Outgrow the Human Language Network [3] | Application | EMNLP |
| The Transfer Neurons Hypothesis [156] | Application | EMNLP |
| Mechanistic Understanding and Mitigation of Language Confusion in English-Centric Large Language Models [123] | Application | EMNLP |
| The Rise and Down of Babel Tower: Investigating the Evolution Process of Multilingual Code Large Language Model [22] | Application | ICLR |
| The LLM Language Network: A Neuroscientific Approach [2] | Application | NAACL |
| Language-Specific Neurons Do Not Facilitate Cross-Lingual Transfer [114] | Application | Workshop |
| Language-Specific Neurons: The Key to Multilingual Capabilities [154] | Application | ACL |
| Unveiling Linguistic Regions in Large Language Models [182] | Application | ACL |
| Converging to a Lingua Franca: Evolution of Linguistic Regions [175] | Application | COLING |
| Unveiling Language Competence Neurons: A Psycholinguistic Approach [37] | Application | COLING |
| Linguistic Minimal Pairs Elicit Linguistic Similarity [188] | Application | COLING |
| Neuron-Level Knowledge Attribution in Large Language Models [174] | Application | EMNLP |
| Neuron Specialization: Leveraging Intrinsic Task Modularity for Multilingual Machine Translation [153] | Application | EMNLP |
| Decoding Probing: Revealing Internal Linguistic Structures [65] | Application | LREC |
| On the Multilingual Ability: Finding and Controlling Language-Specific Neurons [88] | Application | NAACL |
| How do Large Language Models Handle Multilingualism? [184] | Application | NIPS |
| Universal Neurons in GPT2 Language Models [60] | Application | TMLR |
| Same Neurons, Different Languages [148] | Application | NAACL |
| Importance-based Neuron Allocation for Multilingual Neural Machine Translation [173] | Application | ACL |
| Circuit Discovery | ||
| Circuit Tracing: Revealing Computational Graphs in Language Models [5] | Method | Anthropic |
| Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs [107] | Method | ICLR |
| Scaling Sparse Feature Circuits For Studying In-Context Learning [86] | Method | ICML |
| Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPT [66] | Method | Arxiv |
| Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small [161] | Method | ICLR |
| Towards Automated Circuit Discovery for Mechanistic Interpretability [32] | Method | NIPS |
| Between Circuits and Chomsky: Pre-pretraining on Formal Languages [74] | Application | ACL |
| The Same But Different: Structural Similarities in Multilingual LM [180] | Application | ICLR |
| Circuit Component Reuse Across Tasks in Transformer Language Models [112] | Application | ICLR |
B. Details of Findings
B.1. Emergence of Formal Linguistic Competence

B.2. Partial Alignment with Human Language Processing

B.4. Language-Selective Components

B.5. Cross-Lingual Transfer

B.6. Mechanisms of Semantic Alignment

References
- Adeeba, F.; Dillon, B.; Sajjad, H.; Bhatt, R. UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu. Urblimp: A benchmark for evaluating the linguistic competence of large language models in urdu. 2025. Available online: https://arxiv.org/abs/2508.01006.
- AlKhamissi, B.; Tuckute, G.; Bosselut, A.; Schrimpf, M. The LLM Language Network: A Neuroscientific Approach for Identifying Causally Task-Relevant Units. The llm language network: A neuroscientific approach for identifying causally task-relevant units. 2025. Available online: https://arxiv.org/abs/2411.02280.
- AlKhamissi, B.; Tuckute, G.; Tang, Y.; Binhuraib, T.; Bosselut, A.; Schrimpf, M. From Language to Cognition: How LLMs Outgrow the Human Language Network. From language to cognition: How llms outgrow the human language network. 2025. Available online: https://arxiv.org/abs/2503.01830.
- Alrajhi, W. A.; Al-Khalifa, H.; AlSalman, A. Assessing the Linguistic Knowledge in Arabic Pre-trained Language Models Using Minimal Pairs. H. Bouamor et al. (), Proceedings of the Seventh Arabic Natural Language Processing Workshop (WANLP) Proceedings of the seventh arabic natural language processing workshop (wanlp) ( 185–193). Abu Dhabi, United Arab Emirates (Hybrid)Association for Computational Linguistics, 202212; Available online: https://aclanthology.org/2022.wanlp-1.17/. [CrossRef]
- Ameisen, E.; Lindsey, J.; Pearce, A.; Gurnee, W.; Turner, N. L.; Chen, B.; Citro, C.; Abrahams, D.; Carter, S.; Hosmer, B.; Marcus, J.; Sklar, M.; Templeton, A.; Bricken, T.; McDougall, C.; Cunningham, H.; Henighan, T.; Jermyn, A.; Jones, A.; Batson, J. Circuit Tracing: Revealing Computational Graphs in Language Models. Anthropic Technical Report. March 2025M. Available online: https://transformer-circuits.pub/2025/attribution-graphs/methods.html.
- Andrylie, L. M.; Rahmanisa, I.; Ihsani, M. K.; Wicaksono, A. F.; Wibowo, H. A.; Aji, A. F. Sparse Autoencoders Can Capture Language-Specific Concepts Across Diverse Languages. Sparse autoencoders can capture language-specific concepts across diverse languages. 2025. Available online: https://arxiv.org/abs/2507.11230.
- Arora, A.; Jurafsky, D.; Potts, C. CausalGym: Benchmarking causal interpretability methods on linguistic tasks. Causalgym: Benchmarking causal interpretability methods on linguistic tasks. 2024. Available online: https://arxiv.org/abs/2402.12560.
- Banayeeanzade, A.; Tak, A. N.; Bahrani, F.; Bolourani, A.; Blas, L.; Ferrara, E.; Gratch, J.; Karimireddy, S. P. Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness. Psychological steering in llms: An evaluation of effectiveness and trustworthiness. 2025. Available online: https://arxiv.org/abs/2510.04484.
- Barbini, M.; Piccini Bianchessi, M. L.; Bressan, V.; Fusco, A.; Neri, S.; Rossi, S.; Sgrizzi, T.; Chesi, C. BLiMP-IT: Harnessing Automatic Minimal Pair Generation for Italian Language Model Evaluation. In Proceedings of the Eleventh Italian Conference on Computational Linguistics (CLiC-it 2025) Proceedings of the eleventh italian conference on computational linguistics (clic-it 2025); Cagliari, ItalyCEUR Workshop Proceedings, Bosco, C., Jezek, E., Polignano, M., Sanguinetti (), M., Eds.; 09 2025; pp. 64–71. Available online: https://aclanthology.org/2025.clicit-1.8/.
- Bayazit, D.; Mueller, A.; Bosselut, A. Crosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM Pretraining. Crosscoding through time: Tracking emergence & consolidation of linguistic representations throughout llm pretraining. 2025. Available online: https://arxiv.org/abs/2509.05291.
- Başar, E.; Padovani, F.; Jumelet, J.; Bisazza, A. TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing Proceedings of the 2025 conference on empirical methods in natural language processing ( 16506–16521). Association for Computational Linguistics, Available online. 2025. [Google Scholar] [CrossRef]
- Beauchemin, D.; Veilleux, P.-L.; Khoury, R.; Roy, J.-P. QFrBLiMP: a Quebec-French Benchmark of Linguistic Minimal Pairs. Qfrblimp: a quebec-french benchmark of linguistic minimal pairs. 2025. Available online: https://arxiv.org/abs/2509.25664.
- Belinkov, Y. Computational Linguistics481207–219; Probing Classifiers: Promises, Shortcomings, and Advances. 03 2022. Available online: https://aclanthology.org/2022.cl-1.7/. [CrossRef]
- Belrose, N.; Ostrovsky, I.; McKinney, L.; Furman, Z.; Smith, L.; Halawi, D.; Biderman, S.; Steinhardt, J. Eliciting Latent Predictions from Transformers with the Tuned Lens. Eliciting latent predictions from transformers with the tuned lens. 2025. Available online: https://arxiv.org/abs/2303.08112.
- Bolhuis, J. J.; Crain, S.; Fong, S.; Moro, A. Three reasons why AI doesn’t model human language. Nature6278004489. Available online. 2024. [CrossRef]
- Bricken, T.; Templeton, A.; Batson, J.; Chen, B.; Jermyn, A.; Conerly, T.; Turner, N.; Anil, C.; Denison, C.; McCauley, H.; et al. Towards Monosemanticity: Decomposing Language Models With Dictionary Learning. Transformer Circuits Thread. 2023. Available online: https://transformer-circuits.pub/2023/monosemantic-features.
- Brinkmann, J.; Wendler, C.; Bartelt, C.; Mueller, A. Large Language Models Share Representations of Latent Grammatical Concepts Across Typologically Diverse Languages. L. Chiruzzo, A. Ritter, L. Wang (), Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies. In Long Papers) Proceedings of the 2025 conference of the nations of the americas chapter of the association for computational linguistics: Human language technologies (volume 1: Long papers) ( 6131–6150), 202504; Albuquerque, New MexicoAssociation for Computational Linguistics; Volume 1. Available online: https://aclanthology.org/2025.naacl-long.312/. [CrossRef]
- Cahyawijaya, S.; Zhang, R.; Cruz, J. C. B.; Lovenia, H.; Gilbert, E.; Nomoto, H.; Aji, A. F. Thank You, Stingray: Multilingual Large Language Models Can Not (Yet) Disambiguate Cross-Lingual Word Senses. L. Chiruzzo, A. Ritter, L. Wang (), Findings of the Association for Computational Linguistics: NAACL 2025 Findings of the association for computational linguistics: Naacl 2025 ( 3228–3250); Albuquerque, New MexicoAssociation for Computational Linguistics, 04 2025; Available online: https://aclanthology.org/2025.findings-naacl.178/. [CrossRef]
- Cai, Z. G.; Duan, X.; Haslett, D. A.; Wang, S.; Pickering, M. J. Do large language models resemble humans in language use? Do large language models resemble humans in language use? 2024. Available online: https://arxiv.org/abs/2303.08014.
- Chang, T. A.; Arnett, C.; Tu, Z.; Bergen, B. When Is Multilinguality a Curse? Language Modeling for 250 High- and Low-Resource Languages. Y. Al-Onaizan, M. Bansal, Y.-N. Chen (), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing Proceedings of the 2024 conference on empirical methods in natural language processing ( 4074–4096). Miami, Florida, USAAssociation for Computational Linguistics, 202411; Available online: https://aclanthology.org/2024.emnlp-main.236/. [CrossRef]
- Chang, T. A.; Tu, Z.; Bergen, B. K. Characterizing Learning Curves During Language Model Pre-Training: Learning, Forgetting, and Stability. Trans. Assoc. Comput. Linguist. 2024, 121346–1362. Available online: https://doi.org/10.1162/tacl_a_00708. [CrossRef]
- Chen, J.; Chen, W.; Su, J.; Xu, J.; Lin, H.; Ren, M.; Lu, Y.; Han, X.; Sun, L. The Rise and Down of Babel Tower: Investigating the Evolution Process of Multilingual Code Large Language Model. The Thirteenth International Conference on Learning Representations. The thirteenth international conference on learning representations., 2025; Available online: https://openreview.net/forum?id=eznTVIM3bs.
- Chi, E. A.; Hewitt, J.; Manning, C. D. Finding Universal Grammatical Relations in Multilingual BERT. Finding universal grammatical relations in multilingual bert. 2020. Available online: https://arxiv.org/abs/2005.04511.
- Chi, Z.; Huang, H.; Mao, X.-L. Can Cross-Lingual Transferability of Multilingual Transformers Be Activated Without End-Task Data? A. Rogers, J. Boyd-Graber, N. Okazaki (), Findings of the Association for Computational Linguistics: ACL 2023 Findings of the association for computational linguistics: Acl 2023 ( 12572–12584). Toronto, CanadaAssociation for Computational Linguistics. 07 2023. Available online: https://aclanthology.org/2023.findings-acl.796/. [CrossRef]
- Chlapanis, O. S.; Mastromichalakis, O. M.; Papadimitriou, C. H. The Grounding Gap: How LLMs Anchor the Meaning of Abstract Concepts Differently from Humans. The grounding gap: How llms anchor the meaning of abstract concepts differently from humans. 2026. Available online: https://arxiv.org/abs/2605.08837.
- Cho, H.; Yang, H.; Kurkoski, B. M.; Inoue, N. Binary Autoencoder for Mechanistic Interpretability of Large Language Models. Binary autoencoder for mechanistic interpretability of large language models. 2025. Available online: https://arxiv.org/abs/2509.20997.
- Cho, I.; Hockenmaier, J. Analyzing Multilingualism in Large Language Models with Sparse Autoencoders. Second Conference on Language Modeling. Second conference on language modeling., 2025; Available online: https://openreview.net/forum?id=NmGSvZoU3K.
- Chomsky, N. Noam Chomsky: The False Promise of ChatGPT. The New York Times; Opinion, 2023M. Available online: https://www.nytimes.com/2023/03/08/opinion/noam-chomsky-chatgpt-ai.html.
- Choshen, L.; Hacohen, G.; Weinshall, D.; Abend, O. The Grammar-Learning Trajectories of Neural Language Models. The grammar-learning trajectories of neural language models. 2022. Available online: https://arxiv.org/abs/2109.06096.
- Chou, C.-T.; Liu, G.; Sun, J.; Blondin, C.; Zhu, K.; Sharma, V.; O’Brien, S. Causal Language Control in Multilingual Transformers via Sparse Feature Steering. Causal language control in multilingual transformers via sparse feature steering. 2025. Available online: https://arxiv.org/abs/2507.13410.
- Chuang, Y.-S.; Xie, Y.; Luo, H.; Kim, Y.; Glass, J.; He, P. DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models. Dola: Decoding by contrasting layers improves factuality in large language models. 2024. Available online: https://arxiv.org/abs/2309.03883.
- Conmy, A.; Mavor-Parker, A. N.; Lynch, A.; Heimersheim, S.; Garriga-Alonso, A. Towards Automated Circuit Discovery for Mechanistic Interpretability. Thirty-seventh Conference on Neural Information Processing Systems. Thirty-seventh conference on neural information processing systems., 2023; Available online: https://openreview.net/forum?id=89ia77nZ8u.
- Deng, B.; Wan, Y.; Yang, B.; Zhang, Y.; Feng, F. Unveiling Language-Specific Features in Large Language Models via Sparse Autoencoders; Che, W., Nabende, J., Shutova, E., Pilehvar (), M. T., Eds.; Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 07 2025; Volume 1, Available online: https://aclanthology.org/2025.acl-long.229/. [CrossRef]
- Diego-Simón, P. J.; Chemla, E.; King, J.-R.; Lakretz, Y. Probing Syntax in Large Language Models: Successes and Remaining Challenges. Probing syntax in large language models: Successes and remaining challenges. 2025. Available online: https://arxiv.org/abs/2508.03211.
- Duan, X.; Xiao, B.; Tang, X.; Cai, Z. G. HLB: Benchmarking LLMs’ Humanlikeness in Language Use. Hlb: Benchmarking llms’ humanlikeness in language use. 2024. Available online: https://arxiv.org/abs/2409.15890.
- Duan, X.; Yao, Z.; Zhang, Y.; Wang, S.; Cai, Z. G. How Syntax Specialization Emerges in Language Models. How syntax specialization emerges in language models. 2025. Available online: https://arxiv.org/abs/2505.19548.
- Duan, X.; Zhou, X.; Xiao, B.; Cai, Z. Unveiling Language Competence Neurons: A Psycholinguistic Approach to Model Interpretability; Rambow, O., Wanner, L., Apidianaki, M., Al-Khalifa, H., Eugenio, B. D., Schockaert (), S., Eds.; Proceedings of the 31st International Conference on Computational Linguistics Proceedings of the 31st international conference on computational linguistics ( 10148–10157): Abu Dhabi, UAEAssociation for Computational Linguistics, 01 2025; Available online: https://aclanthology.org/2025.coling-main.677/.
- Dunefsky, J.; Chlenski, P.; Nanda, N. Transcoders Find Interpretable LLM Feature Circuits. Transcoders find interpretable llm feature circuits. 2024. Available online: https://arxiv.org/abs/2406.11944.
- Elhage, N.; Hume, T.; Olsson, C.; Schiefer, N.; Henighan, T.; Kravec, S.; Hatfield-Dodds, Z.; Lasenby, R.; Drain, D.; Chen, C.; Grosse, R.; McCandlish, S.; Kaplan, J.; Amodei, D.; Wattenberg, M.; Olah, C. Toy Models of Superposition. Toy models of superposition. 2022. Available online: https://arxiv.org/abs/2209.10652.
- Elhage, N.; Nanda, N.; Olsson, C.; Henighan, T.; Joseph, N.; Mann, B.; Askell, A.; Bai, Y.; Chen, A.; Conerly, T.; DasSarma, N.; Drain, D.; Ganguli, D.; Hatfield-Dodds, Z.; Hernandez, D.; Jones, A.; Kernion, J.; Lovitt, L.; Ndousse, K.; Olah, C. A Mathematical Framework for Transformer Circuits. Transformer Circuits Thread. 2021. Available online: https://transformer-circuits.pub/2021/framework/index.html.
- Evanson, L.; Lakretz, Y.; King, J.-R. Language acquisition: do children and language models follow similar learning stages? Language acquisition: do children and language models follow similar learning stages? 2023. Available online: https://arxiv.org/abs/2306.03586.
- Fang, J.; Jiang, H.; Wang, K.; Ma, Y.; Shi, J.; Wang, X.; He, X.; Chua, T.-S. AlphaEdit: Null-Space Constrained Model Editing for Language Models. The Thirteenth International Conference on Learning Representations. The thirteenth international conference on learning representations., 2025; Available online: https://openreview.net/forum?id=HvSytvg3Jh.
- Ferrando, J.; Voita, E. Information Flow Routes: Automatically Interpreting Language Models at Scale. Information flow routes: Automatically interpreting language models at scale. 2024. Available online: https://arxiv.org/abs/2403.00824.
- Feucht, S.; Todd, E.; Wallace, B.; Bau, D. The Dual-Route Model of Induction. The dual-route model of induction. 2025. Available online: https://arxiv.org/abs/2504.03022.
- Finlayson, M.; Mueller, A.; Gehrmann, S.; Shieber, S.; Linzen, T.; Belinkov, Y. Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models. C. Zong, F. Xia, W. Li, R. Navigli (), Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing. In Long Papers) Proceedings of the 59th annual meeting of the association for computational linguistics and the 11th international joint conference on natural language processing (volume 1: Long papers) ( 1828–1843); OnlineAssociation for Computational Linguistics, 08 2021; Volume 1, Available online: https://aclanthology.org/2021.acl-long.144/. [CrossRef]
- Freenor, M.; Alvarez, L. Mapping Semantic & Syntactic Relationships with Geometric Rotation. Mapping semantic & syntactic relationships with geometric rotation. 2026. Available online: https://arxiv.org/abs/2510.09790.
- Gao, C.; Ma, Z.; Chen, J.; Li, P.; Huang, S.; Li, J. Increasing alignment of large language models with language processing in the human brain. Nature computational science1–11. 2025. Available online: https://www.nature.com/articles/s43588-025-00863-0.
- Gao, Y.; Meng, Q.; Zhou, Y.; Pan, L. Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures. Towards intrinsic interpretability of large language models:a survey of design principles and architectures. 2026. Available online: https://arxiv.org/abs/2604.16042.
- Gauthier, J.; Hu, J.; Wilcox, E.; Qian, P.; Levy, R. SyntaxGym: An Online Platform for Targeted Evaluation of Language Models. A. Celikyilmaz T.-H. Wen (), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations Proceedings of the 58th annual meeting of the association for computational linguistics: System demonstrations ( 70–76). OnlineAssociation for Computational Linguistics. 07 2020. Available online: https://aclanthology.org/2020.acl-demos.10/. [CrossRef]
- Geva, M.; Caciularu, A.; Wang, K.; Goldberg, Y. Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing Proceedings of the 2022 conference on empirical methods in natural language processing ( 30–45); Abu Dhabi, United Arab EmiratesAssociation for Computational Linguistics, Goldberg, Y., Kozareva, Z., Zhang (), Y., Eds.; 12 2022; Available online: https://aclanthology.org/2022.emnlp-main.3/. [CrossRef]
- Giulianelli, M.; Harding, J.; Mohnert, F.; Hupkes, D.; Zuidema, W. Under the Hood: Using Diagnostic Classifiers to Investigate and Improve how Language Models Track Agreement Information. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP Proceedings of the 2018 EMNLP workshop BlackboxNLP: Analyzing and interpreting neural networks for NLP ( 240–248); Brussels, BelgiumAssociation for Computational Linguistics, Linzen, T., Chrupała, G., Alishahi (), A., Eds.; 11 2018; Available online: https://aclanthology.org/W18-5426/. [CrossRef]
- Goldstein, A.; Ham, E.; Schain, M.; Nastase, S. A.; Aubrey, B.; Zada, Z.; Grinstein-Dabush, A.; Gazula, H.; Feder, A.; Doyle, W.; et al. Temporal structure of natural language processing in the human brain corresponds to layered hierarchy of large language models. Nature communications16110529. Available online. 2025. [CrossRef]
- Graichen, N.; de Dios-Flores, I.; Boleda, G. The Grammar of Transformers: A Systematic Review of Interpretability Research on Syntactic Knowledge in Language Models. The grammar of transformers: A systematic review of interpretability research on syntactic knowledge in language models. 2026. Available online: https://arxiv.org/abs/2601.19926.
- Guo, Y.; Conia, S.; Zhou, Z.; Li, M.; Potdar, S.; Xiao, H. Do Large Language Models Have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMs. Do large language models have an english accent? evaluating and improving the naturalness of multilingual llms. 2025. Available online: https://arxiv.org/abs/2410.15956.
- Guo, Y.; Shang, G.; Vazirgiannis, M.; Clavel, C. The Curious Decline of Linguistic Diversity: Training Language Models on Synthetic Text. K. Duh, H. Gomez, S. Bethard (), Findings of the Association for Computational Linguistics: NAACL 2024 Findings of the association for computational linguistics: Naacl 2024 ( 3589–3604). Mexico City, MexicoAssociation for Computational Linguistics. 06 2024. Available online: https://aclanthology.org/2024.findings-naacl.228/. [CrossRef]
- Gupta, R.; Arcuschin, I.; Kwa, T.; Garriga-Alonso, A. InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques. Interpbench: Semi-synthetic transformers for evaluating mechanistic interpretability techniques. 2025. Available online: https://arxiv.org/abs/2407.14494.
- Gur-Arieh, Y.; Geva, M.; Geiger, A. Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context. Mixing mechanisms: How language models retrieve bound entities in-context. 2025. Available online: https://arxiv.org/abs/2510.06182.
- Gurgurov, D.; Trinley, K.; Ghussin, Y. A.; Baeumel, T.; van Genabith, J.; Ostermann, S. Language Arithmetics: Towards Systematic Language Neuron Identification and Manipulation. Language arithmetics: Towards systematic language neuron identification and manipulation. 2025. Available online: https://arxiv.org/abs/2507.22608.
- Gurgurov, D.; van Genabith, J.; Ostermann, S. Sparse Subnetwork Enhancement for Underrepresented Languages in Large Language Models. Sparse subnetwork enhancement for underrepresented languages in large language models. 2025. Available online: https://arxiv.org/abs/2510.13580.
- Gurnee, W.; Horsley, T.; Guo, Z. C.; Kheirkhah, T. R.; Sun, Q.; Hathaway, W.; Nanda, N.; Bertsimas, D. Universal Neurons in GPT2 Language Models. Transactions on Machine Learning Research. 2024. Available online: https://openreview.net/forum?id=ZeI104QZ8I.
- Gurnee, W.; Nanda, N.; Pauly, M.; Harvey, K.; Troitskii, D.; Bertsimas, D. Finding Neurons in a Haystack: Case Studies with Sparse Probing. Finding neurons in a haystack: Case studies with sparse probing. 2023. Available online: https://arxiv.org/abs/2305.01610.
- Hanna, M.; Mueller, A. Incremental Sentence Processing Mechanisms in Autoregressive Transformer Language Models. L. Chiruzzo, A. Ritter, L. Wang (), Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies. In Long Papers) Proceedings of the 2025 conference of the nations of the americas chapter of the association for computational linguistics: Human language technologies (volume 1: Long papers) ( 3181–3203), 202504; Albuquerque, New MexicoAssociation for Computational Linguistics; Volume 1. Available online: https://aclanthology.org/2025.naacl-long.164/. [CrossRef]
- Hao, S.; Linzen, T. Verb Conjugation in Transformers Is Determined by Linear Encodings of Subject Number. H. Bouamor, J. Pino, K. Bali (), Findings of the Association for Computational Linguistics: EMNLP 2023 Findings of the association for computational linguistics: Emnlp 2023 ( 4531–4539). SingaporeAssociation for Computational Linguistics. 12 2023. Available online: https://aclanthology.org/2023.findings-emnlp.300/. [CrossRef]
- Harrasse, A.; Draye, F.; Jin, Z.; Schölkopf, B. Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders. Tracing multilingual representations in llms with cross-layer transcoders. 2025. Available online: https://arxiv.org/abs/2511.10840.
- He, L.; Chen, P.; Nie, E.; Li, Y.; Brennan, J. R. Decoding Probing: Revealing Internal Linguistic Structures in Neural Language Models using Minimal Pairs. Decoding probing: Revealing internal linguistic structures in neural language models using minimal pairs. 2024. Available online: https://arxiv.org/abs/2403.17299.
- He, Z.; Ge, X.; Tang, Q.; Sun, T.; Cheng, Q.; Qiu, X. Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPT. Dictionary learning improves patch-free circuit discovery in mechanistic interpretability: A case study on othello-gpt. 2024. Available online: https://arxiv.org/abs/2402.12201.
- He, Z.; Zhao, H.; Qiao, Y.; Yang, F.; Payani, A.; Ma, J.; Du, M. SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models. Saif: A sparse autoencoder framework for interpreting and steering instruction following of language models. 2025. Available online: https://arxiv.org/abs/2502.11356.
- Heimersheim, S.; Nanda, N. How to use and interpret activation patching. How to use and interpret activation patching. 2024. Available online: https://arxiv.org/abs/2404.15255.
- Hewitt, J.; Manning, C. D. A Structural Probe for Finding Syntax in Word Representations; Burstein, J., Doran, C., Solorio (), T., Eds.; Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 06 2019; Volume 1, Available online: https://aclanthology.org/N19-1419/. [CrossRef]
- Honarmand, M.; Sharma, A.; AlKhamissi, B.; Mehrer, J.; Schrimpf, M. Inducing Dyslexia in Vision Language Models. Inducing dyslexia in vision language models. 2025. Available online: https://arxiv.org/abs/2509.24597.
- Hosseini, E. A.; Schrimpf, M.; Zhang, Y.; Bowman, S.; Zaslavsky, N.; Fedorenko, E. Artificial Neural Network Language Models Predict Human Brain Responses to Language Even After a Developmentally Realistic Amount of Training. Neurobiol. Lang. Available online. 2024, 5143–63. [Google Scholar] [CrossRef] [PubMed]
- Hu, J.; Gauthier, J.; Qian, P.; Wilcox, E.; Levy, R. A Systematic Assessment of Syntactic Generalization in Neural Language Models. In OnlineAssociation for Computational Linguistics; Jurafsky, D., Chai, J., Schluter, N., Tetreault (), J., Eds.; Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics Proceedings of the 58th annual meeting of the association for computational linguistics ( 1725–1744), 07 2020; Available online: https://aclanthology.org/2020.acl-main.158/. [CrossRef]
- Hu, J.; Levy, R. P. Prompting is not a substitute for probability measurements in large language models. The 2023 Conference on Empirical Methods in Natural Language Processing. The 2023 conference on empirical methods in natural language processing., 2023; Available online: https://openreview.net/forum?id=hMqRphmoM9.
- Hu, M. Y.; Petty, J.; Shi, C.; Merrill, W.; Linzen, T. Between Circuits and Chomsky: Pre-pretraining on Formal Languages Imparts Linguistic Biases; Che, W., Nabende, J., Shutova, E., Pilehvar (), M. T., Eds.; Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 07 2025; Volume 1, Available online: https://aclanthology.org/2025.acl-long.478/. [CrossRef]
- Huang, J.; Geiger, A.; D’Oosterlinck, K.; Wu, Z.; Potts, C. Rigorously Assessing Natural Language Explanations of Neurons. Rigorously assessing natural language explanations of neurons. 2023. Available online: https://arxiv.org/abs/2309.10312.
- Huben, R.; Cunningham, H.; Smith, L. R.; Ewart, A.; Sharkey, L. Sparse Autoencoders Find Highly Interpretable Features in Language Models. The Twelfth International Conference on Learning Representations. The twelfth international conference on learning representations., 2024; Available online: https://openreview.net/forum?id=F76bwRSLeK.
- Huo, J.; Yan, Y.; Hu, B.; Yue, Y.; Hu, X. MMNeuron: Discovering Neuron-Level Domain-Specific Interpretation in Multimodal Large Language Model. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing Proceedings of the 2024 conference on empirical methods in natural language processing ( 6801–6816), Available online. 2024; Association for Computational Linguistics. [Google Scholar] [CrossRef]
- Iqbal, A.; Younas, M.; Iftikhar, S.; Fatima, F.; Saleem, R. Spam detection using hybrid model on fusion of spammer behavior and linguistics features. Egyptian Informatics Journal29100605. 2025. Available online: https://www.sciencedirect.com/science/article/pii/S1110866524001683. [CrossRef]
- Jing, Y.; Yao, Z.; Guo, H.; Ran, L.; Wang, X.; Hou, L.; Li, J. LinguaLens: Towards Interpreting Linguistic Mechanisms of Large Language Models via Sparse Auto-Encoder. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing Proceedings of the 2025 conference on empirical methods in natural language processing ( 28232–28251); Suzhou, ChinaAssociation for Computational Linguistics, Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng (), V., Eds.; 11 2025; Available online: https://aclanthology.org/2025.emnlp-main.1433/. [CrossRef]
- Jiralerspong, T.; Bricken, T. Cross-Architecture Model Diffing with Crosscoders: Unsupervised Discovery of Differences Between LLMs. Cross-architecture model diffing with crosscoders: Unsupervised discovery of differences between llms. 2026. Available online: https://arxiv.org/abs/2602.11729.
- Jumelet, J.; Weissweiler, L.; Nivre, J.; Bisazza, A. MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs. Multiblimp 1.0: A massively multilingual benchmark of linguistic minimal pairs. 2025. Available online: https://arxiv.org/abs/2504.02768.
- Kantamneni, S.; Engels, J.; Rajamanoharan, S.; Tegmark, M.; Nanda, N. Are Sparse Autoencoders Useful? A Case Study in Sparse Probing. Are sparse autoencoders useful? a case study in sparse probing. 2025. Available online: https://arxiv.org/abs/2502.16681.
- Karny, S.; Baez, A.; Pataranutaporn, P. Neural Transparency: Mechanistic Interpretability Interfaces for Anticipating Model Behaviors for Personalized AI. Neural transparency: Mechanistic interpretability interfaces for anticipating model behaviors for personalized ai. 2025. Available online: https://arxiv.org/abs/2511.00230.
- Karvonen, A.; Rager, C.; Lin, J.; Tigges, C.; Bloom, J.; Chanin, D.; Lau, Y.-T.; Farrell, E.; McDougall, C.; Ayonrinde, K.; Till, D.; Wearden, M.; Conmy, A.; Marks, S.; Nanda, N. SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability. Saebench: A comprehensive benchmark for sparse autoencoders in language model interpretability. 2025. Available online: https://arxiv.org/abs/2503.09532.
- Kendiukhov, I. A Review of Developmental Interpretability in Large Language Models. A review of developmental interpretability in large language models. 2025. Available online: https://arxiv.org/abs/2508.15841.
- Kharlapenko, D.; Shabalin, S.; Barez, F.; Conmy, A.; Nanda, N. Scaling sparse feature circuit finding for in-context learning. Scaling sparse feature circuit finding for in-context learning. 2025. Available online: https://arxiv.org/abs/2504.13756.
- Kissane, H.; Schilling, A.; Krauss, P. Probing Internal Representations of Multi-Word Verbs in Large Language Models. A. K. Ojha et al. (), Proceedings of the 21st Workshop on Multiword Expressions (MWE 2025) Proceedings of the 21st workshop on multiword expressions (mwe 2025) ( 7–13). Albuquerque, New Mexico, U.S.A.Association for Computational Linguistics, 202505; Available online: https://aclanthology.org/2025.mwe-1.2/. [CrossRef]
- Kojima, T.; Okimura, I.; Iwasawa, Y.; Yanaka, H.; Matsuo, Y. On the Multilingual Ability of Decoder-based Pre-trained Language Models: Finding and Controlling Language-Specific Neurons. On the multilingual ability of decoder-based pre-trained language models: Finding and controlling language-specific neurons. 2024. Available online: https://arxiv.org/abs/2404.02431.
- Kryvosheieva, D.; de Varda, A.; Fedorenko, E.; Tuckute, G. Different types of syntactic agreement recruit the same units within large language models. Different types of syntactic agreement recruit the same units within large language models. 2025. Available online: https://arxiv.org/abs/2512.03676.
- Kryvosheieva, D.; Levy, R. Controlled Evaluation of Syntactic Knowledge in Multilingual Language Models. Controlled evaluation of syntactic knowledge in multilingual language models. 2024. Available online: https://arxiv.org/abs/2411.07474.
- Kumar, S.; Sumers, T. R.; Yamakoshi, T.; Goldstein, A.; Hasson, U.; Norman, K. A.; Griffiths, T. L.; Hawkins, R. D.; Nastase, S. A. Shared functional specialization in transformer-based language models and the human brain. Nature communications1515523. Available online. 2024. [CrossRef]
- Lakretz, Y.; Kruszewski, G.; Desbordes, T.; Hupkes, D.; Dehaene, S.; Baroni, M. The emergence of number and syntax units in LSTM language models. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies; Burstein, J., Doran, C., Solorio (), T., Eds.; Minneapolis, MinnesotaAssociation for Computational Linguistics, 06 2019; Volume 1, Available online: https://aclanthology.org/N19-1002/. [CrossRef]
- Leong, W. Q.; Ngui, J. G.; Susanto, Y.; Rengarajan, H.; Sarveswaran, K.; Tjhi, W. C. BHASA: A Holistic Southeast Asian Linguistic and Cultural Evaluation Suite for Large Language Models. Bhasa: A holistic southeast asian linguistic and cultural evaluation suite for large language models. 2023. Available online: https://arxiv.org/abs/2309.06085.
- Li, X.; Yong, Z.-X.; Bach, S. H. Preference Tuning For Toxicity Mitigation Generalizes Across Languages. Preference tuning for toxicity mitigation generalizes across languages. 2024. Available online: https://arxiv.org/abs/2406.16235.
- Libovický, J.; Rosa, R.; Fraser, A. On the Language Neutrality of Pre-trained Multilingual Representations. T. Cohn, Y. He, Y. Liu (), Findings of the Association for Computational Linguistics: EMNLP 2020 Findings of the association for computational linguistics: Emnlp 2020 ( 1663–1674). OnlineAssociation for Computational Linguistics. 11 2020. Available online: https://aclanthology.org/2020.findings-emnlp.150/. [CrossRef]
- Lieberum, T.; Rahtz, M.; Kramár, J.; Nanda, N.; Irving, G.; Shah, R.; Mikulik, V. Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla. Does circuit analysis interpretability scale? evidence from multiple choice capabilities in chinchilla. 2023. Available online: https://arxiv.org/abs/2307.09458.
- Lin, J. Neuronpedia: Interactive Reference and Tooling for Analyzing Neural Networks. Neuronpedia: Interactive reference and tooling for analyzing neural networks. 2023. Available online: https://www.neuronpedia.org.
- Lindsey, J.; Gurnee, W.; Ameisen, E.; Chen, B.; Pearce, A.; Turner, N. L.; Citro, C.; Abrahams, D.; Carter, S.; Hosmer, B.; Marcus, J.; Sklar, M.; Templeton, A.; Bricken, T.; McDougall, C.; Cunningham, H.; Henighan, T.; Jermyn, A.; Jones, A.; Batson, J. On the Biology of a Large Language Model. Anthropic Technical Report. March 2025M. Available online: https://transformer-circuits.pub/2025/attribution-graphs/biology.html.
- Linzen, T.; Dupoux, E.; Goldberg, Y. Assessing the Ability of LSTMs to Learn Syntax-Sensitive Dependencies. Transactions of the Association for Computational Linguistics. 2016, pp. 4521–535. Available online: https://aclanthology.org/Q16-1037/. [CrossRef]
- Liu, W.; Xiang, M.; Ding, N. Active use of latent tree-structured sentence representation in humans and large language models. Nature Human Behaviour1–14. Available online. 2025. [CrossRef]
- Liu, W.; Xu, Y.; Xu, H.; Chen, J.; Hu, X.; Wu, J. Unraveling Babel: Exploring Multilingual Activation Patterns of LLMs and Their Applications. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing Proceedings of the 2024 conference on empirical methods in natural language processing ( 11855–11881); Miami, Florida, USAAssociation for Computational Linguistics, Al-Onaizan, Y., Bansal, M., Chen (), Y.-N., Eds.; 11 2024; Available online: https://aclanthology.org/2024.emnlp-main.662/. [CrossRef]
- Liu, X.; Song, Q.; Zhou, Q.; Du, H.; Xu, S.; Jiang, W.; Zhang, W.; Jia, X. Focusing on Language: Revealing and Exploiting Language Attention Heads in Multilingual Large Language Models. Focusing on language: Revealing and exploiting language attention heads in multilingual large language models. 2025. Available online: https://arxiv.org/abs/2511.07498.
- Liu, Y.; Shen, Y.; Zhu, H.; Xu, L.; Qian, Z.; Song, S.; Zhang, K.; Tang, J.; Zhang, P.; Yang, B.; Wang, R.; Hu, H. A Systematic Assessment of Language Models with Linguistic Minimal Pairs in Chinese. A systematic assessment of language models with linguistic minimal pairs in chinese. 2025. Available online: https://arxiv.org/abs/2411.06096.
- Maar, J.; Paperno, D.; McDougall, C. S.; Nanda, N. What’s the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering. What’s the plan? metrics for implicit planning in llms and their application to rhyme generation and question answering. 2026. Available online: https://arxiv.org/abs/2601.20164.
- Makelov, A.; Lange, G.; Geiger, A.; Nanda, N. Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching. The Twelfth International Conference on Learning Representations. The twelfth international conference on learning representations., 2024; Available online: https://openreview.net/forum?id=Ebt7JgMHv1.
- Manning, C. D.; Clark, K.; Hewitt, J.; Khandelwal, U.; Levy, O. Emergent linguistic structure in artificial neural networks trained by self-supervision. In Proceedings of the National Academy of Sciences, Available online. 2020; pp. 1174830046–30054. [Google Scholar] [CrossRef] [PubMed]
- Marks, S.; Rager, C.; Michaud, E. J.; Belinkov, Y.; Bau, D.; Mueller, A. Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models. The Thirteenth International Conference on Learning Representations. The thirteenth international conference on learning representations., 2025; Available online: https://openreview.net/forum?id=I4e82CIDxv.
- McGiff, J.; Tran, K.-T.; Mulcahy, W.; Luinín, D. Ó.; Dalzell, J.; Bhroin, R. N.; Burke, A.; O’Sullivan, B.; Nguyen, H. D.; Nikolov, N. S. Irish-BLiMP: A Linguistic Benchmark for Evaluating Human and Language Model Performance in a Low-Resource Setting. Irish-blimp: A linguistic benchmark for evaluating human and language model performance in a low-resource setting. 2025. Available online: https://arxiv.org/abs/2510.20957.
- Méloux, M.; Maniu, S.; Portet, F.; Peyrard, M. Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable? The Thirteenth International Conference on Learning Representations. The thirteenth international conference on learning representations. 2025. Available online: https://openreview.net/forum?id=5IWJBStfU7.
- Meng, K.; Bau, D.; Andonian, A.; Belinkov, Y. Locating and Editing Factual Associations in GPT. Locating and editing factual associations in gpt. 2023. Available online: https://arxiv.org/abs/2202.05262.
- Meng, K.; Sharma, A. S.; Andonian, A.; Belinkov, Y.; Bau, D. Mass-Editing Memory in a Transformer. Mass-editing memory in a transformer. 2023. Available online: https://arxiv.org/abs/2210.07229.
- Merullo, J.; Eickhoff, C.; Pavlick, E. Circuit Component Reuse Across Tasks in Transformer Language Models. The Twelfth International Conference on Learning Representations. The twelfth international conference on learning representations., 2024; Available online: https://openreview.net/forum?id=fpoAYV6Wsk.
- Mohebbi, H.; Jumelet, J.; Hanna, M.; Alishahi, A.; Zuidema, W. Transformer-specific Interpretability; M. Mesgar S. Loáiciga (), Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Tutorial Abstracts Proceedings of the 18th conference of the european chapter of the association for computational linguistics: Tutorial abstracts ( 21–26). St. Julian’s, MaltaAssociation for Computational Linguistics, 03 2024; Available online: https://aclanthology.org/2024.eacl-tutorials.4/. [CrossRef]
- Mondal, S. K.; Sen, S.; Singhania, A.; Jyothi, P. Language-Specific Neurons Do Not Facilitate Cross-Lingual Transfer. In Shu (), The Sixth Workshop on Insights from Negative Results in NLP The sixth workshop on insights from negative results in nlp ( 46–62); Drozd, A., Sedoc, J., Tafreshi, S., A. Akula, R., Eds.; Albuquerque, New MexicoAssociation for Computational Linguistics, 05 2025; Available online: https://aclanthology.org/2025.insights-1.6/. [CrossRef]
- Mondorf, P.; Wang, M.; Gerstner, S.; Hakimi, A. D.; Liu, Y.; Veloso, L.; Zhou, S.; Schütze, H.; Plank, B. BlackboxNLP-2025 MIB Shared Task: Exploring Ensemble Strategies for Circuit Localization Methods. Blackboxnlp-2025 mib shared task: Exploring ensemble strategies for circuit localization methods. 2025. Available online: https://arxiv.org/abs/2510.06811.
- Mueller, A.; Brinkmann, J.; Li, M.; Marks, S.; Pal, K.; Prakash, N.; Rager, C.; Sankaranarayanan, A.; Sharma, A. S.; Sun, J.; Todd, E.; Bau, D.; Belinkov, Y. The Quest for the Right Mediator: Surveying Mechanistic Interpretability Through the Lens of Causal Mediation Analysis. The quest for the right mediator: Surveying mechanistic interpretability through the lens of causal mediation analysis. 2025. Available online: https://arxiv.org/abs/2408.01416.
- Mueller, A.; Geiger, A.; Wiegreffe, S.; Arad, D.; Arcuschin, I.; Belfki, A.; Chan, Y. S.; Fiotto-Kaufman, J.; Haklay, T.; Hanna, M.; Huang, J.; Gupta, R.; Nikankin, Y.; Orgad, H.; Prakash, N.; Reusch, A.; Sankaranarayanan, A.; Shao, S.; Stolfo, A.; Belinkov, Y. MIB: A Mechanistic Interpretability Benchmark. Mib: A mechanistic interpretability benchmark. 2025. Available online: https://arxiv.org/abs/2504.13151.
- Mueller, A.; Nicolai, G.; Petrou-Zeniou, P.; Talmina, N.; Linzen, T. Cross-Linguistic Syntactic Evaluation of Word Prediction Models. Cross-linguistic syntactic evaluation of word prediction models. 2020. Available online: https://arxiv.org/abs/2005.00187.
- Mueller, A.; Xia, Y.; Linzen, T. Causal Analysis of Syntactic Agreement Neurons in Multilingual Language Models. Causal analysis of syntactic agreement neurons in multilingual language models. 2022. Available online: https://arxiv.org/abs/2210.14328.
- Muller, B.; Elazar, Y.; Sagot, B.; Seddah, D. First Align, then Predict: Understanding the Cross-Lingual Ability of Multilingual BERT. P. Merlo, J. Tiedemann, R. Tsarfaty (), Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume Proceedings of the 16th conference of the european chapter of the association for computational linguistics: Main volume ( 2214–2231). OnlineAssociation for Computational Linguistics, 202104; Available online: https://aclanthology.org/2021.eacl-main.189/. [CrossRef]
- Nanda, N. Attribution Patching: Activation Patching at Industrial Scale. Neel Nanda’s Blog. 2023. Available online: https://www.neelnanda.io/mechanistic-interpretability/attribution-patching.
- Nanda, N.; Bloom, J. TransformerLens. Transformerlens. 2022. Available online: https://github.com/TransformerLensOrg/TransformerLens.
- Nie, E.; Schmid, H.; Schuetze, H. Mechanistic Understanding and Mitigation of Language Confusion in English-Centric Large Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2025 Findings of the association for computational linguistics: Emnlp 2025 ( 690–706). Suzhou, ChinaAssociation for Computational Linguistics.; Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng (), V., Eds.; 11 2025; Available online: https://aclanthology.org/2025.findings-emnlp.37/. [CrossRef]
- Nostalgebraist. Interpreting GPT: The logit lens. Blog Post. 2020. [Google Scholar] [CrossRef]
- Olah, C.; Cammarata, N.; Schubert, L.; Goh, G.; Petrov, M.; Carter, S. Distill53e00024–001; Zoom in: An introduction to circuits. 2020.
- Orhan, P.; Diego-Simón, P.; Chemla, E.; Lakretz, Y.; Boubenec, Y.; King, J.-R. Emergence of Phonemic, Syntactic, and Semantic Representations in Artificial Neural Networks. Emergence of phonemic, syntactic, and semantic representations in artificial neural networks. 2026. Available online: https://arxiv.org/abs/2601.18617.
- Ou, Y.; Yao, Y.; Zhang, N.; Jin, H.; Sun, J.; Deng, S.; Li, Z.; Chen, H. How Do LLMs Acquire New Knowledge? A Knowledge Circuits Perspective on Continual Pre-Training. How do llms acquire new knowledge? a knowledge circuits perspective on continual pre-training. 2025. Available online: https://arxiv.org/abs/2502.11196.
- Pal, K.; Sun, J.; Yuan, A.; Wallace, B.; Bau, D. Future Lens: Anticipating Subsequent Tokens from a Single Hidden State. In Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL) Proceedings of the 27th conference on computational natural language learning (conll), Available online. 2023; Association for Computational Linguistics; pp. 548–560. [Google Scholar] [CrossRef]
- Panigrahi, A.; Saunshi, N.; Zhao, H.; Arora, S. Task-Specific Skill Localization in Fine-tuned Language Models. Task-specific skill localization in fine-tuned language models. 2023. Available online: https://arxiv.org/abs/2302.06600.
- Pires, T.; Schlinger, E.; Garrette, D. How Multilingual is Multilingual BERT? A. Korhonen, D. Traum, L. Màrquez (), Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics Proceedings of the 57th annual meeting of the association for computational linguistics ( 4996–5001); Florence, ItalyAssociation for Computational Linguistics, 07 2019; Available online: https://aclanthology.org/P19-1493/. [CrossRef]
- Potertì, D.; Seveso, A.; Mercorio, F. Can Role Vectors Affect LLM Behaviour? C. Christodoulopoulos, T. Chakraborty, C. Rose, V. Peng (), Findings of the Association for Computational Linguistics: EMNLP 2025 Findings of the association for computational linguistics: Emnlp 2025 ( 17735–17747). Suzhou, ChinaAssociation for Computational Linguistics. 11 2025. Available online: https://aclanthology.org/2025.findings-emnlp.963/. [CrossRef]
- Rai, D.; Zhou, Y.; Feng, S.; Saparov, A.; Yao, Z. A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models. A practical review of mechanistic interpretability for transformer-based language models. 2025. Available online: https://arxiv.org/abs/2407.02646.
- Riemenschneider, F.; Frank, A. Cross-Lingual Generalization and Compression: From Language-Specific to Shared Neurons. Cross-lingual generalization and compression: From language-specific to shared neurons. 2025. Available online: https://arxiv.org/abs/2506.01629.
- Rimsky, N.; Gabrieli, N.; Schulz, J.; Tong, M.; Hubinger, E.; Turner, A. Steering Llama 2 via Contrastive Activation Addition. L.-W. Ku, A. Martins, V. Srikumar (), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. In Long Papers) Proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers) ( 15504–15522), 202408; Bangkok, ThailandAssociation for Computational Linguistics; Volume 1. Available online: https://aclanthology.org/2024.acl-long.828/. [CrossRef]
- Rufail, A.; Rathore, S.; Son, D.; Simon, A.; Dave, S.; Zhang, D.; Blondin, C.; O’Brien, S.; Zhu, K. Semantic Convergence: Investigating Shared Representations Across Scaled LLMs. ACL 2025 Student Research Workshop. Acl 2025 student research workshop. 2025. Available online: https://openreview.net/forum?id=oOxrKNo1lQ.
- Ruzzetti, E. S.; Zanzotto, F. M.; Caselli, T. Lexical Popularity: Quantifying the Impact of Pre-training for LLM Performance. V. Demberg, K. Inui, L. Marquez (), Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics. In Long Papers) Proceedings of the 19th conference of the European chapter of the Association for Computational Linguistics (volume 1: Long papers) ( 1209–1230); Rabat, MoroccoAssociation for Computational Linguistics, 03 2026; Volume 1, Available online: https://aclanthology.org/2026.eacl-long.55/. [CrossRef]
- Sajjad, H.; Durrani, N.; Dalvi, F. Neuron-level Interpretation of Deep NLP Models: A Survey. Trans. Assoc. Comput. Linguist. 2021, 101285–1303. Available online: https://arxiv.org/abs/2108.13138.
- Schulz, L. Y.; Mitropolsky, D.; Poggio, T. Unraveling Syntax: How Language Models Learn Context-Free Grammars. Unraveling syntax: How language models learn context-free grammars. 2025. Available online: https://arxiv.org/abs/2510.02524.
- Schut, L.; Gal, Y.; Farquhar, S. Do Multilingual LLMs Think In English? ICLR 2025 Workshop on Building Trust in Language Models and Applications. Iclr 2025 workshop on building trust in language models and applications, 2025; Available online: https://openreview.net/forum?id=I8BOtOPcOv.
- Shafran, O.; Ronen, S.; Fahn, O.; Ravfogel, S.; Geiger, A.; Geva, M. From Directions to Regions: Decomposing Activations in Language Models via Local Geometry. From directions to regions: Decomposing activations in language models via local geometry. 2026. Available online: https://arxiv.org/abs/2602.02464.
- Sharkey, L.; Chughtai, B.; Batson, J.; Lindsey, J.; Wu, J.; Bushnaq, L.; Goldowsky-Dill, N.; Heimersheim, S.; Ortega, A.; Bloom, J.; Biderman, S.; Garriga-Alonso, A.; Conmy, A.; Nanda, N.; Rumbelow, J.; Wattenberg, M.; Schoots, N.; Miller, J.; Michaud, E. J.; McGrath, T. Open Problems in Mechanistic Interpretability. Open problems in mechanistic interpretability. 2025. Available online: https://arxiv.org/abs/2501.16496.
- Shu, B.; Singh, A.; ElSherief, M. From Syntax to Emotion: A Mechanistic Analysis of Emotion Inference in LLMs. From syntax to emotion: A mechanistic analysis of emotion inference in llms. 2026. Available online: https://arxiv.org/abs/2604.25866.
- Shu, D.; Wu, X.; Zhao, H.; Rai, D.; Yao, Z.; Liu, N.; Du, M. A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models. A survey on sparse autoencoders: Interpreting the internal mechanisms of large language models. 2025. Available online: https://arxiv.org/abs/2503.05613.
- Singh, A. K.; Chan, S. C. Y.; Moskovitz, T.; Grant, E.; Saxe, A. M.; Hill, F. The Transient Nature of Emergent In-Context Learning in Transformers. The transient nature of emergent in-context learning in transformers. 2023. Available online: https://arxiv.org/abs/2311.08360.
- Someya, T.; Oseki, Y. JBLiMP: Japanese Benchmark of Linguistic Minimal Pairs. A. Vlachos I. Augenstein (), Findings of the Association for Computational Linguistics: EACL 2023 Findings of the association for computational linguistics: Eacl 2023 ( 1581–1594). Dubrovnik, CroatiaAssociation for Computational Linguistics. 05 2023. Available online: https://aclanthology.org/2023.findings-eacl.117/. [CrossRef]
- Son, D.; Rathore, S.; Rufail, A.; Simon, A.; Zhang, D.; Dave, S.; Blondin, C.; Zhu, K.; O’Brien, S. Semantic Convergence: Investigating Shared Representations Across Scaled LLMs. Semantic convergence: Investigating shared representations across scaled llms. 2025. Available online: https://arxiv.org/abs/2507.22918.
- Song, Y.; Krishna, K.; Bhatt, R.; Iyyer, M. SLING: Sino Linguistic Evaluation of Large Language Models. Y. Goldberg, Z. Kozareva, Y. Zhang (), Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing Proceedings of the 2022 conference on empirical methods in natural language processing ( 4606–4634); Abu Dhabi, United Arab EmiratesAssociation for Computational Linguistics, 12 2022; Available online: https://aclanthology.org/2022.emnlp-main.305/. [CrossRef]
- Stańczak, K.; Ponti, E.; Hennigen, L. T.; Cotterell, R.; Augenstein, I. Same Neurons, Different Languages: Probing Morphosyntax in Multilingual Pre-trained Models. Same neurons, different languages: Probing morphosyntax in multilingual pre-trained models. 2022. Available online: https://arxiv.org/abs/2205.02023.
- Stickland, A. C.; Lyzhov, A.; Pfau, J.; Mahdi, S.; Bowman, S. R. Steering Without Side Effects: Improving Post-Deployment Control of Language Models. Steering without side effects: Improving post-deployment control of language models. 2024. Available online: https://arxiv.org/abs/2406.15518.
- Suijkerbuijk, M.; Prins, Z.; Kloots, M. d. H.; Zuidema, W.; Frank, S. L. BLiMP-NL: A Corpus of Dutch Minimal Pairs and Acceptability Judgments for Language Model Evaluation. Comput. Linguist. Available online. 2025, 1–35. [Google Scholar] [CrossRef]
- Syed, A.; Rager, C.; Conmy, A. Attribution Patching Outperforms Automated Circuit Discovery; Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP Proceedings of the 7th blackboxnlp workshop: Analyzing and interpreting neural networks for nlp ( 407–416). Miami, Florida, Belinkov, Y., Kim, N., Jumelet, J., Mohebbi, H., Mueller, A., Chen (), H., Eds.; USAssociation for Computational Linguistics, 11 2024; Available online: https://aclanthology.org/2024.blackboxnlp-1.25/. [CrossRef]
- Taktasheva, E.; Bazhukov, M.; Koncha, K.; Fenogenova, A.; Artemova, E.; Mikhailov, V. RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing Proceedings of the 2024 conference on empirical methods in natural language processing ( 9268–9299); Miami, Florida, USAAssociation for Computational Linguistics, Al-Onaizan, Y., Bansal, M., Chen (), Y.-N., Eds.; 11 2024; Available online: https://aclanthology.org/2024.emnlp-main.522/. [CrossRef]
- Tan, S.; Wu, D.; Monz, C. Neuron Specialization: Leveraging Intrinsic Task Modularity for Multilingual Machine Translation. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing Proceedings of the 2024 conference on empirical methods in natural language processing ( 6506–6527); Miami, Florida, USAAssociation for Computational Linguistics, Al-Onaizan, Y., Bansal, M., Chen (), Y.-N., Eds.; 11 2024; Available online: https://aclanthology.org/2024.emnlp-main.374/. [CrossRef]
- Tang, T.; Luo, W.; Huang, H.; Zhang, D.; Wang, X.; Zhao, X.; Wei, F.; Wen, J.-R. Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models. L.-W. Ku, A. Martins, V. Srikumar (), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. In Long Papers) Proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers) ( 5701–5715); Bangkok, ThailandAssociation for Computational Linguistics, 08 2024; Volume 1, Available online: https://aclanthology.org/2024.acl-long.309/. [CrossRef]
- Tang, Y.; Saini, H.; Yao, Z.; Lin, Z.; Liao, Y.; Cui, J.; Wang, Y.; Du, M.; Liu, D. A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious Minima. A unified theory of sparse dictionary learning in mechanistic interpretability: Piecewise biconvexity and spurious minima. 2026. Available online: https://arxiv.org/abs/2512.05534.
- Tezuka, H.; Inoue, N. The Transfer Neurons Hypothesis: An Underlying Mechanism for Language Latent Space Transitions in Multilingual LLMs. The transfer neurons hypothesis: An underlying mechanism for language latent space transitions in multilingual llms. 2025. Available online: https://arxiv.org/abs/2509.17030.
- Todd, E.; Li, M.; Sharma, A. S.; Mueller, A.; Wallace, B. C.; Bau, D. Function Vectors in Large Language Models. The Twelfth International Conference on Learning Representations. The twelfth international conference on learning representations., 2024; Available online: https://openreview.net/forum?id=AwyxtyMwaG.
- Tufanov, I.; Hambardzumyan, K.; Ferrando, J.; Voita, E. LM Transparency Tool: Interactive Tool for Analyzing Transformer Language Models. Lm transparency tool: Interactive tool for analyzing transformer language models. 2024. Available online: https://arxiv.org/abs/2404.07004.
- Turner, A. M.; Thiergart, L.; Leech, G.; Udell, D.; Vazquez, J. J.; Mini, U.; MacDiarmid, M. Steering Language Models With Activation Engineering. Steering language models with activation engineering. 2024. Available online: https://arxiv.org/abs/2308.10248.
- Waldis, A.; Perlitz, Y.; Choshen, L.; Hou, Y.; Gurevych, I. Holmes: A Benchmark to Assess the Linguistic Competence of Language Models. Transactions of the Association for Computational Linguistics121616–1647. 2024. Available online: https://aclanthology.org/2024.tacl-1.88/. [CrossRef]
- Wang, K.; Variengien, A.; Conmy, A.; Shlegeris, B.; Steinhardt, J. Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small. Interpretability in the wild: a circuit for indirect object identification in gpt-2 small. 2022. Available online: https://arxiv.org/abs/2211.00593.
- Wang, M.; Adel, H.; Lange, L.; Liu, Y.; Nie, E.; Strötgen, J.; Schuetze, H. Lost in Multilinguality: Dissecting Cross-lingual Factual Inconsistency in Transformer Language Models. W. Che, J. Nabende, E. Shutova, M. T. Pilehvar (), Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics. In Long Papers) Proceedings of the 63rd annual meeting of the association for computational linguistics (volume 1: Long papers) ( 5075–5094); Vienna, AustriaAssociation for Computational Linguistics, 07 2025; Volume 1, Available online: https://aclanthology.org/2025.acl-long.253/. [CrossRef]
- Wang, M.; Xu, Z.; Mao, S.; Deng, S.; Tu, Z.; Chen, H.; Zhang, N. Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms. Beyond prompt engineering: Robust behavior control in llms via steering target atoms. 2025. Available online: https://arxiv.org/abs/2505.20322.
- Warstadt, A.; Parrish, A.; Liu, H.; Mohananey, A.; Peng, W.; Wang, S.-F.; Bowman, S. R. BLiMP: The Benchmark of Linguistic Minimal Pairs for English. Transactions of the Association for Computational Linguistics. 2020, pp. 8377–392. Available online: https://aclanthology.org/2020.tacl-1.25/. [CrossRef]
- Wei, J.; Garrette, D.; Linzen, T.; Pavlick, E. Frequency Effects on Syntactic Rule Learning in Transformers; Online and Punta Cana, Dominican RepublicAssociation for Computational Linguistics, Moens, M.-F., Huang, X., Specia, L., Yih (), S. W.-t., Eds.; Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing Proceedings of the 2021 conference on empirical methods in natural language processing ( 932–948), 11 2021; Available online: https://aclanthology.org/2021.emnlp-main.72/. [CrossRef]
- Wendler, C.; Veselovsky, V.; Monea, G.; West, R. Do Llamas Work in English? On the Latent Language of Multilingual Transformers. L.-W. Ku, A. Martins, V. Srikumar (), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. In Long Papers) Proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers) ( 15366–15394); Bangkok, ThailandAssociation for Computational Linguistics, 08 2024; Volume 1, Available online: https://aclanthology.org/2024.acl-long.820/. [CrossRef]
- Wilson, M.; Petty, J.; Frank, R. How Abstract Is Linguistic Generalization in Large Language Models? Experiments with Argument Structure. Trans. Assoc. Comput. Linguist. 2023, 111377–1395. Available online: https://doi.org/10.1162/tacl_a_00608. [CrossRef]
- Wu, S.; Dredze, M. Beto, Bentz, Becas: The Surprising Cross-Lingual Effectiveness of BERT. Beto, bentz, becas: The surprising cross-lingual effectiveness of bert. 2019. Available online: https://arxiv.org/abs/1904.09077.
- Wu, Z.; Arora, A.; Geiger, A.; Wang, Z.; Huang, J.; Jurafsky, D.; Manning, C. D.; Potts, C. AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders. Axbench: Steering llms? even simple baselines outperform sparse autoencoders. 2025. Available online: https://arxiv.org/abs/2501.17148.
- Wu, Z.; Arora, A.; Wang, Z.; Geiger, A.; Jurafsky, D.; Manning, C. D.; Potts, C. ReFT: Representation Finetuning for Language Models. The Thirty-eighth Annual Conference on Neural Information Processing Systems. The thirty-eighth annual conference on neural information processing systems., 2024; Available online: https://openreview.net/forum?id=fykjplMc0V.
- Wu, Z.; Yu, X. V.; Yogatama, D.; Lu, J.; Kim, Y. The Semantic Hub Hypothesis: Language Models Share Semantic Representations Across Languages and Modalities. The semantic hub hypothesis: Language models share semantic representations across languages and modalities. 2025. Available online: https://arxiv.org/abs/2411.04986.
- Xiang, B.; Yang, C.; Li, Y.; Warstadt, A.; Kann, K. CLiMP: A Benchmark for Chinese Language Model Evaluation. P. Merlo, J. Tiedemann, R. Tsarfaty (), Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume Proceedings of the 16th conference of the european chapter of the association for computational linguistics: Main volume ( 2784–2790). OnlineAssociation for Computational Linguistics, 202104; Available online: https://aclanthology.org/2021.eacl-main.242/. [CrossRef]
- Xie, W.; Feng, Y.; Gu, S.; Yu, D. Importance-based Neuron Allocation for Multilingual Neural Machine Translation. C. Zong, F. Xia, W. Li, R. Navigli (), Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing. In Long Papers) Proceedings of the 59th annual meeting of the association for computational linguistics and the 11th international joint conference on natural language processing (volume 1: Long papers) ( 5725–5737); OnlineAssociation for Computational Linguistics, 08 2021; Volume 1, Available online: https://aclanthology.org/2021.acl-long.445/. [CrossRef]
- Yu, Z.; Ananiadou, S. Neuron-Level Knowledge Attribution in Large Language Models. Y. Al-Onaizan, M. Bansal, Y.-N. Chen (), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing Proceedings of the 2024 conference on empirical methods in natural language processing ( 3267–3280). Miami, Florida, 202411; USAAssociation for Computational Linguistics. Available online: https://aclanthology.org/2024.emnlp-main.191/. [CrossRef]
- Zeng, H.; Han, S.; Chen, L.; Yu, K. Converging to a Lingua Franca: Evolution of Linguistic Regions and Semantics Alignment in Multilingual Large Language Models; Abu Dhabi, UAEAssociation for Computational Linguistics, Rambow, O., Wanner, L., Apidianaki, M., Al-Khalifa, H., Eugenio, B. D., Schockaert (), S., Eds.; Proceedings of the 31st International Conference on Computational Linguistics Proceedings of the 31st international conference on computational linguistics ( 10602–10617), 01 2025; Available online: https://aclanthology.org/2025.coling-main.707/.
- Zhang, C.; Lu, J.; Tran, V. Q.; Schuster, T.; Metzler, D.; Lin, J. Tomato, Tomahto, Tomate: Do Multilingual Language Models Understand Based on Subword-Level Semantic Concepts? L. Chiruzzo, A. Ritter, L. Wang (), Findings of the Association for Computational Linguistics: NAACL 2025 Findings of the association for computational linguistics: Naacl 2025 ( 1821–1837). Albuquerque, New MexicoAssociation for Computational Linguistics. 04 2025. Available online: https://aclanthology.org/2025.findings-naacl.98/. [CrossRef]
- Zhang, F.; Nanda, N. Towards Best Practices of Activation Patching in Language Models: Metrics and Methods. The Twelfth International Conference on Learning Representations. The twelfth international conference on learning representations., 2024; Available online: https://openreview.net/forum?id=Hf17y6u9BC.
- Zhang, H.; Shang, C.; Wang, S.; Zhang, D.; Yu, Y.; Yao, F.; Sun, R.; Yang, Y.; Wei, F. ShifCon: Enhancing Non-Dominant Language Capabilities with a Shift-based Multilingual Contrastive Framework; Che, W., Nabende, J., Shutova, E., Pilehvar (), M. T., Eds.; Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 07 2025; Volume 1, Available online: https://aclanthology.org/2025.acl-long.239/. [CrossRef]
- Zhang, H.; Zhang, Z.; Wang, M.; Su, Z.; Wang, Y.; Wang, Q.; Yuan, S.; Nie, E.; Duan, X.; Han, F.; Xue, Q.; Yu, Z.; Shang, C.; Liang, X.; Xiong, J.; Shen, H.; Tao, C.; Liu, Z.; Jin, S.; Wong, N. Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models. Locate, steer, and improve: A practical survey of actionable mechanistic interpretability in large language models. 2026. Available online: https://arxiv.org/abs/2601.14004.
- Zhang, R.; Yu, Q.; Zang, M.; Eickhoff, C.; Pavlick, E. The Same but Different: Structural Similarities and Differences in Multilingual Language Modeling. The Thirteenth International Conference on Learning Representations. The thirteenth international conference on learning representations., 2025; Available online: https://openreview.net/forum?id=NCrFA7dq8T.
- Zhang, X.; Liang, Y.; Meng, F.; Zhang, S.; Chen, Y.; Xu, J.; Zhou, J. Multilingual Knowledge Editing with Language-Agnostic Factual Neurons; Abu Dhabi, UAEAssociation for Computational Linguistics, Rambow, O., Wanner, L., Apidianaki, M., Al-Khalifa, H., Eugenio, B. D., Schockaert (), S., Eds.; Proceedings of the 31st International Conference on Computational Linguistics Proceedings of the 31st international conference on computational linguistics ( 5775–5788), 01 2025; Available online: https://aclanthology.org/2025.coling-main.385/.
- Zhang, Z.; Zhao, J.; Zhang, Q.; Gui, T.; Huang, X. Unveiling Linguistic Regions in Large Language Models. L.-W. Ku, A. Martins, V. Srikumar (), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. In Long Papers) Proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers) ( 6228–6247), 202408; Bangkok, ThailandAssociation for Computational Linguistics; Volume 1. Available online: https://aclanthology.org/2024.acl-long.338/. [CrossRef]
- Zhao, N.; Duan, X.; Cai, Z. G. The Missing Half of Language Learning in Current Developmental Language Models: Exogenous and Endogenous Linguistic Input. Open Mind91543-1549 Available online. 2025. [Google Scholar] [CrossRef] [PubMed]
- Zhao, Y.; Zhang, W.; Chen, G.; Kawaguchi, K.; Bing, L. How do Large Language Models Handle Multilingualism? The Thirty-eighth Annual Conference on Neural Information Processing Systems. The thirty-eighth annual conference on neural information processing systems, 2024; Available online: https://openreview.net/forum?id=ctXYOoAgRy.
- Zhong, C.; Cheng, F.; Liu, Q.; Jiang, J.; Wan, Z.; Chu, C.; Murawaki, Y.; Kurohashi, S. Beyond English-Centric LLMs: What Language Do Multilingual Language Models Think in? Beyond english-centric llms: What language do multilingual language models think in? 2024. Available online: https://arxiv.org/abs/2408.10811.
- Zhong, C.; Cheng, F.; Liu, Q.; Murawaki, Y.; Chu, C.; Kurohashi, S. Language Lives in Sparse Dimensions: Toward Interpretable and Efficient Multilingual Control for Large Language Models. Language lives in sparse dimensions: Toward interpretable and efficient multilingual control for large language models. 2025. Available online: https://arxiv.org/abs/2510.07213.
- Zhou, E.; Salhan, S. Extended Abstract for “Linguistic Universals”: Emergent Shared Features in Independent Monolingual Language Models via Sparse Autoencoders. D. I. Adelani et al. (), Proceedings of the 5th Workshop on Multilingual Representation Learning (MRL 2025) Proceedings of the 5th workshop on multilingual representation learning (mrl 2025) ( 128–130). Suzhuo, ChinaAssociation for Computational Linguistics, 11 2025. Available online: https://aclanthology.org/2025.mrl-main.9/. [CrossRef]
- Zhou, X.; Chen, D.; Cahyawijaya, S.; Duan, X.; Cai, Z. Linguistic Minimal Pairs Elicit Linguistic Similarity in Large Language Models. In Proceedings of the 31st International Conference on Computational Linguistics Proceedings of the 31st international conference on computational linguistics ( 6866–6888); Abu Dhabi, UAEAssociation for Computational Linguistics, Rambow, O., Wanner, L., Apidianaki, M., Al-Khalifa, H., Eugenio, B. D., Schockaert (), S., Eds.; 01 2025; Available online: https://aclanthology.org/2025.coling-main.459/.
- Zhu, S.; Pan, L.; Li, B.; Xiong, D. LANDeRMT: Dectecting and Routing Language-Aware Neurons for Selectively Finetuning LLMs to Machine Translation. L.-W. Ku, A. Martins, V. Srikumar (), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. In Long Papers) Proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers) ( 12135–12148); Bangkok, ThailandAssociation for Computational Linguistics, 08 2024; Volume 1, Available online: https://aclanthology.org/2024.acl-long.656/. [CrossRef]
- Üveges, I.; Ring, O. Evaluating the Impact of Synthetic Data on Emotion Classification: A Linguistic and Structural Analysis. Information164. 2025. Available online: https://www.mdpi.com/2078-2489/16/4/330. [CrossRef]
| Paper | Focus |
|---|---|
| [132] | MI methods for Transformer. |
| [141] | Roadmap and future directions. |
| [143] | Feature disentanglement. |
| [116] | Causal Mediation Analysis. |
| [85] | Training dynamics. |
| [48] | Intrinsic interpretability. |
| [53] | Syntactic Knowledge. |
| [179] | Locate, Steer, and Improve. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).