Preprint
Review

This version is not peer-reviewed.

Mechanistic Interpretability for Understanding Language Abilities in Large Language Models: A Survey

Submitted:

18 July 2026

Posted:

20 July 2026

You are already at the latest version

Abstract
Large Language Models (LLMs) exhibit strong performance across language tasks, raising the question of whether this success reflects linguistic abstraction or reliance on surface-level statistical patterns. This survey reviews recent advances in mechanistic interpretability (MI) for analyzing candidate internal mechanisms underlying linguistic competence in LLMs. We trace the shift from model-agnostic analyses to transformer-specific methods and summarize core approaches including probing, neuron- and circuit-level analysis, and causal interventions. Synthesizing empirical findings, we argue that current evidence supports an intermediate view: LLMs appear to use reusable mechanisms for some grammatical behaviors, while these mechanisms remain frequency-conditioned, distributed, and only partially understood. We further review interpretability studies of multilingual models, highlighting evidence for both language-specific and shared representations. Finally, we outline open challenges in scalability, benchmarking, and ethics, and discuss directions toward a more cautious mechanistic account of language in LLMs. The curated list of papers for this work is available at https://github.com/xufengduan/Awesome-Language-MI-Survey.
Keywords: 
;  ;  

1. Introduction

Large Language Models (LLMs) achieve strong performance across a wide range of language tasks, including translation and text generation. This success raises a fundamental question about the nature of their underlying linguistic knowledge: do LLMs internalize systematic grammatical and semantic abstractions akin to human-like competence, or do they primarily exploit surface-level statistical regularities in large-scale data to approximate linguistic performance? While their behavior suggests the emergence of nontrivial linguistic structure [19,100], it remains plausible that such results arise from sophisticated pattern matching without the underlying symbolic or rule-based representations traditionally associated with human language [15,28].
Mechanistic interpretability (MI) aims to clarify this distinction by decomposing a model’s blackbox predictions into the internal computational structures, such as specialized attention heads or individual neurons for linguistic competence, that produce them. Recent studies have identified structured “circuits” or feature subspaces that appear to correlate with specific linguistic phenomena, offering partial analogies to human cognitive processes [98,109]. However, a critical challenge in this field is the conceptual gap between mechanistic identification (e.g., locating a circuit) and linguistic attribution (e.g., confirming that the circuit implements a generalizable rule). By scrutinizing these internal states, researchers seek to determine whether a model truly implements abstract generalizations or merely leverages high-dimensional heuristics that are sufficient for its training distribution.

Scope and Audience

In this survey, linguistic competence refers primarily to internal knowledge of morphosyntax (e.g., agreement, binding, word order) and formal semantics (e.g., argument structure and thematic roles). We discuss discourse, pragmatics, reasoning, and neural alignment only where mechanistic evidence directly connects them to linguistic representations. This focus is intentionally narrower than the full space of modern agentic, multimodal, or tool-using systems: our goal is to ask what MI reveals about structured language knowledge in Transformer-based LLMs. The survey is written for linguists seeking a map of what MI can and cannot establish, MI researchers seeking linguistic grounding, and NLP researchers interested in why models succeed or fail on linguistic evaluations.
Figure 1. The overview of this survey.
Figure 1. The overview of this survey.
Preprints 223887 g001
Prior surveys primarily categorize interpretability methods (Table 1), leaving a critical gap in synthesizing what these techniques reveal about linguistic competence. We argue that mechanistic evidence supports a layered account of linguistic implementation in LLMs: feature-level variables encode local grammatical and semantic distinctions; circuit-level mechanisms reuse computational subroutines across tasks; and multilingual models employ routing-level organization to link surface forms to shared semantic spaces. This thesis suggests that LLM behavior reflects structural organization rather than shallow memorization, without implying the existence of classical symbolic grammars.
This account occupies a middle position between two extremes. While mechanistic findings rule out a pure n-gram or rote-memorization explanation by identifying explicit pathways for linguistic competence (e.g., syntax), they do not justify claims of human-like symbolic competence. The discovered mechanisms remain distributed, frequency-conditioned, and model-specific. Therefore, we treat mechanistic evidence as a window into how specific models implement linguistic behavior, rather than proof of human-like cognitive architecture.
This survey presents a comprehensive overview of MI methods for studying linguistic competence in LLMs. Our contributions are as follows: (1) We review the evolution of interpretability techniques, from early probing approaches to recent advances, and clarify the linguistic claims each method can support. (2) We synthesize empirical evidence on how linguistic competence dynamically emerges in LLMs, uncover the underlying structural mechanisms, and evaluate its neurocognitive alignment with human language processing. (3) We extend the discussion to multilingual models, distinguishing between language-specific mechanisms and shared, language-agnostic representational spaces that support cross-lingual transfer. (4) We identify open challenges in scaling interpretability to increasingly large models, limitations of existing linguistic benchmarks, and tensions between interpretability and performance, and outline directions for future research.

2. Interpretability Methodology

In this paper, we review key methods for analyzing how LLMs represent linguistic knowledge (see Table 3 for a summary and a running example for all methods in Appendix 7.1.4).

2.1. Vocabulary Projection

Vocabulary Projection (or the Logit Lens) [124] interprets intermediate representations by projecting hidden states h directly onto the vocabulary via the unembedding matrix W U , revealing real-time predictions [31]:
Logits = W U · LayerNorm ( h )
This method tracks layer-by-layer prediction evolution [50]. Enhancements include the Tuned Lens [14], which calibrates intermediate representations, and the Future Lens [128], which predicts subsequent states.

Linguistic Leverage

Vocabulary projection helps localize when a model favors specific tokens (e.g., number-marked verbs). However, it does not confirm the causal necessity of these representations.

2.2. Activation Patching

To establish causality, Activation Patching systematically replaces a component’s activation in a corrupted input ( x corrupt ) with the corresponding activation from a clean input ( x clean ) [110,161,177]. If the intervention restores the model’s performance on a target task, that component is considered causally decisive [45,157]. Attribution Patching [121] further refines this by using gradients to approximate these interventions linearly.
The importance of a component i is quantified by its indirect effect (IE), measuring the change in target token probability p r after patching:
p r ( f ( x corrupt ; a i a i ( x clean ) ) ) p r ( f ( x corrupt ) )
While powerful, these methods require caution [68]. [105] warn of an “interpretability illusion,” where interventions might activate dormant pathways rather than revealing the natural mechanism.

Linguistic Leverage

Patching provides direct causal evidence for linguistic tasks. Its validity depends on demonstrating generalizability across diverse lexical items, templates, and distractors to avoid identifying prompt-specific pathways.

2.3. Sparse Autoencoders

A fundamental barrier to interpreting individual units (e.g., neurons) is superposition, where models compress more features than they have dimensions, resulting in polysemantic neurons [39]. For example, a single neuron might fire for both “plural nouns” and “Python code comments,” making it impossible to assign it a unique linguistic label. Sparse Autoencoders (SAEs) address this by decomposing dense, polysemantic activations into an overcomplete, sparse basis where each direction represents a single concept [16,67,76,143,155]. SAEs reconstruct an activation x as a linear combination of features:
x x ^ = b d e c + i f i ( x ) d i
where f i ( x ) is the sparse feature activation (via L 1 regularization) and d i is the decoder direction.
Researchers have used this to identify monosemantic units for specific linguistic phenomena (e.g., "German plural endings") [79]. However, [82] caution that SAEs require rigorous evaluation, as they do not consistently outperform baselines in downstream tasks.

Linguistic Leverage

SAEs improve inspectability by disentangling dense activations. The primary evidential challenge is demonstrating that these extracted features are causally active: does ablating, steering, or tracing a specific feature reliably alter the corresponding linguistic behavior in a model?

2.4. Circuit Discovery

While SAEs disentangle representations into interpretable features, these units can also be composed into higher-level computational mechanisms. Building on activation patching, Circuit Discovery [115,177] aims to identify the complete subgraph of components, such as attention heads and neurons, that are responsible for a specific task. While early work relied on manual hypothesis testing [125], recent approaches employ automated search algorithms. Automated Circuit Discovery [32] automates the process of pruning the computational graph. It iteratively ablates edges in the network, keeping only those whose removal would significantly degrade performance on a specific task. To address the difficulty of analyzing dense MLP layers in circuits, [38] introduce Transcoders, which approximate MLPs with sparse, interpretable components. More recently, researchers have shifted focus from “component circuits” (where nodes are heads/MLPs) to “SAE feature circuits”. [107] and [86] utilize sparse autoencoders to define circuits over fine-grained semantic features rather than coarse model components. This allows for the discovery of circuits that are more interpretable, as edges represent causal dependencies between specific concepts (e.g., a “plural subject” feature causing a “plural verb” prediction) rather than opaque vectors.

Linguistic Leverage

Circuit discovery is most informative when it identifies a reusable algorithmic pathway rather than a single correlated unit. For language, the central test is whether the circuit implements an abstraction across lexical and structural variation, or whether it only solves a narrow benchmark template.

2.5. Activation Steering

Methodology has expanded from observation to control, enabling direct intervention in a model’s linguistic behavior. Activation Steering involves injecting vectors [83,131,157,171] into the residual stream to elicit specific behaviors. [159] introduced Activation Addition (ActAdd), utilizing mean activation differences between positive and negative examples. To improve robustness, Contrastive Activation Addition (CAA) [134] refines this by averaging differences across contrastive pairs (e.g., “Helpful” vs. “Unhelpful”). Building on this, Steering Target Atoms (STA) [163] leverages SAEs to decompose steering vectors into sparse, interpretable features, allowing for more precise control with fewer side effects.

2.5.0.6. Linguistic Leverage

Steering links explanation to intervention: if linguistic properties are predictably shifted, the underlying representation is confirmed as functional rather than merely descriptive. However, these results require extra validation to ensure that steering does not inadvertently degrade unrelated fluency, style, or factual accuracy.

3. MI on Linguistic Competence

This section reviews MI evidence on linguistic competence, focusing on how grammatical knowledge emerges during training, how it is implemented in model internals, and how far these representations align with human language processing. Together, these perspectives shift the analysis from whether models perform well on linguistic tasks to how such competence is formed, organized, and constrained.

3.1. Dynamic View: How It Emerges

Behavioral benchmarks often report strong performance on core grammatical phenomena [72]. However, analyses of internal model states suggest that this competence is not a monolithic body of knowledge acquired instantaneously, but rather the outcome of a complex, frequency-dependent developmental process.
Critical Periods and Structural Evolution Evidence suggests that the timing of language acquisition may be structurally important. [36] identify “critical periods” in training, specific windows where rapid internal syntactic specialization occurs. [127] and [10] trace this evolution using circuit analysis and crosscoders, observing that linguistic features are not merely added but actively consolidated during pre-training. This temporal evolution is complemented by cross-scale findings from [103], who note that syntactic acquisition also saturates at specific model scales.
From Memorization to Abstraction The acquisition of language competence during training appears to follow a hierarchical trajectory shaped by data distribution. In early stages, models preferentially learn high-frequency tokens and local n-grams, while rarer and long-distance dependencies emerge only through gradual refinement [21,165]. This suggests that linguistic abstraction is not purely symbolic but remains anchored to the statistical density of the training data. At the mechanistic level, this temporal progression has been linked to shifts in internal circuitry, including reported transitions from transient in-context strategies to more stable in-weights representations [144]. For a comprehensive review, see [85].

What We Learn

A recurring pattern across these studies is developmental: formal linguistic competence appears to emerge through staged, frequency-dependent trajectories rather than all at once. Mechanistic evidence points to increasingly stable grammatical variables during training, but claims about abstract rules are best supported when they survive lexical, frequency, and structural controls rather than only benchmark-level accuracy.

3.2. Mechanistic View: Specialized vs. Shared Circuits

Following the gradual emergence of linguistic capabilities during training, it remains unclear whether neural networks rely on isolated modules or reconfigure overlapping subsystems to support diverse tasks. While numerous studies have identified language-specific neurons and features, recent circuit-level analyses increasingly challenge the assumption of strict functional independence [89]. [112] report that circuits identified for Indirect Object Identification (IOI) are reused in tasks sharing similar algorithmic structure, with substantial overlap in attention heads. These findings suggest that models rely on reusable circuits rather than strictly task-specific mechanisms [e.g., [32]. Beyond circuit overlaps, researchers have identified functional neurons whose targeted perturbation can degrade or redirect performance. The works [58,77,154,170] report that ablating these neurons yields large drops in downstream accuracy.
However, the degree of localization varies: simple local agreement checks may be handled by a compact set of heads [61,125], whereas complex reasoning requires interactions among numerous polysemantic units [107].

What We Learn

Current circuit evidence favors reusable subroutines over fully isolated linguistic modules. Duplicate-token detection, agreement routing, and mover/inhibition patterns suggest that models can reuse low-level mechanisms while attaching task-specific execution heads.

3.3. Cognitive View: Partial Alignment with Human Language Processing

Recent works [3,70,91] suggest that the relationship between LLMs and human language processing can be characterized as partial alignment: systematic correspondences emerge at specific representational levels, while substantial divergences persist elsewhere.
Alignment is most relevant for this survey at the level of language-selective structure. [2] identify language-selective units across a wide range of LLMs, which exhibit stronger alignment with neural activity in human language regions than randomly sampled areas [91]. At a finer grain, [37] use theory-driven psycholinguistic paradigms to probe neuron-level mechanisms, reporting that when a model exhibits human-like behavior in tasks such as sound–gender association, a small subset of neurons can be causally identified.
At the representational level, hidden states can predict neural responses during naturalistic language comprehension [47,71]. Furthermore, this alignment exhibits a layered correspondence between model depth and the temporal hierarchy of neural responses [52]. However, this alignment is selective. Tracking models across training, [3] report that brain alignment correlates more strongly with formal linguistic competence than with functional abilities such as reasoning.

What We Learn

Cognitive alignment findings are suggestive but bounded. The most relevant evidence concerns formal linguistic competence and language-selective representations; broader pragmatics and human-like cognition remain less directly explained by current MI.

4. Multilingualism

This section reviews MI evidence on multilingualism, focusing on how multilingual LLMs represent language-specific information, enable cross-lingual transfer, and align semantic content across languages. Current evidence points to a hybrid organization of multilingual representations. Language-specific components appear to govern language identity and surface-form processing, whereas shared representations are more closely associated with transfer and semantic alignment.

4.1. Language-Selective Components

A growing body of work reports language-selective components in multilingual decoder-style LLMs at multiple granularities, from coarse region-level localization to fine-grained, steerable features.
At the neuron level, language competence exhibits partial spatial localization. [88] find these language-specific neurons are concentrated in early and late layers, and that perturbing fewer than 1% of neurons can substantially alter output language probabilities. [154] further study this localized sensitivity by introducing Language Activation Probability Entropy to pinpoint these neurons, reporting that deactivating a small subset selectively impairs comprehension or generation in a target language. Beyond individual neurons, region-level analyses show similar patterns. [182] identify distributed monolingual regions associated with the syntactic and lexical properties of particular languages.
Beyond static neuron parameters, dynamic routing mechanisms further isolate language processing. On the attention side, [102] identify both language-specific and language-general attention heads. Circuit-level analyses further support this view: Anthropic’s circuit tracing work on Claude 3.5 Haiku shows that multilingual behavior recruits language-dominant clusters alongside abstract, shared pathways, with the proportion of shared circuitry increasing with model scale [5,98].
Recent SAE-based studies provide additional perspective and fine-grained steering. [33] and [6] identify monosemantic language-specific features whose ablation selectively degrades individual languages. Building on this, [186] propose that language control relies on a sparse, cross-layer set of dimensions, and [30] report zero-shot language control through sparse feature steering.
Overall, these findings suggest that multilingual LLMs, despite extensive parameter sharing, preserve some specialized internal structures that support language-specific processing.

What We Learn

Language-selective components offer a useful entry point to study how multilingual LLMs encode language identity, typological variation, and language-specific form. Current evidence suggests that parameter sharing does not eliminate specialized structure; rather, multilingual models combine localized language-sensitive units with broader shared pathways.

4.2. Cross-Lingual Transfer

We define cross-lingual transfer as a model’s ability to reuse knowledge or task competence learned in one language when processing inputs in another. Mechanistically, recent evidence supports a two-component view: transfer is primarily enabled by language-agnostic computation in intermediate layers, while language-specific units help route information between language-dependent surface forms and that shared computation.
The Multilingual Workflow and Shared Latent Spaces. [184] formalize this as a multilingual workflow, where non-English inputs are progressively mapped into an English-like internal workspace for reasoning, and later projected back to the query language during generation. This remains a model-specific hypothesis rather than a settled account of multilingual computation. The mechanistic explanations for this transfer are becoming more precise: [156] propose the Transfer Neuron Hypothesis, identifying specific MLP neurons associated with transformations between language-specific and shared latent spaces. As evidence for partial functional independence, [184] report that targeted tuning of those neurons can selectively improve a language without broadly affecting others. To map this flow to the entire network, [64] utilize cross-layer transcoders, finding that while internal representations are largely shared, language-specific decoding emerges in the final layers driven by high-frequency features. Furthermore, [17] reports shared feature directions for grammatical concepts even in models trained primarily on English.
Developmental Dynamics of Transfer. [22] posit that during training, an LLM initially allocates its resources to a “primary” or dominant language, only later partitioning its internal representation to accommodate new languages. Their empirical analysis tracks language transferring neurons that mark the point at which the model begins establishing dedicated circuits for additional languages. [133] observe a compression process during training, where models initially form language-specific representations that gradually converge into cross-lingual abstractions.
While the localization of cross-lingual mechanisms is becoming precise, the causal utility of these components remains contested. [114] systematically test neuron-specific interventions and find that manipulating language-specific neurons does not reliably improve downstream cross-lingual performance on benchmarks such as XNLI and XQuAD for low-resource languages.

What We Learn

Cross-lingual transfer appears to depend on an interaction between language-specific routing and shared task-solving computation. Localizing language-specific neurons is therefore not sufficient: the open question is how these components coordinate with shared intermediate representations, and why targeted interventions sometimes change language identity without improving downstream transfer.

4.3. Mechanisms of Semantic Alignment

While multilingual LLMs develop components that exhibit strong language selectivity, a growing body of work shows that these models converge on a shared latent space for semantic content. [185] and [166] find that representations for inputs in different languages often pass through English-like intermediate spaces. Generalizing this observation, [175] propose a hidden-state lingua franca: a shared semantic manifold to which inputs are mapped irrespective of language, with larger models and increased training further strengthening this language-agnostic alignment.
This shared semantic structure is supported by partially overlapping subspaces and token representations. [176] report that multilingual embeddings cluster near-synonymous tokens across languages in overlapping regions, while [101] find that domain knowledge is integrated into these spaces in a way that balances shared semantic representations with language- or domain-specific structure. Together, these results point to a dual organization where localized mechanisms coexist with a common semantic hub.
The shared space itself may be English-centric in some models. Rather than showing that multilingual LLMs literally “think in English,” [139] provide evidence that intermediate states can become English-aligned before generation in the target language. Related evidence extends beyond translation: [17] report that grammatical concepts such as number and gender can be encoded in shared feature directions across typologically diverse languages, and [146] find semantic alignment across models of varying scales.
Overall, these findings suggest that multilingual LLMs implement a hybrid representational strategy: language-specific circuits capture idiosyncratic form, while higher-level semantic representations converge into a shared latent space. Understanding this balance between specialization and convergence has practical implications for model design, enabling targeted control of language-specific behavior while preserving cross-lingual generalization.

What We Learn

The multilingual evidence extends the layered account: local language-specific features and circuits handle surface-form control, while higher-level representations increasingly converge into a shared semantic hub. The English-centric nature of this hub is plausible but not settled; it should be tested across more low-resource, non-Indo-European, and typologically distant languages before being treated as a universal property of multilingual LLMs.

5. Challenges and Open Questions

Despite progress in MI, model scaling complicates the isolation and manipulation of linguistic mechanisms. Simultaneously, existing evaluation paradigms face increasing scrutiny regarding their scope and inclusivity.
MI is also most useful when triangulated with complementary approaches. Behavioral evaluation constrains what a model can do; probing and representational geometry characterize what information is encoded; cognitive modeling tests whether model behavior and representations align with human processing; and MI contributes causal internal evidence about how a behavior is produced. A full account of linguistic competence requires these sources of evidence to be interpreted together.

5.1. Scaling vs. Interpretability Trade-Off

While scaling enhances performance, it challenges fine-grained interpretability. [96] observe that circuit-level analysis becomes increasingly intractable at scale, suggesting emergent behaviors may resist decomposition. Furthermore, [109] argue that extensive polysemanticity in complex models precludes unique interpretations. Conversely, [32] provide evidence that automated circuit discovery can scale beyond fully manual analysis, indicating that methodological advances may offset some complexity. However, the tractability of such approaches, and techniques like SAE [79,107], for massive-scale models remains unproven.

5.2. Evaluation Gaps

Although benchmarks like CausalGym [7], Holmes [160] and BHASA [93] have expanded linguistic coverage, significant gaps persist. Minimal-pair and targeted syntactic tests [7,73] effectively probe local grammar but often neglect higher-level reasoning, discourse, and pragmatics. Moreover, the focus on high-resource languages limits the detection of failures in underrepresented contexts, risking interpretability methods that overfit to narrow linguistic variations. Developing robust, culturally sensitive frameworks for low-resource languages remains a critical challenge [20,175].

5.3. Causality and Model Editing

A core goal of MI is causal intervention. Techniques such as attribution patching [43] and neuron or feature steering [8,98] test whether specific components can be causally linked to linguistic behavior. However, these interventions often lack stability: localized changes can propagate through the network, producing unintended side effects [149].
Superposition further complicates causal control. As shown by [39], polysemantic neurons encode multiple features simultaneously, making it difficult to edit one function without disrupting others. Developing reliable methods to identify modular, independently editable circuits and organize them into a safe, reusable “circuit library” remains an open challenge.

5.4. Cognitive Plausibility and Human-Like Generalization

Alignment between LLMs and human language processing remains partial. [3] observe that while specific units mirror human language networks, divergence increases as models integrate broader reasoning capabilities. Training dynamics also echo human acquisition via staged morphological and syntactic learning [29,36,41], yet generalization remains limited: LLMs transfer argument-structure patterns mainly when the relevant contextual relation was seen in pre-training [167]. Distinguishing shared cognitive implementation from behavioral convergence is critical for interpretability [183]. Despite calls for analyzing critical training transitions [7], the fundamental cognitive plausibility of LLM competence remains an open question.

5.5. Ethical and Societal Dimensions

MI also raises ethical considerations, particularly regarding bias, fairness, and language equity. Neuron-level analyses can expose how models encode stereotypes or toxic associations, enabling targeted mitigation strategies [78,94]. From a global perspective, interpretability highlights persistent linguistic inequities: multilingual models remain strongly English-centric [54], with uneven performance across languages [20]. Developing interpretable and equitable models for underrepresented languages, remains an underexplored but critical challenge.

5.6. Implications of Synthetic Data and Model Collapse

As LLMs are increasingly trained on machine-generated data, interpretability faces new uncertainties. [55] show that iterative training on synthetic text can lead to lexical or syntactic homogenization, potentially simplifying internal circuits. Moreover, large-scale synthetic corpora training may also distort internal representations in ways that obscure the origins of linguistic behavior [20,190]. Whether mechanistic insights derived from current models will generalize to future systems trained predominantly on synthetic data remains an open and pressing question.

6. Conclusion

This survey reviewed MI advances for understanding linguistic abilities in LLMs, from static probing to circuit discovery [32] and causal intervention [110]. Together, these methods provide evidence about candidate mechanisms underlying grammatical and semantic competence and support a hybrid view in which LLMs can exhibit structured abstraction while still relying on surface heuristics.
The reviewed evidence is difficult to reconcile with a pure memorization account: several studies identify reusable circuits for grammatical phenomena such as indirect object identification and subject–verb agreement [112,161]. At the same time, these capabilities are implemented through distributed, polysemantic representations [39,107] and frequency-dependent learning dynamics [21,165], which differ from classical symbolic accounts of grammar. In multilingual settings, this complexity is further reflected in evidence for language-specific components [154,182] and proposed shared semantic representations such as a model-internal lingua franca [171,175].
Looking ahead, progress will hinge on addressing the trade-off between model scale and interpretability [96], developing evaluation benchmarks that extend beyond local syntactic competence [34], expanding multilingual replication beyond high-resource languages, and triangulating MI with behavioral and cognitive evidence.

Limitations

While this survey attempts to provide a critical synthesis of MI as applied to linguistic competence, several constraints limit its scope.
First, our synthesis is inherently restricted by the monolingual bias in existing MI literature. While we dedicate a section to multilingualism, the majority of available research focuses on English or high-resource languages; consequently, our conclusions regarding internal linguistic mechanisms may not generalize to typologically diverse or low-resource languages where data scarcity affects representational development [20,54].
Second, we acknowledge a methodological risk of inflation in our synthesis. Because MI research often relies on “toy” settings or templated prompts to isolate circuits, the significance of these findings might be over-represented in this survey when compared to the model’s behavior on complex, naturalistic language. As argued by [105], what we identify as a “linguistic circuit” may be a compensatory pathway, and our survey is limited by the current field’s inability to definitively distinguish between these two.
Finally, we restricted our scope to Transformer-based architectures. As new paradigms like State Space Models (e.g., Mamba) emerge, the applicability of attention-based interpretability techniques [125] will need to be re-evaluated.

Appendix A

Appendix A.1. Related Works

Appendix A.1.1. Foundational Interpretability Studies

Early interpretability research on RNNs and LSTMs used behavioral diagnostics and hidden-state analysis to show that individual neurons capture long-distance syntactic dependencies [92,99].
As focus shifted to transformers (e.g., BERT, GPT), studies utilized causal and correlational methods to link attention heads and hidden states to linguistic phenomena [45,119,130,168]. These models distribute linguistic information across layers for increased compositional flexibility [63], with early surveys [137] synthesizing these insights to establish a basis for mechanistic interpretability.

Appendix A.1.2. Benchmarks and Linguistic Probing

Researchers evaluate grammatical competence using controlled benchmarks, primarily employing the minimal pairs paradigm to compare model likelihoods for grammatical vs. ungrammatical sentences [145,147,152,164,172] (See Table 2 for a summary).
To minimize surface heuristic reliance, more rigorous approaches include SyntaxGym [49] for standardized evaluation, Holmes [160] for classifier-based hidden-state probing, and HLB [35] for assessing psycholinguistic humanlikeness. Furthermore, BHASA [93] extends these diagnostics to Southeast Asian languages.
However, because these benchmarks measure behavioral outcomes rather than causal mechanisms, they risk rewarding spurious correlations [7]. This limitation drives the shift toward MI, which targets the internal computational structures that causally generate linguistic behavior.

Appendix A.1.3. Shift from Model-Agnostic to Model-Specific Interpretability

To move beyond model behavior, interpretability has shifted from model-agnostic saliency methods to transformer-specific techniques [75,137]. By explicitly targeting structural components like attention heads, feed-forward layers, and the residual stream, these methods enable precise causal interventions and analysis of information flow [40,113,125].
This transition is bolstered by integrated toolkits such as the LM Transparency Tool [158], Neuronpedia [97], and TransformerLens [122].

Appendix A.1.4. Running Example: Subject–Verb Agreement

Subject–verb agreement illustrates how the methods above form a cumulative evidential chain rather than separate tool descriptions. Behavioral minimal pairs first test whether a model prefers The keys are over The keys is, but such success may still reflect local frequency or lexical memorization. Vocabulary projection can then show when number-consistent verb predictions become available across layers. Activation patching tests whether specific heads, MLP activations, or residual-stream states are causally necessary for transferring number information from the subject to the verb position. SAEs and probes can identify candidate number features, while circuit discovery asks whether those features compose into a reusable pathway that survives attractor nouns and lexical substitutions. A strong mechanistic claim therefore requires more than one positive result: it requires causal necessity, generalization beyond templates, and evidence against simpler frequency-based explanations.
Figure 2. A taxonomy of the MI landscape for LLMs. Click on the nodes to jump to the corresponding section.
Figure 2. A taxonomy of the MI landscape for LLMs. Click on the nodes to jump to the corresponding section.
Preprints 223887 g002
Table 2. A comparative summary of minimal pair datasets. The list includes monolingual and multilingual benchmarks, detailing their target language, total number of samples, and the count of specific linguistic paradigms evaluated.
Table 2. A comparative summary of minimal pair datasets. The list includes monolingual and multilingual benchmarks, detailing their target language, total number of samples, and the count of specific linguistic paradigms evaluated.
Dataset Language Size Paradigms
BLiMP [164] English 67k 67
CLiMP [172] Chinese 16k 16
SLING [147] Chinese 38k 38
ZhoBLiMP [103] Chinese 35k 118
JBLiMP [145] Japanese 331 39
RuBLiMP [152] Russian 45k 45
BLiMP-NL [150] Dutch 8.4k 84
QFrBLiMP [12] Quebec-French 1761 20
TurBLiMP [11] Turkish 16k 16
UrBLiMP [1] Urdu 5696 19
Irish-BLiMP [108] Irish 1020 102
Arabic MPs [4] Arabic 3000 9
BLiMP-IT [9] Italian 2899 78
CLAMS [118] English, French, German, Hebrew and Russian 229.9k 7
MultiBLiMP 1.0 [81] 101 languages 128321 2
BHS [90] Basque, Hindi, Swahili 300 3
Table 3. Comparison of core interpretability methods and intervention methods in LLMs. Each method is characterized by its representative studies, key idea, strengths and limitations.
Table 3. Comparison of core interpretability methods and intervention methods in LLMs. Each method is characterized by its representative studies, key idea, strengths and limitations.
Methodology Representative Works Key Idea Strengths and Limitations
Methodology Representative Works Key Idea Strengths and Limitations
Vocabulary Projection Logit Lens [50,124]; Tuned Lens [14]; Future Lens [128]. Decodes intermediate hidden states into vocabulary tokens to reveal the model’s evolving predictions layer-by-layer. Strengths: Maps word vectors across languages rapidly; no model retraining needed. Limitations: Ineffective for distant languages; performance depends on initial dictionary quality; struggles with new words or complex tasks.
Patching Activation Patching/Attribution Patching [7,43,110,121,151,177] Identifies causally decisive components by swapping activations between inputs or using gradient-based approximations. Strengths: Pinpoints layers/heads encoding key information; enables transplant experiments; shows causal roles of activations. Limitations: Sensitive to superposition; computationally costly.
Sparse Autoencoders SAE, Transcoder, Cross-layer Transcoder, Crosscoder and Binary Autoencoder [16,26,39,76] Decomposes polysemantic neurons into sparse, interpretable feature directions, resolving superposition. Strengths: Disentangles polysemantic neurons; reveals interpretable subspaces; quantifies feature entanglement. Limitations: Requires training or specialized modeling; imperfect disentanglement.
Automated Circuit Discovery Circuit Tracing [5,32,107] Automatically finds minimal subgraphs (component- or feature-level) that implement specific tasks. Strengths: Traces information flow via minimal subgraphs; reveals emergent “circuit reuse”; mechanistically grounded. Limitations: Hard to automate fully; scalability remains challenging.
Activation Steering ActAdd, CAA, STA [42,110,111,134,157,159,163] Manipulates model behavior via inference-time activation injection (Steering) or permanent weight modification (Editing). Strengths: Directly alters internal behavior; useful for debiasing and control; connects interpretability with intervention. Limitations: Can disrupt unrelated behaviors; hard to find minimal safe edits.
Table 4. Representative papers in MI. We categorize works into Method (foundational techniques, tools, frameworks, or benchmarks) and Application (studies leveraging these methods to investigate specific linguistic capabilities). The table spans five core domains: Probing, Vocabulary Projection (Logit Lens), Activation Patching, Sparse Autoencoders (SAE), and Circuit Discovery.
Table 4. Representative papers in MI. We categorize works into Method (foundational techniques, tools, frameworks, or benchmarks) and Application (studies leveraging these methods to investigate specific linguistic capabilities). The table spans five core domains: Probing, Vocabulary Projection (Logit Lens), Activation Patching, Sparse Autoencoders (SAE), and Circuit Discovery.
Paper Title Type Venue
Benchmarks
MIB: A Mechanistic Interpretability Benchmark [117] Method ICML
AXBENCH: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders [169] Method ICML
SAEBench: A Comprehensive Benchmark for Sparse Autoencoders [84] Method ICML
InterpBench: Semi-Synthetic Transformers for Evaluating MI Techniques [56] Method NIPS
Thank You, Stingray: Multilingual Large Language Models Can Not (Yet) Disambiguate Cross-Lingual Word Senses [18] Method NAACL
CausalGym: Benchmarking causal interpretability methods on linguistic tasks [7] Method ACL
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models [160] Method TACL
SyntaxGym: An Online Platform for Targeted Evaluation of Language Models [49] Method ACL
Probing
Lexical Popularity: Quantifying the Impact of Pre-training for LLM Performance [136] Method Arxiv
Steering Embedding Models with Geometric Rotation: Mapping Semantic Relationships Across Languages and Models [46] Method Arxiv
Emergence of Phonemic, Syntactic, and Semantic Representations in Artificial Neural Networks [126] Application Arxiv
ShifCon: Enhancing Non-Dominant Language Capabilities with a Shift-based Multilingual Contrastive Framework [178] Method ACL
Lost in Multilinguality: Dissecting Cross-lingual Factual Inconsistency in Transformer Language Models [162] Application ACL
Probing Syntax in Large Language Models: Successes and Remaining Challenges [34] Application COLM
Probing Internal Representations of Multi-Word Verbs in Large Language Models [87] Application MWE
Can Cross-Lingual Transferability of Multilingual Transformers Be Activated Without End-Task Data? [24] Method ACL
Finding Neurons in a Haystack: Case Studies with Sparse Probing [61] Application Arxiv
Probing Classifiers: Promises, Shortcomings, and Advances [13] Method CL
First Align, then Predict: Understanding the Cross-Lingual Ability of Multilingual BERT [120] Application EACL
Finding Universal Grammatical Relations in Multilingual BERT [23] Application ACL
On the Language Neutrality of Pre-trained Multilingual Representations [95] Application EMNLP
Emergent Linguistic Structure in Artificial Neural Networks Trained by Self-supervision [106] Application PNAS
A Structural Probe for Finding Syntax in Word Representations [69] Application NAACL
Under the Hood: Using Diagnostic Classifiers to Investigate and Improve how Language Models Track Agreement Information [51] Application BlackboxNLP
Vocabulary Projection (Logit Lens)
Eliciting Latent Predictions from Transformers with the Tuned Lens [14] Method Arxiv
Future Lens: Anticipating Subsequent Tokens from a Single Hidden State [128] Method CoNLL
Interpreting GPT: the logit lens [124] Method Blog
Unraveling Syntax: How Language Models Learn Context-Free Grammars [138] Application Arxiv
The Semantic Hub Hypothesis: Language Models Share Semantic Representations [171] Application ICLR
Do Multilingual LLMs Think in English? [139] Application Workshop
Do Llamas Work in English? On the Latent Language of Multilingual Transformers [166] Application ACL
DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models [31] Application ICLR
Sparse AutoEncoders (SAE)
Are Sparse Autoencoders Useful? A Case Study in Sparse Probing [82] Method ICML
Binary Autoencoder for Mechanistic Interpretability of Large Language Models [26] Method Arxiv
LinguaLens: Towards Interpreting Linguistic Mechanisms via Sparse Auto-Encoder [79] Method EMNLP
SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models [67] Method Arxiv
A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious Minima [155] Method Arxiv
The Grounding Gap: How LLMs Anchor the Meaning of Abstract Concepts Differently from Humans [25] Application Arxiv
From Syntax to Emotion: A Mechanistic Analysis of Emotion Inference in LLMs [142] Application Arxiv
Cross-Architecture Model Diffing with Crosscoders: Unsupervised Discovery of Differences Between LLMs [80] Application Arxiv
What’s the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering [104] Application ICLR
Crosscoding Through Time: Tracking Emergence of Linguistic Representations [10] Application ICML
Large Language Models Share Representations of Latent Grammatical Concepts [17] Application NAACL
Incremental Sentence Processing Mechanisms in Autoregressive Transformer Language Models [62] Application NAACL
Sparse Autoencoders Can Capture Language-Specific Concepts Across Diverse Languages [6] Application Arxiv
Unveiling Language-Specific Features via Sparse Autoencoders [33] Application ACL
Analyzing Multilingualism in Large Language Models with Sparse Autoencoders [27] Application COLM
Sparse Autoencoders Can Capture Language-Specific Concepts [6] Application Arxiv
Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders [64] Application Arxiv
Causal Language Control in Multilingual Transformers via Sparse Feature Steering [30] Application Arxiv
Semantic Convergence: Investigating Shared Representations Across Scaled LLMs [135] Application SRW
Extended Abstract for “Linguistic Universals”: Emergent Shared Features in Independent Monolingual Language Models via Sparse Autoencoders [187] Application MRL
Sparse Autoencoders Find Highly Interpretable Features in Language Models [76] Method ICLR
Transcoders Find Interpretable LLM Feature Circuits [38] Method NIPS
Activation Patching
Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context [57] Method Arxiv
How to use and interpret activation patching [68] Method Arxiv
Is This the Subspace You Are Looking for? An Interpretability Illusion [105] Method ICLR
Towards Best Practices of Activation Patching in Language Models [177] Method ICLR
Function Vectors in Large Language Models [157] Method ICLR
Attribution Patching Outperforms Automated Circuit Discovery [151] Method Blackbox
Attribution Patching: Activation Patching At Industrial Scale [121] Method BLOG
Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models [45] Method IJCNLP
The Dual-Route Model of Induction [44] Application COLM
Neuron Analysis
From Directions to Regions: Decomposing Activations in Language Models via Local Geometry[140] Method Arxiv
LANDeRMT: Dectecting and Routing Language-Aware Neurons for Selectively Finetuning LLMs to Machine Translation [189] Method ACL
Task-Specific Skill Localization in Fine-tuned Language Models [129] Method ICML
Language Arithmetics: Towards Systematic Language Neuron Identification and Manipulation [58] Application AACL
Cross-Lingual Generalization and Compression [133] Application ACL
How Syntax Specialization Emerges in Language Models [36] Application Arxiv
Language Lives in Sparse Dimensions: Interpretable Multilingual Control [186] Application Arxiv
Inducing Dyslexia in Vision Language Models [70] Application Arxiv
Different types of syntactic agreement recruit the same units [89] Application Arxiv
Sparse Subnetwork Enhancement for Underrepresented Languages in Large Language Models [59] Application Arxiv
Multilingual Knowledge Editing with Language-Agnostic Factual Neurons [181] Application COLING
From Language to Cognition: How LLMs Outgrow the Human Language Network [3] Application EMNLP
The Transfer Neurons Hypothesis [156] Application EMNLP
Mechanistic Understanding and Mitigation of Language Confusion in English-Centric Large Language Models [123] Application EMNLP
The Rise and Down of Babel Tower: Investigating the Evolution Process of Multilingual Code Large Language Model [22] Application ICLR
The LLM Language Network: A Neuroscientific Approach [2] Application NAACL
Language-Specific Neurons Do Not Facilitate Cross-Lingual Transfer [114] Application Workshop
Language-Specific Neurons: The Key to Multilingual Capabilities [154] Application ACL
Unveiling Linguistic Regions in Large Language Models [182] Application ACL
Converging to a Lingua Franca: Evolution of Linguistic Regions [175] Application COLING
Unveiling Language Competence Neurons: A Psycholinguistic Approach [37] Application COLING
Linguistic Minimal Pairs Elicit Linguistic Similarity [188] Application COLING
Neuron-Level Knowledge Attribution in Large Language Models [174] Application EMNLP
Neuron Specialization: Leveraging Intrinsic Task Modularity for Multilingual Machine Translation [153] Application EMNLP
Decoding Probing: Revealing Internal Linguistic Structures [65] Application LREC
On the Multilingual Ability: Finding and Controlling Language-Specific Neurons [88] Application NAACL
How do Large Language Models Handle Multilingualism? [184] Application NIPS
Universal Neurons in GPT2 Language Models [60] Application TMLR
Same Neurons, Different Languages [148] Application NAACL
Importance-based Neuron Allocation for Multilingual Neural Machine Translation [173] Application ACL
Circuit Discovery
Circuit Tracing: Revealing Computational Graphs in Language Models [5] Method Anthropic
Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs [107] Method ICLR
Scaling Sparse Feature Circuits For Studying In-Context Learning [86] Method ICML
Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPT [66] Method Arxiv
Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small [161] Method ICLR
Towards Automated Circuit Discovery for Mechanistic Interpretability [32] Method NIPS
Between Circuits and Chomsky: Pre-pretraining on Formal Languages [74] Application ACL
The Same But Different: Structural Similarities in Multilingual LM [180] Application ICLR
Circuit Component Reuse Across Tasks in Transformer Language Models [112] Application ICLR

B. Details of Findings

B.1. Emergence of Formal Linguistic Competence

Figure 3. Overview of the analysis in [10]. The study investigates how formal linguistic competence changes during model pretraining. It first identifies performance phase transitions on benchmarks such as BLiMP across training scales, then uses Crosscoder to learn a shared feature space for selected checkpoints. The analysis suggests that some syntactic features, including language-specific features acquired early in training, become more consolidated as scale and cross-lingual exposure increase, while other features become less salient. We treat this as evidence for training-dependent representational change, not as proof of symbolic rule acquisition.
Figure 3. Overview of the analysis in [10]. The study investigates how formal linguistic competence changes during model pretraining. It first identifies performance phase transitions on benchmarks such as BLiMP across training scales, then uses Crosscoder to learn a shared feature space for selected checkpoints. The analysis suggests that some syntactic features, including language-specific features acquired early in training, become more consolidated as scale and cross-lingual exposure increase, while other features become less salient. We treat this as evidence for training-dependent representational change, not as proof of symbolic rule acquisition.
Preprints 223887 g003

B.2. Partial Alignment with Human Language Processing

Figure 4. Overview of the analysis in [3]. The study relates neural alignment to formal linguistic competence and functional abilities during training. In the reported results, brain alignment rises early and then plateaus, while functional performance follows a different trajectory. We interpret this as evidence that alignment is selective and bounded: some language-related representations become more similar to human neural responses, but this does not imply broad human-like cognition or a shared implementation of semantic and pragmatic abilities.
Figure 4. Overview of the analysis in [3]. The study relates neural alignment to formal linguistic competence and functional abilities during training. In the reported results, brain alignment rises early and then plateaus, while functional performance follows a different trajectory. We interpret this as evidence that alignment is selective and bounded: some language-related representations become more similar to human neural responses, but this does not imply broad human-like cognition or a shared implementation of semantic and pragmatic abilities.
Preprints 223887 g004

B.3. Task-Specific Mechanisms vs. Shared Circuits

Figure 5. Overview of the circuit analysis in [112] for Colored Objects and Indirect Object Identification (IOI). The study reports overlap in heads involved in duplicate-token detection across the two tasks, while downstream heads use the duplicate signal differently. This supports a limited “shared subroutine, task-specific use” interpretation: some components can be reused across tasks with related structure, but the result should not be generalized to all linguistic phenomena or all models without further validation.
Figure 5. Overview of the circuit analysis in [112] for Colored Objects and Indirect Object Identification (IOI). The study reports overlap in heads involved in duplicate-token detection across the two tasks, while downstream heads use the duplicate signal differently. This supports a limited “shared subroutine, task-specific use” interpretation: some components can be reused across tasks with related structure, but the result should not be generalized to all linguistic phenomena or all models without further validation.
Preprints 223887 g005

B.4. Language-Selective Components

Figure 6. Overview of the language-selective neuron analysis in [88]. The study identifies neurons with stronger responses to particular languages (e.g., Chinese or German) and tests whether enhancing or suppressing them changes output language probabilities. We treat these interventions as evidence for language-sensitive components, while noting that steering output language is not the same as fully explaining multilingual competence.
Figure 6. Overview of the language-selective neuron analysis in [88]. The study identifies neurons with stronger responses to particular languages (e.g., Chinese or German) and tests whether enhancing or suppressing them changes output language probabilities. We treat these interventions as evidence for language-sensitive components, while noting that steering output language is not the same as fully explaining multilingual competence.
Preprints 223887 g006

B.5. Cross-Lingual Transfer

Figure 7. Overview of the Transfer Neurons Hypothesis proposed by [156]. The hypothesis models cross-lingual transfer as movement between language-specific latent spaces and a shared semantic latent space, mediated by specialized MLP neurons. This provides a candidate mechanism for cross-lingual routing, but its scope and causal sufficiency remain open questions across models, languages, and tasks.
Figure 7. Overview of the Transfer Neurons Hypothesis proposed by [156]. The hypothesis models cross-lingual transfer as movement between language-specific latent spaces and a shared semantic latent space, mediated by specialized MLP neurons. This provides a candidate mechanism for cross-lingual routing, but its scope and causal sufficiency remain open questions across models, languages, and tasks.
Preprints 223887 g007

B.6. Mechanisms of Semantic Alignment

Figure 8. Overview of the “Lingua Franca” hypothesis in [175]. The study proposes that multilingual models can map surface forms from different languages into a partially shared latent semantic space, with stronger alignment observed in larger or more extensively trained models. We present this as a useful hypothesis about semantic alignment, while leaving open how consistently it holds for low-resource languages, typologically distant languages, and non-translation tasks.
Figure 8. Overview of the “Lingua Franca” hypothesis in [175]. The study proposes that multilingual models can map surface forms from different languages into a partially shared latent semantic space, with stronger alignment observed in larger or more extensively trained models. We present this as a useful hypothesis about semantic alignment, while leaving open how consistently it holds for low-resource languages, typologically distant languages, and non-translation tasks.
Preprints 223887 g008

References

  1. Adeeba, F.; Dillon, B.; Sajjad, H.; Bhatt, R. UrBLiMP: A Benchmark for Evaluating the Linguistic Competence of Large Language Models in Urdu. Urblimp: A benchmark for evaluating the linguistic competence of large language models in urdu. 2025. Available online: https://arxiv.org/abs/2508.01006.
  2. AlKhamissi, B.; Tuckute, G.; Bosselut, A.; Schrimpf, M. The LLM Language Network: A Neuroscientific Approach for Identifying Causally Task-Relevant Units. The llm language network: A neuroscientific approach for identifying causally task-relevant units. 2025. Available online: https://arxiv.org/abs/2411.02280.
  3. AlKhamissi, B.; Tuckute, G.; Tang, Y.; Binhuraib, T.; Bosselut, A.; Schrimpf, M. From Language to Cognition: How LLMs Outgrow the Human Language Network. From language to cognition: How llms outgrow the human language network. 2025. Available online: https://arxiv.org/abs/2503.01830.
  4. Alrajhi, W. A.; Al-Khalifa, H.; AlSalman, A. Assessing the Linguistic Knowledge in Arabic Pre-trained Language Models Using Minimal Pairs. H. Bouamor et al. (), Proceedings of the Seventh Arabic Natural Language Processing Workshop (WANLP) Proceedings of the seventh arabic natural language processing workshop (wanlp) ( 185–193). Abu Dhabi, United Arab Emirates (Hybrid)Association for Computational Linguistics, 202212; Available online: https://aclanthology.org/2022.wanlp-1.17/. [CrossRef]
  5. Ameisen, E.; Lindsey, J.; Pearce, A.; Gurnee, W.; Turner, N. L.; Chen, B.; Citro, C.; Abrahams, D.; Carter, S.; Hosmer, B.; Marcus, J.; Sklar, M.; Templeton, A.; Bricken, T.; McDougall, C.; Cunningham, H.; Henighan, T.; Jermyn, A.; Jones, A.; Batson, J. Circuit Tracing: Revealing Computational Graphs in Language Models. Anthropic Technical Report. March 2025M. Available online: https://transformer-circuits.pub/2025/attribution-graphs/methods.html.
  6. Andrylie, L. M.; Rahmanisa, I.; Ihsani, M. K.; Wicaksono, A. F.; Wibowo, H. A.; Aji, A. F. Sparse Autoencoders Can Capture Language-Specific Concepts Across Diverse Languages. Sparse autoencoders can capture language-specific concepts across diverse languages. 2025. Available online: https://arxiv.org/abs/2507.11230.
  7. Arora, A.; Jurafsky, D.; Potts, C. CausalGym: Benchmarking causal interpretability methods on linguistic tasks. Causalgym: Benchmarking causal interpretability methods on linguistic tasks. 2024. Available online: https://arxiv.org/abs/2402.12560.
  8. Banayeeanzade, A.; Tak, A. N.; Bahrani, F.; Bolourani, A.; Blas, L.; Ferrara, E.; Gratch, J.; Karimireddy, S. P. Psychological Steering in LLMs: An Evaluation of Effectiveness and Trustworthiness. Psychological steering in llms: An evaluation of effectiveness and trustworthiness. 2025. Available online: https://arxiv.org/abs/2510.04484.
  9. Barbini, M.; Piccini Bianchessi, M. L.; Bressan, V.; Fusco, A.; Neri, S.; Rossi, S.; Sgrizzi, T.; Chesi, C. BLiMP-IT: Harnessing Automatic Minimal Pair Generation for Italian Language Model Evaluation. In Proceedings of the Eleventh Italian Conference on Computational Linguistics (CLiC-it 2025) Proceedings of the eleventh italian conference on computational linguistics (clic-it 2025); Cagliari, ItalyCEUR Workshop Proceedings, Bosco, C., Jezek, E., Polignano, M., Sanguinetti (), M., Eds.; 09 2025; pp. 64–71. Available online: https://aclanthology.org/2025.clicit-1.8/.
  10. Bayazit, D.; Mueller, A.; Bosselut, A. Crosscoding Through Time: Tracking Emergence & Consolidation Of Linguistic Representations Throughout LLM Pretraining. Crosscoding through time: Tracking emergence & consolidation of linguistic representations throughout llm pretraining. 2025. Available online: https://arxiv.org/abs/2509.05291.
  11. Başar, E.; Padovani, F.; Jumelet, J.; Bisazza, A. TurBLiMP: A Turkish Benchmark of Linguistic Minimal Pairs. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing Proceedings of the 2025 conference on empirical methods in natural language processing ( 16506–16521). Association for Computational Linguistics, Available online. 2025. [Google Scholar] [CrossRef]
  12. Beauchemin, D.; Veilleux, P.-L.; Khoury, R.; Roy, J.-P. QFrBLiMP: a Quebec-French Benchmark of Linguistic Minimal Pairs. Qfrblimp: a quebec-french benchmark of linguistic minimal pairs. 2025. Available online: https://arxiv.org/abs/2509.25664.
  13. Belinkov, Y. Computational Linguistics481207–219; Probing Classifiers: Promises, Shortcomings, and Advances. 03 2022. Available online: https://aclanthology.org/2022.cl-1.7/. [CrossRef]
  14. Belrose, N.; Ostrovsky, I.; McKinney, L.; Furman, Z.; Smith, L.; Halawi, D.; Biderman, S.; Steinhardt, J. Eliciting Latent Predictions from Transformers with the Tuned Lens. Eliciting latent predictions from transformers with the tuned lens. 2025. Available online: https://arxiv.org/abs/2303.08112.
  15. Bolhuis, J. J.; Crain, S.; Fong, S.; Moro, A. Three reasons why AI doesn’t model human language. Nature6278004489. Available online. 2024. [CrossRef]
  16. Bricken, T.; Templeton, A.; Batson, J.; Chen, B.; Jermyn, A.; Conerly, T.; Turner, N.; Anil, C.; Denison, C.; McCauley, H.; et al. Towards Monosemanticity: Decomposing Language Models With Dictionary Learning. Transformer Circuits Thread. 2023. Available online: https://transformer-circuits.pub/2023/monosemantic-features.
  17. Brinkmann, J.; Wendler, C.; Bartelt, C.; Mueller, A. Large Language Models Share Representations of Latent Grammatical Concepts Across Typologically Diverse Languages. L. Chiruzzo, A. Ritter, L. Wang (), Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies. In Long Papers) Proceedings of the 2025 conference of the nations of the americas chapter of the association for computational linguistics: Human language technologies (volume 1: Long papers) ( 6131–6150), 202504; Albuquerque, New MexicoAssociation for Computational Linguistics; Volume 1. Available online: https://aclanthology.org/2025.naacl-long.312/. [CrossRef]
  18. Cahyawijaya, S.; Zhang, R.; Cruz, J. C. B.; Lovenia, H.; Gilbert, E.; Nomoto, H.; Aji, A. F. Thank You, Stingray: Multilingual Large Language Models Can Not (Yet) Disambiguate Cross-Lingual Word Senses. L. Chiruzzo, A. Ritter, L. Wang (), Findings of the Association for Computational Linguistics: NAACL 2025 Findings of the association for computational linguistics: Naacl 2025 ( 3228–3250); Albuquerque, New MexicoAssociation for Computational Linguistics, 04 2025; Available online: https://aclanthology.org/2025.findings-naacl.178/. [CrossRef]
  19. Cai, Z. G.; Duan, X.; Haslett, D. A.; Wang, S.; Pickering, M. J. Do large language models resemble humans in language use? Do large language models resemble humans in language use? 2024. Available online: https://arxiv.org/abs/2303.08014.
  20. Chang, T. A.; Arnett, C.; Tu, Z.; Bergen, B. When Is Multilinguality a Curse? Language Modeling for 250 High- and Low-Resource Languages. Y. Al-Onaizan, M. Bansal, Y.-N. Chen (), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing Proceedings of the 2024 conference on empirical methods in natural language processing ( 4074–4096). Miami, Florida, USAAssociation for Computational Linguistics, 202411; Available online: https://aclanthology.org/2024.emnlp-main.236/. [CrossRef]
  21. Chang, T. A.; Tu, Z.; Bergen, B. K. Characterizing Learning Curves During Language Model Pre-Training: Learning, Forgetting, and Stability. Trans. Assoc. Comput. Linguist. 2024, 121346–1362. Available online: https://doi.org/10.1162/tacl_a_00708. [CrossRef]
  22. Chen, J.; Chen, W.; Su, J.; Xu, J.; Lin, H.; Ren, M.; Lu, Y.; Han, X.; Sun, L. The Rise and Down of Babel Tower: Investigating the Evolution Process of Multilingual Code Large Language Model. The Thirteenth International Conference on Learning Representations. The thirteenth international conference on learning representations., 2025; Available online: https://openreview.net/forum?id=eznTVIM3bs.
  23. Chi, E. A.; Hewitt, J.; Manning, C. D. Finding Universal Grammatical Relations in Multilingual BERT. Finding universal grammatical relations in multilingual bert. 2020. Available online: https://arxiv.org/abs/2005.04511.
  24. Chi, Z.; Huang, H.; Mao, X.-L. Can Cross-Lingual Transferability of Multilingual Transformers Be Activated Without End-Task Data? A. Rogers, J. Boyd-Graber, N. Okazaki (), Findings of the Association for Computational Linguistics: ACL 2023 Findings of the association for computational linguistics: Acl 2023 ( 12572–12584). Toronto, CanadaAssociation for Computational Linguistics. 07 2023. Available online: https://aclanthology.org/2023.findings-acl.796/. [CrossRef]
  25. Chlapanis, O. S.; Mastromichalakis, O. M.; Papadimitriou, C. H. The Grounding Gap: How LLMs Anchor the Meaning of Abstract Concepts Differently from Humans. The grounding gap: How llms anchor the meaning of abstract concepts differently from humans. 2026. Available online: https://arxiv.org/abs/2605.08837.
  26. Cho, H.; Yang, H.; Kurkoski, B. M.; Inoue, N. Binary Autoencoder for Mechanistic Interpretability of Large Language Models. Binary autoencoder for mechanistic interpretability of large language models. 2025. Available online: https://arxiv.org/abs/2509.20997.
  27. Cho, I.; Hockenmaier, J. Analyzing Multilingualism in Large Language Models with Sparse Autoencoders. Second Conference on Language Modeling. Second conference on language modeling., 2025; Available online: https://openreview.net/forum?id=NmGSvZoU3K.
  28. Chomsky, N. Noam Chomsky: The False Promise of ChatGPT. The New York Times; Opinion, 2023M. Available online: https://www.nytimes.com/2023/03/08/opinion/noam-chomsky-chatgpt-ai.html.
  29. Choshen, L.; Hacohen, G.; Weinshall, D.; Abend, O. The Grammar-Learning Trajectories of Neural Language Models. The grammar-learning trajectories of neural language models. 2022. Available online: https://arxiv.org/abs/2109.06096.
  30. Chou, C.-T.; Liu, G.; Sun, J.; Blondin, C.; Zhu, K.; Sharma, V.; O’Brien, S. Causal Language Control in Multilingual Transformers via Sparse Feature Steering. Causal language control in multilingual transformers via sparse feature steering. 2025. Available online: https://arxiv.org/abs/2507.13410.
  31. Chuang, Y.-S.; Xie, Y.; Luo, H.; Kim, Y.; Glass, J.; He, P. DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models. Dola: Decoding by contrasting layers improves factuality in large language models. 2024. Available online: https://arxiv.org/abs/2309.03883.
  32. Conmy, A.; Mavor-Parker, A. N.; Lynch, A.; Heimersheim, S.; Garriga-Alonso, A. Towards Automated Circuit Discovery for Mechanistic Interpretability. Thirty-seventh Conference on Neural Information Processing Systems. Thirty-seventh conference on neural information processing systems., 2023; Available online: https://openreview.net/forum?id=89ia77nZ8u.
  33. Deng, B.; Wan, Y.; Yang, B.; Zhang, Y.; Feng, F. Unveiling Language-Specific Features in Large Language Models via Sparse Autoencoders; Che, W., Nabende, J., Shutova, E., Pilehvar (), M. T., Eds.; Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 07 2025; Volume 1, Available online: https://aclanthology.org/2025.acl-long.229/. [CrossRef]
  34. Diego-Simón, P. J.; Chemla, E.; King, J.-R.; Lakretz, Y. Probing Syntax in Large Language Models: Successes and Remaining Challenges. Probing syntax in large language models: Successes and remaining challenges. 2025. Available online: https://arxiv.org/abs/2508.03211.
  35. Duan, X.; Xiao, B.; Tang, X.; Cai, Z. G. HLB: Benchmarking LLMs’ Humanlikeness in Language Use. Hlb: Benchmarking llms’ humanlikeness in language use. 2024. Available online: https://arxiv.org/abs/2409.15890.
  36. Duan, X.; Yao, Z.; Zhang, Y.; Wang, S.; Cai, Z. G. How Syntax Specialization Emerges in Language Models. How syntax specialization emerges in language models. 2025. Available online: https://arxiv.org/abs/2505.19548.
  37. Duan, X.; Zhou, X.; Xiao, B.; Cai, Z. Unveiling Language Competence Neurons: A Psycholinguistic Approach to Model Interpretability; Rambow, O., Wanner, L., Apidianaki, M., Al-Khalifa, H., Eugenio, B. D., Schockaert (), S., Eds.; Proceedings of the 31st International Conference on Computational Linguistics Proceedings of the 31st international conference on computational linguistics ( 10148–10157): Abu Dhabi, UAEAssociation for Computational Linguistics, 01 2025; Available online: https://aclanthology.org/2025.coling-main.677/.
  38. Dunefsky, J.; Chlenski, P.; Nanda, N. Transcoders Find Interpretable LLM Feature Circuits. Transcoders find interpretable llm feature circuits. 2024. Available online: https://arxiv.org/abs/2406.11944.
  39. Elhage, N.; Hume, T.; Olsson, C.; Schiefer, N.; Henighan, T.; Kravec, S.; Hatfield-Dodds, Z.; Lasenby, R.; Drain, D.; Chen, C.; Grosse, R.; McCandlish, S.; Kaplan, J.; Amodei, D.; Wattenberg, M.; Olah, C. Toy Models of Superposition. Toy models of superposition. 2022. Available online: https://arxiv.org/abs/2209.10652.
  40. Elhage, N.; Nanda, N.; Olsson, C.; Henighan, T.; Joseph, N.; Mann, B.; Askell, A.; Bai, Y.; Chen, A.; Conerly, T.; DasSarma, N.; Drain, D.; Ganguli, D.; Hatfield-Dodds, Z.; Hernandez, D.; Jones, A.; Kernion, J.; Lovitt, L.; Ndousse, K.; Olah, C. A Mathematical Framework for Transformer Circuits. Transformer Circuits Thread. 2021. Available online: https://transformer-circuits.pub/2021/framework/index.html.
  41. Evanson, L.; Lakretz, Y.; King, J.-R. Language acquisition: do children and language models follow similar learning stages? Language acquisition: do children and language models follow similar learning stages? 2023. Available online: https://arxiv.org/abs/2306.03586.
  42. Fang, J.; Jiang, H.; Wang, K.; Ma, Y.; Shi, J.; Wang, X.; He, X.; Chua, T.-S. AlphaEdit: Null-Space Constrained Model Editing for Language Models. The Thirteenth International Conference on Learning Representations. The thirteenth international conference on learning representations., 2025; Available online: https://openreview.net/forum?id=HvSytvg3Jh.
  43. Ferrando, J.; Voita, E. Information Flow Routes: Automatically Interpreting Language Models at Scale. Information flow routes: Automatically interpreting language models at scale. 2024. Available online: https://arxiv.org/abs/2403.00824.
  44. Feucht, S.; Todd, E.; Wallace, B.; Bau, D. The Dual-Route Model of Induction. The dual-route model of induction. 2025. Available online: https://arxiv.org/abs/2504.03022.
  45. Finlayson, M.; Mueller, A.; Gehrmann, S.; Shieber, S.; Linzen, T.; Belinkov, Y. Causal Analysis of Syntactic Agreement Mechanisms in Neural Language Models. C. Zong, F. Xia, W. Li, R. Navigli (), Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing. In Long Papers) Proceedings of the 59th annual meeting of the association for computational linguistics and the 11th international joint conference on natural language processing (volume 1: Long papers) ( 1828–1843); OnlineAssociation for Computational Linguistics, 08 2021; Volume 1, Available online: https://aclanthology.org/2021.acl-long.144/. [CrossRef]
  46. Freenor, M.; Alvarez, L. Mapping Semantic & Syntactic Relationships with Geometric Rotation. Mapping semantic & syntactic relationships with geometric rotation. 2026. Available online: https://arxiv.org/abs/2510.09790.
  47. Gao, C.; Ma, Z.; Chen, J.; Li, P.; Huang, S.; Li, J. Increasing alignment of large language models with language processing in the human brain. Nature computational science1–11. 2025. Available online: https://www.nature.com/articles/s43588-025-00863-0.
  48. Gao, Y.; Meng, Q.; Zhou, Y.; Pan, L. Towards Intrinsic Interpretability of Large Language Models:A Survey of Design Principles and Architectures. Towards intrinsic interpretability of large language models:a survey of design principles and architectures. 2026. Available online: https://arxiv.org/abs/2604.16042.
  49. Gauthier, J.; Hu, J.; Wilcox, E.; Qian, P.; Levy, R. SyntaxGym: An Online Platform for Targeted Evaluation of Language Models. A. Celikyilmaz T.-H. Wen (), Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics: System Demonstrations Proceedings of the 58th annual meeting of the association for computational linguistics: System demonstrations ( 70–76). OnlineAssociation for Computational Linguistics. 07 2020. Available online: https://aclanthology.org/2020.acl-demos.10/. [CrossRef]
  50. Geva, M.; Caciularu, A.; Wang, K.; Goldberg, Y. Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing Proceedings of the 2022 conference on empirical methods in natural language processing ( 30–45); Abu Dhabi, United Arab EmiratesAssociation for Computational Linguistics, Goldberg, Y., Kozareva, Z., Zhang (), Y., Eds.; 12 2022; Available online: https://aclanthology.org/2022.emnlp-main.3/. [CrossRef]
  51. Giulianelli, M.; Harding, J.; Mohnert, F.; Hupkes, D.; Zuidema, W. Under the Hood: Using Diagnostic Classifiers to Investigate and Improve how Language Models Track Agreement Information. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP Proceedings of the 2018 EMNLP workshop BlackboxNLP: Analyzing and interpreting neural networks for NLP ( 240–248); Brussels, BelgiumAssociation for Computational Linguistics, Linzen, T., Chrupała, G., Alishahi (), A., Eds.; 11 2018; Available online: https://aclanthology.org/W18-5426/. [CrossRef]
  52. Goldstein, A.; Ham, E.; Schain, M.; Nastase, S. A.; Aubrey, B.; Zada, Z.; Grinstein-Dabush, A.; Gazula, H.; Feder, A.; Doyle, W.; et al. Temporal structure of natural language processing in the human brain corresponds to layered hierarchy of large language models. Nature communications16110529. Available online. 2025. [CrossRef]
  53. Graichen, N.; de Dios-Flores, I.; Boleda, G. The Grammar of Transformers: A Systematic Review of Interpretability Research on Syntactic Knowledge in Language Models. The grammar of transformers: A systematic review of interpretability research on syntactic knowledge in language models. 2026. Available online: https://arxiv.org/abs/2601.19926.
  54. Guo, Y.; Conia, S.; Zhou, Z.; Li, M.; Potdar, S.; Xiao, H. Do Large Language Models Have an English Accent? Evaluating and Improving the Naturalness of Multilingual LLMs. Do large language models have an english accent? evaluating and improving the naturalness of multilingual llms. 2025. Available online: https://arxiv.org/abs/2410.15956.
  55. Guo, Y.; Shang, G.; Vazirgiannis, M.; Clavel, C. The Curious Decline of Linguistic Diversity: Training Language Models on Synthetic Text. K. Duh, H. Gomez, S. Bethard (), Findings of the Association for Computational Linguistics: NAACL 2024 Findings of the association for computational linguistics: Naacl 2024 ( 3589–3604). Mexico City, MexicoAssociation for Computational Linguistics. 06 2024. Available online: https://aclanthology.org/2024.findings-naacl.228/. [CrossRef]
  56. Gupta, R.; Arcuschin, I.; Kwa, T.; Garriga-Alonso, A. InterpBench: Semi-Synthetic Transformers for Evaluating Mechanistic Interpretability Techniques. Interpbench: Semi-synthetic transformers for evaluating mechanistic interpretability techniques. 2025. Available online: https://arxiv.org/abs/2407.14494.
  57. Gur-Arieh, Y.; Geva, M.; Geiger, A. Mixing Mechanisms: How Language Models Retrieve Bound Entities In-Context. Mixing mechanisms: How language models retrieve bound entities in-context. 2025. Available online: https://arxiv.org/abs/2510.06182.
  58. Gurgurov, D.; Trinley, K.; Ghussin, Y. A.; Baeumel, T.; van Genabith, J.; Ostermann, S. Language Arithmetics: Towards Systematic Language Neuron Identification and Manipulation. Language arithmetics: Towards systematic language neuron identification and manipulation. 2025. Available online: https://arxiv.org/abs/2507.22608.
  59. Gurgurov, D.; van Genabith, J.; Ostermann, S. Sparse Subnetwork Enhancement for Underrepresented Languages in Large Language Models. Sparse subnetwork enhancement for underrepresented languages in large language models. 2025. Available online: https://arxiv.org/abs/2510.13580.
  60. Gurnee, W.; Horsley, T.; Guo, Z. C.; Kheirkhah, T. R.; Sun, Q.; Hathaway, W.; Nanda, N.; Bertsimas, D. Universal Neurons in GPT2 Language Models. Transactions on Machine Learning Research. 2024. Available online: https://openreview.net/forum?id=ZeI104QZ8I.
  61. Gurnee, W.; Nanda, N.; Pauly, M.; Harvey, K.; Troitskii, D.; Bertsimas, D. Finding Neurons in a Haystack: Case Studies with Sparse Probing. Finding neurons in a haystack: Case studies with sparse probing. 2023. Available online: https://arxiv.org/abs/2305.01610.
  62. Hanna, M.; Mueller, A. Incremental Sentence Processing Mechanisms in Autoregressive Transformer Language Models. L. Chiruzzo, A. Ritter, L. Wang (), Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies. In Long Papers) Proceedings of the 2025 conference of the nations of the americas chapter of the association for computational linguistics: Human language technologies (volume 1: Long papers) ( 3181–3203), 202504; Albuquerque, New MexicoAssociation for Computational Linguistics; Volume 1. Available online: https://aclanthology.org/2025.naacl-long.164/. [CrossRef]
  63. Hao, S.; Linzen, T. Verb Conjugation in Transformers Is Determined by Linear Encodings of Subject Number. H. Bouamor, J. Pino, K. Bali (), Findings of the Association for Computational Linguistics: EMNLP 2023 Findings of the association for computational linguistics: Emnlp 2023 ( 4531–4539). SingaporeAssociation for Computational Linguistics. 12 2023. Available online: https://aclanthology.org/2023.findings-emnlp.300/. [CrossRef]
  64. Harrasse, A.; Draye, F.; Jin, Z.; Schölkopf, B. Tracing Multilingual Representations in LLMs with Cross-Layer Transcoders. Tracing multilingual representations in llms with cross-layer transcoders. 2025. Available online: https://arxiv.org/abs/2511.10840.
  65. He, L.; Chen, P.; Nie, E.; Li, Y.; Brennan, J. R. Decoding Probing: Revealing Internal Linguistic Structures in Neural Language Models using Minimal Pairs. Decoding probing: Revealing internal linguistic structures in neural language models using minimal pairs. 2024. Available online: https://arxiv.org/abs/2403.17299.
  66. He, Z.; Ge, X.; Tang, Q.; Sun, T.; Cheng, Q.; Qiu, X. Dictionary Learning Improves Patch-Free Circuit Discovery in Mechanistic Interpretability: A Case Study on Othello-GPT. Dictionary learning improves patch-free circuit discovery in mechanistic interpretability: A case study on othello-gpt. 2024. Available online: https://arxiv.org/abs/2402.12201.
  67. He, Z.; Zhao, H.; Qiao, Y.; Yang, F.; Payani, A.; Ma, J.; Du, M. SAIF: A Sparse Autoencoder Framework for Interpreting and Steering Instruction Following of Language Models. Saif: A sparse autoencoder framework for interpreting and steering instruction following of language models. 2025. Available online: https://arxiv.org/abs/2502.11356.
  68. Heimersheim, S.; Nanda, N. How to use and interpret activation patching. How to use and interpret activation patching. 2024. Available online: https://arxiv.org/abs/2404.15255.
  69. Hewitt, J.; Manning, C. D. A Structural Probe for Finding Syntax in Word Representations; Burstein, J., Doran, C., Solorio (), T., Eds.; Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 06 2019; Volume 1, Available online: https://aclanthology.org/N19-1419/. [CrossRef]
  70. Honarmand, M.; Sharma, A.; AlKhamissi, B.; Mehrer, J.; Schrimpf, M. Inducing Dyslexia in Vision Language Models. Inducing dyslexia in vision language models. 2025. Available online: https://arxiv.org/abs/2509.24597.
  71. Hosseini, E. A.; Schrimpf, M.; Zhang, Y.; Bowman, S.; Zaslavsky, N.; Fedorenko, E. Artificial Neural Network Language Models Predict Human Brain Responses to Language Even After a Developmentally Realistic Amount of Training. Neurobiol. Lang. Available online. 2024, 5143–63. [Google Scholar] [CrossRef] [PubMed]
  72. Hu, J.; Gauthier, J.; Qian, P.; Wilcox, E.; Levy, R. A Systematic Assessment of Syntactic Generalization in Neural Language Models. In OnlineAssociation for Computational Linguistics; Jurafsky, D., Chai, J., Schluter, N., Tetreault (), J., Eds.; Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics Proceedings of the 58th annual meeting of the association for computational linguistics ( 1725–1744), 07 2020; Available online: https://aclanthology.org/2020.acl-main.158/. [CrossRef]
  73. Hu, J.; Levy, R. P. Prompting is not a substitute for probability measurements in large language models. The 2023 Conference on Empirical Methods in Natural Language Processing. The 2023 conference on empirical methods in natural language processing., 2023; Available online: https://openreview.net/forum?id=hMqRphmoM9.
  74. Hu, M. Y.; Petty, J.; Shi, C.; Merrill, W.; Linzen, T. Between Circuits and Chomsky: Pre-pretraining on Formal Languages Imparts Linguistic Biases; Che, W., Nabende, J., Shutova, E., Pilehvar (), M. T., Eds.; Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 07 2025; Volume 1, Available online: https://aclanthology.org/2025.acl-long.478/. [CrossRef]
  75. Huang, J.; Geiger, A.; D’Oosterlinck, K.; Wu, Z.; Potts, C. Rigorously Assessing Natural Language Explanations of Neurons. Rigorously assessing natural language explanations of neurons. 2023. Available online: https://arxiv.org/abs/2309.10312.
  76. Huben, R.; Cunningham, H.; Smith, L. R.; Ewart, A.; Sharkey, L. Sparse Autoencoders Find Highly Interpretable Features in Language Models. The Twelfth International Conference on Learning Representations. The twelfth international conference on learning representations., 2024; Available online: https://openreview.net/forum?id=F76bwRSLeK.
  77. Huo, J.; Yan, Y.; Hu, B.; Yue, Y.; Hu, X. MMNeuron: Discovering Neuron-Level Domain-Specific Interpretation in Multimodal Large Language Model. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing Proceedings of the 2024 conference on empirical methods in natural language processing ( 6801–6816), Available online. 2024; Association for Computational Linguistics. [Google Scholar] [CrossRef]
  78. Iqbal, A.; Younas, M.; Iftikhar, S.; Fatima, F.; Saleem, R. Spam detection using hybrid model on fusion of spammer behavior and linguistics features. Egyptian Informatics Journal29100605. 2025. Available online: https://www.sciencedirect.com/science/article/pii/S1110866524001683. [CrossRef]
  79. Jing, Y.; Yao, Z.; Guo, H.; Ran, L.; Wang, X.; Hou, L.; Li, J. LinguaLens: Towards Interpreting Linguistic Mechanisms of Large Language Models via Sparse Auto-Encoder. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing Proceedings of the 2025 conference on empirical methods in natural language processing ( 28232–28251); Suzhou, ChinaAssociation for Computational Linguistics, Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng (), V., Eds.; 11 2025; Available online: https://aclanthology.org/2025.emnlp-main.1433/. [CrossRef]
  80. Jiralerspong, T.; Bricken, T. Cross-Architecture Model Diffing with Crosscoders: Unsupervised Discovery of Differences Between LLMs. Cross-architecture model diffing with crosscoders: Unsupervised discovery of differences between llms. 2026. Available online: https://arxiv.org/abs/2602.11729.
  81. Jumelet, J.; Weissweiler, L.; Nivre, J.; Bisazza, A. MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs. Multiblimp 1.0: A massively multilingual benchmark of linguistic minimal pairs. 2025. Available online: https://arxiv.org/abs/2504.02768.
  82. Kantamneni, S.; Engels, J.; Rajamanoharan, S.; Tegmark, M.; Nanda, N. Are Sparse Autoencoders Useful? A Case Study in Sparse Probing. Are sparse autoencoders useful? a case study in sparse probing. 2025. Available online: https://arxiv.org/abs/2502.16681.
  83. Karny, S.; Baez, A.; Pataranutaporn, P. Neural Transparency: Mechanistic Interpretability Interfaces for Anticipating Model Behaviors for Personalized AI. Neural transparency: Mechanistic interpretability interfaces for anticipating model behaviors for personalized ai. 2025. Available online: https://arxiv.org/abs/2511.00230.
  84. Karvonen, A.; Rager, C.; Lin, J.; Tigges, C.; Bloom, J.; Chanin, D.; Lau, Y.-T.; Farrell, E.; McDougall, C.; Ayonrinde, K.; Till, D.; Wearden, M.; Conmy, A.; Marks, S.; Nanda, N. SAEBench: A Comprehensive Benchmark for Sparse Autoencoders in Language Model Interpretability. Saebench: A comprehensive benchmark for sparse autoencoders in language model interpretability. 2025. Available online: https://arxiv.org/abs/2503.09532.
  85. Kendiukhov, I. A Review of Developmental Interpretability in Large Language Models. A review of developmental interpretability in large language models. 2025. Available online: https://arxiv.org/abs/2508.15841.
  86. Kharlapenko, D.; Shabalin, S.; Barez, F.; Conmy, A.; Nanda, N. Scaling sparse feature circuit finding for in-context learning. Scaling sparse feature circuit finding for in-context learning. 2025. Available online: https://arxiv.org/abs/2504.13756.
  87. Kissane, H.; Schilling, A.; Krauss, P. Probing Internal Representations of Multi-Word Verbs in Large Language Models. A. K. Ojha et al. (), Proceedings of the 21st Workshop on Multiword Expressions (MWE 2025) Proceedings of the 21st workshop on multiword expressions (mwe 2025) ( 7–13). Albuquerque, New Mexico, U.S.A.Association for Computational Linguistics, 202505; Available online: https://aclanthology.org/2025.mwe-1.2/. [CrossRef]
  88. Kojima, T.; Okimura, I.; Iwasawa, Y.; Yanaka, H.; Matsuo, Y. On the Multilingual Ability of Decoder-based Pre-trained Language Models: Finding and Controlling Language-Specific Neurons. On the multilingual ability of decoder-based pre-trained language models: Finding and controlling language-specific neurons. 2024. Available online: https://arxiv.org/abs/2404.02431.
  89. Kryvosheieva, D.; de Varda, A.; Fedorenko, E.; Tuckute, G. Different types of syntactic agreement recruit the same units within large language models. Different types of syntactic agreement recruit the same units within large language models. 2025. Available online: https://arxiv.org/abs/2512.03676.
  90. Kryvosheieva, D.; Levy, R. Controlled Evaluation of Syntactic Knowledge in Multilingual Language Models. Controlled evaluation of syntactic knowledge in multilingual language models. 2024. Available online: https://arxiv.org/abs/2411.07474.
  91. Kumar, S.; Sumers, T. R.; Yamakoshi, T.; Goldstein, A.; Hasson, U.; Norman, K. A.; Griffiths, T. L.; Hawkins, R. D.; Nastase, S. A. Shared functional specialization in transformer-based language models and the human brain. Nature communications1515523. Available online. 2024. [CrossRef]
  92. Lakretz, Y.; Kruszewski, G.; Desbordes, T.; Hupkes, D.; Dehaene, S.; Baroni, M. The emergence of number and syntax units in LSTM language models. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies; Burstein, J., Doran, C., Solorio (), T., Eds.; Minneapolis, MinnesotaAssociation for Computational Linguistics, 06 2019; Volume 1, Available online: https://aclanthology.org/N19-1002/. [CrossRef]
  93. Leong, W. Q.; Ngui, J. G.; Susanto, Y.; Rengarajan, H.; Sarveswaran, K.; Tjhi, W. C. BHASA: A Holistic Southeast Asian Linguistic and Cultural Evaluation Suite for Large Language Models. Bhasa: A holistic southeast asian linguistic and cultural evaluation suite for large language models. 2023. Available online: https://arxiv.org/abs/2309.06085.
  94. Li, X.; Yong, Z.-X.; Bach, S. H. Preference Tuning For Toxicity Mitigation Generalizes Across Languages. Preference tuning for toxicity mitigation generalizes across languages. 2024. Available online: https://arxiv.org/abs/2406.16235.
  95. Libovický, J.; Rosa, R.; Fraser, A. On the Language Neutrality of Pre-trained Multilingual Representations. T. Cohn, Y. He, Y. Liu (), Findings of the Association for Computational Linguistics: EMNLP 2020 Findings of the association for computational linguistics: Emnlp 2020 ( 1663–1674). OnlineAssociation for Computational Linguistics. 11 2020. Available online: https://aclanthology.org/2020.findings-emnlp.150/. [CrossRef]
  96. Lieberum, T.; Rahtz, M.; Kramár, J.; Nanda, N.; Irving, G.; Shah, R.; Mikulik, V. Does Circuit Analysis Interpretability Scale? Evidence from Multiple Choice Capabilities in Chinchilla. Does circuit analysis interpretability scale? evidence from multiple choice capabilities in chinchilla. 2023. Available online: https://arxiv.org/abs/2307.09458.
  97. Lin, J. Neuronpedia: Interactive Reference and Tooling for Analyzing Neural Networks. Neuronpedia: Interactive reference and tooling for analyzing neural networks. 2023. Available online: https://www.neuronpedia.org.
  98. Lindsey, J.; Gurnee, W.; Ameisen, E.; Chen, B.; Pearce, A.; Turner, N. L.; Citro, C.; Abrahams, D.; Carter, S.; Hosmer, B.; Marcus, J.; Sklar, M.; Templeton, A.; Bricken, T.; McDougall, C.; Cunningham, H.; Henighan, T.; Jermyn, A.; Jones, A.; Batson, J. On the Biology of a Large Language Model. Anthropic Technical Report. March 2025M. Available online: https://transformer-circuits.pub/2025/attribution-graphs/biology.html.
  99. Linzen, T.; Dupoux, E.; Goldberg, Y. Assessing the Ability of LSTMs to Learn Syntax-Sensitive Dependencies. Transactions of the Association for Computational Linguistics. 2016, pp. 4521–535. Available online: https://aclanthology.org/Q16-1037/. [CrossRef]
  100. Liu, W.; Xiang, M.; Ding, N. Active use of latent tree-structured sentence representation in humans and large language models. Nature Human Behaviour1–14. Available online. 2025. [CrossRef]
  101. Liu, W.; Xu, Y.; Xu, H.; Chen, J.; Hu, X.; Wu, J. Unraveling Babel: Exploring Multilingual Activation Patterns of LLMs and Their Applications. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing Proceedings of the 2024 conference on empirical methods in natural language processing ( 11855–11881); Miami, Florida, USAAssociation for Computational Linguistics, Al-Onaizan, Y., Bansal, M., Chen (), Y.-N., Eds.; 11 2024; Available online: https://aclanthology.org/2024.emnlp-main.662/. [CrossRef]
  102. Liu, X.; Song, Q.; Zhou, Q.; Du, H.; Xu, S.; Jiang, W.; Zhang, W.; Jia, X. Focusing on Language: Revealing and Exploiting Language Attention Heads in Multilingual Large Language Models. Focusing on language: Revealing and exploiting language attention heads in multilingual large language models. 2025. Available online: https://arxiv.org/abs/2511.07498.
  103. Liu, Y.; Shen, Y.; Zhu, H.; Xu, L.; Qian, Z.; Song, S.; Zhang, K.; Tang, J.; Zhang, P.; Yang, B.; Wang, R.; Hu, H. A Systematic Assessment of Language Models with Linguistic Minimal Pairs in Chinese. A systematic assessment of language models with linguistic minimal pairs in chinese. 2025. Available online: https://arxiv.org/abs/2411.06096.
  104. Maar, J.; Paperno, D.; McDougall, C. S.; Nanda, N. What’s the plan? Metrics for implicit planning in LLMs and their application to rhyme generation and question answering. What’s the plan? metrics for implicit planning in llms and their application to rhyme generation and question answering. 2026. Available online: https://arxiv.org/abs/2601.20164.
  105. Makelov, A.; Lange, G.; Geiger, A.; Nanda, N. Is This the Subspace You Are Looking for? An Interpretability Illusion for Subspace Activation Patching. The Twelfth International Conference on Learning Representations. The twelfth international conference on learning representations., 2024; Available online: https://openreview.net/forum?id=Ebt7JgMHv1.
  106. Manning, C. D.; Clark, K.; Hewitt, J.; Khandelwal, U.; Levy, O. Emergent linguistic structure in artificial neural networks trained by self-supervision. In Proceedings of the National Academy of Sciences, Available online. 2020; pp. 1174830046–30054. [Google Scholar] [CrossRef] [PubMed]
  107. Marks, S.; Rager, C.; Michaud, E. J.; Belinkov, Y.; Bau, D.; Mueller, A. Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language Models. The Thirteenth International Conference on Learning Representations. The thirteenth international conference on learning representations., 2025; Available online: https://openreview.net/forum?id=I4e82CIDxv.
  108. McGiff, J.; Tran, K.-T.; Mulcahy, W.; Luinín, D. Ó.; Dalzell, J.; Bhroin, R. N.; Burke, A.; O’Sullivan, B.; Nguyen, H. D.; Nikolov, N. S. Irish-BLiMP: A Linguistic Benchmark for Evaluating Human and Language Model Performance in a Low-Resource Setting. Irish-blimp: A linguistic benchmark for evaluating human and language model performance in a low-resource setting. 2025. Available online: https://arxiv.org/abs/2510.20957.
  109. Méloux, M.; Maniu, S.; Portet, F.; Peyrard, M. Everything, Everywhere, All at Once: Is Mechanistic Interpretability Identifiable? The Thirteenth International Conference on Learning Representations. The thirteenth international conference on learning representations. 2025. Available online: https://openreview.net/forum?id=5IWJBStfU7.
  110. Meng, K.; Bau, D.; Andonian, A.; Belinkov, Y. Locating and Editing Factual Associations in GPT. Locating and editing factual associations in gpt. 2023. Available online: https://arxiv.org/abs/2202.05262.
  111. Meng, K.; Sharma, A. S.; Andonian, A.; Belinkov, Y.; Bau, D. Mass-Editing Memory in a Transformer. Mass-editing memory in a transformer. 2023. Available online: https://arxiv.org/abs/2210.07229.
  112. Merullo, J.; Eickhoff, C.; Pavlick, E. Circuit Component Reuse Across Tasks in Transformer Language Models. The Twelfth International Conference on Learning Representations. The twelfth international conference on learning representations., 2024; Available online: https://openreview.net/forum?id=fpoAYV6Wsk.
  113. Mohebbi, H.; Jumelet, J.; Hanna, M.; Alishahi, A.; Zuidema, W. Transformer-specific Interpretability; M. Mesgar S. Loáiciga (), Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: Tutorial Abstracts Proceedings of the 18th conference of the european chapter of the association for computational linguistics: Tutorial abstracts ( 21–26). St. Julian’s, MaltaAssociation for Computational Linguistics, 03 2024; Available online: https://aclanthology.org/2024.eacl-tutorials.4/. [CrossRef]
  114. Mondal, S. K.; Sen, S.; Singhania, A.; Jyothi, P. Language-Specific Neurons Do Not Facilitate Cross-Lingual Transfer. In Shu (), The Sixth Workshop on Insights from Negative Results in NLP The sixth workshop on insights from negative results in nlp ( 46–62); Drozd, A., Sedoc, J., Tafreshi, S., A. Akula, R., Eds.; Albuquerque, New MexicoAssociation for Computational Linguistics, 05 2025; Available online: https://aclanthology.org/2025.insights-1.6/. [CrossRef]
  115. Mondorf, P.; Wang, M.; Gerstner, S.; Hakimi, A. D.; Liu, Y.; Veloso, L.; Zhou, S.; Schütze, H.; Plank, B. BlackboxNLP-2025 MIB Shared Task: Exploring Ensemble Strategies for Circuit Localization Methods. Blackboxnlp-2025 mib shared task: Exploring ensemble strategies for circuit localization methods. 2025. Available online: https://arxiv.org/abs/2510.06811.
  116. Mueller, A.; Brinkmann, J.; Li, M.; Marks, S.; Pal, K.; Prakash, N.; Rager, C.; Sankaranarayanan, A.; Sharma, A. S.; Sun, J.; Todd, E.; Bau, D.; Belinkov, Y. The Quest for the Right Mediator: Surveying Mechanistic Interpretability Through the Lens of Causal Mediation Analysis. The quest for the right mediator: Surveying mechanistic interpretability through the lens of causal mediation analysis. 2025. Available online: https://arxiv.org/abs/2408.01416.
  117. Mueller, A.; Geiger, A.; Wiegreffe, S.; Arad, D.; Arcuschin, I.; Belfki, A.; Chan, Y. S.; Fiotto-Kaufman, J.; Haklay, T.; Hanna, M.; Huang, J.; Gupta, R.; Nikankin, Y.; Orgad, H.; Prakash, N.; Reusch, A.; Sankaranarayanan, A.; Shao, S.; Stolfo, A.; Belinkov, Y. MIB: A Mechanistic Interpretability Benchmark. Mib: A mechanistic interpretability benchmark. 2025. Available online: https://arxiv.org/abs/2504.13151.
  118. Mueller, A.; Nicolai, G.; Petrou-Zeniou, P.; Talmina, N.; Linzen, T. Cross-Linguistic Syntactic Evaluation of Word Prediction Models. Cross-linguistic syntactic evaluation of word prediction models. 2020. Available online: https://arxiv.org/abs/2005.00187.
  119. Mueller, A.; Xia, Y.; Linzen, T. Causal Analysis of Syntactic Agreement Neurons in Multilingual Language Models. Causal analysis of syntactic agreement neurons in multilingual language models. 2022. Available online: https://arxiv.org/abs/2210.14328.
  120. Muller, B.; Elazar, Y.; Sagot, B.; Seddah, D. First Align, then Predict: Understanding the Cross-Lingual Ability of Multilingual BERT. P. Merlo, J. Tiedemann, R. Tsarfaty (), Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume Proceedings of the 16th conference of the european chapter of the association for computational linguistics: Main volume ( 2214–2231). OnlineAssociation for Computational Linguistics, 202104; Available online: https://aclanthology.org/2021.eacl-main.189/. [CrossRef]
  121. Nanda, N. Attribution Patching: Activation Patching at Industrial Scale. Neel Nanda’s Blog. 2023. Available online: https://www.neelnanda.io/mechanistic-interpretability/attribution-patching.
  122. Nanda, N.; Bloom, J. TransformerLens. Transformerlens. 2022. Available online: https://github.com/TransformerLensOrg/TransformerLens.
  123. Nie, E.; Schmid, H.; Schuetze, H. Mechanistic Understanding and Mitigation of Language Confusion in English-Centric Large Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2025 Findings of the association for computational linguistics: Emnlp 2025 ( 690–706). Suzhou, ChinaAssociation for Computational Linguistics.; Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng (), V., Eds.; 11 2025; Available online: https://aclanthology.org/2025.findings-emnlp.37/. [CrossRef]
  124. Nostalgebraist. Interpreting GPT: The logit lens. Blog Post. 2020. [Google Scholar] [CrossRef]
  125. Olah, C.; Cammarata, N.; Schubert, L.; Goh, G.; Petrov, M.; Carter, S. Distill53e00024–001; Zoom in: An introduction to circuits. 2020.
  126. Orhan, P.; Diego-Simón, P.; Chemla, E.; Lakretz, Y.; Boubenec, Y.; King, J.-R. Emergence of Phonemic, Syntactic, and Semantic Representations in Artificial Neural Networks. Emergence of phonemic, syntactic, and semantic representations in artificial neural networks. 2026. Available online: https://arxiv.org/abs/2601.18617.
  127. Ou, Y.; Yao, Y.; Zhang, N.; Jin, H.; Sun, J.; Deng, S.; Li, Z.; Chen, H. How Do LLMs Acquire New Knowledge? A Knowledge Circuits Perspective on Continual Pre-Training. How do llms acquire new knowledge? a knowledge circuits perspective on continual pre-training. 2025. Available online: https://arxiv.org/abs/2502.11196.
  128. Pal, K.; Sun, J.; Yuan, A.; Wallace, B.; Bau, D. Future Lens: Anticipating Subsequent Tokens from a Single Hidden State. In Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL) Proceedings of the 27th conference on computational natural language learning (conll), Available online. 2023; Association for Computational Linguistics; pp. 548–560. [Google Scholar] [CrossRef]
  129. Panigrahi, A.; Saunshi, N.; Zhao, H.; Arora, S. Task-Specific Skill Localization in Fine-tuned Language Models. Task-specific skill localization in fine-tuned language models. 2023. Available online: https://arxiv.org/abs/2302.06600.
  130. Pires, T.; Schlinger, E.; Garrette, D. How Multilingual is Multilingual BERT? A. Korhonen, D. Traum, L. Màrquez (), Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics Proceedings of the 57th annual meeting of the association for computational linguistics ( 4996–5001); Florence, ItalyAssociation for Computational Linguistics, 07 2019; Available online: https://aclanthology.org/P19-1493/. [CrossRef]
  131. Potertì, D.; Seveso, A.; Mercorio, F. Can Role Vectors Affect LLM Behaviour? C. Christodoulopoulos, T. Chakraborty, C. Rose, V. Peng (), Findings of the Association for Computational Linguistics: EMNLP 2025 Findings of the association for computational linguistics: Emnlp 2025 ( 17735–17747). Suzhou, ChinaAssociation for Computational Linguistics. 11 2025. Available online: https://aclanthology.org/2025.findings-emnlp.963/. [CrossRef]
  132. Rai, D.; Zhou, Y.; Feng, S.; Saparov, A.; Yao, Z. A Practical Review of Mechanistic Interpretability for Transformer-Based Language Models. A practical review of mechanistic interpretability for transformer-based language models. 2025. Available online: https://arxiv.org/abs/2407.02646.
  133. Riemenschneider, F.; Frank, A. Cross-Lingual Generalization and Compression: From Language-Specific to Shared Neurons. Cross-lingual generalization and compression: From language-specific to shared neurons. 2025. Available online: https://arxiv.org/abs/2506.01629.
  134. Rimsky, N.; Gabrieli, N.; Schulz, J.; Tong, M.; Hubinger, E.; Turner, A. Steering Llama 2 via Contrastive Activation Addition. L.-W. Ku, A. Martins, V. Srikumar (), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. In Long Papers) Proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers) ( 15504–15522), 202408; Bangkok, ThailandAssociation for Computational Linguistics; Volume 1. Available online: https://aclanthology.org/2024.acl-long.828/. [CrossRef]
  135. Rufail, A.; Rathore, S.; Son, D.; Simon, A.; Dave, S.; Zhang, D.; Blondin, C.; O’Brien, S.; Zhu, K. Semantic Convergence: Investigating Shared Representations Across Scaled LLMs. ACL 2025 Student Research Workshop. Acl 2025 student research workshop. 2025. Available online: https://openreview.net/forum?id=oOxrKNo1lQ.
  136. Ruzzetti, E. S.; Zanzotto, F. M.; Caselli, T. Lexical Popularity: Quantifying the Impact of Pre-training for LLM Performance. V. Demberg, K. Inui, L. Marquez (), Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics. In Long Papers) Proceedings of the 19th conference of the European chapter of the Association for Computational Linguistics (volume 1: Long papers) ( 1209–1230); Rabat, MoroccoAssociation for Computational Linguistics, 03 2026; Volume 1, Available online: https://aclanthology.org/2026.eacl-long.55/. [CrossRef]
  137. Sajjad, H.; Durrani, N.; Dalvi, F. Neuron-level Interpretation of Deep NLP Models: A Survey. Trans. Assoc. Comput. Linguist. 2021, 101285–1303. Available online: https://arxiv.org/abs/2108.13138.
  138. Schulz, L. Y.; Mitropolsky, D.; Poggio, T. Unraveling Syntax: How Language Models Learn Context-Free Grammars. Unraveling syntax: How language models learn context-free grammars. 2025. Available online: https://arxiv.org/abs/2510.02524.
  139. Schut, L.; Gal, Y.; Farquhar, S. Do Multilingual LLMs Think In English? ICLR 2025 Workshop on Building Trust in Language Models and Applications. Iclr 2025 workshop on building trust in language models and applications, 2025; Available online: https://openreview.net/forum?id=I8BOtOPcOv.
  140. Shafran, O.; Ronen, S.; Fahn, O.; Ravfogel, S.; Geiger, A.; Geva, M. From Directions to Regions: Decomposing Activations in Language Models via Local Geometry. From directions to regions: Decomposing activations in language models via local geometry. 2026. Available online: https://arxiv.org/abs/2602.02464.
  141. Sharkey, L.; Chughtai, B.; Batson, J.; Lindsey, J.; Wu, J.; Bushnaq, L.; Goldowsky-Dill, N.; Heimersheim, S.; Ortega, A.; Bloom, J.; Biderman, S.; Garriga-Alonso, A.; Conmy, A.; Nanda, N.; Rumbelow, J.; Wattenberg, M.; Schoots, N.; Miller, J.; Michaud, E. J.; McGrath, T. Open Problems in Mechanistic Interpretability. Open problems in mechanistic interpretability. 2025. Available online: https://arxiv.org/abs/2501.16496.
  142. Shu, B.; Singh, A.; ElSherief, M. From Syntax to Emotion: A Mechanistic Analysis of Emotion Inference in LLMs. From syntax to emotion: A mechanistic analysis of emotion inference in llms. 2026. Available online: https://arxiv.org/abs/2604.25866.
  143. Shu, D.; Wu, X.; Zhao, H.; Rai, D.; Yao, Z.; Liu, N.; Du, M. A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models. A survey on sparse autoencoders: Interpreting the internal mechanisms of large language models. 2025. Available online: https://arxiv.org/abs/2503.05613.
  144. Singh, A. K.; Chan, S. C. Y.; Moskovitz, T.; Grant, E.; Saxe, A. M.; Hill, F. The Transient Nature of Emergent In-Context Learning in Transformers. The transient nature of emergent in-context learning in transformers. 2023. Available online: https://arxiv.org/abs/2311.08360.
  145. Someya, T.; Oseki, Y. JBLiMP: Japanese Benchmark of Linguistic Minimal Pairs. A. Vlachos I. Augenstein (), Findings of the Association for Computational Linguistics: EACL 2023 Findings of the association for computational linguistics: Eacl 2023 ( 1581–1594). Dubrovnik, CroatiaAssociation for Computational Linguistics. 05 2023. Available online: https://aclanthology.org/2023.findings-eacl.117/. [CrossRef]
  146. Son, D.; Rathore, S.; Rufail, A.; Simon, A.; Zhang, D.; Dave, S.; Blondin, C.; Zhu, K.; O’Brien, S. Semantic Convergence: Investigating Shared Representations Across Scaled LLMs. Semantic convergence: Investigating shared representations across scaled llms. 2025. Available online: https://arxiv.org/abs/2507.22918.
  147. Song, Y.; Krishna, K.; Bhatt, R.; Iyyer, M. SLING: Sino Linguistic Evaluation of Large Language Models. Y. Goldberg, Z. Kozareva, Y. Zhang (), Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing Proceedings of the 2022 conference on empirical methods in natural language processing ( 4606–4634); Abu Dhabi, United Arab EmiratesAssociation for Computational Linguistics, 12 2022; Available online: https://aclanthology.org/2022.emnlp-main.305/. [CrossRef]
  148. Stańczak, K.; Ponti, E.; Hennigen, L. T.; Cotterell, R.; Augenstein, I. Same Neurons, Different Languages: Probing Morphosyntax in Multilingual Pre-trained Models. Same neurons, different languages: Probing morphosyntax in multilingual pre-trained models. 2022. Available online: https://arxiv.org/abs/2205.02023.
  149. Stickland, A. C.; Lyzhov, A.; Pfau, J.; Mahdi, S.; Bowman, S. R. Steering Without Side Effects: Improving Post-Deployment Control of Language Models. Steering without side effects: Improving post-deployment control of language models. 2024. Available online: https://arxiv.org/abs/2406.15518.
  150. Suijkerbuijk, M.; Prins, Z.; Kloots, M. d. H.; Zuidema, W.; Frank, S. L. BLiMP-NL: A Corpus of Dutch Minimal Pairs and Acceptability Judgments for Language Model Evaluation. Comput. Linguist. Available online. 2025, 1–35. [Google Scholar] [CrossRef]
  151. Syed, A.; Rager, C.; Conmy, A. Attribution Patching Outperforms Automated Circuit Discovery; Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP Proceedings of the 7th blackboxnlp workshop: Analyzing and interpreting neural networks for nlp ( 407–416). Miami, Florida, Belinkov, Y., Kim, N., Jumelet, J., Mohebbi, H., Mueller, A., Chen (), H., Eds.; USAssociation for Computational Linguistics, 11 2024; Available online: https://aclanthology.org/2024.blackboxnlp-1.25/. [CrossRef]
  152. Taktasheva, E.; Bazhukov, M.; Koncha, K.; Fenogenova, A.; Artemova, E.; Mikhailov, V. RuBLiMP: Russian Benchmark of Linguistic Minimal Pairs. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing Proceedings of the 2024 conference on empirical methods in natural language processing ( 9268–9299); Miami, Florida, USAAssociation for Computational Linguistics, Al-Onaizan, Y., Bansal, M., Chen (), Y.-N., Eds.; 11 2024; Available online: https://aclanthology.org/2024.emnlp-main.522/. [CrossRef]
  153. Tan, S.; Wu, D.; Monz, C. Neuron Specialization: Leveraging Intrinsic Task Modularity for Multilingual Machine Translation. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing Proceedings of the 2024 conference on empirical methods in natural language processing ( 6506–6527); Miami, Florida, USAAssociation for Computational Linguistics, Al-Onaizan, Y., Bansal, M., Chen (), Y.-N., Eds.; 11 2024; Available online: https://aclanthology.org/2024.emnlp-main.374/. [CrossRef]
  154. Tang, T.; Luo, W.; Huang, H.; Zhang, D.; Wang, X.; Zhao, X.; Wei, F.; Wen, J.-R. Language-Specific Neurons: The Key to Multilingual Capabilities in Large Language Models. L.-W. Ku, A. Martins, V. Srikumar (), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. In Long Papers) Proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers) ( 5701–5715); Bangkok, ThailandAssociation for Computational Linguistics, 08 2024; Volume 1, Available online: https://aclanthology.org/2024.acl-long.309/. [CrossRef]
  155. Tang, Y.; Saini, H.; Yao, Z.; Lin, Z.; Liao, Y.; Cui, J.; Wang, Y.; Du, M.; Liu, D. A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious Minima. A unified theory of sparse dictionary learning in mechanistic interpretability: Piecewise biconvexity and spurious minima. 2026. Available online: https://arxiv.org/abs/2512.05534.
  156. Tezuka, H.; Inoue, N. The Transfer Neurons Hypothesis: An Underlying Mechanism for Language Latent Space Transitions in Multilingual LLMs. The transfer neurons hypothesis: An underlying mechanism for language latent space transitions in multilingual llms. 2025. Available online: https://arxiv.org/abs/2509.17030.
  157. Todd, E.; Li, M.; Sharma, A. S.; Mueller, A.; Wallace, B. C.; Bau, D. Function Vectors in Large Language Models. The Twelfth International Conference on Learning Representations. The twelfth international conference on learning representations., 2024; Available online: https://openreview.net/forum?id=AwyxtyMwaG.
  158. Tufanov, I.; Hambardzumyan, K.; Ferrando, J.; Voita, E. LM Transparency Tool: Interactive Tool for Analyzing Transformer Language Models. Lm transparency tool: Interactive tool for analyzing transformer language models. 2024. Available online: https://arxiv.org/abs/2404.07004.
  159. Turner, A. M.; Thiergart, L.; Leech, G.; Udell, D.; Vazquez, J. J.; Mini, U.; MacDiarmid, M. Steering Language Models With Activation Engineering. Steering language models with activation engineering. 2024. Available online: https://arxiv.org/abs/2308.10248.
  160. Waldis, A.; Perlitz, Y.; Choshen, L.; Hou, Y.; Gurevych, I. Holmes: A Benchmark to Assess the Linguistic Competence of Language Models. Transactions of the Association for Computational Linguistics121616–1647. 2024. Available online: https://aclanthology.org/2024.tacl-1.88/. [CrossRef]
  161. Wang, K.; Variengien, A.; Conmy, A.; Shlegeris, B.; Steinhardt, J. Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small. Interpretability in the wild: a circuit for indirect object identification in gpt-2 small. 2022. Available online: https://arxiv.org/abs/2211.00593.
  162. Wang, M.; Adel, H.; Lange, L.; Liu, Y.; Nie, E.; Strötgen, J.; Schuetze, H. Lost in Multilinguality: Dissecting Cross-lingual Factual Inconsistency in Transformer Language Models. W. Che, J. Nabende, E. Shutova, M. T. Pilehvar (), Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics. In Long Papers) Proceedings of the 63rd annual meeting of the association for computational linguistics (volume 1: Long papers) ( 5075–5094); Vienna, AustriaAssociation for Computational Linguistics, 07 2025; Volume 1, Available online: https://aclanthology.org/2025.acl-long.253/. [CrossRef]
  163. Wang, M.; Xu, Z.; Mao, S.; Deng, S.; Tu, Z.; Chen, H.; Zhang, N. Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms. Beyond prompt engineering: Robust behavior control in llms via steering target atoms. 2025. Available online: https://arxiv.org/abs/2505.20322.
  164. Warstadt, A.; Parrish, A.; Liu, H.; Mohananey, A.; Peng, W.; Wang, S.-F.; Bowman, S. R. BLiMP: The Benchmark of Linguistic Minimal Pairs for English. Transactions of the Association for Computational Linguistics. 2020, pp. 8377–392. Available online: https://aclanthology.org/2020.tacl-1.25/. [CrossRef]
  165. Wei, J.; Garrette, D.; Linzen, T.; Pavlick, E. Frequency Effects on Syntactic Rule Learning in Transformers; Online and Punta Cana, Dominican RepublicAssociation for Computational Linguistics, Moens, M.-F., Huang, X., Specia, L., Yih (), S. W.-t., Eds.; Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing Proceedings of the 2021 conference on empirical methods in natural language processing ( 932–948), 11 2021; Available online: https://aclanthology.org/2021.emnlp-main.72/. [CrossRef]
  166. Wendler, C.; Veselovsky, V.; Monea, G.; West, R. Do Llamas Work in English? On the Latent Language of Multilingual Transformers. L.-W. Ku, A. Martins, V. Srikumar (), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. In Long Papers) Proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers) ( 15366–15394); Bangkok, ThailandAssociation for Computational Linguistics, 08 2024; Volume 1, Available online: https://aclanthology.org/2024.acl-long.820/. [CrossRef]
  167. Wilson, M.; Petty, J.; Frank, R. How Abstract Is Linguistic Generalization in Large Language Models? Experiments with Argument Structure. Trans. Assoc. Comput. Linguist. 2023, 111377–1395. Available online: https://doi.org/10.1162/tacl_a_00608. [CrossRef]
  168. Wu, S.; Dredze, M. Beto, Bentz, Becas: The Surprising Cross-Lingual Effectiveness of BERT. Beto, bentz, becas: The surprising cross-lingual effectiveness of bert. 2019. Available online: https://arxiv.org/abs/1904.09077.
  169. Wu, Z.; Arora, A.; Geiger, A.; Wang, Z.; Huang, J.; Jurafsky, D.; Manning, C. D.; Potts, C. AxBench: Steering LLMs? Even Simple Baselines Outperform Sparse Autoencoders. Axbench: Steering llms? even simple baselines outperform sparse autoencoders. 2025. Available online: https://arxiv.org/abs/2501.17148.
  170. Wu, Z.; Arora, A.; Wang, Z.; Geiger, A.; Jurafsky, D.; Manning, C. D.; Potts, C. ReFT: Representation Finetuning for Language Models. The Thirty-eighth Annual Conference on Neural Information Processing Systems. The thirty-eighth annual conference on neural information processing systems., 2024; Available online: https://openreview.net/forum?id=fykjplMc0V.
  171. Wu, Z.; Yu, X. V.; Yogatama, D.; Lu, J.; Kim, Y. The Semantic Hub Hypothesis: Language Models Share Semantic Representations Across Languages and Modalities. The semantic hub hypothesis: Language models share semantic representations across languages and modalities. 2025. Available online: https://arxiv.org/abs/2411.04986.
  172. Xiang, B.; Yang, C.; Li, Y.; Warstadt, A.; Kann, K. CLiMP: A Benchmark for Chinese Language Model Evaluation. P. Merlo, J. Tiedemann, R. Tsarfaty (), Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume Proceedings of the 16th conference of the european chapter of the association for computational linguistics: Main volume ( 2784–2790). OnlineAssociation for Computational Linguistics, 202104; Available online: https://aclanthology.org/2021.eacl-main.242/. [CrossRef]
  173. Xie, W.; Feng, Y.; Gu, S.; Yu, D. Importance-based Neuron Allocation for Multilingual Neural Machine Translation. C. Zong, F. Xia, W. Li, R. Navigli (), Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing. In Long Papers) Proceedings of the 59th annual meeting of the association for computational linguistics and the 11th international joint conference on natural language processing (volume 1: Long papers) ( 5725–5737); OnlineAssociation for Computational Linguistics, 08 2021; Volume 1, Available online: https://aclanthology.org/2021.acl-long.445/. [CrossRef]
  174. Yu, Z.; Ananiadou, S. Neuron-Level Knowledge Attribution in Large Language Models. Y. Al-Onaizan, M. Bansal, Y.-N. Chen (), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing Proceedings of the 2024 conference on empirical methods in natural language processing ( 3267–3280). Miami, Florida, 202411; USAAssociation for Computational Linguistics. Available online: https://aclanthology.org/2024.emnlp-main.191/. [CrossRef]
  175. Zeng, H.; Han, S.; Chen, L.; Yu, K. Converging to a Lingua Franca: Evolution of Linguistic Regions and Semantics Alignment in Multilingual Large Language Models; Abu Dhabi, UAEAssociation for Computational Linguistics, Rambow, O., Wanner, L., Apidianaki, M., Al-Khalifa, H., Eugenio, B. D., Schockaert (), S., Eds.; Proceedings of the 31st International Conference on Computational Linguistics Proceedings of the 31st international conference on computational linguistics ( 10602–10617), 01 2025; Available online: https://aclanthology.org/2025.coling-main.707/.
  176. Zhang, C.; Lu, J.; Tran, V. Q.; Schuster, T.; Metzler, D.; Lin, J. Tomato, Tomahto, Tomate: Do Multilingual Language Models Understand Based on Subword-Level Semantic Concepts? L. Chiruzzo, A. Ritter, L. Wang (), Findings of the Association for Computational Linguistics: NAACL 2025 Findings of the association for computational linguistics: Naacl 2025 ( 1821–1837). Albuquerque, New MexicoAssociation for Computational Linguistics. 04 2025. Available online: https://aclanthology.org/2025.findings-naacl.98/. [CrossRef]
  177. Zhang, F.; Nanda, N. Towards Best Practices of Activation Patching in Language Models: Metrics and Methods. The Twelfth International Conference on Learning Representations. The twelfth international conference on learning representations., 2024; Available online: https://openreview.net/forum?id=Hf17y6u9BC.
  178. Zhang, H.; Shang, C.; Wang, S.; Zhang, D.; Yu, Y.; Yao, F.; Sun, R.; Yang, Y.; Wei, F. ShifCon: Enhancing Non-Dominant Language Capabilities with a Shift-based Multilingual Contrastive Framework; Che, W., Nabende, J., Shutova, E., Pilehvar (), M. T., Eds.; Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 07 2025; Volume 1, Available online: https://aclanthology.org/2025.acl-long.239/. [CrossRef]
  179. Zhang, H.; Zhang, Z.; Wang, M.; Su, Z.; Wang, Y.; Wang, Q.; Yuan, S.; Nie, E.; Duan, X.; Han, F.; Xue, Q.; Yu, Z.; Shang, C.; Liang, X.; Xiong, J.; Shen, H.; Tao, C.; Liu, Z.; Jin, S.; Wong, N. Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models. Locate, steer, and improve: A practical survey of actionable mechanistic interpretability in large language models. 2026. Available online: https://arxiv.org/abs/2601.14004.
  180. Zhang, R.; Yu, Q.; Zang, M.; Eickhoff, C.; Pavlick, E. The Same but Different: Structural Similarities and Differences in Multilingual Language Modeling. The Thirteenth International Conference on Learning Representations. The thirteenth international conference on learning representations., 2025; Available online: https://openreview.net/forum?id=NCrFA7dq8T.
  181. Zhang, X.; Liang, Y.; Meng, F.; Zhang, S.; Chen, Y.; Xu, J.; Zhou, J. Multilingual Knowledge Editing with Language-Agnostic Factual Neurons; Abu Dhabi, UAEAssociation for Computational Linguistics, Rambow, O., Wanner, L., Apidianaki, M., Al-Khalifa, H., Eugenio, B. D., Schockaert (), S., Eds.; Proceedings of the 31st International Conference on Computational Linguistics Proceedings of the 31st international conference on computational linguistics ( 5775–5788), 01 2025; Available online: https://aclanthology.org/2025.coling-main.385/.
  182. Zhang, Z.; Zhao, J.; Zhang, Q.; Gui, T.; Huang, X. Unveiling Linguistic Regions in Large Language Models. L.-W. Ku, A. Martins, V. Srikumar (), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. In Long Papers) Proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers) ( 6228–6247), 202408; Bangkok, ThailandAssociation for Computational Linguistics; Volume 1. Available online: https://aclanthology.org/2024.acl-long.338/. [CrossRef]
  183. Zhao, N.; Duan, X.; Cai, Z. G. The Missing Half of Language Learning in Current Developmental Language Models: Exogenous and Endogenous Linguistic Input. Open Mind91543-1549 Available online. 2025. [Google Scholar] [CrossRef] [PubMed]
  184. Zhao, Y.; Zhang, W.; Chen, G.; Kawaguchi, K.; Bing, L. How do Large Language Models Handle Multilingualism? The Thirty-eighth Annual Conference on Neural Information Processing Systems. The thirty-eighth annual conference on neural information processing systems, 2024; Available online: https://openreview.net/forum?id=ctXYOoAgRy.
  185. Zhong, C.; Cheng, F.; Liu, Q.; Jiang, J.; Wan, Z.; Chu, C.; Murawaki, Y.; Kurohashi, S. Beyond English-Centric LLMs: What Language Do Multilingual Language Models Think in? Beyond english-centric llms: What language do multilingual language models think in? 2024. Available online: https://arxiv.org/abs/2408.10811.
  186. Zhong, C.; Cheng, F.; Liu, Q.; Murawaki, Y.; Chu, C.; Kurohashi, S. Language Lives in Sparse Dimensions: Toward Interpretable and Efficient Multilingual Control for Large Language Models. Language lives in sparse dimensions: Toward interpretable and efficient multilingual control for large language models. 2025. Available online: https://arxiv.org/abs/2510.07213.
  187. Zhou, E.; Salhan, S. Extended Abstract for “Linguistic Universals”: Emergent Shared Features in Independent Monolingual Language Models via Sparse Autoencoders. D. I. Adelani et al. (), Proceedings of the 5th Workshop on Multilingual Representation Learning (MRL 2025) Proceedings of the 5th workshop on multilingual representation learning (mrl 2025) ( 128–130). Suzhuo, ChinaAssociation for Computational Linguistics, 11 2025. Available online: https://aclanthology.org/2025.mrl-main.9/. [CrossRef]
  188. Zhou, X.; Chen, D.; Cahyawijaya, S.; Duan, X.; Cai, Z. Linguistic Minimal Pairs Elicit Linguistic Similarity in Large Language Models. In Proceedings of the 31st International Conference on Computational Linguistics Proceedings of the 31st international conference on computational linguistics ( 6866–6888); Abu Dhabi, UAEAssociation for Computational Linguistics, Rambow, O., Wanner, L., Apidianaki, M., Al-Khalifa, H., Eugenio, B. D., Schockaert (), S., Eds.; 01 2025; Available online: https://aclanthology.org/2025.coling-main.459/.
  189. Zhu, S.; Pan, L.; Li, B.; Xiong, D. LANDeRMT: Dectecting and Routing Language-Aware Neurons for Selectively Finetuning LLMs to Machine Translation. L.-W. Ku, A. Martins, V. Srikumar (), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics. In Long Papers) Proceedings of the 62nd annual meeting of the association for computational linguistics (volume 1: Long papers) ( 12135–12148); Bangkok, ThailandAssociation for Computational Linguistics, 08 2024; Volume 1, Available online: https://aclanthology.org/2024.acl-long.656/. [CrossRef]
  190. Üveges, I.; Ring, O. Evaluating the Impact of Synthetic Data on Emotion Classification: A Linguistic and Structural Analysis. Information164. 2025. Available online: https://www.mdpi.com/2078-2489/16/4/330. [CrossRef]
Table 1. Representative surveys in MI and their focus.
Table 1. Representative surveys in MI and their focus.
Paper Focus
[132] MI methods for Transformer.
[141] Roadmap and future directions.
[143] Feature disentanglement.
[116] Causal Mediation Analysis.
[85] Training dynamics.
[48] Intrinsic interpretability.
[53] Syntactic Knowledge.
[179] Locate, Steer, and Improve.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings