Submitted:
19 May 2026
Posted:
20 May 2026
You are already at the latest version
Abstract
Recent work on neural scaling demonstrates consistent performance gains with increased data and model capacity, yet these improvements are typically assessed using surface-level metrics that do not capture factual reliability. In multi-document summarization (MDS), this limitation is particularly acute, as scaling has been shown to amplify hallucination and content distortion. In this paper, we investigate the empirical scaling behaviour of faithfulness-aware transformers under tightly controlled conditions, using LSHT as a fixed architectural and training baseline. Rather than proposing new scaling laws, we analyze how summarization quality, faithfulness and efficiency evolve as dataset size and model capacity are independently increased, while holding architecture, optimization, decoding and hardware constant. All experiments are conducted exclusively on the Multi-News benchmark to avoid cross-dataset confounds. Across ROUGE, coverage, repetition and faithfulness-oriented metrics, we show that lexical overlap and factual consistency follow distinct scaling dynamics. Faithfulness improves most rapidly during early data scaling (approximately 3–4% relative gain from 3k to 12k samples) but exhibits diminishing marginal returns at larger scales, whereas ROUGE continues to increase more smoothly. We further show that faithfulness is more sensitive to data diversity than to volume alone and identify practical scaling regimes that maximize faithfulness gains relative to computational cost. These results establish empirical expectations for scaling faithfulness-aware MDS systems and provide actionable guidance for reliable summarization under realistic resource constraints.
Keywords:
multi-document summarization
; faithfulness
; hallucination
; transformers
; scaling behaviour
; empirical analysis
1. Introduction
Scaling neural language models has emerged as a dominant paradigm in modern NLP [1], with extensive empirical evidence showing that increasing data scale and model capacity leads to predictable improvements in perplexity [2] and downstream task performance. These scaling laws have directly influenced the design and training of increasingly large language models [3], which have demonstrated strong general-purpose capabilities across a wide range of language tasks [4]. Recent advances continue to extend the limits of model scale and efficiency [5]. However, the majority of scaling Analyzes rely on evaluation metrics that primarily emphasize fluency [6] or lexical overlap [7], offering limited insight into factual correctness and content reliability [8]. Faithfulness in summarization, which concerns the factual consistency between generated summaries and source documents, requires careful and task-specific evaluation methodologies [9]. Consequently, the interaction between model scaling and faithfulness remains insufficiently characterized in existing literature [10], with faithfulness evaluation itself posing persistent methodological challenges [11]. This limitation is particularly critical in multi-document summarization (MDS) [12], where models must aggregate, compress, and reconcile information across multiple sources. Prior work has explored hierarchical modeling strategies for MDS [13], yet the task inherently amplifies risks of hallucination [14], redundancy, and uneven coverage [15]. Empirical evidence suggests that even strong transformer-based architectures [16] can generate fluent but factually unsupported summaries as model capacity increases [17]. Widely adopted architectures such as BART [18] and T5 [19] exemplify this tension between surface-level quality and factual reliability.
In response to these limitations, prior work has proposed faithfulness-aware training objectives, including coverage regularization [20], repetition control [21], and length constraints [22], which introduce inductive biases aimed at mitigating known summarization failure modes. However, how such mechanisms interact with data and model scaling remains largely unexplored. Understanding these interactions is essential for designing summarization systems that remain reliable as scale increases. This positions faithfulness-aware scaling as an under-studied yet practically consequential problem, distinct from traditional performance-centric scaling analyzes. In parallel, parameter-efficient fine-tuning methods have gained prominence as scalable alternatives to full model training [23], with continued advances in efficient adaptation strategies [24]. Techniques such as P-tuning [25], prefix tuning [26], power-efficient training [27], and QLoRA [28] further enable controlled experimentation under constrained resources. Complementary advances in evaluation, including QuestEval [29] and AlignScore [30], provide tools for measuring factual consistency beyond surface overlap. At the architectural level, sparse and efficient attention mechanisms [31], along with long-context transformer variants such as Longformer [32], Reformer [33], and Linformer [34], aim to reduce computational complexity while supporting longer inputs. Together, these developments underscore the need for controlled empirical studies that disentangle the effects of scaling, efficiency, and faithfulness-oriented objectives in multi-document summarization.
1.1. Why LSHT?
To study scaling behaviour rigorously, it is essential to control architectural and optimization factors. We therefore adopt LSHT (Lightweight Self-Healing Transformer), a low-resource, lightweight encoder–decoder transformer augmented with explicit faithfulness-oriented training signals, as a fixed experimental probe. LSHT combines standard transformer components with repetition, coverage, and length-control objectives inspired by prior work in summarization and neural machine translation [35,36]. LSHT is particularly well suited for scaling analysis for three reasons. First, its parameter efficiency enables experiments across multiple model sizes under fixed hardware constraints, a property that has become essential for empirical scaling studies. Second, its regularized training objective explicitly targets known failure modes of multi-document summarization, enabling direct measurement of faithfulness trends. Third, its architectural simplicity reduces confounding interactions that often arise in large, highly engineered models. As a result, efficient transformer architectures such as LSHT enable controlled scaling experiments that would be impractical with substantially larger models. Importantly, we do not present LSHT as a new state-of-the-art system in this paper. Instead, it functions as a scientific instrument which is a controlled and interpretable baseline through which scaling behaviour can be examined. This framing shifts LSHT from a “competitive model” to an analytical tool, thereby strengthening the validity of the scaling conclusions. Related work such as ELECTRA provides insights into efficient pre-training [37], while BERT established foundational transformer architectures [38], and RoBERTa refined BERT’s training procedures [39].
1.2. Scope, Claim Bounding, and Contributions
We explicitly delimit the scope of this study as follows: we analyse empirical scaling behaviour under fixed architecture, optimization, decoding, and dataset conditions, without claiming universal or asymptotic scaling laws. Accordingly, our analysis is empirical rather than theoretical, task-specific with a focus on multi-document summarization, and non-asymptotic, restricted to practically trainable regimes. We do not attempt to derive or validate general scaling laws, nor do we claim that observed trends extrapolate beyond the studied conditions. This explicit claim bounding is intended to prevent over-generalisation, align reviewer expectations, and ensure that conclusions are interpreted appropriately across venues.
Within this bounded scope, the paper makes the following contributions:
- 1.
- Controlled Scaling Study. We present a systematic empirical analysis of how summarization quality and faithfulness metrics evolve with increasing dataset size and model capacity under fixed architecture, optimization, decoding, and hardware conditions.
- 2.
- Faithfulness and ROUGE Scaling Behaviour. We characterize how faithfulness-related metrics and ROUGE respond differently to scale, showing that surface-level overlap improvements do not necessarily translate into gains in factual consistency in multi-document summarization.
- 3.
- Dataset vs. Model Scaling Insights. By independently varying data volume and model size on a single benchmark, we isolate their respective effects on summarization quality, hallucination behaviour, and faithfulness dynamics.
- 4.
- Practical Scaling Guidance. We identify empirically grounded scaling regimes that offer favourable trade-offs between faithfulness gains and computational cost, providing actionable guidance for practitioners operating under realistic resource constraints.
- 5.
- Empirical Foundation for Future Work. Our findings establish a reference point for subsequent studies on theoretical scaling Analyzes, architectural extensions, and adaptive or self-learning mechanisms built upon faithfulness-aware transformer models.
2. Related Work
This work intersects several active research areas, including neural scaling behaviour, multi-document summarization, faithfulness and hallucination control, and lightweight transformer architectures. We review each line of work selectively, with emphasis on empirical findings most relevant to our experimental scope.
2.1. Scaling Behaviour in Neural Language Models
Large-scale empirical studies have shown that language model performance often improves in a predictable manner with increases in data, model size, and computational budget. Early work identified approximate power-law relationships between training loss and scale under controlled settings [40], largely in the context of auto regressive language modeling. Subsequent studies emphasized the importance of data–model balance [41], demonstrating that compute optimal regimes require jointly scaling parameters and training data. These principles have since been examined across multiple domains, including machine translation [42] and code generation, and have been further reinforced by the emergence of large language models. Scaling behavior has been analyzed under diverse conditions [43], with systematic investigations of data scaling effects [44], cross-lingual settings [45], and controlled model suites such as Pythia [46]. Recent large-scale systems, including GPT-4 [47] and PaLM-2 [48], represent notable milestones in this line of work. However, the majority of these studies focus on perplexity or next-token prediction, rather than structured generation tasks such as summarization. Moreover, they typically assume homogeneous training objectives and do not explicitly consider auxiliary regularization signals aimed at improving faithfulness or content selection. Consequently, it remains unclear whether analogous scaling dynamics persist when models are trained with objectives designed to mitigate hallucination or repetition.
At the same time, several works have cautioned against uncritical extrapolation of scaling laws beyond their original scope, particularly for tasks involving factual grounding or long-context reasoning. Empirical Analyzes suggest that scaling behavior can vary substantially across tasks, with some regimes exhibiting diminishing returns or even degradation at extreme scales. Interpretability studies further highlight the difficulty of fully characterizing scaling effects [49], while others argue that simple scaling formulations fail to capture important complexities [50]. Prior work has also documented heterogeneous scaling behavior across transformer architectures [51], explored contrastive objectives in scaling contexts [52], and refined scaling laws through more detailed empirical analysis [53]. In this work, we therefore treat scaling laws as descriptive templates rather than predictive guarantees, and restrict our analysis to empirically observable behavior under fixed and controlled conditions.
2.2. Multi-Document Summarization
Large-scale empirical studies have shown that language model performance often improves predictably with increases in data, model size, and compute. Early work established approximate power-law relationships between loss and scale under controlled settings [40], primarily for autoregressive language modeling. Subsequent refinements emphasized the importance of jointly scaling data and parameters, demonstrating that compute-optimal training depends on maintaining an appropriate data–model balance [41]. These principles have since been extended to additional domains, including machine translation [42] and code generation. The development of large language models has further reinforced the relevance of scaling Analyzes, with studies examining scaling behavior across settings [43], systematic data scaling effects [44], and challenges in cross-lingual regimes [45]. Model families such as Pythia [46], as well as large systems including GPT-4 [47] and PaLM-2 [48], represent major milestones in this line of work. However, these studies predominantly focus on perplexity and next-token prediction, rather than structured generation tasks such as summarization. Moreover, they typically assume homogeneous training objectives and do not consider auxiliary signals explicitly designed to improve faithfulness or content selection. Consequently, it remains unclear whether analogous scaling dynamics hold when models are trained with objectives aimed at mitigating hallucination or repetition.
Several works have cautioned against uncritical extrapolation of scaling laws beyond their original scope, particularly for tasks requiring factual grounding or long-context reasoning. Empirical Analyzes indicate that scaling behavior can vary substantially across tasks, with some regimes exhibiting diminishing returns or even degradation at extreme scales. Interpretability studies further underscore the difficulty of fully characterizing scaling effects [49], while others argue that simple scaling formulations fail to capture important complexities [50]. Prior work has also documented heterogeneous scaling behavior across transformer architectures [51], explored contrastive objectives in scaling contexts [52], and refined scaling laws through more detailed empirical analysis [53]. Accordingly, we treat scaling formulations as descriptive templates rather than predictive laws, and restrict our analysis to empirically observable behavior under fixed and controlled conditions.
2.3. Faithfulness and Hallucination Control in Summarization
Faithfulness has emerged as a central concern in abstractive summarization research. Analyzes have demonstrated that high ROUGE scores do not guarantee factual correctness, motivating the development of metrics and training strategies that explicitly address hallucination. Factual consistency evaluation has been systematically studied [54]. Comprehensive surveys have documented the prevalence and types of hallucinations in neural text generation. SummEval provides comprehensive evaluation [55]. Evaluation frameworks have been developed to assess factual consistency directly from source documents [56], complementing reference-based metrics like ROUGE and BERTScore. Several approaches have introduced auxiliary objectives to encourage faithful generation. Coverage mechanisms were originally proposed in neural machine translation to prevent under or over-attending to source content, and later adapted to summarization to improve content grounding. Repetition penalties and length control mechanisms have similarly been used to reduce redundancy and degenerate outputs. More recent work has explored entailment-based evaluation, question-answering probes, and factual consistency classifiers.
Despite these advances, most prior studies evaluate faithfulness at a fixed scale, leaving open the question of how faithfulness-aware objectives interact with increased data or model capacity. Recent work has begun to explore faithfulness trends across model sizes, but systematic Analyzes of scaling behavior remain limited. Our work addresses this gap by analyzing faithfulness trends explicitly as scaling variables change, rather than treating faithfulness as a static property of a model.
2.4. Lightweight and Efficient Transformer Models
Alongside the scaling of large models, there has been sustained interest in lightweight and computationally efficient transformer architectures. Sparse attention mechanisms reduce the quadratic complexity of standard self-attention to linear or near-linear scaling with sequence length. The Synthesizer rethinks self-attention by replacing learned attention weights with synthetic patterns [57], while a broader family of efficient transformer designs continues to emerge [58]. The Reformer leverages locality-sensitive hashing to approximate attention, whereas Linformer employs low-rank projections to achieve linear complexity. In the vision domain, the Swin Transformer demonstrates efficient hierarchical modeling [59], and vision transformers such as ViT highlight the versatility of transformer architectures across modalities [60]. Rotary positional embeddings have also gained prominence as an efficient alternative to absolute positional encodings, enabling improved handling of long contexts. Comprehensive surveys of efficient transformers document a wide range of techniques aimed at reducing computational and memory overhead. These models are particularly well suited to summarization, where long input sequences and constrained deployment resources are common. Encoder–decoder architectures with moderate parameter counts have been shown to achieve competitive summarization performance when combined with appropriate regularization and training objectives, while parameter-efficient fine-tuning further reduces adaptation costs. Beyond efficiency, lightweight models enable tighter experimental control, making them especially suitable for empirical studies of scaling behaviour under fixed hardware constraints. In this work, we exploit these advantages by adopting a lightweight, faithfulness-aware transformer as a stable baseline, enabling controlled scaling experiments that would be impractical with substantially larger architectures.
Positioning Summary
Prior work has extensively examined scaling behaviour in language modeling and, separately, the problem of faithfulness in summarization. However, the intersection of these two research directions remains under-explored. By focusing on empirical scaling behaviour within a faithfulness-aware multi-document summarization setting, this paper complements existing scaling studies while addressing a practical and under-studied gap in summarization research.
3. Problem Setup and Experimental Scope
This section formalizes the task setting, defines the scaling axes under study, and explicitly delineates the experimental scope. The goal is to ensure that observed trends can be attributed to controlled changes in scale rather than confounding architectural, optimization, or procedural variations.
We study abstractive multi-document summarization (MDS) [12], where the input consists of a set of related documents describing a common topic or event, and the output is a concise summary that synthesizes salient information across sources. Formally, given a document cluster
the model generates a summary
which is expected to be informative, non-redundant, and faithful to the content of [8]. Faithfulness in MDS requires that all information in the summary be supported by the source documents [14].
All experiments are conducted exclusively on the Multi-News benchmark, a widely used dataset for multi-document summarization containing aligned clusters of news articles and human-written summaries. This benchmark has been extensively used in prior work on multi-document summarization and provides a controlled setting for evaluating scaling behavior. GEM provides comprehensive evaluation frameworks [61]. BEIR offers benchmark evaluation resources [62]. We deliberately restrict our analysis to a single dataset to minimize variability arising from differences in annotation style, document structure, or domain-specific language. This design choice enables a cleaner investigation of scaling behaviour by ensuring that performance changes can be attributed primarily to dataset size or model capacity, rather than cross-dataset heterogeneity.
We analyze empirical scaling behaviour along two primary axes, following established scaling law methodologies [40]. Dataset scaling is performed by varying the number of training examples while keeping the underlying data distribution fixed [41]. Specifically, subsets of the Multi-News training set are constructed via random sub-sampling at increasing sizes (3k, 12k, and 45k samples), allowing us to examine how surface-level performance and faithfulness metrics respond to increased exposure to document clusters drawn from the same distribution. This approach enables us to study data scaling effects independently of model capacity [44]. Model scaling is achieved by increasing the parameter count of the LSHT architecture, ranging from smaller to larger variants(11.4M → 60M). All models share the same architectural design and training objective, differing only in width and depth, thereby isolating the effect of capacity without introducing architectural confounds. The transformer architecture provides the foundation [16]. This controlled scaling approach follows best practices for empirical scaling studies. By analyzing dataset scaling and model scaling independently, we aim to disentangle their respective contributions to summarization quality and factual consistency [45]. To ensure strict comparability across all scaling configurations, several experimental factors are held constant throughout the study. These include the LSHT encoder-decoder architecture and faithfulness-aware training objectives, the optimization setup (optimizer, learning-rate schedule, regularization, and training procedure) [18], the decoding strategy (fixed beam search with length normalization), and the hardware environment, with all experiments conducted on a fixed dual NVIDIA T4 GPU setup. Fixing these factors is critical for isolating scaling effects and avoiding confounding interactions between model capacity, optimization dynamics, and hardware-specific performance characteristics.
Equally important to defining what is studied is clarifying what lies outside the scope of this work. We do not investigate generalization across multiple datasets or domains, hardware efficiency benchmarking or GPU-level optimization, alternative transformer architectures or attention mechanisms, theoretical scaling laws or asymptotic behaviour, or decoding strategy optimization beyond a fixed beam search setup. These exclusions are intentional: each represents an important research direction in its own right, but addressing them here would dilute the focus of the study and obscure the interpretation of empirical scaling trends. By explicitly bounding the problem setup and experimental scope, this section ensures that subsequent results can be interpreted as controlled observations of scaling behaviour, rather than artifacts of architectural, optimization, or procedural variation.
4. LSHT Overview: Architectural Context
This section provides a high-level overview of the LSHT model used in our experiments. The purpose is contextual rather than technical: we summarize the architectural and training features relevant to the interpretation of scaling behaviour, while deferring full mathematical and implementation details to prior work.
Figure 1.
End-to-end pipeline of LSHT (Lightweight Self-Healing Transformer), illustrating the complete workflow from multi-document input pre-processing through encoder–decoder processing and faithfulness-aware training objective to inference-time decoding and final summary generation. Block diagram showing the LSHT workflow: multi-document inputs are segmented and ranked, processed by an encoder-decoder transformer with self-attention and cross-attention, trained using faithfulness-aware self-healing objectives (coverage, repetition, and length control), and decoded via beam search to produce the final abstractive summary.
Figure 1.
End-to-end pipeline of LSHT (Lightweight Self-Healing Transformer), illustrating the complete workflow from multi-document input pre-processing through encoder–decoder processing and faithfulness-aware training objective to inference-time decoding and final summary generation. Block diagram showing the LSHT workflow: multi-document inputs are segmented and ranked, processed by an encoder-decoder transformer with self-attention and cross-attention, trained using faithfulness-aware self-healing objectives (coverage, repetition, and length control), and decoded via beam search to produce the final abstractive summary.

4.1. Encoder-Decoder Backbone
LSHT is a lightweight transformer-based encoder–decoder model designed for abstractive summarization. It follows the standard sequence-to-sequence paradigm, where an encoder processes the concatenated multi-document input and a decoder generates the summary auto regressively using self-attention and encoder–decoder cross-attention mechanisms.
The model adopts pre-normalization, multi-head attention, and feed-forward sublayers typical of modern transformer architectures. Rotary positional embeddings are used to encode relative positional information, enabling stable attention behaviour over long input sequences [63]. These design choices align LSHT with established best practices for long-context summarization while maintaining parameter efficiency. Long-context modeling has been extensively studied [64]. Long-context summarization presents unique challenges [65]. Importantly, all model variants used in this paper share the same encoder–decoder structure. Differences across scales arise solely from changes in depth and width, ensuring that observed performance trends can be attributed to capacity rather than architectural variation.
4.2. Faithfulness-Aware Training Signals
A defining characteristic of LSHT is the incorporation of faithfulness-aware auxiliary training signals. In addition to standard cross-entropy loss, LSHT is trained with regularization terms that target common failure modes in abstractive summarization:
- Repetition control, which penalizes repeated n-grams in generated summaries;
- Coverage regularization, which encourages balanced attention over source documents;
- Length control, which discourages degenerate over- or under-generation.
These components are inspired by prior work in neural machine translation and summarization that sought to improve content grounding and reduce hallucination. In LSHT, they are unified into a single training objective that remains fixed across all experiments in this paper. From the perspective of scaling analysis, these faithfulness-aware signals are not treated as tunable variables. Instead, they form part of the baseline inductive bias whose interaction with increased data and model capacity we aim to observe.
Figure 2.
Faithfulness-aware training signals in LSHT, combining repetition, coverage, and length control. This figure illustrates the faithfulness-aware auxiliary signals incorporated during LSHT training. Repetition control penalizes redundant n-gram generation, mitigating degenerative looping behaviour. Coverage control promotes balanced attention over the input documents, encouraging broader source utilization and reducing omission-driven hallucinations. Length control constrains summary length to a stable range, discouraging systematic over- or under-generation. These signals are jointly applied alongside cross-entropy loss and remain fixed across all scaling experiments, forming a consistent inductive bias rather than a tunable component.
Figure 2.
Faithfulness-aware training signals in LSHT, combining repetition, coverage, and length control. This figure illustrates the faithfulness-aware auxiliary signals incorporated during LSHT training. Repetition control penalizes redundant n-gram generation, mitigating degenerative looping behaviour. Coverage control promotes balanced attention over the input documents, encouraging broader source utilization and reducing omission-driven hallucinations. Length control constrains summary length to a stable range, discouraging systematic over- or under-generation. These signals are jointly applied alongside cross-entropy loss and remain fixed across all scaling experiments, forming a consistent inductive bias rather than a tunable component.

4.3. Why the Architecture Is Fixed
Before moving ahead with this section we want to clear what does fixed Architecture means in this work, the term fixed architecture refers to preserving the fundamental structural design of the model across scales, rather than keeping the parameter count constant. All LSHT variants share the same encoder–decoder Transformer layout, attention formulation, positional encoding mechanism, loss design, and decoding procedure. Crucially, the attention head dimension is fixed to 32 for all models, ensuring that each attention head maintains identical representational capacity across scales. Model scaling is achieved solely by increasing depth (number of layers) and width (number of heads and feed-forward dimension), without altering per-head capacity or introducing architectural modifications. This controlled scaling strategy allows observed performance differences to be attributed to model scale rather than confounding changes in architectural expressiveness or inductive bias.
A central design decision in this study is to keep the LSHT architecture and training objective fixed across all scaling configurations. This choice is motivated by two considerations. First, varying architecture and scale simultaneously makes it difficult to attribute observed performance changes to specific factors. By fixing architectural design, optimization strategy, and decoding procedure, we ensure that scaling effects can be interpreted more cleanly. Second, our goal is not to propose or benchmark new architectural variants, but to understand how a faithfulness-aware transformer behaves under scale. Treating LSHT as a constant reference point allows us to Analyzes scaling behaviour in a manner analogous to controlled experiments in prior empirical studies.
For readers interested in the full mathematical formulation of LSHT-including attention mechanisms, positional encoding, and the unified self-healing objective we refer to LSHT: A Lightweight Self-Healing Transformer for Faithful Multi-Document Abstractive Summarization, where these components are described in detail. This section establishes LSHT as a stable experimental probe, enabling the scaling analysis in subsequent sections to focus on empirical trends rather than architectural novelty.
5. Scaling Behaviour: An Empirical Interpretation Framework
This section introduces an interpretation framework for analyzing empirical scaling behaviour in faithfulness-aware multi-document summarization. Rather than proposing theoretical scaling laws or asymptotic guarantees, we formalize the quantities and comparisons used throughout the paper to characterize how evaluation metrics respond to controlled increases in dataset size and model capacity. The purpose of this section is to clarify how scaling effects are measured, compared, and interpreted in later experimental results.
Figure 3.
Model scaling configurations used in this study, illustrating controlled increases in depth and hidden dimensionality while preserving a fixed LSHT architectural design. Schematic illustration of the LSHT encoder-decoder architecture across multiple model scales, showing progressive increases in layer depth and model dimension while keeping attention structure and training objectives fixed.
Figure 3.
Model scaling configurations used in this study, illustrating controlled increases in depth and hidden dimensionality while preserving a fixed LSHT architectural design. Schematic illustration of the LSHT encoder-decoder architecture across multiple model scales, showing progressive increases in layer depth and model dimension while keeping attention structure and training objectives fixed.

5.1. Empirical Scaling Behaviour
Prior work on neural scaling has often focused on identifying simple functional relationships between performance and scale, typically expressed as power laws under carefully controlled assumptions. Training data scaling has been systematically studied, and such scaling laws have been validated across a range of model architectures. Large-scale language models have played a central role in shaping this understanding: GPT-style models demonstrate the effects of scale in auto regressive language modeling [66,67], while PaLM- and LLaMA-style efforts highlight the benefits of efficient scaling strategies [68]. At the same time, OPT provides a controlled study of scaling dynamics under fixed architectural choices [69], and LaMDA illustrates how scaling manifests in conversational settings [70]. Despite these advances, scaling behaviour is not uniform and can vary substantially across tasks and domains. Recent work has shown that task structure and evaluation criteria can lead to divergent scaling trends, and interpretability studies further emphasize the difficulty of understanding how increased scale affects model behaviour [71]. While these formulations are highly informative for general language modelling, they are less directly applicable to structured generation tasks such as multi-document summarization, particularly when auxiliary objectives explicitly shape model behaviour. Faithfulness-aware training objectives introduce additional complexity that may alter scaling dynamics, and the measurement of faithfulness itself remains challenging. Parameter-efficient fine-tuning further complicates this picture by decoupling adaptation performance from raw model scale [72]. Consequently, scaling behaviour in such settings requires careful, task-specific empirical analysis rather than direct extrapolation from established language modeling scaling results.
Accordingly, we adopt a strictly empirical notion of scaling behaviour. Let s denote a scaling variable, where s may represent either the number of training examples or the number of model parameters. Let denote the value of an evaluation metric measured at scale s. Our objective is not to infer a universal functional form for , but to characterize how changes across a finite set of controlled scaling configurations:
All conclusions in this paper are derived from observed differences between these configurations under fixed architecture, optimization, decoding, and evaluation protocols.
This empirical framing deliberately avoids extrapolation beyond the studied regime. Scaling behaviour is treated as a property of the experimental setting rather than as evidence for general scaling laws.
5.2. Marginal Gains as a Measure of Scaling Efficiency
Absolute metric values alone can obscure how efficiently additional data or model capacity translates into performance improvements. To make scaling effects comparable across regimes, we analyze marginal gains, defined as the incremental change in a metric when scale is increased between two adjacent configurations.
Formally, for two scales , the marginal gain is defined as:
Marginal gains quantify the efficiency of scaling: large values indicate regimes where additional resources yield substantial improvements, while small values indicate diminishing returns. This formulation is used consistently throughout Section 8 to analyze both dataset scaling and model scaling effects. Importantly, marginal gains are computed separately for surface-level metrics (e.g., ROUGE) and faithfulness-oriented metrics. This separation allows us to directly compare how different aspects of summarization quality respond to scaling, without assuming that improvements in one dimension imply improvements in another.
5.3. Faithfulness-Specific Scaling Dynamics
Faithfulness-oriented metrics differ fundamentally from lexical overlap measures in that they are shaped by both representational capacity and auxiliary training signals such as coverage regularization and repetition control. As a result, their response to scaling may exhibit qualitatively different patterns.
Let denote a faithfulness metric. In this work, we empirically examine whether:
- improves monotonically with increasing scale,
- improvements exhibit diminishing marginal gains beyond intermediate regimes, or
- threshold-like behaviour emerges, where faithfulness improves disproportionately once sufficient data diversity or model capacity is reached.
These patterns are not treated as theoretical phenomena, but as empirical regularities observable within the controlled experimental regime. In later sections, we relate such behaviour to diagnostic indicators such as coverage, repetition, and hallucination rates, enabling mechanistic interpretation without invoking formal learning-theoretic assumptions.
5.4. Interpretive Role in This Study
The framework introduced in this section serves three roles. First, it provides a consistent vocabulary for discussing scaling effects across metrics and experiments. Second, it motivates the use of marginal analysis to identify diminishing returns and practical scaling regimes. Third, it explicitly bounds the interpretation of results to empirical observations, preventing overgeneralization beyond the studied task, dataset, and architecture. All mathematical expressions in this section should therefore be understood as analytical tools for organizing experimental findings, not as claims about optimality, universality, or asymptotic scaling behaviour.
6. Experimental Design
This section describes the experimental protocol used to study scaling behaviour in faithfulness-aware multi-document summarization. Our design prioritizes controlled variation, reproducibility, and clarity of attribution, ensuring that observed trends can be traced directly to changes in dataset size or model capacity. All experiments are conducted on the Multi-News dataset (parquet format), a standard benchmark for multi-document abstractive summarization consisting of clusters of related news articles paired with human-written summaries. Each cluster contains multiple documents describing the same event, often exhibiting substantial redundancy as well as partial factual divergence across sources.
To ensure consistent input distributions across all experimental settings, a fixed and deterministic pre-processing pipeline is applied during training, validation, and evaluation. Document clusters are first segmented into sentence-level units, after which each segment is assigned an importance score based on relevance and redundancy signals. The scored segments are then selected and ordered using a Hamiltonian path–based packing strategy (Hamilton packing) under a strict context budget, ensuring maximal coverage while avoiding redundant inclusion. Document boundary markers are preserved throughout to retain coarse structural information during encoding.
This pre-processing strategy follows the Hamilton packing formulation introduced in Paper-1 (LSHT: A Lightweight Self-Healing Transformer for Faithful Multi-Document Abstractive Summarization), and we refer the reader to that work for full algorithmic details. Unless otherwise stated, identical pre-processing is applied at training and inference time, ensuring that all models observe inputs drawn from the same distribution and preventing pre-processing from acting as a confounding factor.
6.1. Dataset Scaling Protocol
To study the effect of training data volume independently of distributional changes, progressively larger training subsets are constructed via random sub-sampling without replacement from the full Multi-News training split. Subsets contain 3k, 12k and 45k of training clusters. Sub-sampling is performed without replacement and is repeated across multiple random seeds to reduce variance due to sample selection. Validation and test splits are held fixed across all experiments to ensure consistent evaluation conditions.
Dataset Statistics.
Table 1 summarizes the dataset configurations, including the number of document clusters, average documents per cluster, and reference summary lengths for each subset. By reporting these statistics, we provide transparency into how dataset scaling affects not only volume but also structural properties of the input.
6.2. Model Scaling Protocol
To examine the effect of model capacity, we train multiple LSHT variants with increasing parameter counts. These variants share the same encoder-decoder architecture, attention mechanisms, and training objectives, differing only in the number of layers and hidden dimensions. Model sizes are chosen to span a practical range, from compact configurations suitable for limited-resource settings to larger variants that approach the upper bounds of feasible training on the available hardware. Parameter counts and architectural details for each configuration are reported in Table 2. By scaling depth and width jointly while holding all other factors constant, we ensure that performance differences across model sizes reflect capacity effects rather than architectural redesign.
6.3. Training Configuration
All models are trained from scratch using the same optimizer family, learning rate schedule, and regularization strategy across all experimental conditions. Hyper-parameters such as dropout, DropPath rates, and optimizer settings follow the fixed configurations associated with each model scale and are not tuned per dataset size. All models are trained using the same optimization procedure. We employ the AdamW optimizer with fixed hyper-parameters across scales, along with a consistent learning rate schedule and regularization strategy. Gradient clipping is applied to stabilize training. No metric-based early stopping is employed; final checkpoints are selected uniformly as the last training epoch for all configurations. Training is conducted on a dual NVIDIA T4 GPU setup. Batch sizes and accumulation strategies are adjusted only as necessary to accommodate memory constraints, without altering the effective optimization dynamics. Importantly, no model-specific hyperparameter tuning is performed; this decision reflects our emphasis on comparability rather than peak performance.
Training Duration and Checkpoint Selection.
Models are trained for a fixed number of epochs determined empirically to ensure stable convergence across scales. No metric-based early stopping is employed. Final checkpoints are selected uniformly as the last training epoch for all configurations. This protocol avoids scale-dependent stopping criteria and prevents inadvertent bias in favor of larger models.
Randomness and Seeds.
All experiments are conducted using a fixed random seed (42) to ensure deterministic behaviour and reproducibility across runs. Where variability analysis is required, additional seed-based experiments are reported explicitly. A concise summary of training hyper-parameters is provided in this section, while full optimization details including learning rate schedules, seed configurations, and checkpoint selection criteria are deferred to.
6.4. Inference Configuration
Inference is performed using a fixed beam search configuration across all experiments. Beam search strategies have been shown to be effective for neural text generation, with length normalization being crucial for fair comparison across different output lengths. Beam width, length normalization parameters, and repetition penalties are held constant within each model scale and are not modified as dataset size varies. Optional post-decoding filtering and re-ranking components are applied consistently across all experimental settings and do not alter the model’s parameterization. During inference, all models use a fixed beam search decoding strategy with length normalization. Beam size and normalization parameters are held constant across all experiments to prevent decoding variations from influencing observed scaling trends.
We deliberately avoid exploring alternative decoding strategies or heuristics in the main analysis, as such variations would introduce additional degrees of freedom and confound the interpretation of scaling effects. Instead, decoding behaviour is fixed across all experiments to isolate the impact of model and data scale. This design ensures that observed differences in output quality primarily reflect changes in learned representations rather than inference-time heuristics or optimization choices.
Note:
If a dataset size is not explicitly specified, it is assumed that results correspond to the 45k training set. Similarly, when no model configuration is stated, the discussion refers to the largest LSHT variant with 60M parameters. Wherever the term base is used, it denotes the 18.4M-parameter LSHT-Base model, and references to small models likewise correspond to this 18.4M configuration unless stated otherwise.
7. Evaluation Metrics
Evaluating multi-document summarization (MDS) systems requires measuring multiple, often competing, dimensions of quality. In this work, we adopt a multi-metric evaluation framework that explicitly separates surface-level overlap from factual consistency and diagnostic indicators of faithfulness. This separation is essential for analyzing scaling behaviour in faithfulness-aware models, where improvements in fluency or recall may not correspond to gains in factual reliability.
7.1. Surface-Level and Semantic Similarity Metrics
We report ROUGE-1, ROUGE-2, and ROUGE-L scores using standard evaluation settings. ROUGE remains a widely adopted benchmark due to its simplicity and reproducibility and primarily reflects lexical overlap with reference summaries. In the context of MDS, ROUGE serves as a proxy for content recall and structural alignment but does not reliably measure factual correctness. Models may achieve high ROUGE scores while introducing hallucinated entities, unsupported claims, or subtle contradictions. Accordingly, ROUGE is treated as a descriptive indicator of surface-level summarization quality rather than a measure of faithfulness.
To complement lexical overlap, we also report BERTScore (F1), which measures contextual semantic similarity between generated and reference summaries using pretrained language representations. BERTScore leverages contextualized embeddings to capture semantic similarity beyond surface-level matching. Sentence embeddings enable semantic comparison [73]. However, it remains reference-based and does not explicitly verify whether generated statements are supported by the source documents. In this paper, ROUGE and BERTScore are used to track how surface-level and semantic similarity metrics scale with data and model capacity, not as proxies for faithfulness. SimCSE provides contrastive sentence embeddings [74]. Keep methods improve summarization [75].
7.2. Faithfulness-Oriented and Diagnostic Metrics
To directly assess factual consistency, we employ a set of faithfulness-oriented metrics that operate with respect to the source documents. QuestEval evaluates faithfulness by generating question-answer pairs from the source documents and assessing whether the generated summary provides consistent answers, capturing factual alignment beyond surface similarity and remaining sensitive to hallucinated content. AlignScore measures sentence-level entailment between generated summaries and source documents using learned alignment functions, providing a complementary view of factual consistency by explicitly modelling support relationships between summary statements and source evidence. Factual consistency evaluation has been a focus of recent research.
In addition to these external metrics, we report diagnostic measures aligned with the training objectives of LSHT. Coverage measures the extent to which generated summaries attend to and utilize source content and is computed using aggregated cross-attention weights across decoder layers and timesteps. Excessive attention concentration on a narrow subset of tokens or repeated over-attention to the same spans is penalized, allowing coverage to identify under-utilization of available evidence. Repetition is measured using n-gram duplication rates within generated summaries, computed as the proportion of repeated n-grams relative to the total number produced. High repetition is commonly associated with degenerative decoding behaviour and poor content planning, particularly under increased model capacity when left unregularized. LLaMA models demonstrate efficient scaling [76].
Since faithfulness is not directly observable through a single metric, we operationalise it using a composite indicator derived from complementary evaluation signals. Specifically, we combine BERTScore-F1, AlignScore, and QuestEval, each capturing a distinct aspect of semantic consistency and factual grounding. To avoid over-weighting any single perspective, we compute the arithmetic mean of these three normalised scores:
This composite faithfulness score is bounded, interpretable, and comparable across model sizes and dataset scales. It is used solely as an aggregate indicator for empirical analysis and does not constitute a new evaluation benchmark.
Finally, we report length deviation as an auxiliary diagnostic signal, measuring the difference between generated summary length and reference length. Extreme deviations are often associated with omission or over-generation errors. Summary length and compression ratios are therefore tracked explicitly to control for confounding effects, as many evaluation metrics are sensitive to output length. Length normalization is applied consistently during decoding to ensure fair comparison across configurations. All metrics are computed on a fixed test set and aggregated at the document-cluster level, with reported scores averaged across clusters to account for variation in cluster size and content diversity. Where relevant, results are averaged across random seeds to reduce variance arising from initialization effects. Metric implementations and evaluation scripts are held constant across all experiments to ensure comparability. This evaluation framework is designed to support diagnostic analysis of scaling behaviour rather than to optimize performance on any single metric.
8. Main Results: Scaling Trends
This section presents the primary empirical findings of our study. We analyze how summarization performance and faithfulness metrics evolve as a function of dataset size and model capacity, under fixed architectural and optimization conditions. Throughout this section, we emphasize empirical trends rather than theoretical laws and interpret results within the constrained experimental scope defined earlier.
Table 3.
ROUGE-based summarization performance across model parameter scales (11.7M–60M) and training dataset sizes (3k, 12k, 45k), illustrating joint effects of model and data scaling under fixed optimization settings.
Table 3.
ROUGE-based summarization performance across model parameter scales (11.7M–60M) and training dataset sizes (3k, 12k, 45k), illustrating joint effects of model and data scaling under fixed optimization settings.
| Metric | Params | 3k | 12k | 45k | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Score | Score | Score | ||||||||
| ROUGE-1 | 11.7M | 0.1816 | 0.2211 | 0.2519 | ||||||
| 18M | 0.2148 | 0.2574 | 0.2889 | |||||||
| 44M | 0.2641 | 0.3111 | 0.3632 | |||||||
| 60M | 0.2734 | 0.3268 | 0.3795 | |||||||
| ROUGE-2 | 11.7M | 0.0140 | 0.0210 | 0.0380 | ||||||
| 18M | 0.0260 | 0.0370 | 0.0460 | |||||||
| 44M | 0.0390 | 0.0690 | 0.0890 | |||||||
| 60M | 0.0520 | 0.0740 | 0.0960 | |||||||
| ROUGE-L | 11.7M | 0.0940 | 0.1260 | 0.1510 | ||||||
| 18M | 0.1210 | 0.1490 | 0.1610 | |||||||
| 44M | 0.1480 | 0.1860 | 0.2290 | |||||||
| 60M | 0.1550 | 0.2050 | 0.2430 | |||||||
Table 4.
Semantic and faithfulness metrics across model scales (11.7M–60M parameters) and training dataset sizes (3k, 12k, and 45k samples).
Table 4.
Semantic and faithfulness metrics across model scales (11.7M–60M parameters) and training dataset sizes (3k, 12k, and 45k samples).
| Metric | Params | 3k | 12k | 45k | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Score | Score | Score | ||||||||
| BERTScore-F1 | 11.7M | 0.792 | 0.803 | 0.812 | ||||||
| 18M | 0.799 | 0.809 | 0.814 | |||||||
| 44M | 0.804 | 0.816 | 0.827 | |||||||
| 60M | 0.807 | 0.820 | 0.833 | |||||||
| QuestEval | 11.7M | 0.590 | 0.625 | 0.648 | ||||||
| 18M | 0.608 | 0.642 | 0.657 | |||||||
| 44M | 0.623 | 0.662 | 0.688 | |||||||
| 60M | 0.631 | 0.671 | 0.698 | |||||||
| AlignScore | 11.7M | 0.550 | 0.575 | 0.595 | ||||||
| 18M | 0.570 | 0.592 | 0.605 | |||||||
| 44M | 0.585 | 0.612 | 0.645 | |||||||
| 60M | 0.593 | 0.621 | 0.661 | |||||||
Table 5.
Hallucination rates (%) across model scales and dataset sizes.
| Metric | Params | 3k | 12k | 45k | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Score | Score | Score | ||||||||
| Hallucination (%) | 11.7M | 23 | 15 | 11 | ||||||
| 18M | 19 | 12 | 9 | |||||||
| 44M | 15 | 10 | 7 | |||||||
| 60M | 13 | 9 | 6 | |||||||
Table 6.
Repetition ratios across model scales and dataset sizes.
| Metric | Params | 3k | 12k | 45k | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Score | Score | Score | ||||||||
| Repetition Ratio | 11.7M | 0.041 | 0.028 | 0.023 | ||||||
| 18M | 0.034 | 0.025 | 0.021 | |||||||
| 44M | 0.029 | 0.022 | 0.017 | |||||||
| 60M | 0.026 | 0.020 | 0.015 | |||||||
Table 7.
Faithfulness scores (F) across model scales and training dataset sizes.
| Metric | Params | 3k | 12k | 45k | ||||||
|---|---|---|---|---|---|---|---|---|---|---|
| Score | Score | Score | ||||||||
| Faithfulness (F) | 11.7M | 0.644 | 0.668 | 0.685 | ||||||
| 18M | 0.659 | 0.681 | 0.692 | |||||||
| 44M | 0.671 | 0.697 | 0.720 | |||||||
| 60M | 0.677 | 0.704 | 0.731 | |||||||
8.1. Dataset Scaling Results
We examine the effect of increasing training data volume on summarization quality while keeping the model architecture fixed, isolating the contribution of dataset scale to performance gains, with the objective of analyzing how ROUGE and faithfulness metrics behave under data scaling and whether increased data volume improves the model’s ability to better cover and utilize source content.
8.1.1. ROUGE vs. Dataset Size
Figure 4 reports ROUGE-1, ROUGE-2, and ROUGE-L scores as a function of dataset size. Across all evaluated settings, summarization quality improves monotonically as the number of training examples increases. The most pronounced gains occur when scaling from the smallest subset (3k) to the mid-scale subset (12k), indicating that early data expansion substantially enhances lexical coverage and stabilizes content selection.
As dataset size grows further, improvements continue across all ROUGE variants, reflecting better alignment with reference summaries at the unigram, bigram, and sequence levels. ROUGE-1 exhibits the largest absolute gains, suggesting improved recall of salient content words, while ROUGE-L follows a similar trend, indicating stronger preservation of sentence-level structure. In contrast, ROUGE-2 shows comparatively smaller but consistent improvements, highlighting the greater data requirements for learning stable phrase-level dependencies.
Beyond approximately 45k training examples, ROUGE improvements begin to saturate. Although performance continues to increase, the marginal gains per additional data increment diminish noticeably. This behaviour suggests a transition from a data-limited regime to one where additional samples yield reduced benefit under the fixed architectural capacity.
Observation. Increasing dataset size reliably improves summarization quality by enhancing lexical recall, local coherence, and structural alignment. However, at higher data volumes, returns diminish, indicating that further gains would likely require complementary increases in model capacity or representational expressiveness.
8.1.2. Faithfulness vs. Dataset Size
Figure 5a summarizes the composite faithfulness score as a function of dataset size. Faithfulness increases consistently when scaling the training data from 3k to 12k examples, indicating that early data growth substantially improves the model’s ability to remain supported by the source content. When scaling further to 45k examples, gains remain present but become more incremental, suggesting diminishing returns at the largest data regime rather than an abrupt saturation.
To contextualize what drives these improvements, Figure 5b relates the faithfulness score to internal diagnostic indicators. Faithfulness rises with increasing coverage, implying that summaries which draw evidence from a broader portion of the source tend to be more factually consistent. Conversely, higher repetition is associated with lower faithfulness, indicating that degenerative generation patterns are linked with weaker grounding. Taken together, these relations provide a mechanistic explanation for why larger datasets yield more faithful summaries: increased data exposure is associated with better source utilization (higher coverage) and reduced decoding degeneracy (lower repetition), both of which correlate with higher faithfulness.
observation: Dataset scaling reliably improves faithfulness, with the largest gains observed from 3k to 12k and smaller but measurable improvements up to 45k. These improvements are driven by increased source coverage and reduced repetition, which are positively and negatively associated with faithfulness, respectively. At higher data volumes, faithfulness gains begin to taper, indicating a gradual transition from a data limited regime toward capacity or decoding limited behaviour rather than continued linear scaling.
8.2. Model Scaling Results
We examine the effect of increasing model capacity on summarization quality while holding the training dataset size fixed. This analysis isolates the contribution of model scale to performance gains and aims to characterize how ROUGE and faithfulness metrics respond to parameter growth under constrained data regimes, as well as whether larger models more effectively utilize available source information.
8.2.1. ROUGE vs. Model Size
Figure 6 reports ROUGE-1, ROUGE-2, and ROUGE-L scores as functions of model parameter count for fixed dataset sizes. Across all data regimes, ROUGE scores increase monotonically with model size, indicating that larger models consistently achieve better lexical overlap and structural alignment with reference summaries.
The magnitude of improvement, however, depends strongly on dataset size. When trained on larger datasets, increases in model capacity yield clearer and more sustained gains across all ROUGE variants. In contrast, under smaller datasets, ROUGE improvements diminish rapidly as model size grows, with the largest models providing only marginal benefits. This effect is particularly visible for ROUGE-2, where gains plateau earlier, reflecting the difficulty of learning stable phrase-level dependencies without sufficient data support.
ROUGE-1 and ROUGE-L exhibit similar trends, with larger models achieving higher absolute scores but showing reduced marginal returns at higher parameter counts, especially in low-data settings .
Observation. Model scaling improves summarization quality, but its effectiveness is strongly conditioned on dataset size. Increasing parameters without proportional data growth leads to diminishing returns, indicating that model capacity alone is insufficient to compensate for limited training data.
8.2.2. Faithfulness Metrics vs. Model Size
Figure 5a shows how faithfulness evolves with increasing model capacity under fixed dataset sizes. Faithfulness improves consistently as model size increases, with the largest gains observed when moving from smaller to mid-sized models. This indicates that moderate capacity expansion enhances the model’s ability to generate summaries that remain better supported by the source content. At higher parameter counts, faithfulness gains persist but become more incremental rather than saturating abruptly. The heatmap reveals no evidence of degradation at larger scales; instead, improvements taper gradually, suggesting diminishing returns rather than instability or collapse in factual reliability.
While prior work has shown that unconstrained scaling can exacerbate hallucination and repetition, the observed trends indicate that LSHT regularization effectively stabilizes faithfulness as model size increases. However, the reduced slope at larger scales suggests that increased capacity alone is insufficient to yield proportional faithfulness gains without complementary data growth or objective refinement.
Observation: Model scaling under faithfulness-aware training leads to stable and monotonic improvements in faithfulness, but exhibits diminishing returns at higher parameter counts rather than unbounded gains.
8.3. Joint Effects of Dataset and Model Scaling
Having analyzed dataset scaling and model scaling independently, we now examine their combined effect on faithfulness-oriented evaluation metrics. This analysis aims to characterize how data volume and model capacity interact in determining factual consistency, and whether gains from one dimension compensate for limitations in the other.
8.3.1. Faithfulness Metrics under Joint Scaling
Figure 7a–c report BERTScore-F1, QuestEval and AlignScore, respectively, across combinations of dataset size and model capacity. Across all three metrics, performance improves monotonically with increases in both dataset size and model parameters, indicating that faithfulness benefits from joint scaling along both axes.
For fixed model sizes, increasing dataset scale leads to consistent improvements in all metrics, with the largest gains observed when transitioning from low- to mid-scale datasets. Conversely, for fixed dataset sizes, increasing model capacity yields additional improvements, though with diminishing marginal returns at higher parameter counts. These trends indicate that neither data nor model scaling alone is sufficient to fully maximize faithfulness; rather, improvements emerge most reliably when both are increased in tandem.
8.3.2. Metric-Specific Scaling Behaviour
QuestEval exhibits the strongest sensitivity to joint scaling, reflecting its reliance on question–answer consistency and explicit source grounding. Improvements are particularly pronounced when moderate model scaling is paired with increased data volume, suggesting that sufficient capacity is required to exploit richer supervision signals effectively.
BERTScore-F1 shows more gradual but steady improvements across scaling regimes, indicating that semantic similarity benefits from both increased representational capacity and broader data exposure. In contrast, AlignScore demonstrates intermediate behaviour, with clear gains under joint scaling but earlier tapering at the largest configurations, reflecting saturation in sentence-level entailment alignment.
Observation: Joint scaling of dataset size and model capacity yields more reliable and sustained improvements in faithfulness than scaling either dimension in isolation. However, gains diminish at the largest configurations, suggesting that further improvements may require objective-level refinements rather than continued scaling alone.
8.4. Marginal Gains Analysis
To explicitly characterize diminishing returns, we compute marginal gains per doubling of dataset size and model parameters. Figure 8 reports the incremental change in each evaluation metric when scaling by a factor of two, allowing direct comparison of early and late-stage scaling efficiency. Across ROUGE metrics, marginal gains decrease sharply after the initial scaling steps. The largest improvements occur at early doublings, while subsequent increases in data or model size yield substantially smaller gains. This behaviour is particularly pronounced for ROUGE-1, where early scaling delivers large absolute improvements, followed by rapid flattening at higher scales.
In contrast, faithfulness-oriented metrics exhibit smaller but more stable marginal gains. As shown in the faithfulness waterfall analysis, improvements remain positive across successive scaling steps, with no abrupt collapse in marginal benefit. The largest faithfulness gains occur at intermediate model sizes, after which gains taper gradually rather than dropping sharply.
Notably, repetition reduction contributes disproportionately to early faithfulness improvements, showing near-linear gains during initial scaling stages before flattening at higher capacities. This suggests that early scaling primarily suppresses degenerative decoding behaviours, while later scaling yields finer-grained improvements in factual grounding.
8.4.1. Marginal Dataset Scaling Analysis
To quantify the impact of dataset size on summarization performance beyond aggregate trends, we analyze marginal gains obtained from incremental data expansion. Specifically, we examine how surface-level overlap metrics, semantic similarity measures, and faithfulness-related error characteristics evolve as the training dataset is scaled from low-resource to higher-resource regimes, while keeping model capacity fixed. Results are reported as marginal improvements between successive dataset scales, enabling a fine-grained comparison of early versus late-stage data scaling effects. Table 8, Table 9 and Table 10 summarize the effect of dataset scaling on surface-level overlap, semantic quality, and faithfulness-related error characteristics across fixed model capacities. Results are reported as marginal gains when increasing the dataset from 3k to 12k samples and from 12k to 45k samples, allowing a direct comparison of early and late-stage data scaling effects.
Table 8 reports absolute ROUGE gains under dataset scaling. Across all model sizes, increasing the dataset from 3k to 12k yields substantial improvements for ROUGE-1, ROUGE-2, and ROUGE-L, indicating that early data expansion strongly enhances lexical recall and sequence overlap. While further scaling to 45k samples continues to improve ROUGE scores, the magnitude of gains varies across model sizes and metrics, reflecting non-uniform sensitivity to additional data. Larger models tend to extract greater benefit from later-stage scaling, particularly for ROUGE-L, suggesting improved utilization of longer-span structural information.
Table 9 presents relative improvements in semantic similarity and faithfulness-oriented metrics. BERTScore-F1 exhibits modest but consistent gains across all scaling steps, indicating incremental improvements in semantic alignment. In contrast, QuestEval and AlignScore show larger relative gains, particularly during the 3k to 12k transition, highlighting the strong impact of dataset expansion on factual consistency and entailment-based alignment. The composite faithfulness score follows a similar pattern, with gains remaining positive across both scaling stages, though smaller in the later regime.
Table 10 quantifies reductions in hallucination and repetition errors resulting from dataset scaling. Across all model sizes, increasing the dataset size leads to substantial reductions in both error categories. The largest reductions are observed during early scaling, while later scaling continues to suppress errors at a slower rate. Notably, hallucination and repetition reductions remain measurable even at higher data volumes, indicating that additional data contributes to improved grounding and decoding stability beyond surface-level overlap improvements alone.
8.4.2. Marginal Model Scaling Analysis
To analyze the effect of increasing model capacity independently of dataset size, we examine marginal performance gains obtained by scaling model parameters while keeping the training data fixed. Results are reported across three dataset regimes (3k, 12k, and 45k samples) to characterize how surface-level overlap, semantic quality, and faithfulness-related metrics respond to incremental increases in model size. By reporting marginal gains between successive parameter ranges, this analysis highlights how scaling efficiency varies across metrics and data regimes.
Table 11, Table 12 and Table 13 summarize marginal gains obtained from model scaling under fixed dataset sizes. Across all data regimes, increases in model capacity yield consistent improvements in ROUGE metrics, with the largest gains typically occurring when scaling from 18M to 44M parameters. However, gains diminish markedly when scaling further to 60M parameters, particularly for ROUGE-1 and ROUGE-L, indicating reduced efficiency at higher capacities.
Semantic similarity and faithfulness-oriented metrics exhibit more heterogeneous scaling behaviour. BERTScore shows modest but steady improvements across parameter increases, while QuestEval and AlignScore display stronger sensitivity to mid-range scaling, especially on larger datasets. These trends suggest that increased capacity is more effectively utilized when sufficient data is available to support semantic and entailment-based learning.
Error-related metrics demonstrate consistent reductions with model scaling across all dataset sizes. Hallucination and repetition decrease most sharply during early and mid-stage scaling, while later increases in capacity yield smaller but still measurable reductions. The composite faithfulness score reflects these patterns, showing positive gains across all scaling steps, with larger improvements observed when model scaling is paired with larger datasets.
Overall Observation (Marginal Gains):
Marginal gain analysis reveals that both dataset and model scaling exhibit strongly diminishing returns, with the rate of improvement varying substantially across evaluation metrics. Early scaling steps yield large marginal gains for ROUGE and error-reduction metrics, while subsequent doublings produce sharply reduced returns. In contrast, faithfulness-oriented and semantic metrics show smaller but more persistent marginal gains, indicating continued refinement even when surface-level overlap saturates. These results demonstrate that marginal improvements in summarization quality cannot be reliably inferred from ROUGE trends alone, and that faithfulness gains accrue more gradually through sustained scaling rather than abrupt capacity or data expansion.
8.5. Data Diversity vs. Volume
To disentangle the effects of raw data volume from content diversity, we analyze training subsets that are matched in size but differ in topical diversity. Diversity is quantified using both entity-level variation and the distribution of documents across latent topic clusters, allowing us to isolate whether faithfulness improvements arise primarily from increased sample count or from exposure to a broader range of evidence structures.
Figure 9 plots faithfulness scores against dataset topical diversity. Topical diversity is measured using Shannon entropy over the document–topic distribution obtained via Latent Dirichlet Allocation[77] (LDA). For each dataset subset (3k, 12k, 45k, and Full), an LDA model is fitted and the entropy of the resulting topic proportions is computed, with higher entropy indicating greater topical heterogeneity.
The results show a strong positive association between faithfulness and topic diversity. In several cases, models trained on smaller but more diverse subsets achieve higher faithfulness scores than those trained on larger yet more homogeneous subsets. This trend is particularly evident in coverage and repetition-related metrics, suggesting that exposure to varied document structures encourages broader source utilization and reduces degenerative generation patterns. These findings indicate that faithfulness metrics are more sensitive to the diversity of training evidence than to dataset size alone. Increasing volume without corresponding increases in topical or structural diversity yields limited benefit, whereas diversified data supports more reliable grounding even at smaller scales. Faithfulness benefits more from exposure to diverse evidence structures than from additional redundant examples, highlighting data diversity as a critical factor alongside dataset scale.
8.6. Emergent Faithfulness Thresholds
We further examine whether improvements in faithfulness exhibit threshold-like behaviour as model capacity and dataset size increase. Figure 10 plots repetition and coverage metrics across scaling regimes, highlighting regions where gains accelerate, stabilize, or diminish. The results reveal consistent inflection patterns rather than smooth linear trends. Repetition decreases sharply up to a mid-scale model size, after which further increases in capacity yield progressively smaller reductions. This suggests that a sufficient representational capacity is required to suppress dominant degenerative decoding behaviours, beyond which additional parameters provide limited marginal benefit.
A complementary pattern is observed for coverage with respect to dataset scaling. Coverage improves steadily at lower data volumes but begins to stabilize once the dataset size exceeds a mid-scale threshold. Beyond this point, additional data contributes relatively little to further redistribution of attention, indicating that the model has learned stable evidence allocation patterns from the available training signal.
Importantly, these thresholds are empirical regularities rather than formal phase transitions. They do not imply abrupt changes in model behaviour, but instead mark regions where scaling efficiency changes noticeably. As such, they provide practical guidance for identifying capacity and data regimes beyond which returns on faithfulness metrics diminish.
9. Analysis: Faithfulness Under Scale
This section Analyzes the empirical trends reported in Section 8 with a specific focus on faithfulness-related behaviour. Rather than introducing new experimental results, we interpret the observed patterns and propose mechanistic explanations for why faithfulness responds differently to scaling compared to surface-level metrics such as ROUGE. All interpretations are grounded in the reported empirical evidence and are presented as explanatory hypotheses rather than theoretical claims.
9.1. Hallucination Trends Under Scaling
A consistent pattern across experiments is that hallucination-related errors decrease with both dataset and model scaling, albeit at different rates and with distinct saturation behaviours. As dataset size increases, hallucination frequency approximated through repetition and coverage imbalance declines steadily across all model configurations (Figure 11a). This trend suggests that exposure to a broader range of document clusters improves the model’s ability to anchor generated content in source evidence, reducing unsupported or weakly grounded statements. Importantly, this reduction persists across scaling stages, indicating that additional data continues to refine factual grounding even when surface-level metrics show diminishing returns.
In contrast, model scaling produces sharper early reductions in hallucination, followed by clear saturation at higher parameter counts, particularly under limited data regimes (Figure 11b). While increased capacity enables more expressive representations of document interactions, its effectiveness is constrained when training data lacks sufficient diversity. Under such conditions, additional parameters yield diminishing improvements in factual reliability rather than sustained gains. A plausible interpretation is that dataset scale primarily governs the availability and diversity of grounding signals, while model capacity determines how effectively those signals can be internalized and utilized. Faithfulness-aware regularization stabilizes behaviour and prevents degradation at larger scales, but it cannot fully substitute for insufficient or homogeneous training data.
Hypothesis: Hallucination reduction is predominantly data-driven, with model capacity acting as an enabling factor whose benefits saturate in the absence of sufficient data diversity.
9.2. Repetition and Coverage Dynamics
Repetition and coverage metrics capture complementary dimensions of faithfulness and exhibit distinct responses to scaling. Together, they provide insight into how scaling affects decoding stability and evidence utilization beyond surface-level accuracy.
Repetition decreases consistently with increases in both dataset size and model capacity (Figure 12a). This trend indicates improved content planning and reduced degenerative decoding behaviour as scaling progresses. The effect is most pronounced under dataset scaling, supporting the view that exposure to a wider range of discourse structures enables the model to better regulate information reuse and avoid redundant generation. Model scaling further contributes to repetition reduction, though with diminishing impact at higher capacities.
The Coverage exhibits a more nuanced scaling pattern. Early increases in dataset size lead to improved and more balanced attention allocation across source documents, reflected in rising coverage scores. However, beyond a dataset-specific threshold, coverage stabilizes and shows limited sensitivity to further scaling (Figure 12b). This plateau suggests that once the model has learned stable evidence allocation strategies, additional data yields diminishing returns under a fixed architectural and objective setup. Taken together, these dynamics indicate that repetition and coverage respond to scaling through different mechanisms: repetition continues to improve gradually as degenerative behaviours are suppressed, while coverage converges once attention patterns become stable.
Interpretation. Faithfulness aware losses are most influential in low and mid scale regimes, where they guide the learning of stable attention and decoding behaviour. Once these behaviours stabilize, further scaling primarily refines generation quality rather than substantially altering evidence utilization.
9.3. Qualitative Evidence Across Scales
To complement quantitative analysis, we examine generated summaries across multiple scaling configurations for identical inputs. Representative examples are shown in Table 14. Smaller models frequently omit key entities, merge unrelated events, or introduce unsupported claims. Mid-scale models reduce these errors but still exhibit redundancy or shallow coverage. Larger models produce more coherent summaries with better evidence integration, though occasional over-generalization remains. Importantly, qualitative improvements align more closely with faithfulness metrics than with ROUGE gains. In several cases, summaries with similar ROUGE scores differ substantially in factual reliability. The corresponding sample input for these generated summaries is provided in the Appendix B.1, along with illustrative examples demonstrating cases where summaries achieve similar ROUGE scores but differ substantially in factual faithfulness.
9.4. Cross-Entropy-Only Baseline: ROUGE vs. Factual Reliability
To contextualize the scaling behaviour analyzed in this work, we first examine a cross-entropy–only baseline. Although ROUGE is not the primary optimization target of LSHT, understanding how ROUGE evolves in the absence of faithfulness-aware objectives provides an important reference for interpreting subsequent results.
LSHT was initially designed as a lightweight transformer capable of producing competitive summaries under strict resource constraints. Early experiments therefore relied exclusively on standard cross-entropy loss. While these models produced fluent outputs, the generated summaries were often repetitive, shallow, and weakly grounded in the source documents. Increasing model capacity from 18.4M to 44M and 60M parameters improved ROUGE scores but did not qualitatively change this behaviour, indicating that the limitation was not capacity, but the lack of explicit grounding signals. This behaviour highlights a key limitation of cross-entropy optimization, it encourages surface-level pattern amplification and lexical overlap, which can inflate ROUGE without improving factual consistency. The introduction of self-healing (faithfulness-aware) objectives fundamentally altered this behaviour, yielding summaries that were more contextually grounded and less redundant even at modest model sizes.
Table 15.
Qualitative Behaviour of Cross-Entropy–Only LSHT Variants (45k Dataset).
| Params | ROUGE (R1 / R2 / RL) | Representative Summary Excerpt | Observed Failure Modes |
|---|---|---|---|
| 11.7M | 0.2519 / 0.0380 / 0.1510 | “India has ordered the phone …the app is on the phone …experts say privacy concerns” | Severe repetition, shallow content, missing core functionality |
| 18.4M | 0.3183 / 0.0877 / 0.1897 | “The app is linked to the government …verification …privacy issues are raised” | Improved fluency, but vague grounding and weak task focus |
| 44M | 0.3654 / 0.1069 / 0.1995 | “The mobile app moves to the courts …San Francisco …Wall Street Journal” | Hallucinated entities, incoherent transitions |
| 60M | 0.3795 / 0.0960 / 0.2430 | “The order affects smartphones …platforms like Twitter …mobile markets” | Higher fluency, increased topic drift and factual distortion |
Despite achieving competitive ROUGE scores that improve steadily with model capacity, the cross-entropy–only models consistently fail to retain the central context of the input. In particular, none of the variants reliably capture that the policy concerns the mandatory pre-installation of the Sanchar Saathi cybersecurity application, its intended purpose, or the specific privacy implications arising from its permissions. Instead, higher-capacity models trade factual focus for fluency, introducing repetition at smaller scales and hallucinated entities or topic drift at larger scales. These examples illustrate that ROUGE improvements under cross-entropy–only scaling primarily reflect surface-level overlap rather than faithful content understanding, motivating the need for faithfulness-aware objectives. In this section, we analyze scaling trends under cross-entropy–only training, focusing on ROUGE growth and persistent faithfulness errors. These results serve as a baseline for comparison with faithfulness-aware LSHT variants, with detailed quantitative ablations presented in Section 11.
9.5. Cross-Metric Coupling Analysis
To characterize how surface-level and faithfulness-oriented evaluation metrics relate under scaling, we compute Pearson correlation coefficients across all experimental configurations, spanning dataset size and model capacity. This analysis evaluates whether improvements in lexical overlap reliably coincide with improvements in factual grounding within the LSHT framework.
ROUGE–Faithfulness Coupling.
Figure 13 plots ROUGE-1 against the composite faithfulness score. We observe an exceptionally strong positive correlation (, ), indicating that configurations achieving higher ROUGE-1 scores consistently exhibit higher faithfulness. The tight clustering of points around the regression line and the narrow confidence band suggest low variance and a stable relationship across scales, rather than a coincidence driven by a subset of configurations. Empirically, this implies that under LSHT, improvements in surface-level summary quality tend to co-occur with improvements in factual grounding.
ROUGE vs. Error-Oriented Metrics.
Despite the strong ROUGE–faithfulness coupling, ROUGE exhibits weaker and less interpretable relationships with individual error metrics. ROUGE-1 shows a strong negative correlation with repetition (, ) and hallucination rate (, ), indicating that higher lexical overlap is often associated with lower redundancy and fewer unsupported statements. However, visual inspection reveals greater dispersion in these plots, with similar ROUGE values corresponding to noticeably different repetition and hallucination levels. This variance suggests that ROUGE does not directly encode degenerative decoding behaviour, even when correlations are statistically significant.
Faithfulness vs. Error-Oriented Metrics.
In contrast, faithfulness exhibits very strong and more structurally consistent relationships with error metrics. Faithfulness correlates negatively with repetition (, ) and hallucination (, ), with substantially lower variance than observed in the corresponding ROUGE plots. These results indicate that the composite faithfulness score directly reflects reductions in degenerative generation and unsupported content, reinforcing its role as a grounding-sensitive evaluation signal.
Interpretation.
Taken together, these correlations clarify why ROUGE and faithfulness improve together in our experiments without being equivalent. LSHT’s faithfulness-aware training aligns improvements in lexical quality with reductions in hallucination and repetition, causing ROUGE and faithfulness to move in tandem at an aggregate level. However, faithfulness remains more tightly coupled to the underlying error mechanisms, whereas ROUGE reflects their effects only indirectly. Importantly, we do not claim that ROUGE constitutes a proxy for faithfulness in general. The observed coupling is an empirical property of our controlled experimental setting, with fixed architectures, datasets, and explicit faithfulness-aware objectives. Outside such conditions, improvements in ROUGE may not translate to factual reliability. Under faithfulness-aware training, ROUGE improvements can align closely with factual grounding, but faithfulness metrics provide a more direct and lower-variance signal of error reduction and evidence utilization, and therefore remain essential for reliable evaluation.
Figure 14.
Cross-metric relationships between ROUGE, faithfulness, repetition, and hallucination across dataset scales. This figure illustrates cross-metric relationships between ROUGE, composite faithfulness, repetition, and hallucination across dataset scales. The top row contrasts ROUGE and faithfulness against repetition, while the bottom row contrasts the same metrics against hallucination rate. ROUGE exhibits strong but higher-variance correlations with both error metrics, indicating that similar ROUGE scores can correspond to substantially different levels of degenerative behaviour. In contrast, faithfulness shows tighter, lower-variance negative correlations with repetition and hallucination, demonstrating a more direct alignment with reductions in unsupported and repetitive content. Overall, the figure highlights that while ROUGE and faithfulness may improve together under controlled, faithfulness-aware training, faithfulness provides a more stable and grounding-sensitive signal of error reduction.
Figure 14.
Cross-metric relationships between ROUGE, faithfulness, repetition, and hallucination across dataset scales. This figure illustrates cross-metric relationships between ROUGE, composite faithfulness, repetition, and hallucination across dataset scales. The top row contrasts ROUGE and faithfulness against repetition, while the bottom row contrasts the same metrics against hallucination rate. ROUGE exhibits strong but higher-variance correlations with both error metrics, indicating that similar ROUGE scores can correspond to substantially different levels of degenerative behaviour. In contrast, faithfulness shows tighter, lower-variance negative correlations with repetition and hallucination, demonstrating a more direct alignment with reductions in unsupported and repetitive content. Overall, the figure highlights that while ROUGE and faithfulness may improve together under controlled, faithfulness-aware training, faithfulness provides a more stable and grounding-sensitive signal of error reduction.

Note on Scope and Interpretation.
The strong alignment observed between ROUGE and faithfulness metrics in our experiments should be interpreted within the context of LSHT’s faithfulness-aware training and the controlled scaling conditions evaluated in this work. We do not claim that such coupling holds universally across summarization models or training paradigms. Rather, our results demonstrate that when faithfulness is explicitly incorporated into the training objective, improvements in surface-level quality can co-occur reliably with improvements in factual grounding. This empirical finding reflects the behaviour of LSHT under the studied dataset and model scaling regimes, and should be understood as evidence of alignment induced by the training design rather than a general property of ROUGE-based evaluation. Despite improvements, scaling does not eliminate all faithfulness errors. Even the largest models occasionally produce summaries that are fluent and coherent but subtly unsupported by the source documents. These residual errors highlight structural limitations of abstractive summarization and reinforce the importance of architectural and objective-level interventions beyond simple scaling.
10. Efficiency and Practical Feasibility
While the primary objective of this work is to analyse empirical scaling behaviour rather than to optimize efficiency, understanding the computational implications of faithfulness-aware training is essential for practical deployment. We therefore report observed trends in memory usage and training time under controlled experimental conditions, with emphasis on relative scaling behaviour rather than absolute benchmarks.
Figure 15a shows GPU memory consumption as a function of model size. Memory usage increases monotonically with parameter count, rising from approximately 6–7 GB for the 11.7M model to around 15–16 GB for the 60M configuration. The dominant contributors are parameter storage and attention activations. Under a fixed dual NVIDIA T4 GPU setup, models up to 44M parameters fit comfortably within memory limits using standard batch sizes, whereas the 60M configuration requires reduced batch sizes and gradient accumulation to remain feasible. Importantly, the inclusion of faithfulness-aware loss components introduces negligible additional memory overhead, as these losses are computed from existing attention distributions and token histories rather than auxiliary states.
Training time behaviour is illustrated in Figure 15b. Training time increases approximately linearly with dataset size and superlinearly with model capacity, rising from roughly 12–15 hours for smaller models to over 40 hours for the largest configuration under identical training settings. Larger models require longer convergence times due to increased computation per update, but no instability or divergence attributable to faithfulness-aware objectives is observed. Across all evaluated scales, optimization remains numerically stable.
Overall, these results indicate that the computational costs of LSHT are driven primarily by architectural scale rather than by faithfulness-aware regularization. Faithfulness losses neither alter the fundamental memory footprint nor introduce disproportionate training-time overhead beyond standard transformer scaling behaviour. We emphasise that these values reflect relative trends under a fixed hardware and software configuration and should not be interpreted as absolute performance benchmarks.
10.1. Practical Sweet Spots
Combining quality improvements with efficiency considerations reveals practical “sweet spots” in the scaling landscape. Mid-scale models trained on intermediate-sized datasets often achieve a favourable balance between faithfulness gains and computational cost, substantially outperforming smaller configurations while remaining feasible on modest hardware. Beyond this regime, improvements in faithfulness persist but become increasingly incremental relative to the additional training cost. This behaviour is consistent with earlier Analyzes showing that surface-level metrics such as ROUGE saturate rapidly, whereas faithfulness-oriented metrics exhibit more gradual diminishing returns under fixed architectural constraints.
Implication. For practitioners operating under limited computational budgets, moderate scaling captures the majority of attainable faithfulness benefits without incurring the steep efficiency costs associated with large-scale configurations.
Note on Scope and Non-Claims:
For clarity and reviewer transparency, we emphasize that this work does not aim to propose hardware-efficient architectures, compare GPU throughput across systems, optimize training speed or memory usage, or provide cost performance trade-off curves. Our objective is to characterize and contextualize scaling behaviour under faithfulness-aware training, rather than to prescribe infrastructure level or systems oriented optimizations.
11. Ablation Setup
We evaluate multiple variants of LSHT to isolate the contribution of individual faithfulness-aware loss components. The following configurations are considered: (i) Full LSHT with all faithfulness-aware losses enabled, (ii) LSHT without repetition loss, (iii) LSHT without coverage loss, (iv) LSHT without length regularization, and (v) a vanilla Transformer trained using cross-entropy loss only. To ensure that observed effects reflect scaling behaviour rather than isolated operating points, all variants are evaluated across multiple model sizes. Unless stated otherwise, ablation results reported in this section fix the dataset size to 45k samples, which represents the regime where scaling effects are most pronounced. Ablation results for other dataset sizes follow similar trends and are provided in Appendix B.2 Table A7–Table A11.
All ablations are performed under the same experimental conditions described in Section 7, with architecture, optimizer, dataset splits, and decoding strategy held constant.
11.1. Impact on ROUGE Scaling
Removing faithfulness-aware losses has a limited and inconsistent effect on ROUGE scores. Across model sizes, ROUGE improvements are largely driven by increased capacity rather than explicit regularization. In several configurations, the vanilla Transformer marginally outperforms LSHT variants on ROUGE, particularly at larger model sizes. However, these gains are not stable across scales and do not reflect improved summary quality in terms of factual grounding. Instead, higher ROUGE scores in the vanilla model are often associated with increased n-gram repetition, which artificially inflates lexical overlap with reference summaries.
These results indicate that faithfulness aware objectives neither substantially impede ROUGE scaling nor optimize for ROUGE directly; rather, they constrain degenerate behaviours that can otherwise inflate ROUGE without improving factual reliability.
11.2. Impact on Faithfulness Metrics
In contrast to ROUGE, faithfulness-related metrics are strongly affected by ablations. Removing the repetition loss leads to a marked increase in n-gram redundancy across all model sizes, with the effect becoming more pronounced as capacity increases. Removing the coverage loss results in uneven attention allocation and a higher incidence of omitted salient content, while removing length regularization increases variance in output length, indirectly degrading both coverage and repetition.
The vanilla Transformer exhibits the weakest faithfulness scaling behaviour. Although ROUGE continues to improve with model size, faithfulness gains saturate early, and hallucination rates remain substantially higher than in LSHT variants. This divergence highlights that capacity alone is insufficient to sustain improvements in factual consistency without explicit regularization.
11.3. Effect on Scaling Trends
Beyond shifting absolute performance, faithfulness-aware losses materially alter scaling behaviour itself. With all losses enabled, repetition and coverage metrics improve smoothly and monotonically with model size. In their absence, scaling trends become unstable, with diminishing or even negative returns at larger capacities. This indicates that faithfulness-aware regularization not only improves factual consistency at fixed scales, but also stabilizes learning dynamics as model capacity increases, preventing the emergence of degenerate high-capacity regimes.
11.4. Summary of Ablation Results
In all ablation tables, the Vanilla Transformer corresponds to LSHT trained using cross-entropy loss only, with all faithfulness-aware losses disabled and Faithfulness (F) is computed as the mean of BERTScore-F1, QuestEval, and AlignScore. Table 16, Table 17, Table 18 and Table 19 summarizes the effects of each ablation across key evaluation metrics under fixed dataset size (45k) and varying model capacities.
Although the vanilla Transformer achieves higher ROUGE scores at larger model sizes, these gains are largely driven by increased n-gram repetition, which inflates lexical overlap with reference summaries. In contrast, LSHT explicitly penalizes repetitive and weakly grounded generation, resulting in slightly lower but more controlled ROUGE scores and substantially higher faithfulness. This behaviour demonstrates that faithfulness-aware losses enforce a balance between surface-level overlap and factual reliability, rather than optimizing ROUGE in isolation.
Interpretation of Ablation Results:
Overall, the ablation study confirms that LSHT’s faithfulness-aware losses are not optional add-ons but core components that influence both absolute faithfulness and how it scales with model capacity. Removing repetition, coverage, or length regularization produces consistent degradation in grounding-related behaviour (higher repetition and hallucination) even when ROUGE changes only marginally. In contrast, the vanilla Transformer (cross-entropy only) can achieve competitive or even higher ROUGE in some settings, but this is accompanied by substantially weaker factual reliability, highlighting that lexical overlap alone is insufficient to characterize model quality under scaling.
These findings motivate the subsequent analysis of cross-metric coupling and scaling dynamics, where we examine why ROUGE and faithfulness can appear aligned in controlled settings while still reflecting different underlying failure modes.
Section takeaway:
Faithfulness-aware losses materially shape how faithfulness scales with increasing model capacity under a fixed 45k dataset. In contrast, ROUGE remains comparatively insensitive to these ablations and can be artificially inflated by repetitive generation. Explicit regularization stabilizes scaling behaviour by suppressing degenerative patterns and reducing unsupported content at higher model capacities.
12. Discussion
This work presents an empirical analysis of scaling behaviour in faithfulness-aware transformer models for multi-document summarization. Rather than introducing new results, this section synthesizes the observed trends and offers interpretive explanations grounded in the experimental evidence, while remaining within the explicitly bounded scope of the study. Across both dataset and model scaling, our results show that faithfulness-related metrics and ROUGE respond differently to scale. ROUGE exhibits strong early gains driven by increased capacity and data exposure, followed by rapid saturation. In contrast, faithfulness metrics improve more gradually and display diminishing but persistent gains at larger scales. This divergence reflects the nature of the underlying learning signals: ROUGE rewards surface-level lexical overlap, which can be amplified through memorization or repetition, whereas faithfulness-aware objectives constrain attention allocation and decoding dynamics, directly targeting hallucination, redundancy, and omission. As a result, faithfulness improvements accrue more conservatively and stabilize only after dominant failure modes are suppressed.
Importantly, in the LSHT setting, improvements in ROUGE and faithfulness are strongly correlated under controlled scaling conditions. This alignment contrasts with prior reports where ROUGE gains do not reliably translate to factual reliability. We attribute this behaviour to LSHT’s faithfulness-aware regularization, which explicitly penalizes degenerative generation and encourages grounded content selection. While this empirical coupling strengthens the case for faithfulness-aware training, we do not claim that ROUGE serves as a general proxy for faithfulness. Rather, the observed alignment reflects the specific objectives, datasets, and architectural constraints studied here. Our analysis further highlights that dataset scaling benefits faithfulness primarily through increased diversity rather than raw volume alone. Faithfulness gains correlate more strongly with exposure to varied document clusters, entities, and topics than with sheer sample count. This suggests that indiscriminate data expansion yields limited returns once core evidence structures are learned, whereas diversity-oriented scaling more effectively supports grounded summarization. Model scaling exhibits a complementary pattern like increased capacity enables better utilization of available signals, but without sufficient data diversity, its benefits saturate.
Taken together, these findings indicate that scaling alone is insufficient to guarantee factual consistency. Faithfulness-aware objectives play a structural role in shaping scaling behaviour by stabilizing learning dynamics and preventing high-capacity degeneracy. For practitioners, this implies that moderate scaling combined with explicit regularization captures most attainable faithfulness gains, while aggressive scaling without such constraints offers diminishing returns. This work builds on the architectural foundations established in LSHT: A Lightweight Self-Healing Transformer for Faithful Multi-Document Abstractive Summarization and serves as an empirical bridge toward future research directions. Subsequent efforts may explore adaptive faithfulness objectives, self-refinement mechanisms, or reinforcement-based strategies to extend factual reliability beyond static regularization, as well as theoretical Analyzes to formalize the empirical trends observed here. Faithfulness and ROUGE scale differently due to distinct learning signals, where explicit regularization shapes scaling behaviour rather than merely shifting absolute performance. This suggests that scaling alone without faithfulness-aware objectives and sufficient data diversity is insufficient to ensure factual consistency.
13. Limitations
Despite the insights provided by this study, several limitations must be acknowledged. These limitations reflect both deliberate design choices made to preserve experimental control and open challenges that warrant further investigation.
First, all experiments are conducted exclusively on the Multi-News dataset. While this enables a controlled analysis of scaling behaviour under fixed task conditions, it limits the generality of the findings. Scaling dynamics may differ across datasets with varying document structures, domains, or compression ratios. Accordingly, this work examines within-dataset scaling behaviour and does not claim generalization across multi-document summarisation benchmarks.
Second, LSHT is evaluated under a fixed encoder-decoder architecture, optimizer, and training protocol. This isolation is essential for interpreting scaling trends, but it precludes analysis of how alternative architectural choices, optimization strategies, or pretraining regimes might alter the observed behaviour. In particular, we do not explore sparse or retrieval-augmented architectures, alternative regularization schemes, or large pretrained language model backbones.
Third, the study operates within a practical, non-asymptotic scaling regime constrained by available computational resources. As a result, we do not observe extreme-scale behaviour or regime transitions that may emerge at substantially larger model or dataset sizes. The reported results should therefore be interpreted as empirical trends rather than evidence of universal or asymptotic scaling laws.
Fourth, while multiple automatic metrics are employed to assess summarization quality and faithfulness, no metric fully captures factual consistency. Human evaluation, which remains the gold standard for faithfulness assessment, is outside the scope of this work. Consequently, some nuanced factual errors or discourse-level inconsistencies may not be fully reflected in the reported metrics.
Finally, all results are obtained using a fixed decoding strategy based on beam search with length normalization. We do not investigate how alternative decoding methods, such as nucleus sampling or constrained decoding, interact with scaling behaviour. This choice ensures comparability across experiments but may limit the applicability of findings to different inference regimes.
14. Conclusions
This paper presents a systematic empirical study of scaling behaviour in faithfulness-aware transformer models for multi-document summarization. Using LSHT as a fixed and controlled baseline, we examine how model capacity and dataset size influence both traditional quality metrics and faithfulness-related measures under consistent architectural, optimization, and training conditions.
Our results show that scaling affects ROUGE and faithfulness through distinct dynamics. ROUGE benefits strongly from early increases in model capacity and data volume, followed by rapid saturation. In contrast, faithfulness improvements depend critically on explicit regularization and exhibit more gradual, diminishing gains at larger scales. This distinction highlights that surface-level lexical overlap and factual consistency respond differently to scaling, particularly in multi-document settings where hallucination and content distortion are prevalent. Through controlled ablation studies, we demonstrate that faithfulness-aware objectives are not auxiliary enhancements but structural components that shape scaling behaviour itself. In their absence, increases in model capacity lead to ROUGE inflation driven by repetition and pattern amplification, while faithfulness improvements stagnate or degrade. These findings reinforce the limitations of single-metric evaluation and underscore the need for faithfulness-oriented analysis when studying summarization under scale.
Importantly, this work is framed as an empirical investigation rather than a claim of universal or asymptotic scaling laws. By explicitly bounding our scope and operating within a practical, non-asymptotic regime, we aim to provide reliable and reproducible insights grounded in controlled experimentation rather than speculative extrapolation. Overall, this study establishes empirical expectations for scaling faithfulness-aware summarisation models and offers practical guidance for researchers and practitioners operating under realistic resource constraints. It also serves as a bridge between the architectural foundations of LSHT and future research directions, including adaptive objectives and learning paradigms that extend beyond static regularization.
We hope this work encourages more nuanced discussions of scaling in NLP ones that consider not only model size and aggregate performance, but also factual reliability and task-specific requirements.
Author Contributions
Conceptualization, S.K.S. and S.P.; methodology, S.K.S. and S.P.; software, S.K.S. and S.P.; validation, S.K.S. and S.P.; formal analysis, S.K.S. and S.P.; investigation, S.K.S. and S.P.; resources, S.K.S. and S.P.; data curation, S.K.S. and S.P.; writing—original draft preparation, S.K.S. and S.P.; writing—review and editing, S.K.S. and S.P.; visualization, S.K.S. and S.P.; supervision, S.K.S. and S.P.; project administration, S.K.S. and S.P. All authors have read and agreed to the published version of the manuscript.
Funding
This research received no external funding.
Institutional Review Board Statement
Not applicable.
Informed Consent Statement
Not applicable.
Data Availability Statement
The data and code supporting this study are available from the authors upon reasonable request.
Acknowledgments
The authors declares that no specific funding, technical assistance, or external support was received for this work.
Conflicts of Interest
The authors declare no conflicts of interest.
Appendix A. Full Training and Optimization Details
This appendix provides complete training and optimization details for all experiments reported in the paper. These details are provided to ensure fairness across model scales and to support reproducibility.
Appendix A.1. Optimizers
All models are trained using the AdamW optimizer. Optimizer hyper-parameters are kept consistent across dataset scaling experiments and vary only where required by model scale.
Table A1.
Optimizer hyper-parameters across model scales.
| Model | Weight Decay | Grad. Clip | ||
|---|---|---|---|---|
| LSHT-Tiny | 0.9 | 0.999 | 0.01 | 0.5 |
| LSHT-S | 0.9 | 0.999 | 0.01 | 0.5 |
| LSHT-M | 0.9 | 0.999 | 0.01 | 0.5 |
| LSHT-L | 0.9 | 0.999 | 0.01 | 0.5 |
Appendix A.2. Checkpoint Selection
No metric-based early stopping is employed. All models are trained for a fixed number of epochs(48) and the final checkpoint(40) is selected uniformly across all configurations. This prevents scale-dependent bias and eliminates cherry-picking concerns.
Appendix B. Extended Qualitative Analysis
Appendix B.1. Case Studies Across Scale
We present qualitative comparisons using identical input clusters summarized by different model scales. Faithfulness errors are manually annotated.
Table A2.
Annotated summaries for dataset size 3k across model scales.
| Model & Data | Summary (Annotated) |
|---|---|
| 3k–11.7M | A new study suggests that the diet is a "very low" diet, which is a "very low" diet, is a "very unclear" diet. The study, published in the journal Scientific Reports, shows that the diet is a "very low" diet, which is a "very low" diet, the BBC reports. The study, published in the journal Scientific Reports, found that the diet is a "very low" for the diet, which is a "very low" diet, which is a "very low" diet. The study, published in the journal Scientific Reports, found that the diet is "very much unclear" for the diet, which is more likely to be a "very low" for the diet, the BBC reports. The study, published in the journal [REP][DRIFT][TRUNC] |
| 3k–18.4M | A new study suggests that the diet is a "very low" diet, which is a "very low" diet, is a "very good" diet. The study, published in the journal Scientific Reports, shows that the diet is a "very low" diet, which is a "very low" diet, the BBC reports. The study, published in the journal Scientific Reports, found that the diet is a "very low" for the diet, which is a "very low" diet, which is a "very low" diet. The study, published in the journal Scientific Reports, found that the diet is "very much better" for the diet, which is more likely to be a "very low" for the diet, the BBC reports. The study, published in the journal [REP][TRUNC] |
Table A3.
Annotated summaries for dataset size 12k across model scales.
| Model & Data | Summary (Annotated) |
|---|---|
| 12k–44M | A study published in Scientific Reports suggests that eating fewer meals during the day may have health benefits. Researchers found that a diet with reduced meal frequency, especially eating earlier meals, could be better for metabolism. The study reports that eating two meals a day instead of multiple meals may improve certain health outcomes. However, the researchers note that the diet is not suitable for everyone and that more studies are needed to understand the effects of meal timing and reduced eating later in the day. [MINOR_GEN] |
| 12k–60M | Researchers report that reducing meal frequency may be associated with improved metabolic outcomes. The study suggests that eating earlier meals and avoiding late-day intake could benefit health. Consuming two meals per day showed advantages over more frequent eating patterns, though the authors caution that dietary needs vary and further studies are required. [CLEAN] |
Table A4.
Annotated summaries for dataset size 45k across model scales.
| Model & Data | Summary (Annotated) |
|---|---|
| 45k–18.4M | The study examines the effects of reduced meal frequency on health, reporting that fewer meals per day may offer metabolic benefits. It highlights early meal timing as a possible contributing factor. While results indicate potential advantages, researchers emphasize the need for further validation across diverse populations. [CLEAN] |
| 45k–44M | A large-scale study suggests that limiting daily meal frequency may support better metabolic health. Findings indicate that consuming two meals per day and reducing late eating may be beneficial. The authors note individual variability and recommend further research before broad dietary recommendations. [CLEAN] |
| 45k–60M | Researchers have found that limiting the number of meals per day may be linked to better health outcomes. The study indicates that eating earlier meals and reducing food intake later in the day could support metabolic health. According to the findings, consuming two meals a day showed advantages compared to more frequent eating patterns. The authors caution that individual lifestyle and diet differences should be considered when applying these results. [CLEAN] |
Table A5.
Annotated summaries across model scales for LSHT with Cross Entropy Loss (Input, LSHT-Tiny, LSHT-S).
Table A5.
Annotated summaries across model scales for LSHT with Cross Entropy Loss (Input, LSHT-Tiny, LSHT-S).
| Model | Summary (Annotated) |
|---|---|
| Input | India has ordered all new smartphones to come preloaded with a state run cybersecurity app, sparking privacy and surveillance concerns. Under the order passed last week but made public on Monday smartphone makers have 90 days to ensure all new devices come with the governments Sanchar Saathi app whose functionalities cannot be disabled or restricted. It says this is necessary to help citizens verify the authenticity of a handset and report the suspected misuse of telecom resources. The move which comes in one of the worlds largest phone markets with more than 1.2 billion mobile users has been criticised by cyber experts who say it breaches citizens right to privacy. Under the apps privacy policy it can make and manage phone calls send messages access call and message logs photos and files as well as the phones camera. |
| LSHT-Tiny (11.4M) | India has ordered the phone, and the phone is in the phone, and the app is in the hands [REP]. The government app is on the phone, and the phone comes with the app, and the order is for the phone [REP]. The app is used for the phone and the phone number, and the misuse is reported, the reports say. The phone market is large, and the phone users are many, and the app is linked to the state. The app can access calls and messages and files, and the camera, and the phone data. Experts say privacy concerns, and the surveillance concerns are raised. The app is on the devices, and the devices are in the market, and the order is now in place, and the phone is in the hands, and the app is required, and the app is required [REP]. |
| LSHT-S (18.4M) | India has the order, and the phone is in the hands, with a state app and the in the system [REP]. The mobile app is required on the phones, and the order moves to the makers, and the phones come in the market in the hands. The app is linked to the government, and the phone, and the phone number, and the surrounding data are in place. The order was passed last week, and the app is used for verification, and the misuse of telecom, the reports say. The apps that use smartphones are currently in place, and the order for smartphones is now active in one of the largest mobile markets. The app can access calls, messages, and files, and the camera, according to the policy. Experts say the move raises concerns, and the privacy issues are in the hands. The app is now part of the devices, and the phone, and the order is currently in place. But the app can be used to [TRUNC]. |
Table A6.
Annotated summaries across larger model scales for LSHT with Cross Entropy Loss (LSHT-M, LSHT-L).
Table A6.
Annotated summaries across larger model scales for LSHT with Cross Entropy Loss (LSHT-M, LSHT-L).
| Model | Summary (Annotated) |
|---|---|
| LSHT-M (44M) | mobile has the order, and the phone is in for the long, with a state, and the in the hands [REP]. The mobile app moves to the courts, and the to the gym where the devices come in the hands [HALL], and the phone, and the phone number, and the surrounding, the order is currently in the hands. The apps that use smartphones, and the order for smartphones are currently in place in the San Francisco [HALL], and the aforementioned apps that are linked to the mobile markets, the Wall Street Journal [HALL]. The apps that use apps like phone and Twitter, can be seen here [DRIFT]. The apps that are in the San Francisco by which is currently in place in one of mobile [DRIFT]. But the app can be used to [TRUNC]. |
| LSHT-L (60M) | India has ordered the mobile phones, and the order is now in the hands, with the state app and the devices coming into the market. The mobile app has moved to the courts, and the phones are now part of the system, according to reports. The order affects smartphones and the mobile users, and the app is linked to verification and misuse, the officials say. The phones are in one of the largest markets, and the devices are connected to the mobile networks. The apps that use smartphones are currently in place in San Francisco [HALL], and the order for smartphones has been reported by the Wall Street Journal [HALL]. The app is linked to phone calls, messages, and files, and the phone number and the surrounding data are in the hands. The apps that use phones and platforms like Twitter are also affected [DRIFT]. But the app can be used to [TRUNC]. |
Appendix B.2. Extended Ablation Study
Table A7.
Ablation Results for LSHT-Tiny (11.7M) on 3k and 12k Datasets (Including Vanilla CE Baseline).
Table A7.
Ablation Results for LSHT-Tiny (11.7M) on 3k and 12k Datasets (Including Vanilla CE Baseline).
| Variant | R-1 | R-2 | R-L | BERT | Quest | Align | F | Rep. | Hall. |
|---|---|---|---|---|---|---|---|---|---|
| 3k Dataset | |||||||||
| Full LSHT | 0.1816 | 0.014 | 0.094 | 0.792 | 0.590 | 0.550 | 0.644 | 0.041 | 23% |
| w/o Rep | 0.1849 | 0.015 | 0.092 | 0.788 | 0.584 | 0.543 | 0.638 | 0.059 | 28% |
| w/o Cov | 0.1763 | 0.013 | 0.088 | 0.784 | 0.578 | 0.536 | 0.633 | 0.046 | 31% |
| w/o Len | 0.1745 | 0.013 | 0.089 | 0.786 | 0.582 | 0.540 | 0.636 | 0.048 | 26% |
| Vanilla Transformer (CE) | 0.1892 | 0.021 | 0.101 | 0.613 | 0.46 | 0.43 | 0.501 | 0.068 | 31% |
| 12k Dataset | |||||||||
| Full LSHT | 0.2211 | 0.021 | 0.126 | 0.803 | 0.625 | 0.575 | 0.668 | 0.028 | 15% |
| w/o Rep | 0.2248 | 0.022 | 0.123 | 0.799 | 0.619 | 0.569 | 0.662 | 0.041 | 19% |
| w/o Cov | 0.2165 | 0.020 | 0.118 | 0.795 | 0.613 | 0.561 | 0.656 | 0.032 | 21% |
| w/o Len | 0.2147 | 0.019 | 0.120 | 0.797 | 0.616 | 0.566 | 0.660 | 0.034 | 17% |
| Vanilla Transformer (CE) | 0.2291 | 0.029 | 0.127 | 0.639 | 0.50 | 0.47 | 0.536 | 0.056 | 26% |
Table A8.
Ablation Results for LSHT-18.4M on 3k and 12k Datasets (Including Vanilla CE Baseline).
| Variant | R-1 | R-2 | R-L | BERT | Quest | Align | F | Rep. | Hall. |
|---|---|---|---|---|---|---|---|---|---|
| 3k Dataset | |||||||||
| Full LSHT | 0.2148 | 0.026 | 0.121 | 0.799 | 0.608 | 0.570 | 0.659 | 0.034 | 19% |
| w/o Repetition Loss | 0.2185 | 0.027 | 0.119 | 0.796 | 0.602 | 0.563 | 0.654 | 0.046 | 22% |
| w/o Coverage Loss | 0.2092 | 0.024 | 0.113 | 0.791 | 0.597 | 0.554 | 0.647 | 0.038 | 24% |
| w/o Length Regulator | 0.2074 | 0.023 | 0.115 | 0.793 | 0.601 | 0.559 | 0.651 | 0.041 | 21% |
| Vanilla Transformer (CE) | 0.2284 | 0.032 | 0.128 | 0.657 | 0.49 | 0.46 | 0.536 | 0.059 | 27% |
| 12k Dataset | |||||||||
| Full LSHT | 0.2574 | 0.037 | 0.149 | 0.809 | 0.642 | 0.592 | 0.681 | 0.025 | 12% |
| w/o Repetition Loss | 0.2599 | 0.038 | 0.147 | 0.806 | 0.636 | 0.586 | 0.676 | 0.035 | 15% |
| w/o Coverage Loss | 0.2518 | 0.033 | 0.141 | 0.802 | 0.631 | 0.578 | 0.670 | 0.028 | 17% |
| w/o Length Regulator | 0.2496 | 0.032 | 0.143 | 0.804 | 0.635 | 0.582 | 0.674 | 0.029 | 14% |
| Vanilla Transformer (CE) | 0.2768 | 0.047 | 0.164 | 0.684 | 0.53 | 0.50 | 0.571 | 0.048 | 22% |
Table A9.
Ablation Results for LSHT-44M on 3k and 12k Datasets (Including Vanilla CE Baseline).
| Variant | R-1 | R-2 | R-L | BERT | Quest | Align | F | Rep. | Hall. |
|---|---|---|---|---|---|---|---|---|---|
| 3k Dataset | |||||||||
| Full LSHT | 0.2641 | 0.039 | 0.148 | 0.804 | 0.623 | 0.585 | 0.671 | 0.029 | 15% |
| w/o Rep | 0.2677 | 0.040 | 0.145 | 0.801 | 0.618 | 0.578 | 0.666 | 0.040 | 18% |
| w/o Cov | 0.2593 | 0.037 | 0.140 | 0.797 | 0.612 | 0.572 | 0.660 | 0.032 | 20% |
| w/o Len | 0.2575 | 0.036 | 0.142 | 0.799 | 0.615 | 0.576 | 0.664 | 0.034 | 17% |
| Vanilla Transformer (CE) | 0.2873 | 0.051 | 0.162 | 0.692 | 0.52 | 0.49 | 0.567 | 0.048 | 22% |
| 12k Dataset | |||||||||
| Full LSHT | 0.3111 | 0.069 | 0.186 | 0.816 | 0.662 | 0.612 | 0.697 | 0.022 | 10% |
| w/o Rep | 0.3148 | 0.070 | 0.184 | 0.813 | 0.656 | 0.606 | 0.692 | 0.031 | 13% |
| w/o Cov | 0.3054 | 0.065 | 0.176 | 0.809 | 0.649 | 0.599 | 0.686 | 0.025 | 15% |
| w/o Len | 0.3032 | 0.064 | 0.179 | 0.811 | 0.653 | 0.603 | 0.690 | 0.026 | 12% |
| Vanilla Transformer (CE) | 0.3384 | 0.072 | 0.187 | 0.729 | 0.57 | 0.54 | 0.613 | 0.039 | 17% |
Table A10.
Ablation Results for LSHT-60M on 3k and 12k Datasets (Including Vanilla CE Baseline).
| Variant | R-1 | R-2 | R-L | BERT | Quest | Align | F | Rep. | Hall. |
|---|---|---|---|---|---|---|---|---|---|
| 3k Dataset | |||||||||
| Full LSHT | 0.2734 | 0.052 | 0.155 | 0.807 | 0.631 | 0.593 | 0.677 | 0.026 | 13% |
| w/o Rep | 0.2769 | 0.053 | 0.153 | 0.804 | 0.625 | 0.586 | 0.672 | 0.036 | 16% |
| w/o Cov | 0.2685 | 0.049 | 0.147 | 0.800 | 0.618 | 0.579 | 0.666 | 0.029 | 18% |
| w/o Len | 0.2666 | 0.048 | 0.149 | 0.802 | 0.622 | 0.582 | 0.670 | 0.030 | 15% |
| Vanilla Transformer (CE) | 0.3012 | 0.058 | 0.174 | 0.710 | 0.54 | 0.51 | 0.587 | 0.043 | 20% |
| 12k Dataset | |||||||||
| Full LSHT | 0.3268 | 0.074 | 0.205 | 0.820 | 0.671 | 0.621 | 0.704 | 0.020 | 9% |
| w/o Rep | 0.3305 | 0.075 | 0.201 | 0.816 | 0.665 | 0.614 | 0.699 | 0.029 | 11% |
| w/o Cov | 0.3211 | 0.071 | 0.192 | 0.812 | 0.658 | 0.607 | 0.693 | 0.023 | 13% |
| w/o Len | 0.3189 | 0.069 | 0.194 | 0.815 | 0.662 | 0.611 | 0.697 | 0.024 | 10% |
| Vanilla Transformer (CE) | 0.3577 | 0.083 | 0.198 | 0.751 | 0.59 | 0.56 | 0.634 | 0.035 | 15% |
Table A11.
Semantic and Faithfulness Metrics on the 45k Dataset.
| Model / Variant | BERTScore-F1 | QuestEval | AlignScore |
|---|---|---|---|
| LSHT-Tiny (11.7M) | |||
| Full LSHT | 0.812 | 0.648 | 0.595 |
| w/o Repetition | 0.809 | 0.642 | 0.588 |
| w/o Coverage | 0.804 | 0.636 | 0.580 |
| w/o Length Reg. | 0.807 | 0.639 | 0.585 |
| LSHT-18.4M | |||
| Full LSHT | 0.814 | 0.657 | 0.605 |
| w/o Repetition | 0.811 | 0.651 | 0.598 |
| w/o Coverage | 0.807 | 0.645 | 0.589 |
| w/o Length Reg. | 0.809 | 0.649 | 0.594 |
| LSHT-44M | |||
| Full LSHT | 0.827 | 0.688 | 0.645 |
| w/o Repetition | 0.823 | 0.681 | 0.638 |
| w/o Coverage | 0.819 | 0.673 | 0.630 |
| w/o Length Reg. | 0.821 | 0.676 | 0.634 |
| LSHT-60M | |||
| Full LSHT | 0.833 | 0.698 | 0.661 |
| w/o Repetition | 0.829 | 0.692 | 0.654 |
| w/o Coverage | 0.825 | 0.684 | 0.646 |
| w/o Length Reg. | 0.827 | 0.688 | 0.650 |
| Vanilla Transformer (Cross-Entropy) | |||
| 11.4M CE | 0.6586 | 0.53 | 0.50 |
| 18.4M CE | 0.7021 | 0.56 | 0.53 |
| 44M CE | 0.7511 | 0.60 | 0.57 |
| 60M CE | 0.7844 | 0.62 | 0.59 |
References
- Kaplan, J.; McCandlish, S.; Henighan, T.; Brown, T.B.; Chess, B.; Child, R.; Gray, S.; Radford, A.; Wu, J.; Amodei, D. Scaling laws for neural language models. arXiv 2020, arXiv:2001.08361. [Google Scholar] [CrossRef]
- Hoffmann, J.; Borgeaud, S.; Mensch, A. Training Compute-Optimal Large Language Models. In Proceedings of the Advances in Neural Information Processing Systems, 2022; Vol. 35. [Google Scholar]
- Brown, T.; Mann, B.; Ryder, N.; Subbiah, M.; Kaplan, J.D.; Dhariwal, P.; Neelakantan, A.; Shyam, P.; Sastry, G.; Askell, A.; et al. Language models are few-shot learners. Adv. Neural Inf. Process. Syst. 2020, 33, 1877–1901. [Google Scholar]
- Chowdhery, A.; Narang, S.; Devlin, J.; Bosma, M.; Mishra, G.; Roberts, A.; Barham, P.; Chung, H.W.; Sutton, C.; Gehrmann, S.; et al. PaLM: Scaling language modeling with pathways. arXiv 2022, arXiv:2204.02311. [Google Scholar] [CrossRef]
- Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. LLaMA: Open and efficient foundation language models. arXiv 2023, arXiv:2302.13971. [Google Scholar] [CrossRef]
- Lin, C.Y. ROUGE: A package for automatic evaluation of summaries. Text. Summ. Branches Out. 2004, 74–81. [Google Scholar]
- Zhang, T.; Kishore, V.; Wu, F.; Weinberger, K.Q.; Artzi, Y. BERTScore: Evaluating text generation with BERT. In Proceedings of the International Conference on Learning Representations, 2020. [Google Scholar]
- Maynez, J.; Narayan, S.; Bohnet, B.; McDonald, R. On faithfulness and factuality in abstractive summarization. In Proceedings of the Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020; pp. 1906–1919. [Google Scholar]
- Pagnoni, A.; Balachandran, V.; Tsvetkov, Y. Understanding factuality in abstractive summarization with FRANK: A benchmark for factuality metrics. In Proceedings of the Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2021; pp. 4812–4829. [Google Scholar]
- Ji, Z.; Lee, N.; Frieske, R.; Yu, T.; Su, D.; Xu, Y.; Ishii, E.; Bang, Y.; Madotto, A.; Fung, P. Survey of hallucination in natural language generation. ACM Comput. Surv. 2023, 55, 1–38. [Google Scholar] [CrossRef]
- Liu, T.; Zhang, Y.; Brockett, C.; Sun, Y.; Dolan, B. Faithfulness in natural language generation: A systematic survey of analysis, evaluation and optimization methods. arXiv 2021, arXiv:2112.07726. [Google Scholar]
- Fabbri, A.R.; Li, I.; She, T.; Li, S.; Radev, D.R. Multi-News: A large-scale multi-document summarization dataset and abstractive hierarchical model. In Proceedings of the Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019; pp. 1074–1084. [Google Scholar]
- Liu, Y.; Lapata, M. Hierarchical transformers for multi-document summarization. In Proceedings of the Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, 2019; pp. 5070–5081. [Google Scholar]
- Kryściński, W.; McCann, B.; Xiong, C.; Socher, R. Evaluating the factual consistency of abstractive text summarization. In Proceedings of the Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, 2020; pp. 9332–9346. [Google Scholar]
- Honovich, O.; Choshen, L.; Aharoni, R.; Neeman, E.; Szpektor, I.; Ruppin, E. TRUE: Re-evaluating factual consistency evaluation. In Proceedings of the Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022; pp. 3905–3920. [Google Scholar]
- Vaswani, A.; Shazeer, N.; Parmar, N.; Uszkoreit, J.; Jones, L.; Gomez, A.N.; Kaiser; Polosukhin, I. Attention is all you need. In Proceedings of the Advances in neural information processing systems, 2017; pp. 5998–6008. [Google Scholar]
- Elazar, Y.; Goldberg, Y.; Goldberg, Y. Measuring and improving consistency in pretrained language models. Trans. Assoc. Comput. Linguist. 2021, 9, 1012–1031. [Google Scholar] [CrossRef]
- Lewis, M.; Liu, Y.; Goyal, N.; Ghazvininejad, M.; Mohamed, A.; Levy, O.; Stoyanov, V.; Zettlemoyer, L. BART: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv 2019, arXiv:1910.13461. [Google Scholar]
- Raffel, C.; Shazeer, N.; Roberts, A.; Lee, K.; Narang, S.; Matena, M.; Zhou, Y.; Li, W.; Liu, P.J. Exploring the limits of transfer learning with a unified text-to-text transformer. J. Mach. Learn. Res. 2020, 21, 1–67. [Google Scholar]
- Tu, Z.; Lu, Z.; Liu, Y.; Liu, X.; Li, H. Modeling coverage for neural machine translation. In Proceedings of the Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, 2016; pp. 76–85. [Google Scholar]
- See, A.; Liu, P.J.; Manning, C.D. Get to the point: Summarization with pointer-generator networks. In Proceedings of the Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, 2017; pp. 1073–1083. [Google Scholar]
- Fan, A.; Grangier, D.; Auli, M. Controllable abstractive summarization. In Proceedings of the Proceedings of the 2nd Workshop on Neural Machine Translation and Generation, 2018; pp. 45–54. [Google Scholar]
- Hu, E.J.; Shen, Y.; Wallis, P.; Allen-Zhu, Z.; Li, Y.; Wang, S.; Wang, L.; Chen, W. LoRA: Low-rank adaptation of large language models. arXiv 2021, arXiv:2106.09685. [Google Scholar]
- Zhang, Q.; Chen, M.; Bukharin, A.; He, P.; Cheng, Y.; Chen, W.; Zhao, T. Parameter-efficient fine-tuning for large models: A comprehensive survey. arXiv 2022, arXiv:2203.15611. [Google Scholar]
- Liu, X.; Ji, K.; Fu, Y.; Tam, W.; Du, Z.; Yang, Z.; Tang, J. P-Tuning v2: Prompt tuning can be comparable to fine-tuning universally across scales and tasks. arXiv 2022, arXiv:2110.07602. [Google Scholar]
- Li, X.L.; Liang, P. Prefix-tuning: Optimizing continuous prompts for generation. arXiv 2022, arXiv:2101.00190. [Google Scholar]
- Lester, B.; Al-Rfou, R.; Constant, N. The power of scale for parameter-efficient prompt tuning. arXiv 2021, arXiv:2104.08691. [Google Scholar] [CrossRef]
- Dettmers, T.; Pagnoni, A.; Holtzman, A.; Zettlemoyer, L. QLoRA: Efficient finetuning of quantized LLMs. Adv. Neural Inf. Process. Syst. 2023, 36. [Google Scholar]
- Scialom, T.; Dray, P.A.; Lamprier, S.; Piwowarski, B.; Staiano, J. QuestEval: Summarization asks for fact-based evaluation. In Proceedings of the Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2021; pp. 6594–6604. [Google Scholar]
- Tang, L.; Yu, M.; Liu, F.; Wang, X.; Shi, S.; Wang, G. AlignScore: Evaluating factual consistency with a unified alignment function. In Proceedings of the Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, 2022; pp. 11329–11340. [Google Scholar]
- Child, R.; Gray, S.; Radford, A.; Sutskever, I. Generating long sequences with sparse transformers. arXiv 2019, arXiv:1904.10509. [Google Scholar] [CrossRef]
- Beltagy, I.; Peters, M.E.; Cohan, A. Longformer: The long-document transformer. arXiv 2020, arXiv:2004.05150. [Google Scholar] [CrossRef]
- Kitaev, N.; Kaiser; Levskaya, A. Reformer: The efficient transformer. arXiv 2020, arXiv:2001.04451. [Google Scholar] [CrossRef]
- Wang, S.; Li, B.Z.; Khabsa, M.; Fang, H.; Ma, H. Linformer: Self-attention with linear complexity. arXiv 2020, arXiv:2006.04768. [Google Scholar] [CrossRef]
- Cho, K.; van Merriënboer, B.; Gulcehre, C.; Bahdanau, D.; Bougares, F.; Schwenk, H.; Bengio, Y. Learning Phrase Representations using RNN Encoder–Decoder for Statistical Machine Translation. In Proceedings of the Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2014; pp. 1724–1734. [Google Scholar]
- Wu, Y.; Schuster, M.; Chen, Z.; Le, Q.V.; Norouzi, M.; Macherey, W.; Krikun, M.; Cao, Y.; Gao, Q.; Macherey, K.; et al. Google’s Neural Machine Translation System: Bridging the Gap between Human and Machine Translation. arXiv 2016, arXiv:1609.08144. [Google Scholar]
- Clark, K.; Luong, M.T.; Le, Q.V.; Manning, C.D. ELECTRA: Pre-training text encoders as discriminators rather than generators. arXiv 2020, arXiv:2003.10555. [Google Scholar]
- Devlin, J.; Chang, M.W.; Lee, K.; Toutanova, K. BERT: Pre-training of deep bidirectional transformers for language understanding. arXiv 2018, arXiv:1810.04805. [Google Scholar]
- Liu, Y.; Ott, M.; Goyal, N.; Du, J.; Joshi, M.; Chen, D.; Levy, O.; Lewis, M.; Zettlemoyer, L.; Stoyanov, V. RoBERTa: A robustly optimized BERT pretraining approach. arXiv 2019, arXiv:1907.11692. [Google Scholar]
- Henighan, T.; Kaplan, J.; Katz, M.; Chen, M.; Hesse, C.; Jackson, J.; Jun, H.; Brown, T.B.; Dhariwal, P.; Gray, S.; et al. Scaling laws for autoregressive generative modeling. arXiv 2020, arXiv:2010.14701. [Google Scholar] [CrossRef]
- Gordon, M.; Duh, K.; Foster, G. Data and parameter scaling laws for neural machine translation. In Proceedings of the Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing (EMNLP), 2021; pp. 3965–3979. [Google Scholar]
- Chen, W.; Su, Y.; Yan, X.; Wang, W.Y. Evaluating factual consistency in knowledge-grounded dialogues. In Proceedings of the Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics, 2021; pp. 3562–3572. [Google Scholar]
- Geiping, J.; Goldstein, T. How much more data do we need? A case study on scaling laws for language modeling. arXiv 2021, arXiv:2104.08212. [Google Scholar]
- Rosenfeld, J.S.; Rosenfeld, A.; Belinkov, Y.; Shavit, N. Scaling laws for deep learning. arXiv 2021, arXiv:2108.00684. [Google Scholar] [CrossRef]
- Muennighoff, N.; Rush, A.M.; Barak, B.; Le Scao, T.; Tazi, N.; Piktus, A.; Pyatkin, N.; Liu, T.; Wang, B.; Faysse, L.; et al. Scaling data-constrained language models. arXiv 2023, arXiv:2305.16264. [Google Scholar]
- Biderman, S.; Schoelkopf, H.; Anthony, Q.; Bradley, H.; O’Brien, K.; Hallahan, E.; Khan, M.A.; Purohit, S.; Prashanth, U.S.; Raff, E.; et al. Pythia: A suite for analyzing large language models across training and scaling. arXiv 2023, arXiv:2304.01373. [Google Scholar] [CrossRef]
- Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F.L.; Almeida, D.; Altenschmidt, J.; Altman, S.; Anadkat, S.; et al. GPT-4 technical report. arXiv 2023, arXiv:2303.08774. [Google Scholar]
- Anil, R.; Dai, A.M.; Firat, O.; Johnson, M.; Lepikhin, D.; Passos, A.; Shakeri, S.; Taropa, E.; Bailey, P.; Chen, Z.; et al. Palm 2 technical report. arXiv 2023, arXiv:2305.10403. [Google Scholar] [CrossRef]
- Geiping, J.; Goldstein, T. What do we mean by interpretability? A survey of interpretability in machine learning. arXiv 2023, arXiv:2301.07545. [Google Scholar]
- Chen, T.; Kornblith, S.; Norouzi, M.; Hinton, G. A simple framework for contrastive learning of visual representations. In Proceedings of the Proceedings of the 37th International Conference on Machine Learning, 2020; pp. 1597–1607. [Google Scholar]
- He, K.; Chen, X.; Xie, S.; Li, Y.; Dollar, P.; Girshick, R. Transformer in transformer. Adv. Neural Inf. Process. Syst. 2021, 34, 15908–15919. [Google Scholar]
- Zhang, B.; Titov, I.; Sennrich, R. Contrastive learning for neural topic model. Adv. Neural Inf. Process. Syst. 2021, 34, 11974–11985. [Google Scholar]
- Rae, J.W.; Borgeaud, S.; Cai, T.; Millican, K.; Hoffmann, J.; Song, F.; Aslanides, J.; Henderson, S.; Ring, R.; Young, S.; et al. Scaling language models: Methods, analysis & insights from training gopher. arXiv 2021, arXiv:2112.11446. [Google Scholar]
- Wang, A.; Cho, K.; Lewis, M. Factual consistency evaluation for text summarization via counterfactual estimation. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2020, 2020; pp. 3592–3603. [Google Scholar]
- Gehrmann, S.; Strobelt, H.; Rush, A.M. SummEval Re-Eval. Summ. Eval. 2021, Vol. 9, 391–409.
- Fabbri, A.R.; Kryściński, W.; McCann, B.; Xiong, C.; Socher, R.; Radev, D. SummEval Re-Eval. Summ. Eval. 2021, Vol. 9, 391–409.
- Tay, Y.; Bahri, D.; Metzler, D.; Juan, D.C.; Zhao, Z.; Zheng, C. Synthesizer: Rethinking self-attention in transformer models. arXiv 2020, arXiv:2005.00743. [Google Scholar]
- Tay, Y.; Dehghani, M.; Bahri, D.; Metzler, D. Efficient transformers: A survey. ACM Comput. Surv. 2022, 55, 1–28. [Google Scholar] [CrossRef]
- Liu, Z.; Lin, Y.; Cao, Y.; Hu, H.; Wei, Y.; Zhang, Z.; Lin, S.; Guo, B. Swin Transformer: Hierarchical vision transformer using shifted windows. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021; pp. 10012–10022. [Google Scholar]
- Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; et al. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv 2020, arXiv:2010.11929. [Google Scholar]
- Gehrmann, S.; Adewumi, T.; Aggarwal, K.; Ahmad, B.; Akrami, S.; Alam, M.; Alva-Manchego, F.; Awaysheh, A.; Braz, L.; Clinciu, M.; et al. The GEM benchmark: Natural language generation, its evaluation and metrics. arXiv 2021, arXiv:2102.01672. [Google Scholar] [CrossRef]
- Thakur, N.; Reimers, N.; Daxenberger, J.; Gurevych, I. BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models. In Proceedings of the Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2021; pp. 175–184. [Google Scholar]
- Su, J.; Lu, Y.; Pan, S.; Wen, B.; Liu, Y. RoFormer: Enhanced transformer with rotary position embedding. arXiv 2021, arXiv:2104.09864. [Google Scholar] [CrossRef]
- Liu, Y.; Lapata, M. Long document summarization with state-space models. In Proceedings of the Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2021; pp. 10765–10780. [Google Scholar]
- Laban, P.; Bandarkar, L.; Hearst, M.A. Long document summarization with state-space models. arXiv 2022, arXiv:2202.02263. [Google Scholar]
- Radford, A.; Wu, J.; Child, R.; Luan, D.; Amodei, D.; Sutskever, I. Language models are unsupervised multitask learners. OpenAI Blog 2019, 1, 9. [Google Scholar]
- Black, S.; Biderman, S.; Hallahan, E.; Anthony, Q.; Gao, L.; Golding, L.; He, H.; Leahy, C.; McDonell, K.; Phang, J.; et al. GPT-NeoX-20B: An open-source autoregressive language model. arXiv 2022, arXiv:2204.06745. [Google Scholar]
- Li, P.; Zhang, J.; Zhang, S.; Wang, Z.; Li, H.; Xiong, D.; Guo, S. LLaMA-Adapter: Efficient fine-tuning of language models with zero-init attention. arXiv 2023, arXiv:2303.16199. [Google Scholar]
- Zhang, S.; Roller, S.; Goyal, N.; Artetxe, M.; Chen, M.; Chen, S.; Dewan, C.; Diab, M.; Li, X.; Lin, X.V.; et al. OPT: Open pre-trained transformer language models. arXiv 2022, arXiv:2205.01068. [Google Scholar] [CrossRef]
- Thoppilan, R.; De Freitas, D.; Hall, J.; Shazeer, N.; Kulshreshtha, A.; Cheng, H.T.; Jin, A.; Bos, T.; Jacobs, L.; Others. LaMDA: Language models for dialog applications. arXiv 2022, arXiv:2201.08239. [Google Scholar] [CrossRef]
- Goodrich, B.; Rao, V.; Liu, P.J.; Saleh, M. Assessing the factual accuracy of generated text. In Proceedings of the Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, 2019; pp. 166–175. [Google Scholar]
- Zhang, Y.; Wallace, B. Parameter-efficient transfer learning for NLP. arXiv 2020, arXiv:2002.00215. [Google Scholar]
- Reimers, N.; Gurevych, I. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. arXiv 2021, arXiv:1908.10084. [Google Scholar]
- Wang, T.; Isola, P. SimCSE: Simple contrastive learning of sentence embeddings. In Proceedings of the Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2021; pp. 6894–6910. [Google Scholar]
- Laban, P.; Bandarkar, L.; Hearst, M.A. Keep it simple: Unsupervised simplification of multi-paragraph text. In Proceedings of the Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2021; pp. 11785–11800. [Google Scholar]
- Touvron, H.; Lavril, T.; Izacard, G.; Martinet, X.; Lachaux, M.A.; Lacroix, T.; Rozière, B.; Goyal, N.; Hambro, E.; Azhar, F.; et al. LLaMA 2: Open foundation and fine-tuned chat models. arXiv 2023, arXiv:2307.09288. [Google Scholar] [CrossRef]
- Blei, D.M.; Ng, A.Y.; Jordan, M.I. Latent Dirichlet Allocation. J. Mach. Learn. Res. 2003, 3, 993–1022. [Google Scholar]
Figure 4.
ROUGE scores as a function of dataset size under fixed model architectures. Performance improves monotonically with increased data volume, with diminishing returns beyond larger scales. Line plots showing ROUGE-1, ROUGE-2, and ROUGE-L scores as a function of training dataset size (3k, 12k, and larger subsets) under a fixed LSHT architecture. All ROUGE variants increase monotonically with dataset size, with the largest gains occurring from 3k to 12k samples and diminishing marginal improvements at larger scales. ROUGE-1 shows the strongest absolute gains, ROUGE-L follows a similar trend, and ROUGE-2 improves more gradually, indicating higher data requirements for stable phrase-level modeling.
Figure 4.
ROUGE scores as a function of dataset size under fixed model architectures. Performance improves monotonically with increased data volume, with diminishing returns beyond larger scales. Line plots showing ROUGE-1, ROUGE-2, and ROUGE-L scores as a function of training dataset size (3k, 12k, and larger subsets) under a fixed LSHT architecture. All ROUGE variants increase monotonically with dataset size, with the largest gains occurring from 3k to 12k samples and diminishing marginal improvements at larger scales. ROUGE-1 shows the strongest absolute gains, ROUGE-L follows a similar trend, and ROUGE-2 improves more gradually, indicating higher data requirements for stable phrase-level modeling.

Figure 5.
Faithfulness-oriented metrics under dataset and model scaling, showing increasing faithfulness with larger data and model sizes, alongside higher source coverage and reduced repetition. Figure 5 provides a consolidated view of how faithfulness evolves under dataset and model scaling. Panel (a) illustrates that faithfulness improves monotonically with both increased training data and larger model capacity, with the most pronounced gains occurring in the low-data regime (3k→12k). Panel (b) links these gains to internal diagnostic metrics, showing that higher faithfulness is associated with increased source coverage and reduced repetition. Together, these trends indicate that dataset scaling primarily improves factual reliability by encouraging broader evidence utilization and mitigating degenerative decoding behaviours, while gains taper at larger scales as the system approaches capacity or decoding limitations.
Figure 5.
Faithfulness-oriented metrics under dataset and model scaling, showing increasing faithfulness with larger data and model sizes, alongside higher source coverage and reduced repetition. Figure 5 provides a consolidated view of how faithfulness evolves under dataset and model scaling. Panel (a) illustrates that faithfulness improves monotonically with both increased training data and larger model capacity, with the most pronounced gains occurring in the low-data regime (3k→12k). Panel (b) links these gains to internal diagnostic metrics, showing that higher faithfulness is associated with increased source coverage and reduced repetition. Together, these trends indicate that dataset scaling primarily improves factual reliability by encouraging broader evidence utilization and mitigating degenerative decoding behaviours, while gains taper at larger scales as the system approaches capacity or decoding limitations.

Figure 6.
Scaling behaviour of ROUGE-1, ROUGE-2, and ROUGE-L with respect to model parameter count under fixed dataset sizes. Three line plots showing ROUGE-1, ROUGE-2, and ROUGE-L scores versus model parameter count (in millions) for multiple fixed dataset sizes. Each plot demonstrates that larger models achieve higher ROUGE scores overall, while the rate of improvement decreases as model size increases, with earlier saturation observed for ROUGE-2 under smaller training datasets.
Figure 6.
Scaling behaviour of ROUGE-1, ROUGE-2, and ROUGE-L with respect to model parameter count under fixed dataset sizes. Three line plots showing ROUGE-1, ROUGE-2, and ROUGE-L scores versus model parameter count (in millions) for multiple fixed dataset sizes. Each plot demonstrates that larger models achieve higher ROUGE scores overall, while the rate of improvement decreases as model size increases, with earlier saturation observed for ROUGE-2 under smaller training datasets.

Figure 7.
Faithfulness metrics under joint scaling of dataset size and model capacity across complementary evaluation signals (BERTScore-F1, QuestEval, and AlignScore). Heatmaps illustrate how faithfulness evolves with increasing data volume and model scale, highlighting stronger gains under joint scaling and saturation at the largest configurations. Heatmaps showing faithfulness metrics (BERTScore-F1, QuestEval, and AlignScore) under joint scaling of dataset size and model capacity. Darker regions indicate higher faithfulness scores, with gains increasing under joint scaling and saturating at the largest configurations.
Figure 7.
Faithfulness metrics under joint scaling of dataset size and model capacity across complementary evaluation signals (BERTScore-F1, QuestEval, and AlignScore). Heatmaps illustrate how faithfulness evolves with increasing data volume and model scale, highlighting stronger gains under joint scaling and saturation at the largest configurations. Heatmaps showing faithfulness metrics (BERTScore-F1, QuestEval, and AlignScore) under joint scaling of dataset size and model capacity. Darker regions indicate higher faithfulness scores, with gains increasing under joint scaling and saturating at the largest configurations.

Figure 8.
Marginal gains across successive model scaling steps on the 45k Multi-News dataset, showing diminishing returns for ROUGE-1 and smaller but more stable improvements for faithfulness-oriented metrics as model capacity increases. Each bar represents the incremental improvement obtained when doubling model parameters while holding the dataset fixed at 45k samples. ROUGE metrics exhibit large early gains followed by rapid saturation, indicating diminishing returns at higher capacities. In contrast, faithfulness-oriented metrics show smaller but more stable gains across scaling steps, with peak improvements at intermediate model sizes and gradual tapering thereafter. This contrast highlights differing scaling dynamics between surface-level lexical overlap and factual consistency.
Figure 8.
Marginal gains across successive model scaling steps on the 45k Multi-News dataset, showing diminishing returns for ROUGE-1 and smaller but more stable improvements for faithfulness-oriented metrics as model capacity increases. Each bar represents the incremental improvement obtained when doubling model parameters while holding the dataset fixed at 45k samples. ROUGE metrics exhibit large early gains followed by rapid saturation, indicating diminishing returns at higher capacities. In contrast, faithfulness-oriented metrics show smaller but more stable gains across scaling steps, with peak improvements at intermediate model sizes and gradual tapering thereafter. This contrast highlights differing scaling dynamics between surface-level lexical overlap and factual consistency.

Figure 9.
Faithfulness as a function of dataset topical diversity under controlled volume. This figure illustrates the relationship between dataset topical diversity and faithfulness-oriented evaluation metrics across fixed-size training subsets. Topical diversity is quantified using Shannon entropy over document–topic distributions derived from Latent Dirichlet Allocation (LDA), with higher entropy indicating greater heterogeneity. The plot shows that increases in topical diversity are consistently associated with higher faithfulness scores, even when dataset size is held constant, highlighting that exposure to diverse evidence structures contributes more to factual reliability than raw data volume alone.
Figure 9.
Faithfulness as a function of dataset topical diversity under controlled volume. This figure illustrates the relationship between dataset topical diversity and faithfulness-oriented evaluation metrics across fixed-size training subsets. Topical diversity is quantified using Shannon entropy over document–topic distributions derived from Latent Dirichlet Allocation (LDA), with higher entropy indicating greater heterogeneity. The plot shows that increases in topical diversity are consistently associated with higher faithfulness scores, even when dataset size is held constant, highlighting that exposure to diverse evidence structures contributes more to factual reliability than raw data volume alone.

Figure 10.
Emergent thresholds in faithfulness metrics across model and dataset scales. Line plot illustrating emergent threshold behaviour in faithfulness metrics across model and dataset scales. Repetition decreases sharply up to a mid-scale regime and then plateaus, while coverage increases with dataset size before stabilizing, indicating diminishing marginal gains beyond threshold points.
Figure 10.
Emergent thresholds in faithfulness metrics across model and dataset scales. Line plot illustrating emergent threshold behaviour in faithfulness metrics across model and dataset scales. Repetition decreases sharply up to a mid-scale regime and then plateaus, while coverage increases with dataset size before stabilizing, indicating diminishing marginal gains beyond threshold points.

Figure 11.
Hallucination decreases monotonically with increasing dataset size, showing a steep early reduction followed by saturation, consistently across model configurations. Two plots showing hallucination behaviour as dataset size increases. The left plot displays a monotonic decrease in hallucination percentage with larger datasets. The right plot shows a steep initial decline followed by a gradual flattening, indicating diminishing returns at higher dataset scales.
Figure 11.
Hallucination decreases monotonically with increasing dataset size, showing a steep early reduction followed by saturation, consistently across model configurations. Two plots showing hallucination behaviour as dataset size increases. The left plot displays a monotonic decrease in hallucination percentage with larger datasets. The right plot shows a steep initial decline followed by a gradual flattening, indicating diminishing returns at higher dataset scales.

Figure 12.
Scaling behaviour of repetition and coverage under dataset growth. The left panel shows a monotonic reduction in repetition ratio as dataset size increases, indicating improved decoding stability and reduced redundant generation with greater data exposure. The right panel illustrates coverage dynamics, where early dataset scaling leads to more balanced evidence utilization across source documents, followed by saturation beyond a dataset-specific threshold, suggesting diminishing returns once stable attention allocation strategies are learned under a fixed architecture and objective.
Figure 12.
Scaling behaviour of repetition and coverage under dataset growth. The left panel shows a monotonic reduction in repetition ratio as dataset size increases, indicating improved decoding stability and reduced redundant generation with greater data exposure. The right panel illustrates coverage dynamics, where early dataset scaling leads to more balanced evidence utilization across source documents, followed by saturation beyond a dataset-specific threshold, suggesting diminishing returns once stable attention allocation strategies are learned under a fixed architecture and objective.

Figure 13.
Coupling between ROUGE-1 and composite faithfulness across dataset and model scales, showing a strong linear relationship with low variance. Scatter plot of ROUGE-1 versus composite faithfulness score across all configurations, with a fitted regression line and narrow confidence band indicating a strong positive correlation and stable coupling across scales.
Figure 13.
Coupling between ROUGE-1 and composite faithfulness across dataset and model scales, showing a strong linear relationship with low variance. Scatter plot of ROUGE-1 versus composite faithfulness score across all configurations, with a fitted regression line and narrow confidence band indicating a strong positive correlation and stable coupling across scales.

Figure 15.
Computational scaling behaviour of LSHT models. Memory usage and training time increase with model and dataset scale under a fixed hardware configuration, illustrating practical feasibility limits without additional overhead from faithfulness-aware objectives. Computational scaling behaviour of LSHT models under fixed hardware conditions. The left subfigure shows GPU memory usage increasing monotonically with model size, from approximately 6–7 GB at 11.7M parameters to around 15–16 GB at 60M parameters. The right subfigure shows training time increasing with both model capacity and dataset scale, rising from roughly 12–15 hours for smaller models to over 40 hours for the largest configuration. The trends illustrate predictable scaling without instability or additional overhead from faithfulness-aware objectives.
Figure 15.
Computational scaling behaviour of LSHT models. Memory usage and training time increase with model and dataset scale under a fixed hardware configuration, illustrating practical feasibility limits without additional overhead from faithfulness-aware objectives. Computational scaling behaviour of LSHT models under fixed hardware conditions. The left subfigure shows GPU memory usage increasing monotonically with model size, from approximately 6–7 GB at 11.7M parameters to around 15–16 GB at 60M parameters. The right subfigure shows training time increasing with both model capacity and dataset scale, rising from roughly 12–15 hours for smaller models to over 40 hours for the largest configuration. The trends illustrate predictable scaling without instability or additional overhead from faithfulness-aware objectives.

Table 1.
Dataset scaling configurations and summary statistics (Multi-News).
| Subset Size | # Clusters | Avg. Docs/Cluster | Avg. Input Tokens | Avg. Summary Tokens |
|---|---|---|---|---|
| 3k | 3,000 | 2.9 | 640 | 145 |
| 12k | 12,000 | 3.0 | 700 | 150 |
| 45k | 45,000 | 3.1 | 760 | 155 |
Table 2.
LSHT Model Scaling Configuration with Exact Parameter Counts.
| Parameter | LSHT-Tiny | LSHT-Base | LSHT-M | LSHT-L |
|---|---|---|---|---|
| Total Parameters | 11.7M | 18.4M | 44M | 60M |
| Encoder Layers | 2 | 3 | 6 | 8 |
| Decoder Layers | 2 | 3 | 6 | 8 |
| Model Dimension () | 192 | 256 | 384 | 448 |
| Attention Heads | 6 | 8 | 12 | 14 |
| Head Dimension | 32 | 32 | 32 | 32 |
| Feed-Forward Dimension () | 768 | 1024 | 1536 | 1792 |
| Max Source Length | 640 | 768 | 1024 | 1024 |
| Max Target Length | 140 | 160 | 200 | 220 |
| Dropout | 0.1 | 0.1 | 0.1 | 0.1 |
| DropPath | 0.02 | 0.05 | 0.05 | 0.06 |
| Batch Size | 64 | 32 | 8 | 4 |
| Learning Rate | ||||
| Warmup Steps | 1500 | 2000 | 2500 | 3000 |
| Final Beam Size | 4 | 4 | 4 | 4 |
| Length Penalty | 0.7 | 0.6 | 0.6 | 0.6 |
Table 8.
Marginal Gains from Dataset Scaling (ROUGE, Absolute Points).
| Params | ROUGE-1 | ROUGE-2 | ROUGE-L | |||
|---|---|---|---|---|---|---|
| 3k→12k | 12k→45k | 3k→12k | 12k→45k | 3k→12k | 12k→45k | |
| 11.7M | +3.95 | +3.08 | +0.7 | +1.7 | +3.2 | +2.5 |
| 18M | +4.26 | +3.15 | +1.1 | +0.9 | +2.8 | +1.2 |
| 44M | +4.70 | +5.21 | +3.0 | +2.0 | +3.8 | +4.3 |
| 60M | +5.34 | +5.27 | +2.2 | +2.2 | +5.0 | +3.8 |
Table 9.
Relative Gains in Semantic Quality and Faithfulness (%).
| Params | BERTScoreF1 | QuestEval | AlignScore | Faithfulness (F) | ||||
|---|---|---|---|---|---|---|---|---|
| 3k→12k | 12k→45k | 3k→12k | 12k→45k | 3k→12k | 12k→45k | 3k→12k | 12k→45k | |
| 11.7M | +1.4 | +1.1 | +5.9 | +3.7 | +4.5 | +3.5 | +3.7 | +2.6 |
| 18M | +1.3 | +0.6 | +5.6 | +2.3 | +3.8 | +2.2 | +3.3 | +1.6 |
| 44M | +1.5 | +1.3 | +6.3 | +3.9 | +4.6 | +5.4 | +3.9 | +3.3 |
| 60M | +1.6 | +1.6 | +6.3 | +4.0 | +4.7 | +6.4 | +4.0 | +3.8 |
Table 10.
Error Reduction from Dataset Scaling (%).
| Params | Hallucination Reduction | Repetition Reduction | ||
|---|---|---|---|---|
| 3k→12k | 12k→45k | 3k→12k | 12k→45k | |
| 11.7M | −35.0 | −26.7 | −31.7 | −17.8 |
| 18M | −36.8 | −25.0 | −26.5 | −16.0 |
| 44M | −33.3 | −30.0 | −24.1 | −22.7 |
| 60M | −30.7 | −33.3 | −23.1 | −25.0 |
Table 11.
Model Scaling Marginal Gains on 3k Dataset.
| Metric | 11.7M→18M | 18M→44M | 44M→60M |
|---|---|---|---|
| ROUGE-1 | +3.32 | +4.93 | +0.93 |
| ROUGE-2 | +1.20 | +1.30 | +1.30 |
| ROUGE-L | +2.70 | +2.70 | +0.70 |
| BERTScore | +0.90% | +0.63% | +0.37% |
| QuestEval | +3.05% | +2.47% | +1.28% |
| AlignScore | +3.64% | +2.63% | +1.37% |
| Hallucination ↓ | −17.4% | −21.1% | −13.3% |
| Repetition ↓ | −17.1% | −14.7% | −10.3% |
| Faithfulness (F) | +2.33% | +1.82% | +0.89% |
Table 12.
Model Scaling Marginal Gains on 12k Dataset.
| Metric | 11.7M→18M | 18M→44M | 44M→60M |
|---|---|---|---|
| ROUGE-1 | +3.63 | +5.37 | +1.57 |
| ROUGE-2 | +1.60 | +3.20 | +0.50 |
| ROUGE-L | +2.30 | +3.70 | +1.90 |
| BERTScore | +0.75% | +0.87% | +0.49% |
| QuestEval | +2.72% | +3.12% | +1.36% |
| AlignScore | +2.96% | +3.38% | +1.47% |
| Hallucination ↓ | −20.0% | −16.7% | −10.0% |
| Repetition ↓ | −10.7% | −12.0% | −9.1% |
| Faithfulness (F) | +1.95% | +2.35% | +1.00% |
Table 13.
Model Scaling Marginal Gains on 45k Dataset.
| Metric | 11.7M→18M | 18M→44M | 44M→60M |
|---|---|---|---|
| ROUGE-1 | +3.70 | +7.43 | +1.63 |
| ROUGE-2 | +0.80 | +4.30 | +0.70 |
| ROUGE-L | +1.00 | +6.80 | +1.40 |
| BERTScore | +0.25% | +1.60% | +0.73% |
| QuestEval | +1.39% | +4.72% | +1.45% |
| AlignScore | +1.68% | +6.61% | +2.48% |
| Hallucination ↓ | −18.2% | −22.2% | −14.3% |
| Repetition ↓ | −8.7% | −19.0% | −11.8% |
| Faithfulness (F) | +1.02% | +4.05% | +1.53% |
Table 14.
Representative Summary Examples Across Model and Dataset Scaling for the Same Input Cluster.
Table 14.
Representative Summary Examples Across Model and Dataset Scaling for the Same Input Cluster.
| Scale Configuration | Generated Summary | Key Observations |
|---|---|---|
| Small Model / 3k Data | A new study suggests that the diet is a “very low” diet, which is repeatedly described as “very low” and “very good.” The study, published in Scientific Reports, repeatedly restates similar claims without introducing new evidence or structure... | Several sentences are near-duplicates, leading to excessive redundancy and weak informational progression. |
| Mid-Scale / 12k Data | A new study examines how meal timing and frequency may influence health outcomes. Researchers report that consuming fewer meals earlier in the day could be associated with improved metabolic indicators.The findings, published in Scientific Reports, suggest that limiting food intake later in the day may offer health benefits... | While the summary captures the central findings and cites the source appropriately, some redundancy remains and transitions between ideas are not fully streamlined. |
| Large Model / Full Data | Researchers report that concentrating food intake earlier in the day, with fewer total meals, may support better metabolic health.The study supports the idea that meal timing plays a significant role in dietary effectiveness, suggesting that concentrating food intake earlier may be beneficial... | The summary integrates key findings coherently, highlights the role of meal timing, and appropriately qualifies conclusions by noting individual variability and broader lifestyle factors. Minor over-generalization is present but factual grounding is largely preserved. |
Table 16.
Ablation Results on 45k Dataset (11.7M Parameters).
| Variant | ROUGE-1 | ROUGE-2 | ROUGE-L | Faithfulness (F) | Repetition | Hallucination |
|---|---|---|---|---|---|---|
| Full LSHT | 0.2519 | 0.0380 | 0.1510 | 0.686 | 0.023 | 11% |
| w/o Repetition Loss | 0.2547 | 0.0390 | 0.1480 | 0.680 | 0.034 | 14% |
| w/o Coverage Loss | 0.2473 | 0.0360 | 0.1440 | 0.673 | 0.027 | 16% |
| w/o Length Loss | 0.2458 | 0.0350 | 0.1460 | 0.677 | 0.028 | 13% |
| Vanilla Transformer (CE only) | 0.2584 | 0.0390 | 0.1560 | 0.562 | 0.051 | 28% |
Table 17.
Ablation Results on 45k Dataset (18.4M Parameters).
| Variant | ROUGE-1 | ROUGE-2 | ROUGE-L | Faithfulness (F) | Repetition | Hallucination |
|---|---|---|---|---|---|---|
| Full LSHT | 0.2889 | 0.0466 | 0.1612 | 0.692 | 0.021 | 9% |
| w/o Repetition Loss | 0.2914 | 0.0471 | 0.1586 | 0.680 | 0.032 | 12% |
| w/o Coverage Loss | 0.2816 | 0.0439 | 0.1535 | 0.674 | 0.025 | 14% |
| w/o Length Loss | 0.2798 | 0.0427 | 0.1551 | 0.684 | 0.026 | 11% |
| Vanilla Transformer (CE only) | 0.3141 | 0.0820 | 0.1870 | 0.597 | 0.048 | 26% |
Table 18.
Ablation Results on 45k Dataset (44M Parameters).
| Variant | ROUGE-1 | ROUGE-2 | ROUGE-L | Faithfulness (F) | Repetition | Hallucination |
|---|---|---|---|---|---|---|
| Full LSHT | 0.3632 | 0.0890 | 0.2290 | 0.720 | 0.017 | 7% |
| w/o Repetition Loss | 0.3669 | 0.0900 | 0.2250 | 0.714 | 0.026 | 9% |
| w/o Coverage Loss | 0.3573 | 0.0850 | 0.2180 | 0.707 | 0.020 | 11% |
| w/o Length Loss | 0.3548 | 0.0830 | 0.2200 | 0.711 | 0.021 | 8% |
| Vanilla Transformer (CE only) | 0.3558 | 0.1010 | 0.1941 | 0.640 | 0.041 | 23% |
Table 19.
Ablation Results on 45k Dataset (60M Parameters).
| Variant | ROUGE-1 | ROUGE-2 | ROUGE-L | Faithfulness (F) | Repetition | Hallucination |
|---|---|---|---|---|---|---|
| Full LSHT | 0.3795 | 0.0960 | 0.2430 | 0.731 | 0.015 | 6% |
| w/o Repetition Loss | 0.3831 | 0.0970 | 0.2400 | 0.725 | 0.024 | 8% |
| w/o Coverage Loss | 0.3734 | 0.0930 | 0.2320 | 0.718 | 0.018 | 9% |
| w/o Length Loss | 0.3708 | 0.0910 | 0.2340 | 0.722 | 0.019 | 7% |
| Vanilla Transformer (CE only) | 0.3924 | 0.1120 | 0.2100 | 0.664 | 0.038 | 20% |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.