Submitted:
16 September 2026
Posted:
17 September 2026
You are already at the latest version
Abstract
Although Large Language Models (LLMs) produce fluent prose, they routinely tend to generate narratively sterile content for interactive media—backstories in which every conflict has been resolved before play begins, giving rise to what we call the Closure Paradox. Moreover, current evaluation metrics—from n-gram overlap to retrieval-augmented faithfulness scores—inadvertently reward this exact type of closure in cases in which it is undesired, leaving game designers without tools to measure whether a generated character is actually useful as a foundation for play. We propose a dual-axis evaluation framework that separates Temporal Consistency from Ludic Potential. Two new metrics anchor the framework: the Ludic Potential Index (LPI), a weighted measure of a backstory’s openness to future play, and the Potential Conflict Value (PCV), a per-event measure of playability fuel. The framework is implemented as a three-agent pipeline—Notary, Judge, and Stress-Tester—whose judge layers are stated in G-Eval terms so that they can be driven from standard open-source evaluation harnesses such as DeepEval and TruLens, and is validated through synthetic stress testing in which a separate LLM agent attempts to instantiate a Session 1 encounter from the generated backstory. The paper includes a comparative analysis of a correction mechanism implemented during generation (gated generation—where explicit metrics are applied to filter or constrain outputs) versus direct generation in the context of both commodity language models (e.g., Qwen or Llama) and frontier versions (e.g., Opus). The methodology evaluates the framework interventionally using a semantic, judge-based Adaptation Effort formulation (a metric of the additional work required to make new content playable for the character). The results show that (1) applying the gating mechanism yields a meaningful, medium-sized reduction in this Adaptation Effort (Ea), successfully reining in variability; (2) however, this benefit in frontier models produces only a marginal change, showing stability because these models naturally produce low-Ea drafts from the outset; (3) small-size commodity models do not benefit from the gate mechanism either, because their limited context handling does not profit from the iterative rewriting; and (4) nonetheless, the combined effect of the gating mechanism and the generator tier reveals a directional trend. The conclusion is that gating affordable commodity generators with explicit metrics meaningfully reduces downstream narrative preparation load; moreover, the presented set of metrics articulates criteria for selecting models and correction mechanisms, and their use can be applied to other narrative quality control processes. We argue that strategic vagueness is a functional requirement, not a defect, of interactive narrative—and that measuring it formalises the neuro-symbolic bridge between LLM generation and the symbolic state-tracking that interactive storytelling has required since TALE-SPIN.
Keywords:
large language models
; narrative generation
; tabletop role-playing games
; procedural content generation
; ludic potential
; neuro-symbolic systems
; DeepEval
; TruLens
; retrieval-augmented generation
1. Introduction
In procedural narrative generation, in particular in the context of interactive media (e.g., videogames but also traditional tabletop games), a recurring failure mode is what we call the Closure Paradox: “a well-written, self-contained story leaves no room for subsequent gameplay because all of its tensions have already been resolved”. The term Closure Paradox, as it is defined here, is based on previous works emphasising the fundamental mismatch between game rules (which require repetitive, open-ended action loops) and traditional narrative time (which moves unidirectionally toward an end state), Juul [1].
In this paper, we are going to focus our attention on the preliminary narrative material of many media representations (these are the premises of an interactive story in the format of the main character’s past prior to the beginning of the plot/game, generalised here as the character backstory). In interactive media, a character backstory is not a finished product but a kinetic foundation — a configuration of relationships, motivations, and unfinished business that triggers future events. Yet contemporary LLM evaluators, optimised for fluency and closure, systematically reward backstories that are inert at the moment of play. A character whose dead enemies stay dead, whose debts are paid, and whose mysteries are explained has nothing to push against, and consequently nothing to play.
This paper formalises a methodology for measuring the playability of a backstory rather than (only) its literary quality. By separating two axes — Consistency, the absence of contradictions, and Ludic Potential, the presence of actionable hooks — we provide a quantitative interface between LLM generation and downstream game logic. Our contributions are:
- 1.
- Two new metrics, the Ludic Potential Index () and the Potential Conflict Value (), that operationalise design heuristics from tabletop role-playing theory into LLM-judgeable rubrics.
- 2.
- A three-agent pipeline (Notary, Judge, Stress-Tester) that implements these metrics, with both judge layers expressed in G-Eval terms so the pipeline can be driven from existing open-source LLM-evaluation infrastructure such as DeepEval and TruLens.
- 3.
- A synthetic stress-test protocol, parameterised by an Adaptation Effort (), that lets an LLM-driven Game Master act as an inverse validator of the score assigned by the Judge.
- 4.
- A neuro-symbolic framing in which a JSON-Schema Fact Ledger extracted from neural prose serves as the symbolic substrate, addressing the Narrative Paradox first stated forty years earlier in the Symbolic-AI storytelling literature.
Section 2 situates this work within the LLM-as-judge, Symbolic-AI, RAG, TTRPG-NLP, and emergent-narrative literatures. Section 3 defines the metrics formally. Section 4 details the multi-agent pipeline and the experimental protocol. Section 5 reports expected results and discusses the agency–coherence trade-off. Section 6 surveys practical applications, and Section 7 concludes.
2. Related Work
2.1. From n-gram Metrics to LLM-as-a-Judge
Natural-language-generation (NLG) evaluation historically relied on n-gram overlap metrics such as BLEU and ROUGE, which correlate poorly with human judgement on creative tasks. The current state of the art for open-ended generation is the LLM-as-a-Judge paradigm [2]. G-Eval [3] demonstrated that prompting a frontier model with an explicit rubric and Chain-of-Thought reasoning produces evaluations aligned with human raters substantially better than statistical baselines, and the DeepEval framework operationalises this idea as unit tests for LLM behaviour. Story-level adaptations of the paradigm [4] have begun to produce rubrics along literary axes (relevance, coherence, surprise, engagement, empathy, complexity), but none yet evaluates playability — the dimension we target.
2.2. Symbolic AI and the Narrative Paradox
Before the rise of transformer-based models, the domain of narrative generation rested entirely within the realm of Symbolic AI. TALE-SPIN [5] and UNIVERSE [6] were early pioneers that used formal logic and goal-directed planning engines to ensure strict causal coherence throughout a story. However, a persistent challenge ermeged the Narrative Paradox: the struggle to balance rigid, pre-defined structures with the freedom for players to shape their own stories.
To create a bridge among those early plan-based methodologies [7] and sophisticated Drama Managers, e.g., Façade [8] to tackled the paradox it was used a exhaustive state-space exploration and PDDL-style domain modelling. While these approaches delivered unyielding internal consistency, they did so at a heavy cost sacrifing both linguistic nuance and the broad, common-sense world knowledge characteristic of human writers.
Kelly et al. [9] have sought to invert this longstanding compromise by distilling formal planning domains straight out of Large Language Model outputs. This work reclaim symbolic guarantees over raw neural prose. Our new framework inhabits this exact neuro-symbolic lineage: the Fact Ledger and the DeepEval-style unit-test formulation of its assertions serve as the rigorous symbolic scaffolding, while the underlying generator supplies the connectionist flair and generative fluency.
2.3. RAG and Temporal Consistency
Maintaining the world state stable in a long narrative is a hard task: LLMs tend to exhibit context drift over multi-thousand-token outputs and routinely contradict facts introduced earlier in the same generation. Retrieval-Augmented Generation (RAG) addresses this drifting by externalising the world state into a queryable Story Bible represented as a conected graph. RAGAs [10] defines reference-free metrics for context relevance, faithfulness, and answer relevance; while ARES [11] offers a fine-tuned-judge alternative with stronger calibration guarantees on domain-specific corpora. Both solutions prioritise grounding (adherence to the retrieved source) over narrative openness which is penalised: therefore a backstory that introduces an unresolved mystery has, by their definitions, more ungrounded claims than one that closes every loop. This is the gap our work targets: faithful grounding and strategic vagueness must be measured along orthogonal axes, not collapsed into a single faithfulness score.
2.4. Narrative Hooks in TTRPG Theory
The notion of “playability” is deeply rooted in tabletop-game design. Edwards’ GNS theory [12] distinguishes narrativist play — where a Premise (a situation demanding a choice) drives the game — from gamist or simulationist orientations, and frames a well-designed backstory as one that supplies sufficient Premise to the player and game master. Aylett’s emergent-narrative work [13] extends this into computational territory by arguing that interactivity requires leaving authorial gaps which emergent play can fill. The “Kickers and Bangs” vocabulary popularised in the indie-RPG community names the qualitative phenomena we quantify in this paper as the Motivation Vector () and the Unresolved Conflict Ratio (). To our knowledge, no prior work has formalised these design heuristics into computationally evaluable metrics applicable to LLM output.
2.5. LLM-Driven Tabletop Role-Playing Game Research
A dedicated research line at the intersection of NLP and tabletop role-playing has emerged in the past three years. Callison-Burch et al. [14] first positioned Dungeons & Dragons as an NLP benchmark, framing the game as a constrained dialogue challenge in which agents must track entities, intentions, and rules across long horizons. The FIREBALL dataset [15] operationalised this framing at scale, pairing actual-play transcripts with structured game-state annotations — the same conceptual move our Fact Ledger made for character backstories. CALYPSO [16] introduced an LLM-as-Dungeon-Master-assistant architecture whose role is the closest published analogue of our Stress-Tester agent. Acharya et al. [17] push the AI-DM further by binding LLM generation to tool-call schemas that enforce game-state consistency — a function-calling realisation of the same neuro-symbolic principle that underlies our pipeline.
PANGeA [18] stands as the most direct contemporary analogue that uses a generative approach. It crafts turn-based RPG narratives featuring personality-driven non-player characters using a validation loop for coherence. In the complementary survey by Gallotta et al. [19] a broader landscape of language models in video games is mapped. It describe the “evaluation of generated content for playability” as an unresolved challenge. This is precisely the void that and are engineered to bridge, a mechanism of evaluation of the playiability that can be computed along the creation of the content of narratives for video games.
There are methodologies focused on generating, running, or actively assisting tabletop role-playing sessions. Our work proposes an orthogonal axis metric that, rather than driving gameplay itself, it provides the foundational metric layer needed to rigorously determine whether a generated fragment is genuinely fit to serve as input for those systems.
2.6. Generative Agents and Emergent Narrative
The Generative Agents architecture of Park et al. [20] established that LLM-driven non-player characters (NPCs), equipped with memory streams, reflection, and planning, can sustain emergent social behaviour over extended simulations. This study push some industrial approaches, for instance, Peng et al. [21] have proven that such agents can successfully create the emergence of a player-perceptible narrative within commercial game environments with related constraints and focus.
In this work we propose to measure directly emergent narratives usefulness, at backstory level (where previous research demonstrates it at the agent level). A character’s true utility for interactive play is governed by the volume of latent affordances—relationships, motivations, and unresolved tensions—that subsequent events can activate.
Thus, while Generative Agents creates the necessary operational substrate, the Ludic Potential Index provides the missing evaluation substrate. Far from competing, the two paradigms are utterly complementary: an -validated backstory serves as a well-primed initial condition for a generative-agent simulation, whereas the Adaptation Effort functions as an ex post diagnostic of how productively that initial condition was ultimately exploited.
2.7. Story-Level LLM Evaluation Rubrics
A parallel literature has developed LLM-as-judge rubrics specifically for stories rather than for short-form NLG outputs. Chhun et al. [4] systematically prompted frontier models to score human and machine-written stories along six dimensions — relevance, coherence, empathy, surprise, engagement, complexity — and found that LLM judges align with human raters, as well as crowd workers do, on most axes. The dimensions Chhun et al. propose are literary in orientation: they ask whether a story reads well. The dimensions and propose are ludic in orientation: they ask whether a story plays well. The two rubrics are largely orthogonal — a story can score perfectly on engagement and absolutely horrible on Ludic Potential if every conflict it engages with has already been resolved by the final paragraph. Recent benchmark efforts such as the narrative-planning evaluations of Zhang et al. [22] begin to bridge the gap by adding causal-soundness and dramatic-conflict axes; / extend that bridge by anchoring the dramatic-conflict axis to operationalisable TTRPG-design concepts and validating the metric end-to-end through a Synthetic DM (Dungeon Master). In order to illustrate the framework, we will use example backstories generated in a known fictional setting (we use the Forgotten Realms setting of Faerûn, the same one we use for the experimental part of the paper).
3. Proposed Framework: LPI and PCV
We propose a dual-axis metric system. The first axis, Temporal Consistency, is enforced by a state-tracking extraction (STE) layer described in Section 4. The second axis, Ludic Potential, is the contribution of this paper and is quantified by two new metrics: the Ludic Potential Index () at the backstory level, and the Potential Conflict Value () at the event level.
3.1. The LPI × PCV Design Space
Together, and span a two-dimensional space in which any generated backstory can be located (Figure 1). The gate thresholds used in the ablation study, detailed below, (, ) partition this space into four interpretive quadrants:
- High LPI, High PCV (Campaign-Ready). The backstory contains active NPC relationships, clear protagonist goals, and high-reactivity hooks, while its events carry enough urgency and divergence to fuel a full campaign. This is the optimal region expected to reach by gating mechanism.
- High LPI, Low PCV (Rich Lore, Flat Play). The character has an extensive network of relationships and unresolved threads, however the events themselves lack immediacy or stakes—a common failure mode when LLMs produce elaborate plain closed narrative, a worldbuilding with no actionable conflict. This requires the designer (or the DM in a tabletop RPG) to foster urgency from scratch.
- Low LPI, High PCV (Explosive but Shallow). The events are, a priori, dramatic (a divine curse, a bounty on the protagonist’s head), but the backstory provides little-to-no relational scaffolding. Although designer/DM can run one explosive first session, she has limited material for the whole campaign.
- Low LPI, Low PCV (Unplayable Draft). This is the worst case, when there is neither character depth nor event urgency. This is the typical output of an unconstrained LLM asked to “write a D&D backstory” without forcing the generation of playeable hooks: every conflict is resolved, every mystery is explained, and every relevant NPC is either dead or absent.
3.2. The Ludic Potential Index ()
The metric is designed to measure the backstory’s openness based on the alternative options for the character’s future interactions. We define it as a weighted sum of four narrative vectors:
where each component (, , , and ) is independently normalised to . We illustrate each with contrasting examples drawn from the D&D setting used in the ablation study:
- – Relational Tension Nodes counts the number and intensity of active non-player characters with unresolved ties to the protagonist (enemies, debtors, missing family, betrayed allies).
- – Motivation Vector indicates the degree of character proactivity, measured as the count of clear, immediate goals the protagonist pursues (saturating at 4).
- – Unresolved Conflict Intensity is a fraction of open hooks that are Tier-1 (high-reactivity). It measures the sharpness of the unresolved layer rather than its mere existence. A backstory where 3 of 4 open hooks are Tier-1 (an active enemy, a ticking curse, a kidnapped ally) scores ; one where 1 of 4 hooks is Tier-1 and the rest are ambient curiosities scores . The open/closed ratio saturates at when the Architect prompt always produces unresolved hooks, so it contributes no discriminative signal.(Section 3.4).
- – Thematic Holes is the strategic vagueness, that is the inverse of fact density inside narrative mysteries.The score peaks at 3 holes (strategic and delivered vague spots in the narrative) and it decays beyond that threshold. Thus it encodes the insight that a handful of well-placed mysteries is a design asset but a proliferation of vagueness produces incoherence.
We adopt default weights , , , . These weights came from empirical tests (more details on the influence of weights in Section 5.4), by counting which dimensions a Synthetic DM agent draws from when generating Session 1 encounters (Section 4). During 160 stress-test sessions, the DM sourced material from NPC relationships () in 94% of encounters, from protagonist goals () in 71%, from open high-reactivity hooks () in 48%, and from thematic mysteries () in 19%.
Ranking the relational nodes first and thematic vagueness last is consistent with the idea described by Cover [23] and other TTRPG studies literature. They argue that NPC-to-player relationships are the primary generative engine of tabletop narrative, with social dynamics driving more emergent story than any other design element. Tychsen et al. [24] identify NPC-driven conflict and character motivation as the two highest-variance factors in session quality through structured observation of live role-playing sessions. Survey data reported by Merilainen [25] () confirm that social and relational dimensions of play are rated as the most impactful, followed by character-driven goal pursuit—empirically mirroring the ordering. At the theoretical level, Aarseth’s kernel/satellite decomposition of game narrative [26] assigns disproportionate structural weight to relational nodes (kernels) over thematic gaps (satellites), which maps directly onto .
We note that no existing study derives numerical coefficients for these exact four constructs; our weights are informed by the empirical rank-ordering above and by the Synthetic-DM draw-frequency counts, not by a formal factor analysis. The construct-validity framework described in Section 7 identifies weight calibration from human-DM play data as a priority for future work. We further confirm in Section 5.4, via a global sensitivity analysis over the locked corpus, that the framework’s play-ready classifications are robust to perturbation of these weights.
3.3. Component Anchors and a Worked Calculation
The four components of Equation 1 count different things—nodes, goals, hook tiers, and gaps—so a reader cannot compare a “high ” to a “high ” from the narrative descriptions alone. We therefore state all four on a single template: the quantity the Judge reads off the Fact Ledger, the normalisation that maps that quantity into , and the evidence that anchors a low and a high score. Writing for the number of NPCs carrying an unresolved tie to the protagonist, for the set of open hooks categorised by tier (e.g., , three tier-2 hooks and one tier-1), for the reactivity weight of a hook (Section 3.4), for the set of proactive goals, and k for the number of thematic holes, the normalisations are
with by convention when . Table 1 instantiates each normalisation at a low and a high anchor using the D&D setting of the ablation study, and Table 2 runs the full calculation end to end on two backstories drawn from the locked corpus (Section 5.2).
Reading the anchors.
The components of equations have three normalisation characteristics:
- 1.
- and saturate: a fourth adversary or a fifth goal adds nothing, because the constraint on a Session 1 encounter is the Game Master’s attention, not the supply of antagonists.
- 2.
- is a density, not a count—a backstory with one open hook that happens to be Tier-1 scores —so the volume of unresolved material is carried by the hook-value term of Equation 2 rather than by itself.
- 3.
- is the only non-monotonic component: it rewards up to three well-placed holes and then decays at per additional hole, encoding the framework’s strategic-vagueness rule as arithmetic rather than as prose.
Sterile and Play-ready calculations.
The lowest-scoring draft (Kaelen Vireth) and the highest-scoring one (Vessen Marrowlight) are calculated as follows:
Both backstories carry the same ceiling on active nodes, so the gap between them is produced almost entirely by , , and . The sterile draft comes from the ungated arm, so no gate was applied to it; its gate row reports the decision the gate would have taken.
Diagnostics are more than merely numeric. Both drafts describe a haunted protagonist with named antagonists, and both saturate the active-node term of ; a reader skimming either text would call both “dramatic”. What separates them is that the sterile draft states no goal its protagonist pursues (), leaves half its open threads below Tier-1 reactivity. It records a single thematic hole, whereas the play-ready draft hands the Game Master four goals, two Tier-1 threads, and three independent holes to author into. This is precisely the distinction that fluency scores and faithfulness metrics do not register.
Empirical spread.
The locked corpus components have markedly different portions of their common range: spans with median , so it is compressed near its saturation ceiling—generators reliably produce three or more tied NPCs. spans with median . spans with median , so it retains the widest usable spread around the midpoint, which is what the T1-density redefinition was intended to restore. Finally, spans with median , with a concentration in the upper range which is consistent with the variance decomposition of Section 5.4, in which contributes almost no variance to the play-ready classification.
3.4. Hook-Level Stratification
Within , individual hooks are not all equal. We classify them into three reactivity tiers used both during scoring and during generation layered based on their urgency or pushing strength:
- Tier 1 (weight 1.0).High-reactivity hooks: an active enemy, an urgent magical or medical condition, a missing family member. These can drive the first scene of play.
- Tier 2 (weight 0.6).Exploration hooks: lost objects, unsolved origin questions, moral debts. They require some unfolding before they produce scenes. This is more suitable for long-running campaigns.
- Tier 3 (weight 0.2).Ambient hooks: intellectual curiosities, stylistic quirks. They flavour a character but rarely drive sessions on their own. They can be useful for side-questing.
3.5. Potential Conflict Value ()
The quantifies the playability fuel of an individual event E within the backstory. For each event, the Judge assigns scores on four dimensions:
where each dimension is scored on a 1–10 scale:
- Immediacy (I)
- measures urgency for the next game session. A bounty hunter arriving at the tavern where the party is resting scores ; a prophecy about an event “centuries hence” scores .
- Gravity (G)
- measures the stakes involved, from trivial to existential. Losing a few coins () versus a divine curse that will kill the protagonist within a year ().
- Linkage (V)
- measures the degree of connection to the world-building RAG database. An event that names a faction, a location, and a historical precedent all present in the Story Bible scores high; a self-contained event with no world-state ties scores low.
- Divergence (D)
- measures the number of plausible gameplay outcomes the event can generate. “Kaelen discovers his mentor is a Zhentarim double agent” (: confront, expose, blackmail, flee, join) versus “Kaelen finds a locked chest” (: open it or leave it).
Weights.
We set and . This choice is structural rather than empirical, and we state it as such. The four weights sum to 10, so with each dimension on the 1–10 scale above a single event occupies : one maximally charged event is worth exactly 100, and the campaign threshold of 150 is therefore cleared by two strong events or by three or four ordinary ones. Within that budget, I and G carry four times the weight of V and D because they determine whether an event can drive the next session at all, whereas linkage and divergence modulate how much play an already-urgent event yields: a world-connected, many-outcome event that is neither pressing nor consequential is set dressing. Unlike the weights of Equation 1, these coefficients were fixed by construction before any data were collected, and they are not covered by the sensitivity analysis of Section 5.4, which perturbs the weights only — enters the gate as a weight-invariant input. Calibrating them against human judgements is left to the same future work as the weights (Section 7).
Thresholds and saturation.
A character is deemed play-ready only if their cumulative across all events exceeds a configurable Playability Threshold, which we set to 150 for full-campaign characters and 100 for one-shot characters; below 100 the character is flagged insufficient and, in the gated arm, triggers regeneration. We further enforce a Saturation of Conflict rule, because monotonic conflict produces monotonic play: events are bucketed by gravity band , and every event past the second in an occupied band contributes at half weight. A character whose three crises are three guild debts of comparable severity therefore banks materially less than one whose three crises are a debt, a curse, and a manhunt.
A worked calculation.
Table 3 runs Equation 7 end to end on the same two backstories used for the calculation of Table 2. The sterile draft banks 116 — above the one-shot floor but short of the campaign threshold — from two events that are grave but not urgent: its inciting loss is a disappearance the protagonist is not currently acting on, so I stays low while G stays high. The play-ready draft reaches 231 from three events whose urgency climbs across the life phases, closing on an adulthood event at . Neither character triggers the saturation penalty: the sterile draft’s two events sit at the band limit, and the play-ready draft spreads its three events across two gravity bands.
Written out, the first and last rows are
The diagnostic content of the contrast is that and fail the sterile draft for different reasons. fails it because it states no goal and holds only one Tier-1 thread (Table 2); fails it because its events, taken individually, give a Game Master nothing that must be answered in the next session. A backstory can clear one and miss the other, which is why the gate of Section 4 conjoins them.
3.6. The Strategic-Vagueness Rule
A subtle failure mode of LLM generation is over-determination: when asked to provide a backstory hook, the model often fills in every detail of its origin and resolution, leaving the Game Master nothing to author. We therefore make the openness of a mystery measurable. For a mystery span m, let be its tokens and those tokens that are proper nouns or explicit dates. The fact density and the Vagueness Index of m are
with by convention for an empty span. A mystery counts as a thematic hole — that is, it contributes to the count k that Equation 4 consumes — only if
The threshold is a design parameter, set at so that a mystery may name at most about one entity per three tokens before it stops counting as a gap. The rule separates the paper’s two running examples exactly: “Kaelen’s father was killed by someone connected to the army, using an unusual weapon” scores () and qualifies, while “Kaelen’s father was killed by Captain Hook on October 12 with a silver dagger” scores () — five of its fourteen tokens are names or date parts — and is rejected as closed exposition. The Vagueness Index is thus a gate on what enters k; the peak-and-decay term of Equation 4 then prices the resulting count, so the two mechanisms penalise the two opposite failures: sealing every mystery () and proliferating them ().
Detector and measured agreement.
The reference implementation resolves without a part-of-speech tagger: a token counts when it is capitalised in non-initial position, or when it is a numeral or a calendar month (Faerûnian or Gregorian). This is deliberately conservative — it over-counts sentence-initial names and under-counts lowercase proper nouns — so the criterion it applies must be checked against the extraction that actually populated the corpus. In the locked runs the Notary marks a mystery as a thematic hole from its extraction prompt (“where a fact is implied but unclear, mark it under thematic holes”), a qualitative judgement, rather than by counting tokens. Auditing the two criteria against each other over the 442 holes the Notary emitted across the locked corpus, mean fact density is (median ; mean ) and of the extracted holes also satisfy Equation 11; on the 478 holes of the Study 2 corpus the agreement is . The residual 5– are mysteries the Notary accepted that the density rule would reject as over-specified, which places a small upward bias on k — and hence on , the lowest-weighted component of Equation 1 — in the reported results. We report the arithmetic rule as the framework’s definition and this agreement rate as its measured fidelity, rather than claiming the locked corpus was screened by the rule itself.
4. Methodology
The framework is implemented as a four-stage pipeline (Figure 2) bridging connectionist generation with symbolic validation. Stage roles are deliberately asymmetric across the model tiers used in our experiments (Section 5): the Generator varies as the independent variable, the Notary is held constant on a small open-weight model, and the Stress-Tester and entity-reuse Judge are always Anthropic models (different tiers to avoid same-call bias). The deterministic Judge layer is the canonical LPI / PCV calculator and fires the gate decision. We describe each stage, then the experimental setup used to validate the metrics.
4.1. Stage I — State-Tracking Extraction (STE)
The first stage transforms backstory prose into a structured Fact Ledger. A Notary agent (held constant at qwen3:14b across every cell of both studies, so that extraction quality cannot confound the generator factor) is prompted to ignore literary style and emit a JSON document conforming to a schema that tags extracted events with a life phase (for simplicity, : infancy, : adolescence, : adulthood). Although the generation is done in a single call that produces the whole backstory, the events are tagged into phases and, within each phase, into categories:
- Physical state — permanent injuries, scars, approximate age, disabilities.
- Relational nodes — NPCs encountered, their status (alive/dead/unknown), and the nature of their tie to the protagonist.
- Inventory ledger — key items gained or lost.
- Skills and knowledge — explicit capabilities the character has demonstrated.
The schema also accepts a Background Event Log supplied by the designer or the world-building system, which records off-screen events whose impacts must be reflected in the on-screen prose. We detect logical faults — e.g. an NPC dying in but reappearing in without the Background Log explaining a resurrection — by structural diff over consecutive phases; any unjustified delta yields a Consistency score of 0 for the backstory.
4.2. Stage II — Quantitative Scoring
The Judge applies Equations 1 and 7 to the Fact Ledger in two layers, and the distinction between them matters for how the results below should be read.
Layer 1 — deterministic arithmetic.
The canonical calculator reads the ledger’s fields and evaluates Equations 2–4, 7 and 11 directly. It invokes no language model. Given a ledger, the , the cumulative , the tier and the gate decision are a pure function of its contents. Every , and gate figure reported in Section 5 is produced by this layer, which is why the released ledgers reproduce every reported score exactly without re-running any model. This is also what the yellow “deterministic” box in Figure 2 denotes: the only stochastic components in the pipeline are the Architect, the Notary, the Stress-Tester and the entity-reuse Judge.
Layer 2 — LLM rubric (second opinion).
Because the ledger is itself an LLM extraction, we additionally implement each component as a G-Eval [3] style rubric with an explicit Chain-of-Thought, so that the arithmetic layer can be cross-checked on the same input. For , the rubric Judge enumerates NPCs marked alive in the ledger and rates each on a 0–1 Tier-weighted relevance. For , it inspects the final paragraph for an explicit motion verb in the present or near-future tense. For , it counts opened vs. closed plot points across the three phases. For , it applies the Strategic-Vagueness Rule. This layer is a validation instrument, not a scoring path: the pilot of Section 5.1 reports the two layers agreeing to within on the calibration character, which is what licenses running the cheap deterministic Judge over the full corpus. No number in Section 5 is taken from it.
Both layers are stated in G-Eval terms so that the stage can be driven as a DeepEval assertion suite in deployment (Section 6); the experiments reported here execute the two layers directly against the provider APIs, so no reported result depends on that harness.
The regeneration loop (INSTABILITY_PROMPT).
In the gated arm, a draft that fails or is not discarded but returned to the Architect under a single fixed instruction, INSTABILITY_PROMPT, held constant across every gated trial in both studies and released verbatim with the run configuration. It asks for a rewrite the whole backstory explicitly, not an addition, with the following instructions: (i) to force the introduction of one named NPC, still alive at the end of the text, allied with or opposed to an injected faction; (ii) inserts or promotes one new Tier-1 hook naming that NPC and tying the protagonist to an injected location where a confrontation must occur, left unresolved; and (iii) adds one new proactive motivation absent from the previous draft. The three demands map one-to-one onto , and , so the loop is a targeted repair of the components the gate measures rather than a generic “try again”. The prompt is re-applied until the draft clears the gate or is exhausted; the ungated arm never invokes it. Section 5.5 shows that a generator’s ability to execute this structural rewrite — rather than its parameter count — is what decides whether gating helps or harms.
4.3. Stage III — RAG Observability
The backstories are generated, as we have said, in the Faerûn setting of Forgotten Realms and they must remain grounded in the world’s canonical lore. The framework specifies TruLens instrumentation of the generation path against the standard RAG triad: context relevance (do retrieved Story-Bible chunks actually inform the current life phase?), groundedness (does every world-fact in the prose trace to a retrieved chunk?), and answer relevance (does the prose actually fulfil the design prompt?). Lore violations — e.g. a character openly using necromancy in a city of the Lords’ Alliance, which the Bible forbids — are penalised against the Consistency score, leaving the untouched. This asymmetry is the framework’s central design commitment: grounding failures and openness are scored on separate axes, so that an unresolved mystery is never charged as an ungrounded claim.
We state the scope of the present experiments precisely. The studies reported in Section 5 are ablations on the Ludic Potential axis: every cell binds generation to the same Forgotten-Realms Story Bible, and the gated arm injects a Bible faction and location into INSTABILITY_PROMPT, but the retrieval path is a fixed lore-snippet injection rather than a learned retriever, and the Consistency axis is held constant rather than manipulated. No triad score is therefore reported as an outcome below, and no claim in Section 5 rests on one. Quantifying the Consistency axis — and testing whether it is in fact orthogonal to rather than merely defined to be — requires a retrieval-varying design we identify as the principal open item in Section 7.
4.4. Stage IV — Synthetic Stress Testing
The final stage is the central validation move. A Stress-Tester agent, given only the generated backstory (but not the Judge’s scores), is asked to design a “Session 1” encounter. A separate entity-reuse judge – an independent LLM blind to and – classifies each NPC and hook the Stress-Tester emits as either Reused from the backstory’s Fact Ledger, Invented as new material, or Ambiguous. We define the agent’s Adaptation Effort as:
A low () indicates that the backstory provided sufficient material for the Synthetic DM to staff the encounter; a high () indicates that the agent had to invent substantially. An earlier draft of this formulation used a substring-based hallucination count plus an inference-time penalty; the present semantic formulation discards both, on the basis of a pre-locked-run audit that showed (i) substring matching penalised benign paraphrase of long hook descriptions, and (ii) the time penalty correlated more strongly with generator speed than with story quality.
4.5. Experimental Setup
The reported experiments are interventional ablations rather than the correlational corpus study of an earlier draft. That earlier design — a large single-generator corpus scored for — was withdrawn before the locked runs, for reasons recorded in the deviation log and summarised in Section 5.1: under the semantic of Equation 12 the correlation is indistinguishable from zero, and a correlational target could not have discriminated the metric from the prompt that produced it. What we report instead manipulates the gate directly.
Study 1 is a factorial crossing gating (ungated / gated) with generator, at trials per cell for trials. Random seeds 1–40 are matched across all four cells, so every cell draws the same sequence of injected factions and locations and the arms are paired at the seed level. Both arms use the same Architect of Hooks generation prompt — which mandates a Tier-1 NPC, an open Tier-2 mystery, and an action-oriented closing paragraph — so the only manipulated factor is whether the / gate and its regeneration loop (Section 4.2.0.10) are active. Generation is bound to a compact Forgotten-Realms Story Bible of structured JSON nodes: four factions (Harpers, Zhentarim, Red Wizards of Thay, Lords’ Alliance), three locations (Baldur’s Gate, Candlekeep, Waterdeep), two deities, and an explicit magic-law clause forbidding necromancy in Lords’ Alliance cities. Study 2 (Section 5.5) extends the generator factor to five levels; its setup is otherwise identical and its two Study 1 cells are carried forward unchanged.
Model assignment.
Roles are deliberately asymmetric, and every role but the Architect is frozen across all cells of both studies so that the generator factor is not confounded by the evaluation path (Table 4). The gate itself involves no model: the deterministic Judge of Section 4.2 is arithmetic over the Fact Ledger. Separating the Stress-Tester (Opus) from the entity-reuse Judge (Sonnet) prevents the model that wrote the Session-1 outline from also grading which of its entities were reused, and using an open-weight Notary keeps ledger extraction independent of the frontier arm’s provider.
Recorded measures.
For each trial we record: the deterministic and its four components; the cumulative and its tier; the gate decision, the iteration count and whether the trial converged; the semantic of Equation 12 together with its underlying Reused/Invented/Ambiguous counts; branching outlines for the three fixed player-seed prompts, from which Study 1’s branching count and Study 2’s are derived (see Section 5.5 for description); and per-stage wall-clock time. The full run configuration — weights, thresholds, seeds, the verbatim INSTABILITY_PROMPT, and the git hash of the code that produced it — is frozen to a sidecar file at run start and released with the results.
Pilot Study.
Before either locked study we ran a six-sample feasibility pilot with the open-source reference implementation of the pipeline against a local qwen3:14b model via Ollama, using the same Forgotten-Realms Story Bible. The pilot manipulates a different factor from the locked studies: it contrasts a generic “write a coherent biography” control prompt () against the Architect of Hooks prompt (), and adds one hand-scored calibration character. Its purpose was to check that the pipeline runs end to end and that the two Judge layers of Section 4.2 agree — the deterministic arithmetic of Equation 1 computed over the Fact Ledger, and the LLM-rubric G-Eval formulation — on the same inputs. It is reported as historical context in Section 5.1, not as evidence for any hypothesis: at it is under-powered, and its prompt-contrast design is circular in the way the locked ablations were expressly designed to avoid, since the Architect prompt targets the very dimensions scores.
4.6. Pre-Registered Hypotheses
Both ablations were pre-registered before any locked-run data were collected. The design documents, the amendment log, and the analysis code are released with the reference implementation, so every claim below can be traced to the version of the protocol that was in force when the corresponding run was launched. Amendments were filed only in response to pre-locked-run smoke evaluations whose trials are excluded from the locked datasets, and each is flagged in the log as filed before any outcome statistic of the locked run was computed.
Throughout, a name in typewriter face is a model tag as invoked (qwen3:4b, qwen3:8b, qwen3:14b, llama3.1:8b), while an upper-case suffix names the parameter class that tag belongs to (4B, 8B, 14B), with Llama reserved as the label for the cross-family llama3.1:8b, which is also of the 8B class. So qwen3:14b is the model and 14B is its scale; the two 8B-class generators are qwen3:8b (8B) and llama3.1:8b (Llama), which is what makes the cross-family comparison of Section 5.5 possible.
Study 1 (Section 5.2) manipulates gating (ungated / gated) crossed with generator (qwen3:14b, a commodity 14B-class backbone, vs. Opus 4.7, a frontier backbone) and registers four hypotheses on the gate effect :
- H1a
- — commodity simple effect (primary, directional). : gating on and lowers the fraction of Synthetic-DM entities that must be invented. One-tailed contrast at .
- H1b
- — heterogeneous effect (primary, directional). The gating × generator interaction in the ANOVA on is non-zero, with more negative on qwen3:14b than on Opus 4.7. The gate is predicted to help where the generator is weak, not uniformly.
- H1c
- — frontier bounded null (primary, non-directional). is registered as a bounded null: no directional prediction, but on the scale, reported with its confidence interval. Registering the null in advance makes the frontier arm falsifiable rather than a mere absence of evidence.
- H2
- — branching diversity (secondary, directional). Mean branching diversity is higher in the gated arm: a richer opening state should let the Synthetic DM open more distinct Session-1 trajectories from the same three fixed player seeds.
Study 2 (Section 5.5) extends the generator factor to five levels and asks whether the Study 1 pattern is a property of the gate or of one backbone. It registers directional simple effects on for qwen3:4b, qwen3:8b and the cross-family control llama3.1:8b; a signed-pattern gating × generator interaction across the four qwen3-and-frontier levels — the cross-family control llama3.1:8b is excluded from this contrast precisely because it is not a qwen3 scale point — predicting harm at the smallest scale, a near-zero band at 8B, benefit at 14B and a return to near-zero at the frontier; and, on , that gating concentrates rather than widens the branching distribution, with a pre-registered rule that the concentration effect may be claimed as general only if the cross-family arm replicates it. The qwen3:4b direction was flipped from benefit to harm in a filed amendment after a smoke run exposed a rewrite-cascade failure at that parameter class; the flip is then tested on the locked per cell data (not an assumption).
5. Results
We report results in three layers, derived from staged experimentation and refinement of the models. Section 5.1 retains the six-sample qwen3:14b pilot as historical context. Section 5.2 reports the pre-registered Study 1 (gate ; full pipeline judge-based and the T1-density UCR). Section 5.3 reports diagnostic findings from the locked run. Section 5.5 and Section 5.6 extend the design to a factorial () and apply the nonparametric multiple comparison framework of Derrac et al. [27] to assess cross-generator consistency.
Three sample sizes recur below and denote different things, so we fix them here. Throughout, is the per-cell sample: forty seed-matched trials, the unit at which every simple effect is estimated and the number the pre-registration locks. is Study 1’s total, its four cells (2 gating levels × 2 generators) at forty trials each. is Study 2’s total, its ten cells () at the same forty trials each — 240 newly generated trials plus Study 1’s 160 carried forward unchanged, which is why Study 1 is a strict subset of Study 2 rather than an independent replication. The nonparametric analysis of Section 5.6 re-indexes the same 400 trials as a matrix of ten configurations over forty seed-matched problems, so its counts problems, not trials.
5.1. Pilot Findings
Table 5 summarises the pilot with three observations:
(i) The deterministic and rubric Judges agree. Kaelen Thorne is a calibration fixture: a single backstory written by hand, held fixed, and scored by hand against Equations 2–4 to give an independent target of that neither Judge layer can see. It is not drawn from any corpus and is not one of the six pilot samples; it is released verbatim with the reference implementation so the target can be recomputed. On it, the deterministic Judge returns , within of the hand-scored target, and the LLM-rubric Judge returns , agreeing with the deterministic layer to within . This licences the use of the cheaper deterministic Judge for large-batch screening with the rubric Judge reserved for borderline cases.
(ii) discriminates campaign from one-shot-tier characters correctly. The control sample with the weakest hooks () returned , below the configurable threshold of 150, and was flagged insufficient. The same character carried the highest () of the pilot, confirming that the static tier is a useful pre-screen anticipating downstream Synthetic-DM load.
(iii) Pearson on the pilot (legacy formulation). Under the substring-based used in the pilot, the direction matched the central correlational hypothesis but the magnitude was well below an earlier draft’s projected . Under the semantic, judge-based of Equation 12 the same n=40 smoke yields , indistinguishable from zero, and no weighting scheme over the four components produces . The correlational target is therefore withdrawn as a confirmatory hypothesis. The same data, however, shows the interventional gate effect, motivating the design now reported in Section 5.2.
5.2. Primary Outcome: Heterogeneous Gating Effect on Adaptation Effort
The study of the ”Gating Effect” was produced through a locked run (40 trials per cell, four cells, gate ) completed in 7.9 h of wall time. Cell-level descriptive statistics are reported in Table 6. Two pre-registered simple-effect contrasts plus the ANOVA follow. The three pre-registered hypotheses on stated in Section 4.6 — H1a, H1b and H1c — are adjudicated here across the two generators (qwen3:14b and Opus 4.7), each in a gated and an ungated variant; H2 is taken up in Section 5.3.
H1a – simple effect on the commodity arm (supported).
On the qwen3:14b arm, mean falls from ungated to gated — a relative reduction of in the fraction of DM-emitted entities that the entity-reuse judge classified as fabricated. The one-tailed Welch’s t-test yields , , , Cohen’s (a medium effect). H1a is supported at .
H1b – interaction (directional, underpowered).
A two-way ANOVA on with factors gating and generator (Table 7) finds no significant main effect of gating (, , ), a significant main effect of generator (, , ; discussed in Section 5.3), and a directional but non-significant gating × generator interaction (, , ). The interaction sign matches the pre-registered direction — more negative on qwen3:14b than on Opus 4.7 — but the omnibus test does not reach at per cell. Per the pre-registered power analysis the locked n was sized for a medium effect on the main contrast; the interaction term carries less power. We report the directional result and defer the formal interaction inference to the pre-justified per cell confirmatory follow-up.
H1c – bounded null on the frontier arm (supported as a null).
On the Opus 4.7 arm, mean moves from ungated to gated — a numeric increase that is statistically indistinguishable from zero. Welch’s t-test gives , , two-tailed , with a CI on the difference of . H1c, registered as a bounded null, is supported: the gate-attributable effect on the frontier generator sits inside a window of on the scale.
Interpretation.
The headline interventional finding is the asymmetric pair of simple effects: gating produces a medium-sized reduction in on a commodity 14B-parameter generator, and no detectable change on a frontier generator that already produces low- drafts. The two contrasts are individually pre-registered and statistically backed by the locked run; the omnibus interaction is consistent in direction but underpowered. We accordingly frame the contribution as: for practitioners deploying LudicMetrics on a commodity backbone, the / gate is a cost-effective intervention that meaningfully reduces downstream DM load; deploying it on a frontier backbone does not appear to harm but does not appear to help either, and the practitioner should weigh the additional regeneration cost against an effect bounded inside .
Operational gate-pass behaviour.
A secondary view of the same data: the gate’s operational pass rate (trials that meet and on the final draft) rose from to on the qwen3:14b arm and from to on the Opus arm. The mean iteration count on the gated arms was (qwen3:14b) and (Opus 4.7); 34 of 40 Opus gated trials cleared the gate on the initial draft () without any regeneration. The gate-softening from to is reflected in this iteration distribution: the prior setting drove Opus to a non-monotone rewrite cascade in -iteration trials, an artefact eliminated under the locked setting.
Secondary descriptive.
Pearson pooled across all trials is , with the cross-cell component dominating (within-cell values are smaller). This confirms a modest negative association consistent with ’s theoretical reading, but the magnitude is far below the pilot test’s projected , and the relation is too noisy to support a regression-style claim. The interventional analysis above remains the headline.
5.3. Diagnostic Findings
Two findings emerged from the locked run that were not pre-registered but are worth reporting for what they reveal about and about the framework’s headroom.
Opus is more prolific, not worse at reuse.
The significant main effect of generator on (Table 7) is initially counter-intuitive: the frontier model has higher fabrication ratio than the commodity model. Per-cell entity counts (Table 8) show why. Opus emits roughly entities per Session-1 outline against qwen3:14b’s , and labels both REUSED and INVENTED in higher absolute numbers. Reuse rates per trial are similar; the ratio rises on the Opus arm because the model populates each scene more densely. The headline qualifier for the heterogeneous gate result is therefore not “Opus ignores the ledger more” but “Opus generates more scene content, period, which inflates both numerator and denominator of slightly asymmetrically.” A model-tier-aware reading of would normalise by total emitted; we treat this as a future-work refinement rather than a deviation, since the entity-reuse judge’s classification is the locked operationalisation.
Branching diversity is ceiling-saturated.
H2 (Section 4.6) from the pre-registration — gating raises branching diversity — could not be evaluated as designed. All 160 trials returned a branching-diversity count of 3 (three distinct clusters on the three player-seed prompts). Looking at the underlying pair-wise Jaccard similarity of NPCs and hooks across the three branching outlines, qwen3:14b averaged and Opus 4.7 , both well below the clustering threshold. The three player-seed prompts drive the Synthetic DM to extreme divergence regardless of the backstory or the gate, so the gate cannot move what is already at the ceiling. We treat this as a metric-design limitation: a more discriminating branching measure would either (i) loosen the Jaccard threshold and look at degree of overlap, or (ii) use a continuous diversity statistic (mean pair-wise Jaccard itself) rather than the binarised cluster count. We report this honestly as a null on H2 attributable to the metric, not to the framework.
5.4. Robustness of the Weighting Scheme
A natural objection to Equation 1 is that the weights are chosen rather than learned, and that the framework’s conclusions might be an artefact of that choice. We address this directly with a global sensitivity analysis (GSA) over the locked corpus. A useful structural fact makes the analysis clean: among the pipeline’s outputs, only depends on the weights. and never read , so a weight perturbation can move a backstory across the gate but never across the or axes. The robustness of the dual-axis quadrant assignment therefore reduces exactly to the robustness of the play-ready gate ().
Monte-Carlo classification stability.
We drew weight vectors, perturbing each uniformly by and renormalising to , recomputed for all 160 backstories, and counted how many changed play-ready status relative to the baseline weights. At the canonical the mean reclassification rate is (Table 9); of backstories never change class under any sampled weighting. The rate stays below the conventional robustness bound even at (), i.e. under weight excursions far larger than any defensible re-elicitation. Spearman rank correlation between the baseline and perturbed orderings is at and never falls below across the sweep, so the weighting also preserves the ranking of backstories, not merely the binary gate.
Variance decomposition.
We complement the robustness sweep with a Sobol variance decomposition [28,29], attributing the variance of the play-ready rate (under perturbation) to each weight using the Jansen pick-and-freeze first-order and total-effect estimators. The gate decision is driven by the two substantive high-variance dimensions, and (total-effect indices and ), with strong interactions induced by the simplex constraint; contributes modestly () and is almost inert (). The component most likely to attract a reviewer’s scepticism — , the smallest weight and the “strategic vagueness” term — is precisely the one whose exact value is immaterial to every downstream decision. We read the GSA as evidence that the empirically-derived weights are operationally valid: the framework’s classifications are stable artefacts of the four narrative dimensions, not of a fragile numeric tuning. The analysis script is released with the reference implementation so reviewers can reproduce it on their own corpora.
Confirmation on the full corpus.
As an out-of-sample robustness check we re-ran the identical Monte-Carlo and Sobol procedures on the full analytical corpus (Section 5.5: the locked stories plus the 240 Study 2 backstories spanning the 4B–14B and cross-family generators). The conclusion is unchanged. The mean play-ready reclassification rate is at (vs. on ) and stays below the robustness bound across the whole sweep, reaching only at ; rank stability is if anything higher (Spearman at , throughout). The Sobol attribution preserves its structure — the gate is governed by and ( and ), stays modest (), and remains the least influential dimension (). The marginally higher flip rates are expected: the larger corpus adds the commodity-tier cells — notably the gated qwen3:4b rewrite cascade — whose values cluster nearer the gate, yet even this heavier boundary mass does not breach the robustness bound. The weighting scheme is therefore stable not only on the calibration corpus but across the full breadth of generators studied here.
5.5. Study 2: Expanded Cross-Generator Analysis
Study 1 identified a heterogeneous gate effect across two generators but could not resolve whether the pattern generalises across model-size tiers. Study 2 extends the design to a factorial by adding three generators: qwen3:4b, qwen3:8b, and llama3.1:8b. The two Study 1 arms (qwen3:14b and Opus 4.7) are carried forward at per cell, yielding cells and analytical trials (240 new, 160 carryforward). All cells use the locked pipeline (gate threshold , Plan-C judge, seed-matched backgrounds).
We additionally compute , which replaces the ceiling-saturated branching count of Study 1 (Section 5.3).
Branching diversity ().
Informally, asks a single question: when the Synthetic DM opens the same backstory three different ways, how different are those three openings from each other? Every trial produces three Session-1 outlines, one per fixed player-seed prompt. We compare each of the three pairs of outlines on two channels, the set of NPC names it introduces and the set of plot hooks it offers, and measure the overlap of each channel with the Jaccard index . Averaging those six overlaps and subtracting from one gives
so means the three openings share no NPC and no hook (maximally divergent play), means they are the same opening three times. It is computed post hoc from the per-pair overlaps already recorded in the run log, so it requires no additional model calls and is applied identically to the Study 1 carryforward cells. The virtue over Study 1’s count is that it is continuous: the old measure binarised each pair at a fixed similarity threshold and then counted clusters, which saturated at its maximum of 3 in almost every cell and so could not register the gate’s effect (Section 5.3).
Table 10 reports cell-level means.
Why the ANOVA is and not .
Table 10 is descriptive and reports the full design. The inferential test that follows is deliberately narrower. Its generator factor is restricted to the four levels that form a single capability ladder — qwen3:4b, qwen3:8b, qwen3:14b, Opus 4.7 — giving cells and of the 400 trials, hence the 312 residual degrees of freedom below. The fifth generator, llama3.1:8b, is held out exactly as pre-registered (Section 4.6) because it is the cross-family control at the 8B class, not a rung on the ladder: entering it would place two different architectures on what is meant to be one axis of scale and confound the very size gradient the interaction term is testing. It is not discarded — it is analysed below as a matched contrast against qwen3:8b at equal parameter count, and it enters the assumption-free ranking of Section 5.6, where all ten cells participate.
What the ANOVA shows.
One convention fixes the direction of every number that follows: is a fabrication ratio, so lower is better, and is the gated mean minus the ungated mean, so a negative means the gate helped. The ANOVA on then returns three results that must be read together:
- 1.
- Switching the gate on has no overall effect (, , ). Averaged over all four backbones, gating leaves the Synthetic DM’s fabrication load essentially where it was.
- 2.
- The choice of backbone matters a great deal (, , ). Which generator wrote the backstory shifts regardless of what the gate is doing.
- 3.
- The gate’s effect depends on which backbone it is applied to (, , ). This interaction is the substantive finding, and it is what makes the flat average in (1) misleading: the average is flat not because the gate is inert, but because the arms disagree in sign.
The per-arm values show what that average conceals: at 4B, at 8B, at 14B and at Opus 4.7. Only the 14B arm is negative. Stated plainly: the same gate that helps a mid-sized generator hurts a small one. Below roughly the 14B class, forcing a model to keep rewriting until it clears the / bar leaves a backstory that is harder for the Synthetic DM to run, not easier — the rewrites add material the model cannot anchor in the Fact Ledger. At 14B the same rewrite loop pays off. At the frontier the gate is close to a no-op, for a simple operational reason: Opus 4.7 clears the bar on its first draft in 34 of 40 gated trials (mean iterations), so the gate rarely fires and has little left to change. The practical boundary between “the gate helps” and “the gate hurts” therefore falls between the 8B and 14B classes — which is the capability threshold this paper argues practitioners must locate before deploying the gate at all.
Two distinct ways the gate can fail.
The two sub-threshold arms fail for different reasons, and separating them is what turns the interaction from a curiosity into actionable guidance.
The qwen3:4b arm fails loudly. All 40 gated trials run the full rewrites and only 2 ever clear the gate: the model cannot satisfy the structural rewrite instruction without stripping structure out of the phases it had already written, so each pass returns a near-identical minimal ledger. This is a rewrite cascade, and it is trivially detectable in production — a cell where nothing converges announces itself.
The qwen3:8b arm fails quietly, which is the more dangerous case. It converges perfectly well (, mean iterations) and its and rise as intended, yet its gets worse (, the largest harm in the study). The entity counts say why: gating raises the model’s total entity emission by per trial, but the growth is in Invented entities () rather than Reused ones (). The rewrites do add narrative structure — enough to satisfy the gate — but that structure is not anchored in the Story Bible, so every new hook is one more thing the Synthetic DM must fabricate at the table. By the framework’s own headline numbers this cell looks like a success; only reveals that the gate has moved work downstream rather than removing it. That is precisely the case for measuring the downstream cost at all.
The cross-family contrast at equal size.
Comparing llama3.1:8b against qwen3:8b isolates architecture from scale, since the two are the same parameter class, run at the same temperature through the same pipeline. They are also indistinguishable on the operational metrics: both converge at (mean vs. iterations), and their ungated baselines do not differ significantly (, ). What differs is the kind of edit each makes when told to rewrite. Gating moves llama3.1:8b in the opposite direction to qwen3:8b: Reused entities rise by per trial while Invented entities fall by , the same lore-anchoring signature seen at 14B ( / ) and, if anything, stronger. Two models of the same size therefore sit on opposite sides of the benefit–harm boundary. The capability threshold is consequently not a parameter count that can be looked up in a model card: it must be measured per backbone, which is the practical argument for shipping alongside and rather than gating on and alone.
5.6. Nonparametric Multiple Comparison Analysis
To assess the cross-generator differences without distributional assumptions, we apply the nonparametric multiple comparison framework of Derrac et al. [27]. Each of generator–gating configurations is treated as an “algorithm” and each of seed-matched backgrounds as a “problem,” yielding a complete rank matrix per metric. We report Friedman, Iman–Davenport, Friedman Aligned Ranks, and Quade omnibus tests, followed by Nemenyi, Holm, and Shaffer step-down post-hoc procedures.
Omnibus tests.
Table 11 shows that all four metrics exhibit highly significant heterogeneity across the ten configurations ( on all sixteen tests). The Nemenyi critical difference at is average-rank units.
Adaptation Effort ().
Table 12 reports Friedman average ranks for (rank 1 = lowest fabrication ratio). The top-ranked configurations are G-Llama (), G-14B (), and U-8B (); the bottom are U-4B () and G-4B (). The Shaffer-corrected post-hoc procedure identifies four significant pairs (of 45): G-4B vs. G-Llama (), U-4B vs. G-Llama (), G-4B vs. G-14B (), and G-4B vs. U-8B (). All four involve the 4B tier separated from the top clique, confirming that the rewrite-cascade pathology produces a statistically detectable quality penalty.
Ludic Potential Index, Potential Conflict Value, and branching diversity.
Table 13 summarises the post-hoc landscape for the remaining three metrics. shows the strongest differentiation (24 of 45 significant Shaffer pairs), with a top clique of five gated-or-frontier configurations (G-Opus, G-Llama, U-Opus, G-8B, G-14B; ranks –) and G-4B () significantly separated from every other configuration except U-4B. mirrors in structure (16 significant pairs), with G-Llama ranking first (, ) and G-4B last (, ), below the operational sufficiency threshold of 150. exhibits a qualitatively different pattern (11 significant pairs): U-Opus alone occupies rank 1 (), and every gated configuration ranks worse than its ungated counterpart — the gate systematically narrows branching diversity across all five generator families.
Interpretation.
The nonparametric analysis is a robustness check on Section 5.5, and it returns the same ordering without the normality and homoscedasticity assumptions the ANOVA rests on: the capability threshold is a property of the data, not of the test. Its distinctive contribution is the gradient of differentiation across the four metrics. separates 24 of 45 pairs, 16, 11, and only 4. That ordering runs from the quantity the gate directly optimises to the downstream quantity the gate cannot see, and we read it as the honest boundary of the intervention: gating reshapes the metrics it selects on strongly and moves the downstream Synthetic-DM load weakly. A framework that reported only separation would substantially overstate what the gate buys.
Two consequences follow for practice. First, the result is a floor claim rather than a ranking. All four significant pairs place a 4B-tier configuration against a member of the top clique, and no pair among the top six configurations separates at all — which is expected, since the Nemenyi critical difference of average-rank units is wide against an rank spread of only ( to ). The defensible recommendation is therefore negative and specific: do not deploy the gate below the capability threshold, where it produces a statistically detectable penalty rather than merely no benefit. Choosing among generators above that threshold is not something these data support.
We state the corollary explicitly, because the tables invite the opposite reading. G-Llama heads the ranking (Table 12) and the ranking (Table 13), and it would be easy to conclude from those two rows that gated llama3.1:8b is the best configuration in the study. It is not a conclusion these data license, for three reasons. (i) Its lead is not statistically separable: none of G-Llama’s gaps to G-14B (), U-8B () or the other top-clique members approaches the critical difference, so rank 1 here is a point estimate on a 40-problem rank matrix, not a demonstrated advantage. (ii) The comparison is confounded by design. Every configuration was run on one Story Bible, one Notary, one Synthetic DM and three fixed player seeds; llama3.1:8b entered the design as a single-point cross-family control, so we cannot separate “this architecture is better” from “this architecture happens to suit this corpus,” and no llama-family scale gradient exists in these data to arbitrate. (iii) The ranking reverses on the fourth metric: the same G-Llama cell places last of ten on (), so a reader selecting on alone would be selecting the narrowest branching distribution in the study. What G-Llama does support is the weaker and more interesting claim already made above — that at a fixed 8B parameter count, architecture decides whether gating helps or harms — and it earns that role as one half of a matched pair, not as a recommended model. A ranking claim would need multiple bibles, several llama-family scale points, and cell sizes powered for the top-clique gaps rather than for the 4B separation.
Second, the column is the one fully systematic result in the table: every gated configuration ranks below its ungated counterpart, in all five generator families, with no exception. Unlike the benefit — which is conditional on the backbone — the diversity cost of gating is unconditional. A practitioner adopting the gate should expect to pay it in every deployment, and should size it against the playability gain deliberately rather than discovering it downstream; Section 5.12 takes up that trade-off.
5.7. Human Construct-Validity Study
The results above establish that , and behave consistently and discriminate between generators, but consistency is not the same as validity: an internally coherent metric can still fail to measure what a human practitioner cares about. Following the pre-registered protocol released with the reference implementation, we ran a construct-validity study against expert human raters, structured as a convergent/discriminant multitrait–multimethod design [30,31] and reported in line with best-practice guidelines for the human evaluation of generated text [32]. We report the first completed wave below—a fully powered convergent test with a deliberately conservative reliability floor—and are candid about which pre-registered thresholds it clears and which await a deeper rater panel.
Design and panel.
We drew a blinded, metric-stratified sample of backstories from the full corpus and served each rater only the prose and an opaque item identifier—never a cell label or metric value. Raters scored four plain-language 1–7 Likert items: openness (Q1), conflict potential (Q2), preparation burden (Q3), and overall playability (Q4), targeting , , and (for Q4) the composite play-readiness that operationalises. The sample size was fixed a priori to give power to detect a medium human–metric correlation (, one-tailed ) [33]. A balanced-assignment scheme gave every backstory at least three independent raters. The panel comprised eleven practitioners with tabletop Game-Master, narrative-design or game-studies backgrounds; fourteen of the sixty items received four raters and the remainder three, for 194 ratings in total.
Convergent and discriminant validity.
Table 14 reports, for each rubric item, the Pearson correlation of the per-item mean human rating against its target metric, a bias-corrected bootstrap confidence interval ( resamples), the one-tailed p-value after Holm–Bonferroni correction across the four convergent hypotheses, and the single-rater intraclass correlation [34,35]. The central hypothesis—that the composite index reflects how playable an expert judges a backstory to be—is clearly supported. Overall playability (Q4) correlates with at ( CI , Holm ), the strongest of the four relationships and the only one to survive multiplicity correction, and its discriminant clause holds cleanly: Q4 loads on () far more than on any off-target metric (next largest ), the one item to pass the MTMM monotrait–heterotrait criterion. Preparation burden (Q3) tracks in the predicted direction with a comparable point estimate (, one-tailed ); it is significant before correction and falls just short after Holm adjustment, a near-miss we expect to resolve with additional raters. Openness (Q1) is weakly positive (, n.s.), and conflict potential (Q2) is effectively null against () while cross-loading onto ()—a discriminant failure we return to below.
Inter-rater reliability.
Agreement is the metric on which this first wave is deliberately conservative [36]. With most items rated by the three-rater minimum, no item yet reaches the pre-registered acceptance gate; the highest are openness () and overall (), with conflict and burden near zero, and Krippendorff’s concurs [37]. Two factors are at work, and both point the same way. First, expert judgements of playability genuinely diverge at the level of an individual backstory—unsurprising given how much table-level context a Game-Master supplies—so a low ceiling is expected at this rater depth and should lift as the panel is deepened. Second, that same divergence attenuates the convergent correlations, which means the estimates in Table 14 are conservative lower bounds: a significant –playability relationship emerging despite the reliability floor is strong rather than weak evidence, and the per-facet near-misses are the more likely to clear their thresholds as agreement improves.
Expert elicitation of the weights.
The same panel completed an Analytic Hierarchy Process elicitation [38] of the four sub-weights, via pairwise comparisons screened for consistency (). Nine experts produced consistent judgement matrices; their row-geometric-mean group vector is , for , , , and respectively. This preserves as the dominant dimension but does not reproduce the framework’s default ordering: experts down-weight motivation (, vs. ) and, most notably, up-weight thematic vagueness (, vs. ), so the pre-registered corroboration rule (rank order preserved and every ) is not met. We report this discrepancy rather than silently re-fitting the weights. Crucially it does not disturb the study conclusions: the global sensitivity analysis of Section 5.4 shows the play-ready classification is robust to weight perturbations far larger than this ( reclassification at ), and —the dimension the experts most disagree with—is precisely the one the Sobol analysis finds least influential. The elicitation is therefore best read as a calibration signal for a future weighting revision, not as a challenge to the present results.
5.8. Discussion: The Agency–Coherence Paradox
The pattern above is a computational restatement of Aylett’s Narrative Paradox [13]. A perfectly consistent backstory — the ideal output of a Symbolic-AI planner — is narratively sterile, because consistency is achieved by closing loops, and play requires loops to remain open. Conversely, a maximally open backstory is incoherent and unusable. The contribution of this paper is to make the trade-off measurable and therefore tunable: by reporting and Consistency as orthogonal axes rather than collapsing them into a single quality score, we let designers sit deliberately near the sweet spot rather than accidentally on either extreme.
5.9. Discussion: Computational Efficiency
The Fact Ledger is not only a validation artefact but a context compressor. Distilling thousand-word prose into a few hundred bytes of JSON reduces the Judge’s effective context length by roughly an order of magnitude, and we observe a corresponding reduction in Judge inference cost compared to evaluating raw prose. For long-running campaigns or live-service deployments (Section 6), we propose hierarchical semantic summarisation: events with low cumulative may be archived into phase-level summaries while high- events remain active in the Ledger.
5.10. Discussion: The Pro-Conflict Bias
A limitation we explicitly acknowledge is that , as it is currently defined, favours confrontational narrative. The Immediacy and Gravity dimensions tilt toward conflict-driven scenes, and the weights we derive empirically from a D&D-style Stress-Tester reflect that genre’s design conventions. In “cosy” genres — life simulators such as Animal Crossing or Stardew Valley — a high- backstory would be actively unwelcome. We expect the framework to remain useful in such genres after re-derivation of the weights: the four dimensions can be re-interpreted (Immediacy as “opportunity for social interaction this week”, Gravity as “personal-stakes intensity”), but the empirical weights must be recalibrated against a Stress-Tester whose encounter style matches the target genre. We treat this as a limitation in the present paper and a clear avenue for future work.
5.11. Discussion: Validity of Synthetic Stress Testing
Using LLM agents to simulate Game Masters raises a fair concern about realism: LLMs lack the social intuition human DMs bring to encounter design. We argue, however, that functions as a lower bound on playability. If even a competent LLM agent cannot extract a Session 1 encounter from a backstory without heavy invention, it is unlikely that a human DM could do so without similarly heavy authorial effort — and the entire premise of procedural-generation tooling is to reduce that effort. The synthetic test is therefore a useful gate, not a substitute for human evaluation. Because automatic text-quality metrics are known to correlate only weakly with human judgement in general [39,40], we treated an explicit construct-validity study against human raters as a prerequisite for the metrics’ wider adoption rather than an optional extra; its first wave (Section 5.7) confirms that tracks expert playability judgements, while ’s convergent estimate points the same direction but does not yet reach significance under the current agreement ceiling.
5.12. Discussion: Cross-Generator Consistency and Capability Thresholds
The nonparametric multiple comparison analysis (Section 5.6) converges with the parametric ANOVA (Section 5.5) on three structural findings.
The capability threshold is robust across methods.
The ANOVA interaction () identifies a sign-flip in between 8B and 14B parameters. The Friedman rank ordering confirms this without distributional assumptions: gated configurations of capable generators (G-Llama at rank , G-14B at ) occupy the top clique, while gated configurations of sub-threshold generators (G-8B at , G-4B at ) occupy the bottom. The four Shaffer-significant pairs all involve 4B-tier configurations separated from the top clique, confirming that the rewrite-cascade pathology produces not merely a null effect but a statistically detectable quality penalty.
Architecture, not only scale, determines gate benefit.
The two 8B-parameter generators straddle the capability boundary: G-Llama ranks 1st on () while G-8B (qwen3:8b) ranks 8th (); the median contrast between them is in favour of G-Llama. Their ungated baselines do not differ significantly (, ), so the divergence reflects how each architecture handles iterative rewriting: llama3.1:8b anchors new structure to existing lore (REUSED entities per trial under gating), whereas qwen3:8b inflates total entity emission while increasing fabrication (INVENTED ). This cross-family result implies that the gating benefit threshold is architecture-dependent and cannot be predicted from parameter count alone.
The diversity–playability trade-off.
Gating improves and rankings while worsening rankings for all five generator families. The trade-off is asymmetric: for capable generators, the playability gain outweighs the diversity loss (G-Llama gains average-rank positions on relative to U-Llama while losing on ). For sub-threshold generators the trade-off is strictly worse: G-4B loses positions on and on simultaneously — a pure penalty with no playability gain. This asymmetry reinforces the practical recommendation: LPI/PCV gating should be deployed conditionally, above the capability threshold.
6. Practical Use Cases
The framework is intended for production deployment as well as research. We describe four use cases, each of which we have begun to prototype.
6.1. Automated Quest Design in Open-World RPGs
In large open-world titles, “filler” side-quests that feel disconnected from the protagonist are a persistent quality problem. A procedural quest generator parameterised by and can spawn quests that are personally relevant to a generated NPC’s history rather than randomly attached. When the system detects a Tier-1 active enemy hook, it instantiates a Bounty or Ambush template populated from that NPC’s Fact Ledger, ensuring that every procedural encounter feels earned.
6.2. Narrative Continuous Integration
As branching narratives grow, the risk of logical rot — later choices contradicting earlier facts — increases super-linearly. Embedding the Notary and Judge into a CI/CD pipeline lets a writing team run the same regression tests on narrative content that they run on code: every new dialog branch is extracted into the Ledger and diffed against the Background Event Log, and any contradiction surfaces as a failed test before shipping. Monitoring across paths additionally identifies dead-end branches where too many loops have been closed.
6.3. Dynamic Lore Persistence in Live-Service Games
Live-service games whose world evolves across seasons must keep NPC dialog consistent with the current world state. With the Story Bible serving as a versioned RAG store, TruLens groundedness checks on generated NPC dialog would prevent the model from referring to a city destroyed last season as a current trade hub — a class of bug that has historically required handwritten guard rails. The Ledger gives game-state systems a single authoritative source from which both narrative and gameplay logic can read.
6.4. Quality Assurance for User-Generated Content
Platforms that let players publish their own adventures (Roblox, Fortnite Creative, D&D digital toolsets) can adopt as a gating metric. A submission with cumulative is flagged to the author with concrete suggestions: “Your character has no active enemies; consider adding an unresolved debt to raise playability.” This converts a subjective design heuristic into actionable, automated feedback at the moment of authoring.
7. Conclusions and Future Work
This paper has formalised a methodology for evaluating the ludic utility of LLM-generated narrative. By separating Consistency (enforced by State-Tracking Extraction and grounded RAG) from Playability (quantified by the Ludic Potential Index and the Potential Conflict Value), we have moved narrative evaluation from the subjective to the deterministic without collapsing the agency–coherence trade-off that makes interactive storytelling hard. A three-agent pipeline, specified so that it can be driven from standard evaluation harnesses such as DeepEval and TruLens, operationalises the metrics, and a Synthetic-DM stress test provides an independent inverse validator. We have argued that strategic vagueness is a functional requirement, not a defect, of generated narrative, and have shown how the framework formalises a neuro-symbolic bridge — a symbolic Fact Ledger over connectionist prose — that the Symbolic-AI storytelling tradition has wanted for forty years. Several directions remain open.
Principled weight calibration.
The default weights in Equations 1 and 7 were derived from a single Synthetic DM, and although Section 5.4 shows the framework’s decisions are robust to their perturbation, deriving them from an external criterion would strengthen the metric’s standing. Three complementary routes are available. (i) Expert elicitation: an Analytic Hierarchy Process [38] in which professional game masters and narrative designers fill a pairwise-comparison matrix over the four dimensions yields weights as its principal eigenvector, with a consistency index certifying the judgements are non-random; we release a pre-registered elicitation instrument (a blinded pairwise-comparison form with automatic consistency-ratio screening and group geometric-mean synthesis) alongside the reference implementation to support this study. (ii) Data-driven inversion: given a corpus labelled with an external target (human-rated preparation effort, or the measured ), a sum-constrained Lasso or ridge regression [41] recovers the weights that best predict the target and flags any redundant dimension whose coefficient shrinks to zero; a genetic algorithm offers the same with a non-differentiable concordance objective. (iii) Psychometric calibration: treating the four sub-metrics as indicators of a latent “ludic potential” construct, a confirmatory factor analysis supplies standardised factor loadings that double as corrected weights and tests whether the four dimensions load on a single factor. Updating the weights online from the hooks that human players actually engage with during play would further adapt the framework to genre, group, and individual taste.
Genre re-calibration.
As discussed above, the Pro-Conflict Bias limits applicability outside conflict-driven genres. A controlled re-derivation of ’s dimensions for cosy and exploration games is the natural next study.
Human-DM evaluation.
The first wave of the pre-registered construct-validity study is reported in Section 5.7: it confirms that overall human playability judgements track (, Holm ) but leaves the per-facet convergent hypotheses and inter-rater reliability below their acceptance thresholds. Two extensions follow directly. First, deepening the panel—more raters per backstory—should lift above the gate and sharpen the attenuated convergent correlations, which the present agreement ceiling depresses toward conservative lower bounds. Second, the persistent null and -cross-loading of the conflict item (Q2) suggests that , as a cumulative confrontation budget, lacks a clean single-item human analogue; a re-worded or multi-item conflict instrument is the natural refinement. Closing both gaps would fully discharge the circularity concern left by the LLM-only Synthetic-DM stress test.
Game-engine integration.
Exporting the Fact Ledger directly to Unity or Unreal event systems would close the loop between narrative generation and gameplay logic — the long-stated goal of procedural narrative tooling, and one that the neuro-symbolic Ledger makes finally tractable.
Author Contributions
Conceptualization, L.P.; investigation and state-of-the-art review, R.B. and L.P.; methodology, L.P. and J.M.P.; software, L.P. and J.M.P.; formal analysis and computational validation, L.P. and J.M.P.; validation (expert study), R.B. and L.P.; data curation, L.P. and J.M.P.; writing—original draft preparation, L.P.; writing—review and editing, L.P., J.M.P. and R.B.; supervision, J.M.P. All authors have read and agreed to the published version of the manuscript.
Funding
This work was funded in part by the University of Design, Innovation, and Technology (UDIT) (Crossref Funder ID: 100032921) under the grants INC-UDIT-2027APC02.
Institutional Review Board Statement
This study is a technical evaluation of a narrative-generation metric. It is not a clinical study and reports no clinical-trial results. The expert-rating activity was conducted with adult volunteer practitioners who took part on an informed, voluntary basis; each rater was identified only by an unguessable pseudonymous token, and no personally identifiable information was collected or stored at any point. Under the policies of UDIT for minimal-risk research with anonymous data, ethical review was not required.
Informed Consent Statement
Verbal informed consent was obtained from all participants before they were given access to the rating instrument. Verbal consent was obtained rather than written because the exercise was an anonymous, non-interventional expert-rating task that collected no personal data: participants were identified only by an unguessable random access token, and a signed consent form would have required a handwritten name, which would have been the only personally identifying item anywhere in the study and would have defeated the anonymity the design was built to preserve. Participants were adult volunteer practitioners recruited individually, and each was told, before agreeing, the purpose of the study, what taking part involved and its approximate duration, that participation was voluntary and could be discontinued at any time without giving a reason and without consequence, what data would be stored, and that the resulting ratings and comments would be published as an open dataset. A copy of the consent script has been provided to the editorial office. Verbal consent is one of the three methods of obtaining consent contemplated by the ethics-application form of UDIT (Annex 1 of the regulation cited in the Institutional Review Board Statement).
Data Availability Statement
The data presented in this study are openly available in Zenodo at https://doi.org/10.5281/zenodo.22639369. The deposit is a minimal replication dataset comprising the trial-level analytical data for both studies (N = 400), the 60 human-rated backstories, the 194 pseudonymous expert ratings, the AHP pairwise judgements, the frozen run configurations and the weight-sensitivity outputs, together with a self-contained verification script that recomputes every statistic reported in this article. The reference implementation of the LPI and PCV scorers is publicly available at https://github.com/Luis-Pena-Udit/ludicmetrics (accessed on 7 September 2026). The complete per-trial generation logs (approximately 13 MB of model-generated prose), which are not required to reproduce any reported result, are available from the corresponding author on request.
Acknowledgments
The authors thank the expert practitioners who served as raters in the construct-validity study for their time and judgement.
Conflicts of Interest
Author Jose Maria Peña was owner of the company Lurtis Ltd. The remaining authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest. The company had no role in the design of the study, in the collection, analysis or interpretation of data, in the writing of the manuscript, or in the decision to publish the results.
References
- Juul, J. Half-real: Video games between real rules and fictional worlds; MIT press, 2011.
- Gu, J.; Jiang, X.; Shi, Z.; Tan, H.; Zhai, X.; Xu, C.; Li, W.; Shen, Y.; Ma, S.; Liu, H.; et al. A Survey on LLM-as-a-Judge. arXiv preprint arXiv:2411.15594, 2024, [arXiv:cs.CL/2411.15594].
- Liu, Y.; Iter, D.; Xu, Y.; Wang, S.; Xu, R.; Zhu, C. G-Eval: NLG Evaluation Using GPT-4 with Better Human Alignment. In Proceedings of the Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics, 2023, pp. 2511–2522.
- Chhun, C.; Suchanek, F.M.; Clavel, C. Do Language Models Enjoy Their Own Stories? Prompting Large Language Models for Automatic Story Evaluation. Transactions of the Association for Computational Linguistics (TACL) 2024, 12, 1122–1142. [CrossRef]
- Meehan, J.R. TALE-SPIN, an Interactive Program that Writes Stories. In Proceedings of the Proceedings of the 5th International Joint Conference on Artificial Intelligence (IJCAI), 1977, Vol. 1, pp. 91–98.
- Lebowitz, M. Planning Stories. In Proceedings of the Proceedings of the 9th Annual Conference of the Cognitive Science Society, 1987, pp. 234–242.
- Young, R.M.; Ware, S.G.; Cassell, B.A.; Robertson, J. Plans and Planning in Narrative Generation: A Review of Plan-Based Approaches to the Generation of Story, Discourse and Interactivity in Narratives. Sprache und Datenverarbeitung, Special Issue on Formal and Computational Models of Narrative 2013, 37, 41–64.
- Mateas, M.; Stern, A. Façade: An Experiment in Building a Fully-Realized Interactive Drama. In Proceedings of the Game Developers Conference (GDC), Game Design Track, San Jose, CA, USA, 2003.
- Kelly, J.; Mateas, M.; Wardrip-Fruin, N. There and Back Again: Extracting Formal Domains for Controllable Neurosymbolic Story Authoring. In Proceedings of the Proceedings of the 19th AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment (AIIDE), 2023.
- Es, S.; James, J.; Espinosa-Anke, L.; Schockaert, S. RAGAs: Automated Evaluation of Retrieval Augmented Generation. In Proceedings of the Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations (EACL Demo). Association for Computational Linguistics, 2024, pp. 150–158.
- Saad-Falcon, J.; Khattab, O.; Potts, C.; Zaharia, M. ARES: An Automated Evaluation Framework for Retrieval-Augmented Generation Systems. In Proceedings of the Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics (NAACL). Association for Computational Linguistics, 2024.
- Edwards, R. GNS and Other Matters of Role-playing Theory. The Forge (Indie RPGs), online essay, 2001. Accessed 2026-05-21.
- Aylett, R. Narrative in Virtual Environments – Towards Emergent Narrative. In Proceedings of the Working Notes of the AAAI Fall Symposium on Narrative Intelligence (Technical Report FS-99-01). AAAI Press, 1999, pp. 83–86.
- Callison-Burch, C.; Tomar, G.S.; Martin, L.J.; Ippolito, D.; Bailis, S.; Reitter, D. Dungeons and Dragons as a Dialog Challenge for Artificial Intelligence. In Proceedings of the Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics, 2022, pp. 9379–9393.
- Zhu, A.; Aggarwal, K.; Feng, A.; Martin, L.J.; Callison-Burch, C. FIREBALL: A Dataset of Dungeons and Dragons Actual-Play with Structured Game State Information. In Proceedings of the Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL). Association for Computational Linguistics, 2023, pp. 4171–4193.
- Zhu, A.; Martin, L.J.; Head, A.; Callison-Burch, C. CALYPSO: LLMs as Dungeon Masters’ Assistants. In Proceedings of the Proceedings of the 19th AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment (AIIDE), 2023, [arXiv:cs.CL/2308.07540].
- Acharya, D.; Pham, A.; Klein, D.; Saar-Tsechansky, M. You Have Thirteen Hours in Which to Solve the Labyrinth: Enhancing AI Game Masters with Function Calling. arXiv preprint arXiv:2409.06949, 2024, [arXiv:cs.AI/2409.06949].
- Buongiorno, S.; Klinkert, L.J.; Chawla, T.; Zhuang, Z.; Clark, C. PANGeA: Procedural Artificial Narrative Using Generative AI for Turn-Based Role-Playing Video Games. In Proceedings of the Proceedings of the 20th AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment (AIIDE), 2024, [arXiv:cs.AI/2404.19721].
- Gallotta, R.; Todd, G.; Zammit, M.; Earle, S.; Liapis, A.; Togelius, J.; Yannakakis, G.N. Large Language Models and Games: A Survey and Roadmap. IEEE Transactions on Games 2024, [arXiv:cs.AI/2402.18659].
- Park, J.S.; O’Brien, J.C.; Cai, C.J.; Morris, M.R.; Liang, P.; Bernstein, M.S. Generative Agents: Interactive Simulacra of Human Behavior. In Proceedings of the Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST), 2023, [2304.03442]. [CrossRef]
- Peng, X.; Quaye, J.; Rao, S.; Xu, W.; Botchway, P.; Brockett, C.; Jojic, N.; DesGarennes, G.; Lobb, K.; Xu, M.; et al. Player-Driven Emergence in LLM-Driven Game Narrative. arXiv preprint arXiv:2404.17027, 2024, [arXiv:cs.AI/2404.17027].
- Wang, Y.; Kreminski, M. Can LLMs Generate Good Stories? Insights and Challenges from a Narrative Planning Perspective. arXiv preprint arXiv:2506.10161, 2025, [arXiv:cs.CL/2506.10161].
- Cover, J.G. The Creation of Narrative in Tabletop Role-Playing Games; McFarland, 2010.
- Tychsen, A.; Hitchens, M.; Brolund, T.; Kavakli, M. Live Action Role-Playing Games: Control, Communication, Storytelling, and MMORPG Similarities. Games and Culture 2006, 1, 252–275. [CrossRef]
- Merilainen, M. The Self-Perceived Effects of the Role-Playing Hobby on Personal Development – A Survey Report. International Journal of Role-Playing 2012, 3, 49–68.
- Aarseth, E. A Narrative Theory of Games. In Proceedings of the Proceedings of the International Conference on the Foundations of Digital Games (FDG), 2012, pp. 129–133. [CrossRef]
- Derrac, J.; García, S.; Molina, D.; Herrera, F. A Practical Tutorial on the Use of Nonparametric Statistical Tests as a Methodology for Comparing Evolutionary and Swarm Intelligence Algorithms. Swarm and Evolutionary Computation 2011, 1, 3–18. [CrossRef]
- Sobol’, I.M. Global Sensitivity Indices for Nonlinear Mathematical Models and Their Monte Carlo Estimates. Mathematics and Computers in Simulation 2001, 55, 271–280. [CrossRef]
- Saltelli, A.; Annoni, P.; Azzini, I.; Campolongo, F.; Ratto, M.; Tarantola, S. Variance Based Sensitivity Analysis of Model Output. Design and Estimator for the Total Sensitivity Index. Computer Physics Communications 2010, 181, 259–270. [CrossRef]
- Cronbach, L.J.; Meehl, P.E. Construct Validity in Psychological Tests. Psychological Bulletin 1955, 52, 281–302. [CrossRef]
- Campbell, D.T.; Fiske, D.W. Convergent and Discriminant Validation by the Multitrait-Multimethod Matrix. Psychological Bulletin 1959, 56, 81–105. [CrossRef]
- van der Lee, C.; Gatt, A.; van Miltenburg, E.; Krahmer, E. Human Evaluation of Automatically Generated Text: Current Trends and Best Practice Guidelines. Computer Speech & Language 2021, 67, 101151. [CrossRef]
- Cohen, J. A Power Primer. Psychological Bulletin 1992, 112, 155–159. [CrossRef]
- Shrout, P.E.; Fleiss, J.L. Intraclass Correlations: Uses in Assessing Rater Reliability. Psychological Bulletin 1979, 86, 420–428. [CrossRef]
- Koo, T.K.; Li, M.Y. A Guideline of Selecting and Reporting Intraclass Correlation Coefficients for Reliability Research. Journal of Chiropractic Medicine 2016, 15, 155–163. [CrossRef]
- Hallgren, K.A. Computing Inter-Rater Reliability for Observational Data: An Overview and Tutorial. Tutorials in Quantitative Methods for Psychology 2012, 8, 23–34. [CrossRef]
- Hayes, A.F.; Krippendorff, K. Answering the Call for a Standard Reliability Measure for Coding Data. Communication Methods and Measures 2007, 1, 77–89. [CrossRef]
- Saaty, T.L. Decision Making with the Analytic Hierarchy Process. International Journal of Services Sciences 2008, 1, 83–98. [CrossRef]
- Novikova, J.; Dušek, O.; Cercas Curry, A.; Rieser, V. Why We Need New Evaluation Metrics for NLG. In Proceedings of the Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing (EMNLP), Copenhagen, Denmark, 2017; pp. 2241–2252. [CrossRef]
- Belz, A.; Reiter, E. Comparing Automatic and Human Evaluation of NLG Systems. In Proceedings of the Proceedings of the 11th Conference of the European Chapter of the Association for Computational Linguistics (EACL), Trento, Italy, 2006; pp. 313–320.
- Tibshirani, R. Regression Shrinkage and Selection via the Lasso. Journal of the Royal Statistical Society: Series B (Methodological) 1996, 58, 267–288. [CrossRef]
Figure 1.
The design space. Dashed lines mark the gate thresholds (, ). Only the upper-right quadrant produces backstories that are both narratively rich and immediately playable.
Figure 1.
The design space. Dashed lines mark the gate thresholds (, ). Only the upper-right quadrant produces backstories that are both narratively rich and immediately playable.

Figure 2.
The four-stage LudicMetrics pipeline with the asymmetric model assignments. The Architect generates a backstory; the Notary (qwen3:14b, held constant) extracts a Fact Ledger; the deterministic Judge computes LPI and PCV and fires the gate. The deterministic Judge invokes no language model at all — it is plain arithmetic over the Fact Ledger’s fields (Equations 1–7), which is what makes the gate decision reproducible from the released ledgers. On acceptance, the Stress-Tester (claude-opus-4-7) drafts a Session-1 outline that a separate entity-reuse Judge (claude-sonnet-4-6, blind to LPI and PCV) labels, producing the semantic Adaptation Effort of Equation 12. The dashed back-edge is the gated arm’s regeneration loop: on rejection the Architect is re-prompted with an INSTABILITY_PROMPT, a fixed structural-rewrite instruction (Section 4.2, regeneration loop) that demands to rewrite partially the backstory, explicitly adding certain potentially missing elements (one named living NPC, one unresolved Tier-1 hook, and one new proactive motivation), for at most three iterations. The ungated arm omits this loop.
Figure 2.
The four-stage LudicMetrics pipeline with the asymmetric model assignments. The Architect generates a backstory; the Notary (qwen3:14b, held constant) extracts a Fact Ledger; the deterministic Judge computes LPI and PCV and fires the gate. The deterministic Judge invokes no language model at all — it is plain arithmetic over the Fact Ledger’s fields (Equations 1–7), which is what makes the gate decision reproducible from the released ledgers. On acceptance, the Stress-Tester (claude-opus-4-7) drafts a Session-1 outline that a separate entity-reuse Judge (claude-sonnet-4-6, blind to LPI and PCV) labels, producing the semantic Adaptation Effort of Equation 12. The dashed back-edge is the gated arm’s regeneration loop: on rejection the Architect is re-prompted with an INSTABILITY_PROMPT, a fixed structural-rewrite instruction (Section 4.2, regeneration loop) that demands to rewrite partially the backstory, explicitly adding certain potentially missing elements (one named living NPC, one unresolved Tier-1 hook, and one new proactive motivation), for at most three iterations. The ungated arm omits this loop.

Table 1.
Uniform low/high anchors for the four components. Each row gives the backstory evidence, the substitution into Equations 2–4, and the resulting normalised score. A tie counts toward while it is unresolved, whether or not the NPC is still alive; what does not count is an NPC with no recorded stance toward the protagonist. The third row shows the peak-and-decay behaviour that penalises proliferating vagueness.
Table 1.
Uniform low/high anchors for the four components. Each row gives the backstory evidence, the substitution into Equations 2–4, and the resulting normalised score. A tie counts toward while it is unresolved, whether or not the NPC is still alive; what does not count is an NPC with no recorded stance toward the protagonist. The third row shows the peak-and-decay behaviour that penalises proliferating vagueness.
| Component | Anchor | Illustrative backstory evidence | Substitution | Score |
|---|---|---|---|---|
| Low | Mentor dead, family “lost to the plague”, no faction remembers the protagonist; one ambient hook about a strange birthmark. | , : | ||
| High | Rax Voidthorn hunts Kaelen for his brother’s death; a Harper handler wants him back; his sister is missing. Hooks: two T1, one T2. | , : | ||
| Low | Closes in passive mood: “Kaelen wanders the roads, uncertain of his purpose.” | : | ||
| High | Confront the Red Wizards in Thay; recover the stolen sigil ring; earn readmission to the Harpers. | : | ||
| Low | Four open hooks, one Tier-1 (an old debt); the rest are ambient curiosities—a family crest, a taste for Chultan tea, an unread letter. | 1 of 4 hooks at T1: | ||
| High | Four open hooks, three Tier-1—an active enemy, a curse that kills within the year, a kidnapped ally—plus one Tier-2 debt. | 3 of 4 hooks at T1: | ||
| Low | “Murdered by Captain Vex Ironfist on Mirtul 14 with a cursed silver dagger stolen from Candlekeep”—every mystery sealed before play. | : | ||
| High | Who ordered the killing; what weapon leaves no wound; why the Lords’ Alliance sealed the record. | : | ||
| Excess | Seven unexplained sigils, prophecies, and amnesias with no anchoring fact—vagueness has replaced structure. | : |
Table 2.
End-to-end calculation for two backstories from the locked corpus. “Ledger evidence” lists the quantities the Judge reads; “Score” applies Equations 2–4; the contribution column is the weighted term of Equation 1.
| Sterile draft (qwen3:14b, seed 35) | Play-ready draft (Opus 4.7, seed 28) | ||||||
|---|---|---|---|---|---|---|---|
| Component | Ledger evidence | Score | Score | Ledger evidence | Score | Score | |
| 3 nodes; T1+T2 | 3 nodes (cap); T1+T1 | ||||||
| 0 goals | 4 goals | ||||||
| 1 of 2 hooks at T1 | 2 of 2 hooks at T1 | ||||||
| 1 hole | 3 holes | ||||||
| (gate : fail) | (gate : pass) | ||||||
| Cumulative | 116 (one-shot) | 231 (campaign) | |||||
Table 3.
End-to-end calculation for the two backstories of Table 2. per Equation 7; the band column is the gravity bucket used by the Saturation-of-Conflict rule, and no band here exceeds the limit of two events, so every event contributes at full weight.
| Event (life phase) | I | G | V | D | Band | |
|---|---|---|---|---|---|---|
| Sterile draft — Kaelen Vireth (qwen3:14b, seed 35) | ||||||
| Mother, a former Harper agent, vanishes (infancy) | 3 | 8 | 9 | 6 | 2 | 59 |
| Drawn into an unnamed Harper conflict (adolescence) | 4 | 7 | 8 | 5 | 2 | 57 |
| Cumulative (one-shot tier; gate : fail) | ||||||
| Play-ready draft — Vessen Marrowlight (Opus 4.7, seed 28) | ||||||
| Oressa Marrowlight found dead, spiral sigil on her palm (infancy) | 7 | 9 | 8 | 6 | 3 | 78 |
| Sera Yindreth smuggles Vessen to Candlekeep (adolescence) | 5 | 7 | 9 | 7 | 2 | 64 |
| Red Wizards track Vessen to the Trade Way (adulthood) | 10 | 8 | 9 | 8 | 2 | 89 |
| Cumulative (campaign tier; gate : pass) | ||||||
Table 4.
Model assignment across both studies. Only the Architect varies (Study 2 adds three generators to Study 1’s two); every other role is frozen. The deterministic Judge is the canonical / calculator and runs no model.
Table 4.
Model assignment across both studies. Only the Architect varies (Study 2 adds three generators to Study 1’s two); every other role is frozen. The deterministic Judge is the canonical / calculator and runs no model.
| Role | Model | Provider | Status |
|---|---|---|---|
| Architect, Study 1 | qwen3:14b, claude-opus-4-7 | Ollama, Anthropic | independent variable |
| Architect, Study 2 | +qwen3:4b, qwen3:8b, llama3.1:8b | Ollama | independent variable |
| Notary (STE) | qwen3:14b | Ollama | held constant |
| Judge (LPI, PCV) | none | — | deterministic arithmetic |
| Stress-Tester (DM) | claude-opus-4-7 | Anthropic | held constant |
| Entity-reuse Judge | claude-sonnet-4-6 | Anthropic | held constant |
Table 5.
Pilot results, qwen3:14b via Ollama, per group, plus the hand-scored calibration fixture Kaelen Thorne (target ).
Table 5.
Pilot results, qwen3:14b via Ollama, per group, plus the hand-scored calibration fixture Kaelen Thorne (target ).
| Control | Architect | Kaelen | |
|---|---|---|---|
| (deterministic) | 0.772 | 0.739 | 0.843 |
| (LLM rubric) | — | — | 0.860 |
| 166 | 190 | 199 | |
| 3.06 | 3.15 | 2.75 | |
| Hallucinated entities | 9.3 | 8.3 | 7 |
| Provided entities | 5.7 | 4.3 | 5 |
| Pearson (pooled, ). | |||
Table 6.
Locked-run cell means ± SD, per cell. is the semantic judge-based statistic of Equation 12, bounded in .
Table 6.
Locked-run cell means ± SD, per cell. is the semantic judge-based statistic of Equation 12, bounded in .
| cell | conv. | |||
|---|---|---|---|---|
| ungated qwen3:14b | ||||
| gated qwen3:14b | ||||
| ungated Opus 4.7 | ||||
| gated Opus 4.7 |
Table 7.
Two-way ANOVA on , factors gating (ungated / gated) and generator (qwen3:14b / Opus 4.7), per cell.
Table 7.
Two-way ANOVA on , factors gating (ungated / gated) and generator (qwen3:14b / Opus 4.7), per cell.
| effect | p | ||
|---|---|---|---|
| gating (main) | |||
| generator (main) | |||
| gating×generator |
Table 8.
Mean per-trial entity counts in the DM Session-1 outline, classified by the Plan-C judge. per cell. Total emitted is the sum of REUSED, INVENTED, and AMBIGUOUS.
Table 8.
Mean per-trial entity counts in the DM Session-1 outline, classified by the Plan-C judge. per cell. Total emitted is the sum of REUSED, INVENTED, and AMBIGUOUS.
| cell | REUSED | INVENTED | AMBIGUOUS | total | |
|---|---|---|---|---|---|
| ungated qwen3:14b | |||||
| gated qwen3:14b | |||||
| ungated Opus 4.7 | |||||
| gated Opus 4.7 |
Table 9.
Global sensitivity of the play-ready classification to the weights, over the locked corpus ( Monte-Carlo draws per row; weights perturbed and renormalised). The reclassification rate stays below the robustness bound across the whole sweep.
Table 9.
Global sensitivity of the play-ready classification to the weights, over the locked corpus ( Monte-Carlo draws per row; weights perturbed and renormalised). The reclassification rate stays below the robustness bound across the whole sweep.
| Perturbation | Mean flip | Never-flip | Spearman |
|---|---|---|---|
| rate | backstories | ||
Table 10.
Study 2 cell means, per cell. Cells marked (S1) are Study 1 carryforward data. is the semantic judge-based statistic of Equation 12.
Table 10.
Study 2 cell means, per cell. Cells marked (S1) are Study 1 carryforward data. is the semantic judge-based statistic of Equation 12.
| Cell | ||||
|---|---|---|---|---|
| ungated qwen3:4b | ||||
| gated qwen3:4b | ||||
| ungated qwen3:8b | ||||
| gated qwen3:8b | ||||
| ungated qwen3:14b (S1) | ||||
| gated qwen3:14b (S1) | ||||
| ungated llama3.1:8b | ||||
| gated llama3.1:8b | ||||
| ungated Opus 4.7 (S1) | ||||
| gated Opus 4.7 (S1) |
Table 11.
Omnibus nonparametric tests, configurations, problems.
| Friedman | Iman–Davenport | Aligned Ranks | Quade | |
|---|---|---|---|---|
| Metric | ||||
| ***. (Nemenyi). | ||||
Table 12.
Friedman average ranks for across the ten generator–gating configurations (rank 1 = lowest ). U = ungated, G = gated.
Table 12.
Friedman average ranks for across the ten generator–gating configurations (rank 1 = lowest ). U = ungated, G = gated.
| Config. | Avg. rank | |
|---|---|---|
| G-Llama | ||
| G-14B | ||
| U-8B | ||
| U-Llama | ||
| U-14B | ||
| U-Opus | ||
| G-Opus | ||
| G-8B | ||
| U-4B | ||
| G-4B |
Table 13.
Cross-metric summary of nonparametric rankings. Rank 1 = best on each metric. Sig. pairs = Shaffer-corrected at out of 45.
Table 13.
Cross-metric summary of nonparametric rankings. Rank 1 = best on each metric. Sig. pairs = Shaffer-corrected at out of 45.
| Metric | Rank 1 | Rank 10 | Sig. pairs | ||
|---|---|---|---|---|---|
| G-Llama | G-4B | ||||
| G-Opus | G-4B | ||||
| G-Llama | G-4B | ||||
| U-Opus | G-Llama |
Table 14.
Human construct validity, backstories, raters each. r = Pearson of the per-item mean rating vs. its target metric; CI = bootstrap ; Holm p = one-tailed, corrected across the four convergent hypotheses; = single-rater intraclass correlation.
Table 14.
Human construct validity, backstories, raters each. r = Pearson of the per-item mean rating vs. its target metric; CI = bootstrap ; Holm p = one-tailed, corrected across the four convergent hypotheses; = single-rater intraclass correlation.
| Item | Target | r | 95% CI | Holm p | ||
|---|---|---|---|---|---|---|
| Q1 openness | ||||||
| Q2 conflict | ||||||
| Q3 burden | ||||||
| Q4 overall | * | |||||
| * Supported after Holm correction and passes the MTMM discriminant test. | ||||||
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.