Preprint
Review

This version is not peer-reviewed.

Beyond Utility Evaluation for LLM-Based and Generative Recommender Systems: A Focused Survey of Offline Protocols

Submitted:

18 September 2026

Posted:

20 September 2026

You are already at the latest version

Abstract
Evaluation validity can change which recommender appears better. Ranking a relevant item supplied by an evaluator differs from finding it in a catalog; generating a valid identifier differs from identifying one item. This focused survey examines 20 offline protocol studies within a 49-paper coded corpus and develops two findings. First, reported model or history-strategy orderings change when studies vary candidate pools, credit for ambiguous identifiers, or the benefit measured by fairness. Second, apparently conflicting results can concern different objects: selecting long-tail recommendations differs from recovering popular benchmark records, and reducing attribute recoverability differs from distributing user benefit. We explain these findings through the items a model can access, the outputs receiving credit, the held-out population, and the outcome being measured. The synthesis derives evaluation controls for interpreting utility, diversity, exposure, and fairness in LLM-based and generative recommendation.
Keywords: 
;  ;  ;  ;  ;  

1. Introduction

Consider two movie recommenders: one ranks a short list containing a relevant film, while the other must find that film in the catalog. Their hit rates answer different questions. A title-generating recommender introduces a further decision: whether its output names exactly one catalog item. These choices affect both accuracy and the diversity or exposure attributed to its recommendations. Surveys of evaluation practice and trustworthy recommendation establish the importance of these distinct outcomes [1,2,3,4].
The limitation of accuracy alone predates language-model recommendation. McNee et al. [5] question the use of accuracy as the dominant design objective, while Ge et al. [6] distinguish coverage and serendipity from predictive accuracy. The present problem is how to preserve these distinctions when a recommender generates the object that the evaluator must identify and score.
LLM-based and generative recommenders also change what receives a score. A model may classify a supplied item, rank a candidate list, generate a title, or emit an item identifier. P5, TALLRec, GenRec, RecRanker, and token-level generative methods illustrate these different uses of language modeling [7,8,9,10,11]. Evaluating them requires a connection from the raw output to the credited recommendation.
That connection can change a conclusion. Jiang et al. [12] observe sensitivity to history, candidate position, and profile generation, and discuss unmatched generated titles under their hallucination label. Di Palma et al. [13] recover public benchmark records from LLMs, raising questions about what familiar-dataset accuracy establishes. These issues compound known effects of baseline tuning, sampled metrics, implementation choices, and incomplete artifacts [14,15,16,17,18]. High utility is interpretable only after the comparison task, credited item, and evaluated population are defined.
Existing LLM4Rec and generative recommendation surveys organize models and tasks [19,20,21]; trustworthy recommendation surveys organize stakeholder objectives. Our focus is the link between these two perspectives: how validity choices change the utility or beyond-utility claim supported by an experiment. Jannach and Chen’s methodological guidance on comparison rigor and reproducibility provides the closest TORS foundation [22]. Table 1 identifies the connections developed here.
We examine 20 detailed protocol comparisons within a 49-paper coded corpus. The synthesis yields two recurring findings:
  • Protocol choices change reported orderings. Candidate-pool selection, credit for ambiguous identifiers, and the reference used to measure fairness each change which model or history strategy appears preferable in source-reported comparisons.
  • Different measured objects can explain apparent disagreement. Long-tail candidate selection and popular-record recovery describe different operations. Attribute recoverability, user benefit, and item exposure likewise express different objectives. Reading these results together clarifies what each establishes.
We use candidate access, item identity, generalization, and outcome reference to organize the comparisons. Section 2 and Section 4 establish the review method and evaluated components. Section 5 distinguishes beyond-utility outcomes, and Section 6 connects their interpretation to the protocol comparisons. The evaluation controls in Section 7 are practical consequences of this synthesis.

2. Scope and Survey Methodology

We conduct a focused review of offline evaluation in LLM-based preference prediction, candidate ranking, and item generation. The review asks how the evaluated output determines the required evidence, which protocol choices change comparisons between systems, and how those choices affect beyond-utility interpretation. Language-based representation methods provide contrasts with direct generation. The empirical synthesis addresses utility, diversity, consumer fairness, item exposure, and generalization; interaction, explanation, and privacy motivate extensions beyond this core.
  • Identification and selection.
The search was completed on 15 September 2026. We followed citation trails from LLM recommendation evaluation, generative recommendation, trustworthy recommendation, and reproducibility studies, consulting proceedings, publisher, DBLP, and arXiv records. Searches combined recommendation with evaluation, protocols, generative models, and LLMs, followed by title-based retrieval. The target publication window is 2019–2026; earlier work is included for established evaluation concepts. Selection is purposive: a study enters the detailed comparison when its reported task, candidate construction, output resolution, evaluated population, or outcome reference supports a concrete protocol contrast. Studies describing adjacent models or general principles provide contextual support. Preprint and published versions are associated with the same work, with the version used for each extracted result recorded separately.
  • Extraction and comparison.
The corpus contains 49 records: 40 with full-text coding and nine with bibliographic or source-record coding. The 23-field schema records system role, task, output, data, split, candidates, metrics, validity checks, prompts, baselines, tuning, artifacts, user evaluation, and protocol threats. Source descriptions are distinguished from our interpretations. Table 2 assigns each record one primary contribution for corpus accounting, while the analysis permits a study to inform several questions.
Twenty full-text studies support the detailed protocol comparison. They cover supplied-target prediction, sampled and retrieved candidates, full-catalog scoring, generated names and identifiers, representation and memorization probes, held-out entities and prompts, and four fairness formulations. Figure 1 organizes these contrasts by the question they resolve. We compare the experimental objects and controls before interpreting reported scores; heterogeneous metrics are not pooled. The supplement links each selected study to its inclusion rationale, extracted fields, PDF locations, and the comparison it supports.
  • Evidence provenance and scope.
An AI assistant performed the extraction and source rechecks; independent human coding agreement has not been established. The detailed comparison retains unspecified and conflicting source fields, and numerical contrasts were checked against the source tables. Results for CFaiRLLM and the provider-fairness study refer to the archived arXiv versions identified in the supplement, with journal metadata recorded separately. The review compares reported protocols and results rather than rerunning model code.
The initial snowballing search has no complete retrieved-and-excluded ledger. Accordingly, corpus counts describe this collection and do not estimate the prevalence of practices across the field. Privacy attacks and end-to-end agent behavior are outside the detailed empirical comparison. The broader bibliography supplies definitions, methodological foundations, and related work; it is distinct from the coded analytical sample.
Table 2. Evidence clusters used to build the survey taxonomy. The n column reports the number of papers in the conservative coded corpus, not a field-wide prevalence estimate.
Table 2. Evidence clusters used to build the survey taxonomy. The n column reports the number of papers in the conservative coded corpus, not a field-wide prevalence estimate.
Evidence cluster n Representative sources Role in this survey
LLM/generative protocols 19 LLM evaluation, memorization, direct generation, reranking, identifier validity, and cold-start protocols [9,10,11,12,13,25,26] Connects history and position sensitivity, generated-item identity, contamination, and unseen-entity evaluation to the meaning of reported performance.
Classical evaluation methodology 6 Reproducibility analyses, sampled-metric critiques, evaluation frameworks, and TORS methodological standards [14,15,16,17,18,22] Provides the baseline for rigorous evidence: tuned comparisons, exact metric definitions, split validity, artifact completeness, and reproducible evaluation pipelines.
Beyond-utility and trustworthy recommendation 8 Evaluation landscape, trustworthy recommendation, popularity bias, and LLM consumer/provider fairness studies [1,2,27,28,29,30] Supplies mature dimensions beyond top-k utility, including diversity, novelty, coverage, calibration, popularity bias, fairness, privacy, robustness, and stakeholder value.
LLM4Rec and agentic landscapes 4 Broad LLM4Rec, generative recommendation, and agentic recommender surveys [19,20,21,31] Serves as scope control: these papers organize methods and tasks, while this survey reorganizes the same space by evidence needs and protocol failure modes.
Interaction and judge evaluation 2 Conversational recommender evaluation and LLM-as-judge surveys [32,33] Extends evaluation from ranked lists to dialogues, subjective judgments, rubric calibration, judge bias, and agreement with human or user-centered evidence.
Multimodal and group boundary cases 9 Multimodal survey, robustness/generalization, fusion, efficiency, consensus, and group recommendation work [34,35,36,37,38,39] Keeps the paper recommender-specific: these boundary cases show that robustness, modality contribution, efficiency, consensus, and uncertainty are not captured by utility alone.
Additional frontier examples 1 Recent generalist embedding evaluation for zero-shot recommendation and search [40] Marks an adjacent evaluation question, whether broad pretrained representations should be judged as recommenders, search systems, or transfer models.

3. Background: From Utility to Evaluation Validity

Utility measures the recovery of relevant items or the prediction of held-out preferences. Beyond-utility outcomes concern what else the system delivers, such as diverse recommendations or long-tail exposure. Validity concerns whether the observations and comparison support the intended interpretation. These are different questions: a score may accurately describe a restricted candidate task while providing weak evidence about full-catalog recommendation.
Evaluation frameworks start with the user task and experimental setting, rather than with a universally preferred metric [41,42,43]. Gunawardana and Shani [44] explain why metric choice can select a different algorithm for a different recommendation task. This distinction is essential when a binary preference predictor, a reranker, and an item generator all report improvements under the broad label of recommendation accuracy.
A metric specifies a calculation; a protocol supplies its inputs and population. Generative outputs add an interpretation step between prediction and scoring: titles must resolve to items, identifiers may denote collision groups, and factual explanations require reference evidence.
Figure 2 traces these dependencies. Protocol choices define the task and experimental conditions. Raw outputs are retained, mapped to catalog records where possible, and scored under explicit credit rules. Utility and beyond-utility measurements are derived from a shared set of traceable records, with each metric specifying its statistical unit, evaluated population, and denominator. Request-level utility, catalog coverage, and provider fairness can therefore use different denominators.
  • Metric names do not fully specify the comparison.
Cumulated gain measures ranked retrieval with graded relevance [45]; top-N recommendation evaluation concerns which relevant items appear in the returned list [46]. The observed labels and candidate construction still determine what either calculation estimates. Analyses of statistical bias in recommendation metrics [47] and of alternative splits, sampling, and domains [48] make those choices part of the scientific comparison. Tamm et al. [49] additionally show that implementations sharing a metric name can disagree. Thus the audit records the relevance rule, candidate universe, discount, normalization, and aggregation unit alongside the score. This task-specific view is consistent with the broader challenges of offline evaluation identified by Castells and Moffat [50].
Output validity also affects aggregation. Section 7 distinguishes quality over all attempted requests from quality conditional on a valid output, and shows how the two can favor different systems.

4. System Roles in LLM-based and Generative Recommendation

The input and output of an evaluated component determine its role. A preference predictor scores a supplied user–item pair; a reranker orders a supplied set; an enhancer supplies features to another recommender; a generator emits an item name or identifier. A conversational or agentic system composes such components into an interaction. Roles describe operations, so a complete system may contain several. Adaptation by prompting or fine-tuning, sequential or multimodal context, and targets such as utility or exposure are separate facets.
TALLRec and RecRanker illustrate the distinction. Both use tuning, but the former evaluates binary preference prediction with AUC, whereas the latter evaluates candidate reranking [8,10]. Figure 3 illustrates the input–output contracts and the question each exposes.
The same language-model technology supports different contracts. LLMRank orders supplied candidates [51]; TIGER and LETTER generate semantic identifiers [52,53]. UniSRec and Recformer instead use item descriptions to learn representations for scoring [54,55]. OpenP5 supports alternative indexing methods and task configurations [56]. These distinctions determine whether evaluation measures preference discrimination, conditional ordering, or access to a relevant catalog item.
An LLM-enhanced recommender supplies profiles, embeddings, or other features to a backbone. Its contribution is assessed by holding the backbone fixed and varying the added information, encoder, or feature-generation step. A cheaper text encoder and an ablation without generated features help separate useful information from additional capacity or reformatting.
Direct generation exposes item resolution as an evaluation operation. GenRec generates names, while semantic-ID methods generate tokens that may identify multiple items [9,25]. Constrained decoding and post-hoc matching determine which items receive credit. The paper-level comparison in Section 6 examines these choices.
Interaction shifts the output from a list to a trajectory, bringing preference elicitation and task completion into evaluation [32]. Multimodal recommendation adds the contribution of different input sources [34]; group recommendation adds member-level outcomes [38,39]. These extensions require observations beyond a single ranked list.

5. Evaluation Dimensions Beyond Utility

Beyond-utility results depend on both the outcome being valued and the operation being observed. Candidate selection, free generation, and record extraction can produce different popularity patterns. Fairness measures can disagree because they compare different benefits. Table 3 summarizes the selection and extraction findings; Table 4 defines the measurement targets used throughout this section.
  • Exposure and memorized knowledge measure different behavior.
Jiang et al. [12] report stronger long-tail exposure for several LLM configurations in Beauty and LastFM, and better reported serendipity despite lower overall accuracy for the best LLM configurations in Beauty (Section 4.3, Figure 2). Their ranking task supplies one positive item and 19 sampled negatives. Di Palma et al. [13], by contrast, recover popular MovieLens items more readily than rarely interacted items in record-extraction probes (Section 3.4, Figure 4). These observations concern different operations: selecting among presented candidates versus recalling a benchmark record. They can coexist. A popularity-skewed knowledge store does not by itself determine the exposure distribution of a candidate-conditioned recommender.
  • Input representation changes beyond-utility outcomes.
In Jiang et al.’s Beauty and Sports comparisons, profile-only inputs retain similar accuracy while improving the reported popularity measure; profile-plus-history improves the reported dimensions (Section 4.6). This intervention supplies evidence about recommendation behavior. Explanation faithfulness requires a separate test of whether the profile reflects the recommendation mechanism.
Metric definitions explain what transfers across comparisons. Jiang et al. define long-tail items as the bottom 80% by item frequency and appendix coverage relative to the union of candidate sets. Their serendipity prose invokes usefulness, whereas the displayed Eq. 6 measures non-overlap with MostPop without an explicit relevance indicator. We therefore retain the source-defined metric name without inferring verified useful novelty from that formula. Appendix user-group disparity uses activity groups; those definitions are not demographic-fairness results.
Table 4. Measurement objects and their reference information. Table 3 and Table 5 compare empirical findings; privacy is an extension, and output validity is a condition on scoring.
Table 4. Measurement objects and their reference information. Table 3 and Table 5 compare empirical findings; privacy is an extension, and output validity is a condition on scoring.
Target or check Measured object Reference information needed
Diversity Differences or category proportions within returned lists. Resolved item identities, item features or taxonomy, list length, and aggregation rule.
Novelty and serendipity Unfamiliarity or unexpectedness relative to prior knowledge or expectations. User or popularity reference; an explicit relevance component when claiming useful unexpectedness.
Catalog coverage Distinct catalog items reached across recommendations. Eligible catalog, evaluated requests, output length, and treatment of unresolved items.
Calibration Agreement between recommendation and preference distributions. Reference preferences, attribute categories, distance measure, and user-level aggregation.
Popularity bias Recommendation exposure relative to item frequency or user preference. Popularity counts and time window, candidate popularity, and the chosen exposure or preference baseline.
Fairness Distribution of benefit, sensitivity, or exposure across defined stakeholders. Beneficiary, groups, benefit reference, allocation rule, statistical unit, and time horizon.
Robustness and generalization Outcome stability under perturbations or transfer to held-out entities. Fixed task, perturbed input, held-out unit, available metadata, and repeated observations.
Privacy (extension) Protected information recoverable under an adversary’s access. Protected-data definition, attack capability, success measure, or formal privacy guarantee.
Output validity (check) Catalog resolution and factual support for generated assertions. Matching rules, ambiguous identities, unresolved requests, and evidence for item or attribute existence.

5.1. Novelty, Diversity, and Calibration

TIGER provides a decoding-level contrast with prompted ranking [52]. On Beauty, increasing sampling temperature from 1.0 to 2.0 raises category Entropy@10 from 0.76 to 1.38 (Table 3). This measures the category distribution of returned items, whereas Jiang et al.’s long-tail exposure and candidate-union coverage measure item frequency and reach. TIGER’s table does not pair these temperature settings with accuracy values, so it establishes a diversity response to decoding rather than a quantified accuracy–diversity frontier. The shared reporting requirement is to name the output distribution being diversified and retain the decoding configuration that produced it.
Beyond-accuracy objectives differ in their statistical objects. Kaminskas and Bridge [57] distinguish diversity, novelty, serendipity, and coverage; Vargas and Castells [58] make rank and relevance explicit in novelty and diversity measurement. Topic diversification changes the composition of a recommendation list [59], while maximal marginal relevance is an earlier retrieval mechanism for balancing relevance and redundancy [60]. A generated list can therefore become less redundant without becoming more useful, more surprising to its recipient, or broader in catalog coverage. These outcomes require separate operational definitions.
Calibration adds a reference distribution. Steck [61] aligns recommendation proportions with the user’s historical interests, and Kaya and Bridge [62] compare calibrated and intent-aware recommendations. The calibration survey further distinguishes technical objectives from evidence of effectiveness [24]. For an LLM-generated profile, the audit must consequently state whether the reference distribution comes from observed interactions, explicit preferences, or the generated summary itself. Measuring agreement with the summary tests consistency with that representation; it does not independently establish preservation of the user’s preferences.
Popularity mitigation presents a related choice of reference. Controlling popularity in learning-to-rank [63] and evaluating popularity bias from the user’s perspective [64] address different aspects of the problem. Increasing aggregate long-tail exposure does not establish that each user’s preference for popular or less popular items is respected. Results should therefore distinguish an item-frequency statistic, a catalog-wide exposure distribution, and user-level preference mismatch. This separation also explains why a representation-spectrum diagnostic cannot establish exposure diversity.

5.2. Fairness Requires a Beneficiary and an Allocation Rule

Multistakeholder recommendation makes consumer and provider interests explicit [65]. Exposure fairness specifies how positions allocate visibility [66]; equity of attention accumulates that allocation across rankings [67]. Pairwise fairness instead evaluates comparisons between items [68], and FairRec formulates a two-sided allocation problem [69]. These are distinct targets, not alternative names for a single fairness score.
For generated recommendations, this distinction has a concrete consequence. A provider cannot receive catalog-level exposure credit until an output is resolved to its item and provider. Conversely, counting only successfully resolved outputs can conceal unequal service failure across user groups. A fairness report should identify the protected or stakeholder groups, the exposure model, the accumulation period, and the relation between group membership and unresolved requests. It can share the traceable request records used for utility while retaining a different statistical unit and denominator.
  • The benefit reference can reverse a fairness comparison.
FaiRLLM [27] compares generated lists under neutral and sensitive-attribute prompts, measuring how their similarity varies across groups. Its constructed music and movie preferences permit controlled prompt interventions, but supply no held-out relevance labels. CFaiRLLM [28] adds a different reference: test interactions used to ground preference alignment. Table 5 places both beside attribute-recoverability and item-exposure evaluations. The choice is substantive: equal sensitivity to demographic wording, equal recommendation benefit, hidden sensitive information, and equal item exposure are different objectives.
Table 5. Four fairness claims and their operational definitions. Source versions and field-level evidence are recorded in the supplement.
Table 5. Four fairness claims and their operational definitions. Source versions and field-level evidence are recorded in the supplement.
Study Measured object Protocol Interpretation
FaiRLLM [27] Group variation in similarity to neutral-prompt lists. Constructed preferences; 20 titles; Jaccard, SERP*, PRAG*, SNSR/SNSV (Sections 2–4). Demographic-prompt sensitivity, without held-out relevance labels.
CFaiRLLM [28] Group disparities under list or test-preference alignment. Temporal MovieLens/LastFM profiles; history selection and title matching (Sections 3–4). Preferred history strategy changes with the reference (Table 4, arXiv v3).
UP5 [29] Sensitive attributes recoverable by a probe. MovieLens/Insurance; attribute AUC and sampled Hit@k (Section 5). Probe-specific recoverability alongside utility, distinct from benefit parity.
Provider-fairness study [30] Item-count concentration and catalog coverage. Prompt and ICL variants; title/artist matching; Gini, HHI, entropy, coverage (Sections 3–5). ICL can concentrate exposure; item equality serves as a provider-side proxy.
The difference changes a reported ordering. In CFaiRLLM’s Table 4 (arXiv v3), the Jaccard-based sex-group SNSR is 0.0010 for random history sampling and 0.0424 for recent sampling under item-list alignment; lower is better. Under preference-grounded alignment, the corresponding values are 0.0210 and 0.0086, favoring recent sampling. Comparing strategies within each reference shows that the preferred history depends on the benefit being measured. The absolute fairness levels remain specific to their respective metric definitions.
  • A probe tests recoverability, not the distribution of benefit.
UP5 [29] evaluates sensitive-attribute prediction alongside recommendation accuracy. In its MovieLens gender comparison for matching-based models (Table 3), CFP reports attribute AUC of 54.19%, versus 56.62% for C-PMF and 70.80% for C-SX, while Hit@10 is 65.82%, 65.32%, and 56.02%, respectively. CFP combines lower attribute recoverability with higher recommendation Hit@10 in this comparison. This is evidence about information accessible to the tested probe; CFaiRLLM instead measures differences in recommendation benefit. The two evaluations can therefore complement one another without agreeing on a single fairness ranking.
  • More context need not broaden exposure.
Deldjoo’s provider-fairness study [30] measures recommendation-count concentration and catalog coverage after matching generated names to items. In its sequential MovieLens experiment, the no-demographics, frequency-profile setting has coverage 0.02621 with zero-shot prompting and 0.01685 with ICL-2; Gini rises from 0.57739 to 0.65835 (arXiv v3, Table 10). The added demonstrations coincide with narrower, more concentrated exposure in this comparison. Read alongside CFaiRLLM, this result shows why history and demonstration policies belong in a beyond-utility protocol: they change both whose preferences are represented and which items receive exposure. These item-frequency statistics use equality among items as a provider-side proxy; provider ownership and merit require additional definitions.

5.3. Privacy and Explanation Need Different Evidence

Privacy evaluation specifies protected information and adversary access. Re-identification from ratings [70], disclosure through collaborative-filtering outputs [71], membership inference [72], and language-model extraction [73] concern different information and access routes. Differential privacy supplies a formal guarantee for a randomized mechanism [74]. These distinctions identify what a recommendation-specific privacy claim must protect; public-item recall alone is not a disclosure test.
Explanation research distinguishes factual support, user understanding, and the reason for a recommendation [75,76]. Scrutable user models additionally expose preferences for correction [77]. For generated explanations, accurate item descriptions and faithful accounts of the ranking mechanism are separate evaluation targets. Section 8 considers how to connect them to user-observed outcomes.

5.4. Output Validity and Robustness

  • Matching failure and hallucination require different checks.
A matching failure occurs when an output cannot be resolved to an eligible catalog item. Causes include malformed output, unresolved aliases, collisions, and real items outside the catalog. We reserve item hallucination for an asserted fabricated item and attribute hallucination for an unsupported factual assertion. An existing but irrelevant item is a relevance error. Jiang et al.’s prose defines hallucination through titles unmatched after case, space, and symbol normalization, including omissions and combined titles. Their displayed Eqs. 10 and 24 instead use an in-catalog membership indicator, reversing this interpretation. We follow the prose definition here; whether the implementation follows that definition or the printed equations remains unresolved. Neither a non-match nor the source’s metric label establishes fabrication.
This distinction also matters to exposure. A failed exact match may denote a real long-tail item; nearest-item repair may instead assign it to a popular item. Novelty and coverage should consequently be computed over resolved item identities with unresolved outputs retained in the audit, rather than treating unfamiliar wording as novel content.
Robustness tests specify both the changed input and the outcome expected to remain stable. Candidate permutations preserve eligibility; missing-modality tests change available information. Multimodal work motivates checking modality contribution and sparse-data behavior [35,36,78], with efficiency a separate practical objective [37]. These extensions follow the same principle of defining the evaluated object before interpreting a score change.

6. How Protocol Choices Change Comparative Findings

The fairness comparisons in Section 5 show that the outcome reference can change which history strategy appears preferable. We now examine how candidates, item credit, and held-out populations determine the recommendation problem being scored. Table 6, Table 7 and Table 8 organize these contrasts by evaluation question.

6.1. Candidate Access Changes the Ranking Problem

Dai et al. use one positive and four negatives; Jiang et al.’s ranking task uses one positive and 19 negatives. RecRanker instead describes retrieval followed by reranking [10,12,79]. The first two designs test ordering when the target is available. The third also depends on whether retrieval finds it. Candidate-order robustness addresses presentation, whereas changing candidate membership changes the task itself.
With one held-out relevant item per request, end-to-end Hit Rate at k factors into the probability of retrieving that item and the probability of ranking it in the top k conditional on retrieval. Inserting the target into every candidate set estimates the conditional term under a constructed set. A pipeline comparison therefore retains retrieval misses and reports conditional reranking separately.
LLMRank separately examines supplied-positive sets, hard negatives, and retrieved candidates [51]; UniSRec and Recformer specify all-item test ranking even though their training objectives use negative examples [54,55]. Thus negative sampling must be located in training or evaluation before a comparison is classified. A model’s language backbone does not determine its candidate protocol.
Complete relevance judgments do not remove candidate-selection effects. Sanner et al. [80] collect judgments from 153 raters over pools combining EASE, text retrieval, and popularity-stratified random selection. In their Table 3, EASE exceeds the language-only three-shot LLM on the full pool (NDCG@10: 0.673 versus 0.640), but falls below it on the random subpool (0.592 versus 0.650). This is a reversal of reported point estimates, not an established paired significant difference. EASE also contributes items to the full pool. The contrast changes the candidate task and shows why label completeness and selection neutrality require separate checks; the random subpool is drawn from specified popularity strata, not uniformly from the entire catalog.
Table 7. Generated outputs and item-level credit. Name resolution, identifier membership, and one-to-one item identity address different sources of scoring ambiguity.
Table 7. Generated outputs and item-level credit. Name resolution, identifier membership, and one-to-one item identity address different sources of scoring ambiguity.
Study Output and evaluation Resolution or decoding rule Consequence for interpretation
GenRec [9], Sections 3–4 Generated names; MovieLens 25M and Amazon Toys; HR/NDCG@5,10. Last/penultimate interactions held out; matching and failure aggregation not fully specified. Reproducing a name-level score requires its catalog matching and failure-counting rules.
TIGER [52], Sections 4.3, 4.5 Four-token semantic IDs; generated top-10 lists. Disambiguating token for known items; approximately 0.1–1.6% invalid IDs across three datasets. Unique registered IDs do not guarantee that every generated sequence is registered.
LETTER [53], Sections 3.2, 4 Learned item tokenization; tokenizer, regularizer, and backbone comparisons. A trie restricts generation to valid successor tokens. Membership in the identifier language does not alone establish one-to-one item credit.
Zhang et al. [25], Sections 3.1, 4 Scientific, Cell, Beauty, Yelp; shared generator across tokenizers; three seeds. SID Hit/NDCG versus collision-corrected item metrics. Correcting identifier-level credit changes tokenizer ordering without changing native outputs.

6.2. Valid Outputs Can Receive Ambiguous Item Credit

Dai et al.’s compliance rate checks membership in the supplied set; it does not specify how invalid answers enter NDCG or MRR. GenRec reports generated-name HR and NDCG without a complete matching and failure-aggregation rule in the inspected paper [9,79]. Semantic identifiers pose a different problem: an output can match a valid identifier but refer to several items.
Zhang et al. [25] show the consequence using the same native tokenizer outputs. On Scientific, RK-Means scores 0.1330 against LETTER’s 0.0804 under SID Hit@10, but 0.0654 against 0.0767 under collision-corrected ItemHit@10 (Table 4). The ordering reverses when credit reflects ambiguous item identity. The correction expands collision groups and averages over unresolved within-group order; an actual disambiguator needs its own selection rule. Together, these studies distinguish response compliance, item resolution, and score aggregation as separate operations.
TIGER and LETTER expose the two sides of identifier validity [52,53]. TIGER appends a disambiguating token to known-item semantic IDs, yet reports approximately 0.1–1.6% invalid IDs among top-10 predictions across its three datasets (Section 4.5). LETTER constrains generation with a trie of valid successor tokens (Section 3.2.2). The first controls collisions among registered items; the second controls membership in the registered identifier language. Neither property substitutes for the other. A protocol should document both the decoding constraint and the map from each credited identifier to an item.
Table 8. Generalization and intermediate evidence. The held-out object and the observed behavior determine whether a result concerns transfer, record recoverability, or delivered recommendation quality.
Table 8. Generalization and intermediate evidence. The held-out object and the observed behavior determine whether a result concerns transfer, record recoverability, or delivered recommendation quality.
Study Evaluation object Generalization or control Consequence for comparison
UniSRec [54], Section 3.1 Full-catalog item scoring from sequence representations. Text-only inductive versus text-plus-ID transductive adaptation. Training negatives are not the test pool; representation transfer differs from a temporal cold-item test.
Recformer [55], Section 3.4.1 Full-catalog text-based scoring; in-set and unseen-item tests. Cold-token handling for SASRec; text encodings for unseen items. Baseline access to unseen items is part of the generalization comparison.
OpenP5 [56], Tables 2–4 Indexed-item generation; random, sequential, or collaborative indices. Seen/unseen prompt templates; T5 and LLaMA backbones. Held-out prompts do not establish held-out-user or held-out-item performance.
Zhang et al. [26], Sections 3–4 Toys, MicroLens, Steam; temporal item split and separately held-out users. Full-catalog item-cold targets; Recall/NDCG@10; scale, ID, and training comparisons. Short history, unseen user, and unseen item are distinct generalization targets.
Di Palma et al. [13], Sections 2–3 MovieLens-1M; recommendation and record-extraction probes. 80/20 leave-n-out; prompt requests 50 titles; HR/nDCG and record coverage. Recovering benchmark records informs exposure; it does not isolate its causal effect on accuracy.
CLLMR [81], Sections 4–5 Amazon Books, Yelp, Steam; generated user/item side information; 3:1:1 split. Recall/NDCG at 10, 30, 50; five initializations; singular-value diagnostics and ablations. Representation collapse is distinct from catalog exposure and stakeholder fairness.

6.3. Generalization Depends on What Is Held Out

TALLRec predicts binary preference for a supplied target. Its Movie data use preceding interactions, but Book histories are sampled because BookCrossing lacks timestamps. RecRanker also states leave-one-out while reconstructing BookCrossing histories randomly [8,10]. Thus a shared split label does not establish a shared temporal prediction problem.
The cold-start study by Zhang et al. [26] makes entity exposure explicit: item-cold targets first appear after a temporal cutoff, while user-cold evaluation holds out users and supplies short histories. These populations differ from shortening an existing user’s context. TALLRec’s few-shot tuning additionally varies training examples, not the same quantity as inference-history length. Comparisons of generalization need the held-out entity, chronology, and available context, rather than the cold-start label alone.
Holding out an item also raises the question of how it becomes eligible at test time. TIGER’s separate cold-item experiment removes 5% of test items from training, then adds unseen items sharing the generated three-token prefix under a mixture cap. Recformer encodes unseen items from text and gives its SASRec baseline a trained cold-token representation [52,55]. These experiments test different routes into the eligible item set. OpenP5’s unseen results instead hold out prompt templates [56]; interpreting them as unseen-item results would change the generalization claim without changing any score.

6.4. Knowledge and Representations Are Intermediate Objects

Record-extraction probes address a different object from the ranking tasks above. Cross-model associations between extraction success and recommendation accuracy leave model capability and training exposure confounded [13]. The candidate-conditioned exposure findings in Section 5 consequently cannot be explained by extraction success alone. Linking the two would require observing how recoverable knowledge is used when selecting recommendations.
CLLMR supplies a complementary example from the enhancer role [81]. Its singular-value analyses and module ablations examine representation collapse and the contribution of generated side information. These feature-level diagnostics complement the list-level exposure studies and knowledge-extraction probes: the three evaluate representations, accessible model knowledge, and delivered recommendations, respectively. Connecting them requires measuring how changes in features or recoverable knowledge affect the returned lists.

6.5. Selection Begins Before Model Evaluation

The observations used for evaluation are already selected: MovieLens is a collected interaction resource [82], and missing-not-at-random analysis [83] and treatment-based recommendation [84] address the resulting bias. Offline bandit evaluation [85], counterfactual risk minimization [86], and unbiased learning-to-rank [87] provide corrections under assumptions about observation and exposure. Generative recommendation adds selection at two later points: which items become eligible and which outputs survive parsing. Correcting the logged-feedback process leaves these later decisions to be specified separately.

6.6. Two Findings Across the Protocols

  • Protocol choices change reported orderings.
The candidate-pool comparison changes which alternatives are evaluated; identifier correction changes which item receives credit; the fairness comparison changes the benefit reference. Each produces a source-reported ordering change through a different operation. The practical question is therefore which comparison matches the intended use: selecting among available candidates, retrieving an identifiable item, or distributing a specified benefit.
  • Different measured objects explain apparent disagreement.
Long-tail selection among presented candidates and popular-record recovery concern recommendation behavior and accessible knowledge, respectively. Attribute recoverability and group benefit concern information in a representation and the distribution of an outcome. The distinctions reconcile findings without averaging unlike measurements. They also locate the missing link: a knowledge or representation diagnostic needs additional evidence connecting it to delivered recommendations. The controls below follow from these two findings.

7. Implications for Evaluation Practice

The findings in Section 6.6 lead to two practical tasks: hold the comparison conditions fixed when attributing an improvement to a model, and retain distinct measurements when they concern different objects. Table 9 collects the reporting fields. We focus here on the design and accounting choices needed to use them.

7.1. Fix the Comparison Before Scoring

Framework comparisons [88], rigorous benchmarking guidance [89], and re-evaluations of collaborative-filtering [90] and session-based models [91] motivate strong baselines configured for the task being tested. The same requirement applies when the new component is an LLM.
A reranker comparison needs the same held-out requests, candidate membership, item metadata, and candidate-order permutations for each model. An end-to-end comparison instead allows each system to retrieve its own candidates, but retains the same request population and reports retrieval failures. These two experiments answer different questions. Prompts, matching thresholds, and hyperparameters should be selected on validation data before the test set is scored. Tuning reports should give search spaces, selection criteria, trial counts, and computational cost; equal trial counts alone do not establish equally effective tuning across model families.
For fairness, fix the benefit reference as well as the protected groups: a neutral-prompt list, held-out preferences, probe recoverability, or an exposure allocation. A comparison can then vary history selection while preserving users, history budget, output resolution, and the metric definition. When several references are scientifically relevant, report the strategy ordering under each, as in the CFaiRLLM contrast, rather than collapsing the results into one fairness score. Generalization reports should similarly identify whether the held-out unit is a prompt, user, item, or domain.

7.2. Keep Validity in the Denominator

Consider a single-item recommendation task. Let N be the number of attempted requests, V the number yielding a valid catalog item, and H the number yielding a valid item judged relevant. Validity rate is V / N , and conditional hit rate is H / V for V > 0 . If an invalid output counts as a failed request, end-to-end hit rate is
H N = V N H V .
When V = 0 , conditional hit rate is undefined and end-to-end hit rate is zero. This accounting identity makes the denominator choice explicit. Conditional quality measures successful outputs; end-to-end quality measures all attempted requests.
Figure 4 gives an illustrative comparison, not a result from an evaluated model. System A returns 50 valid items and 40 hits across 100 requests; system B returns 100 valid items and 60 hits. A leads on conditional hit rate (80% versus 60%), but B leads on end-to-end hit rate (60% versus 40%). Dropping invalid requests reverses the ordering. The figure motivates reporting both rates with their counts.
Figure 4. Illustrative model-order reversal from the denominator choice. Restricting evaluation to valid requests favors A; counting all attempts, with invalid outputs receiving no hit credit, favors B. Both panels share a zero-based percentage scale. Counts are constructed, not empirical benchmark results.
Figure 4. Illustrative model-order reversal from the denominator choice. Restricting evaluation to valid requests favors A; counting all attempts, with invalid outputs receiving no hit credit, favors B. Both panels share a zero-based percentage scale. Counts are constructed, not empirical benchmark results.
Preprints 233950 g004
For top-k lists, the evaluator must additionally define missing slots, duplicate items, and ranks after parsing. Removing an invalid first item and shifting later items upward changes discounted gain. A protocol can preserve the original slots and assign zero gain to invalid entries, then report post-repair utility separately when repair is part of the deployed system. Equation 1 should not be transferred directly to listwise NDCG, whose normalization and partial-list validity require their own definitions.
Validity also conditions beyond-utility metrics. Exposure fairness computed only over matched items omits users for whom generation failed. A useful report therefore pairs exposure or novelty with validity by the relevant user or item stratum. The catalog snapshot and popularity reference must stay fixed across systems so that changes in the metric reflect their outputs.

7.3. Test Stability on Paired Requests

Candidate-order tests should reuse the same permutations across models, and history-length tests should reuse the same users. This pairing allows the analysis to compare performance differences under each perturbation. Selecting different users for each history length confounds input truncation with user composition. Report the mean paired difference and its uncertainty, together with the settings under which the sign changes. Repeated generations for one user measure generation variability, not additional independent users. Uncertainty calculations should preserve that grouping, for example by resampling users while retaining their repeated outputs.
Repeated requests within a fixed model version measure generation variability; a later version introduces possible service drift. LLM judges add their own reliability and bias concerns [33]. Judge-swap comparisons should preserve item evidence and use recommendation-specific human or catalog-grounded calibration cases. These checks separate changes in the recommender from changes in the instrument measuring it.

7.4. Make Each Score Traceable

LensKit [92], RiVal [93], and Elliot [94] support reusable recommendation experiments. Accountability guidance [95] and the reproducibility survey [23] emphasize inspectable artifacts. For generative outputs, the crucial addition is a link from each stored prediction to its parsing, catalog resolution, and failure treatment, enabling the same outputs to be rescored.
An inspectable evaluation record links the request and candidate identifiers to the raw output, parsed catalog matches, validity decisions, per-request metrics, prompt, model version, and decoding configuration. These records expose whether a score changed because of generation, matching, or aggregation. When data or API restrictions prevent complete release, authors can still publish the evaluation code, configuration, permissible examples, and aggregate failure counts, stating which portions can be rerun.

8. Open Challenges

The protocol contrasts leave several questions unresolved. They concern the interaction between measurement choices and model behavior, rather than the availability of another metric.

8.1. How Does Item Resolution Redistribute Exposure?

The identifier studies establish that credit rules can change utility rankings. The corresponding effect on beyond-utility outcomes is less well characterized. Exact title matching may disproportionately reject aliases or rare items, while nearest-item repair may concentrate their credit on popular alternatives. A useful comparison would rescore fixed raw outputs under several documented resolution policies, measuring utility, long-tail exposure, and user-group failure rates together. It would distinguish a model’s recommendation behavior from the distribution induced by its evaluator. Catalog editions, translations, and many-to-one identifiers make this question relevant even when every generated string denotes an existing item.

8.2. When Does Benchmark Knowledge Help Generalization?

Extraction probes reveal recoverable benchmark content, but the observed association with accuracy does not isolate its contribution. Comparing familiar and newly released catalogs also changes domain, item difficulty, and metadata quality. A more discriminating design would match these properties while varying access to interaction records and catalog descriptions. Cold-item experiments introduce a related distinction: learning to use an unseen item’s attributes differs from making that item reachable through an index. Separating knowledge access, representation transfer, and retrieval eligibility would clarify when performance reflects generalization rather than benchmark-specific exposure.

8.3. Which Offline Outcomes Predict User Benefit?

User-centered frameworks distinguish system properties from perceived quality and experience [96,97,98]. The fairness studies make a related distinction between list similarity, preference alignment, and attribute recoverability. Whether improvements in these offline measurements translate into perceived relevance, control, or equitable benefit remains a separate empirical question.
Conversational recommendation provides a setting in which to examine that connection. Surveys organize preference elicitation and feedback [32,99]; ReDial supplies recommendation dialogues [100], and CRSLab provides conversational-system tooling [101]. Holding recommended items fixed while varying explanation or dialogue behavior would isolate changes in factual understanding, perceived control, and task completion. Varying items under a fixed dialogue policy would test a different pathway to user benefit.
LLM judges introduce a measurement question within such studies. MT-Bench and Chatbot Arena examine model judgments and human preferences [102]; long-context evaluations document sensitivity to information position [103]; HELM combines multiple metrics with recorded prompts and completions [104]. Recommendation evaluation needs to determine when a judge’s preference tracks catalog-grounded or user-observed quality and when it tracks explanation style. Separate interventions on item relevance, factual support, and wording would help identify that boundary.

8.4. How Do Generated Profiles Alter Feedback Loops?

Algorithmic confounding [105], feedback-loop bias amplification [106], and degenerating recommendation dynamics [107] show why repeated exposure differs from a static held-out test. Generated profiles add a possible feedback path: earlier recommendations can enter a later summary and be interpreted as evidence of user preference. The open question is whether this representation step amplifies or corrects the exposure effects already present in recommendation logs. A sequential comparison would hold the interaction stream and update opportunities fixed while varying how summaries distinguish observed preferences from prior system outputs. Intent-aware recommendation [108] and RS4Good [109] motivate assessing adaptation against changing user purposes and practical outcomes.

8.5. How Stable Are Conclusions Across Evaluators?

The reviewed reversals arise from particular candidate, credit, or reference choices. Their frequency and joint effect across system families remain unknown. Shared raw-output records would permit a crossed comparison: multiple recommenders evaluated under the same alternative candidate protocols, matching policies, and failure-accounting rules. Such a study could identify conclusions that survive evaluator changes and those driven by a specific policy. Fixed request panels and versioned outputs would further separate evaluator effects from stochastic generation and service drift. This moves beyond checking whether a result can be reproduced under one implementation to testing whether its interpretation is stable under scientifically justified alternatives.

9. Conclusion

The reviewed evidence leads to two conclusions. Protocol choices can change reported orderings: candidate pools, identifier credit, and fairness references each provide a concrete instance. Apparent disagreement can also reflect different measured objects: long-tail selection can coexist with popularity-skewed record recovery, while attribute recoverability and user benefit describe separate outcomes. Together, these findings make comparability central to evaluating LLM-based and generative recommenders. Progress is established by an improvement on a defined recommendation problem, with traceable item credit and an explicit outcome reference, rather than by a larger collection of favorable scores.

References

  1. Bauer, C.; Zangerle, E.; Said, A. Exploring the Landscape of Recommender Systems Evaluation: Practices and Perspectives. ACM Transactions on Recommender Systems 2024, 2, 1–31. [CrossRef]
  2. Ge, Y.; Liu, S.; Fu, Z.; Tan, J.; Li, Z.; Xu, S.; Li, Y.; Xian, Y.; Zhang, Y. A Survey on Trustworthy Recommender Systems. ACM Transactions on Recommender Systems 2025, 3, 1–68. [CrossRef]
  3. Klimashevskaia, A.; Jannach, D.; Elahi, M.; Trattner, C. A Survey on Popularity Bias in Recommender Systems. User Modeling and User-Adapted Interaction 2024, 34, 1777–1834. [CrossRef]
  4. Wu, Y.; Cao, J.; Xu, G. Fairness in Recommender Systems: Evaluation Approaches and Assurance Strategies. ACM Transactions on Knowledge Discovery from Data 2024, 18, 1–37. [CrossRef]
  5. McNee, S.M.; Riedl, J.; Konstan, J.A. Being accurate is not enough: how accuracy metrics have hurt recommender systems. In Proceedings of the CHI ’06 Extended Abstracts on Human Factors in Computing Systems. ACM, 2006, pp. 1097–1101. [CrossRef]
  6. Ge, M.; Delgado-Battenfeld, C.; Jannach, D. Beyond accuracy: evaluating recommender systems by coverage and serendipity. In Proceedings of the Proceedings of the fourth ACM conference on Recommender systems. ACM, 2010, pp. 257–260. [CrossRef]
  7. Geng, S.; Liu, S.; Fu, Z.; Ge, Y.; Zhang, Y. Recommendation as Language Processing (RLP): A Unified Pretrain, Personalized Prompt & Predict Paradigm (P5), 2022, [arXiv:cs.IR/2203.13366].
  8. Bao, K.; Zhang, J.; Zhang, Y.; Wang, W.; Feng, F.; He, X. TALLRec: An Effective and Efficient Tuning Framework to Align Large Language Model with Recommendation. In Proceedings of the Proceedings of the 17th ACM Conference on Recommender Systems, New York, NY, USA, 2023; pp. 1007–1014. [CrossRef]
  9. Ji, J.; Li, Z.; Xu, S.; Hua, W.; Ge, Y.; Tan, J.; Zhang, Y. GenRec: Large Language Model for Generative Recommendation. In Proceedings of the Advances in Information Retrieval, Cham, 2024; pp. 494–502. [CrossRef]
  10. Luo, S.; He, B.; Zhao, H.; Shao, W.; Qi, Y.; Huang, Y.; Zhou, A.; Yao, Y.; Li, Z.; Xiao, Y.; et al. RecRanker: Instruction Tuning Large Language Model as Ranker for Top-k Recommendation, 2024, [arXiv:cs.IR/2312.16018]. Version 3, 31 March 2024.
  11. Lin, F.; Hu, B.; Zheng, Z.; Zhu, X.; Liu, Z.; Zhang, Z.; Zhou, J.; Xu, T. Token-level Collaborative Alignment for LLM-based Generative Recommendation, 2026, [arXiv:cs.IR/2601.18457].
  12. Jiang, C.; Wang, J.; Ma, W.; Clarke, C.L.A.; Wang, S.; Wu, C.; Zhang, M. Beyond Utility: Evaluating LLM as Recommender. In Proceedings of the Proceedings of the ACM on Web Conference 2025, New York, NY, USA, 2025; pp. 3850–3862. [CrossRef]
  13. Di Palma, D.; Merra, F.A.; Sfilio, M.; Anelli, V.W.; Narducci, F.; Di Noia, T. Do LLMs Memorize Recommendation Datasets? A Preliminary Study on MovieLens-1M. In Proceedings of the Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, New York, NY, USA, 2025; pp. 2582–2586. [CrossRef]
  14. Ferrari Dacrema, M.; Cremonesi, P.; Jannach, D. Are We Really Making Much Progress? A Worrying Analysis of Recent Neural Recommendation Approaches. In Proceedings of the Proceedings of the 13th ACM Conference on Recommender Systems, New York, NY, USA, 2019; pp. 101–109. [CrossRef]
  15. Ferrari Dacrema, M.; Boglio, S.; Cremonesi, P.; Jannach, D. A Troubling Analysis of Reproducibility and Progress in Recommender Systems Research. ACM Transactions on Information Systems 2021, 39, 1–49. [CrossRef]
  16. Krichene, W.; Rendle, S. On Sampled Metrics for Item Recommendation. In Proceedings of the Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, New York, NY, USA, 2020; pp. 1748–1757. [CrossRef]
  17. Zhao, W.X.; Mu, S.; Hou, Y.; Lin, Z.; Chen, Y.; Pan, X.; Li, K.; Lu, Y.; Wang, H.; Tian, C.; et al. RecBole: Towards a Unified, Comprehensive and Efficient Framework for Recommendation Algorithms. In Proceedings of the Proceedings of the 30th ACM International Conference on Information and Knowledge Management, New York, NY, USA, 2021; pp. 4653–4664. [CrossRef]
  18. Shehzad, F.; Breuer, T.; Maistro, M.; Jannach, D. "We Share Our Code Online": Why This Is Not Enough to Ensure Reproducibility and Progress in Recommender Systems Research. In Proceedings of the Proceedings of the 19th ACM Conference on Recommender Systems, New York, NY, USA, 2025; pp. 884–893. [CrossRef]
  19. Wu, L.; Zheng, Z.; Qiu, Z.; Wang, H.; Gu, H.; Shen, T.; Qin, C.; Zhu, C.; Zhu, H.; Liu, Q.; et al. A Survey on Large Language Models for Recommendation. World Wide Web 2024, 27. [CrossRef]
  20. Zhao, Z.; Fan, W.; Li, J.; Liu, Y.; Mei, X.; Wang, Y.; Wen, Z.; Wang, F.; Zhao, X.; Tang, J.; et al. Recommender Systems in the Era of Large Language Models (LLMs). IEEE Transactions on Knowledge and Data Engineering 2024, 36, 6889–6907. [CrossRef]
  21. Li, L.; Zhang, Y.; Liu, D.; Chen, L. Large Language Models for Generative Recommendation: A Survey and Visionary Discussions. In Proceedings of the Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), Torino, Italia, 2024; pp. 10146–10159.
  22. Jannach, D.; Chen, L. Improving Methodological Standards in Recommender Systems Offline Evaluation. ACM Transactions on Recommender Systems 2026, 4. [CrossRef]
  23. Said, A.; Bellogín, A. Reproducibility in Recommender Systems: A Survey. ACM Transactions on Recommender Systems 2026. Advance online publication, . [CrossRef]
  24. da Silva, D.C.; Jannach, D. Calibrated Recommendations: Survey and Future Directions. ACM Transactions on Recommender Systems 2026, 4, 1–32. [CrossRef]
  25. Zhang, Q.; Szymanski, L.; Deng, J.D.; Zhang, H. Faithful Evaluation of Semantic-ID Tokenizers for Generative Recommendation, 2026, [arXiv:cs.IR/2605.25330]. Version 2, 22 August 2026.
  26. Zhang, Z.; Zhao, J.; Ma, X.; Xin, X.; de Rijke, M.; Ren, Z. Cold-Starts in Generative Recommendation: A Reproducibility Study, 2026, [arXiv:cs.IR/2603.29845]. Version 2, 6 April 2026.
  27. Zhang, J.; Bao, K.; Zhang, Y.; Wang, W.; Feng, F.; He, X. Is ChatGPT Fair for Recommendation? Evaluating Fairness in Large Language Model Recommendation. In Proceedings of the Proceedings of the 17th ACM Conference on Recommender Systems. ACM, 2023, pp. 993–999. [CrossRef]
  28. Deldjoo, Y.; di Noia, T. CFaiRLLM: Consumer Fairness Evaluation in Large-Language Model Recommender System. ACM Transactions on Intelligent Systems and Technology 2025, 16, 1–31. [CrossRef]
  29. Hua, W.; Ge, Y.; Xu, S.; Ji, J.; Zhang, Y. UP5: Unbiased Foundation Model for Fairness-aware Recommendation. In Proceedings of the Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, 2024, pp. 1899–1912. [CrossRef]
  30. Deldjoo, Y. Understanding Biases in ChatGPT-based Recommender Systems: Provider Fairness, Temporal Stability, and Recency. ACM Transactions on Recommender Systems 2026, 4, 1–35. [CrossRef]
  31. Zhu, X.; Wang, Y.; Gao, H.; Xu, W.; Wang, C.; Liu, Z.; Wang, K.; Jin, M.; Pang, L.; Wen, Q.; et al. Recommender Systems Meet Large Language Model Agents: A Survey. Foundations and Trends in Privacy and Security 2025, 7, 247–396. [CrossRef]
  32. Jannach, D. Evaluating Conversational Recommender Systems: A Landscape of Research. Artificial Intelligence Review 2023, 56, 2365–2400. [CrossRef]
  33. Gu, J.; Jiang, X.; Shi, Z.; Tan, H.; Zhai, X.; Xu, C.; Li, W.; Shen, Y.; Ma, S.; Liu, H.; et al. A Survey on LLM-as-a-Judge, 2024, [arXiv:cs.CL/2411.15594]. [CrossRef]
  34. Xu, J.; Chen, Z.; Yang, S.; Li, J.; Wang, W.; Hu, X.; Hoi, S.; Ngai, E.C.H. A Survey on Multimodal Recommender Systems: Recent Advances and Future Directions, 2025, [arXiv:cs.IR/2502.15711].
  35. Xu, J.; Chen, Z.; Li, J.; Yang, S.; Wang, W.; Hu, X.; Wong, R.C.W.; Ngai, E.C.H. Enhancing Robustness and Generalization Capability for Multimodal Recommender Systems via Sharpness-Aware Minimization. IEEE Transactions on Knowledge and Data Engineering 2025, 37, 6406–6419. [CrossRef]
  36. Xu, J.; Chen, Z.; Wang, W.; Hu, X.; Kim, S.W.; Ngai, E.C.H. COHESION: Composite Graph Convolutional Network with Dual-Stage Fusion for Multimodal Recommendation. In Proceedings of the Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, New York, NY, USA, 2025; pp. 1830–1839. [CrossRef]
  37. Xu, J.; Chen, Z.; Yang, S.; Li, J.; Ngai, E.C.H. The Best is Yet to Come: Graph Convolution in the Testing Phase for Multimodal Recommendation, 2025, [arXiv:cs.IR/2507.18489].
  38. Xu, J.; Chen, Z.; Li, J.; Yang, S.; Wang, H.; Ngai, E.C.H. AlignGroup: Learning and Aligning Group Consensus with Member Preferences for Group Recommendation. In Proceedings of the Proceedings of the 33rd ACM International Conference on Information and Knowledge Management, New York, NY, USA, 2024; pp. 2682–2691. [CrossRef]
  39. Xu, J.; Chen, Z.; Li, J.; Yang, S.; Wang, W.; Wang, H.; Li, Y.; Hu, X.; Ngai, E.C.H. DGGVAE: Dual-Granularity Graph Variational Auto-Encoder for Group Recommendation. ACM Transactions on Information Systems 2026, 44, 1–30. [CrossRef]
  40. Attimonelli, M.; De Bellis, A.; Pomo, C.; Jannach, D.; Di Sciascio, E.; Di Noia, T. Do We Really Need Specialization? Evaluating Generalist Text Embeddings for Zero-Shot Recommendation and Search. In Proceedings of the Proceedings of the 19th ACM Conference on Recommender Systems, New York, NY, USA, 2025; pp. 575–580. [CrossRef]
  41. Herlocker, J.L.; Konstan, J.A.; Terveen, L.G.; Riedl, J.T. Evaluating collaborative filtering recommender systems. ACM Transactions on Information Systems 2004, 22, 5–53. [CrossRef]
  42. Shani, G.; Gunawardana, A. Evaluating Recommendation Systems. In Recommender Systems Handbook; Springer US, 2011; pp. 257–297. [CrossRef]
  43. Zangerle, E.; Bauer, C. Evaluating Recommender Systems: Survey and Framework. ACM Computing Surveys 2023, 55, 1–38. [CrossRef]
  44. Gunawardana, A.; Shani, G. A Survey of Accuracy Evaluation Metrics of Recommendation Tasks. Journal of Machine Learning Research 2009, 10, 2935–2962.
  45. Järvelin, K.; Kekäläinen, J. Cumulated gain-based evaluation of IR techniques. ACM Transactions on Information Systems 2002, 20, 422–446. [CrossRef]
  46. Cremonesi, P.; Koren, Y.; Turrin, R. Performance of recommender algorithms on top-n recommendation tasks. In Proceedings of the Proceedings of the fourth ACM conference on Recommender systems. ACM, 2010, pp. 39–46. [CrossRef]
  47. Bellogín, A.; Castells, P.; Cantador, I. Statistical biases in Information Retrieval metrics for recommender systems. Information Retrieval Journal 2017, 20, 606–634. [CrossRef]
  48. Zhao, W.X.; Chen, J.; Wang, P.; Gu, Q.; Wen, J.R. Revisiting Alternative Experimental Settings for Evaluating Top-N Item Recommendation Algorithms. In Proceedings of the Proceedings of the 29th ACM International Conference on Information & Knowledge Management. ACM, 2020, pp. 2329–2332. [CrossRef]
  49. Tamm, Y.M.; Damdinov, R.; Vasilev, A. Quality Metrics in Recommender Systems: Do We Calculate Metrics Consistently? In Proceedings of the Fifteenth ACM Conference on Recommender Systems. ACM, 2021, pp. 708–713. [CrossRef]
  50. Castells, P.; Moffat, A. Offline recommender system evaluation: Challenges and new directions. AI Magazine 2022, 43, 225–238. [CrossRef]
  51. Hou, Y.; Zhang, J.; Lin, Z.; Lu, H.; Xie, R.; McAuley, J.; Zhao, W.X. Large Language Models are Zero-Shot Rankers for Recommender Systems. In Advances in Information Retrieval; Lecture Notes in Computer Science, Springer Nature Switzerland, 2024; pp. 364–381. [CrossRef]
  52. Rajput, S.; Mehta, N.; Singh, A.; Hulikal Keshavan, R.; Vu, T.; Heldt, L.; Hong, L.; Tay, Y.; Tran, V.; Samost, J.; et al. Recommender Systems with Generative Retrieval. In Proceedings of the Advances in Neural Information Processing Systems, 2023, Vol. 36, pp. 10299–10315. [CrossRef]
  53. Wang, W.; Bao, H.; Lin, X.; Zhang, J.; Li, Y.; Feng, F.; Ng, S.K.; Chua, T.S. Learnable Item Tokenization for Generative Recommendation. In Proceedings of the Proceedings of the 33rd ACM International Conference on Information and Knowledge Management. ACM, 2024, pp. 2400–2409. [CrossRef]
  54. Hou, Y.; Mu, S.; Zhao, W.X.; Li, Y.; Ding, B.; Wen, J.R. Towards Universal Sequence Representation Learning for Recommender Systems. In Proceedings of the Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. ACM, 2022, pp. 585–593. [CrossRef]
  55. Li, J.; Wang, M.; Li, J.; Fu, J.; Shen, X.; Shang, J.; McAuley, J. Text Is All You Need: Learning Language Representations for Sequential Recommendation. In Proceedings of the Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining. ACM, 2023, pp. 1258–1267. [CrossRef]
  56. Xu, S.; Hua, W.; Zhang, Y. OpenP5: An Open-Source Platform for Developing, Training, and Evaluating LLM-based Recommender Systems. In Proceedings of the Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 2024, pp. 386–394. [CrossRef]
  57. Kaminskas, M.; Bridge, D. Diversity, Serendipity, Novelty, and Coverage: A Survey and Empirical Analysis of Beyond-Accuracy Objectives in Recommender Systems. ACM Transactions on Interactive Intelligent Systems 2017, 7, 1–42. [CrossRef]
  58. Vargas, S.; Castells, P. Rank and relevance in novelty and diversity metrics for recommender systems. In Proceedings of the Proceedings of the fifth ACM conference on Recommender systems. ACM, 2011, pp. 109–116. [CrossRef]
  59. Ziegler, C.N.; McNee, S.M.; Konstan, J.A.; Lausen, G. Improving recommendation lists through topic diversification. In Proceedings of the Proceedings of the 14th international conference on World Wide Web - WWW ’05. ACM Press, 2005, pp. 22–32. [CrossRef]
  60. Carbonell, J.; Goldstein, J. The use of MMR, diversity-based reranking for reordering documents and producing summaries. In Proceedings of the Proceedings of the 21st annual international ACM SIGIR conference on Research and development in information retrieval. ACM, 1998, pp. 335–336. [CrossRef]
  61. Steck, H. Calibrated recommendations. In Proceedings of the Proceedings of the 12th ACM Conference on Recommender Systems. ACM, 2018, pp. 154–162. [CrossRef]
  62. Kaya, M.; Bridge, D. A comparison of calibrated and intent-aware recommendations. In Proceedings of the Proceedings of the 13th ACM Conference on Recommender Systems. ACM, 2019, pp. 151–159. [CrossRef]
  63. Abdollahpouri, H.; Burke, R.; Mobasher, B. Controlling Popularity Bias in Learning-to-Rank Recommendation. In Proceedings of the Proceedings of the Eleventh ACM Conference on Recommender Systems. ACM, 2017, pp. 42–46. [CrossRef]
  64. Abdollahpouri, H.; Mansoury, M.; Burke, R.; Mobasher, B.; Malthouse, E. User-centered Evaluation of Popularity Bias in Recommender Systems. In Proceedings of the Proceedings of the 29th ACM Conference on User Modeling, Adaptation and Personalization. ACM, 2021, pp. 119–129. [CrossRef]
  65. Abdollahpouri, H.; Adomavicius, G.; Burke, R.; Guy, I.; Jannach, D.; Kamishima, T.; Krasnodebski, J.; Pizzato, L. Multistakeholder recommendation: Survey and research directions. User Modeling and User-Adapted Interaction 2020, 30, 127–158. [CrossRef]
  66. Singh, A.; Joachims, T. Fairness of Exposure in Rankings. In Proceedings of the Proceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. ACM, 2018, pp. 2219–2228. [CrossRef]
  67. Biega, A.J.; Gummadi, K.P.; Weikum, G. Equity of Attention: Amortizing Individual Fairness in Rankings. In Proceedings of the The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval. ACM, 2018, pp. 405–414. [CrossRef]
  68. Beutel, A.; Chen, J.; Doshi, T.; Qian, H.; Wei, L.; Wu, Y.; Heldt, L.; Zhao, Z.; Hong, L.; Chi, E.H.; et al. Fairness in Recommendation Ranking through Pairwise Comparisons. In Proceedings of the Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining. ACM, 2019, pp. 2212–2220. [CrossRef]
  69. Patro, G.K.; Biswas, A.; Ganguly, N.; Gummadi, K.P.; Chakraborty, A. FairRec: Two-Sided Fairness for Personalized Recommendations in Two-Sided Platforms. In Proceedings of the Proceedings of The Web Conference 2020. ACM, 2020, pp. 1194–1204. [CrossRef]
  70. Narayanan, A.; Shmatikov, V. Robust De-anonymization of Large Sparse Datasets. In Proceedings of the 2008 IEEE Symposium on Security and Privacy (sp 2008). IEEE, 2008, pp. 111–125. [CrossRef]
  71. Calandrino, J.A.; Kilzer, A.; Narayanan, A.; Felten, E.W.; Shmatikov, V. "You Might Also Like:" Privacy Risks of Collaborative Filtering. In Proceedings of the 2011 IEEE Symposium on Security and Privacy. IEEE, 2011, pp. 231–246. [CrossRef]
  72. Shokri, R.; Stronati, M.; Song, C.; Shmatikov, V. Membership Inference Attacks Against Machine Learning Models. In Proceedings of the 2017 IEEE Symposium on Security and Privacy (SP). IEEE, 2017, pp. 3–18. [CrossRef]
  73. Carlini, N.; Tramèr, F.; Wallace, E.; Jagielski, M.; Herbert-Voss, A.; Lee, K.; Roberts, A.; Brown, T.; Song, D.; Erlingsson, Ú.; et al. Extracting Training Data from Large Language Models. In Proceedings of the 30th USENIX Security Symposium (USENIX Security 21). USENIX Association, 2021, pp. 2633–2650.
  74. Dwork, C.; McSherry, F.; Nissim, K.; Smith, A. Calibrating Noise to Sensitivity in Private Data Analysis. In Theory of Cryptography; Lecture Notes in Computer Science, Springer Berlin Heidelberg, 2006; pp. 265–284. [CrossRef]
  75. Tintarev, N.; Masthoff, J. A Survey of Explanations in Recommender Systems. In Proceedings of the 2007 IEEE 23rd International Conference on Data Engineering Workshop. IEEE, 2007, pp. 801–810. [CrossRef]
  76. Zhang, Y.; Chen, X. Explainable Recommendation: A Survey and New Perspectives. Foundations and Trends in Information Retrieval 2020, 14, 1–101. [CrossRef]
  77. Balog, K.; Radlinski, F.; Arakelyan, S. Transparent, Scrutable and Explainable User Models for Personalized Recommendation. In Proceedings of the Proceedings of the 42nd International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 2019, pp. 265–274. [CrossRef]
  78. Xu, J.; Chen, Z.; Li, J.; Yang, S.; Wang, H.; Li, Y.; Li, M.; Wu, P.; Ngai, E.C.H. MDVT: Enhancing Multimodal Recommendation with Model-Agnostic Multimodal-Driven Virtual Triplets, 2025, [arXiv:cs.IR/2505.16665].
  79. Dai, S.; Shao, N.; Zhao, H.; Yu, W.; Si, Z.; Xu, C.; Sun, Z.; Zhang, X.; Xu, J. Uncovering ChatGPT’s Capabilities in Recommender Systems. In Proceedings of the Proceedings of the 17th ACM Conference on Recommender Systems, New York, NY, USA, 2023; pp. 1126–1132. [CrossRef]
  80. Sanner, S.; Balog, K.; Radlinski, F.; Wedin, B.; Dixon, L. Large Language Models are Competitive Near Cold-start Recommenders for Language- and Item-based Preferences. In Proceedings of the Proceedings of the 17th ACM Conference on Recommender Systems, New York, NY, USA, 2023; pp. 890–896. [CrossRef]
  81. Zhang, G.; Yuan, G.; Cheng, D.; Liu, L.; Li, J.; Zhang, S. Mitigating Propensity Bias of Large Language Models for Recommender Systems. ACM Transactions on Information Systems 2025, 43. [CrossRef]
  82. Harper, F.M.; Konstan, J.A. The MovieLens Datasets: History and Context. ACM Transactions on Interactive Intelligent Systems 2016, 5, 1–19. [CrossRef]
  83. Steck, H. Training and testing of recommender systems on data missing not at random. In Proceedings of the Proceedings of the 16th ACM SIGKDD international conference on Knowledge discovery and data mining. ACM, 2010, pp. 713–722. [CrossRef]
  84. Schnabel, T.; Swaminathan, A.; Singh, A.; Chandak, N.; Joachims, T. Recommendations as Treatments: Debiasing Learning and Evaluation. In Proceedings of the International Conference on Machine Learning. PMLR, 2016, Vol. 48, Proceedings of Machine Learning Research, pp. 1670–1679.
  85. Li, L.; Chu, W.; Langford, J.; Wang, X. Unbiased offline evaluation of contextual-bandit-based news article recommendation algorithms. In Proceedings of the Proceedings of the fourth ACM international conference on Web search and data mining. ACM, 2011, pp. 297–306. [CrossRef]
  86. Swaminathan, A.; Joachims, T. Counterfactual Risk Minimization: Learning from Logged Bandit Feedback. In Proceedings of the International Conference on Machine Learning. PMLR, 2015, Vol. 37, Proceedings of Machine Learning Research, pp. 814–823.
  87. Joachims, T.; Swaminathan, A.; Schnabel, T. Unbiased Learning-to-Rank with Biased Feedback. In Proceedings of the Proceedings of the Tenth ACM International Conference on Web Search and Data Mining. ACM, 2017, pp. 781–789. [CrossRef]
  88. Said, A.; Bellogín, A. Comparative recommender system evaluation: benchmarking recommendation frameworks. In Proceedings of the Proceedings of the 8th ACM Conference on Recommender systems. ACM, 2014, pp. 129–136. [CrossRef]
  89. Sun, Z.; Yu, D.; Fang, H.; Yang, J.; Qu, X.; Zhang, J.; Geng, C. Are We Evaluating Rigorously? Benchmarking Recommendation for Reproducible Evaluation and Fair Comparison. In Proceedings of the Fourteenth ACM Conference on Recommender Systems. ACM, 2020, pp. 23–32. [CrossRef]
  90. Rendle, S.; Krichene, W.; Zhang, L.; Anderson, J. Neural Collaborative Filtering vs. Matrix Factorization Revisited. In Proceedings of the Fourteenth ACM Conference on Recommender Systems. ACM, 2020, pp. 240–248. [CrossRef]
  91. Ludewig, M.; Jannach, D. Evaluation of session-based recommendation algorithms. User Modeling and User-Adapted Interaction 2018, 28, 331–390. [CrossRef]
  92. Ekstrand, M.D.; Ludwig, M.; Konstan, J.A.; Riedl, J.T. Rethinking the recommender research ecosystem: reproducibility, openness, and LensKit. In Proceedings of the Proceedings of the fifth ACM conference on Recommender systems. ACM, 2011, pp. 133–140. [CrossRef]
  93. Said, A.; Bellogín, A. Rival: a toolkit to foster reproducibility in recommender system evaluation. In Proceedings of the Proceedings of the 8th ACM Conference on Recommender systems. ACM, 2014, pp. 371–372. [CrossRef]
  94. Anelli, V.W.; Bellogin, A.; Ferrara, A.; Malitesta, D.; Merra, F.A.; Pomo, C.; Donini, F.M.; Di Noia, T. Elliot: A Comprehensive and Rigorous Framework for Reproducible Recommender Systems Evaluation. In Proceedings of the Proceedings of the 44th International ACM SIGIR Conference on Research and Development in Information Retrieval. ACM, 2021, pp. 2405–2414. [CrossRef]
  95. Bellogín, A.; Said, A. Improving accountability in recommender systems research through reproducibility. User Modeling and User-Adapted Interaction 2021, 31, 941–977. [CrossRef]
  96. Knijnenburg, B.P.; Willemsen, M.C.; Gantner, Z.; Soncu, H.; Newell, C. Explaining the user experience of recommender systems. User Modeling and User-Adapted Interaction 2012, 22, 441–504. [CrossRef]
  97. Pu, P.; Chen, L.; Hu, R. A user-centric evaluation framework for recommender systems. In Proceedings of the Proceedings of the fifth ACM conference on Recommender systems. ACM, 2011, pp. 157–164. [CrossRef]
  98. Konstan, J.A.; Riedl, J. Recommender systems: from algorithms to user experience. User Modeling and User-Adapted Interaction 2012, 22, 101–123. [CrossRef]
  99. Jannach, D.; Manzoor, A.; Cai, W.; Chen, L. A Survey on Conversational Recommender Systems. ACM Computing Surveys 2022, 54, 1–36. [CrossRef]
  100. Li, R.; Ebrahimi Kahou, S.; Schulz, H.; Michalski, V.; Charlin, L.; Pal, C. Towards Deep Conversational Recommendations. In Proceedings of the Advances in Neural Information Processing Systems, 2018, Vol. 31.
  101. Zhou, K.; Wang, X.; Zhou, Y.; Shang, C.; Cheng, Y.; Zhao, W.X.; Li, Y.; Wen, J.R. CRSLab: An Open-Source Toolkit for Building Conversational Recommender System. In Proceedings of the Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing: System Demonstrations. Association for Computational Linguistics, 2021, pp. 185–193. [CrossRef]
  102. Zheng, L.; Chiang, W.L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.; et al. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. In Proceedings of the Advances in Neural Information Processing Systems 36. Neural Information Processing Systems Foundation, Inc. (NeurIPS), 2023, pp. 46595–46623. [CrossRef]
  103. Liu, N.F.; Lin, K.; Hewitt, J.; Paranjape, A.; Bevilacqua, M.; Petroni, F.; Liang, P. Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics 2024, 12, 157–173. [CrossRef]
  104. Liang, P.; Bommasani, R.; Lee, T.; Tsipras, D.; Soylu, D.; Yasunaga, M.; Zhang, Y.; Narayanan, D.; Wu, Y.; Kumar, A.; et al. Holistic Evaluation of Language Models, 2022, [2211.09110].
  105. Chaney, A.J.B.; Stewart, B.M.; Engelhardt, B.E. How algorithmic confounding in recommendation systems increases homogeneity and decreases utility. In Proceedings of the Proceedings of the 12th ACM Conference on Recommender Systems. ACM, 2018, pp. 224–232. [CrossRef]
  106. Mansoury, M.; Abdollahpouri, H.; Pechenizkiy, M.; Mobasher, B.; Burke, R. Feedback Loop and Bias Amplification in Recommender Systems. In Proceedings of the Proceedings of the 29th ACM International Conference on Information & Knowledge Management. ACM, 2020, pp. 2145–2148. [CrossRef]
  107. Jiang, R.; Chiappa, S.; Lattimore, T.; György, A.; Kohli, P. Degenerate Feedback Loops in Recommender Systems. In Proceedings of the Proceedings of the 2019 AAAI/ACM Conference on AI, Ethics, and Society. ACM, 2019, pp. 383–390. [CrossRef]
  108. Jannach, D.; Zanker, M. A Survey on Intent-aware Recommender Systems. ACM Transactions on Recommender Systems 2025, 3, 1–32. [CrossRef]
  109. Jannach, D.; Said, A.; Tkalčič, M.; Zanker, M. Recommender Systems for Good (RS4Good): Survey of Use Cases and a Call to Action for Research that Matters: Opinion Article. ACM Transactions on Recommender Systems 2026. Advance online publication, . [CrossRef]
Figure 1. Analytical questions organizing the protocol comparison. A study can inform several questions: candidates and held-out entities define access, matching defines item credit, and the outcome reference defines the benefit being compared.
Figure 1. Analytical questions organizing the protocol comparison. A study can inform several questions: candidates and held-out entities define access, matching defines item credit, and the outcome reference defines the benefit being compared.
Preprints 233950 g001
Figure 2. Validity conditions shape the evaluation claim. Protocol choices and output-credit rules affect parallel utility and beyond-utility measurements, each with an explicit population and denominator. The schematic item-output record retains failed attempts. Arrows denote measurement dependencies, not estimated causal effects.
Figure 2. Validity conditions shape the evaluation claim. Protocol choices and output-credit rules affect parallel utility and beyond-utility measurements, each with an explicit population and denominator. The schematic item-output record retains failed attempts. Arrows denote measurement dependencies, not estimated causal effects.
Preprints 233950 g002
Figure 3. Input–output contracts of evaluated components. The schematic examples show preference prediction, candidate ordering, feature enhancement, item generation, and interaction. The identifier collision illustrates the distinction between a valid output and a uniquely identified item.
Figure 3. Input–output contracts of evaluated components. The schematic examples show preference prediction, candidate ordering, feature enhancement, item generation, and interaction. The identifier collision illustrates the distinction between a valid output and a uniquely identified item.
Preprints 233950 g003
Table 1. Prior contributions and the evaluation decisions connected by this survey.
Table 1. Prior contributions and the evaluation decisions connected by this survey.
Source and focus Established contribution Connection developed in this survey
Jannach and Chen [22]: TORS methodology Editorial guidance emphasizing tuned comparisons and complete reproducibility materials. Connects comparison rigor to candidate eligibility, item resolution, and request retention; identifies which decision changes the estimand before tuning is compared.
Bauer et al. [1]: evaluation landscape Systematic review of 57 papers; codes experiment types, datasets, metrics, and contributions (Sections 2–3). Relates concrete protocol differences to the claim a generative recommender can support, rather than estimating practice frequencies.
Said and Bellogín [23]: reproducibility Collects 51 track papers and analyzes 50 after excluding one non-recommender study; examines evaluation and artifact practices. Connects reusable artifacts to the candidate, output-resolution, and aggregation decisions needed to reconstruct a comparison.
da Silva and Jannach [24]: calibration Separates calibration mechanisms from empirical and analytical evidence of their effectiveness. Identifies which preference distribution a generated profile preserves and how unresolved items affect the measured distribution.
Ge et al. [2]: trustworthiness Organizes explainability, fairness, privacy, robustness, controllability, and their relationships. Distinguishes delivered exposure from model knowledge and representation diagnostics, and specifies the item identities and request populations needed for outcome comparisons.
Li et al. [21]: generative recommendation Defines generative tasks and item representations; discusses factual item existence among the challenges. Separates item existence, unique item identity, relevance, and failure aggregation as distinct evaluation decisions.
Jiang et al. [12]: LLM evaluation Empirically evaluates history sensitivity, candidate position, generation, and hallucination (Sections 3–4). Separates candidate membership from presentation and connects both to conditional versus pipeline claims.
Ferrari Dacrema et al. [15]: reproducibility Re-examines neural recommenders with tuned baselines and source-specific protocols (Sections 3–4). Extends comparison validity to prompts, parsing, generation failures, and the population entering each score.
Zhang et al. [25]: identifier evaluation Re-evaluates semantic-ID matches at item level and reports changed tokenizer orderings (Section 4.3). Connects identifier ambiguity to catalog matching and denominator effects across different output formats.
Table 3. Selection, decoding, and knowledge-extraction findings. Differences in the observed object explain why popularity and diversity results need not agree across protocols. Fairness references are compared in Table 5.
Table 3. Selection, decoding, and knowledge-extraction findings. Differences in the observed object explain why popularity and diversity results need not agree across protocols. Fairness references are compared in Table 5.
Study and operation Source-reported finding Cross-paper interpretation
Jiang et al. [12], Section 4.3: candidate selection LLM configurations can favor long-tail items; better source-defined serendipity can coexist with lower accuracy in Beauty. Sampled-candidate outcomes. The displayed serendipity formula (Eq. 6) has no explicit relevance term; it does not establish useful novelty.
Jiang et al., Section 4.6: profile generation Generated-profile inputs alter reported accuracy and popularity behavior; profile-plus-history improves the reported dimensions. Changing the user representation changes recommendation behavior; explanation faithfulness is a separate question.
TIGER [52], Table 3: decoding On Beauty, temperature 1.0 versus 2.0 gives category Entropy@10 of 0.76 versus 1.38. Category diversity responds to sampling; the table does not supply paired accuracy for a trade-off frontier.
Di Palma et al. [13], Section 3.4: record extraction Popular MovieLens items are more readily recovered than rarely interacted items. Recoverability concerns model knowledge, whereas the ranking studies measure selection among available alternatives.
Table 6. Candidate access and the prediction task. A supplied target, a constructed candidate set, and a retrieved set support different comparisons even when their scores share a metric name.
Table 6. Candidate access and the prediction task. A supplied target, a constructed candidate set, and a retrieved set support different comparisons even when their scores share a metric name.
Study Task and population Candidate and scoring protocol Consequence for interpretation
TALLRec [8], Section 3 MovieLens100K and BookCrossing derivatives; 8:1:1 samples. Supplied target; ten context items; binary AUC; few-shot tuning. Preference discrimination is distinct from finding a relevant item in a catalog.
Dai et al. [79], Section 4.1 MovieLens-1M, Amazon Books/CDs, MIND-small; 500 sampled records per dataset. One positive + four negatives; NDCG/MRR@3; compliance rate. Compliance concerns the supplied set; five-item ranking does not measure catalog retrieval.
Jiang et al. [12], Sections 3.2–4.8 Beauty, Sports, MovieLens-1M, LastFM; leave-one-out ranking and reranking. Sampled ranking: one positive + 19 negatives; history and position interventions. Candidate membership defines eligibility; candidate order tests presentation sensitivity.
LLMRank [51], Sections 2–3 Sequential recommendation with history and candidate-order interventions. Main sampled test: one positive + 19 negatives; hard-negative and retrieved-candidate tests. Sampled and retrieved sets must be distinguished within the same study.
RecRanker [10], Sections 4.3–5.1 MovieLens-100K/1M and BookCrossing; leave-one-out. Top ten retrieved candidates; HR/NDCG@3,5; position shifting and hybrid ranking. Retrieval success and conditional ordering are separate components of performance.
Sanner et al. [80], Sections 3–5 153 movie raters; elicited preferences and pool-specific relevance judgments. 40-item combined pool; 20-item popularity-stratified random subpool; NDCG@10. Complete judgments do not make pool selection neutral; EASE/LLM ordering changes across pools.
Table 9. Reporting checklist for LLM-based and generative recommender-system evaluation.
Table 9. Reporting checklist for LLM-based and generative recommender-system evaluation.
Category Minimum report Strong report
Data Dataset release, filtering, and split protocol. Temporal rationale, contamination probe, and processed-data artifact.
Candidates Candidate source, size, and ordering policy. Multiple candidate sizes and shuffled-order variance.
History Length, order, and selection rule for interactions. Paired length and selection sweeps on fixed users.
Prompt Full prompt template and examples. Prompt ablations and prompt repository.
Model Model name and access date. Versioned model identifier, repeated-run logs, and open-model comparison.
Decoding Temperature, top-p, max tokens, and parser. Deterministic and stochastic sensitivity analysis.
Baselines Implementations, search spaces, budgets, and validation criterion. Tuning-sensitivity and cost reports under a common evaluation protocol.
Metrics Definitions, outcome references, units, and denominators. Strategy-order sensitivity to justified alternative benefit references.
Generated output Matching rule, invalid-output counts, and metric denominators. Conditional and end-to-end quality, matching audit, and faithfulness checks.
Human or judge evaluation Participants or judge model and rubric. Calibration, agreement, failure cases, and judge-swap sensitivity.
Artifacts Code, processed data, and main configuration. Exact splits, prompts, logs, model/config files, and supplement coding sheet.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.