Preprint
Review

This version is not peer-reviewed.

Generative Recommendation in Dynamic Catalogs: A Survey of Item Admission, Semantic ID Evolution, and Model Adaptation

Submitted:

18 September 2026

Posted:

20 September 2026

You are already at the latest version

Abstract
Generative recommendation must connect predicted identifiers to a catalog that changes after training. New items precede feedback, identifier assignments evolve, and request rules constrain which outputs can be served. This survey compares item admission, semantic-ID evolution, and model adaptation through their update dependencies and dynamic operating conditions. A coupled-state framework connects catalog events to item assignments, model parameters, histories, and access structures. We use it to explain when metadata retrieval, targeted editing, continual learning, recoding, and request filtering address the relevant bottleneck. Three findings synthesize interface dependencies and their operating conditions: access and feedback learning require separate tests; recoding requires coordinated model and resolver handoffs; candidate access, item resolution, and eligibility expose different failures. From these findings and the heterogeneous cost evidence, we derive an evaluation principle: assess update quality against complete cost within the item's available lifetime. Six source-located comparison groups show when cohort choice reverses a baseline preference, mapping repair leaves learning unresolved, and evaluation or filtering changes the meaning of a gain. A conditional method-selection table connects these contrasts to the information, migration work, and service window required by each route. The resulting evaluation guidance preserves early access-only intervals, measures quality by item age, and accounts for preparation, updates, and serving. Open questions concern when to switch or combine update routes, how to migrate identifier versions, and how admission translates into useful exposure.
Keywords: 
;  ;  ;  ;  ;  

1. Introduction

A newly listed product can be available for purchase before a recommender has observed any interaction with it. A live room can retain its platform identity while its content changes from conversation to product demonstration. An established product can acquire a different audience without changing its description. These events change different parts of the recommendation problem. Generative recommenders must connect their predicted token sequences to the items available at the time of a request, while learning from behavior recorded under earlier catalog states. Recent cold-start studies and live-streaming deployments make these requirements concrete [1,2,3].
Semantic identifiers provide a compact interface between items and autoregressive generation. A tokenizer maps an item representation to a sequence of discrete codes, and a recommender predicts that sequence from user history. TIGER established this approach for recommendation, and subsequent tokenization work incorporated collaborative information into identifier learning [4,5]. The modularity of this interface has also motivated practical frameworks such as GRID [6]. Its operational consequence is a dependency: a model learns the meaning of the identifier assignments on which it is trained. Changing the assignments can therefore change the prediction problem even when the model architecture and token vocabulary remain fixed.
Several research lines now address different consequences of catalog evolution. SpecGR uses inductive candidates to support new-item recommendation [7]. LIGER combines generative and dense retrieval [8]. DACT selectively adapts identifiers as collaborative signals drift [9]. GenRecEdit modifies model behavior for cold items through targeted edits [10], while PESO studies continual adaptation of low-rank parameters [11]. Live-streaming systems refresh representations and integrate generation into time-sensitive serving pipelines [2,3,12]. Understanding their relationship requires asking which state changes, what information is available when it changes, and which dependent states remain compatible.
Existing surveys organize item identifiers by construction, alignment, and generation [13], generative recommendation by data, models, and tasks [14], and search and recommendation through shared stages [15]. Continual recommendation and generative retrieval examine adaptation, retention, and incremental indexing [16,17]. SIDInspector and SIDScope directly establish mapping diagnostics, path resolution, and generator handoff checks [18,19]. The remaining selection problem spans these interfaces: after locating a failure, should a service change candidate access, adapt a model with fixed codes, migrate its representation, or update request filtering?
The answer depends on both the failing interface and the time available to act. A repaired mapping can still require learning; a ranking improvement can leave candidate coverage unchanged; and a more effective update after training can arrive too late for a short-lived item. These cases lead to different controls and different choices. Our synthesis connects within-study contrasts to the information each route needs, the dependent states it must update, and the portion of an item’s lifetime in which it can contribute.
The survey makes three contributions. First, it organizes update mechanisms by the failure they can address and the state changes required to compose them. Second, it uses six source-located comparison groups to derive conditional selection criteria, including cases where a preferred model, metric interpretation, or intervention changes. Third, it specifies a common event-cohort evaluation that measures quality before and after update activation and charges preparation, adaptation, and serving to the same resource budget. The method-selection table identifies candidate routes, the evidence synthesis explains the conditions behind the choice, and the evaluation specification makes that choice testable.
Section 2 describes the evidence base, and Section 3 establishes the technical distinctions. Section 4 and Section 5 develop the state framework and method comparison. Section 6 and Section 7 examine protocols and cross-study findings. Section 8 and Section 9 turn the synthesis into reporting guidance and research questions.

2. Scope and Survey Methodology

2.1. Object of the Survey

We examine recommenders that use generated item identifiers or generated semantic representations to select existing catalog items, including hybrid pipelines in which a generative component supports retrieval or ranking. The primary questions concern item arrival, changes to existing item content or collaborative relations, and adaptation across successive time periods. Static tokenization and inductive representation methods serve as conceptual anchors. Generating entirely new media is a different task because it changes the object being recommended as well as the means of selection.
We distinguish three groups of research evidence, alongside neighboring surveys and tutorials. Direct studies examine catalog change, previously unseen items, or access and resolution mechanisms that determine usable catalog outputs. Context studies establish the relevant representation or adaptation mechanism without demonstrating the complete dynamic setting. Adjacent retrieval studies examine changing document collections and inform the comparison of update strategies. Their retrieval outcomes are not treated as recommendation outcomes. This separation matters because predicting an explicit query’s relevant document and predicting a user’s next interaction impose different conditioning and feedback requirements [1,17].

2.2. Retrieval and Evidence Extraction

This focused literature synthesis uses public literature located through 16 September 2026. Discovery combined public web queries for generative recommendation with cold-start, inductive, incremental, continual, dynamic catalog, live-streaming, semantic identifier, and survey terms. References in directly relevant papers were followed to connect the new-item, tokenizer, and continual-learning branches. Primary records were checked through arXiv, publisher or conference pages, and author-maintained repositories. Versioned primary-source URLs identify the preprints used for the synthesis. The working collection contains 70 references: 17 direct studies, 34 conceptual or evaluation anchors, 13 adjacent studies, and six neighboring syntheses or tutorials. These roles identify how each source is used rather than assigning a quality rank.
Inclusion requires a traceable primary source and a contribution to one of the survey’s comparison questions. A paper is not treated as evidence for dynamic adaptation merely because it describes a sequential model. For each inspected study, we extract the changing object, updated state, fixed dependencies, information used for adaptation, evaluation setting, and reported outcome. A field absent from an inspected location is recorded as unresolved rather than interpreted as an absent capability. Publication versions and preprints are linked where possible and are not counted as separate demonstrations.
The extraction covered 22 targeted protocol records. Across the collection, 30 sources were inspected at selected full-text locations, four neighboring syntheses at selected sections, and four foundational papers through supplied local texts. Other sources provide background from primary abstracts or proceedings summaries. Extraction was assisted by an AI system and has not undergone independent human coding.
The empirical basis of the synthesis is presented in the article. Section 7 develops six comparison groups, naming the interventions, observed directions, source tables or sections, and limits of each inference. Table 6 indexes all 22 records and 17 direct studies, and Table 7 connects the observed contrasts to decisions. The groups are selected because a controlled change distinguishes candidate routes or changes the interpretation of an outcome. The records cover 16 direct studies and six contextual sources; the remaining direct study, Reformer, is represented by its formal publication abstract. A study may inform more than one decision; these uses do not constitute independent replications. The collection supports conditional mechanism comparisons rather than estimates of field-wide prevalence.
Table 1. Relationship to neighboring syntheses. The final column identifies the decision developed by the present synthesis.
Table 1. Relationship to neighboring syntheses. The final column identifies the decision developed by the present synthesis.
Reference Organizing object Comparison in this survey
Multimodal recommendation [20] Feature extraction, encoders, fusion, and losses When content representations or inferred supervision can be used during catalog updates
Item identifiers [13] Construction, alignment, and generation Dependencies when assignments or the catalog change after training
Generative recommendation [14] Data, model, and task opportunities The update state and temporal contract of each adaptation route
Continual RS tutorial [16] Replay, regularization, and practical continual settings Coupling of parameter adaptation with generated item representations
Generative IR [17] Document retrieval and response generation, including incremental indexing Recommendation-specific behavior drift and current item availability
SID diagnostics [18,19] Mapping artifacts, path resolution, and model handoff Which update route to choose after diagnosis, given available feedback and the time remaining to serve the item
GRID [6] Modular semantic-ID implementation and comparisons Compatibility and update cost across successive deployed states

3. Technical Foundations and Boundaries

3.1. Sequence Modeling, Generative Retrieval, and Catalog Change

Sequential recommendation models the dependence of the next interaction on earlier behavior. SASRec uses self-attention over item histories, while BERT4Rec learns from masked items with bidirectional context [21,22]. Their sequence representations can capture changing preferences within a history. A dynamic-catalog problem additionally asks how the deployed system represents and accesses an item that was absent when its parameters or index were last constructed. Sequence order, parameter updates, and catalog admission are therefore separate design choices.
The term generative recommendation covers several output interfaces. P5 expresses recommendation tasks in a common text-to-text format [23]. TIGER generates a discrete identifier for an existing catalog item [4]. HSTU formulates recommendation as sequential transduction of behavioral events [24]. OneRec develops an industrial end-to-end generative system [25]. These systems share sequence modeling ideas but need not share the same identifier space, candidate mechanism, or update procedure. We center the survey on item access under changing catalogs and include broader sequential generators when they provide evidence about adaptation resources.
Generative item retrieval also inherits ideas from document retrieval. DSI encodes an index in a sequence-to-sequence model, and NCI combines structured identifiers with a decoder adapted to their prefixes [26,27]. The shared abstraction is to predict an address rather than independently score every object. In recommendation, the conditioning signal is usually a user history, and the target distribution reflects exposure and behavior. Consequently, preserving a document’s query associations and preserving an item’s usefulness to a user are related maintenance problems with different supervision.

3.2. What a Semantic Identifier Encodes

Vector quantization replaces a continuous representation with a discrete code [28]. Residual quantization refines this representation through successive codebooks [29]. In semantic-ID recommendation, the input can contain item text, multimodal content, collaborative information, or their combination. The resulting code sequence is then used as a prediction target and often as the representation of past interactions. TIGER established a widely used content-to-code-to-generation pipeline [4]. LC-Rec, LETTER, and CoST introduce different ways of incorporating recommendation information into the representation or its alignment with the generator [5,30,31].
Four objects must be distinguished. The codebook contains representative vectors. The token vocabulary contains the discrete symbols predicted by the model. The assignment maps each item to a sequence of those symbols. The resolver maps a generated sequence back to one or more concrete items. Codebook vectors can move without adding tokens. An item’s assignment can change even when the vocabulary size stays fixed. Two items can also share the same complete code unless uniqueness is enforced. These differences determine what has to be updated when the catalog changes.
A textual identifier offers another representation. IDGenRec learns concise identifiers composed of natural-language tokens and studies transfer to held-out domains [32]. Such identifiers connect items to the language model’s existing vocabulary. They still require item identification, duplicate handling, and a current mapping from generated text to the catalog. A natural-language tokenization scheme changes the available prior knowledge; it does not eliminate the catalog-management problem.
Table 2. Representation choices and their maintenance consequences. Consequences are derived from the representation, while references provide the mechanism anchors.
Table 2. Representation choices and their maintenance consequences. Consequences are derived from the representation, while references provide the mechanism anchors.
Representation Mechanism anchors Consequence under catalog change
Atomic item ID SASRec; HSTU [21,24] A new item needs an address and usable model representation; it cannot inherit meaning from the symbol alone.
Content-based dense vector UniSRec; Recformer [33,34] An encoder can represent an unseen item; the vector must enter the current retrieval structure.
Quantized semantic ID TIGER; LETTER; LC-Rec [4,5,30] Shared codes enable composition, but assignment changes affect both model inputs and prediction targets.
Textual ID IDGenRec [32] Language priors aid transfer; uniqueness, resolution, and temporal evidence remain separate requirements.
Sparse–dense pair LIGER; COBRA [8,35] Candidate generation and dense refinement divide access responsibility; the cold-item insertion rule matters.
Static plus dynamic SID SSRLive [12] Persistent entity information and current content can change on different schedules.

3.3. Coldness Has Multiple Dimensions

We use unseen item for an identity with no training interactions at the declared cutoff, and low-frequency item for an identity with few training occurrences. Domain shift and content change can coexist with either condition.
SpecGR studies unseen items under timestamp-based splits [7]; centroid initialization defines its reported cold cohort by at most one training-target occurrence [36]. IDGenRec transfers to unseen datasets [32], while a live room can retain its streamer identity as its content changes [12]. When discussing a source’s warm/cold results, we retain its cohort definition. These distinctions determine which representations and feedback an update can reuse.
The distinction also applies inside an identifier. A new item can consist entirely of familiar tokens but combine them in a rarely supported path. Alternatively, one of its symbols can lack training support. The temporal reachability study separates token support from item novelty and uses diagnostic variants to investigate where generation fails [37]. A model’s ability to assign a valid code is thus weaker evidence than its ability to place that item in a useful recommendation list.

3.4. Why Continual Learning Is Necessary but Not Sufficient

Streaming recommendation already studies historical retention, recent adaptation, and bounded training resources. GAG uses a reservoir of sessions, and DEGC expands graph-model capacity to separate evolving preferences [38,39]. Recent LLM-enhanced methods such as TRACER and SCoRD address semantic guidance and the coordination of retrievers and rerankers [40,41]. These works establish that replay, expansion, and coordinated updates are not unique to semantic-ID systems.
Generated identifiers introduce an additional dependency: the symbols used to express old behavior may themselves change. A replay buffer containing integer items can be re-encoded with a new mapping; a buffer containing token sequences requires its mapping version or an explicit migration rule. A parameter-preservation objective protects a different object from an assignment-preservation objective. The next sections organize methods by which state they modify and which dependencies they retain.

4. Catalog Events and Coupled System States

4.1. What Changes When the Catalog Changes?

We use item admission to mean registering an item so that the deployed recommendation procedure can consider it. Admission is distinct from having observed user feedback about that item. We use identifier assignment for the mapping from a concrete item to tokens, and vocabulary for the set of tokens the model can output. Changing an assignment need not expand the vocabulary. Likewise, an unchanged assignment does not imply that the content or collaborative meaning represented by it has remained stable.
Four events help separate the adaptation problems. An arrival introduces a new item identity. Content evolution changes the attributes of an existing item. Collaborative drift changes the observed relationships among users and items. Retirement removes an item from the currently eligible set. Arrival is central to SpecGR and the cold-start studies, collaborative drift to DACT, and rapidly changing content to the live-streaming systems [1,2,7,9,12]. Retirement is included as an operational consequence of a changing catalog; the inspected evidence does not establish an equally mature generative recommendation literature for it.
A live session can combine a new session identity, an established streamer, changing content, and a short lifetime. The predicted entity determines which information transfers across these events.

4.2. A State-Based View

At request time t, let I t be the universe of registered item identities, including retained records of retired items. Let C t ⊆ I t denote the active catalog and H u < t the history available for user u. Let q t map item information to an identifier, θ t parameterize the recommendation model, and A t denote the access structures used in serving, such as a valid-prefix trie, an item resolver, or an auxiliary retrieval index. The returned list is
L ^ t ( u ) = Serve ( H u < t , C t , q t , θ t , A t ) .
This representation exposes states beyond the model parameters. For example, an arrival can update C t and A t while leaving θ t unchanged. A tokenizer update can alter q t while requiring a corresponding update to historical training sequences or the resolver.
For direct autoregressive generation, the score of an item’s identifier z = q t ( i ) depends on the successive token probabilities,
p θ t ( z ∣ H u < t ) = ∏ ℓ = 1 | z | p θ t ( z ℓ ∣ z < ℓ , H u < t ) .
The expression clarifies why the availability of the individual tokens does not establish that the complete path will be selected by a bounded decoder. A path may be valid but have little learned support or be pruned during search. Conversely, a hybrid system can score or verify an externally supplied item without requiring that the item first emerge from unrestricted generation [7,8,37].
Figure 1 places feedback learning and catalog maintenance within a shared recommendation loop. An item can enter an access structure before feedback updates the generator, while a changed assignment introduces compatibility dependencies across both branches.

4.3. Three Different Senses of Reachability

An item is representable if the tokenizer can assign it a usable code. It is structurally accessible if the deployed procedure admits the code or retrieves the concrete item through some route. It is selected if it actually reaches the returned list for a request. These distinctions avoid equating a theoretical vocabulary capacity with observed recommendation coverage. The token-support analysis in temporal cold-start work concerns the relationship between learned generation and item access [37]; our framework places that relationship beside updates to the other system states.
A metadata record can precede trie insertion or an eligibility refresh. The active serving state therefore determines which of these reachability conditions holds at a request.
Table 3. Catalog events require different information and interventions. Proposed controls in the last column follow from the state model.
Table 3. Catalog events require different information and interventions. Proposed controls in the last column follow from the state model.
Event Information at issue Example evidence Relevant control
New item Metadata before feedback SpecGR; GenRecEdit Catalog insertion and index admission time
Content evolution Current attributes of an existing entity OneLive; SSRLive; TAGR Content snapshot and assignment version
Collaborative drift Recent interactions and co-occurrence DACT; PESO Feedback cutoff and jointly updated states
Retirement Current availability Derived operating requirement Removal propagation and stale outputs

4.4. Resolution, Eligibility, and Compatible Updates

We distinguish the active catalog C t from the request-eligible subset
E t ( u ) = { i ∈ C t : e t ( u , i ) = 1 } ,
where e t expresses request-specific rules within the active catalog. Let R t ( z ) ⊆ I t be the identities returned by the deployed resolver for code z, before active-catalog and request filtering. An empty set denotes an unresolved path, and multiple identities denote ambiguous resolution. Stale resolution is also expressible: R t ( z ) ∖ C t contains inactive identities still returned by the resolver. A path resolving uniquely to a retired item is therefore resolvable but currently ineligible. The set R t ( z ) ∩ E t ( u ) identifies eligible resolved items. Collision-corrected evaluation, trace diagnostics, and request-aware decoding motivate these separate checks [19,42,43].
For an active item with an unambiguous code, identity consistency requires R t ( q t ( i ) ) = { i } . A system that deliberately uses group codes can instead resolve several items and rank them explicitly. Both interfaces are meaningful, but their top-K units differ. Evaluating K code paths as though they were K uniquely identified items can change the recommendation metric [43]. We therefore treat the resolver as part of the access state and report whether the output budget counts paths, candidates, or final items.
Compatibility has two levels. Identity compatibility preserves the connection among an interaction, its encoded representation, and its resolved item. Behavioral compatibility concerns whether a reused model still gives useful predictions under the current mapping. A bijective relabeling can preserve identity perfectly and still invalidate learned token associations. Conversely, an unchanged mapping can remain interpretable while user preferences drift. DACT and SIDScope already investigate this separation through adaptation and model-handoff experiments [9,19]. Here, compatibility determines which components an update route must carry with it; Section 7 uses controlled contrasts to decide when that extra work is justified.
The minimal deployed version is thus a tuple of mutually interpretable states, rather than a model checkpoint alone. A change can involve new metadata, recoded histories, an updated resolver, new model parameters, and revised eligibility rules. An evaluation that replaces one component should say which of the others are regenerated or retained. This allows readers to distinguish a representation intervention from a complete system update.

4.5. One Item Across the Update Boundaries

Consider a constructed product lifecycle. Metadata and the catalog entry become public at 09:00, and a retrieval index admits the item at 09:02. The first feedback becomes available for learning at 09:10. At 09:20, a candidate assignment changing ( a , b , c ) to ( a , d , e ) is ready in a staging area. It is not yet used for serving. A compatible state becomes active at 09:30, and the product retires at 10:00. These times illustrate an ordering, not a measured deployment trace.
The example uses a staged migration with a single activation boundary. Between 09:20 and 09:30, requests continue to use the old assignment, histories encoded under it, resolver, and model. In staging, histories are re-encoded and the new resolver and model are prepared and checked. At 09:30, requests switch to the compatible tuple of assignment, history encoding, resolver, and model. Each request uses one tuple throughout; in-flight requests finish on the old tuple. The catalog and eligibility state reflect retirement at 10:00.
Thus the 09:05 request can use metadata-based access, but neither the later feedback nor the staged assignment. Index insertion, migration, training, and verification consume resources before their respective benefits become available. The example’s history re-encoding policy is one concrete choice; alternative migration policies must state their own preparation and activation boundaries.

5. Methods Across Update Boundaries

Read Figure 2 horizontally to identify coupled components within a route and vertically to compare interventions at the same interface. Figure 3 distinguishes three mechanism classes by where concrete item candidates enter.

5.1. Admitting New Items Through Retrieval and Verification

Content representations offer a route to unseen items because a new item can be encoded from its attributes. UniSRec and Recformer provide relevant representation-based foundations [33,34]. In a generative system, however, encoding an item and assigning its target sequence do not necessarily make that sequence likely under the trained decoder. New-item admission methods address the gap between these operations.
SpecGR places inductive drafting before generative verification [7]. A drafter proposes concrete candidate items, including unseen items, and the generative model evaluates their identifier likelihoods. Guided re-drafting steers further candidates when too few have been accepted. The verifier also distinguishes semantic tokens from an extra item-disambiguation token: in the inspected implementation, the latter is excluded from the verification score for unseen items. This adjustment matters because a token introduced to distinguish registered items need not carry transferable semantics. The method changes how candidates reach the generator and how their likelihoods are interpreted, rather than relying on the generator to discover each new identifier unassisted.
SpecGR++ additionally learns the drafting representation. Its training and inference ablations separate this preparation from the serving procedure ([7] Table 3).
LIGER connects the same two representation families through a different scoring interface [8]. In its inspected inference algorithm, beam-generated candidates are supplemented with the cold-item set and ranked using the encoder’s dense output. The design assumes that unseen items are relatively few between periodic model updates. SpecGR instead proposes candidates through an inductive drafter before verification. This difference changes the work required as the number of new items grows: adding an entire cold set and retrieving a bounded subset have different candidate-cost dependencies. Describing both as hybrid methods leaves this operating assumption unspecified.

5.2. Updating Identifier Assignments

Static tokenizer design determines which item properties the generator sees. LETTER makes item tokenization learnable and incorporates recommendation-relevant signals [5]. Once collaborative information is encoded in the identifier, new interactions can change the representation that the tokenizer ought to preserve. The item description may remain constant while its behavioral neighborhood changes.
Incremental tokenization work studies updates to both tokenizers and recommender language models [46]. DACT more specifically distinguishes drifting items from relatively stationary items and constrains their updates differently [9]. Its design reflects two competing needs: assignments should respond to changing collaborative information, while previously learned token meanings should remain useful. The relevant unit of stability is consequently the assignment of an existing item and its compatibility with the generator, not merely the existence of a fixed-size codebook.
DACT’s ablations on the Amazon Tools and Home Improvement category (Tools) isolate drift identification, differentiated updates, global stability, and reassignment [9]. The frozen and jointly updated controls are compared in Section 7.
ChronoID examines where explicit temporal information enters semantic-ID learning, including embedding, fusion, and quantization choices [47]. It extends the design space beyond a single time-invariant item representation. Time conditioning and continual codebook adaptation nevertheless remain distinct: a fixed parameterized mapping can produce different outputs at different times, while an updated mapping can change without receiving time as an explicit input. The distinction helps separate temporal expressiveness from the maintenance procedure used during deployment.

5.3. Adapting Model Parameters and Editing Unseen-Item Behavior

Parameter adaptation changes how the recommender maps its inputs to predictions. LSAT separates long- and short-term adaptation through low-rank modules, with a short-term module fitted to newly collected data and a long-term module updated more slowly [48]. The original study concerns LLM-based recommendation and provides a mechanism anchor; its existence alone does not demonstrate continual semantic-ID maintenance. PESO maintains an evolving adapter and regularizes it relative to its preceding state [11]. This makes preservation a local constraint on adaptation rather than an instruction to reconstruct every historical preference.
GenRecEdit takes a more targeted route for items unseen in training interactions [10]. It uses semantic similarity to warm items to construct surrogate histories, converts these into position-wise editing requests, and controls which edit is activated while generating each identifier position. The method thus uses transferred context to alter the model’s response before ordinary retraining would absorb the new item. Its information source differs from observed feedback about that unseen item: the surrogate context is inferred from other items. An evaluation should preserve this distinction when comparing editing with supervised continual updates.

5.3.0.1. Upstream representations and inferred supervision.

MENTOR aligns modalities with interaction-trained ID embeddings; MDVT and VI-MMRec construct virtual supervision, while SG-URInit initializes user representations from semantics [49,50,51,52]. For dynamic use, these upstream choices determine what context can transfer and what preparation must be charged. Construct features and surrogate histories at the same information cutoff, distinguish inferred supervision from observed target-item feedback, and include initialization or refresh in the update cost.

5.4. Refreshing Representations in Live-Streaming Systems

Live-streaming provides direct evidence that item identity and current content need different representations. OneLive encodes changing content and immediate behavior through a dynamic tokenizer [2]. Its content encoder uses a sliding window, illustrating that freshness can enter upstream of quantization. The model must represent what a room is currently showing while retaining useful information about the creator and audience.
SSRLive explicitly separates static and dynamic semantic IDs and combines a generative component with discriminative prediction [12]. Its inference uses generated representations in downstream multi-task scoring. Its generated representation is therefore an input to a hybrid ranking pipeline.
TAGR makes the distinction between vocabulary and assignment concrete [3]. Its deployment description reports periodic re-encoding of active live ads, propagation to a bidirectional identifier-to-item index, and separate synchronization of model parameters. Its fixed hierarchical vocabulary coexists with changing assignments and separately refreshed model parameters.
OneMall provides broader context for sharing a generative recommendation family across e-commerce scenarios [53]. The central comparison here is the responsibility of each stage, including any downstream filtering or ranking. A deployed configuration couples content encoding, user modeling, objectives, and serving. Isolating identifier maintenance therefore requires a control that holds the other components fixed.

5.5. Lessons from Dynamic Generative Retrieval

Changing document collections motivate several transferable mechanisms: generative replay in DSI++, constrained insertion in IncDSI, continual pretraining in CorpusBrain++, and incremental quantization with memory in CLEVER [54,55,56,57]. DynamicIR examines adaptation and resource costs; replication work tests identifier generalization, and MixLoRA-DSI selectively expands low-rank experts [58,59,60]. These mechanisms preserve access to old entities while incorporating new ones. Their transfer to recommendation depends on the relevance signal: user preferences, exposure, and item availability evolve alongside the corpus. Shared search–recommendation identifiers further make the representation’s training task relevant [61]. Retrieval retention consequently motivates an update design; recommendation outcomes must establish its utility.

5.6. Maintaining Codebooks Under Exposure Drift

The dynamic large-codebook study separates the depth of a semantic identifier from the schedule on which it is refreshed [45]. Its representation uses one large semantic code and a separate disambiguation code. The dynamic component maintains decayed exposure weights, updates centers with an exponential moving average, and penalizes changes relative to the previous assignment. The penalty is weighted by exposure, making a reassignment of a heavily exposed item more costly. The underlying concern is that changing many frequently used prediction targets can disrupt subsequent optimization.
DACT and this approach place stability constraints at different points. DACT adapts representations and assignments learned from collaborative signals [9]. Exposure-aware center updates organize the pressure to preserve assignments around traffic [45]. Both use previous representation state, but the evidence that triggers change differs. A comparison that fixes only the generator would therefore be incomplete: the content encoder, interaction window, exposure weights, assignment rule, and generator update must be recorded together.
The fixed-date KuaiRec experiment also clarifies the treatment being evaluated. The dynamic arm updates the codebook using a historical traffic window and fine-tunes a matched checkpoint on re-encoded targets; validation and test days follow that window [45]. Its outcome measures this combined update. Separating assignment maintenance from further optimization requires an additional control. The appropriate baseline includes a static-codebook model given the same update data and optimization budget. This control is particularly useful when deciding whether a more complex representation-maintenance policy is worthwhile.

5.7. Allocating Training Data and Preserving Semantic Priors

The data-selection study adapts HSTU on longitudinal music and podcast histories [44]. A recent reference subset guides selection through gradient or hidden-state representations, and evaluation uses the following interval. Selected incoming histories augment existing training data. Thus, the incoming fraction alone does not specify training cost: gradient extraction, replay, and optimization remain in the budget. Section 7 compares the selection controls.
Centroid initialization maps codebook centers into new token embeddings through dimension matching and a mean shift [36]. It preserves semantic information at model initialization rather than changing item assignments. Its evaluated cohort includes targets with one training occurrence, so its cold-cohort result does not isolate targets with zero training interactions. Variable-length identifiers instead allocate different code lengths [62]; using this capacity choice dynamically requires a policy for migrating histories and decoding constraints when lengths change.

5.8. Enforcing Request-Time Eligibility

GRACE incorporates user-dependent targeting into constrained SID decoding [42]. Bitmasks handle low-cardinality attributes and Bloom filters handle higher-cardinality attributes. A prefix survives if its descendants may contain an eligible ad; an exact matcher then filters the resolved ads. This final step matters because an eligible descendant does not make every item sharing a code eligible. The study evaluates ad-level pass rates using synthetic user locations; Section 7 separates the effects of its targeting and partitioning controls.

5.9. Separating Search Support from Post-Search Ranking

TAAL uses joint-prefix training to improve early path support and applies a PMI-based calibration after beam search [63]. These operations affect different sets. Training can alter which complete paths survive; post-beam calibration can reorder only those already retained. Its reported experiments use leave-one-out evaluation and a fixed seed, so the term temporal refers to sequence transitions rather than demonstrating a rolling catalog-update deployment.
Hybrid interfaces can be classified by the ordering of candidate access and generative scoring (Figure 3). Propose, then verify supplies concrete items to a generator; SpecGR instantiates this interface with inductive drafting and likelihood verification [7]. Merge candidates, then rank joins generated candidates with an external candidate source; LIGER uses the cold-item set and a dense ranker [8]. Generation-conditioned dense retrieval uses generated codes to condition a continuous retrieval query; COBRA provides a sparse-conditioned dense refinement example [35].
These classes describe information flow, and a system may combine them. Their admission conditions differ: proposal and merging need an external route that supplies new items, whereas generation-conditioned retrieval depends on both generated support and the continuous search domain. The presence of dense vectors alone therefore does not identify the route by which a new item becomes selectable.

5.10. Choosing a Route from the Available Evidence

Table 4 organizes route choice around an observed bottleneck. First identify whether the target is missing from candidates, incorrectly resolved, rejected by current rules, or poorly ranked. Then restrict the alternatives to routes whose inputs are already available and whose dependent updates can complete during the service window. This yields a comparison between feasible interventions at the same interface.
Table 4. Conditional route selection. Source observations motivate each route; the final column states what would justify choosing it for the current workload.
Table 4. Conditional route selection. Source observations motivate each route; the final column states what would justify choosing it for the current workload.
Route and available input Source observation or scope State dependency and cost Condition for selection
Inductive drafting; item metadata (SpecGR) [7] Removing drafting or unseen-score adjustment weakens induction ([7] Table 3) Current retrieval index and resolver; drafting and verification work Use when bounded proposals reach missing targets within the refresh and verification budget
Cold-set union; item metadata (LIGER) [8] Algorithm unions generated items with the cold set; source assumes few arrivals between updates Aligned item IDs and dense features; scoring grows with cold set Retain while scoring the accumulated cold set meets the request budget
Targeted editing; metadata and warm-item context (GenRecEdit) [10] Surrogate histories transfer warm-item context; cold targets and edit controls No target-item interactions required for surrogate construction; edits use current codes Require useful transferred context and acceptable warm-item retention and preparation delay
Parameter adaptation; recent feedback (PESO; LSAT anchor) [11,48] PESO: blockwise updates; LSAT: parameter-adaptation mechanism Fixed history/target mapping; selection, replay and training cost Prefer if the fixed-code update meets the quality target and joint migration adds no useful gain
Selective recoding; collaborative evidence (DACT) [9] Tokenizer-only update harms ranking; selective updates are not best in every period ([9] Table 2) Stage recoded histories, resolver and model; encoding/migration cost Select only with quality benefit over fixed-code adaptation and a compatible, timely handoff
Dynamic representation; fresh content/behavior [2,3,12] Live-stream offline/online treatments Entity continuity and synchronized assignments; recurring refresh Require fresh-content benefit before expiry, including propagation delay
Eligibility decoding; current rules (GRACE) [42] Decode-time targeting improves final ad pass; finer partitioning adds little ([42] Table 3) Consistent prefix and exact-item filters; index and serving cost Use when eligibility limits yield; require final quality and latency within budget
Two choices require different evidence. When new-item candidate coverage is low, compare bounded drafting with cold-set union at the observed arrival load; their candidate-scoring costs scale differently. When coverage and resolution are already adequate, start with a frozen-code model update. Escalate to joint recoding only if it adds useful quality under matched data and a budget that includes migration. For metadata-based editing, semantically matched and shuffled surrogate histories test whether the transferred context justifies the edit. These are conditional choices: Section 7 provides the source contrasts, and Section 8 defines their common evaluation window.

6. Datasets and the Questions Their Protocols Can Answer

6.1. A Dataset Name Does Not Identify an Experiment

Amazon datasets connect many of the reviewed methods, but their versions and uses differ. LIGER uses categories from Amazon 2014 with per-user leave-one-out evaluation [8]. SpecGR uses Amazon Reviews 2023 with timestamp cutoffs [7]. DACT and PESO study successive interaction blocks with evaluation inside those blocks [9,11]. The shared product domain does not make their scores directly comparable. Catalog size, item metadata, the definition of an unseen target, and the information available for tokenization can all change.
Amazon Reviews 2023 provides a large content-rich resource and is associated with semantic-encoder benchmarking [64]. Its review times support chronological interaction partitions. A first review is still an observation of behavior, rather than proof of a product’s original publication time. Similarly, a description present in a current metadata file does not establish the exact description available at an earlier request. Studies that infer arrival from first observation should name that proxy and restrict their conclusions to it.
MicroLens provides video content and interaction histories, making it useful for content-based item representations [65]. MIND provides news text and impression records, supporting comparisons conditioned on displayed candidate sets [66]. KuaiRec offers a densely observed subset that helps examine the effect of missing feedback and exposure [67]. These resources resolve different data problems. None should be assumed to contain a complete sequence of publication, metadata revision, serving admission, and retirement events without checking the particular release.
Table 5. Data resources and their use in dynamic-catalog research. The final column identifies information needed for the proposed lifecycle claim, rather than asserting that every release lacks it.
Table 5. Data resources and their use in dynamic-catalog research. The final column identifies information needed for the proposed lifecycle claim, rather than asserting that every release lacks it.
Resource Useful observed structure Additional information to verify
Amazon reviews [7,8,64] Product metadata and timestamped interactions; multiple historical releases Publication time versus first review; historical content snapshots; filtering before or after time split
MicroLens [65] Micro-video interactions with raw content modalities Content version at prediction; upload and retirement events
MIND [66] News text, user histories, displayed candidates and click labels Full eligible catalog outside an impression; article freshness and expiry rule
KuaiRec [67] Dense feedback subset alongside larger interaction data Which subset and exposure process are used; actual availability events
Industrial live streams [2,3,12] Rapid content changes and online engagement outcomes Reproducible event interface; state synchronization and precise experimental treatment
Music/podcast stream [44] Longitudinal item/action histories and rolling adaptation Selection and replay cost; item-admission rule beyond the modeling window

6.2. Four Protocol Families

Held-out identity or domain.

A withheld-item experiment asks whether a model can act on identities excluded from training. A held-out-domain experiment additionally changes the domain’s content and interaction distribution. LIGER excludes its cold items from tokenizer training, whereas IDGenRec’s zero-shot study transfers across datasets [8,32]. These controls test generalization to withheld objects. They need a global information cutoff before they can establish what would have been available in a historical deployment.

Global-time future evaluation.

SpecGR and the cold-start study use global-time splits; the latter trains on the first 90% of chronologically ordered interactions and defines cold items by first appearance after that cutoff [1,7]. The temporal reachability analysis additionally examines the training support of tokens and paths associated with future items [37]. This family directly tests future-item access under the specified information boundary. Its validity also depends on when item features, tokenizers, and eligibility structures were constructed.

Blockwise continual adaptation.

DACT and PESO organize updates into chronological blocks, with per-user holdouts inside each block [9,11]. Such experiments reveal how a system adapts across successive data distributions. Within a block, however, another user’s later interaction can contribute to the model evaluated on an earlier target. That design answers an after-block learning question unless the implementation enforces finer request-time causality. The global-timeline leakage literature explains why per-user ordering alone is insufficient [68]. A survey should report the actual unit of causality rather than assigning every timestamped dataset the same temporal label.

Rolling prediction and online serving.

The data-selection study updates from an interval and evaluates in the next interval [44]. RecNextEval explicitly separates data release, prediction, and result release [69]. Industrial A/B tests evaluate complete deployed treatments, potentially including representation, ranking, objectives, and serving changes [2,3,12]. These designs provide complementary evidence. Rolling tests measure forward generalization under a logged update policy, while online tests capture responses to actual exposure. Neither automatically isolates the contribution of the tokenizer.

6.3. Filtering and Candidate Sets Are Part of the Treatment

A k-core filter removes users or items with too few interactions. Applying it over the complete observation horizon can use future behavior to decide which early entities exist in the benchmark. This does not erase the usefulness of a controlled benchmark, but it changes the target population from all arrivals to entities that eventually meet an activity condition. For admission research, the difference is central: short-lived or rarely exposed items may be precisely those needing support. Report when the filter is applied and how many test targets fall outside its retained catalog.
A second boundary is the candidate universe. Full-catalog item ranking, generation of a fixed number of code paths, scoring an externally supplied candidate list, and reranking displayed impressions impose different opportunities for success. Sampled recommendation metrics can change model comparisons [70]. Under semantic-ID collisions, even a fixed path budget can correspond to a variable number of items [43]. Both the candidate-construction rule and the reverse mapping must therefore accompany Hit@K or NDCG@K.
Table 6 links every extracted protocol record and core study to its inspected source. The protocol families above explain the different information boundaries. Each empirical contrast there identifies its source table or section, so its intervention and outcome can be checked against the cited version.
Table 6. Source index for the 22 extracted protocol records (P01–P22) and 17 core studies (*). The inspected text is linked separately from the formal publication in the bibliography. Reformer contributes an abstract-level mechanism description and has no extracted protocol record. Section and table locators refer to the inspected text.
Table 6. Source index for the 22 extracted protocol records (P01–P22) and 17 core studies (*). The inspected text is linked separately from the formal publication in the bibliography. Reformer contributes an abstract-level mechanism description and has no extracted protocol record. Section and table locators refer to the inspected text.
ID Study Inspected arXiv text Comparison focus Location
P01 SpecGR* [7] 2410.02939v2 Future-item drafting Setup; App. A2, A7
P02 LIGER* [8] 2411.18814v2 Withheld-item union Alg. 1; §4.1
P03 Cold-start study* [1] 2603.29845v2 Seen/unseen cohorts §4.1; Table 4
P04 Reachability* [37] 2607.21101v1 Token/path support §§3, 4.3, 6
P05 DACT* [9] 2603.29705v1 Tokenizer/model updates §§2.3, 4.1; Table 2
P06 PESO* [11] 2510.25093v3 Blockwise adapters §5.1; App. C.1
P07 LSAT [48] 2312.15599v2 Two adaptation scales §5
P08 GenRecEdit* [10] 2603.14259v1 Surrogate-context editing §§4.1–4.3, 5.1; Table 2
P09 OneLive* [2] 2602.08612v1 Live-content updates §3.2; App. B.1
P10 SSRLive* [12] 2606.06970v1 Static/dynamic IDs §§4.4, 6.1; App. C.2
P11 TAGR* [3] 2608.24034v1 Live-ad synchronization §§3.2–3.3, 4.1
P12 ChronoID* [47] 2606.14260v1 Temporal ID learning §§4, 5.1; App. B
P13 Data selection* [44] 2604.07739v1 Rolling data selection §§2–5; Table 1
P14 Dynamic codebook* [45] 2608.21012v1 Exposure-weighted updates §§3.4, 4.3.2, 4.5
P15 GRACE* [42] 2608.00938v2 Eligibility controls §§6.1–6.2; Table 3
P16 SIDScope* [19] 2608.18779v1 Mapping/model handoff §§5.3, 6.3; Tables 6–7
P17 Faithful evaluation* [43] 2605.25330v2 Item-resolution credit §§3.1, 4.1–4.3
P18 Centroid initialization [36] 2608.07816v1 Low-frequency targets §§4.1, 5.2
P19 TAAL [63] 2608.29179v1 Survival versus ranking §§4.1, 4.4; App. Table 6
P20 IDGenRec [32] 2403.19021v2 Cross-domain transfer §3.5
P21 RecNextEval [69] 2604.13665v1 Release/prediction clocks §2
P22 CoST [31] 2404.14774v2 Contrastive quantization §3; abstract
– Reformer* [46] CIKM 2025 Incremental tokenization Publisher abstract

7. Evidence Synthesis for Choosing an Update Strategy

Six comparison groups connect the mechanisms to decisions. Each changes a defined part of an experiment: the target cohort, the mapping/model state, the credited output, the filtering policy, or the selected update data. The first three findings describe observed interface distinctions. The final subsection derives a lifetime quality–cost principle from the operations those distinctions require. Table 7 summarizes the resulting choices.
Table 7. Observed contrasts and their implications for route choice. Source table and section locators identify the comparisons directly. Conditions in the final column are the survey’s synthesis of those observations.
Table 7. Observed contrasts and their implications for route choice. Source table and section locators identify the comparisons directly. Conditions in the final column are the survey’s synthesis of those observations.
Choice at issue Within-study comparison Observation and boundary Consequence for selection
Which arrival baseline? Warm versus cold targets at one cutoff ([1] Table 4) The TIGER/textual-ID ordering reverses on Toys; model differences remain coupled. Choose on the intended arrival cohort; test metadata access before feedback adaptation.
Repair or learn? Mapping-only repair versus repair with adaptation ([19] Table 7) Repair restores mappings but not new-item retrieval. Common-item harm remains uncertain in this case. Retain the repaired mapping; test learning separately on the same targets.
Adapt or recode? Frozen, tokenizer-only, model-only and joint updates ([9] Table 2) Tokenizer-only updating hurts Beauty; selective updating is not best in every period. Require an incremental benefit over fixed-code adaptation before paying migration cost.
Recover or rerank? Credit rule on fixed outputs; calibration on fixed beams ([43] Table 4); ([63] Table 6) Resolution changes model ordering; calibration changes rank without survival. Tie rules and static protocols bound the claims. Fix item credit first; improve candidate support only when the target is absent.
Which eligibility filter? Catalog-only, targeted and partitioned decoding ([42] Table 3) Targeting improves pass rate; finer partitioning adds little. Synthetic locations constrain interpretation. Select on eligible yield and final relevance, including matcher cost.
Which adaptation budget? Random versus guided incoming-data selection ([44] Table 1, Appendix C) Guided selection recovers more quality but adds preparation; earlier training data remain. Compare complete cost and activation delay, not selected-data fraction alone.

7.1. Separate Candidate Access from Feedback Learning

Target population can reverse the preferred baseline.

In the cold-start reproducibility study, TIGER has higher Recall@10 than textual IDs on warm Amazon-Toys targets, whereas the ordering reverses on cold targets under the same global-time split ([1] Table 4). The comparison changes the evaluated cohort, not the training cutoff or full candidate catalog. It supports separate route selection for targets with and without training feedback. Because the compared models differ in representation and training, this reversal does not isolate tokenization as its sole cause. A warm-dominated aggregate can nonetheless favor the wrong baseline for an arrival-focused service.

Repair can restore addressability without restoring prediction.

SIDScope compares an inherited generator before and after a mapping repair on DACT’s Tools category, then tests adapted models on the same target population ([19] Section 6.3, Table 7). Repair closes missing-item mappings, but the inherited model still fails to retrieve the new-item targets; adaptation recovers new-item retrieval. The mapping-only comparison leaves it uncertain whether the handoff harms ranking quality for common items. These are different conclusions from one handoff case: coverage repair can be necessary, learning can still be needed, and harm to existing items must be assessed separately.
Together, these contrasts favor a staged comparison. Before feedback is available, test metadata admission routes such as inductive drafting or candidate union [7,8]. Once the target is accessible and feedback arrives, compare the same access route with and without adaptation. A gain in this second comparison can support learning; a gain obtained by changing the candidate route requires access to remain part of the explanation. The two interventions become complementary when both coverage and scoring limit new-item quality.

7.2. Coordinate Recoding with Histories, the Model, and the Resolver

The component updated determines the conclusion.

DACT compares frozen tokenizer/generator states, tokenizer-only fine-tuning, model-only adaptation, and joint updates on the same periodwise tasks ([9] Table 2). On Beauty, updating only the tokenizer reduces ranking quality relative to keeping both components frozen. Joint updates recover useful quality, but the selective method is not uniformly best: tokenizer fine-tuning with generator retraining has higher Hit@10 in Beauty’s third update period, while DACT retains the NDCG advantage. The preferred update therefore depends on the outcome being optimized as well as the period, within this blockwise protocol.
Read alongside the uncertainty about common-item harm in the SIDScope case, DACT establishes a conditional dependence: recoding can disrupt an inherited generator, while the need and benefit of adaptation depend on the changed interface and target population. Assignment churn alone therefore cannot select the update policy. Frozen-code model adaptation is the informative first control because it learns from new feedback without incurring a representation migration. A joint update is justified when it adds useful quality after including history recoding, resolver synchronization, and model preparation. This criterion connects a handoff diagnostic to a choice between parameter adaptation and representation maintenance.

7.3. Diagnose Candidate Access, Resolution, and Eligibility Separately

A better reported score can arise at different stages.

The faithful-tokenizer study re-evaluates the same native outputs under SID-level and collision-corrected item-level credit. On Scientific, Cell, and Beauty, a tokenizer leading under SID-level hit rate falls behind under item-level credit ([43] Section 4.3, Table 4). The change is in evaluation, not in the predictions. Its correction uses a uniform within-group tie convention, so a deployed resolver with a different tie policy needs its own item accounting. TAAL supplies a complementary contrast: adding post-beam calibration changes the ranking of surviving targets while leaving their survival unchanged ([63] Appendix, Table 6). Its joint-prefix training changes survival, under leave-one-out evaluation with a fixed seed. Resolution correction and post-search ranking thus answer different questions from candidate recovery.

Eligibility improves without establishing better relevance.

GRACE compares catalog-constrained decoding with the same decoding augmented by request targeting, followed by exact ad-level filtering ([42] Section 6.2, Table 3). Targeting improves the final pass rate. Finer matcher partitioning makes local matchers more selective but provides little additional final-pass benefit, while adding storage and checking work. This supports the simpler targeting configuration for that filtering objective. Synthetic user locations and treatment-dependent generated-volume buckets limit the comparison: it measures eligibility, and the buckets are not fixed populations for attributing gains to request subgroups.
The synthesis changes the intervention selected after a failure. Missing candidates call for admission or search-support changes. Shared or stale resolutions call for item-level resolution controls. Low eligible yield calls for current rule enforcement. Poor ordering among eligible candidates calls for scoring. A larger final metric supports the corresponding claim only when the upstream candidate population and credit rule are held fixed. This also explains why combining more selective filters with a better ranker may add cost without improving the useful output.

7.4. Compare Complete Cost Within the Item’s Useful Window

Data efficiency can shift work into preparation.

The continual data-selection study compares random incoming histories with representation- and gradient-guided selection under matched adaptation settings ([44] Table 1). Guided selection recovers more of the drift-related quality gap to full retraining at both reported horizons. The selected histories augment the existing training data; reducing incoming examples therefore does not proportionally reduce the entire training workload. The source’s efficiency analysis also charges gradient selection for backward computation that representation selection avoids ([44] Appendix C). The quality ordering at a fixed selected-data budget is consequently insufficient to select a route under a fixed total-resource budget.
The same accounting issue occurs at other interfaces: drafting incurs request-time verification, editing prepares transferable contexts, and recoding propagates changed assignments [3,7,10]. These costs have different timing. Preparation delays the first request that benefits, whereas repeated retrieval or filtering consumes resources during serving. A route that adds post-update quality is useful only for the requests it reaches before the item retires, and its benefit must justify the complete workload.
This yields an evaluation principle rather than a measured ranking of routes: compare all procedures on the same eligible event cohort, include requests before activation, and charge preparation, updating, and serving over one observation horizon. Section 8 defines the accounting. It makes a proposed combination testable by asking whether its additional component improves quality within the common window after its added costs and delay.

8. Evaluating Quality and Cost During the Available Lifetime

A comparison should ask how much useful recommendation quality each procedure delivers while the same items are available, under a common resource budget. This section specifies the event population, clocks, outcomes, and accounting needed to answer that question. Figure 4 separates information retention, serving eligibility, and activation of a compatible state.

8.1. A Matched Two-Route Comparison

Consider a constructed trace over 09:00–10:00, extending Figure 4. Item A is published at 09:00 and retires at 10:00; item B arrives at 09:20 and remains eligible when observation ends. Both have no interactions in the initial training set. The external eligibility snapshot includes each item from publication until retirement, independently of a method’s admission time. Both routes use the same initial checkpoint, request histories, candidate budget, and exact item resolver.
Route R admits each item through metadata retrieval two minutes after publication and keeps its generator fixed. Route J uses the same admission route, then jointly updates assignments and the model. Its mapping is prepared at 09:20; recoded histories, the resolver, and the compatible model activate together at 09:30. Until then, J serves the old state. Both routes can use the same separate adaptation stream, with feedback ingested by 09:10; scored requests are held out from fitting and selection. Metadata through 09:20 can enter the staged mapping. Thus, the comparison asks whether adaptation adds value after access is held fixed.
Table 8 assigns binary Hit@1 outcomes to illustrate the calculation. These are constructed outcomes, not experimental measurements. The request cohort and a 10:15 label freeze are fixed before scoring. All six labeled targets are externally eligible at request time. The 09:01 target remains in the denominator despite being unadmitted; it is an access failure for both routes. The later missed targets are admitted and correctly resolved, separating ranking from access.

8.2. Age, Aggregation, and Incomplete Observations

The primary summary pools utility over labeled requests in the fixed window: R obtains 3 / 6 and J obtains 4 / 6 , a paired difference of 1 / 6 . With prespecified age bins [ 0 , 10 ) , [ 10 , 30 ) , and [ 30 , 60 ) minutes, the respective counts are three, one, and two. Both routes obtain 2 / 3 in the first bin and zero in the second; R obtains 1 / 2 and J obtains 2 / 2 in the third. Scoring only after activation would report 1 / 2 versus 2 / 2 while discarding four labeled requests that belong to the service window.
Equal item weighting answers a different question. A contributes 1 / 4 for R and 2 / 4 for J; B contributes 2 / 2 for each. Their item-weighted means are 5 / 8 and 3 / 4 . Report this secondary summary with the per-item counts, rather than choosing the weighting after seeing which favors a route. At larger scale, the same calculation applies to Hit@K or NDCG@K, with declared candidate and final-item budgets. A path match receives item-level credit only through the specified resolver [43].
The 09:50 request is scored using its label received at 10:10, but that label cannot affect an earlier prediction or update. The unlabeled 09:40 request is excluded from quality for both routes and reported separately, yielding label coverage of six out of seven requests. It is not a negative example. These scores estimate quality among labeled requests; extending them to all requests requires an observation model or additional judgments. B’s lifetime is right-censored at age 40 minutes. Its observed requests remain included, with no imputed utility after 10:00. A’s retirement is observed. If arrival is proxied by first interaction, the resulting age must instead be named interaction age.

8.3. Complete Cost and the Route Decision

For a common horizon T, account for preparation, updates, and every served request:
C ( T ) = C init + ∑ j ∈ U ( T ) C update , j + ∑ r ∈ R ( T ) C serve , r .
Use a declared common charging unit, such as monetary cost under fixed resource rates. Keep GPU time, CPU time, memory, and end-to-end latency separately observable. Concurrent component durations determine activation through their actual completion path, rather than a sum of runtimes.
For the trace, let P charge the shared checkpoint preparation and let E charge content encoding and the two initial admissions. Let U charge J’s data selection, retained-data processing, tokenizer/model training, history recoding, resolver/index replacement, validation, and state transfer. Include storage and overlapping old/new-state residency in the corresponding terms. If s r R and s r J are complete request costs, including drafting, verification, resolution, and final filtering, the two totals are
C R = P + E + ∑ r = 1 7 s r R ,
C J = P + E + U + ∑ r = 1 4 s r R + ∑ r = 5 7 s r J .
Thus Δ C = U + ∑ r = 5 7 ( s r J − s r R ) is the additional cost for the constructed gain of one hit. The unlabeled request is included in both totals. Shared preparation is charged once to each alternative service, with no repeated per-item charge. Any reusable artifact is amortized over its actual reuse horizon; work discarded or completed after retirement remains a cost even when it earns no utility in this window.
Fix a total budget B and request-latency bound L before comparison. If both routes satisfy them, J has the higher request-weighted utility in this trace. If C R ≤ B < C J , only R is affordable. Under a prespecified trade-off Q − λ C , J is preferred exactly when λ Δ C < 1 / 6 , subject to the latency bound. This boundary connects the utility gained during availability to the measured incremental cost. To attribute an observed advantage specifically to recoding, add a third route with the same feedback and activation time but a frozen tokenizer; R versus J identifies the value of the combined update.

8.4. Match Uncertainty to the Comparison Unit

Seeds measure optimization variability, user-level resampling measures variation across users, and repeated update windows measure temporal heterogeneity. SIDScope uses paired, stratified user bootstrap for a mapping handoff, while the dynamic-codebook study reports repeated seeds [19,45]. These address different uncertainty sources.
For a shared event stream, use paired differences and preserve user or session clusters where they contribute repeated targets. Report update-window, unseen-target, and retired-item counts separately from prediction counts. A large interaction sample does not establish that an update policy generalizes across many independent catalog changes.
Table 9. Minimum fields for an interpretable dynamic-catalog comparison. These are proposed reporting requirements derived from the synthesis, not fields claimed to be reported by every reviewed paper.
Table 9. Minimum fields for an interpretable dynamic-catalog comparison. These are proposed reporting requirements derived from the synthesis, not fields claimed to be reported by every reviewed paper.
Field group Required distinction Comparison enabled
Entity and event Item, creator, session; arrival, content change, drift Same evolving object
Information boundary Event time, availability time, feedback cutoff No use of unavailable evidence
Catalog and candidates Eligible snapshot; admission/removal; retrieval route Comparable opportunity to select the target
Tokenizer state Fit data, vocabulary, assignment version, history encoding Stable meaning of model inputs and outputs
Model state Checkpoint, training data, update trigger, active time Matched adaptation information
Outcome population Training-seen/unseen, item age, drift stratum Benefit in the intended operating regime
Resource boundary Encoding, editing/training, index work, serving Quality under a stated total budget

9. Open Research Questions

9.1. When Should a System Switch or Combine Update Routes?

The unresolved choice is when metadata access should be supplemented by editing or feedback learning. Feedback volume, representation uncertainty, retrieval coverage, and expected remaining lifetime are candidate switching signals. A discriminating study would compare individual routes with staged combinations on the same event stream and resource budget, measuring candidate coverage and item-age quality before and after each transition. This would reveal whether a slower update adds value within the available window or repeats work already performed by another route.
Combination tests should inspect the interface carrying the benefit. Recoding plus replay requires recording which replay inputs change representation. Prefix filtering plus ranking requires measuring eligible candidate yield at a fixed final-item budget. Data selection plus a small adapter requires charging selection and retained-data processing to the update. The unresolved issue is whether these combinations add useful quality after accounting for interactions among their preparation, migration, and serving costs.

9.2. How Should Identifier Versions Be Migrated?

A migration policy can retain old assignments, re-encode histories, or preserve a stable identity channel alongside dynamic state, as motivated by SSRLive [12]. What remains unresolved is how to choose among these policies under recurring change while maintaining useful model behavior.
Synthetic code permutations can isolate relabeling from semantic change. Actual catalog streams are needed to test meaningful representation drift. Compare identity preservation, model quality, migration delay, and coherent rollback across the same state versions. This would distinguish a compatible migration from a merely low-churn assignment update.

9.3. What Happens at Retirement and Re-Entry?

Evidence is stronger for arrival and representation change than for retirement and re-entry. A benchmark with actual availability records could measure removal propagation, stale outputs after removal, and restored usefulness after re-entry. Serving exclusion and removing learned influence are separate objectives and require different outcomes. Inferring retirement from missing interactions would confound availability with demand or exposure.

9.4. When Does Better Access Produce Useful Exposure?

Cold-start benchmarks isolate selected offline populations, while live-streaming studies report outcomes for deployed configurations [1,2,3,7,12]. Connecting them requires changing admission or update policy and measuring both access and downstream response. More new-item interactions can arise from increased exposure even when per-exposure matching worsens. Logged exposure or randomized interventions are therefore needed to establish which part of an observed benefit comes from improved matching, and whether it persists across the item’s available lifetime.

10. Conclusion

The survey connects dynamic-catalog diagnostics to conditional update choices. The inspected comparisons show why warm-item preference need not select an arrival baseline, repaired mappings need not restore new-item prediction, and improvements in resolution, search survival, ranking, and eligibility require different interpretations. Together, these distinctions identify which state to change and which competing update must serve as its control.
The resulting selection logic starts with the failing interface, restricts alternatives to available information, and requires joint updates to justify their dependent migration work. Comparing these alternatives over a common item-availability window brings preparation delay and recurring serving cost into the same decision. The next step is to test when staged access and learning, or coupled representation and model updates, add useful exposure within that window.

References

  1. Zhang, Z.; Zhao, J.; Ma, X.; Xin, X.; de Rijke, M.; Ren, Z. Cold-Starts in Generative Recommendation: A Reproducibility Study. arXiv 2026, arXiv:2603.29845v2. [Google Scholar]
  2. Wang, S.; Huang, Y.; Yang, R.; Wen, S.; Xu, P.; Cao, J.; Liu, Y.; Cai, K.; Guo, C.; Wang, S.; et al. OneLive: Dynamically Unified Generative Framework for Live-Streaming Recommendation. arXiv 2026, arXiv:2602.08612v1. [Google Scholar]
  3. Ye, W.; Liu, G.; Wang, C.; Luo, W.; Wang, S.; Sun, M.; Wang, P.; Yao, Q.; Wu, W.; Jiang, P. TAGR: Temporally Adaptive Generative Recommendation for Industrial Live-Streaming Advertising. arXiv 2026, arXiv:2608.24034v1. [Google Scholar]
  4. Rajput, S.; Mehta, N.; Singh, A.; Hulikal Keshavan, R.; Vu, T.; Heldt, L.; Hong, L.; Tay, Y.; Tran, V.; Samost, J.; et al. Recommender Systems with Generative Retrieval. In Proceedings of the Advances in Neural Information Processing Systems, 2023, Vol. 36, pp. 10299–10315. [CrossRef]
  5. Wang, W.; Bao, H.; Lin, X.; Zhang, J.; Li, Y.; Feng, F.; Ng, S.K.; Chua, T.S. Learnable Item Tokenization for Generative Recommendation. In Proceedings of the Proceedings of the 33rd ACM International Conference on Information and Knowledge Management, 2024; ACM; pp. 2400–2409. [Google Scholar] [CrossRef]
  6. Ju, C.M.; Collins, L.; Neves, L.; Kumar, B.; Wang, L.Y.; Zhao, T.; Shah, N. Generative Recommendation with Semantic IDs: A Practitioner’s Handbook. arXiv 2025, arXiv:2507.22224v1. [Google Scholar]
  7. Ding, Y.; Li, J.; McAuley, J.; Hou, Y. Inductive Generative Recommendation via Retrieval-based Speculation. Proc. AAAI Conf. Artif. Intell. 2026, 40, 14675–14683. [Google Scholar] [CrossRef]
  8. Yang, L.; Paischer, F.; Hassani, K.; Li, J.; Shao, S.; Li, Z.G.; He, Y.; Feng, X.; Noorshams, N.; Park, S.; et al. Unifying Generative and Dense Retrieval for Sequential Recommendation. Trans. Mach. Learn. Res. 2025. [Google Scholar] [CrossRef]
  9. Feng, Y.; Liu, J.; Han, M.; Li, D.; Gu, H.; Zhang, P.; Lu, T.; Gu, N. Drift-Aware Continual Tokenization for Generative Recommendation. arXiv 2026, arXiv:2603.29705v1. [Google Scholar]
  10. Shen, C.; Shi, T.; Yu, W.; Zhang, X.; Xu, J. Bringing Model Editing to Generative Recommendation in Cold-Start Scenarios. arXiv 2026, arXiv:2603.14259v1. [Google Scholar]
  11. Yoo, H.; Li, T.W.; Kang, S.; Liu, Z.; Xu, C.; Qi, Q.; Tong, H. Continual Low-Rank Adapters for LLM-based Generative Recommender Systems. arXiv 2025, arXiv:2510.25093v3. [Google Scholar]
  12. Shi, T.; Li, Z.; Qu, Y.; Liu, Y.; Lai, L.; Jiang, Y. SSRLive: Live Streaming Recommendation with Dynamic Semantic ID. arXiv 2026, arXiv:2606.06970v1. [Google Scholar]
  13. Zhang, T.; Wang, H.; Fan, Y.; Yang, K.; Zeng, J.; Yang, R. A Survey of Item Identifiers in Generative Recommendation: Construction, Alignment, and Generation. TechRxiv 2026. [Google Scholar] [CrossRef]
  14. Hou, M.; Wu, L.; Liao, Y.; Yang, Y.; Zhang, Z.; Wang, Y.; Zheng, C.; Wu, H.; Hong, R. A Survey on Generative Recommendation: Data, Model, and Tasks. AI Open 2026, 7, 169–193. [Google Scholar] [CrossRef]
  15. Li, Y.; Lin, X.; Wang, W.; Feng, F.; Pang, L.; Li, W.; Nie, L.; He, X.; Chua, T.S. A Survey of Generative Search and Recommendation in the Era of Large Language Models. arXiv 2024, arXiv:2404.16924. [Google Scholar]
  16. Yoo, H.; Kang, S.; Tong, H. Continual Recommender Systems. In Proceedings of the Proceedings of the 34th ACM International Conference on Information and Knowledge Management, 2025; pp. 6857–6860. [Google Scholar] [CrossRef]
  17. Li, X.; Jin, J.; Zhou, Y.; Zhang, Y.; Zhang, P.; Zhu, Y.; Dou, Z. From Matching to Generation: A Survey on Generative Information Retrieval. ACM Trans. Inf. Syst. 2025, 43, 1–62. [Google Scholar] [CrossRef]
  18. Ding, J.; Chang, H.; Qin, H.; Liu, T. SIDInspector: A Mapping-First Diagnostic Resource for Semantic-ID Tokenizers. arXiv 2026, arXiv:2606.10375. [Google Scholar]
  19. Ding, J.; Qin, H.; Wu, T.; Cao, Y. SIDScope: A Diagnostic Resource for Semantic-ID Interfaces in Generative Recommendation. arXiv 2026, arXiv:2608.18779. [Google Scholar]
  20. Xu, J.; Chen, Z.; Yang, S.; Li, J.; Wang, W.; Hu, X.; Hoi, S.; Ngai, E.C.H. A Survey on Multimodal Recommender Systems: Recent Advances and Future Directions. IEEE Trans. Multimed. 2026, 28, 7189–7203. [Google Scholar] [CrossRef]
  21. Kang, W.C.; McAuley, J. Self-Attentive Sequential Recommendation. In Proceedings of the 2018 IEEE International Conference on Data Mining (ICDM), 2018; pp. 197–206. [Google Scholar] [CrossRef]
  22. Sun, F.; Liu, J.; Wu, J.; Pei, C.; Lin, X.; Ou, W.; Jiang, P. BERT4Rec: Sequential Recommendation with Bidirectional Encoder Representations from Transformer. In Proceedings of the Proceedings of the 28th ACM International Conference on Information and Knowledge Management, 2019; pp. 1441–1450. [Google Scholar] [CrossRef]
  23. Geng, S.; Liu, S.; Fu, Z.; Ge, Y.; Zhang, Y. Recommendation as Language Processing (RLP): A Unified Pretrain, Personalized Prompt & Predict Paradigm (P5). In Proceedings of the Proceedings of the 16th ACM Conference on Recommender Systems, 2022; pp. 299–315. [Google Scholar] [CrossRef]
  24. Zhai, J.; Liao, L.; Liu, X.; Wang, Y.; Li, R.; Cao, X.; Gao, L.; Gong, Z.; Gu, F.; He, J.; et al. Actions Speak Louder than Words: Trillion-Parameter Sequential Transducers for Generative Recommendations. In Proceedings of the Proceedings of the 41st International Conference on Machine Learning, 2024, Vol. 235, Proceedings of Machine Learning Research, pp. 58484–58509.
  25. Zhou, G.; Deng, J.; Zhang, J.; Cai, K.; Ren, L.; Luo, Q.; Wang, Q.; Hu, Q.; Huang, R.; Wang, S.; et al. OneRec Technical Report. arXiv 2025, arXiv:2506.13695. [Google Scholar]
  26. Tay, Y.; Tran, V.Q.; Dehghani, M.; Ni, J.; Bahri, D.; Mehta, H.; Qin, Z.; Hui, K.; Zhao, Z.; Gupta, J.; et al. Transformer Memory as a Differentiable Search Index. In Proceedings of the Advances in Neural Information Processing Systems, 2022. [Google Scholar]
  27. Wang, Y.; Hou, Y.; Wang, H.; Miao, Z.; Wu, S.; Sun, H.; Chen, Q.; Xia, Y.; Chi, C.; Zhao, G.; et al. A Neural Corpus Indexer for Document Retrieval. In Proceedings of the Advances in Neural Information Processing Systems, 2022; Vol. 35. [Google Scholar]
  28. Oord, A.v.d.; Vinyals, O.; Kavukcuoglu, K. Neural Discrete Representation Learning. In Proceedings of the Advances in Neural Information Processing Systems; 2017. [Google Scholar]
  29. Lee, D.; Kim, C.; Kim, S.; Cho, M.; Han, W.S. Autoregressive Image Generation using Residual Quantization. In Proceedings of the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022; pp. 11523–11532. [Google Scholar]
  30. Zheng, B.; Hou, Y.; Lu, H.; Chen, Y.; Zhao, W.X.; Chen, M.; Wen, J.R. Adapting Large Language Models by Integrating Collaborative Semantics for Recommendation. In Proceedings of the 2024 IEEE 40th International Conference on Data Engineering (ICDE), 2024; pp. 1435–1448. [Google Scholar] [CrossRef]
  31. Zhu, J.; Jin, M.; Liu, Q.; Qiu, Z.; Dong, Z.; Li, X. CoST: Contrastive Quantization based Semantic Tokenization for Generative Recommendation. In Proceedings of the 18th ACM Conference on Recommender Systems, 2024; pp. 969–974. [Google Scholar] [CrossRef]
  32. Tan, J.; Xu, S.; Hua, W.; Ge, Y.; Li, Z.; Zhang, Y. IDGenRec: LLM-RecSys Alignment with Textual ID Learning. In Proceedings of the Proceedings of the 47th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2024; pp. 355–364. [Google Scholar] [CrossRef]
  33. Hou, Y.; Mu, S.; Zhao, W.X.; Li, Y.; Ding, B.; Wen, J.R. Towards Universal Sequence Representation Learning for Recommender Systems. In Proceedings of the Proceedings of the 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2022; ACM; pp. 585–593. [Google Scholar] [CrossRef]
  34. Li, J.; Wang, M.; Li, J.; Fu, J.; Shen, X.; Shang, J.; McAuley, J. Text Is All You Need: Learning Language Representations for Sequential Recommendation. In Proceedings of the Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2023; ACM; pp. 1258–1267. [Google Scholar] [CrossRef]
  35. Yang, Y.; Ji, Z.; Li, Z.; Li, Y.; Mo, Z.; Ding, Y.; Chen, K.; Zhang, Z.; Li, J.; Li, S.; et al. Sparse Meets Dense: Unified Generative Recommendations with Cascaded Sparse-Dense Representations. arXiv 2025, arXiv:2503.02453. [Google Scholar]
  36. Loveland, D.; Collins, L.; Kumar, B.; Koutra, D.; Shah, N. Preserving Item Semantics for Free: Rethinking Token Initialization in LLM-Based Generative Recommendation. arXiv 2026, arXiv:2608.07816. [Google Scholar]
  37. Peng, J.; Zheng, Y.; Zhe, Z.; Tong, B.; Wang, G.; Zheng, B. Can Generative Recommendation Reach Cold Items? A Temporal Perspective on Semantic-ID Generation. arXiv 2026, arXiv:2607.21101v1. [Google Scholar]
  38. Qiu, R.; Yin, H.; Huang, Z.; Chen, T. GAG: Global Attributed Graph Neural Network for Streaming Session-based Recommendation. In Proceedings of the Proceedings of the 43rd International ACM SIGIR Conference on Research and Development in Information Retrieval, 2020; pp. 669–678. [Google Scholar] [CrossRef]
  39. He, B.; He, X.; Zhang, Y.; Tang, R.; Ma, C. Dynamically Expandable Graph Convolution for Streaming Recommendation. In Proceedings of the Proceedings of the ACM Web Conference 2023, 2023; pp. 1457–1467. [Google Scholar] [CrossRef]
  40. Kim, W.; Yoo, H.; Kim, J.; Lim, J.; Kang, S.; Yu, H. TRACER: Balancing Stability-Plasticity-Cognitivity Trilemma for LLM Enhanced Continual Recommendation. arXiv 2026, arXiv:2608.16075. [Google Scholar]
  41. Baek, S.; Lee, G.; Lee, S.; Kweon, W.; Wang, D.; Kang, S. SCoRD: Semantic-Assisted Continual Retriever-Reranker Distillation for LLM-Based Recommendation. arXiv 2026, arXiv:2608.19998. [Google Scholar]
  42. Fang, Z.; Huang, Y.; Zhang, A.; He, Y.; Xiao, R.; Li, C.; Yetim, Y.; Yang, S.; Wei, X.; Tian, F.; et al. GRACE: Generative Recommender Acceleration Engine for Real-Time Ads Retrieval. arXiv 2026, arXiv:2608.00938. [Google Scholar]
  43. Zhang, Q.; Szymanski, L.; Zhang, H.; Deng, J.D. Faithful Evaluation of Semantic-ID Tokenizers for Generative Recommendation. arXiv 2026, arXiv:2605.25330. [Google Scholar]
  44. Jiao, C.; Elenter, J.; Ravichandran, P.; Huber, B.; Cauteruccio, J.; Wasson, T.; Heath, T.; Xiong, C.; Lalmas, M.; Bennett, P. Efficient Dataset Selection for Continual Adaptation of Generative Recommenders. arXiv 2026, arXiv:2604.07739. [Google Scholar]
  45. Xie, T.; Ku, X.; Sun, M.; Sha, Y.; Wang, L.; Wang, P.; Wang, Y.; Wu, W.; Liu, Z.; Jiang, P.; et al. From a Static Multi-Level Small Semantic Codebook to a Dynamic Single-Level Large Semantic Codebook for Generative Recommendation. arXiv 2026, arXiv:2608.21012. [Google Scholar]
  46. Shi, H.; Lin, X.; Wang, W.; Shi, W.; Pan, J.; Jie, J.; Feng, F. Incremental Learning for LLM-based Tokenization and Recommendation. In Proceedings of the Proceedings of the 34th ACM International Conference on Information and Knowledge Management, 2025; pp. 2643–2652. [Google Scholar] [CrossRef]
  47. Nian, D.; Fu, D.; Xu, C.; Xia, Y.; Li, H.; Yan, H.; Kang, J. ChronoID: Infusing Explicit Temporal Signals into Semantic IDs for Generative Recommendation. arXiv 2026, arXiv:2606.14260v1. [Google Scholar]
  48. Shi, T.; Zhang, Y.; Xu, Z.; Chen, C.; Feng, F.; He, X.; Tian, Q. Preliminary Study on Incremental Learning for Large Language Model-based Recommender Systems. In Proceedings of the Proceedings of the 33rd ACM International Conference on Information and Knowledge Management, 2024; pp. 4051–4055. [Google Scholar] [CrossRef]
  49. Xu, J.; Chen, Z.; Yang, S.; Li, J.; Wang, H.; Ngai, E.C.H. MENTOR: Multi-level Self-supervised Learning for Multimodal Recommendation. Proc. AAAI Conf. Artif. Intell. 2025, 39, 12908–12917. [Google Scholar] [CrossRef]
  50. Xu, J.; Chen, Z.; Li, J.; Yang, S.; Wang, H.; Li, Y.; Li, M.; Wu, P.; Ngai, E.C.H. MDVT: Enhancing Multimodal Recommendation with Model-Agnostic Multimodal-Driven Virtual Triplets. In Proceedings of the Proceedings of the 31st ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.2, 2025; pp. 3378–3389. [Google Scholar] [CrossRef]
  51. Xu, J.; Chen, Z.; Yang, S.; Li, J.; Wan, Z.; Wang, H.; Liu, W.; Li, Y.; Ngai, E.C.H. VI-MMRec: Similarity-Aware Training Cost-free Virtual User-Item Interactions for Multimodal Recommendation. Proceedings of the Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining 2026, V.1, 1683–1692. [Google Scholar] [CrossRef]
  52. Xu, J.; Chen, Z.; Yang, S.; Li, J.; Wang, H.; Tang, J.; Wang, W.; Hu, X.; Ngai, E.C.H. Well Begun is Half Done: Training-Free and Model-Agnostic Semantically Guaranteed User Representation Initialization for Multimodal Recommendation. In Proceedings of the Proceedings of the 49th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2026; pp. 2096–2106. [Google Scholar] [CrossRef]
  53. Zhang, K.; Zhang, J.; Cheng, W.; Cheng, Y.; Zhang, J.; Lu, H.; Zhang, X.; Gan, H.; Cao, J.; Wang, T.; et al. OneMall: One Model, More Scenarios – End-to-End Generative Recommender Family at Kuaishou E-Commerce. arXiv 2026, arXiv:2601.21770v1. [Google Scholar]
  54. Mehta, S.V.; Gupta, J.; Tay, Y.; Dehghani, M.; Tran, V.Q.; Rao, J.; Najork, M.; Strubell, E.; Metzler, D. DSI++: Updating Transformer Memory with New Documents. In Proceedings of the Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023. [Google Scholar]
  55. Kishore, V.; Wan, C.; Lovelace, J.; Artzi, Y.; Weinberger, K.Q. IncDSI: Incrementally Updatable Document Retrieval. In Proceedings of the Proceedings of the 40th International Conference on Machine Learning, 2023. [Google Scholar]
  56. Guo, J.; Zhou, C.; Zhang, R.; Chen, J.; de Rijke, M.; Fan, Y.; Cheng, X. CorpusBrain++: A Continual Generative Pre-Training Framework for Knowledge-Intensive Language Tasks. ACM Trans. Inf. Syst. 2026, 44, 1–35. [Google Scholar] [CrossRef]
  57. Chen, J.; Zhang, R.; Guo, J.; de Rijke, M.; Chen, W.; Fan, Y.; Cheng, X. Continual Learning for Generative Retrieval over Dynamic Corpora. In Proceedings of the Proceedings of the 32nd ACM International Conference on Information and Knowledge Management, 2023; pp. 306–315. [Google Scholar] [CrossRef]
  58. Kim, C.; Yoon, S.; Lee, H.; Jang, J.; Yang, S.; Seo, M. Exploring the Practicality of Generative Retrieval on Dynamic Corpora. In Proceedings of the Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing; Al-Onaizan, Y., Bansal, M., Chen, Y.N., Eds.; Miami, Florida, USA, 2024; pp. 13616–13633. [Google Scholar] [CrossRef]
  59. Zhang, Z.; Ma, X.; Sun, W.; Ren, P.; Chen, Z.; Wang, S.; Yin, D.; de Rijke, M.; Ren, Z. Replication and Exploration of Generative Retrieval over Dynamic Corpora. In Proceedings of the Proceedings of the 48th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2025; pp. 3325–3334. [Google Scholar] [CrossRef]
  60. Huynh, T.L.; Vu, T.T.; Wang, W.; Le, T.; Gasevic, D.; Li, Y.F.; Do, T.T. MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing; Christodoulopoulos, C., Chakraborty, T., Rose, C., Peng, V., Eds.; Suzhou, China, 2025; pp. 380–396. [Google Scholar] [CrossRef]
  61. Penha, G.; D’Amico, E.; De Nadai, M.; Palumbo, E.; Tamborrino, A.; Vardasbi, A.; Lefarov, M.; Lin, S.; Heath, T.; Fabbri, F.; et al. Semantic IDs for Joint Generative Search and Recommendation. In Proceedings of the Proceedings of the Nineteenth ACM Conference on Recommender Systems; 2025; pp. 1296–1301. [Google Scholar] [CrossRef]
  62. Khrylchenko, K. Variable-Length Semantic IDs for Recommender Systems. arXiv 2026, arXiv:2602.16375. [Google Scholar]
  63. Li, L.; Tu, Z.; Chu, D.; Sun, H. TAAL: Mitigating Early Beam Pruning in Generative Recommendation via Temporal Autoregressive Alignment. arXiv 2026, arXiv:2608.29179. [Google Scholar]
  64. Hou, Y.; Li, J.; Fu, X.; He, Z.; Yan, A.; Chen, X.; McAuley, J. Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders. arXiv 2024, arXiv:2403.03952. [Google Scholar]
  65. Ni, Y.; Cheng, Y.; Liu, X.; Fu, J.; Li, Y.; He, X.; Zhang, Y.; Yuan, F. A Content-Driven Micro-Video Recommendation Dataset at Scale. In Proceedings of the Proceedings of the 34th ACM International Conference on Information and Knowledge Management, 2025; pp. 6486–6491. [Google Scholar] [CrossRef]
  66. Wu, F.; Qiao, Y.; Chen, J.H.; Wu, C.; Qi, T.; Lian, J.; Liu, D.; Xie, X.; Gao, J.; Wu, W.; et al. MIND: A Large-scale Dataset for News Recommendation. In Proceedings of the Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics;Online; Jurafsky, D., Chai, J., Schluter, N., Tetreault, J., Eds.; 2020; pp. 3597–3606. [Google Scholar] [CrossRef]
  67. Gao, C.; Li, S.; Lei, W.; Chen, J.; Li, B.; Jiang, P.; He, X.; Mao, J.; Chua, T.S. KuaiRec: A Fully-observed Dataset and Insights for Evaluating Recommender Systems. In Proceedings of the Proceedings of the 31st ACM International Conference on Information & Knowledge Management, 2022; pp. 540–550. [Google Scholar] [CrossRef]
  68. Ji, Y.; Sun, A.; Zhang, J.; Li, C. A Critical Study on Data Leakage in Recommender System Offline Evaluation. ACM Trans. Inf. Syst. 2023, 41, 1–27. [Google Scholar] [CrossRef]
  69. Ng, T.K.; Khoo, J.T.K.; Sun, A. RecNextEval: A Reference Implementation for Temporal Next-Batch Recommendation Evaluation. arXiv 2026, arXiv:2604.13665. [Google Scholar]
  70. Krichene, W.; Rendle, S. On Sampled Metrics for Item Recommendation. In Proceedings of the Proceedings of the 26th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, New York, NY, USA, 2020; pp. 1748–1757. [Google Scholar] [CrossRef]
Figure 1. Feedback learning and catalog maintenance share serving dependencies. The outer loop links interaction history, generator updates, recommendation, and subsequent feedback. The inner branch distinguishes catalog membership, code assignment, and candidate access; the lower admission arc represents an alternative content-index route. The dashed link denotes code compatibility with the generator. These are logical interfaces that can be combined, rather than a mandatory architecture or update order. Eligibility and ranking still determine the returned list. The framework synthesizes admission, adaptation, mapping, and eligibility concerns [7,9,19,42].
Figure 1. Feedback learning and catalog maintenance share serving dependencies. The outer loop links interaction history, generator updates, recommendation, and subsequent feedback. The inner branch distinguishes catalog membership, code assignment, and candidate access; the lower admission arc represents an alternative content-index route. The dashed link denotes code compatibility with the generator. These are logical interfaces that can be combined, rather than a mandatory architecture or update order. Eligibility and ranking still determine the returned list. The framework synthesizes admission, adaptation, mapping, and eligibility concerns [7,9,19,42].
Preprints 233953 g001
Figure 2. Intervention map derived from the inspected mechanisms [3,7,8,9,10,11,12,42,44,45]. Filled squares locate the main intervention discussed here; open circles identify supporting components. Dashes indicate no role assigned in this comparison, rather than absent capability.
Figure 2. Intervention map derived from the inspected mechanisms [3,7,8,9,10,11,12,42,44,45]. Filled squares locate the main intervention discussed here; open circles identify supporting components. Dashes indicate no role assigned in this comparison, rather than absent capability.
Preprints 233953 g002
Figure 3. Three hybrid interface classes organized by candidate entry and scoring order. The orange item record follows the same illustrative new item through each route. Explicit branches show proposal, external-set union, and admission to the dense search domain; generated support remains an additional condition in the third route. Points in the search domain are schematic item representations, not measured embeddings. The classes can be combined. Representative implementations are SpecGR for proposal–verification, LIGER for candidate merging, and COBRA for generation-conditioned continuous retrieval [7,8,35]. The diagram abstracts their shared interface roles rather than reproducing each implementation.
Figure 3. Three hybrid interface classes organized by candidate entry and scoring order. The orange item record follows the same illustrative new item through each route. Explicit branches show proposal, external-set union, and admission to the dense search domain; generated support remains an additional condition in the third route. Points in the search domain are schematic item representations, not measured embeddings. The classes can be combined. Representative implementations are SpecGR for proposal–verification, LIGER for candidate merging, and COBRA for generation-conditioned continuous retrieval [7,8,35]. The diagram abstracts their shared interface roles rather than reproducing each implementation.
Preprints 233953 g003
Figure 4. Two clocks in the running lifecycle. (a) Metadata and stored feedback may remain after retirement, whereas serving eligibility ends. Admission at 09:02 is separate from eligibility and from later feedback. (b) The new assignment, encoded histories, resolver, and model remain staged until their joint activation at 09:30. New requests then use v1; in-flight requests retain v0. The switch policy is illustrative, not an attributed implementation or measured trace. Event spacing does not encode duration.
Figure 4. Two clocks in the running lifecycle. (a) Metadata and stored feedback may remain after retirement, whereas serving eligibility ends. Admission at 09:02 is separate from eligibility and from later feedback. (b) The new assignment, encoded histories, resolver, and model remain staged until their joint activation at 09:30. New requests then use v1; in-flight requests retain v0. The switch policy is illustrative, not an attributed implementation or measured trace. Event spacing does not encode duration.
Preprints 233953 g004
Table 8. Constructed request trace for the two-route example. Age is request time minus target publication time, in minutes. The unlabeled request has no known target or target age; its serving cost is still counted. R and J share all requests and labels.
Table 8. Constructed request trace for the two-route example. Age is request time minus target publication time, in minutes. The unlabeled request has no known target or target age; its serving cost is still counted. R and J share all requests and labels.
Request time Target Age Label usable State in J R hit J hit
09:01 A 1 09:10 Old; unadmitted 0 0
09:05 A 5 09:10 Old 1 1
09:15 A 15 09:18 Old 0 0
09:25 B 5 09:26 Old 1 1
09:35 A 35 09:36 Updated 0 1
09:40 Unknown – Unlabeled Updated – –
09:50 B 30 10:10 Updated 1 1
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.