Submitted:
09 August 2026
Posted:
11 August 2026
You are already at the latest version
Abstract
Human feedback is a key mechanism for aligning large language models (LLMs), yet it appears in complex forms that vary in expressive-ness, granularity, and cost. While recent work has explored richer feedback beyond binary preferences, the literature remains fragmented, and the relationships between feedback structure , optimization methods, and evaluation protocols are often unclear. This survey provides a feedback-centric overview of LLM alignment. We propose a taxonomy that organizes complex feedback by supervision form and orthogonal dimensions, use it to categorize alignment methods into reward-model-free and reward-model-based approaches, and survey evaluation practices, distinguishing direct verification from proxy measures derived from human feedback. By synthesizing results across feedback modalities, optimization strategies, and evaluation designs, we clarify their interdependen-cies, identify recurring mismatches in current pipelines, and highlight underexplored settings for future feedback-driven alignment research.
Keywords:
feedback-driven alignment
; complex human feedback
; large language models
; model alignment
; evaluation metrics
1. Introduction
Human feedback is central to aligning LLMs with human preferences, values, and task-specific goals, yet it is not a homogeneous signal: it takes multiple forms that differ in expressiveness, annotation cost, and how directly they specify desired improvements. Most large-scale alignment pipelines nonetheless rely mainly on binary preference labels, which are easy to collect but information-sparse (Wu et al., 2023), ambiguous across factual, stylistic, and safety dimensions (Casper et al., 2023; Chaudhari et al., 2025), and outcome-focused, offering little guidance for multi-step reasoning or structured decision-making tasks.
Recent work has demonstrated the benefits of richer feedback: linguistic feedback can be more efficient than scalar rewards (Scheurer et al., 2023), no single modality is optimal across tasks (Yu et al., 2025), and supervision on intermediate reasoning steps improves generalization (Uesato et al., 2022; Lightman et al., 2024). To compare these roles empirically, we conduct a quantitative meta-analysis on HelpSteer3 Wang et al. (2025), quantifying the mutual information contributed by binary preferences, scores, and language toward recovering the aggregated preference (Figure 1; details and a domain-specific breakdown in Appendix C).
Our results show that binary preferences already capture most information needed to recover the preference, while scores and language add only modest value; however, linguistic feedback remains informative for calibration and diagnostics, explaining score magnitude and annotator disagreement and aligning with the semantic direction from rejected to preferred responses. This suggests a trade-off: binary supervision is often efficient for broad alignment, whereas language feedback is more useful for targeted improvement.
We stress that these results do not imply language feedback is uniformly more or less informative than binary preferences; rather, informativeness is objective-dependent: binary comparisons suffice to recover a general preference direction, but more specific objectives, such as localizing where a reasoning chain goes wrong or expressing a stylistic or constraint-based requirement, need richer supervision that scalar or binary labels cannot supply.
To fill this gap, we adopt a feedback-centric perspective, proposing a taxonomy that characterizes feedback by its form and several orthogonal dimensions and shows how these map to corresponding training and evaluation frameworks (Figure 2). Unlike prior reviews organized mainly around algorithms or pipelines, our survey centers complex feedback itself, organizing the literature by its structure and semantics (Table 1).
Beyond offering broader coverage of the alignment pipeline (Table 1), organizing the survey around feedback structure enables a form of analysis that prior, component-centric surveys do not directly support: because feedback connects data collection, training, and evaluation, treating it as the central object exposes mismatches across these stages, as we discuss in Section 5.
Table 1.
Comparison with representative surveys. A checkmark (✓) indicates substantial coverage of the corresponding aspect of the alignment pipeline, while a circle (∘) indicates partial discussion.
Table 1.
Comparison with representative surveys. A checkmark (✓) indicates substantial coverage of the corresponding aspect of the alignment pipeline, while a circle (∘) indicates partial discussion.
| Survey | Main Perspective | Data | Preference Modeling | Optimization | Evaluation |
| Fernandes et al. (2023) | Human feedback in NLG | ✓ | ✓ | ||
| Jiang et al. (2024) | Preference learning | ✓ | ✓ | ||
| Liu et al. (2025) | Optimization algorithms | ✓ | ✓ | ||
| Ji et al. (2025) | Reward design and reward modeling | ∘ | ✓ | ✓ | |
| Lai et al. (2025) | Post-training scaling | ✓ | ✓ | ✓ | ∘ |
| Ours | Feedback structure and semantics | ✓ | ✓ | ✓ | ✓ |
This survey makes the following contributions:
• We introduce a feedback-centric view on LLM alignment and propose a systematic taxonomy of complex feedback by its form and semantics that fundamentally shape learning and evaluation.
• We propose a taxonomy that categorizes existing alignment methods by the feedback types they assume, and introduce a corresponding taxonomy of alignment evaluation based on learned or inferred human preferences.
• We identify mismatches and underexplored regimes among feedback types, optimization objectives, and evaluation designs; and further outline directions for designing more faithful and data-efficient feedback-driven alignment.
We describe our literature search process, inclusion criteria, and scope in Appendix A.
2. A Taxonomy of Complex Feedback
Feedback is central to LLM alignment but varies widely in form, granularity, and cost. We organize feedback by its supervision form—how evaluative information is expressed. Orthogonal dimensions such as interaction structure, feedback source, and granularity further characterize feedback and help situate existing datasets and alignment methods.
Figure 3.
Taxonomy of feedback types based on signal form.

2.1. Ordinal Feedback
Ordinal feedback captures relative order between outputs without specifying preference strength.
2.1.1. Binary Preference
Binary preference feedback is the most common form in RLHF. Given a prompt and two candidate outputs , annotators select the preferred response, which we encode as
where ≻ denotes preference, with ties typically discarded or mapped to a neutral label. Binary preferences are easy to collect, and integrate naturally with pairwise reward modeling and policy optimization. Yet, they encode only relative order, not preference strength, and can be ambiguous when trade-offs across multiple attributes are involved.
Representative Datasets. Binary preference signals are widely used in RLHF, e.g., summarization (Stiennon et al., 2020), UltraFeedback (Cui et al., 2023), and hh-rlhf (Bai et al., 2022).
2.1.2. List Preference
List preference feedback generalizes binary comparisons by allowing annotators to rank a set of candidate outputs for the same prompt. Though such a total ordering can be decomposed into pairwise comparisons, listwise supervision exploits joint judgments in a shared context and encourages globally consistent rankings, which may improve sample efficiency and ranking coherence. Listwise feedback is more informative than binary comparisons but becomes cognitively demanding as m increases. In practice, it is often relaxed to partial orderings, such as top-k or bottom-k selection (marking only the best or worst few outputs) or grouping candidates into quality tiers without ordering within tiers.
Representative Datasets. Few public datasets retain listwise annotations; examples include OpenAssistant Conversations (öpf et al., 2023) and RewardBench v2 (Malik et al., 2025).
2.2. Cardinal Feedback
Cardinal feedback provides absolute assessments of model outputs using discrete labels (e.g., poor–good) or numerical scores (e.g., in ). Unlike ranking-based feedback, it captures preference magnitude and enables smoother rewards and cross-instance comparison. However, cardinal ratings are more subjective and noisy, as annotators may use scales inconsistently and exhibit systematic biases (e.g., central tendency, order effects), requiring careful scale design and calibration at scale.
2.2.1. Scalar Rating
Scalar rating feedback assigns a single numerical or categorical score to each output, typically reflecting overall quality or a specific attribute. For a prompt and output , the annotator provides where with and denoting the lowest and highest possible ratings. Compared to binary or pairwise feedback, scalar ratings capture preference strength and support smoother, potentially more sample-efficient reward modeling.
Representative Datasets. UltraFeedback (Cui et al., 2023) and WebGPT (Nakano et al., 2021) provide scalar human ratings; MultiPref (Miranda et al., 2024) additionally encodes annotator confidence for calibration.
2.2.2. Multi-Aspect Rating
Multi-aspect rating feedback extends scalar ratings to multiple predefined dimensions, providing a vector of scores for each output: where each corresponds to an attribute. A common implementation is rubric-based evaluation, where detailed scoring guidelines specify criteria for each dimension to reduce inter-annotator variance and improve reliability. Attribute-specific supervision decomposes evaluation into interpretable components, enabling multi-objective alignment and fine-grained diagnostics.
Representative Datasets. Examples include HelpSteer2 (Wang et al., 2024), UltraFeedback (Cui et al., 2023), and OpenAssistant Conversations (öpf et al., 2023); other evaluation frameworks Zheng et al. (2023) similarly rely on multi-aspect scales.
2.3. Linguistic Feedback
Linguistic feedback conveys evaluative or corrective information through natural language. We distinguish signals that directly prescribe revised behavior from those that describe and guide it.
2.3.1. Prescriptive Feedback
Prescriptive feedback provides direct supervision on how to improve a model’s response. Given a prompt and output , the feedback yields an improved target , typically in one of two forms: (1) Edits: Localized modifications (insertions, deletions, replacements) that specify how to revise the original response; (2) Demonstrations: Rewrites or example outputs that instantiate the desired behavior and can be used as targets for imitation or distillation. Although prescriptive feedback can in principle be converted into binary preference, they go beyond ranking by encoding interpretable transformations that directly guide alignment.
Representative Datasets. Edit-based examples include AVS Edits (Mishra et al., 2024; Yao et al., 2023), WikiPrefs (Majkutewicz and Szymański 2025), and HelpSteer3 (Wang et al., 2025); demonstration-style supervision is exemplified by Anthropic-HH-Golden (Cai et al., 2023).
2.3.2. Descriptive Feedback
Descriptive feedback communicates evaluation and guidance through natural language rather than directly modifying a model output. Given a prompt and output , we write , where is a linguistic comment on qualitative properties or desired changes (e.g., “the argument lacks evidence,” “make the tone more formal”).
Unlike edits or demonstrations, descriptive feedback does not replace the original output but provides meta-level comments. This supports supervision over high-level attributes such as coherence, politeness, and persuasiveness that are not easily captured through token-level corrections, while requiring the model to interpret and operationalize abstract guidance, which poses challenges.
Representative Datasets. Examples include HelpSteer3 (Wang et al., 2025), TofuEval (Tang et al., 2024), and SLF5K (Scheurer et al., 2023).
2.4. Orthogonal Dimensions of Feedback Taxonomy
The form-centric taxonomy specifies the mathematical representation of supervision but abstracts away operational context. Therefore, it is also useful to analyze signals along several orthogonal dimensions. In practice, feedback with the same form may differ substantially in reliability due to its source. These additional dimensions help disentangle data quality from other factors.
2.4.1. Interaction Structure
This dimension characterizes the temporal scope and dependency of the feedback signal.
•Single-step Feedback: Feedback targets a single turn independently of past or future interactions. This is commonly modeled under an i.i.d. assumption, which simplifies reward learning but may neglect long-term coherence or planning.
•Multi-step and Agentic Feedback: Multi-step feedback captures dependencies across turns, supporting multi-turn consistency and constraint satisfaction, as in OpenAssistant (öpf et al., 2023), LMSYS Chatbot Arena (Chiang et al., 2024), and MT-Bench-101 (Bai et al., 2024). This setting is increasingly relevant to agentic interaction, where feedback covers an entire trajectory of tool calls and actions rather than a single response (Shani et al., 2024; Gao et al., 2024; Shinn et al., 2023).
2.4.2. Feedback Source
This dimension characterizes the origin of the supervision signal, ranging from subjective human judgments to objective verifications.
•Human-Explicit: Human-provided annotations. They remain the gold standard for capturing human values, but are costly to collect at scale and subject to inter-annotator disagreement, which bounds label reliability; Appendix C quantifies this trade-off empirically.
•Human-Implicit: Behavioral signals that indirectly reveal human preferences without requiring explicit annotation, including WildFeedback (Shi et al., 2024) and Liu et al. (2025), which provide supervision from real-world user interaction logs.
•Model-Generated: Signals produced by models rather than by humans, offering improved scalability at the cost of potential bias or drift from human intent. Examples include Constitutional AI (Bai et al., 2022), which generates critiques and preferences guided by human-written constitutions; RFM-RM-as-User (Barreto et al., 2025), which simulates user feedback via trained reward models; and LLM-as-a-Judge approaches.
•Environment-Verifiable: Objective feedback derived from execution environments, test cases, or formal verification. These signals are typically low-noise but are limited to domains with externally checkable correctness, such as coding or instruction following. Examples include EvalPlus (Liu et al., 2023), EvalPerf (Liu et al., 2024), OpenCodeInstruct (Ahmad et al., 2025), IFEval (Dai et al., 2025), and Reveal (Jacovi et al., 2024). Tool-use feedback in agentic settings is a further instance of this source: an agent’s individual tool calls can be checked against environment state, execution results, or task completion criteria.
2.4.3. Feedback Granularity
Feedback granularity specifies the scope of supervision relative to model generation, ranging from holistic judgments to localized annotations.
•Outcome-level: Feedback applies to the final output y and is standard in RLHF. While cost-effective and aligned with user utility, it provides limited localization and exacerbates credit assignment.
•Process-level: Feedback targets intermediate steps or generation trajectories, including PRM800K (Song et al., 2025) and MathShepherd (Wang et al., 2024). In agentic settings, this granularity extends to trajectory-level rewards over an entire episode, per-tool-call feedback on individual actions, and subgoal rewards for intermediate objectives within a longer task, all of which are fine-grained instances of process-level supervision.
2.4.4. Extension to Multimodal Feedback
Although the discussion above centers on text, the same forms and dimensions extend to multimodal generation. Ordinal and cardinal feedback have already been collected over non-textual outputs, e.g., ranked human preferences over generated images (Karthik et al., 2025) and preference-based fine-tuning of diffusion models (Yuan et al., 2024); linguistic feedback in these settings can likewise reference non-textual content, such as visual composition or motion, rather than text alone. Adapting the granularity and source dimensions to multimodal outputs raises distinct challenges, including localization below the level of a whole image or clip and verification with vision-language rather than text-only judges.
Considering these orthogonal dimensions clarifies practical trade-offs central to deploying feedback at scale, including annotation cost, feedback reliability, inter-annotator disagreement, and scalability, alongside the granularity of supervision, and informs the choice of objectives and modeling strategies across alignment settings.
2.5. Scope of the Taxonomy
Category Boundaries: Different types of feedback can convey a judgment about the same concept, but they vary in how that judgment is represented and operationalized for training.
Domains and Languages: The taxonomy is defined by feedback structure rather than task domain or language, so it applies directly to multilingual and domain-specific settings without modification; what varies across settings is which feedback types are practical to collect.
3. Methodology for LLMs Alignment via Complex Human Feedback
After categorizing the feedback data, an important question remains: how can this feedback be effectively leveraged to improve the model?
This section reviews two main categories of alignment methods: (1) Reward-Model-Free Alignment approaches, which directly optimize the language model using feedback without relying on an additional reward model; and (2) Reward-Model-Based Alignment approaches, which either first train a reward model from feedback data, or use a reward model to generate the corresponding feedback and then perform alignment using the reward signal. Our goal is not to provide a deeper algorithmic re-coverage compared with prior studies, but to contribute a feedback-centric view that organizes alignment methods tied to the feedback structure. Figure 4 gives a compact overview of representative methods discussed in this section, organized by feedback type and by whether they are reward-model-free or reward-model-based; Appendix B synthesizes, for each feedback type, the assumptions, suitability, and information exploited by each family of methods.
3.1. Alignment with Ordinal Feedback
Reward-Model-Free Alignment. This line of work (Rafailov et al., 2023; Zhao et al., 2023; Liu et al., 2023; Azar et al., 2024; Meng et al., 2024; Wu et al., 2024; Xiao et al., 2024; Liu et al., 2024; Fisch et al., 2024; Deng et al., 2025) eliminates the need for explicitly training a reward model, and instead directly optimizes the policy from offline preference data by minimizing a pairwise ranking loss defined over preferred and non-preferred responses. Direct Preference Optimization (DPO) (Rafailov et al., 2023), for instance, leverages a closed-form reparameterization that casts the optimal policy as a classifier over preference pairs, enabling stable and efficient fine-tuning without intermediate reward model training. Identity Preference Optimization (IPO) (Azar et al., 2024), derived as a special case of the general -preference optimization framework with instantiated as the identity function, relaxes the Bradley–Terry assumption underlying DPO. Beyond reliance on pre-collected preference datasets, self-play–based methods (Chen et al., 2024; Gao et al., 2025; Wang et al., 2025) treat model-generated responses as dispreferred and human responses as preferred, enabling automatic data construction and self-improvement.
Reward-Model-based Alignment. Preference annotations are first used to train a reward model with a pairwise ranking objective, such as the Bradley–Terry loss (Stiennon et al., 2020; Bradley and Terry 1952) or its variants (Liu et al., 2024); the language model is then optimized against this reward via reinforcement learning (Williams 1992; Schulman et al., 2017; Ahmadian et al., 2024) or inference-time methods such as Best-of-n (Stiennon et al., 2020; Nakano et al., 2021). Several studies extend this to an iterative setting, using the reward model to construct new pairwise samples for successive policy optimization (Dong et al., 2024; Xiong et al., 2024; Zhang et al., 2024; Rosset et al., 2024; Xie et al., 2024); Razin et al. further argues that effectiveness also depends on preserving sufficient reward variance.
From binary to list preferences. Extending alignment methods from binary to list preferences typically entails generalizing pairwise ranking objectives to listwise formulations, such as ListMLE or contrastive losses (Karthik et al., 2025; Zhu et al., 2024; Song et al., 2024; Liu et al., 2025; Jin et al., 2025). Compared to the binary setting, listwise preference learning jointly models the relative ordering among multiple candidate responses, providing richer signals and enabling hard example mining as well as globally consistent preference modeling (Cao et al., 2007; Zhu et al., 2023).
3.2. Alignment with Cardinal Feedback
Reward-Model-Free Alignment. KTO (Ethayarajh et al., 2024) introduces a human-aware loss framework grounded in prospect theory, which maximizes the utility of model generations using only binary desirable/undesirable feedback. More recently, a growing body of work (Rame et al., 2023; Zhou et al., 2024; Guo et al., 2024; Yang et al., 2024; Zhong et al., 2024) has shifted alignment research toward more flexible, multi-objective, and personalized paradigms, aiming to better capture heterogeneous human values and support controllable preferences. For instance, Controllable Preference Optimization (Guo et al., 2024) and Rewards-in-Context (Yang et al., 2024) incorporate user-specified, multi-dimensional reward signals directly into the model input, enabling dynamic preference adjustment at inference time.
Reward-Model-Based Alignment. Sun et al. (2024) proposed an alternative reward modeling objective, which preserves order consistency without requiring pairwise preference annotations. In addition, several studies (Wang et al., 2024; Adler et al., 2024; Wang et al., 2024; Lin et al., 2025) leverage multi-dimensional scalar feedback (e.g., helpfulness, safety) to train multi-objective reward models that capture diverse and potentially competing human preferences.
3.3. Alignment with Linguistic Feedback
Reward-Model-Free Alignment. Unlike ordinal or cardinal feedback, linguistic feedback is high-dimensional, unstructured, and semantically rich, which makes it substantially more challenging to exploit effectively for learning. A prominent line of research (Scheurer et al., 2023; Bai et al., 2022; Chen et al., 2023,2; Mehri et al., 2025) addresses this challenge by leveraging linguistic feedback for data synthesis, where the original training data are augmented or refined into improved examples guided by the feedback, and the resulting synthesized data are used to fine-tune the model. To reduce reliance on expensive human annotations, a growing body of work (Shinn et al., 2023; Dhuliawala et al., 2024; Besta et al., 2024; Madaan et al., 2023) further explores using LLMs themselves to generate textual feedback via reflection and perform self-refinement.
Reward-Model-Based Alignment. Reward models leverage linguistic feedback in at least two complementary ways, distinguished by whether the critique is a target to imitate or a scaffold for the reward model’s own reasoning. The first treats critique generation as an auxiliary output that the reward model is trained to imitate: the model is trained on critiques that are generated at inference time as part of the judgment itself (Ankner et al., 2024), used as an additional self-supervised training signal (Yu et al., 2025), synthesized to augment scarce critique-annotated data (Ye et al., 2025), or, most directly, produced by casting reward assignment as next-token generation over a critique (Zhang et al., 2024); at inference time, these models first generate an explicit critique before producing the final judgment. The second treats reward assignment itself as a reasoning task, allocating additional training- or test-time computation to deliberate before scoring: examples include reward models fine-tuned to reason step-by-step before a judgment (Chen et al., 2025), models trained to plan an evaluation strategy before acting as a judge (Saha et al., 2025), reward models trained with reinforcement learning to reason toward a score (Guo et al., 2025), and approaches that scale inference-time computation for a generalist reward model (Liu et al., 2025). Compared to the critique-imitation paradigm, this second line of work uses linguistic feedback less as a supervised target and more as a scaffold that shapes how the reward model reasons.
4. Evaluation of Aligned LLMs
As illustrated in the right part of Figure 5, feedback types and evaluation metrics have a many-to-many relationship. Since the goal is to evaluate alignment rather than feedback itself, organizing evaluation methods by feedback type can be ambiguous.
Motivated by these considerations, we categorize evaluation methods according to the type of alignment evidence they rely on, as shown in Figure 5. Direct verification evaluates model outputs against ground truths or other objectively verifiable criteria, whereas proxy verification relies on models or metrics derived from human feedback to approximate human preferences.
Figure 5.
Direct and proxy verification metrics and their many-to-many mapping to feedback types.

4.1. Direct Verification
Direct verification assumes the existence of a verifiable checker that maps a prompt–response pair to an explicit signal, such as factual consistency with task rules, agreement with labels, or successful execution of generated code. These signals allow alignment to be assessed against objective criteria. This perspective aligns with Reinforcement Learning from Verifiable Rewards (RLVR) (Wen et al., 2025), where the same checker is used as the reward function during training.
Factuality–Based. Factuality-based methods assess whether outputs align with verified facts, typically via accuracy against trusted references or deterministic solutions (e.g., QA, math) (Lightman et al., 2024; Fan et al., 2024; Lai et al., 2024), or objective task-defined rewards (Shani et al., 2024; Wang et al., 2025).
Classification-Based. When objectives are categorical, such as safety or stylistic constraints, standard classification metrics (F1, precision, recall, NDCG) (Ouyang et al., 2022; Mu et al., 2024) quantify consistency with human annotations.
Executable Verification. For objectively testable responses, such as code generation, correctness is validated via execution or test cases, using metrics like Pass@k or pass rate (Chidambaram et al., 2024; Wong and Tan 2024; Li et al., 2025).
Reliability of Direct Verification. Although direct verification grounds evaluation in objective criteria, verifier scores are not always faithful proxies for the intended objective. First, verifiers can be gamed: models may exploit imperfections in the checker or reward function rather than solve the intended task, a failure mode closely related to reward hacking and specification gaming (Helff et al., 2026). Second, when the checker is implemented as an LLM-as-a-judge rather than a fully objective procedure, its judgments may reflect systematic biases or self-preference rather than true task quality (Chen et al., 2024; Ye et al., 2025). Third, in RLVR, higher verifier scores need not correspond to better reasoning: they may arise from spurious rewards that are unrelated to, or even negatively correlated with, correctness (Shao et al., 2026), or from amplifying reasoning paths already present in the base model rather than inducing genuinely new capability (Yue et al., 2025). Verifier scores can therefore saturate or become misleading proxies for the objective they are meant to measure.
4.2. Proxy Verification
When explicit evaluation signals are unavailable, alignment evaluation relies on proxy feedback that approximates human preferences. Proxy verification therefore uses learned evaluators or automated metrics to assess alignment. We refer to this setting as Reinforcement Learning with Unverifiable Rewards (RLUR), a notion introduced in this work.
Reward Model-Based. These methods use trained reward models to score outputs or assess correlation with human preferences, via accuracy or F1 (Wang et al., 2024; Liu et al., 2025), NDCG (Wang et al., 2025), cross-entropy (Liu et al., 2025), or the reward model score itself (Yuan et al., 2023).
Attribute-Based. Attribute-based evaluators use pretrained classifiers, e.g., toxicity or harmfulness detectors, as proxy judges for particular dimensions, such as the safety-focused Perspective API (Wu et al., 2023; Ouyang et al., 2022; Li et al., 2025).
Text-Based. Text-based methods measure similarity to reference responses: lexical metrics such as BLEU Papineni et al. (2002) and ROUGE Lin (2004), and semantic metrics such as BERTScore Zhang et al. (2020) and COMET Williamson et al. (2017), are widely used for fully automatic evaluation of tasks such as summarization and translation (Wu et al., 2023; Stiennon et al., 2020; Ramos et al., 2024; Song et al., 2025).
LLM-as-a-Judge. LLM-as-a-judge uses a large LLM itself as the evaluator, supporting various formats: ranking-based judgments, including pairwise win rates (Shi et al., 2024; Jin et al., 2023; Tucker et al., 2024; Shaikh et al., 2025) and listwise rankings (Yang et al., 2024); scalar ratings, such as attribute-specific scores (Bai et al., 2023; Zhu et al., 2024), preference similarity (Aroca-Ouellette et al., 2025), or numeric scales (Yu et al., 2025; Cui et al., 2023; Liu et al., 2023; Hashemi et al., 2024); and textual critiques (Li et al., 2024; Xie et al., 2024).
Reward Models versus LLM-as-a-Judge. Although both serve as proxies for human preference, the two differ in operational form: a reward model is a trained scoring function that outputs a scalar or ranking directly, while an LLM-as-a-judge relies on prompted evaluation and can additionally produce a natural-language rationale alongside its judgment. They also differ in stability and reproducibility: LLM judges are comparatively more sensitive to prompt design and to the underlying model version, whereas a trained reward model is typically more standardized once fixed.
5. Discussion
5.1. Mismatches and Gaps
The feedback-centric organization adopted throughout this survey surfaces mismatches between data collection, optimization, and evaluation (Section 1) that a component-centric organization would leave implicit. On the training side, rich feedback such as critiques or process-level supervision is often reduced to scalar rewards or pairwise preferences, losing information it originally conveyed, and multi-aspect signals are often aggregated into a single reward, obscuring trade-offs across dimensions such as helpfulness and factuality. On the evaluation side, generic metrics such as win rates or LLM judges often miss improvements along the specific dimensions models were trained on—e.g., process-level supervision evaluated only by outcome-level accuracy leaves reasoning gains unmeasured—and feedback is growing more complex, with preferences varying across contexts (Kim et al., 2025) and rating scales obscuring demographic differences (Ali et al., 2025).
Two directions remain underexplored: principled task–feedback design—which supervision form suits a given deployment and annotation budget—and extending feedback taxonomies and objectives to multimodal models.
5.2. Conclusions
This survey advocates a feedback-centric perspective on LLM alignment, treating the structure of feedback as a primary lens for organizing alignment methods and evaluation, rather than exhaustively cataloging algorithms or metrics. We hope this perspective helps re-center alignment around the structure and semantics of feedback, and supports next-generation alignment pipelines in which data collection, optimization, and evaluation are designed jointly rather than in isolation.
Limitations
This survey has several limitations. First, the rapidly evolving literature means recent work may be omitted, so our taxonomy is best viewed as a snapshot rather than an exhaustive catalog. Second, our feedback-centric taxonomy abstracts away task- and domain-specific nuances, e.g., in safety-critical or low-resource settings. Third, our discussion primarily reflects English-language LLMs; applicability to multilingual settings requires further consideration. Finally, the mismatches and underexplored regimes we identify are based on qualitative synthesis and should be validated empirically.
Appendix A. Survey Methodology and Scope
We build this survey from prior surveys of LLM alignment and human feedback, followed by a targeted, iterative literature search that combines keyword-based search, citation tracing, and AI-assisted literature discovery tools, with all candidate works manually verified before inclusion. Given the scale and pace of this literature, we prioritize work that is (i) published at top-tier conferences or journals, (ii) produced by well-established research or industrial institutions, or (iii) a widely recognized preprint, as reflected by citation counts. Because the literature we survey overlaps substantially with that of prior surveys, our primary contribution is not exhaustive coverage of this literature, but the feedback-centric taxonomy and organization presented in this survey.
Appendix B. Synthesis of Alignment Methods by Feedback Type
For each feedback type introduced in Section 3, we compare reward-model-free and reward-model-based methods in terms of their assumptions about the feedback signal, the settings they are best suited to, and the information they preserve and exploit.
Ordinal Feedback. Reward-model-free and reward-model-based methods differ in how they handle noisy preference signals: the former optimize directly from offline comparisons, while the latter interpose a reward model to smooth over noise, at the cost of propagating reward-model error into the policy. Both are limited to relative-order information—neither recovers preference strength—and listwise supervision only extends this to more candidates, not to cardinal information.
Cardinal Feedback. Because cardinal feedback specifies magnitude rather than only order, these methods exploit information ordinal methods discard—KTO weights generations asymmetrically around a reference point, while multi-objective reward models keep per-attribute scores rather than one scalar—making them best suited to explicit multi-attribute objectives, at the cost of greater sensitivity to inconsistent annotator scale usage.
Linguistic Feedback. Linguistic-feedback methods differ chiefly in what they assume about the model: reward-model-free approaches assume the policy itself can interpret free-form critiques, while reward-model-based approaches assume a reward model can learn to imitate and score them. Both can in principle preserve attribute-specific information that ordinal and cardinal feedback compress away, but only to the extent that the model reliably parses and grounds natural-language guidance—still an open challenge.
Appendix C. Meta Analysis: Experimental Details
This appendix provides full implementation details for the quantitative feedback analysis reported in Figure 1. All experiments use the HelpSteer3 preference dataset Wang et al. (2025), which provides, for each comparison item, multi-annotator binary preferences P, scalar quality scores S, and free-form linguistic feedback for both responses. Beyond quantifying the incremental information carried by each feedback channel (Appendix C.3), we use the same dataset to probe practical considerations relevant to deploying feedback at scale: how well feedback predicts score magnitude (Appendix C.4), how much of annotator disagreement, i.e., feedback reliability, it can explain (Appendix C.5), and whether it points in an actionable direction for improvement (Appendix C.6).
Appendix C.1. Notation and Data
Let denote the annotator-level dataset, where t indexes individual annotations, is the binary preference sign, is the scalar score, are the linguistic feedback texts for response 1 and response 2 respectively, and is the aggregated majority preference for the corresponding item, which serves as the prediction target throughout. Each annotation carries an item identifier mapping it to the underlying comparison pair.
- Embedding representations.
We encode all linguistic feedback using a frozen Sentence-BERT model (all-MiniLM-L6-v2; Reimers and Gurevych 2019) to obtain dense representations. Specifically, we embed the reasoning text to obtain and the two per-response feedback texts to obtain , where . Response texts are similarly encoded: for item i.
Appendix C.2. Unified Mutual Information Estimation
A central design principle of our analysis is to estimate all mutual information (MI) quantities using the same methodology, so that the resulting numbers are directly comparable across feedback channels.
- Classifier-probe framework.
We estimate MI through the cross-entropy gap of a logistic regression probe. For a generic feature matrix and target Y:
where is the plug-in marginal entropy (in bits) and is the cross-entropy (in bits) of a logistic regression classifier trained on and evaluated on held-out data. For conditional MI we use:
This framework is applied uniformly to all channels: the binary preference P (encoded as a one-dimensional feature), the scalar score S (concatenated with P), and the linguistic feedback embedding (concatenated with P and S). Using the same estimator for every channel ensures that any systematic bias (e.g., under-fitting of logistic regression on high-dimensional inputs) affects all estimates in the same direction, preserving ordinal comparisons.
- Cross-validation and uncertainty.
All cross-entropy estimates are computed via K-fold stratified cross-validation () to reduce variance from a single train/test split. We report the mean and standard deviation over folds. For conditional MI, the standard deviation is propagated as , treating fold-level errors as approximately independent.
Appendix C.3. Mutual Information
- : information in the binary preference sign alone. The feature matrix is .
- : incremental information from the score beyond the preference sign. We compare against .
- : incremental information from the linguistic feedback embedding beyond both P and S. We compare against .
Appendix C.4. Score Magnitude Prediction
To assess whether linguistic feedback carries information about how strongly an annotator prefers one response, we estimate using the same classifier-probe approach with as a multi-class target. This measures the extent to which reasoning text predicts score magnitude beyond what the preference direction already reveals.
- Permutation null.
We randomly permute the reasoning embeddings across annotators, breaking the correspondence between each annotator’s feedback and their score while preserving the marginal distribution of embeddings. The permuted conditional MI is expected to be near zero; the gap between the real and permuted estimates isolates genuine information content.
Appendix C.5. Disagreement Prediction
For each item i, we define annotator disagreement as the binary entropy of the preference distribution:
where is the set of annotators for item i, and when .
- Response-only features.
To avoid circularity—annotator-produced reasoning is generated jointly with the preference and thus reflects disagreement almost by construction—we predict from response-only features:
where are the Sentence-BERT embeddings of the two response texts. A Ridge regression () is trained to predict from , and we report averaged over 10-fold cross-validation.
This setup measures whether the content of the response pair itself is predictive of how much annotators will disagree, without leaking annotator-side information.
- Permutation null.
We randomly permute the response embeddings across items, breaking the correspondence between response content and annotator disagreement while preserving the marginal distribution. A negative permuted (as observed) confirms that the model does not extract structure from mismatched response–disagreement pairings.
Appendix C.6. Actionability
The actionability metric tests whether the direction of linguistic feedback (from rejected to preferred) aligns with the corresponding direction in response embedding space.
- Per-annotator alignment.
For annotator t, we define the signed feedback difference and signed response difference as:
Both vectors point in the “preferred − rejected” direction according to annotator t’s own decision. This per-annotator alignment avoids a directional mismatch that would arise if item-level majority labels were used to sign one vector while annotator-level labels were used to sign the other.
The actionability score is the average cosine similarity:
- Permutation null and effect size.
To assess significance, we construct a null distribution by randomly permuting the annotator indices, breaking the correspondence between feedback and response pairs while preserving marginal distributions. The permuted mean cosine similarity is expected to be near zero; a significant gap confirms that the alignment reflects genuine semantic correspondence rather than distributional artifacts. We report with a 95% bootstrap confidence interval (5,000 resamples) and Cohen’s d using the pooled standard deviation. Because cosine similarity in high-dimensional spaces concentrates near zero for unrelated vectors, the absolute value of A is expected to be small; Cohen’s d provides a scale-free effect size that is more appropriate for assessing practical significance.
Appendix C.7. Domain-Specific Analysis
HelpSteer3 Wang et al. (2025) provides data from multiple domains, which supports a domain-specific analysis. Across the Code, General, Multilingual, and STEM subsets, the results consistently show that language feedback contains signals beyond binary preferences. First, score magnitude retains substantial mutual information with language feedback in every subset (Table A1), indicating that language feedback encodes preference strength regardless of domain. Second, disagreement predictability varies by domain (Table A3): higher-ambiguity settings such as Multilingual show the strongest predictability (), while more structured domains such as Code show a weaker, though still positive, effect (); in every domain the real exceeds its permuted null, so the direction of the effect is consistent even as its magnitude varies. Third, actionability shows a stable positive alignment between feedback and response improvements across all domains, with a non-trivial effect size in every subset (Cohen’s d from 0.298 to 0.754; Table A2). Together, these domain-level results support the robustness of the patterns discussed in the main text.
Table A1.
Score Magnitude MI: .
| Subset | Real | Permuted Null |
|---|---|---|
| Code | 0.716 | -0.013 |
| General | 0.861 | -0.007 |
| Multilingual | 0.834 | -0.009 |
| STEM | 0.727 | -0.016 |
| Overall | 0.892 | -0.004 |
Table A2.
Actionability: .
| Subset | Real | Permuted Null | Cohen’s d |
|---|---|---|---|
| Code | 0.102 | 0.000 | 0.754 |
| General | 0.071 | 0.001 | 0.580 |
| Multilingual | 0.030 | 0.000 | 0.298 |
| STEM | 0.070 | -0.001 | 0.599 |
| Overall | 0.070 | 0.000 | 0.580 |
Table A3.
Disagreement Prediction.
| Subset | Score Variance | (real) | (permuted null) |
|---|---|---|---|
| Code | 0.166 | 1.5% | -9.6% |
| General | 0.198 | 0.6% | -5.9% |
| Multilingual | 0.351 | 10.4% | -10.8% |
| STEM | 0.205 | 4.3% | -18.9% |
| Overall | 0.222 | 10.4% | -2.8% |
References
- Wu, Z.; Hu, Y.; Shi, W.; Dziri, N.; Suhr, A.; Ammanabrolu, P.; Smith, N.A.; Ostendorf, M.; Hajishirzi, H. Fine-Grained Human Feedback Gives Better Rewards for Language Model Training. In Proceedings of the Advances in Neural Information Processing Systems; Oh, A.; Naumann, T.; Globerson, A.; Saenko, K.; Hardt, M.; Levine, S., Eds. Curran Associates, Inc., 2023, Vol. 36, pp. 59008–59033.
- Casper, S.; Davies, X.; Shi, C.; Gilbert, T.K.; Scheurer, J.; Rando, J.; Freedman, R.; Korbak, T.; Lindner, D.; Freire, P.; et al. Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback. Transactions on Machine Learning Research 2023. Survey Certification, Featured Certification.
- Chaudhari, S.; Aggarwal, P.; Murahari, V.; Rajpurohit, T.; Kalyan, A.; Narasimhan, K.; Deshpande, A.; Castro da Silva, B. RLHF Deciphered: A Critical Analysis of Reinforcement Learning from Human Feedback for LLMs. ACM Comput. Surv. 2025, 58. [CrossRef]
- Scheurer, J.; Campos, J.A.; Korbak, T.; Chan, J.S.; Chen, A.; Cho, K.; Perez, E. Training language models with language feedback at scale. arXiv preprint arXiv:2303.16755 2023.
- Yu, T.; Lin, T.E.; Wu, Y.; Yang, M.; Huang, F.; Li, Y. Diverse AI Feedback For Large Language Model Alignment. Transactions of the Association for Computational Linguistics 2025, 13, 392–407. [CrossRef]
- Uesato, J.; Kushman, N.; Kumar, R.; Song, F.; Siegel, N.; Wang, L.; Creswell, A.; Irving, G.; Higgins, I. Solving math word problems with process-and outcome-based feedback. arXiv preprint arXiv:2211.14275 2022.
- Lightman, H.; Kosaraju, V.; Burda, Y.; Edwards, H.; Baker, B.; Lee, T.; Leike, J.; Schulman, J.; Sutskever, I.; Cobbe, K. Let’s Verify Step by Step. In Proceedings of the The Twelfth International Conference on Learning Representations, 2024.
- Wang, Z.; Zeng, J.; Delalleau, O.; Shin, H.C.; Soares, F.; Bukharin, A.; Evans, E.; Dong, Y.; Kuchaiev, O. HelpSteer3-Preference: Open Human-Annotated Preference Data across Diverse Tasks and Languages, 2025, [arXiv:cs.CL/2505.11475].
- Fernandes, P.; Madaan, A.; Liu, E.; Farinhas, A.; Martins, P.H.; Bertsch, A.; de Souza, J.G.; Zhou, S.; Wu, T.; Neubig, G.; et al. Bridging the gap: A survey on integrating (human) feedback for natural language generation. Transactions of the Association for Computational Linguistics 2023, 11, 1643–1668.
- Jiang, R.; Chen, K.; Bai, X.; He, Z.; Li, J.; Yang, M.; Zhao, T.; Nie, L.; Zhang, M. A Survey on Human Preference Learning for Large Language Models, 2024, [arXiv:cs.CL/2406.11191].
- Liu, S.; Fang, W.; Hu, Z.; Zhang, J.; Zhou, Y.; Zhang, K.; Tu, R.; Lin, T.E.; Huang, F.; Song, M.; et al. A survey of direct preference optimization. arXiv preprint arXiv:2503.11701 2025.
- Ji, M.; Wu, Y.; Wu, Z.; Wang, S.; Yang, J.; Dras, M.; Naseem, U. A Survey on Progress in LLM Alignment from the Perspective of Reward Design, 2025, [arXiv:cs.CL/2505.02666].
- Lai, H.; Liu, X.; Gao, J.; Cheng, J.; Qi, Z.; Xu, Y.; Yao, S.; Zhang, D.; Du, J.; Hou, Z.; et al. A Survey of Post-Training Scaling in Large Language Models. In Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Che, W.; Nabende, J.; Shutova, E.; Pilehvar, M.T., Eds., Vienna, Austria, 2025; pp. 2771–2791. [CrossRef]
- Stiennon, N.; Ouyang, L.; Wu, J.; Ziegler, D.; Lowe, R.; Voss, C.; Radford, A.; Amodei, D.; Christiano, P.F. Learning to summarize with human feedback. Advances in neural information processing systems 2020, 33, 3008–3021.
- Bai, Y.; Jones, A.; Ndousse, K.; Askell, A.; Chen, A.; DasSarma, N.; Drain, D.; Fort, S.; Ganguli, D.; Henighan, T.; et al. Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862 2022.
- Cui, G.; Yuan, L.; Ding, N.; Yao, G.; Zhu, W.; Ni, Y.; Xie, G.; Liu, Z.; Sun, M. UltraFeedback: Boosting Language Models with High-quality Feedback, 2023, [arXiv:cs.CL/2310.01377].
- Köpf, A.; Kilcher, Y.; Von Rütte, D.; Anagnostidis, S.; Tam, Z.R.; Stevens, K.; Barhoum, A.; Nguyen, D.; Stanley, O.; Nagyfi, R.; et al. OpenAssistant conversations-democratizing large language model alignment. Advances in neural information processing systems 2023, 36, 47669–47681.
- Malik, S.; Pyatkin, V.; Land, S.; Morrison, J.; Smith, N.A.; Hajishirzi, H.; Lambert, N. RewardBench 2: Advancing Reward Model Evaluation, 2025, [arXiv:cs.CL/2506.01937].
- Nakano, R.; Hilton, J.; Balaji, S.; Wu, J.; Ouyang, L.; Kim, C.; Hesse, C.; Jain, S.; Kosaraju, V.; Saunders, W.; et al. WebGPT: Browser-assisted question-answering with human feedback. arXiv preprint arXiv:2112.09332 2021.
- Miranda, L.J.V.; Wang, Y.; Elazar, Y.; Kumar, S.; Pyatkin, V.; Brahman, F.; Smith, N.A.; Hajishirzi, H.; Dasigi, P. Hybrid Preferences: Learning to Route Instances for Human vs. AI Feedback. arXiv 2024, abs/2410.19133.
- Zheng, L.; Chiang, W.L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E.P.; et al. Judging LLM-as-a-judge with MT-Bench and Chatbot Arena, 2023, [arXiv:cs.CL/2306.05685].
- Wang, Z.; Dong, Y.; Delalleau, O.; Zeng, J.; Shen, G.; Egert, D.; Zhang, J.J.; Sreedhar, M.N.; Kuchaiev, O. HelpSteer2: Open-source dataset for training top-performing reward models. In Proceedings of the Proceedings of the 38th International Conference on Neural Information Processing Systems, 2024, pp. 1474–1501.
- Cai, T.; Song, X.; Jiang, J.; Teng, F.; Gu, J.; Zhang, G. ULMA: Unified language model alignment with human demonstration and point-wise preference. arXiv preprint arXiv:2312.02554 2023.
- Mishra, P.; Yao, Z.; Vashisht, P.; Ouyang, F.; Wang, B.; Mody, V.D.; Yu, H. SYNFAC-EDIT: Synthetic Imitation Edit Feedback for Factual Alignment in Clinical Summarization. In Proceedings of the Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing; Al-Onaizan, Y.; Bansal, M.; Chen, Y.N., Eds., Miami, Florida, USA, 2024; pp. 20061–20083. [CrossRef]
- Majkutewicz, J.; Szymański, J. Aligning large language models with human preferences using historical text edits. Knowledge-Based Systems 2025, 322, 113566. [CrossRef]
- Wang, Z.; Zeng, J.; Delalleau, O.; Egert, D.; Evans, E.; Shin, H.C.; Soares, F.; Dong, Y.; Kuchaiev, O. HelpSteer3: Human-Annotated Feedback and Edit Data to Empower Inference-Time Scaling in Open-Ended General-Domain Tasks, 2025, [arXiv:cs.CL/2503.04378].
- Tang, L.; Shalyminov, I.; Wong, A.; Burnsky, J.; Vincent, J.; Yang, Y.; Singh, S.; Feng, S.; Song, H.; Su, H.; et al. TofuEval: Evaluating hallucinations of LLMs on topic-focused dialogue summarization. In Proceedings of the NAACL 2024, 2024.
- Yao, Z.; Schloss, B.; Selvaraj, S. Improving Summarization with Human Edits. In Proceedings of the Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 2023, pp. 2604–2620.
- Chiang, W.L.; Zheng, L.; Sheng, Y.; Angelopoulos, A.N.; Li, T.; Li, D.; Zhu, B.; Zhang, H.; Jordan, M.; Gonzalez, J.E.; et al. Chatbot Arena: An open platform for evaluating LLMs by human preference. In Proceedings of the Forty-first International Conference on Machine Learning, 2024.
- Bai, G.; Liu, J.; Bu, X.; He, Y.; Liu, J.; Zhou, Z.; Lin, Z.; Su, W.; Ge, T.; Zheng, B.; et al. MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues. In Proceedings of the Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 7421–7454.
- Shani, L.; Rosenberg, A.; Cassel, A.; Lang, O.; Calandriello, D.; Zipori, A.; Noga, H.; Keller, O.; Piot, B.; Szpektor, I.; et al. Multi-turn Reinforcement Learning with Preference Human Feedback. In Proceedings of the The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024.
- Gao, Z.; Zhan, W.; Chang, J.D.; Swamy, G.; Brantley, K.; Lee, J.D.; Sun, W. Regressing the relative future: Efficient policy optimization for multi-turn RLHF. arXiv preprint arXiv:2410.04612 2024.
- Shinn, N.; Cassano, F.; Gopinath, A.; Narasimhan, K.; Yao, S. Reflexion: Language agents with verbal reinforcement learning. Advances in Neural Information Processing Systems 2023, 36, 8634–8652.
- Shi, T.; Wang, Z.; Yang, L.; Lin, Y.C.; He, Z.; Wan, M.; Zhou, P.; Jauhar, S.K.; Xu, X.; Song, X.; et al. WildFeedback: Aligning LLMs With In-situ User Interactions And Feedback. In Proceedings of the NeurIPS 2024 Workshop on Behavioral Machine Learning, 2024.
- Liu, Y.; Zhang, M.J.; Choi, E. User Feedback in Human-LLM Dialogues: A Lens to Understand Users But Noisy as a Learning Signal. In Proceedings of the Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing; Christodoulopoulos, C.; Chakraborty, T.; Rose, C.; Peng, V., Eds., Suzhou, China, 2025; pp. 2666–2681.
- Bai, Y.; Kadavath, S.; Kundu, S.; Askell, A.; Kernion, J.; Jones, A.; Chen, A.; Goldie, A.; Mirhoseini, A.; McKinnon, C.; et al. Constitutional AI: Harmlessness from AI feedback. arXiv preprint arXiv:2212.08073 2022.
- Barreto, A.; Dumoulin, V.; Mao, Y.; Rowland, M.; Perez-Nieves, N.; Shahriari, B.; Dauphin, Y.; Precup, D.; Larochelle, H. Capturing Individual Human Preferences with Reward Features. In Proceedings of the Advances in Neural Information Processing Systems (NeurIPS), 2025.
- Liu, J.; Xia, C.S.; Wang, Y.; Zhang, L. Is your code generated by ChatGPT really correct? rigorous evaluation of large language models for code generation. Advances in Neural Information Processing Systems 2023, 36, 21558–21572.
- Liu, J.; Xie, S.; Wang, J.; Wei, Y.; Ding, Y.; Zhang, L. Evaluating language models for efficient code generation. arXiv preprint arXiv:2408.06450 2024.
- Ahmad, W.U.; Ficek, A.; Samadi, M.; Huang, J.; Noroozi, V.; Majumdar, S.; Ginsburg, B. OpenCodeInstruct: A Large-scale Instruction Tuning Dataset for Code LLMs, 2025, [arXiv:cs.CL/2504.04030].
- Dai, D.; Liu, M.; Li, A.; Cao, J.; Wang, Y.; Wang, C.; Peng, X.; Zheng, Z. FeedbackEval: A Benchmark for Evaluating Large Language Models in Feedback-Driven Code Repair Tasks. arXiv preprint arXiv:2504.06939 2025.
- Jacovi, A.; Bitton, Y.; Bohnet, B.; Herzig, J.; Honovich, O.; Tseng, M.; Collins, M.; Aharoni, R.; Geva, M. A Chain-of-Thought Is as Strong as Its Weakest Link: A Benchmark for Verifiers of Reasoning Chains. In Proceedings of the Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers); Ku, L.W.; Martins, A.; Srikumar, V., Eds., Bangkok, Thailand, 2024; pp. 4615–4634. [CrossRef]
- Song, M.; Su, Z.; Qu, X.; Zhou, J.; Cheng, Y. PRMBench: A fine-grained and challenging benchmark for process-level reward models. arXiv preprint arXiv:2501.03124 2025.
- Wang, P.; Li, L.; Shao, Z.; Xu, R.; Dai, D.; Li, Y.; Chen, D.; Wu, Y.; Sui, Z. Math-shepherd: Verify and reinforce LLMs step-by-step without human annotations. In Proceedings of the Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 9426–9439.
- Karthik, S.; Coskun, H.; Akata, Z.; Tulyakov, S.; Ren, J.; Kag, A. Scalable ranked preference optimization for text-to-image generation. In Proceedings of the Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025, pp. 18399–18410.
- Yuan, H.; Chen, Z.; Ji, K.; Gu, Q. Self-play fine-tuning of diffusion models for text-to-image generation. Advances in Neural Information Processing Systems 2024, 37, 73366–73398.
- Rafailov, R.; Sharma, A.; Mitchell, E.; Manning, C.D.; Ermon, S.; Finn, C. Direct preference optimization: Your language model is secretly a reward model. Advances in neural information processing systems 2023, 36, 53728–53741.
- Zhao, Y.; Joshi, R.; Liu, T.; Khalman, M.; Saleh, M.; Liu, P.J. SLiC-HF: Sequence likelihood calibration with human feedback. arXiv preprint arXiv:2305.10425 2023.
- Liu, T.; Zhao, Y.; Joshi, R.; Khalman, M.; Saleh, M.; Liu, P.J.; Liu, J. Statistical rejection sampling improves preference optimization. arXiv preprint arXiv:2309.06657 2023.
- Azar, M.G.; Guo, Z.D.; Piot, B.; Munos, R.; Rowland, M.; Valko, M.; Calandriello, D. A general theoretical paradigm to understand learning from human preferences. In Proceedings of the International Conference on Artificial Intelligence and Statistics. PMLR, 2024, pp. 4447–4455.
- Chen, Z.; Deng, Y.; Yuan, H.; Ji, K.; Gu, Q. Self-play fine-tuning converts weak language models to strong language models. arXiv preprint arXiv:2401.01335 2024.
- Meng, Y.; Xia, M.; Chen, D. SimPO: Simple preference optimization with a reference-free reward. Advances in Neural Information Processing Systems 2024, 37, 124198–124235.
- Wu, J.; Xie, Y.; Yang, Z.; Wu, J.; Gao, J.; Ding, B.; Wang, X.; He, X. β-DPO: Direct Preference Optimization with Dynamic β. Advances in Neural Information Processing Systems 2024, 37, 129944–129966.
- Xiao, T.; Yuan, Y.; Zhu, H.; Li, M.; Honavar, V.G. Cal-DPO: Calibrated direct preference optimization for language model alignment. Advances in Neural Information Processing Systems 2024, 37, 114289–114320.
- Liu, Z.; Lu, M.; Zhang, S.; Liu, B.; Guo, H.; Yang, Y.; Blanchet, J.; Wang, Z. Provably mitigating overoptimization in RLHF: Your SFT loss is implicitly an adversarial regularizer. Advances in Neural Information Processing Systems 2024, 37, 138663–138697.
- Fisch, A.; Eisenstein, J.; Zayats, V.; Agarwal, A.; Beirami, A.; Nagpal, C.; Shaw, P.; Berant, J. Robust preference optimization through reward model distillation. arXiv preprint arXiv:2405.19316 2024.
- Zhu, M.; Liu, Y.; Zhang, L.; Guo, J.; Mao, Z. LIRE: listwise reward enhancement for preference alignment. arXiv preprint arXiv:2405.13516 2024.
- Deng, X.; Zhong, H.; Ai, R.; Feng, F.; Wang, Z.; He, X. Less is more: Improving LLM alignment via preference data selection. arXiv preprint arXiv:2502.14560 2025.
- Gao, C.; Chen, R.; Yuan, S.; Huang, K.; Yu, Y.; He, X. SPRec: Self-play to debias LLM-based recommendation. In Proceedings of the Proceedings of the ACM on Web Conference 2025, 2025, pp. 5075–5084.
- Wang, Y.; Chen, Q.G.; Xu, Z.; Luo, W.; Zhang, K.; Zhang, L. SPACE: Noise contrastive estimation stabilizes self-play fine-tuning for large language models. arXiv preprint arXiv:2512.07175 2025.
- Bradley, R.A.; Terry, M.E. Rank analysis of incomplete block designs: I. the method of paired comparisons. Biometrika 1952, 39, 324–345.
- Liu, C.Y.; Zeng, L.; Liu, J.; Yan, R.; He, J.; Wang, C.; Yan, S.; Liu, Y.; Zhou, Y. Skywork-Reward: Bag of tricks for reward modeling in LLMs. arXiv preprint arXiv:2410.18451 2024.
- Dong, H.; Xiong, W.; Pang, B.; Wang, H.; Zhao, H.; Zhou, Y.; Jiang, N.; Sahoo, D.; Xiong, C.; Zhang, T. RLHF Workflow: From Reward Modeling to Online RLHF A Comprehensive Practical Alignment Recipe of Iterative Preference Learning. Transactions on Machine Learning Research 2024, 2024.
- Xiong, W.; Dong, H.; Ye, C.; Wang, Z.; Zhong, H.; Ji, H.; Jiang, N.; Zhang, T. Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-constraint. In Proceedings of the International Conference on Machine Learning. PMLR, 2024, pp. 54715–54754.
- Zhang, S.; Yu, D.; Sharma, H.; Zhong, H.; Liu, Z.; Yang, Z.; Wang, S.; Hassan, H.; Wang, Z. Self-exploring language models: Active preference elicitation for online alignment. arXiv preprint arXiv:2405.19332 2024.
- Rosset, C.; Cheng, C.A.; Mitra, A.; Santacroce, M.; Awadallah, A.; Xie, T. Direct Nash optimization: Teaching language models to self-improve with general preferences. arXiv preprint arXiv:2404.03715 2024.
- Xie, T.; Foster, D.J.; Krishnamurthy, A.; Rosset, C.; Awadallah, A.; Rakhlin, A. Exploratory preference optimization: Harnessing implicit q*-approximation for sample-efficient RLHF. arXiv preprint arXiv:2405.21046 2024.
- Gupta, T.; Madhavan, R.; Zhang, X.; Bansal, C.; Rajmohan, S. AMPO: Active Multi Preference Optimization for Self-play Preference Selection. In Proceedings of the Forty-second International Conference on Machine Learning, 2025.
- Ethayarajh, K.; Xu, W.; Muennighoff, N.; Jurafsky, D.; Kiela, D. KTO: Model alignment as prospect theoretic optimization. arXiv preprint arXiv:2402.01306 2024.
- Rame, A.; Couairon, G.; Dancette, C.; Gaya, J.B.; Shukor, M.; Soulier, L.; Cord, M. Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards. Advances in Neural Information Processing Systems 2023, 36, 71095–71134.
- Zhou, Z.; Liu, J.; Shao, J.; Yue, X.; Yang, C.; Ouyang, W.; Qiao, Y. Beyond one-preference-fits-all alignment: Multi-objective direct preference optimization. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2024, 2024, pp. 10586–10613.
- Guo, Y.; Cui, G.; Yuan, L.; Ding, N.; Sun, Z.; Sun, B.; Chen, H.; Xie, R.; Zhou, J.; Lin, Y.; et al. Controllable preference optimization: Toward controllable multi-objective alignment. arXiv preprint arXiv:2402.19085 2024.
- Yang, R.; Pan, X.; Luo, F.; Qiu, S.; Zhong, H.; Yu, D.; Chen, J. Rewards-in-context: Multi-objective alignment of foundation models with dynamic preference adjustment. arXiv preprint arXiv:2402.10207 2024.
- Zhong, Y.; Ma, C.; Zhang, X.; Yang, Z.; Chen, H.; Zhang, Q.; Qi, S.; Yang, Y. Panacea: Pareto alignment via preference adaptation for LLMs. Advances in Neural Information Processing Systems 2024, 37, 75522–75558.
- Sun, H.; Shen, Y.; Ton, J.F. Rethinking Bradley-Terry models in preference-based reward modeling: Foundations, theory, and alternatives. arXiv preprint arXiv:2411.04991 2024.
- Wang, H.; Xiong, W.; Xie, T.; Zhao, H.; Zhang, T. Interpretable Preferences via Multi-Objective Reward Modeling and Mixture-of-Experts. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2024; Al-Onaizan, Y.; Bansal, M.; Chen, Y.N., Eds., Miami, Florida, USA, 2024; pp. 10582–10592. [CrossRef]
- Adler, B.; Agarwal, N.; Aithal, A.; Anh, D.H.; Bhattacharya, P.; Brundyn, A.; Casper, J.; Catanzaro, B.; Clay, S.; Cohen, J.; et al. Nemotron-4 340B technical report. arXiv preprint arXiv:2406.11704 2024.
- Wang, H.; Lin, Y.; Xiong, W.; Yang, R.; Diao, S.; Qiu, S.; Zhao, H.; Zhang, T. Arithmetic control of LLMs for diverse user preferences: Directional preference alignment with multi-objective rewards. arXiv preprint arXiv:2402.18571 2024.
- Gao, Z.; Chang, J.; Zhan, W.; Oertell, O.; Swamy, G.; Brantley, K.; Joachims, T.; Bagnell, D.; Lee, J.D.; Sun, W. REBEL: Reinforcement learning via regressing relative rewards. Advances in Neural Information Processing Systems 2024, 37, 52354–52400.
- Lin, B.; Jiang, W.; Xu, Y.; Chen, H.; Chen, Y.C. PARM: Multi-objective test-time alignment via preference-aware autoregressive reward model. arXiv preprint arXiv:2505.06274 2025.
- Chen, A.; Scheurer, J.; Korbak, T.; Campos, J.A.; Chan, J.S.; Bowman, S.R.; Cho, K.; Perez, E. Improving code generation by training with natural language feedback. arXiv preprint arXiv:2303.16749 2023.
- Dhuliawala, S.; Komeili, M.; Xu, J.; Raileanu, R.; Li, X.; Celikyilmaz, A.; Weston, J. Chain-of-verification reduces hallucination in large language models. In Proceedings of the Findings of the association for computational linguistics: ACL 2024, 2024, pp. 3563–3578.
- Besta, M.; Blach, N.; Kubicek, A.; Gerstenberger, R.; Podstawski, M.; Gianinazzi, L.; Gajda, J.; Lehmann, T.; Niewiadomski, H.; Nyczyk, P.; et al. Graph of thoughts: Solving elaborate problems with large language models. In Proceedings of the Proceedings of the AAAI conference on artificial intelligence, 2024, Vol. 38, pp. 17682–17690.
- Chen, X.; Wen, H.; Nag, S.; Luo, C.; Yin, Q.; Li, R.; Li, Z.; Wang, W. IterAlign: Iterative constitutional alignment of large language models. arXiv preprint arXiv:2403.18341 2024.
- Mehri, S.; Chen, X.; Ji, H.; Hakkani-Tür, D. Beyond Sample-Level Feedback: Using Reference-Level Feedback to Guide Data Synthesis. arXiv preprint arXiv:2502.04511 2025.
- Madaan, A.; Tandon, N.; Gupta, P.; Hallinan, S.; Gao, L.; Wiegreffe, S.; Alon, U.; Dziri, N.; Prabhumoye, S.; Yang, Y.; et al. Self-Refine: Iterative refinement with self-feedback. Advances in Neural Information Processing Systems 2023, 36, 46534–46594.
- Ankner, Z.; Paul, M.; Cui, B.; Chang, J.D.; Ammanabrolu, P. Critique-out-loud reward models. arXiv preprint arXiv:2408.11791 2024.
- Zhang, L.; Hosseini, A.; Bansal, H.; Kazemi, M.; Kumar, A.; Agarwal, R. Generative verifiers: Reward modeling as next-token prediction. arXiv preprint arXiv:2408.15240 2024.
- Yu, Y.; Chen, Z.; Zhang, A.; Tan, L.; Zhu, C.; Pang, R.Y.; Qian, Y.; Wang, X.; Gururangan, S.; Zhang, C.; et al. Self-generated critiques boost reward modeling for language models. In Proceedings of the Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2025, pp. 11499–11514.
- Ye, Z.; Greenlee, F.D.; Bartolo, M.; Blunsom, P.; Campos, J.A.; Gallé, M. Improving reward models with synthetic critiques. In Proceedings of the Findings of the Association for Computational Linguistics: NAACL 2025, 2025, pp. 4506–4520.
- Liu, Z.; Wang, P.; Xu, R.; Ma, S.; Ruan, C.; Li, P.; Liu, Y.; Wu, Y. Inference-time scaling for generalist reward modeling. arXiv preprint arXiv:2504.02495 2025.
- Chen, X.; Li, G.; Wang, Z.; Jin, B.; Qian, C.; Wang, Y.; Wang, H.; Zhang, Y.; Zhang, D.; Zhang, T.; et al. RM-R1: Reward modeling as reasoning. arXiv preprint arXiv:2505.02387 2025.
- Saha, S.; Li, X.; Ghazvininejad, M.; Weston, J.; Wang, T. Learning to plan & reason for evaluation with thinking-LLM-as-a-judge. arXiv preprint arXiv:2501.18099 2025.
- Guo, J.; Chi, Z.; Dong, L.; Dong, Q.; Wu, X.; Huang, S.; Wei, F. Reward reasoning model. arXiv preprint arXiv:2505.14674 2025.
- Williams, R.J. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine learning 1992, 8, 229–256.
- Schulman, J.; Wolski, F.; Dhariwal, P.; Radford, A.; Klimov, O. Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347 2017.
- Ahmadian, A.; Cremer, C.; Gallé, M.; Fadaee, M.; Kreutzer, J.; Pietquin, O.; Üstün, A.; Hooker, S. Back to basics: Revisiting reinforce style optimization for learning from human feedback in LLMs. arXiv preprint arXiv:2402.14740 2024.
- Razin, N.; Wang, Z.; Strauss, H.; Wei, S.; Lee, J.D.; Arora, S. What Makes a Reward Model a Good Teacher? An Optimization Perspective. In Proceedings of the The Thirty-ninth Annual Conference on Neural Information Processing Systems.
- Song, F.; Yu, B.; Li, M.; Yu, H.; Huang, F.; Li, Y.; Wang, H. Preference ranking optimization for human alignment. In Proceedings of the Proceedings of the AAAI Conference on Artificial Intelligence, 2024, Vol. 38, pp. 18990–18998.
- Liu, T.; Qin, Z.; Wu, J.; Shen, J.; Khalman, M.; Joshi, R.; Zhao, Y.; Saleh, M.; Baumgartner, S.; Liu, J.; et al. LiPO: Listwise preference optimization through learning-to-rank. In Proceedings of the Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2025, pp. 2404–2420.
- Jin, B.; Yoon, J.; Qin, Z.; Wang, Z.; Xiong, W.; Meng, Y.; Han, J.; Arik, S.O. LLM alignment as retriever optimization: An information retrieval perspective. arXiv preprint arXiv:2502.03699 2025.
- Cao, Z.; Qin, T.; Liu, T.Y.; Tsai, M.F.; Li, H. Learning to rank: from pairwise approach to listwise approach. In Proceedings of the Proceedings of the 24th international conference on Machine learning, 2007, pp. 129–136.
- Zhu, B.; Jordan, M.; Jiao, J. Principled reinforcement learning with human feedback from pairwise or K-wise comparisons. In Proceedings of the International Conference on Machine Learning. PMLR, 2023, pp. 43037–43067.
- Fan, R.Z.; Li, X.; Zou, H.; Li, J.; He, S.; Chern, E.; Hu, J.; Liu, P. Reformatted Alignment. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2024; Al-Onaizan, Y.; Bansal, M.; Chen, Y.N., Eds., Miami, Florida, USA, 2024; pp. 574–597. [CrossRef]
- Lai, X.; Tian, Z.; Chen, Y.; Yang, S.; Peng, X.; Jia, J. Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs. CoRR 2024, abs/2406.18629.
- Wang, J.; Liu, Y.; Sun, Y.; Ma, X.; Wang, Y.; Ma, H.; Su, Z.; Chen, M.; Gao, M.; Dalal, O.; et al. User Feedback Alignment for LLM-powered Exploration in Large-scale Recommendation Systems. In Proceedings of the Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 6: Industry Track); Rehm, G.; Li, Y., Eds., Vienna, Austria, 2025; pp. 996–1003. [CrossRef]
- Ouyang, L.; Wu, J.; Jiang, X.; Almeida, D.; Wainwright, C.; Mishkin, P.; Zhang, C.; Agarwal, S.; Slama, K.; Ray, A.; et al. Training language models to follow instructions with human feedback. Advances in neural information processing systems 2022, 35, 27730–27744.
- Mu, T.; Helyar, A.; Heidecke, J.; Achiam, J.; Vallone, A.; Kivlichan, I.; Lin, M.; Beutel, A.; Schulman, J.; Weng, L. Rule based rewards for language model safety. Advances in Neural Information Processing Systems 2024, 37, 108877–108901.
- Chidambaram, S.; Li, L.E.; Bai, M.; Li, X.; Lin, K.; Zhou, X.; Williams, A.C. Socratic human feedback (SoHF): Expert steering strategies for LLM code generation. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2024, 2024, pp. 15491–15502.
- Wong, M.F.; Tan, C.W. Aligning crowd-sourced human feedback for reinforcement learning on code generation by large language models. IEEE Transactions on Big Data 2024.
- Li, Q.; Dai, X.; Li, X.; Zhang, W.; Wang, Y.; Tang, R.; Yu, Y. CodePRM: Execution Feedback-enhanced Process Reward Model for Code Generation. In Proceedings of the Findings of the Association for Computational Linguistics: ACL 2025; Che, W.; Nabende, J.; Shutova, E.; Pilehvar, M.T., Eds., Vienna, Austria, 2025; pp. 8169–8182. [CrossRef]
- Liu, S.; Pan, Y.; Chen, G.; Li, X. Reward Modeling with Ordinal Feedback: Wisdom of the Crowd. In Proceedings of the Forty-second International Conference on Machine Learning, 2025.
- Yuan, H.; Yuan, Z.; Tan, C.; Wang, W.; Huang, S.; Huang, F. RRHF: Rank responses to align language models with human feedback. Advances in Neural Information Processing Systems 2023, 36, 10935–10950.
- Li, A.J.; Krishna, S.; Lakkaraju, H. More RLHF, More Trust? On The Impact of Preference Alignment On Trustworthiness. In Proceedings of the The Thirteenth International Conference on Learning Representations, 2025.
- Ramos, M.; Fernandes, P.; Farinhas, A.; Martins, A. Aligning Neural Machine Translation Models: Human Feedback in Training and Inference. In Proceedings of the Proceedings of the 25th Annual Conference of the European Association for Machine Translation (Volume 1); Scarton, C.; Prescott, C.; Bayliss, C.; Oakley, C.; Wright, J.; Wrigley, S.; Song, X.; Gow-Smith, E.; Bawden, R.; Sánchez-Cartagena, V.M.; et al., Eds., Sheffield, UK, 2024; pp. 258–274.
- Song, H.; Yun, T.; Lee, Y.; Oh, J.; Lee, G.; Cai, J.; Su, H. Learning to Summarize from LLM-generated Feedback. In Proceedings of the Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers); Chiruzzo, L.; Ritter, A.; Wang, L., Eds., Albuquerque, New Mexico, 2025; pp. 835–857. [CrossRef]
- Jin, D.; Mehri, S.; Hazarika, D.; Padmakumar, A.; Lee, S.; Liu, Y.; Namazifar, M. Data-Efficient Alignment of Large Language Models with Human Feedback Through Natural Language. In Proceedings of the NeurIPS 2023 Workshop on Instruction Tuning and Instruction Following, 2023.
- Tucker, A.D.; Brantley, K.; Cahall, A.; Joachims, T. Coactive learning for large language models using implicit user feedback. In Proceedings of the Forty-first International Conference on Machine Learning, 2024.
- Shaikh, O.; Lam, M.S.; Hejna, J.; Shao, Y.; Cho, H.J.; Bernstein, M.S.; Yang, D. Aligning Language Models with Demonstrated Feedback. In Proceedings of the The Thirteenth International Conference on Learning Representations, 2025.
- Yang, S.; Bi, K.; Cui, W.; Guo, J.; Cheng, X. LINKAGE: Listwise Ranking among Varied-Quality References for Non-Factoid QA Evaluation via LLMs. In Proceedings of the Findings of the Association for Computational Linguistics: EMNLP 2024; Al-Onaizan, Y.; Bansal, M.; Chen, Y.N., Eds., Miami, Florida, USA, 2024; pp. 6985–7000. [CrossRef]
- Bai, Y.; Ying, J.; Cao, Y.; Lv, X.; He, Y.; Wang, X.; Yu, J.; Zeng, K.; Xiao, Y.; Lyu, H.; et al. Benchmarking foundation models with language-model-as-an-examiner. Advances in Neural Information Processing Systems 2023, 36, 78142–78167.
- Zhu, B.; Frick, E.; Wu, T.; Zhu, H.; Ganesan, K.; Chiang, W.L.; Zhang, J.; Jiao, J. Starling-7B: Improving helpfulness and harmlessness with RLAIF. In Proceedings of the First Conference on Language Modeling, 2024.
- Aroca-Ouellette, S.; Mackraz, N.; Theobald, B.J.; Metcalf, K. Aligning LLMs by Predicting Preferences from User Writing Samples. In Proceedings of the Forty-second International Conference on Machine Learning, 2025.
- Liu, Y.; Iter, D.; Xu, Y.; Wang, S.; Xu, R.; Zhu, C. G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment. In Proceedings of the The 2023 Conference on Empirical Methods in Natural Language Processing, 2023.
- Hashemi, H.; Eisner, J.; Rosset, C.; Van Durme, B.; Kedzie, C. LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts. In Proceedings of the Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 13806–13834.
- Li, J.; Sun, S.; Yuan, W.; Fan, R.Z.; hai zhao.; Liu, P. Generative Judge for Evaluating Alignment. In Proceedings of the The Twelfth International Conference on Learning Representations, 2024.
- Xie, Y.; Zhou, W.; Prakash, P.; Jin, D.; Mao, Y.; Fettes, Q.; Talebzadeh, A.; Wang, S.; Fang, H.; Rose, C.; et al. Improving model factuality with fine-grained critique-based evaluator. arXiv preprint arXiv:2410.18359 2024.
- Wen, X.; Liu, Z.; Zheng, S.; Ye, S.; Wu, Z.; Wang, Y.; Xu, Z.; Liang, X.; Li, J.; Miao, Z.; et al. Reinforcement learning with verifiable rewards implicitly incentivizes correct reasoning in base LLMs. arXiv preprint arXiv:2506.14245 2025.
- Helff, L.; Delfosse, Q.; Steinmann, D.; Härle, R.; Shindo, H.; Schramowski, P.; Stammer, W.; Kersting, K.; Friedrich, F. LLMs Gaming Verifiers: RLVR can Lead to Reward Hacking. In Proceedings of the ICLR 2026 Workshop on Logical Reasoning of Large Language Models, 2026.
- Chen, G.H.; Chen, S.; Liu, Z.; Jiang, F.; Wang, B. Humans or LLMs as the Judge? A Study on Judgement Bias. In Proceedings of the Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing; Al-Onaizan, Y.; Bansal, M.; Chen, Y.N., Eds., Miami, Florida, USA, 2024; pp. 8301–8327. [CrossRef]
- Ye, J.; Wang, Y.; Huang, Y.; Chen, D.; Zhang, Q.; Moniz, N.; Gao, T.; Geyer, W.; Huang, C.; Chen, P.Y.; et al. Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge. In Proceedings of the The Thirteenth International Conference on Learning Representations, 2025.
- Shao, R.; Li, S.S.; Xin, R.; Geng, S.; Wang, Y.; Oh, S.; Du, S.S.; Lambert, N.; Min, S.; Krishna, R.; et al. Spurious Rewards: Rethinking Training Signals in RLVR. In Proceedings of the Forty-third International Conference on Machine Learning, 2026.
- Yue, Y.; Chen, Z.; Lu, R.; Zhao, A.; Wang, Z.; Yue, Y.; Song, S.; Huang, G. Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? In Proceedings of the The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025.
- Papineni, K.; Roukos, S.; Ward, T.; Zhu, W.J. BLEU: a method for automatic evaluation of machine translation. In Proceedings of the Proceedings of the 40th annual meeting of the Association for Computational Linguistics, 2002, pp. 311–318.
- Lin, C.Y. ROUGE: A package for automatic evaluation of summaries. In Proceedings of the Text summarization branches out, 2004, pp. 74–81.
- Zhang, T.; Kishore, V.; Wu, F.; Weinberger, K.Q.; Artzi, Y. BERTScore: Evaluating Text Generation with BERT. In Proceedings of the International Conference on Learning Representations, 2020.
- Williamson, P.R.; Altman, D.G.; Bagley, H.; Barnes, K.L.; Blazeby, J.M.; Brookes, S.T.; Clarke, M.; Gargon, E.; Gorst, S.; Harman, N.; et al. The COMET handbook: version 1.0. Trials 2017, 18, 280.
- Kim, T.S.; Lee, Y.; Park, Y.; Kim, J.; Kim, Y.H.; Kim, J. Cupid: Evaluating personalized and contextualized alignment of llms from interactions. arXiv preprint arXiv:2508.01674 2025.
- Ali, D.; Zhao, D.; Koenecke, A.; Papakyriakopoulos, O. Operationalizing Pluralistic Values in Large Language Model Alignment Reveals Trade-offs in Safety, Inclusivity, and Model Behavior. arXiv preprint arXiv:2511.14476 2025.
- Reimers, N.; Gurevych, I. Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks. In Proceedings of the Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, 11 2019.
Figure 1.
Quantitative meta-analysis of feedback signals in HelpSteer3. , P, S, and F denote aggregated preference, individual preference, score, and language feedback, respectively.
Figure 1.
Quantitative meta-analysis of feedback signals in HelpSteer3. , P, S, and F denote aggregated preference, individual preference, score, and language feedback, respectively.

Figure 2.
A feedback-centric view of LLM alignment. Representative examples illustrating how feedback structure determines training and evaluation choices.
Figure 2.
A feedback-centric view of LLM alignment. Representative examples illustrating how feedback structure determines training and evaluation choices.

Figure 4.
Compact overview of representative alignment methods (Section 3), organized by feedback type (rows) and by whether they are reward-model-free or reward-model-based (columns).
Figure 4.
Compact overview of representative alignment methods (Section 3), organized by feedback type (rows) and by whether they are reward-model-free or reward-model-based (columns).

Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (http://creativecommons.org/licenses/by/4.0/).
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.