Preprint
Hypothesis

This version is not peer-reviewed.

Auditors, Not Oracles: Sparse Human Grounding for Bias-Aware Self-Alignment

Submitted:

20 September 2026

Posted:

21 September 2026

You are already at the latest version

Abstract
Reinforcement learning from human feedback (RLHF) treats annotators as preference oracles. This is expensive and can encode annotator-pool bias into the reward signal: rater demographics, positional artifacts, and misspecified helpfulness criteria can become part of what the model is trained to produce. A growing body of work—including AI-labeled feedback on summarization and dialogue, direct preference optimization, self-rewarding models evaluated on AlpacaEval 2.0, and iterative self-evolution for multimodal models—shows that competitive benchmark performance can, in several evaluated settings, be obtained without dense human pair labels. Whether annotator-induced bias decreases when humans leave the loop is a separate question that these works do not directly measure. We propose a protocol called SPARSE-ALIGN in which humans remain in the alignment loop but perform a different task. A policy generates its own candidate pairs. A frozen LLM judge, operating inside a debate structure intended to reduce reward hacking, evaluates the pairs for consistency. High-consistency pairs are accepted automatically. A small, stratified sample of routed pairs—at most 5% of training pairs in any run—is sent to human auditors. Auditors never directly provide an A/B preference or ranking. Instead, they flag non-factual bias artifacts (length preference, positional bias, sycophancy, demographic stereotyping, self-preference) or, when objective correctness is at stake, provide a sparse factual reference. A factual reference may condition the automated judge and therefore indirectly affect the machine-generated preference label; direct human pairwise preference labels remain excluded from the training process. We hypothesize that this arrangement can match RLHF on instruction-following quality while reducing annotator-induced bias relative to dense RLHF and ungrounded self-rewarding. The paper specifies the experiments, statistical tests, ablations, and failure conditions that would confirm or falsify those predictions. No experimental results are reported here.
Keywords: 
;  ;  ;  ;  ;  ;  ;  
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.