Preprint
Article

This version is not peer-reviewed.

Validating LLM-Generated Expert Judgments in AHP Decision Support: A Preregistered Cross-Study Benchmark and Randomized-Order Audit

A peer-reviewed version of this preprint was published in:
Systems 2026, 14(10), 1213. https://doi.org/10.3390/systems14101213

Submitted:

03 September 2026

Posted:

03 September 2026

You are already at the latest version

Abstract
Large language models (LLMs) are entering decision-support systems as inexpensive synthetic experts, yet their judgments are rarely validated against published human targets. We benchmark six LLMs on a source-audited corpus of 27 published Analytic Hierarchy Process (AHP) studies in English and Korean, with a preregistered holdout of 22 studies and 21 reference-weight tasks, retaining 28,783 valid model calls across a preregistered phase and a prospectively frozen extension; inference is task-level with study-clustered uncertainty. Holdout rank replication was modest and model-dependent, and only HCX-007 clearly beat uniform weights on absolute error. A diagnostic showed that the fixed-source-order rank evaluation is confounded with criterion presentation order: the corpus’s reference vectors largely follow source listing order, and a model-free descending-order rule reaches a mean correlation of 0.735, outscoring every model tested. A prospective randomized-order experiment reversed the point-estimate pattern: HCX-007, the only model that strongly tracks presentation order, was also the only one whose accuracy declined, and a low-yield direct-recall probe found no verbatim reproduction of published weights. Persona conditioning was often weaker than repeated-draw variability. Synthetic panels can support piloting and stress-testing of decision pipelines, but replay benchmarks must randomize presentation order before rank accuracy can be read as reconstructed expert judgment.
Keywords: 
;  ;  ;  ;  ;  ;  ;  

1. Introduction

Structured expert judgment supports technology assessment, risk analysis, policy design, and strategic foresight. These applications often require a panel to compare multiple criteria, explain trade-offs, and preserve disagreement among stakeholders. Recruiting qualified experts is costly, however, and large language models (LLMs) offer an apparently scalable alternative. The broader “silicon sampling” literature has reported that demographically conditioned LLMs can reproduce some aggregate response patterns [1]. This promise has encouraged researchers to use LLMs as simulated participants, role-playing agents, and virtual experts, and to feed their judgments into decision-support systems, where model-generated criterion weights can stand in for an expert panel.
The substitution claim remains unsettled. LLM samples can diverge from target populations [2], flatten identity groups [3], and exhibit psychometric structures that are more internally regular than their human references [4]. These limitations are especially important in expert elicitation. A decision-support system does not merely require a plausible answer; it requires criterion weights that reproduce the direction, magnitude, and heterogeneity of domain judgments.
The Analytic Hierarchy Process (AHP) provides an unusually informative test bed. AHP converts reciprocal pairwise comparisons into criterion weights and quantifies internal coherence using the consistency ratio (CR) [5,6]. Published AHP studies provide two evidentiary levels that must not be conflated: aggregate panel weights can test consensus reconstruction, whereas individual matrices are required to test whether synthetic responses resemble the distribution of human judgments. Recent work has evaluated LLMs as “virtual experts” in individual application domains, both within AHP [7] and in structured expert questionnaires validated against newly collected human samples [8]. What is still needed is a cross-study benchmark with simple baselines, repeated sampling, a holdout design, and explicit separation between evidence about panel means and evidence about panel diversity. Viewed from a systems perspective, a synthetic expert panel is a stochastic judgment-generation subsystem inside an AI-enabled decision-support system: its outputs feed AHP aggregation, the resulting criterion weights steer downstream alternative selection, and its errors therefore propagate to system-level decisions. Validating that subsystem before deployment is the problem this paper addresses.
This study builds that benchmark from 27 source-validated published AHP studies spanning health, environment, engineering, public policy, and organizational decision making, drawn from 29 collected sources. Five studies informed pilot development, and the 24 subsequently collected studies constituted the preregistered holdout; source validation retained 22 of them for scientific analysis. Three preregistered models were crossed with three prompting conditions and three repetitions. A subsequently frozen prospective extension applied the same design to three newer or regionally distinct models. Published AHP decisions were reconstructed from source matrices or reported weights, and all uncertainty intervals were computed at the study level. The evaluation discipline of simple baselines, preregistration with a holdout, and implementation audits carries over from our prior validation of a synthetic persona panel [9].
The paper makes three contributions.
1.
A cross-study external-validation benchmark. It provides a reusable, multilingual benchmark that tests LLM-generated structured expert judgments against 27 published AHP studies with a preregistered 22-study holdout, evaluating rank correlation, top-ranked-criterion agreement, absolute weight error, and consistency.
2.
A baseline discipline for judgment replay. It enforces uniform-weight, chance, random-prior, single-call, and repeated-call baselines, and adds a model-free descending-order rule that exposes criterion presentation order as a first-order confound of fixed-order replay: in the benchmark corpus, published reference vectors largely follow source listing order, and the order rule outscores every model tested.
3.
A pre-deployment audit protocol. It separates aggregate accuracy from individual human-likeness, audits whether nominal persona slots contain stable profile information beyond repeated-draw noise, and probes the order confound prospectively with a randomized-order experiment and a direct-recall probe, distilling the results into a validation checklist for LLM-assisted decision-support pipelines.
The full analysis leads to a more cautious conclusion than an observation-level comparison would suggest. Several LLMs contain a rank signal and one extension model improves on equal weighting, but performance remains strongly model- and metric-dependent. The available published panel descriptions are also often too sparse to instantiate distinct synthetic experts.

3. Materials and Methods

3.1. Pilot and Preregistered Holdout

Search and screening were conducted on 21 July 2026. The systematic English-language open-access route used the Europe PMC REST search endpoint with cursor-based retrieval. Its complete sampling-frame query was:
BODY:"pairwise comparison matrix" AND
BODY:"consistency ratio" AND
BODY:"experts" AND
"analytic hierarchy process" AND
OPEN_ACCESS:Y.
It returned 272 records; 268 had full-text XML and 128 had a matrix-table-caption signal. Ninety-five non-pilot records entered detailed English full-text verification: 18 were initially included, 75 excluded, and two unresolved records were conservatively not included. The subsequent source audit reclassified two initially included records as excluded for reconstruction defects, so the final English counts are 16 included and 77 excluded, as shown in Figure 1.
Korea Citation Index (KCI) full texts were added through a purposive Korean-language diversity supplement, not an exhaustive Korean sampling frame. Ten retrieved papers were verified: six included, three excluded, and one unresolved. The five targeted searches were executed on 21 July 2026 through a commercial web-search API, each requesting ten results. Their literal Korean query strings were recovered verbatim from the archived session log rather than reconstructed from memory, and are reproduced in STUDYSELECTIONAUDIT.md. In translation, they combined “AHP” with (1) pairwise comparison matrix, consistency ratio, experts, weights, and a PDF file-type filter restricted to journal.kci.go.kr; (2) hierarchical analysis, expert questionnaire, pairwise comparison, weights, consistency ratio; (3) expert group, weights, CR, evaluation criteria, industry policy; (4) expert panel, pairwise comparison, priority, weights, consistency; and (5) priority, experts, Delphi, weights, consistency ratio, with a health-care/construction/logistics/education domain disjunction. Because commercial search indices are time-dependent, recovering the queries does not guarantee that the same result set would be returned today; the retrieved records themselves are preserved in the screening log.
Eligibility required (1) an explicit AHP hierarchy and criterion definitions, (2) a numeric aggregate or individual pairwise matrix, or at minimum a weight vector with a consistency ratio, and (3) a described expert panel. Fuzzy AHP and non-Saaty variants were excluded because their scales were not commensurate with the discrete reciprocal generation scale. Full-text verification checked scale type, fuzzy-method terminology, the basis of reported weights, appendices and supplements, and recomputed eigenvectors and consistency statistics.
Across the final 105-record verification log, 80 papers were excluded: 40 used fuzzy or otherwise noncommensurate AHP variants, 30 lacked an adequately described expert panel or used literature-derived weights, four used nonexpert respondents, two used a different method or were methods-only, and four had reporting or task-reconstruction defects that prevented validation. The last category includes two records initially collected but invalidated by the source audit. Three unresolved records were not included. Titles, sources, years, decisions, and recorded reasons are provided row by row in the screening log.
Five studies and 14 tasks were used to develop the corpus schema, prompts, equivalence margins, and analysis code. These observations are treated as pilot data. After preregistration, 24 additional records, each contributing one top-level task, were collected as the holdout. Source validation excluded two records entirely and retained one further task for implementation and CR audits but not accuracy analysis. The validated holdout contains 22 studies and 22 tasks, of which 21 have recoverable reference weights. The full validated corpus contains 27 studies and 36 tasks: 20 studies are in English and seven in Korean; because several pilot studies contain multiple hierarchy levels, task counts are 29 English and seven Korean. The source studies are listed in the released corpus and cited in [21,22,23,24,25,26,27,28,29,30,31,32,33,34,35,36,37,38,39,40,41,42,43,44,45,46,47]. Table 1 summarizes the resulting corpus and analysis sets.

3.2. Two Validation Targets

The design distinguishes aggregate consensus replication from individual-judgment resemblance. The first target asks whether a synthetic panel reconstructs the published aggregate weight vector. It is estimable for 21 holdout tasks. The second asks whether synthetic matrices reproduce human heterogeneity and consistency. It requires hierarchy-matched individual human matrices, available for only two holdout tasks (12 human matrices). Consequently, aggregate-weight accuracy is the primary benchmark, whereas individual human-likeness is descriptive and cannot be confirmed.

3.3. Reference Reconstruction

When a source paper reported a reciprocal matrix, we recomputed the principal eigenvector and CR. When only normalized criterion weights were available, the reported vector was used after checking its sum and dimension. For a pilot study that reported individual weight vectors and an aggregate group vector, the group vector was used as the replication target. Two holdout papers could not support a valid replay task because one source’s published matrix, reported weights, and stated CR were mutually incompatible, and the other conflated two national-development contexts in a degenerate two-criterion task; both records were excluded from all scientific analyses. One further holdout task reported only within-SWOT-domain priorities, not the generated top-level SWOT vector. It remained eligible for implementation and CR summaries but was excluded from accuracy metrics.
Individual human CR values could be matched unambiguously to only two holdout tasks (12 matrices). Three additional matched tasks came from one pilot study, increasing the pooled descriptive comparison to five tasks from three studies and 30 human matrices. A fourth source contained individual CR values, but its generated two-criterion task had C R = 0 by construction while the published CR values referred to another hierarchy level. It was therefore excluded from the matched CR analysis.

3.4. Models and Prompting Conditions

Collection began on 21 July 2026 at 22:51 KST; the first pass ended on 22 July at 01:30 KST and the prespecified recovery pass ended later that morning. The requested identifiers and routes were gemini-3.5-flash through Google GenAI (temperature 1.0, top-p 1.0, low thinking, JSON MIME type); openai/gpt-5-mini through OpenRouter (minimal reasoning, JSON-object format, no temperature sent); and LGAI-EXAONE/K-EXAONE-236B-A23B through FriendliAI (temperature 1.0, top-p 1.0, 4096 maximum tokens, thinking disabled). Every request was stateless and asked for all upper-triangle pairwise comparisons in machine-readable JSON.
Three conditions were crossed with model and task:
  • Persona: source-paper panel information plus the decision context and criterion definitions;
  • Generic expert: the same context and criteria without a panel-profile description;
  • Names only: criterion names with the contextual description removed.
Each nominal panel slot was sampled three times. Parsing removed code fences and text outside the first and last JSON braces, then validated keys and values against the requested Saaty scale. A failed response triggered a guided repair request naming up to eight errors, with at most three attempts; rate-limit backoff was handled separately. One prespecified recovery pass was allowed. Remaining failures were retained as missing.
Exact system prompts, deterministic user-message construction, and parsing/repair rules are reproduced in MODELCALLREPRODUCIBILITY.md. Historical JSONL files preserve parsed judgments and derived statistics, not the original provider payload. The runner did not log the OpenRouter-served provider, immutable backend revision, per-call timestamps and request IDs, token use, latency, or the historical SDK lockfile. These fields cannot be recovered and are not implied by the requested model labels.

3.5. Prospective Post-Registration Model Extension

A Version 1.1 extension protocol was locally time-stamped and hashed on 29 July 2026 before any extension benchmark calls. It was not posted as a new OSF registration; extension results are therefore prospective post-registration evidence and do not enlarge or revise the original confirmatory family. Collection ran from 29 July at 22:52 KST to 30 July at 12:44 KST.
The requested models were google/gemini-3.6-flash through OpenRouter pinned to Google AI Studio, anthropic/claude-sonnet-5 through OpenRouter pinned to Anthropic, and HCX-007 through NAVER Cloud CLOVA Studio. OpenRouter fallbacks were disabled and parameter support was required. The returned endpoint revisions were google/gemini-3.6-flash-20260721 and anthropic/claude-sonnet-5-20260630. All models used a low reasoning/effort setting. Sampling controls unsupported across providers were not forced, and all models used the same JSON-only instruction, parser, Saaty-scale validator, and three-attempt repair rule.
The frozen extension collector used 603 then-current panel slots, three conditions, and three repetitions: 5,427 physical base cells per model and 16,281 overall. The subsequent source audit defined a common scientific analysis scope of 534 historically collected slots, or 4,806 cells per model and 14,418 overall. Exactly one recovery pass was applied to failed HCX-007 cells. Complete prompts, request and response envelopes, request identifiers, timestamps, returned model and provider routes, token use, and latency were retained privately; public records remove prompts, raw envelopes, and returned thinking content.
The primary extension analysis retained the persona condition and the 21 holdout tasks with reference weights. Prespecified paired contrasts compared Gemini 3.6 Flash with Gemini 3.5 Flash, Claude Sonnet 5 with GPT-5-mini, and HCX-007 with K-EXAONE. Spearman correlation and MAE formed one six-test family with Holm adjustment. Kendall τ b , tie-aware top-1, prompt contrasts, language and criterion-count strata, leave-one-study-out deletion, uniform and single-call baselines, consistency, latency, and cost were secondary.
In the secondary names-only condition, one Gemini and one Claude task produced exactly uniform aggregate vectors, for which rank correlation is undefined. Analysis Version 1.1a reports the number defined, summarizes only defined correlations, and supplies a zero-imputation sensitivity mean. This correction was made after the first analysis run, is documented separately, and cannot affect the persona condition or any of the six primary tests.

3.6. Post-Collection Data Audit

The raw registered-model files contained 606 nominal panel slots. Inspection of one pilot source (C21) showed that three published group-summary vectors had been parsed as individual experts. Source verification then corrected panel counts and targets in 15 records, excluded two invalid replay tasks, and removed one further task from accuracy analysis. Two Korean studies described human panels of 22 and 24, but the frozen historical files contained only eight synthetic slots for each; the source panel counts are preserved in the corpus while analysis is capped at the actually collected slots. Raw observations were preserved. These structural rules leave 534 historical slots and 14,418 in-scope registered-model calls, of which 14,374 contained valid AHP outputs (99.69%). No outcome-based exclusion was performed; every inclusion or exclusion decision was made from source-document properties alone, independent of model outputs.
The audit also distinguished nominal slots from unique profile texts. The holdout contained 304 historically collected analysis slots and 38 within-task unique profile strings. Fifteen of 22 holdout tasks repeated one identical profile across all slots. Such tasks can measure repeated synthetic draws, but they cannot identify a between-profile effect.

3.7. Accuracy Metrics and Baselines

For each model, condition, and task, response-level weight vectors were averaged to obtain w ^ . Four metrics were computed; the first three are
ρ t = Spearman ( w ^ t , w t ) ,
τ b , t = error ( w ^ t , w t ) ,
M A E t = 1 n t ∑ j = 1 n t | w ^ t j − w t j | .
Kendall τ b corrects for ties. The fourth metric, top-1 agreement, used an explicit tie rule: if P t is the set of predicted maxima and R t the set of reference maxima, T o p 1 t = | P t ∩ R t | / | P t | and its conditional chance expectation is | R t | / n t . No actual holdout top ties occurred, so this rule did not alter the reported values. The absolute-error baseline was the uniform vector u t = ( 1 / n t , … , 1 / n t ) . This baseline is intentionally simple: an LLM panel should not be considered useful for weight elicitation if it cannot improve on equal weights. As an aggregation sensitivity analysis, we also geometrically averaged pairwise matrices (aggregation of individual judgments) before computing the eigenvector.
Three additional baselines isolate where any apparent value originates. First, a symmetric Dirichlet ( 1 ) baseline samples random simplex weights. Second, a leave-one-study-out empirical-shape prior samples reference-weight components only from other studies, selects n t components, and renormalizes them; criterion labels remain exchangeable because meanings differ across tasks. Both used 20,000 draws per task. Third, a single-generic-call baseline was compared with the aggregate of all valid generic-expert calls for each task. This last contrast tests the value of repeated LLM sampling itself, independently of persona wording.

3.8. Consistency and Profile Differentiation

Human and LLM CRs were summarized within each matched task before comparison. For the five-task pooled set, an exact paired Wilcoxon test and an exact sign-flip test were computed. A study-level sign-flip sensitivity analysis addressed dependence among the three tasks from one pilot paper. The two-task holdout was reported descriptively because it cannot support a meaningful confirmatory test.
For tasks with at least two unique profile texts, we averaged all draws from each text to obtain a profile centroid. The profile-signal ratio was
R profile = mean SD ( profile centroids ) mean SD ( draws within profile ) .
A ratio below one indicates that variation attributable to profile wording is smaller than repeated-draw variation. This analysis is exploratory because only seven holdout tasks contain multiple profile strings and only three draws were obtained per nominal profile slot. It is therefore an audit of the implemented profile design, not a high-powered causal persona experiment.

3.9. Statistical Inference

The primary analysis uses only the 22 source-validated studies from the preregistered holdout. Means and confidence intervals were calculated by resampling studies with replacement (50,000 cluster-bootstrap draws; seed 20260729). Persona effects were estimated as paired task differences relative to the generic-expert condition. Equivalence was declared when the 90% cluster-bootstrap interval fell within the pilot-derived bounds of ± 0.22 for ρ and ± 0.20 for top-1 agreement; the registration stated that these margins would be derived from pilot variance, and the operational rule and numeric bounds were fixed in a locally time-stamped commit before collection (Section 3.11). Results combining pilot and holdout studies are explicitly labeled descriptive or sensitivity analyses.
Robustness analyses added Kendall τ b ; strata for 2–3, 4–5, and 7–9 criteria; equivalence margins narrowed to 75% and 50% of the pilot-fixed bounds; and leave-one-study-out estimates. Exploratory robustness summaries used 10,000 study-cluster bootstrap draws. The registered entropy outcome used two-sided paired tests and Holm adjustment across the three model-specific H5 tests. No unadjusted battery of new significance tests was introduced for the exploratory strata or baseline comparisons; inference emphasizes effect sizes and bootstrap intervals.

3.10. Presentation-Order Diagnostic and Prospective Randomized-Order Audit

Two further analyses concern criterion presentation order. Neither belongs to the registered protocol, and their chronology is stated here so that confirmatory and exploratory claims remain separable: the diagnostic was constructed post hoc, after the extension benchmark had been inspected, whereas the randomized-order experiment it motivated was specified in a frozen protocol before any of its calls were issued. Section 4.8 reports the diagnostic and Section 4.9 the experiment.
The corpus preserves the order in which each source paper lists its criteria. The diagnostic correlates, for each holdout task, the reference weight vector with the reverse presentation index, which ranks criteria from last listed to first, and evaluates a model-free descending-order rule that assigns weights in strictly decreasing order of listing position. Both quantities were computed on all 21 holdout tasks and, because rank correlation is degenerate with two criteria, also on the 19 tasks with at least three criteria. Each model’s order tracking was measured by correlating its aggregate predicted weights with the reverse presentation index.
The randomized-order experiment followed a Version 1.2 protocol that was locally time-stamped and hashed before any new call and then executed prospectively. It crossed the three extension models with the 21 holdout tasks under three conditions: anchor, presenting criteria in the source order; shuffled, permuting the presentation order independently on every call; and recall, asking directly, without any AHP elicitation, for the weights the source paper published. Persona conditioning was excluded by using the generic-expert instruction throughout. Shuffled responses were mapped back to the original criterion order before scoring, so the reference vector was never altered. Eight repetitions per task were collected for the anchor and shuffled conditions and two for recall; the frozen protocol text stated eight for all three conditions, a discrepancy recorded in Table S1 in the Supplementary Materials. Permutations were drawn from a deterministic seed that included the model name, so the three models received different realized schedules, and no task shared an identical schedule across all three models.
The frozen protocol prespecified a benchmark-relative decision rule, HP1, for HCX-007, the model whose fixed-order advantage was under test: position bias is judged the primary cause of that advantage when the study-clustered 95% confidence interval of the model’s shuffled-condition mean ρ lies entirely below the descending-order rule’s fixed-order accuracy. This rule is a benchmark-relative operational screening criterion, not a causal identification on its own; a model without any fixed-order advantage can also satisfy it, so the judgment rests jointly on the manipulation check. Anchor-minus-shuffled degradation was tested per model with two-sided Wilcoxon signed-rank tests on task-level paired differences, Holm-adjusted across the three models; interval estimates used study-clustered bootstrap resampling (10,000 draws; seed 20260801). A recall response counted as a direct-recall match when a parsable weight vector reproduced the reference vector within a mean absolute error of 0.02; unparsable responses were tallied separately, and the condition is reported descriptively. Two exploratory analyses accompany the experiment: a paired contrast of descending-order-rule accuracy minus each model’s shuffled accuracy, which propagates the rule’s own study-level uncertainty through the same bootstrap, and a schedule sensitivity analysis that computes the maximum position-exposure imbalance of each realized permutation schedule and its task-level Spearman correlation with shuffled accuracy. A manipulation check correlated each model’s weights with the order in which criteria were actually presented on each call, in both conditions.

3.11. Preregistration Adherence and Deviations

The protocol was registered on 21 July 2026 (osf.io/39jsa), before collection of the holdout responses. The authors had already inspected the five-study pilot, which was disclosed in the registration. The registration stated that the H4 equivalence bounds would be derived from pilot variance but specified neither the numerical values nor the operational rule; the rule (one half of the standard deviation of pilot task-level paired differences), the numeric bounds, and the analysis code were frozen minutes after registration, and hours before collection began, in a locally time-stamped Git commit (6fbd68c). The freeze record carrying the full 40-character hash is archived in the public OSF project. The registration also asserted that independent, stateless API calls preclude order effects, a claim that conflated between-conversation carryover with within-prompt position effects; Table S1 records this deviation, and Section 4.8 and Section 4.9 test it. The same table maps every material change identified in the post-collection audit to its reason and analytical consequence. Deviations caused by inadequate human reference data were handled by withdrawing the corresponding confirmatory claim rather than substituting an observation-level test.

3.12. Use of Generative AI Tools

The authors used OpenAI Codex [48] for statistical code review, for initial drafting of figure captions and of Sections 3.4–3.7, 4, and 5, and for language editing throughout the manuscript; the service did not expose immutable model version identifiers. The authors reviewed the raw data, reran all analyses, verified citations and numerical results, and take full responsibility for all analytical decisions and claims.

4. Results

Sections 4.1–4.7 report the preregistered and extension benchmarks in the order they were run. Section 4.8 and Section 4.9 then present the study’s central methodological finding: a presentation-order confound in fixed-order replay and its prospective randomized test.

4.1. Aggregate Consensus Weight Replication

Table 2 and Figure 2 show the 21-task holdout results. Gemini had the highest mean rank correlation, ρ = 0.27 [0.05, 0.49]. GPT-5-mini and K-EXAONE had lower point estimates but intervals that included zero. Arithmetic and geometric aggregation gave similar rank estimates for Gemini (0.27 versus 0.28), while GPT-5-mini decreased from 0.16 to 0.14 and K-EXAONE from 0.15 to 0.05. Mean Kendall τ b was 0.24 [0.06, 0.42] for Gemini, 0.17 [ − 0.11 , 0.44] for GPT-5-mini, and 0.14 [ − 0.13 , 0.41] for K-EXAONE.
Top-1 agreement ranged from 29% to 43%, compared with a task-size-adjusted chance expectation of 26.2%. For every model, the confidence interval for the paired advantage included zero. The same was true for MAE relative to uniform weights. There were no predicted or reference top-1 ties in these 21 tasks. Thus, the holdout provides evidence of a modest rank signal, clearest for Gemini, but not of a reliable advantage on absolute weights or selection of the top criterion.

4.2. Persona Implementation Audit and Prompt Contrast

Table S2 in the Supplementary Materials compares persona prompting with both registered baselines. Gemini and GPT-5-mini met both pilot-fixed equivalence criteria relative to the generic-expert condition. For Gemini, persona prompting was directionally worse in rank correlation ( Δ ρ = − 0.103 ), although the entire 90% interval remained within the equivalence bound. GPT-5-mini’s difference was near zero. K-EXAONE did not establish equivalence because its rank-correlation and top-1 intervals crossed the prespecified margins; this is an inconclusive result, not evidence that personas improved performance. Relative to names-only prompting, no model met both equivalence criteria: K-EXAONE met the top-1 criterion only. The equivalence finding was margin-sensitive. At 75% of the pilot-fixed bounds, Gemini no longer passed both generic-expert criteria, GPT-5-mini still did, and K-EXAONE did not; at 50%, none passed. No names-only contrast passed both criteria even at the full pilot-fixed bounds. Thus, equivalence is evidence relative to a prespecified tolerance, not proof of identical prompt effects.
The persona implementation audit limits the interpretation of these comparisons. Fifteen holdout tasks supplied one repeated profile text for every panel slot. Only seven holdout tasks contained multiple profile texts. In the pooled 18 eligible tasks, profile-signal ratios were 0.60 (Gemini), 0.44 (GPT-5-mini), and 0.42 (K-EXAONE): variation among profile centroids was smaller than variation among repeated draws from the same profile (Figure 3). This supports the narrow conclusion that the implemented profile texts did not generate stable differentiation, rather than a general claim that richer, independently validated personas can never matter.

4.3. Individual-Judgment Evidence Is Descriptive

Only two holdout tasks reported individual human CRs that could be matched to the generated hierarchy. Across their 12 human matrices, the pooled median CR was 0.109. Synthetic medians were 0.007 (Gemini), 0.043 (GPT-5-mini), and 0.154 (K-EXAONE). These contrasts are potentially important but cannot support a confirmatory test with two tasks.
Adding the three matched pilot tasks yields five tasks from three studies. The human median becomes 0.125, compared with 0.022 for Gemini, 0.127 for GPT-5-mini, and 0.158 for K-EXAONE. Gemini was below the human median in all five tasks, but the two-sided exact task-level test was p = 0.0625 ; after reducing the analysis to three study-level contrasts, the exact sensitivity test was p = 0.25 . GPT-5-mini ( p = 1.00 ) and K-EXAONE ( p = 0.625 ) did not show a consistent directional shift. Figure S1 in the Supplementary Materials displays the matched contrasts and the pilot/holdout boundary.
The original row-level comparison would have contrasted thousands of LLM responses with a few human matrices and produced spuriously small p-values. The corrected task-level analysis does not support a general confirmed “over-consistency” claim. It supports a model-specific hypothesis for future testing, especially for Gemini.

4.4. Expanded Baselines and Robustness

Table S3 in the Supplementary Materials isolates the effect of repeated generic-expert calls. Averaging approximately 42 valid generic calls per task reduced MAE relative to the expected single call for all three models. The paired improvement was − 0.005 [ − 0.011 , − 0.002 ] for Gemini, − 0.017 [ − 0.030 , − 0.007 ] for GPT-5-mini, and − 0.035 [ − 0.050 , − 0.022 ] for K-EXAONE. Rank-correlation gains were uncertain for Gemini and GPT-5-mini but clear for K-EXAONE. Repeated sampling therefore provides some variance-reduction value, especially in absolute error, without establishing that the resulting aggregate is human-like.
Model-free baselines reinforced the need for simple comparators. Symmetric Dirichlet ( 1 ) and leave-one-study-out empirical-shape priors had expected rank correlations of zero and top-1 agreement of 0.262, because criterion labels are exchangeable across domains. Their mean MAEs were 0.202 [0.170,0.235] and 0.184 [0.153,0.216], respectively, versus 0.143 [0.109,0.182] for uniform weights. Thus, uniform weights remained the strongest naive absolute-error baseline.
Results were heterogeneous by criterion count. Across the 2–3, 4–5, and 7–9 criterion strata, mean ρ was, respectively, 0.313, 0.171, and 0.341 for Gemini; 0.329, 0.071, and 0.054 for GPT-5-mini; and 0.329, 0.214, and − 0.159 for K-EXAONE. MAE declined mechanically as the number of normalized weights increased, so cross-stratum MAE magnitudes should not be interpreted as improved model competence. Leave-one-study-out deletion changed mean ρ by at most 0.044, 0.058, and 0.054 and MAE by at most 0.008, 0.009, and 0.009 for Gemini, GPT-5-mini, and K-EXAONE. No single study overturned the overall weak and uncertain replication conclusion. Full task and interval results are released in analysis_reference/ieee_analysis.json and analysis_reference/robustness_summary.csv in the OSF deposit.

4.5. Registered Entropy Outcome

Table S4 in the Supplementary Materials restores the registered H5 analysis. Positive differences indicate a flatter LLM weight distribution than the published reference. Gemini and GPT-5-mini had small positive point estimates with intervals including zero. K-EXAONE produced substantially higher normalized entropy ( Δ H = 0.116 [0.078, 0.155]), consistent with excessive equalization of criterion weights. Its paired test remained significant after Holm adjustment within the three model-specific H5 tests. Because the complete registered H1–H5 multiplicity family could not be executed, this result is reported as a secondary registered outcome rather than a standalone confirmatory success.

4.6. Combined Benchmark as a Sensitivity Analysis

Combining pilot and holdout data increases coverage to 35 tasks with reference weights. Mean rank correlations are 0.33 [0.16, 0.47] for Gemini, 0.23 [0.02, 0.44] for GPT-5-mini, and 0.09 [ − 0.11 , 0.33] for K-EXAONE. These values are useful as descriptive benchmark estimates but are not independent confirmation because the pilot informed the preregistration. The same distinction applies to human-dispersion comparisons: only two holdout tasks contain individual weights, whereas the combined set contains eight. No confirmatory homogenization claim is made.

4.7. Prospective Model Extension

The physical extension archive recovered 16,271 of 16,281 originally planned cells. In the source-validated scientific scope, 14,409 of 14,418 cells were valid (99.94%): Gemini 3.6 Flash and Claude Sonnet 5 each yielded all 4,806 cells, and HCX-007 yielded 4,797. Every OpenRouter response used the prespecified provider route. Table 3 reports holdout accuracy and in-scope implementation consistency.
Table S5 in the Supplementary Materials compares each extension model with its historical counterpart. HCX-007 showed the largest rank improvement; its raw p = 0.005 became p = 0.031 after Holm adjustment across the six primary tests. Its MAE improvement interval excluded zero under cluster bootstrap, although the multiplicity-adjusted paired-test result did not. Neither Gemini 3.6 Flash nor Claude Sonnet 5 showed a clear improvement. Thus, the extension supports a family-wise rank-correlation improvement for HCX-007, while its absolute-error contrast remains estimation-led and multiplicity-sensitive.
Only HCX-007 clearly improved on uniform weights, with Δ M A E = − 0.042 [ − 0.084 , − 0.006 ]. The corresponding Gemini 3.6 Flash estimate was − 0.003 [ − 0.052 , 0.040], and the Claude Sonnet 5 estimate was 0.010 [ − 0.029 , 0.045]; both intervals included zero. Aggregating generic calls rather than taking the expected single generic call reduced MAE by 0.010 [0.003,0.018], 0.002 [0.001,0.003], and 0.019 [0.011,0.029], respectively. HCX-007 also gained 0.132 [0.050,0.218] in rank correlation from aggregation.
Secondary analyses were stable to deleting one study: the largest absolute change in mean ρ was 0.066, 0.059, and 0.055 for Gemini 3.6 Flash, Claude Sonnet 5, and HCX-007. Persona-minus-generic rank differences were 0.064 [ − 0.083 ,0.257], 0.072 [0.008,0.148], and 0.061 [ − 0.023 ,0.171]. These unadjusted secondary intervals do not establish a general persona effect.
Three exploratory task-level analyses characterize the HCX-007 result more precisely than its mean does. All are descriptive and were computed after the primary tests.
First, the advantage is not confined to Korean tasks. In the 15 English holdout tasks alone, HCX-007 reached mean ρ = 0.619 [0.400, 0.806], an interval excluding zero, against 0.223 [ − 0.099 , 0.521] for Gemini 3.6 Flash and 0.267 [ − 0.005 , 0.513] for Claude Sonnet 5. A purely language-driven account does not explain this English-only separation.
Second, the separation is uneven across tasks and traces to tail behavior. On those 15 English tasks, HCX-007 beat Gemini 3.6 Flash in eight, lost in four, and tied in three; the mean difference of + 0.396 reflects asymmetric magnitudes (a mean gain of 0.890 when ahead, a mean loss of 0.295 when behind), and a paired sign test on this pattern would not be significant. The asymmetry has a concrete source: rank inversions, defined as tasks with ρ < 0 , occurred in 6 of 21 holdout tasks for Gemini 3.6 Flash and 7 of 21 for Claude Sonnet 5, both reaching ρ = − 1.00 , but in only 1 of 21 for HCX-007, whose worst task was ρ = − 0.40 ; HCX-007 also reached ρ ≥ 0.5 in 17 tasks versus 12 and 9. Its advantage is the near-absence of catastrophic ordering failures rather than uniformly higher peak accuracy, and it holds within the English subset.
Third, the Korean results are weaker than their headline values suggest, and their error pattern is informative about memorization. Although HCX-007 matched the top criterion in all six Korean tasks, four of those tasks have only two or three criteria, where the conditional chance expectation is 0.33 to 0.50; the two genuinely demanding cases are seven-criterion tasks with ρ = 0.786 and ρ = 1.000 . HCX-007 also recovered Korean orderings almost exactly while its magnitudes remained wrong, with task MAEs of 0.031 to 0.166. Verbatim recall of a published weight vector would have reduced MAE as well, so this pattern is more consistent with ordering-level familiarity than with numeric memorization. It does not exclude exposure to abstracts or conclusions, and the extension therefore cannot resolve the question of contamination.
Consistency and profile differentiation again separated from accuracy. Median CR was 0.011 for Gemini 3.6 Flash, 0.008 for Claude Sonnet 5, and 0.136 for HCX-007; the in-scope proportions below 0.10 were 99.94%, 99.29%, and 23.97%. Across 18 tasks with multiple profile texts, profile-signal ratios were 0.74, 1.42, and 0.38, respectively. Claude Sonnet 5’s ratio above one is exploratory evidence that profile wording can exceed repeated-draw noise for some models, even though its aggregate accuracy remained weak.
Extension token use, cost, and latency are reported in Note S1 in the Supplementary Materials.

4.8. A Presentation-Order Artifact in the Rank Metric

The post hoc diagnostic of Section 3.10 revealed that the rank evaluation itself, computed under the preserved source ordering, is confounded with criterion presentation order. In the present corpus the published tables frequently list criteria in or near descending order of importance: of the 19 holdout tasks with at least three criteria, eight reference vectors are strictly descending in listing order and one more is descending except for a single tie.
By construction, the accuracy of the model-free descending-order rule equals the rank correlation between the reference vector and the reverse presentation index. This model-free accuracy is a mean ρ = 0.735 across all 21 holdout tasks and 0.707 across the 19 tasks with at least three criteria, for which rank correlation is non-degenerate. It exceeds every model evaluated here: on the 21 tasks, HCX-007 attains 0.694, Gemini 3.6 Flash 0.312, and Claude Sonnet 5 0.186; on the 19 tasks, the rule’s 0.707 compares with 0.661, 0.240, and 0.205. No model exceeded the order rule on more than 3 of 21 tasks.
The models differ markedly in how far they track this order. The order-tracking statistic is 0.751 for HCX-007, with 8 of 19 tasks perfectly descending. Every other model evaluated in this study lies far below that value: 0.180 for Gemini 3.6 Flash, 0.173 for Claude Sonnet 5, and, among the preregistered models, 0.178 for Gemini 3.5 Flash, − 0.125 for GPT-5-mini, and 0.047 for K-EXAONE. Across six models the order-tracking statistic is thus confined to [ − 0.13 , 0.18 ] with a single outlier, and that outlier sits close to the corpus artifact itself. The model whose weights follow the listing order most closely is also the model with the highest measured rank accuracy, and the size of its advantage is close to what the order rule alone delivers.
This artifact also explains results that semantic judgment alone could not produce. In one holdout task the criteria are labeled only E ε 1 through E ε 9, carrying no semantic content. Under the names-only condition, which additionally removes the decision context and all definitions, HCX-007 still attained ρ = 0.90 on that task, whose reference vector tracks the reverse presentation index at 0.900 . Ordering nine meaningless labels correctly is not achievable by judgment; it is achievable by assigning weights that decrease monotonically with list position.
The order rule is not a usable elicitation method, because its ordering information comes from the target paper rather than from the decision problem. It is a diagnostic. What it shows is that rank correlation on this benchmark substantially rewards agreement with criterion listing order, so ρ cannot be read as a pure measure of reconstructed expert judgment unless presentation order is randomized. The original protocol did not randomize it. Section 4.9 reports a prospective experiment that does.

4.9. Randomized-Order Experiment

To test whether the order artifact explains the measured differences, the prospective randomized-order experiment specified in Section 3.10 was executed, crossing the three extension models with the 21 holdout tasks under the anchor, shuffled, and recall conditions. Collection yielded 1,134 of 1,134 cells with no failures.
Table 4 and Figure 4 report the outcome. In point estimates, HCX-007 is the only model whose accuracy declined when presentation order was randomized, and it was degraded on more tasks than either comparator (10 of 21, versus 3 and 2), although not on a majority. Its decline of 0.168 is directional rather than conclusive, with an interval including zero and Holm p = 0.160 . Both comparators’ point estimates instead rose; only Claude Sonnet 5’s improvement of 0.455 survived Holm adjustment across the three tests ( p = 0.005 ), while the Gemini 3.6 Flash improvement did not ( p = 0.160 ).
The prespecified HP1 decision rule (Section 3.10) is met: HCX-007’s shuffled interval [0.05, 0.58] lies entirely below the order-rule baseline of 0.735. The criterion is not by itself specific to order-following models, since Gemini 3.6 Flash’s shuffled interval [0.26, 0.63] also lies below the baseline. Together with the manipulation check below, however, the result is consistent with position bias contributing to HCX-007’s fixed-order advantage, although the paired anchor-minus-shuffled decline itself was not statistically conclusive. The exploratory paired contrast that carries the order rule’s own study-level uncertainty is 0.411 [0.074, 0.769] for order-rule accuracy minus HCX-007 shuffled accuracy under the same study-clustered bootstrap, an interval excluding zero. The corresponding contrasts are 0.285 [0.016, 0.539] for Gemini 3.6 Flash and 0.169 [ − 0.049 , 0.389] for Claude Sonnet 5, so the contrast separates the rule from the models generally rather than isolating HCX-007. The schedule sensitivity analysis found mean maximum position-exposure imbalances of 0.30, 0.33, and 0.29 across the three models’ realized permutations, and task-level imbalance did not significantly predict shuffled accuracy for any model (all p ≥ 0.09 ). Within-model contrasts do not require identical schedules across models, but each shuffled estimate still reflects its realized, incompletely balanced permutation schedule; cross-model comparisons, including the point-estimate reversal, additionally reflect realized-permutation variation and remain exploratory.
The manipulation check (Section 3.10) points to the mechanism. The correlation of each model’s weights with the actually presented order, with study-clustered 95% intervals, is 0.471 [0.262, 0.658] for HCX-007 under anchor and 0.428 [0.298, 0.561] under shuffling: both intervals exclude zero, so the model continues to assign greater weight, on average, to earlier-presented criteria, irrespective of content. The comparators exhibit no comparable positive order dependence in either condition (Gemini 3.6 Flash 0.102 [ − 0.176 , 0.371] and − 0.062 [ − 0.155 , 0.038]; Claude Sonnet 5 0.019 [ − 0.264 , 0.297] and − 0.186 [ − 0.355 , − 0.029 ], the latter a small negative dependence whose interval excludes zero). HCX-007 therefore carries a stable, content-independent position bias, which under the source ordering coincided with the corpus artifact; this coincidence is consistent with upward inflation of its measured point estimate.
In point estimates, the distortion runs in both directions: removing the order cue lowered the model that tracked it and raised the models that did not, a pattern consistent with the fixed-order protocol inflating one model’s measured score and suppressing the other two’s; of these paired changes, only the Claude Sonnet 5 increase is statistically established. Under the randomized protocol, the point-estimate ordering of the three models reverses relative to Table 3, with Claude Sonnet 5 highest at 0.566 and HCX-007 lowest at 0.323, although the confidence intervals overlap and this reversal is not itself established.
The recall condition found no direct-recall match under this probe. Across 126 calls the models claimed to know the source study twice, produced a parsable weight vector twice, and reproduced no reference vector within the prespecified tolerance of 0.02 mean absolute error. This is a weak test of memorization, since a model may know a study without reporting it. The extension results therefore do not require direct numeric recall as an explanation: position bias was directly measured and can account for the pattern. The low-yield probe cannot, however, rule out broader training exposure.

5. Discussion

For decision-support practice, the results bear on one question: would a pipeline that consumes model-generated criterion weights reach defensible decisions? The subsections below interpret the evidence from that perspective, covering what the metric disagreements mean for system evaluation, what model updates change, the gap between aggregate accuracy and human similarity, the requirements a persona implementation must meet, and the deployment implications for foresight and decision support.

5.1. Correlation Alone Overstates Decision Utility

The most important technical result is the disagreement among metrics. Gemini’s holdout rank correlation is positive, but its absolute weight error is indistinguishable from uniform weights and its top-1 advantage is uncertain. A system developer reporting only ρ could therefore claim favorable performance even though the resulting decision weights provide no clear improvement over equal weighting.
This metric separation matters operationally. Rank correlation measures whether high-weight criteria tend to remain high, but many decision systems use the weight magnitudes in downstream scoring. Small magnitude errors can change alternative rankings when criterion performances are close. Validation should consequently require at least one rank metric, one absolute metric, and one decision-level metric, each evaluated against a prespecified baseline.
The expanded baselines add a narrower positive finding: aggregating many generic calls reduced MAE relative to one generic call, most strongly for K-EXAONE. This is evidence for Monte Carlo variance reduction, not for the expert validity of the resulting panel. The primary persona aggregates still did not clearly beat uniform weights. “Multiple calls” and “multiple experts” should therefore be treated as different constructs.

5.2. Model Updates Move Performance, Not the Validation Standard

The prospective extension changes the empirical results without changing the evidentiary standard. HCX-007 reproduced aggregate rankings and top criteria substantially better than K-EXAONE and was the only tested model with a study-clustered MAE advantage over equal weights. Its rank-correlation contrast with K-EXAONE survived the six-test Holm correction, whereas its multiplicity-adjusted MAE contrast did not, and three quarters of its individual matrices exceeded the conventional CR threshold. Strong aggregate accuracy therefore did not imply internally consistent or human-like individual judgments.
Gemini 3.6 Flash was not clearly better than Gemini 3.5 Flash, and Claude Sonnet 5 was not clearly better than GPT-5-mini. Newer model labels alone are not a validation argument.
The randomized-order experiment probes one possible source of that advantage. What it identifies is a stable position bias: HCX-007 assigns greater weight, on average, to earlier-presented criteria in both the source order and randomized orders, while the comparators show no comparable positive dependence. Because the source tables in this corpus largely list criteria in descending importance, that bias coincided with the corpus ordering, a coincidence consistent with upward inflation of the model’s measured rank accuracy.
The correct reading of the extension is therefore methodological rather than substantive. The evidence is consistent with the fixed-order benchmark rewarding a response tendency that is independent of the decision content and, in the comparators’ point estimates, penalizing models whose position effects merely added noise. Removing the cue moved all three point estimates, in opposite directions. A synthetic panel that reproduces published rank orders by assigning monotonically decreasing weights is not reconstructing expert judgment, even when its correlation is high.
Two cautions bound this conclusion in turn. HCX-007’s decline under randomization is directional and not statistically conclusive, and the reordering of the three models under the randomized protocol has overlapping intervals. The claim we can defend is narrower: presentation order is a first-order confound in fixed-source-order AHP replay benchmarks of this kind, one of six tested models is strongly sensitive to it, and rank correlation measured under a fixed source ordering cannot separate judgment from position.

5.3. Aggregate Accuracy Is Not Human Similarity

Twenty-one holdout tasks, drawn from 22 studies, support comparison with published aggregate weights, so they can evaluate consensus reconstruction. They cannot establish that individual synthetic matrices resemble individual experts. That stronger claim is supported by only two holdout tasks and is therefore descriptive. This distinction explains why the paper can report an aggregate benchmark while declining to confirm human-like consistency, dispersion, or heterogeneity.
This also clarifies the contribution relative to the two closest validation studies. Memduhoğlu et al. tested six models against ten experts in one 14-criterion solar-siting decision [7], and Fares et al. validated synthetic pavement-engineering experts against a newly collected sample of 12 experts and 11 students in one domain [8]. Both designs compare synthetic and human judgment within a single application. The present study instead tests transportability across published domains with a preregistered holdout, expanded baselines, an implementation audit, and a randomized-order audit of the benchmark itself; its contribution is cross-study transportability and benchmark-confound auditing, not the first validation of synthetic experts. The designs are complementary rather than competing replications of the same claim.

5.4. A Persona Implementation Requires Stable Differentiation

The generic-expert condition met the pilot-fixed equivalence margins relative to persona prompting for Gemini and GPT-5-mini, but narrower-margin sensitivity was less favorable. When multiple profile texts were available, between-profile signal was also smaller than repeated-draw noise. These findings do not establish that persona engineering is universally ineffective. They show that panel labels and sparse role descriptions should not be assumed to create stable synthetic stakeholders.
The audit revealed a more basic design problem: most published papers do not report expert biographies in enough detail to reconstruct distinct individuals. Repeating one role description across 20 or 50 model calls creates a Monte Carlo sample, not a panel of differentiated experts. Future synthetic panel studies should report the number of unique profile texts separately from the number of generated responses and should demonstrate that profile effects exceed within-profile sampling variance.
The extension reinforces the model-specific nature of this requirement. Claude Sonnet 5 produced a profile-signal ratio above one and a positive persona-minus-generic rank contrast, whereas Gemini 3.6 Flash and HCX-007 did not show the same combination. Because this was secondary, based on only 18 profile-varying tasks, and not multiplicity-adjusted, it is a hypothesis for a future randomized persona experiment rather than evidence that persona conditioning has been validated.

5.5. Implications for Foresight and Decision Support

Foresight methods often value diversity, weak signals, and disagreement rather than consensus alone. A synthetic panel that reproduces a central rank while compressing or randomizing stakeholder differences can remove precisely the information a foresight exercise seeks. The present evidence therefore does not support substituting LLM panels for human experts in policy-facing AHP, Delphi, scenario prioritization, or stakeholder comparison.
Lower-risk uses remain plausible. LLM panels can help test questionnaire wording, identify malformed hierarchies, exercise data pipelines, and generate candidate criteria before human elicitation. In these roles, synthetic outputs are design probes rather than evidence about stakeholder preferences.
In systems terms, the unit these results govern is not a model but a subsystem: a judgment generator embedded upstream of AHP aggregation and alternative selection, whose errors propagate through the weight vector to the final ranking rather than staying local. The validation protocol therefore belongs at the subsystem boundary as a pre-deployment gate. It must be reapplied whenever the underlying model, prompt interface, or corpus changes, because Section 5.2 shows that model updates move performance without moving the validation standard. Table 5 summarizes a minimum validation protocol.

5.6. Limitations

First, the benchmark inherits selective reporting in published AHP studies. Only two holdout tasks contained task-matched individual human matrices or weights, making strong claims about consistency and human dispersion impossible. Second, 15 holdout tasks used one repeated profile because the source papers reported panel composition only at a coarse level. Third, the corpus emphasizes recent open-access studies and is not a probability sample of all AHP applications. The Europe PMC query and English flow are reproducible, and the five literal KCI query strings were recovered from the archived session log. That branch was, however, purposive rather than an exhaustive frame, and commercial search indices are time-dependent, so its result set is not exactly reproducible; screening and extraction also lacked independent double coding. Fourth, source articles may have appeared in model training corpora; nevertheless, low absolute replication performance suggests that exposure did not yield reliable reconstruction. Fifth, the results are specific to model versions, temperature settings, and prompts used at the collection date. Provider-returned immutable revisions, the OpenRouter-served provider, historical SDK versions, and original raw payloads were not logged, limiting exact historical model-call replay. The prospective extension improved this audit trail but was locally time-stamped rather than externally registered, and OpenRouter/NAVER routing remains dependent on third-party infrastructure. Sixth, language-stratified extension results use only six Korean holdout tasks and confound language with domain and model family. Seventh, source validation was completed after response collection. It invalidated two replay tasks and revealed that two larger human panels had only eight historical synthetic slots each. The analysis therefore represents the experiment as implemented, not a fully panel-size-matched re-collection. Eighth, the benchmark evaluates criterion weights, not the complete downstream ranking of alternatives in every source study. Ninth, the randomized-order experiment was conducted later than the registered and extension phases and used the generic-expert instruction throughout, so it is not an exact re-run of those phases; within-batch anchor cells are its only strictly comparable baseline, which is how it is analyzed. The Version 1.2 freeze record carries a date but not a clock time, and the repository commit followed collection, so the prospective status of Versions 1.1 and 1.2 rests on locally time-stamped and hashed records rather than external registration. Tenth, the randomized-order schedules drew eight permutations per task, so realized position exposure was incompletely balanced (mean maximum imbalance 0.29–0.33); shuffled estimates reflect their realized schedules, and future designs should use position-balanced (for example, Latin-square) presentation schedules.
These limitations are also design requirements for follow-up research. A decisive human-likeness study should recruit new experts, preserve individual-level matrices, randomly assign validated profile descriptions, and reserve an untouched task set for final evaluation.

6. Conclusions

Across a source-validated preregistered holdout of 22 published AHP studies, including 21 accuracy tasks, the three registered models showed modest, model-dependent rank replication without a clear absolute-error advantage over uniform weights. A prospective post-registration extension identified a stronger aggregate result for HCX-007, including an advantage over uniform weights, although its individual matrices were often inconsistent. A subsequent diagnostic showed that this benchmark’s fixed-source-order rank evaluation is confounded with criterion presentation order: reference vectors are largely sorted by importance, a model-free descending-order rule outscores every model tested, and HCX-007 is the only one of six models that strongly tracks presentation order. In a second prospective experiment that randomized the order, HCX-007 was the only model whose point-estimate accuracy fell while that of both comparators rose; only the Claude Sonnet 5 improvement was significant after Holm correction, and a low-yield direct-recall probe found no verbatim reproduction of published weights. Persona differentiation was model-dependent, and task-matched individual human data remained too sparse for a confirmatory human-likeness claim.
The practical conclusion is not that LLMs have no role in expert elicitation. It is that their role should remain assistive until external validity, profile differentiation, and decision-level value are demonstrated against transparent baselines. The released corpus, corrected analysis pipeline, and validation checklist provide a reproducible test for future models and a guardrail for LLM-assisted foresight and decision-support research.

Supplementary Materials

The following supporting information can be downloaded at the website of this paper posted on Preprints.org, Table S1: Preregistration adherence and material deviations; Table S2: Persona minus registered baseline accuracy; Table S3: Value of repeated generic-expert calls; Table S4: Registered H5 normalized-entropy difference; Table S5: New minus historical model contrasts; Figure S1: Task-matched median consistency ratios; Note S1: Extension token use, cost, and latency.

Author Contributions

Conceptualization, H.K. and K.C.; methodology, H.K. and K.C.; software, H.K.; validation, H.K.; formal analysis, H.K.; investigation, H.K.; resources, H.K.; data curation, H.K.; writing—original draft preparation, H.K.; writing—review and editing, H.K. and K.C.; visualization, H.K.; supervision, K.C.; project administration, K.C. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Institutional Review Board Statement

Not applicable. The study reanalyzed only previously published, deidentified aggregate and individual expert-judgment data reported in the cited source articles; no new human-subjects data were collected and no individual participants were identifiable.

Data Availability Statement

The public deposit in the OSF project at https://osf.io/wa8bt/ (DOI: 10.17605/OSF.I O/WA8BT) contains the structured corpus records for all 29 collected sources; the complete parsed response records of the three registered models, whose historical runner never stored prompts or provider envelopes; publicly redacted response records for all 16,281 extension cells and all 1,134 randomized-order cells, retaining judgments, weights, consistency statistics, per-call presentation permutations, prompt hashes, and request metadata; the frozen Version 1.1 and 1.2 protocols with their pre-collection SHA-256 freeze records; the row-level screening log; all collection, analysis, and redaction code with an exact dependency lockfile; and a single-command script (run_all.sh) that regenerates every reported analysis, including analysis_reference/probe_analysis.json, from the public records alone. The confirmatory hypotheses and analysis plan were preregistered before the start of data collection in a public OSF registration (https://osf.io/39jsa). The canonical archive is osf_deposit_v2_3_full_repro.zip (SHA-256 6a0d9bf98594812df50b47db98c9b311b59a1c61434b22290a7976d9ba73ff08); it supersedes the earlier archives in the same project, and its end-to-end regeneration was verified in a clean-room directory, reproducing the archived analysis outputs exactly. STUDY_SELECTION_AUDIT.md documents the search and full-text decisions, and MODEL_CALL_REPRODUCIBILITY.md reproduces the prompt and repair protocol and states which provider metadata were not historically retained. Code is released under the MIT license, and our data and documentation under CC BY 4.0; short quoted excerpts from the cited source articles (criterion names, definitions, and panel descriptions) remain under their original copyright and are excluded from the CC BY grant. Private raw envelopes are retained by the authors for audit; the deposits are archive files within the OSF project, not additional OSF registrations.

Acknowledgments

During the preparation of this study, the authors used OpenAI Codex [48] for statistical code review, initial drafting of figure captions and of Sections 3.4–3.7, 4, and 5, and language editing; the service did not expose immutable model version identifiers. The authors have reviewed and edited all output, reran the analyses, verified citations and numerical results, and take full responsibility for the content of this publication.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Argyle, L.P.; Busby, E.C.; Fulda, N.; Gubler, J.R.; Rytting, C.; Wingate, D. Out of One, Many: Using Language Models to Simulate Human Samples. Political Anal. 2023, 31, 337–351. [Google Scholar] [CrossRef]
  2. Bisbee, J.; Clinton, J.D.; Dorff, C.; Kenkel, B.; Larson, J.M. Synthetic Replacements for Human Survey Data? The Perils of Large Language Models. Political Anal. 2024, 32, 401–416. [Google Scholar] [CrossRef]
  3. Wang, A.; Morgenstern, J.; Dickerson, J.P. Large Language Models That Replace Human Participants Can Harmfully Misportray and Flatten Identity Groups. Nat. Mach. Intell. 2025, 7, 400–411. [Google Scholar] [CrossRef]
  4. Wang, Y.; Zhao, J.; Ones, D.S.; He, L.; Xu, X.; et al. Evaluating the Ability of Large Language Models to Emulate Personality. Sci. Rep. 2025, 15, 519. [Google Scholar] [CrossRef] [PubMed]
  5. Saaty, T.L. The Analytic Hierarchy Process; McGraw-Hill: New York, NY, USA, 1980. [Google Scholar]
  6. Saaty, T.L. How to Make a Decision: The Analytic Hierarchy Process. Eur. J. Oper. Res. 1990, 48, 9–26. [Google Scholar] [CrossRef]
  7. Memduhoglu, A.; Fulman, N.; Polat, N.; Atas, T. Large Language Models as Virtual Experts? Evaluating AHP-Based Criteria Weighting Performance for Solar Power Plant Site Selection. Expert Syst. With Appl. 2026, 299, 130171. [Google Scholar] [CrossRef]
  8. Fares, A.; Yu, J.; Taiwo, R.; Faris, N.; Zayed, T.; Miranda-Moreno, L. Leveraging Large Language Models as Synthetic Experts: A Methodology and Validation Framework for Pavement Performance Assessment. Results Eng. 2026, 31, 111630. [Google Scholar] [CrossRef]
  9. Kim, H.; Cho, K. Distributional Validity and Calibration of a Korean Synthetic Persona Panel for Digital and AI Service Use: A Secondary-Data Validation Against the Korea Media Panel Survey. SocArXiv 2026. [Google Scholar] [CrossRef]
  10. Dillion, D.; Tandon, N.; Gu, Y.; Gray, K. Can AI Language Models Replace Human Participants? Trends Cogn. Sci. 2023, 27, 597–600. [Google Scholar] [CrossRef] [PubMed]
  11. Doshi, A.R.; Hauser, O.P. Generative AI Enhances Individual Creativity but Reduces the Collective Diversity of Novel Content. Sci. Adv. 2024, 10, eadn5290. [Google Scholar] [CrossRef] [PubMed]
  12. Kampourakis, V.; Kavallieratos, G.; Spathoulas, G.; Gkioulos, V.; Katsikas, S. LLM-Assisted AHP for Explainable Cyber Range Evaluation. In Proceedings of the International Workshop on Engineering and Cybersecurity of Critical Systems (EnCyCriS), co-located with ICSE, 2026. [Google Scholar] [CrossRef]
  13. Chen, D.F.; Chen, Y.H.; Chen, B.S. Reproducible Expert Weight Elicitation via LLM Multi-Agent Simulation: A Best–Worst Method Decision Support Framework for AI-Driven E-Commerce Platform Evaluation. Appl. Sci. 2026, 16, 6093. [Google Scholar] [CrossRef]
  14. Lee, N.; Kwon, O. Differentiated Effects of Agent Diversity on Collective Decision-Making in LLM-Based Multi-Agent Delphi Systems. Appl. Sci. 2026, 16, 6715. [Google Scholar] [CrossRef]
  15. Zhu, L.; Xie, Y.; Xiang, N.; Chen, G. An Explainable HCI-Based Decision Support Framework for Human–AI Co-Design. Appl. Sci. 2026, 16, 4007. [Google Scholar] [CrossRef]
  16. Reinholtz, K.; Shahroudi, K.E.; Lawrence, S. LLM-Powered, Expert-Refined Causal Loop Diagramming via Pipeline Algebra. Systems 2025, 13, 784. [Google Scholar] [CrossRef]
  17. Parnell, G.S.; Kenley, C.R.; Clark, D.; Smith, J.; Salvatore, F.; Nwobodo, C.; Davis, S. Decision Analysis Data Model for Digital Engineering Decision Management. Systems 2025, 13, 596. [Google Scholar] [CrossRef]
  18. Basak, I. Incorporating Within-Pair Order Effects in the Analytic Hierarchy Process. Math. Comput. Model. 1993, 17, 83–92. [Google Scholar] [CrossRef]
  19. Webber, S.A.; Apostolou, B.; Hassell, J.M. The Sensitivity of the Analytic Hierarchy Process to Alternative Scale and Cue Presentations. Eur. J. Oper. Res. 1997, 96, 351–362. [Google Scholar] [CrossRef]
  20. Pezeshkpour, P.; Hruschka, E. Large Language Models Sensitivity to the Order of Options in Multiple-Choice Questions. Proc. Find. Assoc. Comput. Linguist. NAACL 2024, 2024, 2006–2017. [Google Scholar] [CrossRef]
  21. Manik, M.H. Addressing the Supplier Selection Problem by Using the Analytical Hierarchy Process. Heliyon 2023, 9, e17997. [Google Scholar] [CrossRef] [PubMed]
  22. Hong, J. An AHP Approach for the Importance Weight of Renewable Energy Investment Criterion in the Private Sector. Korean Energy Econ. Rev. 2011, 10, 115–142. [Google Scholar] [CrossRef]
  23. Zhou, C.; Cao, W.; Huang, Q.; Lin, X.; Chen, J.; Zheng, Y.; Li, W.; Wang, Z.; Jiang, M. Identifying Key Factors for Organizational Resilience Among Medical Alliance Using the Analytic Hierarchy Process Method. Sci. Rep. 2025, 15, 29728. [Google Scholar] [CrossRef] [PubMed]
  24. Karlsson, C.S.J.; Kalantari, Z.; Mortberg, U.; Olofsson, B.; Lyon, S.W. Natural Hazard Susceptibility Assessment for Road Planning Using Spatial Multi-Criteria Analysis. Environ. Manag. 2017, 60, 823–851. [Google Scholar] [CrossRef] [PubMed]
  25. Kadoic, N.; Simic, D.; Mesaric, J.; Redep, N.B. Measuring Quality of Public Hospitals in Croatia Using a Multi-Criteria Approach. Int. J. Environ. Res. Public Health 2021, 18, 9984. [Google Scholar] [CrossRef] [PubMed]
  26. Park, J.; Yoo, H.; Park, M. A Study on Assessment Items Analysis for Eco-Corridors’ Area Using the Analytic Hierarchy Process. J. Environ. Impact Assess. 2009, 18, 301–312. [Google Scholar]
  27. Oh, D.G.; Yeo, J.S.; Choi, S.Y. A Study on the Development of Evaluation Measures and Indicators for Foreign Research Information Centers. J. Korean Libr. Inf. Sci. Soc. 2012, 43, 99–116. [Google Scholar] [CrossRef]
  28. Kim, H.; Lee, J. The Role of Social Enterprise in Creating Women-Friendly Safe Cities: An AHP Analysis of Safe-City Design. Crisisonomy 2012, 8, 85–104. [Google Scholar]
  29. Lee, M.; Kim, M.; Lee, S. Weight Setting of Major Environmental Assessment Items Using Analytical Hierarchy Process: Case for the Selection of Railroad Route. J. Environ. Impact Assess. 2014, 23, 517–526. [Google Scholar] [CrossRef]
  30. Park, D. A Study on the Pre-Feasibility Assessment Using the AHP: Focusing on the Case of Public Project. J. Korea Soc. Comput. Inf. 2015, 20, 163–168. [Google Scholar] [CrossRef]
  31. Lim, S.M.; Lim, S.U. An Exploratory Study on the Weighting of Career Level Diagnostic Factors for Employees Using AHP. J. Korean Soc. Qual. Manag. 2026, 54, 423–438. [Google Scholar] [CrossRef]
  32. Shoman, H.; Almeida, N.D.; Tanzer, M. Ranking Decision-Making Criteria for Early Adoption of Innovative Surgical Technologies. JAMA Netw. Open 2023, 6, e2343703. [Google Scholar] [CrossRef] [PubMed]
  33. Kang, J.; Choi, J.; Lee, K. Development of an Evaluation Index for Forest Therapy Environments. Int. J. Environ. Res. Public Health 2024, 21, 136. [Google Scholar] [CrossRef] [PubMed]
  34. Quttainah, M.; Mishra, V.; Madakam, S.; Lurie, Y.; Mark, S. Cost, Usability, Credibility, Fairness, Accountability, Transparency, and Explainability Framework for Safe and Effective Large Language Models in Medical Education. JMIR AI 2024, 3, e51834. [Google Scholar] [CrossRef] [PubMed]
  35. Habsi, T.A.; Al-Khusaibi, S.; Hashmi, D.A.; Al-Jabri, A.; Al-Rahbi, A.; Almamari, N.; Alkindi, T. Sustainability in the Public Healthcare Sector: Insights From an Analytical Hierarchy Process Analysis. Cureus 2024, 16, e61672. [Google Scholar] [CrossRef] [PubMed]
  36. Ahmed, I.; Iswara, A.P.; Abbas, S.; Jamal, F.Q.; Ahmad, I.; Shah, S.T.H.; Naseem, A. Modelling and Optimization of an Existing Onshore Gas Gathering Network Using PIPESIM. Heliyon 2024, 10, e35006. [Google Scholar] [CrossRef] [PubMed]
  37. Sengul, H.; Aydin, S.; Onder, E.; Bulut, A. Developing an Index to Measure the Value of Health Care Provided to Stroke Patients. Ann. Indian Acad. Neurol. 2024, 27, 695–705. [Google Scholar] [CrossRef] [PubMed]
  38. Zeki, O.; Guner, P.; Zeki, M.; Ozkaynak, M. Prioritizing Delirium Risk Factors in Nursing: A Cross-Sectional Study Using the Analytic Hierarchy Process. BMC Nurs. 2025, 24, 996. [Google Scholar] [CrossRef] [PubMed]
  39. Shu, X.; Zhang, P.; Li, M.; Ge, Y.; Tang, H.; Wang, Q. Critical Factors Analysis of Student-Athletes Learning Training Contradiction via AHP. Sci. Rep. 2025, 15, 31812. [Google Scholar] [CrossRef] [PubMed]
  40. Podlesnik, L.; Bernik, I.; Mihelic, A. Integrating CTI and Threat Modeling for Cyber Resilience: An AHP Assessment. PLoS ONE 2025, 20, e0335154. [Google Scholar] [CrossRef] [PubMed]
  41. Singhdeo, A.K.; Tripathy, S.; Singhal, D. Decoding E-Waste Challenges with Hybrid AHP and ISM Model Approach. F1000Research 2025, 14, 856. [Google Scholar] [CrossRef] [PubMed]
  42. Musyafa, A. A Strategic Roadmap for Construction Automation in Indonesian Mass Housing Projects. Sci. Rep. 2026, 16, 15419. [Google Scholar] [CrossRef] [PubMed]
  43. Wang, X.; He, L.; Zhu, K.; Zhang, S.; Xin, L.; Xu, W.; Guan, Y. An Integrated Model to Evaluate the Impact of Social Support on Improving Self-Management of Type 2 Diabetes Mellitus. BMC Med. Inform. Decis. Mak. 2019, 19, 197. [Google Scholar] [CrossRef] [PubMed]
  44. Morgan, C.N.; Wallace, R.M.; Vokaty, A.; Seetahal, J.F.R.; Nakazawa, Y.J. Risk Modeling of Bat Rabies in the Caribbean Islands. Trop. Med. Infect. Dis. 2020, 5, 35. [Google Scholar] [CrossRef] [PubMed]
  45. Li, H.; Chen, X.; Fang, Y. The Development Strategy of Home-Based Exercise in China Based on the SWOT-AHP Model. Int. J. Environ. Res. Public Health 2021, 18, 1224. [Google Scholar] [CrossRef] [PubMed]
  46. Satpathy, S.; Patel, G.; Kumar, K. Identifying and Ranking Techno-Stressors Among IT Employees Due to Work From Home Arrangement During the COVID-19 Pandemic. Decision 2021, 48, 391–402. [Google Scholar] [CrossRef]
  47. Varolgunes, F.K.; Celik, F.; del Rio-Rama, M.; Alvarez-Garcia, J. Reassessment of Sustainable Rural Tourism Strategies After COVID-19. Front. Psychol. 2022, 13, 944412. [Google Scholar] [CrossRef] [PubMed]
  48. OpenAI. Codex. 2026. Available online: https://openai.com/codex/get-started/ (accessed on 29 July 2026).
Figure 1. Study-selection flow. The KCI branch was a purposive language-diversity supplement; its five literal query strings were recovered from the archived session log and are reproduced in the selection audit. The row-level verification decisions and reasons are released in analysis_reference/study_screening_log.csv in the OSF deposit.
Figure 1. Study-selection flow. The KCI branch was a purposive language-diversity supplement; its five literal query strings were recovered from the archived session log and are reproduced in the selection audit. The row-level verification decisions and reasons are released in analysis_reference/study_screening_log.csv in the OSF deposit.
Preprints 231494 g001
Figure 2. Primary accuracy results. Error bars are study-clustered 95% confidence intervals. No top-1 advantage over chance and no MAE advantage over uniform weights were estimated with an interval excluding zero.
Figure 2. Primary accuracy results. Error bars are study-clustered 95% confidence intervals. No top-1 advantage over chance and no MAE advantage over uniform weights were estimated with an interval excluding zero.
Preprints 231494 g002
Figure 3. Persona audit. (a) Most holdout tasks repeated a single profile text. (b) In pooled tasks with multiple profile texts, between-profile signal was smaller than within-profile sampling variation. Panel (b) is exploratory.
Figure 3. Persona audit. (a) Most holdout tasks repeated a single profile text. (b) In pooled tasks with multiple profile texts, between-profile signal was smaller than within-profile sampling variation. Panel (b) is exploratory.
Preprints 231494 g003
Figure 4. Randomized-order experiment. Paired change in mean Spearman rank correlation from source-order (anchor) to randomized-order (shuffled) presentation for each extension model, with study-clustered 95% confidence intervals. p-values are Holm-adjusted two-sided Wilcoxon signed-rank tests on task-level paired differences. The dashed line marks the model-free descending-order rule (0.735).
Figure 4. Randomized-order experiment. Paired change in mean Spearman rank correlation from source-order (anchor) to randomized-order (shuffled) presentation for each extension model, with study-clustered 95% confidence intervals. p-values are Holm-adjusted two-sided Wilcoxon signed-rank tests on task-level paired differences. The dashed line marks the model-free descending-order rule (0.735).
Preprints 231494 g004
Table 1. Corpus and Analysis Sets
Table 1. Corpus and Analysis Sets
Pilot Holdout Combined
Published studies 5 22 27
AHP tasks 14 22 36
Tasks with reference weights 14 21 35
Source-validated panel slots 230 334 564
Historically collected analysis slots 230 304 534
Tasks with multiple profile texts 11 7 18
Table 2. Preregistered Holdout Accuracy (21 Tasks; Study-Clustered 95% CIs)
Table 2. Preregistered Holdout Accuracy (21 Tasks; Study-Clustered 95% CIs)
Model Mean ρ Mean τ b Top-1 Chance top-1 MAE Uniform MAE Δ MAE (LLM−uniform)
Gemini 0.27 [0.05,0.49] 0.24 [0.06,0.42] 0.29 [0.10,0.48] 0.262 0.140 [0.107,0.174] 0.143 [0.109,0.181] − 0.003 [ − 0.052 ,0.039]
GPT-5-mini 0.16 [ − 0.13 ,0.45] 0.17 [ − 0.11 ,0.44] 0.38 [0.19,0.57] 0.262 0.149 [0.112,0.187] 0.143 [0.109,0.181] +0.006 [ − 0.046 ,0.052]
K-EXAONE 0.15 [ − 0.14 ,0.44] 0.14 [ − 0.13 ,0.41] 0.43 [0.24,0.62] 0.262 0.127 [0.096,0.161] 0.143 [0.109,0.181] − 0.016 [ − 0.053 ,0.016]
Table 3. Prospective Extension: Holdout Accuracy and Implementation Audit
Table 3. Prospective Extension: Holdout Accuracy and Implementation Audit
Model Valid cells Mean ρ Mean τ b Top-1 MAE Median CR
Gemini 3.6 Flash 4,806/4,806 0.312 [0.048,0.554] 0.267 [0.030,0.486] 0.429 [0.238,0.619] 0.140 [0.104,0.178] 0.011
Claude Sonnet 5 4,806/4,806 0.186 [ − 0.071 ,0.425] 0.158 [ − 0.066 ,0.368] 0.238 [0.095,0.429] 0.153 [0.113,0.200] 0.008
HCX-007 4,797/4,806 0.694 [0.523,0.841] 0.621 [0.450,0.776] 0.762 [0.571,0.952] 0.101 [0.074,0.129] 0.136
Accuracy columns: persona condition on the 21 reference-weight holdout tasks. Valid cells and
median CR: implementation audit over all 4,806 in-scope cells per model (three conditions, 36 tasks).
Table 4. Randomized-Order Experiment (21 Holdout Tasks; Study-Clustered 95% CIs)
Table 4. Randomized-Order Experiment (21 Holdout Tasks; Study-Clustered 95% CIs)
Model Anchor ρ Shuffled ρ Anchor − shuffled Holm p Tasks degraded by shuffling
Gemini 3.6 Flash 0.269 [0.01,0.52] 0.450 [0.26,0.63] − 0.181 [ − 0.38 ,0.01] 0.160 3 of 21
Claude Sonnet 5 0.110 [ − 0.17 ,0.37] 0.566 [0.36,0.75] − 0.455 [ − 0.76 , − 0.20 ] 0.005 2 of 21
HCX-007 0.491 [0.20,0.75] 0.323 [0.05,0.58] + 0.168 [ − 0.14 ,0.44] 0.160 10 of 21
Positive differences indicate that randomizing presentation order reduced accuracy. p values: two-sided
Wilcoxon signed-rank tests on task-level paired differences (21 tasks), Holm-adjusted across the three model tests.
Table 5. Minimum Validation Checks for Synthetic Expert Panels
Table 5. Minimum Validation Checks for Synthetic Expert Panels
Check Requirement
External target Separate aggregate consensus from individual judgments.
Baselines Include uniform, random/prior, single-call, and generic ensembles.
Analysis unit Aggregate at task level and preserve study/panel clustering.
Presentation order Randomize criterion order per call with position-balanced (Latin-square) schedules; report a model-free descending-order baseline.
Pilot separation Exclude design-stage observations from confirmatory tests.
Profile audit Report nominal slots, unique profile texts, and profile reuse.
Repeatability Decompose profile effects and repeated-draw noise.
Decision impact Test whether errors alter downstream alternative rankings.
Versioning Archive prompts, revisions, routes, timestamps, payloads, and code.
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.