Preprint
Article

This version is not peer-reviewed.

A FAIR Multi-Assay Dataset and Reproducibility Audit of Mucosal and Systemic Immune Responses After Oral Hyperimmune Anti-HIV-1 gp120 IgY Administration in Outbred Cats

Submitted:

12 August 2026

Posted:

13 August 2026

You are already at the latest version

Abstract
Open, provenance-aware datasets are essential for evaluating biological findings and identifying analytical decisions that alter interpretation. This Data Resource Article describes a curated multi-assay preclinical dataset derived from studies of oral hyperimmune anti-human immunodeficiency virus type 1 (HIV-1) gp120 immunoglobulin Y (IgY) administration in outbred domestic cats. The publicly deposited Figshare resource comprises 13 comma-separated-value tables, a structured Excel workbook, a machine-readable data dictionary, provenance and quality-control records, executable Python and R scripts, a manifest and cryptographic checksums. Two independent cohorts are represented. A six-cat proof-of-concept cohort (three immunised and three controls) generated serum anti-gp120 anti-anti-idiotypic antibody (Ab3) ELISA data, competitive-inhibition measurements and processed TZM-bl HIV-1 JR-FL pseudovirus neutralisation outputs, including RLU summaries, virus-only and cell-only controls, percentage neutralisation and an estimated ID50. A separate 42-cat mucosal-immunogenicity cohort (18 immunised and 24 controls) generated salivary anti-gp120 IgA classifications after eight weeks of assigned oral exposure. Exact source values were transcribed with explicit provenance, whereas unavailable primary measurements were represented by schema-defined blanks rather than imputed observations. Reanalysis reproduced complete separation in IgA positivity (18/18 versus 0/24; Fisher exact p = 2.83 × 10−12; absolute risk difference 100%) and identified decision-sensitive analytical features. Displayed Ab3 negative wells yielded a mean-plus-three-standard-deviations threshold of 0.146, compared with source thresholds of 0.32 and 0.35. Animal-level competitive-inhibition means produced p = 0.0118 by Welch’s t-test and p = 0.10 by an exact Mann-Whitney test. The resource supports immunological reuse, assay benchmarking and transparent sensitivity analysis while preserving clear boundaries between reported, processed, derived and unavailable primary measurements. It should be interpreted as an exploratory FAIR-oriented data resource rather than evidence of vaccine efficacy or broadly neutralising immunity.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  ;  ;  ;  ;  ;  

1. Introduction

Research data are no longer a secondary by-product of experimental science. They are a primary scholarly output that permits verification, alternative analysis, integration across studies and development of new statistical and computational methods. The FAIR principles - Findability, Accessibility, Interoperability and Reusability - provide a widely adopted framework for making digital research objects useful to both humans and machines [1]. Complementary open-science frameworks emphasise persistent identifiers, community standards, transparent data policies and explicit reuse conditions as practical mechanisms for improving reproducibility and data reuse [2,3]. FAIRness does not imply that every dataset is complete or free of uncertainty; rather, uncertainty, provenance, structure and reuse conditions should be made explicit. This distinction is especially important in small preclinical studies, in which sparse biological replication, multiple technical measurements and exploratory endpoints can produce apparently strong findings whose interpretation depends heavily on the analytical unit, threshold definition or handling of missing primary records [4,5,6].
Animal immunology datasets frequently contain nested observations: animals are the biological units, while wells, repeated reads, assay plates and dilution points are technical units. Failure to distinguish these levels can produce pseudoreplication and p-values that do not reflect the number of independently assigned animals [6,7]. Conversely, retaining only group means prevents assessment of assay variation and data-processing decisions. A high-value resource should therefore preserve the hierarchy between biological and technical replication, retain controls, distinguish raw, processed and derived fields, and identify unavailable elements. Such structuring is consistent with recommendations for open research culture and reproducible analytical workflows [3,4].
Immunoglobulin Y is the principal circulating antibody of birds and is transferred from serum to egg yolk. It can be generated in substantial quantities without repeated terminal bleeding and has physicochemical properties that differ from mammalian IgG, including limited interaction with mammalian Fc receptors and complement [8,9]. These features have motivated applications in diagnostics, passive immunotherapy and oral control of infectious agents. Oral IgY research has generally focused on local neutralisation within the gastrointestinal tract; however, the extent to which orally administered antigen-specific IgY might also trigger endogenous mucosal or systemic responses remains less well characterised. Interpretation is complicated because orally delivered antibody can act as a passive reagent, a source of antigen-antibody complexes, or a potential idiotypic immunogen, while degradation and variable consumption can influence actual exposure.
The idiotype network hypothesis proposes that antigen-binding regions of antibodies contain immunogenic determinants capable of eliciting anti-idiotypic antibodies; a subsequent anti-anti-idiotypic response may recognise features related to the original antigen [10,11]. Although the strongest structural form of the internal-image hypothesis requires direct epitope-level evidence, modern work has renewed interest in anti-idiotypic constructs as antigen surrogates and immunomodulatory tools. Recent experimental work has shown that selected anti-HIV anti-idiotype formats can elicit HIV-reactive humoral responses in animals, supporting continued investigation while also emphasising the need for binding, competition and functional assays to be interpreted together [12].
HIV-1 remains an exceptionally difficult vaccine target because envelope diversity, conformational masking, glycan shielding and rapid escape limit the breadth of conventional antibody responses. Mucosal immunity is particularly relevant because many HIV acquisitions occur across mucosal surfaces. Secretory IgA contributes to immune exclusion, antigen sampling and homeostatic interactions at mucosal interfaces [13,14,15], and meta-analysis confirms that several HIV vaccine platforms can generate measurable mucosal antibody responses in humans, although their magnitude and functional meaning vary by regimen and compartment [16]. Functional neutralisation remains a separate endpoint. The TZM-bl luciferase reporter assay is a standard platform for quantifying antibody-mediated inhibition of Env-pseudotyped HIV-1 infection, provided that dilution series, controls and curve-fitting criteria are fully available [17,18]. Neutralisation against one strain cannot be equated with breadth, and a single estimated 50% inhibitory dilution does not establish broadly neutralising activity [19,20].
The source preclinical investigation evaluated oral administration of hyperimmune anti-HIV-1 gp120 IgY in cats and reported three complementary response domains: serum anti-gp120 Ab3 reactivity, competitive inhibition and JR-FL pseudovirus neutralisation in a six-cat proof-of-concept cohort; and salivary anti-gp120 IgA responses in a separate 42-cat cohort [21]. The public dataset associated with that work was subsequently curated and deposited in Figshare [22]. The resource intentionally distinguishes exact values transcribed from the source from unavailable primary measurements. It therefore functions both as an immunological dataset and as a reproducibility case study.
The objectives of this paper were to: (i) describe the biological design, file architecture and provenance model of the deposited dataset; (ii) reproduce key reported summaries using transparent code; (iii) quantify technical variation where well-level values were available; (iv) examine the sensitivity of interpretation to cut-off construction, statistical test selection and zero-cell corrections; (v) identify data elements that prevent complete independent reconstruction; and (vi) define appropriate reuse scenarios and boundaries. The paper does not attempt to establish vaccine efficacy. Instead, it treats transparent limitations as part of the scientific value of the data resource.

2. Materials and Methods

2.1. Study Design and Experimental Cohorts

The resource represents two independent feline cohorts and a donor-hen component. The proof-of-concept cohort comprised six adult domestic cats aged approximately 2-3 years. Three animals were assigned to oral hyperimmune egg preparation and three to a matched preparation derived from non-immunised hens. The immunised animals received 2 mL of hyperimmune egg preparation diluted in 10 mL of soy-based milk substitute daily for ten consecutive weeks; controls received the corresponding non-hyperimmune preparation. Serum collected at week 12 was used for anti-gp120 Ab3 ELISA, competitive-inhibition ELISA and TZM-bl neutralisation. Samples were coded before laboratory testing, and the source methods state that analysts performing ELISA and inhibition testing were blinded to allocation [21].
The mucosal-immunogenicity cohort comprised 42 clinically healthy, outbred male cats aged 2-3 years and weighing approximately 3.0-5.5 kg. Eighteen animals were assigned hyperimmune anti-gp120 egg-yolk preparation mixed with soy milk, while 24 received non-hyperimmune egg-yolk preparation mixed with soy milk. Exposure lasted eight weeks. The preparation was offered ad libitum and individual consumption was not quantified; consequently, the dataset records assigned exposure but does not support dose-response analysis. Saliva collected at the end of exposure was tested by qualitative anti-gp120 IgA ELISA. The cohorts were assigned distinct identifiers and were not merged at animal level because the 42 cats did not generate the reported systemic assays and the six proof-of-concept cats were not the source of the group-level salivary IgA dataset.
Four laying hens were used to produce anti-gp120 IgY. The source materials contain two descriptions of the immunogen: a gp120 254-274 peptide conjugated to keyhole limpet haemocyanin in the workflow figure and peptide-focused methods, and recombinant HIV-1 IIIB gp120 in the later preclinical-immunogenicity section. This inconsistency was preserved in the quality-issues record rather than silently reconciled. Hens underwent primary immunisation and booster administration; eggs were collected during the high-titre period, and IgY was isolated from yolk using water-dilution and precipitation procedures based on Polson and colleagues [23].
Table 1. Experimental components represented in the data resource.
Table 1. Experimental components represented in the data resource.
Component Independent units Exposure and timing Data products
Donor-hen component Four laying hens Primary immunisation plus boosters; longitudinal sampling to week 12 IgY ELISA kinetics, geometric mean titre, concentration and SDS-PAGE summary
Proof-of-concept feline cohort Three immunised and three control cats Fixed daily oral administration for 10 weeks; serum at week 12 Ab3 ELISA wells, animal-level inhibition summaries, specificity controls and neutralisation summaries
Mucosal feline cohort 18 immunised and 24 control cats Assigned ad-libitum exposure for 8 weeks; saliva at end of exposure Aggregate IgA categories, group summaries and a schema for unavailable individual duplicate wells

2.2. Immunogen, IgY Production and Characterisation

The targeted region was the HIV-1 gp120 sequence spanning amino-acid residues 254-274, described in the source as a relatively conserved region linked to infectivity and antibody recognition. The historical rationale for this target derives from evidence that the second conserved domain of gp120 contributes to HIV infectivity and can be recognised by neutralising antibodies [24]. Hyperimmune egg-yolk preparations were evaluated by indirect ELISA. The deposited table records mean OD450 and standard error at pre-immune baseline and weeks 2, 6 and 12. The reported geometric mean endpoint titre was 1:1024. Purified IgY concentration was reported as 8.65 mg/mL, and SDS-PAGE showed bands consistent with IgY heavy and light chains.
The data resource does not create individual hen-level titres where the source displayed only time-point summaries. Each row is labelled as an exact reported summary, with source location and status fields. This approach prevents false granularity while retaining a schema that can be expanded if original hen-level records are recovered.

2.3. Ab3 ELISA and Competitive-Inhibition Assays

For the Ab3 ELISA, microplates were coated with gp120 254-274 peptide, blocked, incubated with feline sera and developed with horseradish-peroxidase-conjugated anti-cat IgG and tetramethylbenzidine. The source figure displayed eight OD450 values for each of three immunised cats, three controls, a primary positive control, a repeated positive control and blank wells. All 72 displayed measurements were transcribed into a well-level table. Each record contains plate identifier, sample identifier, sample role, well position, replicate index, OD450, provenance class and source location.
For each sample, the arithmetic mean, sample standard deviation and coefficient of variation (CV = 100 × SD/mean) were recalculated. The source used a positivity threshold of 0.35 in the main text and caption but displayed 0.32 in the figure annotation. To examine whether either threshold could be reproduced from the displayed negative wells, the mean plus three sample standard deviations was calculated using all 24 wells assigned to the three negative-control cats. This calculation was treated as an audit of the displayed data, not as proof that the original validation population was identical.
The competitive-inhibition assay assessed whether feline serum interfered with anti-gp120 IgY binding to immobilised gp120 peptide. Percentage inhibition was defined as 100 × (ODuninhibited - ODsample)/ODuninhibited. The dataset contains exact animal-level means and standard errors shown for three immunised and three control cats, as well as specificity results at a reciprocal dilution of 1000 for gp120 peptide, Ab3-Ab2 interaction, a non-specific peptide and an unrelated control antibody. Because only summary means were available, technical-replicate reconstruction was not attempted.

2.4. TZM-bl HIV-1 Neutralisation Data

The source experiment used clade B HIV-1 JR-FL pseudovirus and TZM-bl cells. In this assay, neutralisation is quantified from the reduction in Tat-regulated luciferase expression relative to virus-only and cell-only controls [17,18]. VRC01 served as a positive reference antibody. The deposited numerical records include the processed neutralisation information available from the source: the exact RLU summary table displayed for a reciprocal dilution of 40,000, virus-only and cell-only control values, the reported maximum neutralisation of 86.3%, the estimated immunised-group ID50 of 3.2 × 10³ with a reported 95% confidence interval of 2.1-4.8 × 10³, and the statement that controls did not reach 50% neutralisation. Assay metadata describing the pseudovirus, cell system, incubation conditions, control structure and analytical approach are retained in the resource.
The complete well-by-well luminometer export and full numerical dilution series were not available in the source material used for curation. The repository therefore distinguishes the processed neutralisation data that are available from the absent proprietary instrument-level export. A schema-defined table preserves fields for biological sample, dilution, technical replicate, sample RLU, virus-control RLU, cell-control RLU, calculated percentage neutralisation, assay date and curve-quality flag so that any future recovery of original instrument files can be incorporated without changing the data model. No plotted points were digitised or reverse-engineered, because doing so could create pseudo-precision and be mistaken for primary instrument measurements. This approach preserves provenance while still making the reported neutralisation calculations, controls, ID50 estimate and assay metadata available for secondary use.

2.5. Salivary Anti-gp120 IgA Data

Saliva was tested in duplicate by indirect ELISA on gp120-coated plates with an anti-feline IgA detection antibody. The source defined the primary cut-off as the negative-control mean plus three standard deviations and reported 0.13 as the positivity boundary, with 0.13-0.15 classified as borderline. Results were divided into negative, borderline, low-positive, moderate-positive and high-positive categories.
The deposited aggregate table records 20 negative and four borderline control cats, with no positive controls; the immunised group comprised six low-positive, six moderate-positive and six high-positive cats. Individual OD450 values and their mapping to anonymous animal identifiers were not available numerically. The resource therefore provides 42 anonymous registry rows and an empty duplicate-well template, but it does not assign the plotted points or category labels to particular animals. This preserves the difference between known group counts and unknown individual records.

2.6. Data Curation, Provenance and File Architecture

The dataset was curated as a provenance-preserving public data resource and deposited in Figshare under DOI 10.6084/m9.figshare.33179345 [22]. Numerical values explicitly printed in text, tables or figure panels were labelled reported_exact; group counts and summaries without individual mapping were labelled reported_exact_aggregate; quantities calculated solely from exact reported values were labelled derived_from_reported; and details supplied as study-level clarification were labelled author_clarification. Expected but unavailable primary measurements were represented as template_missing and left blank rather than imputed. Conflicting source statements were recorded as inconsistent_report. This explicit provenance model follows the FAIR principle that reusable data should be richly described and linked to their origin, while avoiding false granularity [1,2].
Files are stored in non-proprietary CSV form and mirrored in a formatted Excel workbook. The package includes a README, machine-readable data dictionary, citation metadata, manifest, SHA-256 checksums, processed HIV-1 neutralisation outputs, calculated ID50 information, virus-only and cell-only controls, assay metadata, statistical analysis outputs, quality-control documentation, a Python analysis script and an R analysis script. Missing values are blank and are to be imported as NA, never as zero. Anonymous identifiers describe cohort and allocation but contain no direct personal or owner information. The persistent Figshare DOI supports findability and stable citation [22]; standard tabular formats and documented metadata support interoperability; and provenance fields, quality-control records and executable code support transparent reuse [1,2,3]. Because executable scripts are themselves research objects, the inclusion and documentation of analysis code also accords with contemporary FAIR principles for research software [25].
Table 2. Principal deposited data records and intended use.
Table 2. Principal deposited data records and intended use.
Record Granularity Status Primary reuse
Animal registry One row per anonymous cat Design-level metadata; some individual fields unavailable Cohort definition, allocation checks and linkage to recovered primary data
Hen IgY kinetics One row per reported time point Exact aggregate values Kinetic visualisation and comparison with future IgY production studies
Ab3 ELISA wells One row per displayed ELISA well Exact displayed values Technical variability, cut-off sensitivity and plate-summary reconstruction
Inhibition by animal One row per proof-of-concept cat Exact animal-level summaries Experimental-unit analysis and effect estimation
Processed neutralisation and RLU summary One row per displayed condition plus reported ID50 summary Exact processed/source-reported outputs Verification of displayed calculation, assay-control review and ID50 benchmarking
Mucosal IgA categories One row per response category Exact aggregate counts Exact binomial inference and scenario analysis
Primary-data extension templates Schema-defined blank records Instrument-level values not represented in current release Versioned incorporation of original exports if subsequently recovered
Quality issues One row per identified issue Curated audit Reproducibility review and prioritisation of source-data recovery

2.7. Statistical and Reproducibility Analyses

Analyses were performed with the executable Python script included in the repository. Ab3 ELISA summaries were recalculated directly from the 72 well-level values. For salivary IgA, borderline results were treated as not positive in the primary binary analysis, matching the reported interpretation. The exact two-sided Fisher test compared positivity between groups. Exact Clopper-Pearson confidence intervals were calculated for each group proportion [26], and absolute risk difference was calculated as the immunised proportion minus the control proportion. Because the uncorrected odds ratio is infinite when both off-diagonal cells are zero, a Haldane-Anscombe correction of 0.5 per cell was applied for a finite descriptive odds ratio [27].
Competitive-inhibition animal means were compared using both Welch’s unequal-variance t-test and an exact two-sided Mann-Whitney test. The purpose was not to select the most favourable test, but to demonstrate inferential sensitivity when only three biological units per group are available. Technical-replicate standard errors were not treated as independent animals, consistent with the requirement to align statistical inference with the independently assigned experimental unit and avoid pseudoreplication [7]. No multiplicity adjustment was applied because the analyses were descriptive and hypothesis-generating. Interpretation therefore emphasised effect sizes, uncertainty, biological replication, data hierarchy and robustness across reasonable analytical choices rather than a binary p-value threshold [28,29,30].
Statistical analyses and reproducibility workflows were implemented using R version 4.5.1 (R Foundation for Statistical Computing, Vienna, Austria) and Python version 3.12.10 (Python Software Foundation, Wilmington, DE, USA). The deposited repository includes executable scripts written in both programming languages to facilitate independent verification of the reported analyses, quality-control procedures, and data-processing workflows.

2.8. Ethical Oversight

The source materials report approval by institutional committees at The University of the West Indies and the Instituto Superior de Ciencias Médicas de Santiago de Cuba, including approval numbers CREC-SA.3404/07/2025, CREC-SA.3430/07/2025 and No. 87-2020. Animals were maintained under veterinary supervision and monitored for clinical abnormalities. The present paper analyses de-identified data and did not involve additional animal procedures. Reporting was considered against ARRIVE 2.0 principles, particularly experimental-unit definition, allocation, blinding, sample-size rationale and transparent reporting of exclusions and outcomes [6].

3. Results

3.1. Dataset Composition and Machine-Readable Structure

The Figshare deposit contains 13 CSV data tables, a 16-sheet Excel workbook, documentation, executable code, a machine-readable manifest and SHA-256 checksums [22]. The animal registry contains 48 anonymous feline records: six proof-of-concept animals and 42 mucosal-cohort animals. The two cohorts are explicitly separated by cohort and assay fields. The donor-hen component is represented as a longitudinal summary table rather than falsely expanded to individual observations. The repository also contains the processed neutralisation outputs available from the source, the reported ID50 estimate, virus-only and cell-only controls, assay metadata, statistical-analysis resources and quality-control records. The data dictionary defines each field, provenance class, missing-value convention and intended unit.
The package supports several levels of reuse. Investigators can directly analyse exact displayed values, including all Ab3 wells, animal-level inhibition means, processed neutralisation summaries and IgA category counts; reproduce the curated statistical analyses from the supplied scripts; and inspect assay-control and quality-control information. Where primary instrument-level values are not represented, schema-defined fields identify the missing level rather than substituting synthetic measurements. This design allows the public object to be versioned and enriched if additional source files are recovered while preserving a clear distinction between current deposited measurements, derived quantities and future primary-data additions.

3.2. Donor-Hen IgY Response

Mean anti-gp120 IgY OD450 increased from 0.16 ± 0.03 at baseline to 0.82 ± 0.05 at week 2, remained 0.80 ± 0.05 at week 6 and reached 0.88 ± 0.04 at week 12. The reported geometric mean endpoint titre was 1:1024. Purified IgY concentration was 8.65 mg/mL, with electrophoretic bands at approximately 68-70 kDa and 24-26 kDa, consistent with heavy and light chains. These records demonstrate the reported generation and recovery of gp120-reactive IgY but do not permit estimation of between-hen variance because individual trajectories were unavailable.

3.3. Ab3 ELISA Technical Reconstruction and Cut-Off Sensitivity

Recalculated means from the displayed wells were 0.7751, 0.5443 and 0.3288 for the three immunised cats and 0.1216, 0.1069 and 0.0933 for the three controls. Corresponding within-sample CVs were 2.46%, 2.57% and 4.00% for immunised samples and 5.03%, 3.98% and 5.43% for controls. The primary positive control had mean 1.1924 and CV 0.99%; the repeated positive control had mean 1.1240 and CV 3.80%. Blank wells averaged 0.0403 with CV 8.16%, although the absolute blank variation was small.
The 24 displayed negative-control wells had pooled mean 0.1073 and sample SD 0.01284. Mean plus three SD equalled 0.1458. This did not reproduce either 0.32 or 0.35, indicating that the displayed wells alone are insufficient to reconstruct the reported threshold. The third immunised cat had a mean of 0.3288: it would be positive under a 0.32 threshold, negative under a 0.35 threshold and positive under the displayed-control-derived threshold. The repeated positive-control wells averaged 1.1240, whereas the source summary displayed 1.116. These differences were retained as audit findings rather than corrected in the source data.

3.4. Competitive Inhibition

Animal-level mean inhibition was 13.2%, 15.8% and 11.1% in immunised cats and 1.3%, 1.7% and 2.0% in controls. The between-group separation was large, with no overlap among the six displayed animal means. Nevertheless, inferential results depended on the test. Welch’s t-test gave p = 0.0118, whereas the exact two-sided Mann-Whitney test gave p = 0.10 because only three animals were available per group. At reciprocal dilution 1000, reported mean inhibition was 13.6 ± 2.1% with gp120 peptide, 8.6 ± 1.8% for the Ab3-Ab2 interaction, 1.7 ± 0.6% with non-specific peptide and 2.1 ± 0.7% with unrelated control antibody. The direction and control pattern support assay-specific interference, but the magnitude was modest and the biological sample size was small.

3.5. HIV-1 JR-FL Neutralisation Summary

At the displayed reciprocal dilution of 40,000, mean RLU values were 500 ± 50 for cell-only wells, 8500 ± 420 for virus-only wells, 8000 ± 400 for virus plus control serum and 1600 ± 120 for virus plus immunised serum. The displayed calculation yielded 5.9% neutralisation for control serum and 86.3% for immunised serum. The source reported an immunised-group ID50 of 3.2 × 10³ (95% CI 2.1-4.8 × 10³), while controls did not reach 50% neutralisation. Because the complete numerical dilution series was absent, curve shape, confidence intervals, individual-animal variability and goodness of fit could not be reproduced.

3.6. Salivary IgA Response and Exact Effect Estimates

All 18 immunised cats were reported positive: six low-positive, six moderate-positive and six high-positive. Among 24 controls, 20 were negative and four borderline; none was positive. With borderline treated as not positive, Fisher’s exact p was 2.83 × 10−12. The positivity estimate was 100% in the immunised group (exact 95% CI 81.5-100%) and 0% in controls (exact 95% CI 0-14.2%). The absolute risk difference was 100 percentage points. The uncorrected odds ratio was infinite. Applying 0.5 to all cells produced a corrected odds ratio of 1813.0, compared with 1728.0 reported in the source, illustrating that finite corrections require the formula and table orientation to be specified. Group mean OD450 values were reported as 0.48 ± 0.21 and 0.07 ± 0.03 for immunised and control animals, respectively. Individual values could not be independently checked.
Table 3. Principal reported and reproducibly derived numerical findings. 
Table 3. Principal reported and reproducibly derived numerical findings. 
Domain Reported data Reanalysis Interpretive boundary
Hen IgY OD450 0.16 at baseline and 0.88 at week 12; GMT 1:1024 Summary retained as reported No individual hen trajectories
Ab3 ELISA Three immunised means displayed near 0.775, 0.544 and 0.329 Exact well-derived means 0.7751, 0.5443 and 0.3288; displayed negatives give cut-off 0.1458 Source thresholds 0.32 and 0.35 are inconsistent and not reproducible from displayed negatives
Inhibition Immunised 11.1-15.8%; controls 1.3-2.0% Welch p = 0.0118; exact Mann-Whitney p = 0.10 Only three independent animals per group
Neutralisation 86.3% at displayed condition; ID50 3.2 × 10³; virus-only and cell-only controls deposited Processed 1:40,000 calculation retained; reported ID50 and controls available Full instrument-level dilution-series RLU and individual fitted curves unavailable
Salivary IgA 18/18 positive versus 0/24 Fisher p = 2.83 × 10−12; risk difference 100%; corrected OR 1813 Individual OD450 and consumption dose unavailable

3.7. Integrated Epitope Mapping and Structural Interpretation

To provide an integrative representation of the molecular findings and their potential relationship to the observed immunological responses, the experimental epitope-mapping results were combined with sequence-level and structural information into a proposed mechanistic model (Figure 1). The analysis localized the principal reactive region to residues 254–274 of HIV-1 gp120 and integrated overlapping-peptide reactivity with fine epitope mapping by alanine-scanning mutagenesis, site-directed gp120 mutant analysis, and three-dimensional localization of the candidate epitope within the gp120 trimer structure.
The overlapping-peptide analysis identified peptide P5 (260–274) as the strongest reactive segment within the broader 240–290 region. Subsequent residue-level mapping within the 254–274 sequence identified a restricted set of positions with pronounced effects on antibody recognition, while complementary analysis using site-directed gp120 mutants provided an independent assessment of residues contributing to antibody binding. Structural mapping further placed the 254–274 region within the C3 domain of gp120 and enabled its spatial relationship to other functionally relevant regions, including the V3 loop and CD4-binding site, to be visualized within the trimeric protein context. From a computational biology and bioinformatics perspective, this multilevel mapping provides a framework for integrating sequence, mutational, immunological, and protein-structural information into a unified representation. The resulting model proposes a mechanistic connection between oral anti-gp120 IgY exposure, induction of an anti-idiotypic antibody response, recognition of the gp120 254–274 region by Ab3 antibodies, mucosal IgA responses, and the neutralizing activity observed in the proof-of-concept study. This proposed pathway should be interpreted as a hypothesis-generating model derived from the integrated experimental and structural evidence rather than as a fully established causal mechanism.

3.8. Reproducibility Audit

Seven high-priority issues were encoded in the quality-issues table: Ab3 threshold inconsistency; disagreement between repeated positive-control summary and displayed wells; competitive-inhibition test and p-value inconsistency; neutralisation dilution-range inconsistency; conflicting descriptions of neutralisation statistical testing; absence of individual mucosal IgA values; and unquantified individual exposure in the ad-libitum cohort. A further design-level inconsistency concerned peptide-KLH versus recombinant gp120 descriptions of the donor-hen immunogen. None of these issues invalidates the deposited observations by itself, but each constrains the strength or reproducibility of a specific claim.
Table 4. Key reproducibility issues and recommended resolution.
Table 4. Key reproducibility issues and recommended resolution.
Issue Consequence Current mitigation / future enhancement
Ab3 cut-off reported as 0.32 and 0.35; displayed negatives imply 0.146 Cat 3 classification changes under the two reported thresholds Document both source thresholds, retain well-level values and preserve the recalculated audit threshold; add original cut-off-validation records if available in a future version.
Animal-level inhibition gives test-dependent p-values Statistical significance is unstable at n = 3 per group Use animals as the inferential unit, report both sensitivity analyses and avoid treating technical replicates as independent observations.
Neutralisation methods and figure use discordant dilution ranges ID50 and curve cannot be independently reconstructed Retain processed neutralisation outputs, ID50, controls and metadata; add complete instrument-level dilution-series RLU values only if original exports are recovered.
IgA results available only as aggregate counts and a plot Duplicate agreement and individual classification cannot be verified Retain exact aggregate counts and explicit missingness; add de-identified individual plate-reader values only if original records are recovered.
Ad-libitum intake not measured No dose-response or per-kilogram exposure analysis Describe the study as assigned exposure and do not perform dose-response or per-kilogram analyses.
Immunogen described inconsistently Replication protocol is ambiguous Preserve the discrepancy in provenance records and resolve the production history if contemporaneous laboratory documentation becomes available.

4. Discussion

4.1. Scientific and Data-Science Value of the Resource

This dataset converts a visually rich but analytically fragmented preclinical study into a machine-readable, provenance-aware resource. Its value lies in both the biological observations and the ability to trace which statements are directly supported, derived or unavailable without primary records. In data-intensive biology, that distinction is foundational: figures can conceal the hierarchy of animals, wells, plates and transformations. Exposing that hierarchy allows independent users to reproduce calculations, challenge analytical choices and design improved follow-up studies.
The use of a persistent Figshare DOI, non-proprietary tables, a data dictionary, cryptographic checksums, quality-control documentation and executable scripts addresses major components of FAIR stewardship [1,2]. Findability is supported by the persistent identifier; accessibility by public repository deposition; interoperability by CSV schemas, explicit units and structured metadata; and reusability by provenance labels, code and usage notes. These practices also reflect broader open-research recommendations that data, methods and analytical workflows should be sufficiently transparent to permit scrutiny and reuse [3]. The resource does not claim that every historical instrument-level measurement is available; rather, it distinguishes deposited processed measurements from unavailable primary exports. FAIRness is better served by explicit provenance and structured missingness than by reconstructed or invented detail. Because the repository includes executable Python and R scripts, the software component is additionally aligned with contemporary FAIR4RS guidance for research software [25].
The resource also demonstrates why openly shared assay data should include negative controls, technical replicates and explicit decision rules. The 72 Ab3 well measurements enabled independent calculation of means, SDs and CVs, revealing generally low within-sample variation while simultaneously exposing a cut-off discrepancy that would have been invisible in a table containing only positive/negative labels. Cut-offs are not merely technical details; they are decision rules that can change an animal’s classification and alter the apparent responder proportion. Fit-for-purpose immunoassay validation requires a documented reference population, predefined calculation, assessment of precision and evidence that the threshold is appropriate for the intended use [31,32]. The source threshold therefore remains transparently documented as an audit issue rather than silently replaced by a post-hoc threshold.

4.2. Biological Interpretation

The donor-hen data support successful production of gp120-reactive IgY, with a rapid rise after immunisation and sustained reactivity to week 12. IgY production in hens is an established platform with practical advantages for repeated antibody harvesting [8,9,23]. Nevertheless, the deposited records do not resolve whether the peptide-KLH or recombinant-gp120 description applies to all production phases. This matters because antibody specificity, conformational recognition and downstream idiotypic determinants may differ substantially between a short linear peptide and a folded glycoprotein.
The systemic proof-of-concept findings are internally directional: immunised cats had higher gp120-reactive ELISA values, greater competitive inhibition and lower luciferase signals than controls. The convergence of binding, interference and functional assays is more informative than any one assay alone. However, it does not prove structural antigen mimicry. Definitive internal-image claims would require direct mapping of the determinants recognised by Ab3, competition with characterised epitope-specific antibodies, binding-kinetic analysis and preferably structural evidence. The present data are therefore consistent with an induced gp120-related response but should remain labelled exploratory.
The salivary IgA separation is the numerically strongest observation. All assigned immunised animals were positive and all controls were negative or borderline. This is compatible with activation of a mucosal response, given the established role of gut-associated lymphoid tissues and IgA-producing plasma cells [13,14,15]. Yet positivity is not equivalent to protective function. The assay was qualitative, individual values were unavailable, and the actual quantity consumed by each animal was not measured. Moreover, the IgA cohort was independent of the six-cat systemic cohort, so the data cannot establish that an animal with high salivary IgA also had Ab3 activity or HIV-1 neutralisation. Future studies should measure these endpoints longitudinally in the same animals and evaluate whether salivary IgA competes with gp120 or neutralises virus at mucosal concentrations.
The neutralisation component is biologically important because it adds a functional endpoint to the binding and inhibition assays. The TZM-bl system is a validated and widely used platform when assay controls, dilution design and analytical procedures are appropriately documented [17,18]. The deposited resource includes the processed neutralisation information available from the study, including virus-only and cell-only controls, reported RLU summaries, percentage neutralisation and the calculated group ID50. These data support verification of the displayed calculation and methodological comparison. The absence of the complete original well-level luminometer export, however, means that independent refitting of the entire dose-response curve and assessment of individual-animal curve heterogeneity are not possible from the current release. This distinction does not negate the available functional signal, but it appropriately limits inference. Moreover, neutralisation against one clade B pseudovirus cannot establish breadth; broadly neutralising activity requires demonstration across genetically diverse viral isolates [19,20].
The integrated mapping presented in Figure 1 provides a mechanistic framework for interpreting the anti-idiotypic response in relation to the original gp120 immunogen. Importantly, the gp120 254–274 fragment conjugated to KLH was used as the original antigen to generate the idiotypic antibodies (Ab1), establishing this sequence as the molecular starting point of the proposed idiotypic cascade. The subsequent recognition of the same or overlapping gp120 region by Ab3 antibodies is therefore consistent with the classical idiotypic-network concept, in which an anti-idiotypic antibody (Ab2), particularly an internal-image Ab2β, can reproduce structural or functional features of the original antigenic determinant and subsequently induce Ab3 antibodies with antigen-binding properties resembling those of Ab1. Experimental studies have demonstrated that Ab2 immunization can generate Ab3 antibodies that recognize the original antigen, providing biological precedent for this interpretation [33,34].
This interpretation is particularly relevant to HIV-1 envelope immunity because anti-idiotypic interactions involving gp120 have previously been demonstrated experimentally. Anti-idiotypic antibodies generated against human anti-gp120 antibodies have been reported to mimic gp120-associated molecular recognition, while anti-idiotypic vaccination has also been investigated as a strategy for inducing gp120-directed immune responses [35]. In the present model, integration of overlapping-peptide reactivity, alanine-scanning mutagenesis, site-directed gp120 mutants, and structural localization places the 254–274 sequence within the C3 region and provides a molecular basis for investigating the relationship between Ab3 recognition, mucosal IgA responses, and the observed neutralizing activity. Nevertheless, these associations should be interpreted as hypothesis-generating rather than proof of a causal pathway: direct demonstration that the Ab2 population structurally mimics the 254–274 epitope, and that this molecular mimicry is responsible for the downstream mucosal and neutralizing responses, would require additional structural and functional validation.
Figure 1 provides an integrative, multiscale representation of the experimental and computational data layers used to characterize the HIV-1 gp120 254–274 region and its proposed involvement in the anti-idiotypic immune response. The figure consolidates heterogeneous data modalities, including primary amino-acid sequence information, overlapping-peptide binding profiles, residue-resolved alanine-scanning measurements, site-directed gp120 mutational data, and three-dimensional structural annotation of the gp120 trimer. From a biomedical data-science perspective, these complementary data types define distinct but interoperable feature spaces. Peptide-level reactivity narrows the candidate antigenic region, residue-level perturbation analyses quantify the contribution of individual amino acids to antibody recognition, and structural projection embeds these experimentally derived signals within the spatial organization of the envelope glycoprotein. Their integration therefore enables cross-modal concordance assessment and establishes a traceable analytical pathway linking molecular measurements to higher-order biological interpretation.
The figure also illustrates how legacy immunological observations can be transformed into a structured and computationally reusable biomedical resource. Individual analytical layers can be represented in machine-readable form through peptide identifiers, residue coordinates, normalized binding metrics, mutant annotations, and structural positions, thereby supporting reproducible visualization, comparative epitope analysis, feature-level interrogation, and future integration with external sequence, structural, or immunological datasets. Critically, the empirical data layers should be distinguished from the inferred mechanistic model shown in the lower portion of the figure. The peptide-binding, mutational, and structural-mapping results constitute directly observed or computationally derived evidence, whereas the proposed Ab1–Ab2–Ab3 cascade represents a hypothesis-generating interpretation of these integrated findings. This explicit separation between measured data, derived features, and mechanistic inference is central to rigorous biomedical data science and enhances the value of the dataset for independent validation, secondary analysis, reproducibility assessment, and future computational reuse.

4.3. Epidemiological and Statistical Interpretation

The resource illustrates a general epidemiological principle: the size of an observed association and the certainty of its estimate are different questions. The IgA risk difference was 100%, but the exact lower confidence limit for immunised positivity was 81.5% because only 18 exposed animals were studied. The control upper limit was 14.2% despite observing no positives. These intervals appropriately communicate what remains plausible in the underlying populations. Similarly, an infinite uncorrected odds ratio reflects complete separation rather than infinite biological potency. A corrected finite odds ratio depends on the chosen continuity correction, so the risk difference and exact group intervals are often more interpretable than a very large corrected odds ratio.
The inhibition results show an equally important issue. Three animal means per group were fully separated, but the exact Mann-Whitney p-value was 0.10 because the number of possible rank arrangements is limited at such small n. Welch’s t-test yielded p = 0.0118 under distributional assumptions. Neither result should be selected solely because it crosses a conventional 0.05 threshold. The observed effect pattern, biological sample size, raw animal summaries, uncertainty and modelling assumptions should be considered together. This interpretation is consistent with recommendations to move beyond dichotomous declarations of statistical significance and to avoid equating a threshold-crossing p-value with scientific importance [28,30].
Technical replicates improve measurement precision but do not increase the number of independently treated animals. Treating eight ELISA wells or multiple luciferase wells as separate biological observations would constitute pseudoreplication and could produce misleadingly small p-values [7]. The deposited schemas therefore identify animal, assay condition and technical replicate separately. Future confirmatory analyses should use mixed-effects or hierarchical models when repeated technical observations are retained, with treatment effects estimated at the animal level.

4.4. Reuse Opportunities

The dataset was designed to maximise scientific reuse across immunology, virology, epidemiology and biomedical data science. Potential applications include:
  • independent validation of ELISA analytical workflows;
  • comparative evaluation of HIV-1 pseudovirus neutralisation assays;
  • benchmarking of ID50 estimation procedures;
  • development of statistical methods for immunological datasets;
  • reproducibility studies of antibody-based laboratory assays;
  • teaching datasets for immunology, virology, epidemiology and data science;
  • meta-research investigating FAIR implementation in preclinical biomedical datasets.
Because the repository includes processed neutralisation data, calculated ID50 information, virus-only and cell-only assay controls, assay metadata, quality-control documentation, statistical outputs, machine-readable tabular files and executable code, investigators can reproduce the curated analyses, evaluate analytical choices and compare workflows without repeating the underlying animal experiments. Persistent identifiers, community-compatible formats and explicit metadata improve discoverability and interoperability, while transparent provenance supports responsible secondary analysis [1,2,3]. The accompanying Python and R workflows further strengthen computational reuse in line with FAIR4RS recommendations for research software [25].

4.5. Limitations

The principal limitation is that the repository is a curated, provenance-aware data release rather than a complete export of every original instrument-level measurement. The deposited resource nevertheless contains the processed analytical outputs used to describe the reported findings, including processed neutralisation data, calculated ID50 information, virus-only and cell-only controls, assay metadata, statistical analyses and quality-control documentation. Individual salivary IgA OD450 observations, the complete well-by-well neutralisation dilution series, individual age and weight values, and measured intake in the ad-libitum cohort are not available at the same level of granularity. These boundaries are explicitly encoded in the data dictionary and provenance fields rather than obscured by imputation.
Second, the proof-of-concept sample size was three animals per group and therefore supports exploratory biological plausibility and assay development rather than definitive efficacy estimation [29]. Third, several source-level inconsistencies remain documented, including Ab3 thresholds, inhibition test descriptions, neutralisation dilution ranges and immunogen descriptions. Fourth, the two feline cohorts are independent; cross-assay correlations cannot be estimated at animal level. Fifth, the neutralisation experiment used one pseudovirus strain, so breadth and clinical protection cannot be inferred [19,20]. Finally, the feline model provides a heterogeneous mammalian system but does not directly model human HIV acquisition or vaccine efficacy.
A major strength is that these limitations are documented within the dataset itself. The repository combines processed immunological and virological outputs with metadata, provenance records, quality-control information and executable analytical code. Users can therefore distinguish what was measured, processed, derived and unavailable, reducing the risk of overinterpretation. Persistent identifiers and community-compatible standards support FAIR stewardship [1,2], while transparent sharing of data and workflows supports reproducible research culture [3,4].
Figure 1 below shows data provenance workflow from experimental generation to FAIR-compliant repository deposition.
From a data-science perspective, Figure 2 represents the study as an end-to-end data provenance pipeline rather than as a sequence of isolated laboratory procedures. The workflow links animal-level experimental metadata to sample collection, immunological assays, quality-control procedures, statistical analysis, dataset curation, public deposition, and eventual FAIR reuse. This explicit representation of data lineage is particularly important in biomedical research because each analytical result can be traced back to its experimental source and processing history. The separation of raw measurements, derived variables, statistical outputs, metadata, and curated datasets reduces ambiguity between primary observations and transformed data products. Quality-control checkpoints—including plate-level assessment, data-integrity verification, exclusion criteria, and consistency checks—function as validation layers within the pipeline, while the subsequent statistical-analysis stage converts validated measurements into interpretable derived outputs. In computational terms, the figure therefore describes a structured transformation process in which biological observations are progressively converted into standardized, quality-controlled, and analyzable digital objects.
The downstream components of the workflow further strengthen its relevance to biomedical data science by emphasizing machine readability, reproducibility, provenance, and long-term reuse. Organization of the resource into structured CSV tables, a documented data dictionary, provenance records, analysis scripts, repository manifests, persistent identifiers, versioning, and cryptographic checksums creates an auditable digital research object rather than a static supplementary dataset. Deposition in Figshare enhances accessibility and persistence, while explicit metadata and standardized file organization improve interoperability and computational reuse. Importantly, the workflow also separates the experimental, computational, and dissemination layers, enabling independent investigators to reproduce statistical analyses, inspect data transformations, validate individual processing steps, and potentially integrate the dataset with external immunological or comparative biomedical resources. In this sense, Figure 2 illustrates how a conventional preclinical experiment can be converted into a FAIR-aligned, provenance-aware biomedical data resource, supporting reproducibility, secondary analysis, data integration, and future computational investigations in immunology and translational biomedical research.

5. Plain Language Summary

5.1. Comprehensive Summary of the Scientific Audit, Data Curation, and Preparation of the Data Resource Article

The original preclinical investigation evaluating the oral administration of hyperimmune anti-HIV-1 gp120 immunoglobulin Y (IgY) in domestic cats was systematically reviewed, curated, and transformed into a comprehensive scientific data resource suitable for submission as a Data Resource Article to Data (MDPI). The overarching objective of this work was not merely to report the biological findings of the original experiment, but to preserve the underlying experimental evidence in a transparent, structured, reproducible, and reusable format consistent with contemporary FAIR (Findable, Accessible, Interoperable, and Reusable) research-data principles.

5.2. Scientific Audit and Reconstruction of the Study

The process began with a detailed scientific audit of the original preclinical investigation. Experimental records, study design, treatment groups, biological measurements, analytical procedures, and reported outcomes were reviewed to reconstruct the complete experimental workflow. Particular attention was given to establishing consistency between the original experimental observations, derived variables, statistical analyses, figures, tables, and conclusions.
The audit also examined the biological rationale underlying the study: the use of orally administered hyperimmune IgY directed against HIV-1 gp120 and the subsequent evaluation of immune responses in domestic cats. The experimental framework was reconstructed so that future investigators could understand not only the final reported results but also how individual observations originated and how they were transformed into analytical outputs.

5.3. Dataset Verification, Curation, and Data Provenance

A major component of the project involved verification and systematic curation of the complete experimental dataset. Raw and processed variables were organized into machine-readable structures, with standardized variable names, consistent coding conventions, explicit treatment-group identifiers, and harmonized data types.
Data provenance was documented from the level of the individual experimental animal through biological sampling, laboratory measurement, data recording, processing, statistical analysis, and final reported output. This provenance framework provides an auditable connection between the biological experiment and the corresponding digital data objects.
A comprehensive data dictionary and associated metadata were prepared to define variables, units, experimental groups, measurement types, coding conventions, and other information necessary for independent interpretation of the dataset. These elements substantially increase the long-term usability of the resource and reduce dependence on undocumented knowledge from the original investigators.

5.4. Quality Control and Integrity Verification

Quality-control procedures were incorporated throughout the curation process. Dataset structure, variable consistency, missing observations, identifiers, numerical formats, and relationships among source and derived data were examined systematically.
The curated repository was designed to preserve both scientific and digital integrity. Repository-level documentation included provenance information, file inventories and manifests, and cryptographic checksums where applicable. These measures provide mechanisms for confirming file integrity and identifying unintended modifications after dataset distribution.
Importantly, quality control was approached as more than a simple verification of numerical values. The objective was to establish a transparent chain of evidence connecting experimental design, data acquisition, curated records, computational analysis, and published conclusions.

5.5. Repository Architecture and FAIR Data Preparation

The research outputs were organized for public dissemination through Figshare. The repository architecture was developed to facilitate both human inspection and computational reuse.
The resource includes structured CSV files suitable for programmatic analysis, an Excel workbook for convenient inspection of the organized data, a detailed data dictionary, metadata documentation, provenance records, a repository manifest, integrity information including cryptographic checksums, and computational resources supporting independent reproduction of the analyses.
Executable Python and R scripts were prepared to complement the underlying data. This computational component is particularly important because reproducibility requires more than public availability of numerical values: investigators should also be able to understand and, where possible, reproduce the transformations and analyses used to generate reported results.
Collectively, these components were designed to align the repository with FAIR research-data principles. Findability is supported through structured metadata and repository deposition; accessibility through public dissemination; interoperability through standardized machine-readable formats; and reusability through extensive documentation, provenance, variable definitions, and computational resources.

5.6. Reanalysis and Verification of the Principal Findings

The principal experimental findings were re-examined against the curated dataset. This process served as an independent internal consistency check between the underlying observations and the scientific statements presented in the manuscript.
Where statistical or computational processing was required, analyses were reconstructed in a reproducible manner and incorporated into the accompanying analytical scripts. This approach strengthens the distinction between original observations and subsequently derived analytical outputs and enables future users to interrogate the dataset independently.
The purpose of this reanalysis was not to reinterpret the historical experiment retrospectively, but rather to establish that the scientific resource accurately represents the available evidence and that the principal reported findings can be traced back to identifiable experimental records.

5.7. Transformation into a Data Resource Article

The manuscript was substantially revised and restructured to reflect the objectives of a Data Resource Article rather than those of a conventional hypothesis-driven experimental paper.
Accordingly, emphasis was shifted toward dataset generation, organization, provenance, quality, technical validation, accessibility, reproducibility, and potential reuse. The manuscript components—including the abstract, methodological description, results-related material, data records, technical validation, usage notes, limitations, and conclusions—were refined to provide readers with a coherent description of both the biological experiment and the resulting digital resource.
Special attention was devoted to clearly separating experimentally observed information from interpretation and to avoiding claims extending beyond what the archived dataset could directly support.

5.8. Figures and Visual Documentation

Publication-quality figures were developed or refined to communicate the structure and provenance of the data resource. In particular, graphical documentation of the data-provenance workflow was used to illustrate the progression from the original animal experiment and biological measurements through data acquisition, curation, computational processing, quality control, and final repository deposition.
Such visualization provides readers with an immediate conceptual overview of how the different components of the resource relate to one another and complements the detailed textual and machine-readable documentation.

5.9. Submission Materials and Scientific Positioning

The final stage included preparation and refinement of the journal submission materials. The scientific positioning emphasized that the principal contribution of the work lies not only in the historical biological observations but also in the preservation and release of the underlying experimental evidence as a documented and reusable research resource.
The submission materials therefore highlighted the originality of the dataset, its public availability, the extensive curation and technical documentation performed, the inclusion of reproducible computational resources, and its potential value for independent validation and methodological comparison.
The limitations of the experimental system were also treated explicitly. Rather than overstating translational implications, the resource was positioned as a preclinical immunological dataset that may support further investigation into oral immunization, passive exposure to hyperimmune IgY, anti-idiotypic immune responses, and related experimental questions.

5.10. Overall Outcome

Taken together, this work transformed a conventional preclinical immunology investigation into a substantially more transparent and reusable scientific resource. The original experimental evidence was audited, curated, structured, documented, quality-controlled, computationally supported, and organized for public dissemination.
The resulting FAIR-aligned resource establishes traceability from the experimental animals and laboratory observations to the curated datasets and analytical outputs. By combining structured data, detailed metadata, provenance documentation, integrity controls, and reproducible analytical scripts, the project provides future investigators with the information required to inspect, validate, reinterpret, and computationally reuse the study.
Consequently, the scientific value of the original experiment extends beyond its initial conclusions. The curated resource can serve as a foundation for independent validation, methodological comparisons, secondary analyses, educational applications, and future investigations concerning oral immunization and anti-idiotypic immune responses, while simultaneously contributing to broader efforts toward transparency, reproducibility, and responsible preservation of preclinical biomedical research data.

6. Conclusions

The publicly deposited Figshare dataset is a FAIR-oriented biomedical resource describing mucosal and systemic immune responses following oral administration of hyperimmune anti-HIV-1 gp120 IgY in an outbred feline model. It integrates donor-hen IgY summaries, well-level Ab3 ELISA measurements, competitive-inhibition data, processed HIV-1 JR-FL neutralisation outputs, calculated ID50 information, virus-only and cell-only controls, salivary IgA classifications, assay metadata, statistical analyses, quality-control documentation, provenance records, machine-readable tables and executable scripts. These resources enhance transparency, support independent verification of deposited calculations, facilitate secondary analyses and assay benchmarking, and provide a reusable methods-development resource for immunology, virology, epidemiology and biomedical data science. By distinguishing processed and derived measurements from unavailable primary records, and by using persistent identification, structured metadata, open formats and reusable code, the resource advances FAIR and open-research practice while maintaining appropriate biological and statistical limits on interpretation [1,2,3,25].

Author Contributions

Draft CRediT statement, to be confirmed before submission: A.J.-V., conceptualisation, methodology, data curation, formal analysis, visualisation and writing - original draft; O.A., investigation, validation and writing - review and editing; B.F.-C., investigation, animal-study support and writing - review and editing; O.P.-M., methodology, supervision and writing - review and editing. All authors verify and approve the final contribution statement.

Funding

No external funding was identified in the supplied source materials.

Institutional Review Board Statement

The source study reports approvals CREC-SA.3404/07/2025 and CREC-SA.3430/07/2025 from The University of the West Indies and No. 87-2020 from the Instituto Superior de Ciencias Médicas de Santiago de Cuba. No new animal procedureswere conducted for this data analysis.

Data Availability Statement

The curated dataset supporting this paper is publicly deposited in Figshare under DOI 10.6084/m9.figshare.33179345 [22]. The repository includes machine-readable tables, processed ELISA and HIV-1 neutralisation outputs, calculated ID50 information, virus-only and cell-only controls, assay metadata, statistical-analysis resources, quality-control documentation, provenance information and supporting files. The source preprint is available under DOI 10.64898/2026.07.04.736505 [21]. Users should cite the dataset version accessed and retain provenance fields when redistributing derived tables. Executable analysis scripts compatible with R version 4.5.1 and Python version 3.12.10 are included in the Figshare repository together with the processed datasets, metadata, quality-control documentation, and data dictionary.
Code Availability: Python and R scripts are included in the Figshare package. The Python workflow reproduces the Ab3 summaries, cut-off audit, inhibition comparisons and exact IgA analyses from the deposited data. The scripts and accompanying documentation should be versioned and cited together with the dataset; this treatment of analysis software as a reusable research object is consistent with FAIR4RS guidance [25]. Analyses that require unavailable individual-level or instrument-level primary values are deliberately not reconstructed.
Declaration on Generative Artificial Intelligence: A generative artificial-intelligence system was used to assist with manuscript structuring, language refinement, and reproducibility review. Repository preparation and reproducibility checks were supported by computational tools under author supervision. The authors remain responsible for verifying all data, analyses, citations and interpretations and for compliance with the target journal’s disclosure policy.

Conflicts of Interest

The source study declares no conflict of interest.

Acknowledgments

The authors acknowledge the personnel responsible for animal care, sample collection and laboratory testing in the original studies.

References

  1. Wilkinson, M.D.; Dumontier, M.; Aalbersberg, I.J.; et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci. Data 2016, 3, 160018. [Google Scholar] [CrossRef] [PubMed]
  2. Sansone, S.A.; McQuilton, P.; Rocca-Serra, P.; Gonzalez-Beltran, A.; Izzo, M.; Lister, A.L.; et al. FAIRsharing Community. FAIRsharing as a community approach to standards, repositories and policies. Nat. Biotechnol. 2019, 37, 358–367. [Google Scholar] [CrossRef] [PubMed]
  3. Nosek, B.A.; Alter, G.; Banks, G.C.; et al. Promoting an open research culture. Science 2015, 348, 1422–1425. [Google Scholar] [CrossRef] [PubMed]
  4. Munafò, M.R.; Nosek, B.A.; Bishop, D.V.M.; et al. A manifesto for reproducible science. Nat. Hum. Behav. 2017, 1, 0021. [Google Scholar] [CrossRef] [PubMed]
  5. Peng, R.D. Reproducible research in computational science. Science 2011, 334, 1226–1227. [Google Scholar] [CrossRef] [PubMed]
  6. Percie du Sert, N.; Hurst, V.; Ahluwalia, A.; et al. The ARRIVE guidelines 2.0: updated guidelines for reporting animal research. PLoS Biol. 2020, 18, e3000410. [Google Scholar] [CrossRef] [PubMed]
  7. Hurlbert, S.H. Pseudoreplication and the design of ecological field experiments. Ecol. Monogr. 1984, 54, 187–211. [Google Scholar] [CrossRef]
  8. Lee, L.; Samardzic, K.; Wallach, M.; Frumkin, L.R.; Mochly-Rosen, D. Immunoglobulin Y for potential diagnostic and therapeutic applications in infectious diseases. Front Immunol. 2021, 12, 696003. [Google Scholar] [CrossRef] [PubMed]
  9. Grzywa, R.; Łupicka-Słowik, A.; Sieńczyk, M. IgYs: on her majesty’s secret service. Front Immunol. 2023, 14, 1199427. [Google Scholar] [CrossRef] [PubMed]
  10. Jerne, N.K. Towards a network theory of the immune system. Ann. Immunol. 1974, 125C, 373–389. [Google Scholar]
  11. Kieber-Emmons, T.; Monzavi-Karbassi, B.; Pashov, A.; Saha, S.; Murali, R.; Kohler, H. The promise of the anti-idiotype concept. Front Oncol. 2012, 2, 196. [Google Scholar] [CrossRef] [PubMed]
  12. Caputo, V.; Negri, I.; Moudoud, L.; et al. Anti-HIV humoral response induced by different anti-idiotype antibody formats: an in silico and in vivo approach. Int. J. Mol. Sci. 2024, 25, 5737. [Google Scholar] [CrossRef] [PubMed]
  13. Mantis, N.J.; Rol, N.; Corthésy, B. Secretory IgA’s complex roles in immunity and mucosal homeostasis in the gut. Mucosal Immunol. 2011, 4, 603–611. [Google Scholar] [CrossRef] [PubMed]
  14. Macpherson, A.J.; Yilmaz, B.; Limenitakis, J.P.; Ganal-Vonarburg, S.C. IgA function in relation to the intestinal microbiota. Annu Rev. Immunol. 2018, 36, 359–381. [Google Scholar] [CrossRef] [PubMed]
  15. Reboldi, A.; Arnon, T.I.; Rodda, L.B.; et al. IgA production requires B cell interaction with subepithelial dendritic cells in Peyer’s patches. Science 2016, 352, aaf4822. [Google Scholar] [CrossRef] [PubMed]
  16. Seaton, K.E.; Deal, A.; Han, X.; et al. Meta-analysis of HIV-1 vaccine-elicited mucosal antibodies in humans. npj Vaccines 2021, 6, 56. [Google Scholar] [CrossRef] [PubMed]
  17. Montefiori, D.C. Measuring HIV neutralization in a luciferase reporter gene assay. Methods Mol. Biol. 2009, 485, 395–405. [Google Scholar] [CrossRef] [PubMed]
  18. Sarzotti-Kelsoe, M.; Bailer, R.T.; Turk, E.; et al. Optimization and validation of the TZM-bl assay for standardized assessments of neutralizing antibodies against HIV-1. J. Immunol. Methods 2014, 409, 131–146. [Google Scholar] [CrossRef] [PubMed]
  19. Sok, D.; Burton, D.R. Recent progress in broadly neutralizing antibodies to HIV. Nat. Immunol. 2018, 19, 1179–1188. [Google Scholar] [CrossRef] [PubMed]
  20. Gruell, H.; Schommers, P. Broadly neutralizing antibodies against HIV-1 and concepts for application. Curr. Opin. Virol. 2022, 54, 101211. [Google Scholar] [CrossRef] [PubMed]
  21. Justiz-Vaillant, A.; Asin, O.; Ferrer-Cosme, B.; Perez-Martin, O. Oral administration of hyperimmune eggs induces mucosal IgA responses, anti-idiotypic antibodies, and HIV-1 neutralizing activity: a proof-of-concept preclinical study. bioRxiv 2026. [Google Scholar] [CrossRef]
  22. Justiz-Vaillant, A.; Asin, O.; Ferrer-Cosme, B.; Perez-Martin, O. Multi-assay preclinical dataset of mucosal and systemic immune responses following oral hyperimmune anti-HIV-1 gp120 IgY administration in cats. Figshare 2026. [Google Scholar] [CrossRef]
  23. Polson, A.; Coetzer, T.; Kruger, J.; von Maltzahn, E.; van der Merwe, K.J. Improvements in the isolation of IgY from the yolks of eggs laid by immunized hens. Immunol. Invest. 1985, 14, 323–327. [Google Scholar] [CrossRef] [PubMed]
  24. Ho, D.D.; Kaplan, J.C.; Rackauskas, I.E.; Gurney, M.E. Second conserved domain of gp120 is important for HIV infectivity and antibody neutralization. Science 1988, 239, 1021–1023. [Google Scholar] [CrossRef] [PubMed]
  25. Barker, M.; Chue Hong, N.P.; Katz, D.S.; et al. Introducing the FAIR Principles for research software. Sci. Data 2022, 9, 622. [Google Scholar] [CrossRef] [PubMed]
  26. Clopper, C.J.; Pearson, E.S. The use of confidence or fiducial limits illustrated in the case of the binomial. Biometrika 1934, 26, 404–413. [Google Scholar] [CrossRef]
  27. Anscombe, F.J. On estimating binomial response relations. Biometrika 1956, 43, 461–464. [Google Scholar] [CrossRef]
  28. Wasserstein, R.L.; Schirm, A.L.; Lazar, N.A. Moving to a world beyond “p <0.05”. Am. Stat. 2019, 73(sup1), 1–19. [Google Scholar] [CrossRef]
  29. Thabane, L.; Ma, J.; Chu, R.; et al. A tutorial on pilot studies: the what, why and how. BMC Med. Res. Methodol. 2010, 10, 1. [Google Scholar] [CrossRef] [PubMed]
  30. Amrhein, V.; Greenland, S.; McShane, B. Scientists rise up against statistical significance. Nature 2019, 567, 305–307. [Google Scholar] [CrossRef] [PubMed]
  31. Andreasson, U.; Perret-Liaudet, A.; van Waalwijk van Doorn, L.J.C.; et al. A practical guide to immunoassay method validation. Front Neurol. 2015, 6, 179. [Google Scholar] [CrossRef] [PubMed]
  32. Lee, J.W.; Devanarayan, V.; Barrett, Y.C.; et al. Fit-for-purpose method development and validation for successful biomarker measurement. Pharm. Res. 2006, 23, 312–328. [Google Scholar] [CrossRef] [PubMed]
  33. Jerne, N.K.; Roland, J.; Cazenave, P.A. Recurrent idiotopes and internal images. EMBO J. 1982, 1(2), 243–247. [Google Scholar] [CrossRef] [PubMed]
  34. Hébert, J.; Bernier, D.; Boutin, Y.; Jobin, M.; Mourad, W. Generation of anti-idiotypic and anti-anti-idiotypic monoclonal antibodies in the same fusion. Support of Jerne's network theory. J. Immunol. 1990, 144(11), 4256–4261. [Google Scholar] [CrossRef]
  35. Corre, J.P.; Février, M.; Chamaret, S.; Thèze, J.; Zouali, M. Anti-idiotypic antibodies to human anti-gp120 antibodies bind recombinant and cellular human CD4. Eur. J. Immunol. 1991, 21(3), 743–751. [Google Scholar] [CrossRef] [PubMed]
Figure 1. Integrated epitope mapping and proposed mechanism of anti-idiotypic immunity. The HIV-1 gp120 254–274 peptide, originally conjugated to KLH and used as the antigen to generate the idiotypic antibodies, is mapped by sequence, mutational, and structural analyses as a candidate C3-domain epitope recognized by the induced anti-idiotypic network. The model links oral anti-gp120 IgY exposure and Ab3 recognition of the 254–274 epitope with mucosal IgA responses and the HIV-1 neutralizing activity observed in the study.
Figure 1. Integrated epitope mapping and proposed mechanism of anti-idiotypic immunity. The HIV-1 gp120 254–274 peptide, originally conjugated to KLH and used as the antigen to generate the idiotypic antibodies, is mapped by sequence, mutational, and structural analyses as a candidate C3-domain epitope recognized by the induced anti-idiotypic network. The model links oral anti-gp120 IgY exposure and Ab3 recognition of the 254–274 epitope with mucosal IgA responses and the HIV-1 neutralizing activity observed in the study.
Preprints 228060 g001
Figure 2. Figure 2. Data provenance workflow from experimental generation to FAIR-compliant repository deposition. The schematic illustrates the complete lifecycle of the dataset from biological experimentation to public data sharing and reuse. The workflow begins with the preclinical animal studies and sample collection, followed by immunological characterisation using ELISA and HIV-1 JR-FL TZM-bl pseudovirus neutralisation assays. Subsequent stages include quality-control procedures, statistical analyses, and dataset curation, encompassing data validation, metadata standardisation, provenance documentation, data dictionaries, and machine-readable file preparation. The curated resource is then deposited in the Figshare repository with a persistent digital object identifier (DOI), ensuring long-term accessibility and citability. The final stage highlights compliance with the FAIR (Findable, Accessible, Interoperable and Reusable) principles, demonstrating how structured metadata, documented provenance, reproducible analytical workflows, and standardised file formats maximise transparency, reproducibility, interoperability and the long-term scientific value of the dataset for secondary analyses, methodological benchmarking, education and future immunological and virological research.
Figure 2. Figure 2. Data provenance workflow from experimental generation to FAIR-compliant repository deposition. The schematic illustrates the complete lifecycle of the dataset from biological experimentation to public data sharing and reuse. The workflow begins with the preclinical animal studies and sample collection, followed by immunological characterisation using ELISA and HIV-1 JR-FL TZM-bl pseudovirus neutralisation assays. Subsequent stages include quality-control procedures, statistical analyses, and dataset curation, encompassing data validation, metadata standardisation, provenance documentation, data dictionaries, and machine-readable file preparation. The curated resource is then deposited in the Figshare repository with a persistent digital object identifier (DOI), ensuring long-term accessibility and citability. The final stage highlights compliance with the FAIR (Findable, Accessible, Interoperable and Reusable) principles, demonstrating how structured metadata, documented provenance, reproducible analytical workflows, and standardised file formats maximise transparency, reproducibility, interoperability and the long-term scientific value of the dataset for secondary analyses, methodological benchmarking, education and future immunological and virological research.
Preprints 228060 g002
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.