Preprint
Article

This version is not peer-reviewed.

Development of an Evidence-Based Synthetic Dataset for AI-Driven Multi-Condition Health Screening Using IoT Sensors and Symptom Profiles

Submitted:

28 July 2026

Posted:

29 July 2026

You are already at the latest version

Abstract
Artificial-intelligence-driven screening on wearable and IoT-connected physiological sensors is limited by the scarcity of holistic, ethically shareable datasets that jointly capture sensor readings, symptoms, demographics, and clinically grounded condition labels. Randomly generated synthetic data can fill this gap in volume but typically fails to preserve the clinically meaningful correlations that a screening model must learn. This paper presents an evidence-to-data framework for constructing a synthetic multi-condition health-screening dataset from measurable IoT-derived physiological indicators (heart rate, blood pressure, peripheral oxygen saturation, body temperature, and electrocardiogram morphology) combined with structured symptom, demographic, and contextual profiles. Twenty-eight conditions spanning oxygenation, cardiovascular/hemodynamic, thermoregulatory, respiratory-infectious, gastrointestinal, urinary, neurological, psychophysiological, dermatological, and cardiac-rhythm categories are characterized from peer-reviewed literature and clinical guidelines, each with defined sensor thresholds, symptom profiles, demographic modifiers, confounders, and clinical exceptions. A cross-condition analysis quantifies feature overlap and identifies the minimal discriminative feature sets that separate clinically similar presentations. This evidence base is then formalized into a structured knowledge model and dataset schema in which every variable, conditional dependency, and label is traceable to a specific clinical finding, providing a reproducible foundation for generating, evaluating, and benchmarking screening-oriented machine-learning models. The framework was implemented end-to-end: a 200,000-record synthetic generator was built from this schema, and a multi-output Random Forest baseline trained on the resulting dataset achieves macro-F1 of 0.936 (Disease_Label), 0.933 (Screening_Category), and 0.993 (Screening_Outcome), with under 0.5% false-negative escalation on the highest-severity class, confirming that the encoded clinical structure is learnable. The resulting framework, generator, and baseline model are intended to support the development and preliminary validation of IoT-based multi-condition screening systems and are explicitly not intended for definitive clinical diagnosis.
Keywords: 
;  ;  ;  ;  ;  ;  ;  
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.
Prerpints.org logo

Preprints.org is a free preprint server supported by MDPI in Basel, Switzerland.

Subscribe

© 2026 MDPI (Basel, Switzerland) unless otherwise stated

Accessibility

Disclaimer

Terms of Use

Privacy Policy

Privacy Settings