Preprint
Review

This version is not peer-reviewed.

A Systematic Review of Cybersecurity Testbeds for Smart Environments: Architectures, Attack Coverage, and Defensive Evidence

A peer-reviewed version of this preprint was published in:
Future Internet 2026, 18(9), 445. https://doi.org/10.3390/fi18090445

Submitted:

29 July 2026

Posted:

30 July 2026

You are already at the latest version

Abstract
Smart-environment cybersecurity increasingly depends on experimental platforms that can reproduce attacks against buildings, homes, and cities under realistic conditions. However, the literature remains fragmented across testbed design, attack demonstration, and defensive validation. This makes it particularly difficult to judge what kind of security evidence each study actually provides. This review systematically analyzes 28 experimentally grounded studies published from 2020 onwards, focusing on how testbed realism, cyber-physical coupling, and evaluation mode shape the strength of the resulting claims. The corpus spans physical, hybrid, emulated, and dataset-driven environments across smart buildings, smart homes, and smart cities. Through our investigation, we discern a clear asymmetry in the field. Detection-oriented studies dominate, especially those based on emulation or public datasets, while live evidence for prevention, response, containment, and recovery is comparatively scarce. Availability and integrity/control attacks are the most frequently exercised, whereas authentication compromise and software exploitation remain rare because they are harder to stage on real hardware. Moreover, an important observation we arrive at is that physical and hardware-in-the-loop platforms support the strongest cyber-physical evidence, but emulated and replayed environments remain valuable for scale and reproducibility. At the same time, public datasets and offline classification results do not by themselves establish operational resilience in a live smart environment. To make these distinctions explicit, we introduce a cross-domain taxonomy of testbed architectures, attack families, and defensive control coverage, and map the evidence strength of reported mitigations using NIST cybersecurity framework-derived operational functions. Last, we identify open challenges, including weak recovery evaluation, limited reuse of reference testbeds, and the need for live, context-aware datasets, outlining promising future directions.
Keywords: 
;  ;  ;  ;  ;  ;  ;  ;  

1. Introduction

Smart environments increasingly integrate sensing, actuation, networking, edge computing, cloud services, and automated decision making into systems that directly affect occupants, assets, and physical processes [1,2]. Indicatively, wireless sensor networks frequently constitute a fundamental part of the sensing infrastructure in a smart environment; however, their resource constraints, distributed organization, and reliance on wireless communication introduce specific security challenges [3]. Examples include smart buildings that coordinate heating, ventilation, access control, lighting, and energy management systems, smart homes that integrate consumer Internet of Things (IoT) devices, mobile applications, vendor hubs, and automation rules, and smart cities that extend these capabilities across large-scale urban infrastructure and services. Although these systems improve efficiency, comfort, and sustainability, their increased connectivity also expands the attack surface. As a result, adversaries can exploit vulnerabilities to monitor sensitive behaviour, disrupt system availability, inject false data, manipulate control commands, or compromise physical operations [4,5,6].
The consequences of these attacks extend beyond traditional information security risks. In smart environments, compromising a device or communication channel can directly influence the physical processes under control. For example, false sensor readings may cause incorrect Heating, Ventilation, and Air Conditioning (HVAC) control decisions, unauthorised commands may alter lighting, access, or temperature settings, and Denial-of-Service (DoS) attacks may disrupt communication between controllers, gateways, and cloud services. Consequently, evaluating security mechanisms solely through offline classification accuracy, such as benchmark datasets, or generic network traffic analysis is often insufficient. Meaningful validation instead requires experimental environments that can reproduce realistic combinations of devices, protocols, attack paths, control logic, and, where applicable, the resulting cyber-physical consequences.
Cybersecurity testbeds provide this experimental foundation [6,7]. Depending on their design, they can support attacks against physical devices, Hardware-in-the-Loop (HitL) controllers, emulated networks, simulated processes, or replayed traffic datasets. However, these approaches do not provide equivalent evidence. In this respect, fully physical platforms can capture device-specific firmware behaviour, wireless effects, timing variation, and proprietary implementation details, but they are costly and difficult to scale or reproduce. Hybrid and HitL configurations can connect real controllers to simulated plants and thus enable controlled cyber-physical experimentation. In contrast, emulated environments facilitate large-scale, repeatable scenarios, while replayed or public datasets are useful for benchmarking detection algorithms but cannot independently demonstrate that a control prevents, contains, or recovers from a live attack. Therefore, the architecture of a testbed directly determines both the attacks it can credibly reproduce and the strength of the security claims it can support.
To the best of our knowledge, existing surveys provide important but fragmented foundations. Namely, prior work has surveyed general IoT cybersecurity testbeds, broad IoT testing methodologies, model-based security testing, building automation security, cyber-physical security in Building Automation Systems (BASs), and smart-city threats and countermeasures. Nevertheless, these streams usually focus either on the target system and its threats or on the experimental infrastructure and its generic properties. In other words, they do not consistently connect testbed realism, cyber-physical coupling, trust-boundary completeness, and tooling choices to the demonstrated attack surface and the operational evidence for defensive controls. As a result, the literature still lacks a cross-domain account of what smart-environment cybersecurity testbeds actually validate, rather than merely what attacks or controls they describe.
Contribution: This paper aims to address this gap through a systematic review of cybersecurity testbeds for smart buildings, homes, and cities through four key contributions. First, it develops a cross-domain taxonomy of smart-environment cybersecurity testbeds that characterizes realism strategy, cyber-physical coupling, trust boundaries, and enabling tools. Second, it maps the demonstrated attack surface across smart buildings, smart homes, and smart cities, including the protocols, attack families, and evaluation modes represented in the identified corpus. Third, it cross-analyzes attacks and defensive controls using four operational functions derived from the National Institute of Standards and Technology (NIST) Cybersecurity Framework [8]: preventive hardening, detection, response and containment, and recovery or safe-state restoration. This analysis distinguishes implemented and evaluated controls from mitigation recommendations, exposing the concentration of current evidence in detection-oriented studies. Fourth, it identifies a forward-looking research agenda centred on closed-loop resilience evaluation, live and context-aware datasets, explicit cloud and edge trust boundaries, reproducible reference testbeds, and controlled Artificial Intelligence (AI)- and Large Language Models (LLM)-assisted testbed engineering.
The remainder of this paper is structured as follows. The next section presents related surveys in this field. Section 3 details our methodology. Section 4 summarizes the included corpus. Section 5 analyzes testbed architectures and composition. Section 6 examines the enabling tools and technologies. Section 7 maps the attack surface covered by the surveyed testbeds. Section 8 reviews the defensive controls and mitigation mechanisms. Section 9 cross-analyzes attacks and defenses. Section 10 discusses open challenges and future research directions. The last section concludes the paper.

3. Methodology

This work provides a systematic review of cybersecurity testbeds for smart environments, specifically for smart buildings, cities, and homes, conducted according to the framework of [18] and [19], reported in accordance with the Preferred Reporting Items for Systematic reviews and Meta-Analyses (PRISMA) statement [20], which has been widely applied in related cybersecurity reviews [6,21]. The identification, screening, eligibility, and inclusion process is summarised in Figure 1.
Step I – Research Questions. The review is organized around four research questions, in one-to-one correspondence with contributions C1–C4, as defined in Section 1.
  • RQ1: How are cybersecurity testbeds for smart buildings, smart homes, and smart cities architected, particularly with respect to implementation paradigm (physical, hybrid/HitL, emulated, or simulated), cyber-physical coupling, trust boundary, and enabling tooling?
  • RQ2: Which smart-environment domains, protocols, attack families, and attack-evaluation modes are represented in the surveyed testbeds, and which threats are exercised through live, HitL, emulated, or replayed/dataset-based experiments?
  • RQ3: To what extent do the surveyed testbeds provide operational evidence for the NIST cybersecurity framework-derived functions [8], and how does the strength of that evidence depend on testbed realism and evaluation mode?
  • RQ4: What limitations and research gaps arise from the observed trade-offs among fidelity, reproducibility, scalability, trust-boundary completeness, and defensive coverage, and which future directions, including reusable reference testbeds and AI-/LLM-assisted testing, could address them?
Step II – Bibliographic Sources. Three major databases were searched: IEEE Xplore, Scopus, and Web of Science. These sources were selected as they offer broad coverage of computer science, engineering, and security research and are widely acknowledged as appropriate repositories for systematic reviews in this area [22]. Relying on multiple databases reduced retrieval bias and improved the likelihood of capturing relevant studies across publication venues. The duration for the completion of the systematic review was approximately four calendar months, from Mar. 2026 to Jun. 2026. The search was scoped from 2020 onwards to concentrate the review on the state-of-the-art and to capture the recent surge of smart-environment testbed research; some seminal earlier works identified through snowballing are retained for context where they are foundational to the reviewed testbeds [23,24,25,26,27,28].
Step III – Search Terms. An initial broad strategy returned an unmanageably large, off-topic set of records. The strategy was iteratively calibrated against a quasi-gold standard [29] of relevant primary studies identified through preliminary searching and snowballing, and refined until it retrieved the known-relevant set with high precision. The resulting strategy uses three conceptual groups, i.e., domain, testbed/resource, and security terms, combined as (G1 AND G2 AND G3), with synonyms OR-combined within each group (Table 2). A broad methodological group (evaluation, framework, experiment, assessment) was deliberately excluded, as it matched nearly all computer-science records without improving relevance; queries were restricted to the title and abstract fields, and the publication window was constrained to 2020-2026. The complete database-specific search strings and calibration rationale are provided in Appendix A.
Step IV – Practical Screening. Screening was performed in three stages: de-duplication; title/abstract screening against the eligibility criteria of Table 3; and full-text examination. Disagreements were resolved by discussion until consensus was reached. Only studies that explicitly designed, implemented, or employed a cybersecurity testbed, emulation/simulation platform, or experimentally generated dataset for smart buildings, homes, or cities were retained. To further reduce subjectivity, a quality appraisal was applied to each study based on methodological clarity, relevance to the research questions, completeness of reporting, and overall contribution to the field.
Full-text examination initially surfaced 41 candidate records that passed title/abstract screening. Close reading against criteria IC1-IC4 and EC1-EC3 led to the exclusion of ten borderline studies that, on closer inspection, did not satisfy the eligibility criterial; primarily because they contributed methodological or cryptographic artifacts rather than an experimental resource, lacked a genuine smart-environment testbed instantiation, or fell outside the review’s scope [30,31,32,33,34,35,36,37,38,39]. After these exclusions, the process yielded 28 included studies, summarised in Table 4.

4. Corpus Overview

The final corpus comprises 28 experimentally grounded studies published between 2020 and 2026. Table 4 provides a study-level overview, grouped by smart-environment domain, while Figure 2 shows the temporal distribution of the included literature. As observed from both the table and figure, the corpus contains ten smart-building studies, 16 smart-home studies, and two smart-city studies. This distribution reflects both the relative maturity of smart-home security experimentation and the differing practical constraints of the three domains, i.e., consumer smart-home platforms permit comparatively accessible physical deployments, whereas building and city infrastructures more often require hybrid or fully emulated approaches.
The temporal distribution indicates a recent expansion of the field. Although the corpus includes early studies from 2020-2022, most included work was published from 2023 onward, with half of the studies appearing between 2024 and 2026. Earlier contributions predominantly establish device-level attack scenarios, protocol analyses, or proof-of-concept detection mechanisms, whereas more recent studies combine cyber experimentation with physical-process modelling, federated learning, embedded detection, SDN-based mitigation, fuzzing, and large-scale emulation. Examples include HitL building-energy testbeds [40,41], real-device and multi-vendor smart-home deployments [42,43,44], and scalable emulated smart-city environments [45,46].
The remainder of the review analyzes the corpus along three dimensions. Specifically, Section 5 examines how the included studies trade off realism, cyber-physical coupling, and trust-boundary completeness. Section 6 then identifies the hardware, software, and cloud components on which these environments depend. Finally, Section 7, Section 8, and Section 9 assess the attacks that are exercised, the controls that are implemented, and the extent to which current smart-environment testbeds provide evidence beyond detection alone.
Table 4. Overview of the studies included in the final review, grouped by application domain and ordered reverse-chronologically within each group.
Table 4. Overview of the studies included in the final review, grouped by application domain and ordered reverse-chronologically within each group.
Study Year Scope and principal contribution
Smart buildings
Weng et al. [47] 2026 Control-system emulation: SBCSE, an open-source MQTT-based platform for MITM/DDoS case studies and countermeasure validation.
Morales-Gonzalez et al. [48] 2025 BAS fuzzing: survey plus a case study on a real BACnet/SC testbed; a Yabe-based fuzzer found two array-OOB DoS bugs across four devices.
Li et al. [40] 2024 HIL BAS testbed generating labelled attack/fault datasets; joint CRF/WPM-FPCA classifier separates attacks from faults at 90.2%.
Balamurugan et al. [49] 2024 BUILD-SOS HIL testbed: Alfalfa simulation, virtual BACnet devices, emulated OT networks, and remote HIL for attack/fault data.
Li et al. [41] 2024 HIL cyber-physical BAS testbed pairing a real-time Modelica building/HVAC emulator with real controllers; generates operational and network datasets.
Cash [50] 2024 BACnet/KNX analysis: automated BACnet enumeration on a campus BAS, a KNX MITM FDI attack detected via a JSD IAT feature (SVM 100%), and attack scenarios.
Runge et al. [51] 2023 HVAC FDI detection: grey-box heat-balance + CUSUM detection of stealthy attacks on a real VAV/HVAC BAS testbed, documenting deployment pitfalls.
Li et al. [52] 2022 HIL BAS testbed pairing a real BACnet BAS with a Modelica HVAC emulator for attack/fault injection and dataset generation.
Rondon et al. [53] 2020 Enterprise IoT: PoisonIvy attacks abusing EIoT (Crestron/Control4) drivers for DoS, botnet control, and resource abuse; real EIoT testbed.
Wu et al. [54] 2020 CPS FSM detection: finite-state-machine framework mapping sniffed ZigBee clusters to device states to detect control-logic tampering; real lighting testbed.
Smart homes
Rahman et al. [55] 2026 ARP/MITM detection: cross-protocol ARP–DHCP binding validator; 100% detection over four public datasets and two real testbeds.
Kostage et al. [56] 2025 Hierarchical FL IDS + HCI: HiFINS distributes training across router/edge/cloud with a human-centered UI; container-based Chameleon testbed on live traffic + CICIoT2023.
Barman et al. [57] 2025 Two-tier FL IDS: SRF2T-ID couples a secure Wireless-of-Things link and HIDS with an outer FL NIDS using nature-inspired feature selection; nRF24L01/Arduino testbed, 99.9%.
Javed et al. [58] 2024 Embedded two-layer ML IDS: TinyML/XGBoost on an ESP32 thermostat plus a cloud IDS; real testbed with a Pi adversary yields the IDSH dataset (97.6/99.5%).
Karmous et al. [59] 2024 SDN MitM IDPS for MQTT smart homes; CNN reaches 99.96% and deploys on a Ryu-SDN testbed across single/tree/mesh topologies.
Li et al. [42] 2024 Knowledge-graph detection: SeIoT, a KG-based bimodal detector from traffic sniffing; 19-device testbed over two months vs. device/platform attacks (97.7% TPR).
Zou et al. [43] 2023 Wi-Fi privacy snooping: IoTBeholder, a low-cost COTS attack inferring device type/location and habits from encrypted Wi-Fi; 23-device testbed.
Yeboah-Ofori et al. [60] 2023 Evil Twin (assistive IoT): physical testbed (Kali/Pi, camera, lock, TV, Echo, Suricata/Snort) demonstrating a Kill-Chain Evil Twin/MITM attack, with mitigations.
Jiang et al. [61] 2022 Anomaly detection: time-interval-aware event-sequence detection catching delay anomalies; 88–93% accuracy on a real-world testbed.
He et al. [62] 2022 Bi-layer IDS (supervised ML + auto-generated benign rules) on a four-device testbed; 0% false/missing alarms, 99.51% F1.
Chi et al. [44] 2022 Automation interference: seven Delay-based Automation Interference attacks via selective TLS-hijacking delays; validated in two real apartment testbeds.
Dai et al. [63] 2022 Context-based anomaly detection: HomeGuardian fuses temporal and environmental context in a learning classifier; self-configured Home Assistant testbed, F1>0.90.
Liu et al. [64] 2021 Anti-sniffing defense: non-intrusive “phantom user” decoy-packet injection cutting sniffer behaviour-inference accuracy from 94.8% to 3.5% on a real testbed.
Fu et al. [65] 2021 Semantics-aware detection: HAWatcher mines app/device correlations and uses Shadow Execution to flag state inconsistencies; SmartThings, four testbeds, 62 cases.
Akestoridis et al. [66] 2020 Zigbee analysis: Zigator exploits disabled MAC-layer security for selective jamming/spoofing to expose the network key; commercial-device testbed.
Arif et al. [67] 2020 Consortium blockchain: security framework cutting request time, storage, and verification burden; ESP32-based household IoT testbed.
Smart cities
Tariq et al. [45] 2025 DL botnet IDS: Stacked-Autoencoder–GRU (pruned/quantized) on a 4-week dataset from a 9-device emulated smart-city testbed; 98.65% accuracy.
Escolar et al. [46] 2024 Pre-6G/5G honeynet with network-slicing to mitigate volumetric DDoS; emulated testbed scaling to 1M devices and 19 Gbps.

5. Testbed Architectures and Composition

This section characterizes the architecture of the 28 testbeds identified in Section 3. We organize the analysis around three architectural questions that recur throughout the corpus and can jointly bound the external validity of any security claim a testbed can support [68,69]. Namely, we investigate i) how realistically it reproduces the target environment (Section 5.1); ii) how tightly it couples the cyber and physical planes (Section 5.2); and iii) where it draws its trust boundary (Section 5.3). Throughout, we relate these choices to the experiments they enable, foreshadowing the attack surface and the defensive coverage analyzed later in Section 7 and Section 8, respectively. Figure 3 summarizes the resulting architectural taxonomy and the distribution of studies across it.

5.1. Realism Strategy: Physical, Hybrid, and Emulated Testbeds

The most consequential architectural decision is the degree of realism, which governs the fundamental trade-off between behavioural fidelity on the one hand and scale, reconfigurability, and reproducibility on the other. Across the corpus, we observe a spectrum bounded by two edges and a middle ground, and position on this spectrum is strongly predicted by application domain, as shown in Figure 3. Specifically, at one end sit fully physical testbeds, in which Commercial Off-The-Shelf (COTS) devices are deployed and observed directly under live traffic. This strategy dominates the smart-home studies, where authentic timing, radio characteristics and firmware behaviour are indispensable. Representative instances include the consumer-platform homes [64,65,66], the large multi-vendor deployments [43,44], and the assistive-technology scenario [60]. Altogether, physical realism maximizes fidelity but contains scale and impairs repeatability, since every experimental condition should be reproduced on real, stateful devices.
At the opposite end sit fully emulated testbeds, in which no physical endpoint is present, and device behaviour is synthesized in software. This strategy characterizes the network- and detection-oriented studies, where the object of study is traffic rather than device internals. For example, the Software-Defined Networking (SDN) intrusion-prevention testbed [59] emulates hosts and switches entirely in software; smart-city botnet [45] models commercial device profiles without deploying them; 5G honeynet [46] emulates all IoT traffic on virtualized infrastructure; and [49] constructs a wholly virtual building environment through co-simulation. In this context, emulation may sacrifice fidelity, but at the same time it enables scale, repeatability and the on-demand generation of large, labelled datasets. Between these poles lies the largest group: hybrid testbeds that interpose real controllers or edge nodes between emulated or simulated components. This is mostly observed in the smart-building domain, where HitL designs connect genuine BAS controllers to a real-time emulation of the physical plant [40,41,52]. In other words, the hybrid strategy is the natural response to a target that is a heavy industrial plant, namely, one cannot repeatedly attack a live chiller or air-handling unit, yet the controller logic and network behaviour under attack remain authentic. Thus, hybrid designs preserve fidelity where it is security-relevant (e.g., the controller and its network) while emulating what cannot be attacked directly (e.g., the physical process).

5.2. Cyber-Physical Coupling

Orthogonal to the realism strategy is the extent to which a testbed models the physical process under control. Prominently, in a smart-home testbed, an attack typically succeeds or fails in the information domain; in a building- or city-scale testbed, it should ultimately perturb a thermal, electrical or mechanical quantity, and reproducing that consequence requires an explicit process model. The building studies achieve this through co-simulation, coupling building-thermodynamics models to real controllers via HitL interfaces [40,41,52] or calibrating a physics-based model against a live system, as in the grey-box heat-balance formulation [51]. This coupling is not merely a matter of realism, as it is what makes an entire class of defences possible. Particularly, physics-informed detectors exploit the residual between observed and physically expected behaviour; a signal that simply does not exist in a purely network-level testbed and consequently appears only where such coupling is present[40,51,54]. Notably, the work in [54] is the sole study to extend explicit cyber-physical state fusion into the smart-home domain, coupling network traffic with physical device state so that inconsistencies between the two betray an attack.

5.3. Trust Boundary and External Validity

The third architectural question is where a testbed draws its trust boundary, i.e., which components it treats as inside the SUT, and which it assumes are benign or simply inherits from outside the laboratory. Simply put, this choice directly determines the completeness of the modelled attack surface and, ultimately, the external validity of the experimentation results derived from the given testbed. A recurring and under-acknowledged issue is the implicit extension of the trust boundary through dependence on vendor cloud back-ends. That is, several smart-home testbeds cannot operate without an external platform cloud [44,61,64,65]. Because such services are integral to device behaviour yet lie outside the instrumented environment, the testbed’s true trust boundary silently extends beyond the laboratory, and any cloud-mediated attack path is left unmodelled. This stands in contrast to the smart-city and cloud-augmented detection studies, where an external platform is instead an elective compute resource provisioned for scale rather than an intrinsic part of the SUT. Overall, the broader consequence is that, as observed from the examined corpus, there is a predominantly instrumental use of the edge and network tiers of the smart environment stack while leaving the cloud tier and the attacks that traverse it comparatively under-represented. Secondly, the realism strategy interacts with the trust boundary. Specifically, as physical testbeds are costly to scale, they are typically small, which, in turn, constrains the generality of the resulting security claims; emulated testbeds scale freely but risk omitting behaviours, including timing jitter, radio imperfections, and firmware intricacies that are themselves exploitable. Therefore, each architectural family embeds a characteristic threat to external validity.

5.4. Synthesis

Taken together, these three dimensions reveal a field organized less by a shared reference architecture than by domain-imposed constraints, as summarized in Figure 3. Specifically, smart-home testbeds favour physical, moderate-scale deployments whose trust boundary is quietly extended by vendor clouds; smart-building testbeds converge on hybrid HitL designs that foreground cyber-physical coupling; and the nascent smart-city works rely on fully emulated, cloud-provisioned schemes. However, these choices are not neutral. They determine the attacks a testbed can credibly stage, the defences it can validate, later discussed in Section 7 and Section 8, and the confidence with which either can be generalized beyond the laboratory.

6. Testbed Tooling

In this section, we inventory the tooling of the 28 primary studies across three layers: hardware, software, and as-a-service components, consolidated per study in Table 5, summarizing the most prevalent tools in Figure 4. Beyond the inventory itself, we aim to expose the de facto tool stack on which smart-environment security research rests and to assess what its degree of openness and convergence implies for the reproducibility and comparability of results. Moreover, we synthesize cross-cutting patterns rather than restating the table, and defer the architectural consequences of these choices, with reference to Section 5, from which they follow.

6.1. Hardware Tooling

The hardware layer partitions cleanly into three functional roles: target and endpoints, adversary platforms, and instrumentation devices. Targets are the Devices Under Test (DUT), and their selection follows the domain split. Namely, smart-home studies draw on COTS consumer ecosystems, including Samsung SmartThings hubs and sensors [61,64,65,66] and Philips Hue, Aqara, Amazon Echo, and Xiaomi devices [42,43,44]. On the other hand, smart-building studies target industrial controllers such as Siemens field panels [50] and Automated Logic Corporation (ALC) BAS controllers [40,41,48,52]. A distinct sub-class narrows the target to a single constrained endpoint to study embedded defence, for example, the ESP32 microcontroller anchors the thermostat [58] and the blockchain nodes [67], while an Arduino/nRF24L01 pairing forms the wireless target [57]. At the same time, testbed scale varies and is domain-dependent: smart-home testbeds combine moderate device counts with high vendor heterogeneity; for example, in [43] 23 devices were utilized across seven vendors, while a comparably diverse multi-platform was identified in [44], with 19 devices from four manufacturers in [42]. On the other hand, smart-building testbeds pair a handful of high-cost controllers, and the smart-city studies are effectively device-free at the hardware layer [45,46].
Regarding the second functional role, adversary and general-purpose compute platforms appear markedly more standardized. The single-board computer, overwhelmingly the Raspberry Pi, is the workhorse of the corpus, serving interchangeably as gateway, edge detector, or infected host [42,44,50,54,55,64]. The attacker is most often a Kali Linux machine [55,59,60]), a convergence that mirrors the maturity of the offensive software stack, later discussed in Section 6.2. Finally, instrumentation hardware is dominated by radio-capture devices, reflecting the wireless nature of the domain. Indicatively, in the Zigbee analyses, we pinpointed reliance on dedicated sniffers and Software-Defined Radios (SDRs), namely, the Texas Instruments CC2531 [54], ATUSB and USRP N210 [66], TelosB dongle [64], and nRF52840 [42]. Moreover, wideband capture extends to the HackRF One SDR [54] and to monitor-mode Wi-Fi adapters such as the ALFA AWUS036ACH [60]. In the smart-building domain, instrumentation is instead physical-signal-oriented, that is, dSPACE SCALEXIO real-time hardware with analogue/digital I/O boards bridges the cyber and physical planes [41,52]. At the extremes of the corpus, several testbeds report no dedicated hardware, being fully emulated [45,46,49] or using commodity laptops purely as compute simulation [47].

6.2. Software Tooling

The software layer is the richest and most revealing, and it stratifies into four tiers: i) packet-capture instrumentation, ii) offensive tooling, iii) simulation and protocol stacks, and iv) ML frameworks.
  • Packet capture and traffic instrumentation. The single most pervasive dependency in the entire corpus is the Wireshark/tcpdump capture stack, which appears in the majority of testbeds regardless of domain [40,45,50,52,54,55,56,58,60,62,66]. Flow-level feature extraction via CICFlowMeter recurs in the detection-oriented studies [62], producing the tabular inputs consumed by the ML tier. This near-universal reliance establishes packet capture as the effective common language of the field and, encouragingly, one grounded in open tooling.
  • Offensive tooling. Studies that stage live attacks converge on a compact, mature, open-source arsenal. Namely, Ettercap and Bettercap drive Man-in-the-Middle (MitM) and Address Resolution Protocol (ARP)-spoofing attacks [55,59,60]; Scapy is the common frame-crafting library [42,55,66]; Nmap performs reconnaissance [50,60]; and the wireless suite of Aircrack-ng, Airgeddon and Hping3 supports the rogue-access-point and DoS [60]. Moreover, as recognized in the surveyed studies, Zigbee attacks are enabled by the killerbee toolkit [64] and custom tools [66]. For fuzzing-based vulnerability discovery, Boofuzz, AFL++, AFLNet and WinAFL are employed [48] against BACnet/KNX devices.
  • Simulation, emulation, and protocol stacks. Domain divergence is sharpest in this tier. Specifically, smart-building testbeds are commonly built on a building-physics co-simulation stack, including the Modelica Buildings library with EnergyPlus, Dymola and BOPTEST/Alfalfa [40,41,49,52], interfaced to industrial protocols through open BACnet stacks such as BACpypes/BAC0 [40,49,50] and the Calimero KNX library [50]. Network-oriented studies instead build on the Mininet emulator, frequently paired with an SDN controller, including OpenFlow [45,59], and OpenDaylight with Open vSwitch [46]. In the smart home, the application layer is dominated by MQTT [47,58,59] and by home-automation runtimes, including OpenHAB [54,64], Home Assistant [63] and the SmartThings SmartApps platform [65].
  • Machine-learning frameworks. Classical pipelines rest on scikit-learn with gradient-boosted trees, most commonly XGBoost [42,43,56,58,62]; deep-learning work uses TensorFlow and PyTorch [42,45,56], with federated learning coordinated through the Flower framework [56]. A notable enabling pattern is the reuse of public benchmark datasets to supply attack traffic instead of, or alongside, live generation: CICIoT2023 [56,57], IoT-23 [62], CICIDS2017 [57] and several public ARP/IoT capture sets [55]. This practice improves label quality and comparability, but also decouples a subset of detectors from the very testbeds nominally under study.

6.3. As-a-Service Dependency

The as-a-service layer is the sparsest of the three; only 8 of the 28 testbeds incorporate an external service. Intrinsic vendor cloud, integral to normal device operation, appears as the Samsung SmartThings cloud [44,61,64,65], with the latter additionally depending on the Apple HomeKit, Aqara, Hue and Alexa back-ends. Elective cloud compute, provisioned for scale or to host a component under test, appears as Chameleon Cloud and AWS [56], Amazon EC2 [45], AWS Elastic Beanstalk [53], and a generic Infrastructure-as-a-Service (IaaS) virtual machines hosting the cloud-side detector [58].

6.4. Synthesis: Openness, Convergence, and Reproducibility

Three observations emerge from the tooling landscape. First, the field rests overwhelmingly on open-source instrumentation, offensive tooling and machine-learning frameworks; proprietary components are confined to vendor BAS software (e.g. WebCTRL, Desigo CC) and to the intrinsic clouds of consumer platforms; thus supporting reproducibility. Second, there is pronounced convergence within each layer, namely, Wireshark for capture, a small Kali-based offensive suite, Modelica/EnergyPlus for building physics, Mininet/SDN for network emulation, and scikit-learn/XGBoost for detection (Figure 4), which aids comparability but concentrates the community’s methodological assumptions in a handful of tools. Third, counterbalancing the first two, reproducibility remains constrained by the persistent presence of bespoke components, including custom Scapy scripts and one-off knowledge-graph and shadow-execution engines, along with proprietary hardware and clouds that cannot be reconstituted elsewhere.
Table 5. Hardware, software, and as-a-service tooling used to build the testbeds of the 28 included studies, grouped by application domain.
Table 5. Hardware, software, and as-a-service tooling used to build the testbeds of the 28 included studies, grouped by application domain.
Study Hardware Software As-a-service / Cloud
Smart buildings
Weng et al. [47] None (SW emulation; MacBook compute only) Python; Mosquitto MQTT broker; MQTTS
Morales-Gonzalez et al. [48] ALC BAS hardware: G5RE routers, ZN551 zone controller, Belimo actuator, Optiflex, Secure Connect Hub boofuzz; AFL++; AFLNet; WinAFL; Yabe; WebCtrl; Valgrind
Li et al. [40] Real HVAC controllers (chiller, boiler, AHU, VAV); A/D–D/A boards; BAS server Modelica Buildings Library; MATLAB; EnergyPlus; WebCtrl; Wireshark; Docker; Python; Isolation Forest; CRF
Balamurugan et al. [49] None (fully virtual BUILD-SOS testbed) Alfalfa (NREL); OpenStudio; EnergyPlus; Modelica; Spawn-of-EnergyPlus; BOPTEST; Docker; BACpypes; BAC0
Li et al. [41] dSPACE SCALEXIO HIL rig (A/D, D/A, I/O); Dell Precision workstation; ALC controllers; ARCnet router dSPACE ConfigurationDesk/ControlDesk; Dymola/Modelica; WebCTRL; EIKON-Logic; Python; Wireshark
Cash [50] Siemens BACnet field panels; QMX/QAM controllers; 2× Raspberry Pi 3 + KNX HAT MITM rig BACpypes; Raspbian; Calimero (KNX); Dymola/Modelica; Siemens Desigo CC; Nmap/NSE; Wireshark
Runge [51] Real campus VAV HVAC system (temp/airflow sensors) via BAS Web-based BAS platform; grey-box heat-balance model; Levenberg–Marquardt; non-parametric CUSUM
Li et al. [52] BAS server; ALC controllers (chiller, AHU, VAV); real-time HIL emulator; A/D–D/A boards; BACnet-to-ARCnet router dSPACE ControlDesk; Modelica Buildings Library; Python; Wireshark; BACnet/IP
Rondon et al. [53] Control4 EA-1 controller + SR-260; LG hospitality TV; TP-Link router; Razer laptop VS Code; Control4 Driver Editor/Composer; LUA; Jersey JAX-RS; Swagger Amazon AWS Elastic Beanstalk
Wu et al. [54] Raspberry Pi 3; TI CC2531 ZigBee sniffer; Nortek HUSBZB-1 dongle; GE switch; Centralite sensor; HackRF One; Wi-Fi router OpenHAB; Wireshark; MATLAB/Simulink; Postman; deterministic FSM model
Smart homes
Rahman et al. [55] 2× Raspberry Pi 4 testbeds; Philips Hue Zigbee hub; Zigbee bulb/sensor/switch; camera; doorbell; Kali laptop Raspberry Pi OS; Python; Scapy; Bettercap; four public IoT/ARP datasets (offline validation)
Kostage et al. [56] Intel NUC server; OpenWRT router; Raspberry Pi + Jetson Nano edge nodes; smart speakers/cameras/plugs TensorFlow; Flower (FL); gRPC; Flask; Paramiko; scikit-learn; Docker; Ubuntu; CICIoT2023 Chameleon Cloud (bare-metal + GPU); AWS VMs (FL edge servers)
Barman et al. [57] nRF24L01 PA/LNA transceiver; Arduino Uno (Wireless-of-Things testbed) AES-128; GWO/PSO/GA/MayFly feature selection; ensemble ML; FL framework; CICIDS2017; CICIoT2023
Javed et al. [58] ESP32 smart thermostat; Raspberry Pi 4 (MITM/DoS adversary + fog IDS); PC/gateway edge nodes; Wi-Fi router XGBoost (embedded + cloud IDS); ANN/RF/DT; TinyML; lwIP + socket lib (ESP32); Wireshark; tcpdump; self-collected IDSH dataset Cloud VM hosting thermostat site + cloud-side multiclass IDS
Karmous et al. [59] Kali laptop (attacker); Ubuntu VMs; emulated OpenFlow switches/hosts Mininet; Ryu SDN controller; OpenFlow; Python; Mosquitto MQTT; Ettercap; Macof; CNN/NB/kNN/RF; self-generated MQTT/ARP dataset
Li et al. [42] Raspberry Pi 4B; TI nRF52840 BLE sniffer; 19 Mijia IoT devices; Wi-Fi router (port mirror); 2× Pi malware nodes Python; Scapy; scikit-learn; PyTorch (FS-HAN); Mirai/Bashlite samples; custom knowledge-graph framework
Zou et al. [43] MacBook (802.11 sniffer); 23 IoT devices, 7 vendors (Xiaomi, TP-Link, Aqara, Amazon Echo, etc.); Xiaomi gateway; plug/bulb/lock XGBoost; MLP; Random Forest; custom 802.11 packet-parsing code
Yeboah-Ofori et al. [60] OKdo ROCK 4 (Kali); ALFA AWUS036ACH adapter; NetGear router; Foscam camera; Amazon Echo; smart TV/lock Kali Linux; VirtualBox; Airgeddon; Aircrack-ng; Bettercap; Ettercap; Nmap; Hping3; Wireshark; Suricata; Snort
Jiang et al. [61] SmartThings hub; motion/contact sensors; light bulb; ThreeReality smart switch SMOTE; SmartThings app (event-log export) Samsung SmartThings cloud
He et al. [62] 4 commercial devices: XiaoDu speaker, QingPing temp monitor, TP-Link cam, GoSund plug; Wi-Fi adapter Wireshark; tcpdump; CICFlowmeter; scikit-learn; XGBoost + LR/KNN/DT/SVM/GBDT/RF/MLP; IoT-23
Chi et al. [44] Raspberry Pi 4 (ARP relay); 55 devices across 4 platforms (SmartThings, HomeKit, Aqara, Hue); sensors, locks, plugs, cameras iptables; ARP-spoofing script; signature recognizer; DFA rule-inference tool; IoT Inspector SmartThings; HomeKit/iCloud; Aqara; Hue; Alexa; WeMo; Arlo clouds
Dai et al. [63] Real devices (lights, sensors, camera, speaker, HVAC, plug) bound to Home Assistant Home Assistant; virtual-device simulator; word2vec; NLP correlation; classifier; CASAS HH114 dataset
Liu et al. [64] Samsung SmartThings hub; many Zigbee devices (sensors, bulbs, plugs, siren); TelosB dongle; Raspberry Pi 3B (OpenHAB decoy) killerbee (Zigbee sniff/inject); Hidden Markov Model; Levenshtein ratio; OpenHAB Samsung SmartThings cloud
Fu et al. [65] 4 real-home SmartThings testbeds: motion/contact/water/presence sensors; buttons; plugs; switches; bulbs SmartThings smart apps; custom shadow-execution engine; NLP semantic correlation; hypothesis testing Samsung SmartThings cloud
Akestoridis et al. [66] USRP N210 SDR; ATUSB (modified firmware); SmartThings hub; outlet/motion/bulb sensors; Schlage & Yale locks; Centralite outlet Zigator (custom); GNU Radio (gr-ieee802-15-4); Wireshark; forked Scapy; PyCryptodome; SQLite; scikit-learn
Arif et al. [67] 4× ESP32 modules; temp/humidity sensor; buzzer; LED; relay; display; laptop (P2P server) Espruino (JS); custom consortium-blockchain (PoW mining) code
Smart cities
Tariq et al. [45] None (emulated; device profiles model Siemens/Bosch/Honeywell/GE products) Mininet; Wireshark; tcpdump; iperf; tc; NumPy; SciPy; TensorFlow; PyTorch; SAE-GRU; PowerBI Amazon EC2 (emulated cloud)
Escolar et al. [46] 2 physical servers (Xeon) hosting VMs; no physical IoT hardware (traffic emulated) Open vSwitch; QEMU/libvirt; CentOS; OpenDaylight SDN; Snort nIDS; custom IoT traffic emulator; VXLAN/GTP

7. Attacks Classification

This section classifies the attacks exercised across the corpus into the eight families depicted in Figure 5. Our aim is not merely to report aggregate counts but to make the attack-study mapping explicit, so that a reader can trace any category back to the concrete evidence that populates it.
Before enumerating the families, one methodological distinction should be drawn, because it conditions how the counts in Figure 5 should be read. Namely, the corpus divides into two modes of adversarial demonstration. In the first, an attack is actively launched against live testbed components – a rogue device floods a controller, an adversarial host injects forged packets, a radio jams a channel – and its effect is observed directly. In the second, prevalent among detection-oriented studies, the “attack” enters the testbed as labelled malicious traffic, either replayed from a public intrusion-detection dataset or generated by a traffic emulator, so that a classifier can be trained and evaluated rather than compromising a live system. We flag this distinction where it applies, as it materially bounds the causal claims a testbed can support (Section 5). Notably, two studies execute no attack at all. Specifically, the work in [59] conducts a purely formal resistance analysis under Burrows-Abadi-Needham (BAN) logic, ROR and ProVerif, while its testbed measures only computational and energy overhead, and the study in [67] benchmarks blockchain block-mining performance while merely surveying the threats it names.

7.1. Availability Attacks

Availability is the most frequently exercised attack objective in the corpus, appearing in fourteen studies and dominated by DoS and its distributed variant (DDoS). The BAS subset favours a protocol-native form of denial, including rogue BACnet traffic that drives a controller into a reinitialisation or soft-reboot loop, demonstrated on HitL air-handling-unit rigs [40,41,52] and via register flooding on the BUILD-SOS testbed [49]. In the smart-home and smart-city subsets, DoS/DDoS is more often volumetric and botnet-driven. Specifically, the work in [47] floods an MQTT broker with a thousand emulated bots, in [53] the authors turn a compromised Control4 hub into a participant in a remote DDoS campaign, and the study in [58] drives an hping3 flood against an ESP32 thermostat. The detection-oriented studies observe DoS/DDoS as a traffic class rather than a live disruption, i.e., the works in [42,56,57,62,62] evaluated classifiers against DDoS flows drawn from CICIoT2023, CICIDS2017 or IoT-23, while the studies in [45,46] emulated volumetric attacks at smart-city scale up to 100 , 000 synthetic IoT sources in the latter. A single study extends availability denial into the physical layer. Particularly, the authors in [66] performed selective Radio Frequency (RF) jamming on a live Zigbee testbed.

7.2. Integrity and Control Attacks

Thirteen studies target the integrity of data or the legitimacy of control, the family that most directly captures the cyber-physical aspects of these domains. Specifically, False Data Injection (FDI), the stealthy tampering of sensor readings or actuator setpoints, is demonstrated in five studies, four of them in building automation. Precisely, the work in [49] overwrote holding registers to force valve and damper positions, while in [50], the authors injected false temperature readings across a KNX/BACnet testbed. In the same vein, the study in [51] covertly recalibrated zone temperature sensors to trigger spurious overheating, and the work in [52] injected a malicious 95 F setpoint into an air-handling unit. In the smart-home domain, the work in [63] forged sensor POST events to overwrite device states on a Home Assistant testbed. Command injection and unauthorised control appear in three studies: the work in [54] abused an OpenHAB Representational State Transfer (REST) endpoint to inject false occupancy status; in [42] simulated command interception and injection across a nineteen-device testbed; and in [65] injected ghost and stealthy fake commands through a rogue smart application. A broader notion of packet and command injection, i.e., forged frames spoofed onto the wire, recurs in six studies [42,55,57,61,65,66], several of which pair it with the reconnaissance or availability behaviours. Finally, automation and rule interference, that is, subverting the logic of a smart environment rather than a single device, is exercised in [44], where the authors launched delay-based automation-interference attacks by selectively delaying events through Transmission Control Protocol (TCP) hijacking, and in [61], where the authors deleted log events and stretched inter-event intervals to defeat event-integrity assumptions.

7.3. Network and Communication Attacks

Eight studies mount attacks at the network and communication layer, with MitM being the most prevalent, present in seven of them. Specifically, the works in [50,63] used MitM to intercept and rewrite building- and home-automation traffic respectively; in [44,55,58] realized it through ARP-based interception on live testbeds; in [47] intercepted and fabricates Building Operating System (BOS) and Robot Control Subsystem (RPF) messages; and in [60] coupled MitM with a rogue access point. ARP spoofing and poisoning, the enabling technique for much of this interception, is called out explicitly in three studies [44,55,58]. Furthermore, replay attacks appear in two works. Namely, the authors in [50] crafted a KNX replay that overwrites a Desigo CC reading, and in [54] they replayed a HackRF captured radio frame to force a lighting switch off. Finally, the corpus contains a single Evil-Twin/rogue-access-point attack, executed in [60], deploying an Airgeddon-driven rogue Access Point (AP) with a deauthentication attack to capture a Wi-Fi Protected Access (WPA) handshake on a real camera-and-gateway testbed.

7.4. Malware and Botnet Attacks

Seven studies incorporate malware, almost entirely in the form of botnet or malware-infection behaviour (seven studies), with the Mirai and Bashlite families recurring as exemplars. Specifically, two studies stage genuine infections on hardware, with the work in [53] loading a malicious driver onto a Control4 hub to enlist it in a botnet and crypto-mining, and the study in [42] infecting live Raspberry Pi boards with Mirai and Bashlite to produce real DDoS, scanning and exfiltration traffic. The remaining malware studies observed infection as a labelled traffic class for detection. Namely, the works in [56,57,62] evaluated against botnet flows from public datasets, while in [45,46] the authors emulated botnet propagation at smart-city scale. Distinct from botnet infection, the work in [53] demonstrated the corpus’s only firmware/driver-level malware in the strict sense of a malicious device driver acting as a persistent backdoor.

7.5. Reconnaissance Attacks

Six studies include a reconnaissance phase, split between active scanning and fingerprinting (four studies) and protocol-specific device or register enumeration (two studies). Scanning and fingerprinting were observed largely within the detection-oriented pipelines. The studies in [42,56,57] treated port-scanning and device-enumeration traffic as a class to be recognised, while in [43] the authors performed active fingerprinting of a twenty-three-device home from sniffed Wi-Fi traffic. Enumeration in the building-automation sense, polling BACnet registers or objects to map a live installation, is demonstrated in [49,50], in both cases as the precursor to the injection attacks catalogued in Section 7.2.

7.6. Privacy and Eavesdropping Attacks

Three studies treated the residents’ privacy itself as the asset under attack, inferring sensitive information from ostensibly protected traffic. Specifically, the work in [43] sniffed encrypted Wi-Fi on a live twenty-three-device testbed to fingerprint devices and infer habitual behaviour; in [64] inferred device states and user activity with 94.8 % accuracy from ZigBee and Z-Wave traffic on a SmartThings testbed prior to any defence; and in [66] used passive traffic analysis to guide its subsequent jamming and injection. These attacks are notable for succeeding against encrypted channels, exploiting metadata and timing rather than payload contents.

7.7. Access and Authentication Attacks

Two studies compromise authentication directly. The study in [52] brute-forced the credentials of a building-automation platform as the entry point for its setpoint-injection attack, and in [60] the authors captured a WPA handshake via its rogue-AP deauthentication attack. The scarcity of this family in live form is itself a finding, as authentication weaknesses are frequently discussed across the corpus, but only rarely exercised end-to-end on a testbed.

7.8. Software Exploitation Attacks

A single study reaches the software-exploitation layer. That is, the study in [48] fuzzed the BACnet Secure Connect implementation on real Automated Logic hardware and discovered two array-index out-of-bounds defects that crash four device models, the only memory-safety exploitation demonstrated on physical equipment in the corpus. Despite the prevalence of resource-constrained embedded controllers across all three domains, systematic exploitation of their firmware remains almost entirely unexplored on the surveyed testbeds.

7.9. Synthesis

Read across the eight families, the corpus reveals a distinctly network-and-availability-centric attack surface. Availability (fourteen studies) and integrity/control (thirteen) dominate, followed by network-layer interception (eight) and malware (seven), while the objectives that presuppose deeper access, i.e., authentication compromise (two) and software exploitation (one), are strikingly rare. This distribution partly reflects threat priorities, but it is also an artifact of testbed realism as characterized in Section 5. Namely, volumetric, injection and interception attacks are readily staged on hybrid and emulated platforms and abundantly represented in the public datasets that the detection-oriented studies consume, whereas firmware exploitation and end-to-end authentication bypass demand physical devices and low-level access that few testbeds provide. The prominence of dataset-replayed and emulated attacks in the availability, malware and reconnaissance families, as opposed to the live, hardware-executed attacks concentrated in the building-automation integrity work, is therefore not incidental but a direct consequence of how these environments are built.

8. Controls Classification

This section classifies the security controls contributed across the corpus into the eight families depicted in Figure 6 and, as in the preceding Section 7, names the studies that populate each one, so that the reader can trace every category back to its concrete evidence rather than reading counts alone. Every one of the 28 included studies contributes something to the defensive picture, including a detector, a preventive mechanism, an enabling platform, or at minimum a set of recommendations. Therefore, each of them is classified in exactly the family that best captures its primary contribution.
A single observation frames the entire classification. Specifically, the corpus is overwhelmingly detection-oriented. Seventeen of the 28 studies contribute a detection mechanism of some kind, and among these, the learning-based detectors dominate. Genuinely preventive or responsive controls, i.e., mechanisms that block, isolate, encrypt, or otherwise stop an attack rather than merely observe it, are comparatively scarce, and a non-trivial tail of the corpus contributes no implemented mechanism at all. We make this skew explicit family by family below in Section 8.1 through Section 8.8.

8.1. Detection: Machine and Deep Learning

The largest single family, comprising eleven studies, applies Machine Learning (ML) or Deep Learning (DL) to intrusion and anomaly detection. Eight studies frame their contribution as an ML- or anomaly-based Intrusion Detection System (IDS). Specifically, the work in [40] paired an Isolation Forest network analyser with a command validator on a HitL building testbed; in [50] fed a Jensen–Shannon-divergence inter-arrival-time feature to Support Vector Machine (SVM) and decision-tree classifiers, reporting complete detection of KNX MitM and FDI attacks; in [56] and[57] built hierarchical and ensemble intrusion detectors, respectively; and in [58,61,62,63] contributed anomaly detectors tuned to temporal, behavioural-profile, contextual-correlation and embedded-classifier settings. Moreover, four studies contribute deep-learning detectors. Precisely, the work in [58] embedded a lightweight classifier on an ESP32 thermostat alongside a heavier cloud model; in [42] employed a feature-separation heterogeneous graph attention network over a device-interaction knowledge graph; in [45] combined stacked autoencoders with gated recurrent units for smart-city botnet detection; and in [59] used a convolutional network reporting 99.96 % accuracy as the detection stage of its prevention pipeline. Finally, two studies adopted a federated-learning architecture, retaining training data on-premises, namely, the authors in [56] distributed learning across home-router, edge and cloud tiers, and in [57] ran a federated ensemble at its second tier.

8.2. Detection: Statistical and Physics-Based

Five studies detect attacks through statistical or physics-grounded models rather than learned representations, an approach especially suited to the cyber-physical coupling of building and home environments. [51] applies a grey-box heat-balance model with a Cumulative Sum (CUSUM) change statistic to flag sensor false-data injection on a real HVAC system, reporting a complete detection rate; [54] fuses cyber and physical state in a deterministic finite-state-machine model to catch replay and Application Programming Interface (API)-tampering attacks; [55] binds ARP replies against independent Dynamic Host Configuration Protocol (DHCP) lease records to detect and mitigate ARP spoofing in real time; [40] complements its learning components with a physical fault detector; and [42] augments its graph model with statistical anomaly features. This family is notable for pairing detection with strong physical-consistency guarantees, at the cost of the domain-specific modelling each instance requires.

8.3. Detection: Semantic and Formal

Two studies detected attacks by reasoning over the semantics of a smart environment, i.e., the intended relationships among devices, events and automation rules, rather than over traffic statistics. Specifically, the study in [44] constructed a formal cross-rule-interference model and used observation-equivalence checking to expose delay-attack-induced rule conflicts. In the same spectrum, the work in [65] mined semantics-based device and event correlations with its HAWatcher system and used a shadow-execution engine to flag events that violate expected correlations. Notably, the latter two approaches are distinctive in offering interpretable, specification-grounded alarms, but they presuppose an accurate model of intended behaviour that is itself costly to obtain.

8.4. Detection: Signature-Based

A single study contributes a signature-based detector. Namely, the work in [46] inspected traffic redirected into an isolated network slice using a Snort-based signature engine, as one component of its broader honeynet defence. The rarity of signature-based detection in the corpus reflects the field’s decisive tilt toward anomaly- and learning-based methods, which do not depend on pre-enumerated attack signatures.

8.5. Prevention and Response

Four studies contributed mechanisms that act against an attack rather than merely detecting it, underscoring how thinly this space is populated. Three pursue moving-target, obfuscation or decoy strategies: i) the work in [64] injected “phantom user” traffic and morphs packet patterns with its SniffMislead system to defeat traffic-sniffing attackers without modifying devices; ii) in [46] redirected malicious traffic into an isolated honeypot slice; and iii) in [43] discussed, though did not fully implement, the injection of decoy and spoofed-Media Access Control (MAC) traffic to obscure real device patterns. Two studies applied SDN-based segmentation. Particularly, the study in [59] blocked malicious flows at a Ryu SDN controller once its detector fires, while the one in [46] used network slicing to segment traffic. A single study [59] contributed an active intrusion-prevention capability, whose SDN/Ryu-based system combined detection with in-line blocking; the corpus’s only end-to-end detect-and-block pipeline.

8.6. Preventive Hardening

Three studies harden the system before any attack occurs, chiefly through cryptography. Specifically, the work in [47] implemented and validated Message Queuing Telemetry Transport Secure (MQTTS) encryption to defeat a MitM case study on its smart-building emulation platform. Similarly, the study in [57] secured device links with AES-128 encryption, packet division, clock synchronisation and channel randomisation at the first tier of its architecture. Last, the work in [67] took an architectural approach, prototyping a consortium-blockchain smart-home design over ESP32 nodes to assure the confidentiality, integrity and availability of home transactions. Importantly, the fact that preventive hardening is confined to only three studies constitutes a salient finding given the prevalence of interception and injection attacks documented in Section 7.

8.7. Testbed and Tooling Contributions

Six studies make their primary defensive contribution not as a running detector but as an enabling artifact, namely, a testbed, emulation platform or discovery tool that the community can reuse to develop and evaluate defences. Five contribute an emulation or digital-twin platform. Specifically, the work in [49] built the BUILD-SOS testbed, in [41,52] developed HitL rigs that generate attack and fault datasets with detection left as future work, in [40] contributed its HitL testbed alongside its detectors, and in [47] provided its Smart Building Control System Emulator (SBCSE) emulation platform. One study contributed vulnerability-discovery tooling. That said, the work in [48] applied the boofuzz, AFL++ and AFLNet fuzzers to BACnet and KNX devices as a defensive means of surfacing software defects before adversaries do. These contributions are foundational rather than protective; that is, they lower the cost of future defensive research without themselves stopping an attack, which is precisely why we opt to account for them separately from the detection and prevention families.

8.8. Guidelines Only

Four studies stop short of an implemented mechanism, contributing security recommendations alone. The study in [53] discussed driver-verification and certification countermeasures; in [60] tabulated generic mitigations, including strong and multi-factor authentication, firewalls, IDS/IPS and frequency-hopping, for Evil-Twin attacks on assistive devices; in [66] offered hardening recommendations for the Zigbee 3.0 commissioning process alongside its offensive tooling; and in [44] supplemented its formal detector with recommended mitigations such as network isolation and faster Transport Layer Security (TLS) heartbeats. These studies defined the residual boundary of the field’s defensive maturity, with the attack being demonstrated or modelled, but the countermeasure remaining prescriptive rather than operational.

8.9. Synthesis

Aggregated across the eight families, the defensive landscape is emphatically detection-heavy and prevention-light. Namely, detection in its four forms accounted for seventeen distinct studies, of which the learning-based family alone contributed eleven; prevention, response and preventive hardening together accounted for seven; six studies contributed enabling platforms or tooling; and four contributed recommendations only. Read against the attack surface of Section 7, this distribution exposes a structural asymmetry. Put simply, the corpus is far better at recognising that an availability, integrity or network attack is underway than at stopping it, and the preventive controls that do exist, spanning encryption, segmentation, and in-line blocking, cluster in a handful of studies rather than spanning the demonstrated threat space. Moreover, the prominence of learning-based detection is entangled with the testbed-realism observations of Section 5. Specifically, many of these detectors are trained and validated on replayed or emulated traffic rather than against live, hardware-executed attacks, which shapes both their reported accuracy and the confidence with which their protection can be claimed.

9. Cross-Analysis: Attack Surface Versus Defensive Coverage

Section 7 and Section 8 independently classify the attacks exercised and the defensive contributions reported by the surveyed testbeds. Considered together, they reveal a structural asymmetry. Specifically, the corpus demonstrates a comparatively broad attack surface. In contrast, defensive coverage is concentrated principally on identifying attacks rather than preventing them, containing their effects, or restoring trustworthy operation after compromise. This distinction is important in smart environments, where a successful attack may have both cyber and physical consequences [4,5,6]. For example, a detector may correctly identify malicious traffic, yet not assure that an unauthorised command was rejected before reaching an actuator, that a DoS condition was contained before disrupting a controller, or that a manipulated physical process was returned to a safe state. Therefore, we assess coverage against four operational functions derived from the NIST Cybersecurity Framework core functions, as a collapsed testbed-oriented view of the framework [8]: detection (Detect), preventive hardening (Protect), response and containment (Respond), and recovery or safe-state restoration (Recover). Note that NIST’s Identify and Govern functions are deliberately not included because this study focuses on implemented and evaluated defensive capability within testbeds rather than organizational governance or asset-management maturity.
Interestingly, while the first three functions appear, albeit unevenly, in the corpus, recovery is not established as a distinct and systematically evaluated control family. Table 6 and Figure 7 summarize the resulting attack-control mapping. Specifically, it reports qualitative coverage rather than one-to-one numerical counts, since individual studies may demonstrate several attacks, while each study in Section 8 is classified according to its primary defensive contribution. Moreover, a single defensive mechanism, such as encryption or network segmentation, may address more than one attack family. In contrast, a detector evaluated only against replayed traffic does not establish resilience against a live attacker. In simple terms, coverage reflects implemented and evaluated mechanisms rather than recommendations alone. Both in Table 6 and Figure 7, strong denotes several implemented contributions; partial denotes limited, scenario-specific, or detection-dominant evidence; minimal denotes nominal contributions; and none identified denotes the absence of an operationally evaluated security control.
With reference to Section 7 and Figure 5, availability and integrity/control are the dominant attack objectives in the corpus. The first class of attacks are primarily represented by DoS and DDoS scenarios, including live attacks against Message Queuing Telemetry Transport (MQTT) brokers, embedded thermostats, building controllers, and Zigbee networks, as well as emulated large-scale botnets and replayed malicious traffic [40,41,42,45,46,47,49,52,53,56,57,58,62,66]. Nevertheless, as highligted in Section 8 and illustrated in Figure 6, the defensive evidence is predominantly detection-oriented: ML- and DL-based IDSs classify availability-related traffic, but few studies demonstrate that service remains available during an attack. Indicatively, the works in [46,59] combined detection with Ryu-controller blocking and used network slicing and honeypot redirection for volumetric attacks, respectively. We argue that these are important exceptions, but they do not yet constitute broad evidence for resilient availability across the protocols and architectures represented in the reviewed corpus.
In the same vein, integrity and control attacks have stronger cyber-physical grounding, particularly in smart-building testbeds. False sensor readings, malicious setpoints, register manipulation, forged device events, and automation-rule interference are exercised against real or HitL systems [44,50,51,52,54,63,65]. Accordingly, several contributions detect such attacks through physical-process models, statistical features, cyber-physical state consistency, semantic relationships, or contextual correlations [40,44,50,51,54,63,65]. For example, the work in [51] used a grey-box heat-balance model and a CUSUM statistic to identify stealthy sensor manipulation in a real HVAC environment, while in [44] the authors reasoned over automation rules to detect delay-induced interference. However, most controls raise an alarm only after anomalous behaviour is observed, without evaluating whether malicious commands can be rejected before execution, constrained within safe operating bounds, or followed by automatic restoration of the affected process.
As detailed in Section 7 and Figure 5, network and communication attacks receive partial but fragmented defensive coverage. Namely, MitM, ARP spoofing, replay, and Evil Twin scenarios are demonstrated across building and smart-home testbeds [44,47,50,54,55,58,60,63]. Interestingly, several studies provide concrete mechanisms beyond detection. For example, the work in [47] evaluated MQTTS encryption against a MitM case study; in [55] bound ARP replies to DHCP lease information to detect and mitigate ARP spoofing; in [57] protected wireless communications using AES-128, packet division, clock synchronisation, and channel randomisation; and in [59] used an SDN controller to block flows classified as malicious. Nevertheless, these mechanisms are evaluated in isolated protocol settings, and the corpus provides no common defensive architecture that spans BACnet, KNX, Zigbee, MQTT, Wi-Fi, and cloud-mediated smart-home platforms.
In the same spectrum, malware and botnet coverage is likewise detection-heavy. Several studies train or evaluate classifiers on botnet, DDoS, scanning, and malware-related traffic [42,45,46,56,57,62]. In contrast, genuine infection is demonstrated in only a small number of testbeds. Specifically, in [53] the authors showed that a malicious driver can compromise a Control4 hub and support botnet and resource-abuse activity, while in [42] they infected Raspberry Pi nodes with Mirai and Bashlite to generate live malicious traffic. The resulting evidence supports traffic-level recognition but provides little basis for claims concerning prevention of compromise, firmware and driver integrity, trusted update mechanisms, remediation, or secure re-enrolment of compromised devices.
Privacy and eavesdropping attacks expose a further limitation of network-level protection. For instance, the work in [43] inferred device types, locations, and habitual user behaviours from encrypted Wi-Fi traffic, while in [64] inferred device states and user activity from Zigbee and Z-Wave traffic before applying their defence. Similarly, the study in [66] used passive traffic analysis to guide subsequent Zigbee jamming and injection activities. These studies show that payload encryption does not, by itself, prevent inference from metadata, timing, and traffic patterns. SniffMislead provides a notable preventive direction by injecting phantom traffic to reduce an adversary’s behavioural-inference accuracy [64]; however, the corpus does not yet establish its scalability, resource cost, or robustness against adaptive attackers.
Recall from Section 7 and Section 8, and Figure 5 and Figure 6 that reconnaissance, authentication compromise, and software exploitation exhibit the largest gaps between the attack surface and evaluated defences. Namely, reconnaissance is included through scanning, fingerprinting, or protocol-specific enumeration [42,43,49,50,56,57], but these activities are typically used either as a precursor to injection attacks or as a traffic class for an IDS. In other words, the corpus offers little evidence for mechanisms that reduce the discoverability of devices and services, such as authenticated enumeration, service minimisation, rate limiting, deception, or protocol-aware access restrictions.
Similarly, authentication compromise is directly exercised in only two studies. Namely, the work in [52] used credential brute forcing as the entry point to a building-automation setpoint-injection scenario, while the study in [60] captured a WPA handshake through an Evil Twin and deauthentication attack. Both studies discuss relevant countermeasures, but the reviewed corpus does not provide corresponding end-to-end evaluations of authentication and access-control enforcement under attack [60]. This is a significant limitation because authentication and authorisation mechanisms define the trust boundary (Section 5.3) through which many later control-plane and command-injection attacks become possible.
As discussed in Section 7, software exploitation is represented by a single study. Particularly, the authors in [48] applied fuzzing tools to BACnet Secure Connect implementations on real Automated Logic hardware and identify array-index-out-of-bounds defects that crash several device models. This contribution is important because vulnerability discovery can identify defects before their exploitation. However, it does not demonstrate the effectiveness of mitigations once a flaw exists. Secure boot, signed firmware, memory-safety protection, runtime integrity monitoring, coordinated patch deployment, and compensating controls for unpatchable legacy controllers remain absent from the evaluated defensive landscape, as evidently showcased in Section 8.
The observed coverage asymmetry is partly shaped by testbed architecture. Hybrid and HitL building testbeds connect real controllers to simulated physical processes; therefore, particularly suited to studying false-data injection, malicious setpoints, and their cyber-physical effects [40,41,52]. These architectures also enable physics-informed detection, such as the HVAC-oriented approach [51]. Fully emulated testbeds, by contrast, can scale to large populations of synthetic nodes and are consequently well suited to volumetric DoS, botnet, and network-traffic classification experiments [45,46,49,59]. Physical smart-home deployments provide realistic radio, firmware, and vendor-platform behaviour, but device cost, statefulness, and reproducibility constraints limit their scale, and the range of attacks repeatedly exercised [42,43,44,60,64,66].
These architectural choices also condition the strength of defensive claims. For example, a physics-aware detector requires an explicit cyber-physical process model; therefore, appears principally in building-oriented testbeds [40,51,54]. Conversely, learning-based network IDSs can be trained and evaluated using replayed datasets or emulated traffic, as in the federated, embedded, and smart-city detection studies [45,56,57,58,62]. This facilitates controlled benchmarking, but weakens the connection between reported classification accuracy and end-to-end operational resilience. In addition, proprietary clouds, vendor hubs, and opaque device firmware often lie outside the directly instrumented system boundary, limiting the assessment of cloud-mediated attack paths, identity compromise, update security, and control-plane abuse [44,61,64,65].
To wrap up, Table 7 provides the evidence ladder used to interpret defensive claims, while Figure 8 shows how testbed realism and evaluation mode shape that evidence. The two together make the key point of this section explicit. That is, live and HitL experiments support the strongest claims about NIST cybersecurity framework Protect, Respond, and Recover outcomes, whereas emulated and dataset-driven setups are more often sufficient for Detect-oriented evidence. This distinction is important because a mitigation that is only recommended in prose does not carry the same evidentiary weight as one that is implemented and exercised in a testbed. For that reason, the cross-analysis conducted in the present section counts only operationally evaluated controls, building on the attack and control taxonomies of Figure 5 and Figure 6, respectively.

10. Open Challenges and Future Research Directions

The surveyed corpus reveals a field that has matured in breadth, but not yet in balance. Particularly, smart-environment cybersecurity testbeds now cover a broad range of domains, protocols, and attack families, yet the evidence they generate remains concentrated in a narrow subset of experimental modes and defensive objectives. Across the included studies, a dominant pattern is discerned: the literature is rich in detection-oriented experiments, especially those driven by network traces and labelled datasets, but comparatively thin in preventive hardening, active response, and recovery-oriented validation. At the same time, the architectural choices of existing testbeds strongly condition which threats can be staged, which defences can be validated, and how far any resulting claim can be generalized beyond the laboratory, as discussed in Section 5 and Section 7. These observations point to a research agenda that is less about building more testbeds in the abstract than about building more representative, interoperable, and operationally meaningful ones.
  • From Detection Accuracy to Operational Resilience: The most immediate challenge is the field’s emphasis on detection as the primary measure of defensive progress. A large fraction of the corpus evaluates IDS-like mechanisms under emulated traffic or replayed datasets, often reporting high classification performance, but without showing whether malicious actions are actually blocked, contained, or safely reversed once detected, as summarized in Section 7 and Section 8. This imbalance is not merely methodological; it affects the kinds of security claims that can be credibly supported. In other words, a detector trained on public or replayed traffic can establish discriminative performance, but it does not by itself demonstrate resilience under an adaptive adversary, nor does it show that a smart environment can continue operating safely under attack.
    Future Direction: Move from open-loop recognition to closed-loop defensive evaluation. A mature testbed experiment should not end when an alarm is raised. Rather, it should trace the full chain from attack launch to detection and intervention to post-intervention outcome. For example, a DoS scenario should report not only detection accuracy but also service continuity, degradation, containment delay, and recovery time. Likewise, an FDI experiment should measure whether a malicious command is rejected before execution, whether the physical process remains within safe bounds, and whether the environment returns to a trustworthy state afterwards. In other words, the next generation of smart-environment testbeds should be explicitly designed to validate resilience, not only recognition.
  • Strengthening Prevention, Response, and Recovery: Closely related is the scarcity of implemented controls beyond observation. Only a small subset of the corpus contributes mechanisms that actively stop, isolate, or harden against attacks, and even fewer evaluate what happens after compromise. This is striking because the attack surface documented in Section 7 is broader than the defensive surface summarized in Section 8. In particular, the prevalence of MitM, injection, DoS, and privacy attacks is not matched by an equally mature body of preventive architectures, containment strategies, or recovery workflows.
    Future Direction: First, protocol-aware prevention deserves much stronger attention. Second, active containment mechanisms, for example, network segmentation, SDN-based quarantine, traffic redirection, and safe degradation policies, should be validated under realistic conditions rather than discussed only at the design level. Third, recovery should be elevated to a first-class research objective. At present, the literature rarely asks how a compromised controller, gateway, or automation rule set is restored to a trusted state, even though recovery is essential in real buildings, homes, and cities where service continuity matters as much as – if not more – detection. For example, testbeds that incorporate rollback, fail-safe operation, rule sanitisation, rekeying, secure re-enrolment, or post-incident reconfiguration would substantially broaden the field’s evidentiary base.
  • Reducing Dependence on Replayed and Public Datasets: A second major challenge concerns the widespread use of replayed or public datasets in detection-oriented studies. Notably, besides the studies included in this survey that used public datasets as part of a broader experimental setup, almost all the 251 studies classified as full-text exclusions in the PRISMA process of Section 3 relied exclusively on public datasets and did not include a demonstration testbed; they were excluded from this survey, since the review focuses on experimentally grounded smart-environment security studies rather than dataset-only classification work. This practice might have clear advantages of improving label quality, facilitating benchmarking, and making experiments more reproducible; however, it also weakens the connection between the detector and the environment it is supposed to protect. Recent real-world IoT dataset work shows that it is possible to collect attack traffic directly from a live testbed, while many existing datasets remain built from simulated or pre-generated traffic, underscoring the gap between benchmark convenience and operational realism.
    Future Direction: Future research should not abandon datasets; rather, it should integrate them more carefully with realistic experimentation. One direction is the creation of hybrid evaluation pipelines in which public datasets are used for pre-training or stress testing, while final validation occurs on live or HitL deployments. Another is the development of benchmark suites derived from shared, open smart-environment testbeds rather than generic IoT corpora, i.e., utilized datasets including CICIoT-2023 [70], CIC-IDS-2018 [71], CIC-DDoS-2019 [72], BoT-IoT [73], TON_IoT [74], IoT-23 [75], IoTID20 [76], N-BaIoT [77], and others. Once more, it is important to note that the latter datasets, and more, were observed as the sole evaluation benchmark of almost all the 251 studies classified as full-text exclusions in Section 3. In this context, it would be particularly valuable for datasets to preserve cyber-physical context. Specifically, network traces are synchronized with sensor states, control actions, process variables, and intervention logs. Such resources would allow the community to compare detectors without severing the link between traffic anomalies and physical consequences. In this sense, the challenge is not simply to collect more data, but to collect data that remain faithful to the semantics of the environment under attack.
  • Improving Realism Without Sacrificing Reproducibility: This review repeatedly exposes a structural separation between realism and reproducibility. Typically, physical smart-home deployments capture genuine timing, radio effects, and firmware behaviour, however, they are expensive, difficult to scale, and arduous to reproduce exactly. Conversely, fully emulated environments scale well and support rapid experimentation, omitting, however, the very artifacts, such as radio interference, protocol corner cases, timing jitter, or device-specific implementation flaws, that an adversary can exploit. Last, hybrid and HitL designs partially bridge this gap, especially in smart-building research, remaining, nonetheless, domain-specific and often rely on bespoke infrastructure, as discussed in Section 5 and Section 6.
    Future Direction: The creation of modular, layered testbed architectures that expose realism as a configurable dimension rather than a fixed design choice. Instead of treating physical, hybrid, and emulated setups as separate families, future platforms should make it possible to substitute layers systematically. Specifically, real devices for selected endpoints, simulated processes for costly physical plants, emulated traffic generators for scale, and optional cloud replicas for service dependencies. Such modularity would allow the same experimental design to be replayed across multiple realism levels, thereby clarifying which security results are robust and which depend on a narrow implementation context. To support this, future testbeds should report realism assumptions explicitly, including what is real, what is modelled, what is omitted, and which attack paths remain outside the trusted boundary.
  • Broadening Under-Studied Attack Families: The distribution of attack families in the corpus is itself a research signal. Availability, integrity/control, and network interception dominate, whereas authentication compromise, firmware exploitation, and deep software exploitation are rare, as detailed in Section 7. This imbalance partly reflects the practical difficulty of such experiments; for example, brute-forcing credentials, reversing embedded firmware, or staging memory-safety exploits on operational controllers requires more physical access, specialised expertise, and ethical caution than generating DoS traffic or replaying labelled flows. Yet these omissions leave important portions of the attack surface underexplored.
    Future Direction: Expand into precisely those areas that current testbeds represent weakly. In smart buildings, this means firmware analysis, patchability studies, secure boot validation, and exploitation-aware testing for legacy BAS devices. In smart homes, it means end-to-end evaluation of onboarding, identity management, cloud account compromise, and multi-user access control. In smart cities, it means moving beyond volumetric DDoS and synthetic botnets toward multi-layer attacks that affect orchestration, service composition, and edge-cloud coordination. The challenge is not simply to add more attack categories to a taxonomy, but to ensure that testbeds can stage them safely, measure them meaningfully, and connect them to realistic defensive outcomes.
  • Supporting Semantic and Physics-Aware Security: One of the most promising findings is that the most interpretable and operationally meaningful detectors tend to exploit semantic or physical consistency rather than traffic statistics alone. Physics-based models in building environments, semantic rule reasoning in smart homes, and cyber-physical state consistency checks provide forms of evidence that are closer to how these environments actually function. Yet such approaches remain comparatively rare, as they require explicit models of expected behaviour and a testbed architecture that exposes process state rather than only packets, as reflected in Section 5 and Section 8.
    Future Direction: Security testbeds should increasingly expose richer internal semantics. For buildings, this means synchronized process variables, actuator states, occupancy models, and energy-control logic. For homes, it means automation rules, user routines, device correlations, and contextual state. For cities, it may include service workflows, mobility assumptions, and infrastructure interdependencies. Overall, the goal is to make security evaluation less dependent on pattern recognition over network traces and more grounded in whether the observed system behaviour remains physically and semantically plausible. Such a shift would also improve interpretability, a major concern as learning-based IDSs become more complex.
  • AI- and LLM-Assisted Testbed Engineering: The review explicitly motivates future directions involving AI- and LLM-assisted testing; a particularly timely area for expansion. Testbed construction today remains labour-intensive: scenarios are hand-crafted, attack scripts are manually configured, datasets require extensive curation, and experimental variation is often limited by human effort. Recent benchmarking work on LLM-driven offensive security [78,79] highlights that testbeds, metrics, and experimental design remain methodological bottlenecks in this area. That operational realism remains difficult to achieve in simplified or Capture the Flag (CtF)-like settings. However, AI assistance should be treated as an enabling layer, not as a substitute for realism.
    Future Direction: A research direction is controlled AI augmentation. This includes LLMs generating candidate attacks, test cases, misconfiguration scenarios, or protocol dialogues, with these being validated inside instrumented environments whose cyber-physical consequences remain measurable. Likewise, AI could support automated mutation of attack campaigns, adversarial generation of edge-case traffic, or explanation-driven exploration of policy and rule conflicts. Particularly promising is the use of LLMs as experiment orchestration agents that can help instantiate variants of a scenario across different realism levels, thereby supporting reproducibility and systematic sensitivity analysis. However, one important challenge here is trustworthiness, namely, AI-generated scenarios must be constrained, auditable, and benchmarked against expert-designed baselines before they can be accepted as valid experimental evidence.
  • Towards Shared Benchmarks and Reusable Reference Testbeds: Finally, the field still lacks widely adopted reference platforms that enable cumulative comparison [7]. Although there is strong convergence in individual tools, for example, Wireshark, Mininet, Modelica-based co-simulation, and scikit-learn/XGBoost pipelines, the testbeds themselves remain fragmented, frequently bespoke, and only partially reusable, as shown in Section 6. This fragmentation slows progress because each study must rebuild enough infrastructure to answer its local question, while cross-paper comparison remains difficult.
    Future Direction: The development of open reference testbeds and benchmark scenarios for each smart-environment domain. These should include not only code and topology, but also scenario definitions, realism assumptions, protocol configurations, attack scripts, expected process responses, and evaluation metrics. Ideally, the community would converge on a small number of openly maintained baseline environments, for example, a reference smart-home deployment with common consumer ecosystems, a reference BAS HitL building platform, and a reference smart-city network-emulation stack. Such resources would not eliminate the need for bespoke experimentation, but they would provide a common starting point, improve reproducibility, and make it easier to tell whether a new detector, prevention mechanism, or response strategy represents genuine progress.
Taken together, the above-discussed challenges suggest that the next phase of smart-environment cybersecurity testbed research should be defined less by the number of new platforms and more by the quality of the evidence they produce. In other words, more realistic trust boundaries, stronger closed-loop evaluation, richer cyber-physical semantics, broader attack coverage, AI-assisted but auditable experimentation, and openly reusable reference environments would move the field from fragmented proof-of-concept studies toward a more cumulative and operationally meaningful science of smart-environment security.

11. Conclusion

This systematic review examined what smart-environment cybersecurity testbeds actually validate across 28 experimentally testbed-grounded studies. The main finding is a strong asymmetry, namely, the corpus is rich in detection-oriented work, especially using emulated or dataset-driven evaluation, but much weaker in live evidence for prevention, containment, and recovery. As a result, many studies show that an attack is recognizable, yet far fewer can show that a smart environment remains safe and operational under attack. Furthermore, we showcase that testbed realism shapes evidentiary strength. Specifically, physical and HitL platforms support the most credible cyber-physical claims, while emulated and replayed setups are better suited to scale and reproducibility than to operational validation. This is why we opt to distinguish not only attacks and controls, but also evaluation mode and evidence maturity. Overall, availability and integrity/control attacks are the most frequent, whereas attacks on authentication and software vulnerabilities are comparatively uncommon. On the defensive side, detection overwhelmingly outweighs prevention, response, and recovery, leaving a persistent gap between recognizing attacks and demonstrating resilience. Moreover, we recognize that public datasets are valuable for testing purposes, yet they cannot serve as proof of live defence capabilities on their own. In other words, the practical implication is clear: future testbeds should move toward closed-loop resilience evaluation, richer live or HitL data, explicit cloud and edge trust boundaries, reusable reference platforms, and recovery-oriented experimentation. In this regard, AI- and LLM-assisted testing are promising future integrations, but only if they remain constrained, auditable, and grounded in real testbed behaviour.

Author Contributions

Conceptualization, V.K.; methodology, V.K.; validation, V.K.; writing—original draft preparation, V.K.; writing—review and editing, V.K, K.K, M.T, V.G.; supervision, V.G. All authors have read and agreed to the published version of the manuscript.

Funding

This work is supported by the Research Council of Norway through the SFI Norwegian Centre for Cybersecurity in Critical Sectors (NORCICS) project no. 310105.

Acknowledgments

The authors would like to acknowledge the use of Grammarly for language refinement and stylistic improvement of the manuscript. The tool was used exclusively for grammar and readability enhancement; all technical content, conceptual development, modelling decisions, and scientific conclusions remain solely the responsibility of the authors.

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

   The following abbreviations are used in this manuscript:
AI Artificial Intelligence
ALC Automated Logic Corporation
AP Access Point
API Application Programming Interface
ARP Address Resolution Protocol
BAN Burrows-Abadi-Needham
BAS Building Automation System
BACS Building Automation and Control System
BOS Building Operating System
COTS Commercial Off-The-Shelf
CUSUM Cumulative Sum
DHCP Dynamic Host Configuration Protocol
DL Deep Learning
DoS Denial-of-Service
DDoS Distributed Denial-of-Service
DUT Devices Under Test
FDI False Data Injection
HitL Hardware-in-the-Loop
HVAC Heating, Ventilation, and Air Conditioning
IaaS Infrastructure-as-a-Service
IDS Intrusion Detection System
IoT Internet of Things
ITS Intelligent Transport Systems
LLM Large Language Models
MAC Media Access Control
MitM Man-in-the-Middle
ML Machine Learning
MQTT Message Queuing Telemetry Transport
MQTTS Message Queuing Telemetry Transport Secure
NIST National Institute of Standards and Technology
PRISMA Preferred Reporting Items for Systematic reviews and Meta-Analyses
RPF Robot Control Subsystem
REST Representational State Transfer
RF Radio Frequency
SBCSE Smart Building Control System Emulator
SDN Software-Defined Networking
SDR Software-Defined Radio
SUT System Under Test
SVM Support Vector Machine
TCP Transmission Control Protocol
TLS Transport Layer Security
WPA Wi-Fi Protected Access

Appendix A. Database-Specific Search Strings

The search strings below were derived using a quasi-gold-standard strategy calibrated against the relevant primary studies, refined until the query retrieved the known-relevant set with high precision. Three conceptual groups (domain, testbed/resource, security) are combined as (G1 AND G2 AND G3), with synonyms OR-combined within each group. Queries are restricted to title and abstract fields only, excluding author keywords and full indexing metadata, which is the principal precision lever.
Preprints 225564 i001

References

  1. Ahmed, E.; Yaqoob, I.; Gani, A.; Imran, M.; Guizani, M. Internet-of-things-based smart environments: state of the art, taxonomy, and open research challenges. IEEE Wirel. Commun. 2016, 23, 10–16. [Google Scholar] [CrossRef]
  2. Alberti, A.M.; Santos, M.A.S.; Souza, R.; Da Silva, H.D.L.; Carneiro, J.R.; Figueiredo, V.A.C.; Rodrigues, J.J.P.C. Platforms for Smart Environments and Future Internet Design: A Survey. IEEE Access 2019, 7, 165748–165778. [Google Scholar] [CrossRef]
  3. Klaoudatou, E.; Konstantinou, E.; Kambourakis, G.; Gritzalis, S. A Survey on Cluster-Based Group Key Agreement Protocols for WSNs. IEEE Commun. Surv. Tutor. 2011, 13, 429–442. [Google Scholar] [CrossRef]
  4. Demertzi, V.; Demertzis, S.; Demertzis, K. An Overview of Cyber Threats, Attacks and Countermeasures on the Primary Domains of Smart Cities. Appl. Sci. 2023, 13. [Google Scholar] [CrossRef]
  5. Abdi, F.; Chen, C.Y.; Hasan, M.; Liu, S.; Mohan, S.; Caccamo, M. Preserving Physical Safety Under Cyber Attacks. IEEE Internet Things J. 2019, 6, 6285–6300. [Google Scholar] [CrossRef]
  6. Kampourakis, V.; Gkioulos, V.; Katsikas, S. A systematic literature review on wireless security testbeds in the cyber-physical realm. Comput. Secur. 2023, 133, 103383. [Google Scholar] [CrossRef]
  7. Kampourakis, V.; Gkioulos, V.; Katsikas, S. A step-by-step definition of a reference architecture for cyber ranges. J. Inf. Secur. Appl. 2025, 88, 103917. [Google Scholar] [CrossRef]
  8. Pascoe, C.E. Public draft: The NIST cybersecurity framework 2.0. In National Institute of Standards and Technology; 2023. [Google Scholar]
  9. de Santana, K.G.Q.; Schwarz, M.; Wangham, M.S. Cybersecurity testbeds for IoT: A systematic literature review and taxonomy. J. Internet Serv. Appl. 2024, 15, 450–473. [Google Scholar] [CrossRef]
  10. Cintuglu, M.H.; Mohammed, O.A.; Akkaya, K.; Uluagac, A.S. A Survey on Smart Grid Cyber-Physical System Testbeds. IEEE Commun. Surv. Tutor. 2017, 19, 446–464. [Google Scholar] [CrossRef]
  11. Minani, J.B.; Sabir, F.; Moha, N.; Guéhéneuc, Y.G. A Systematic Review of IoT Systems Testing: Objectives, Approaches, Tools, and Challenges. IEEE Trans. Softw. Eng. 2024, 50, 785–815. [Google Scholar] [CrossRef]
  12. Lonetti, F.; Bertolino, A.; Di Giandomenico, F. Model-based security testing in IoT systems: A Rapid Review. Inf. Softw. Technol. 2023, 164, 107326. [Google Scholar] [CrossRef]
  13. Graveto, V.; Cruz, T.; Simöes, P. Security of Building Automation and Control Systems: Survey and future research directions. Comput. Secur. 2022, 112, 102527. [Google Scholar] [CrossRef]
  14. Li, G.; Ren, L.; Fu, Y.; Yang, Z.; Adetola, V.; Wen, J.; Zhu, Q.; Wu, T.; Candan, K.; O’Neill, Z. A critical review of cyber-physical security for building automation systems. Annu. Rev. Control 2023, 55, 237–254. [Google Scholar] [CrossRef]
  15. Panahi Rizi, M.H.; Hosseini Seno, S.A. A systematic review of technologies and solutions to improve security and privacy protection of citizens in the smart city. Internet Things 2022, 20, 100584. [Google Scholar] [CrossRef]
  16. Houichi, M.; Jaidi, F.; Bouhoula, A. Cyber Security within Smart Cities: A Comprehensive Study and a Novel Intrusion Detection-Based Approach. Comput. Mater. Contin. 2024, 81. [Google Scholar] [CrossRef]
  17. Joshi, S.; Baviskar, A.; Rajmane, S. A review of cybersecurity in smart cities and intelligent transport systems. Discov. Internet Things 2026, 6, 41. [Google Scholar] [CrossRef]
  18. Fink, A. Conducting research literature reviews: From the internet to paper; Sage publications, 2019. [Google Scholar]
  19. Okoli, C.; Schabram, K. A guide to conducting a systematic literature review of information systems research. SSRN Electron. J. 2010, 10. [Google Scholar] [CrossRef]
  20. Page, M.J.; McKenzie, J.E.; Bossuyt, P.M.; Boutron, I.; Hoffmann, T.C.; Mulrow, C.D.; Shamseer, L.; Tetzlaff, J.M.; Akl, E.A.; Brennan, S.E.; et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ 2021, 372. [Google Scholar] [CrossRef] [PubMed]
  21. Kannelønning, K.; Katsikas, S.K. A systematic literature review of how cybersecurity-related behavior has been assessed. Inf. Comput. Secur. 2023, 31, 463–477. [Google Scholar] [CrossRef]
  22. Silva, R.; Neiva, F. Systematic literature review in computer science—a practical guide; Federal University of Juiz de Fora Technical Report of Computer Science Department, 2016. [Google Scholar] [CrossRef]
  23. Tong, J.; Sun, W.; Wang, L. A Smart Home Network Simulation Testbed for Cybersecurity Experimentation. In Proceedings of the Testbeds and Research Infrastructure: Development of Networks and Communities, Cham, 2014; pp. 136–145. [Google Scholar]
  24. Eini, R.; Linkous, L.; Zohrabi, N.; Abdelwahed, S. A testbed for a smart building: design and implementation. In Proceedings of the SCOPE ’19, New York, NY, USA, 2019; pp. 1–6. [Google Scholar] [CrossRef]
  25. Hyman, B.T.; Alisha, Z.; Gordon, S. Secure Controls for Smart Cities; Applications in Intelligent Transportation Systems and Smart Buildings. Int. J. Sci. Eng. Appl. 2019, 8, 167–171. [Google Scholar] [CrossRef]
  26. Sivanathan, A.; Loi, F.; Gharakheili, H.H.; Sivaraman, V. Experimental evaluation of cybersecurity threats to the smart-home. In Proceedings of the 2017 IEEE International Conference on Advanced Networks and Telecommunications Systems (ANTS), 2017; pp. 1–6. [Google Scholar] [CrossRef]
  27. Furfaro, A.; Argento, L.; Parise, A.; Piccolo, A. Using virtual environments for the assessment of cybersecurity issues in IoT scenarios. Simul. Model. Pract. Theory;Smart Cities Internet Things 2017, 73, 43–54. [Google Scholar] [CrossRef]
  28. Abu Waraga, O.; Bettayeb, M.; Nasir, Q.; Abu Talib, M. Design and implementation of automated IoT security testbed. Comput. Secur. 2020, 88, 101648. [Google Scholar] [CrossRef]
  29. Tran, H.K.V.; Börstler, J.; Ali, N.B.; Unterkalmsteiner, M. How good are my search strings? Reflections on using an existing review as a quasi-gold standard. arXiv 2024, arXiv:2402.11041. [Google Scholar]
  30. Hasanin, T. A pseudorandom-based blockchain authentication protocol for resource-constrained UAVs in smart cities. Ain Shams Eng. J. 2026, 17, 104296. [Google Scholar] [CrossRef]
  31. Almagrabi, A.O. A provably secure and lightweight authentication protocol for smart home environment. J. Eng. Res. 2026. [Google Scholar] [CrossRef]
  32. Ullah, I.; Saeed, K.; Jan, S.U.; Yahya, K.; Rai, H.M.; Ghani, A. Secure and Lightweight Authentication for IoT-Based Smart-Home Surveillance. IEEE Internet Things J. 2026, 13, 30777–30794. [Google Scholar] [CrossRef]
  33. Waheed, T.; Marchetti, E.; Calabrò, A. Improving Cybersecurity for Smart Home Systems. In Proceedings of the Proceedings of the 21st International Conference on Web Information Systems and Technologies - WEBIST. INSTICC, SciTePress, 2025; pp. 220–227. [Google Scholar] [CrossRef]
  34. Kölsch, J.; Post, S.; Zivkovic, C.; Ratzke, A.; Grimm, C. Model-based development of smart home scenarios for IoT simulation. In Proceedings of the 2020 8th Workshop on Modeling and Simulation of Cyber-Physical Energy Systems, 2020; pp. 1–6. [Google Scholar] [CrossRef]
  35. Agarwal, R.; Meng, N.; Gao, X.; Liu, Y. Graph-Based Simulation for Cyber-Physical Attacks on Smart Buildings. Proc. Constr. Res. Congr. 2022, 2021, 28–37. [Google Scholar] [CrossRef]
  36. Al-Balasmeh, H. Blockchain-Enabled Cybersecurity and Data Privacy Solutions for Smart Cities. In Proceedings of the 2024 IEEE 9th International Conference on Engineering Technologies and Applied Sciences (ICETAS), 2024; pp. 1–9. [Google Scholar] [CrossRef]
  37. Srinivasan, A.; Parmar, V.; Oh, T.; Ryoo, J.; Viglione, M. Anomaly Detection System for Smart Home using Machine Learning. In Proceedings of the 2021 International Conference on Software Security and Assurance (ICSSA), 2021; pp. 52–55. [Google Scholar] [CrossRef]
  38. Natrayan, L.; Kaliappan, S.; Arputharaj, B.S.; Muthukannan, M.; Ramya, M. Federated Learning-Enabled Edge Intelligence for Sustainable Smart Cities. In In Proceedings of the 2025 6th International Conference on Electronics and Sustainable Communication Systems (ICESC), 2025; pp. 1436–1443. [Google Scholar] [CrossRef]
  39. Yan, Q.; Xia, Q.; Wang, Y.; Zhou, P.; Zeng, H. URadio: Wideband Ultrasound Communication for Smart Home Applications. IEEE Internet Things J. 2022, 9, 13113–13125. [Google Scholar] [CrossRef]
  40. Li, G.; Ren, L.; Pradhan, O.; O’Neill, Z.; Wen, J.; Yang, Z.; Fu, Y.; Chu, M.; Huang, J.; Wu, T.; et al. Emulation and detection of physical faults and cyber-attacks on building energy systems through real-time hardware-in-the-loop experiments. Energy Build. 2024, 320, 114596. [Google Scholar] [CrossRef]
  41. Li, G.; Yang, Z.; Fu, Y.; O’Neill, Z.; Ren, L.; Pradhan, O.; Wen, J. A hardware-in-the-loop (HIL) testbed for cyber-physical energy systems in smart commercial buildings. Sci. Technol. Built Environ. 2024, 30, 415–432. [Google Scholar] [CrossRef]
  42. Li, R.; Li, Q.; Huang, Y.; Zou, Q.; Zhao, D.; Zhang, Z.; Jiang, Y.; Zhu, F.; Vasilakos, A.V. SeIoT: Detecting Anomalous Semantics in Smart Homes via Knowledge Graph. IEEE Trans. Inf. Forensics Secur. 2024, 19, 7005–7018. [Google Scholar] [CrossRef]
  43. Zou, Q.; Li, Q.; Li, R.; Huang, Y.; Tyson, G.; Xiao, J.; Jiang, Y. IoTBeholder: A Privacy Snooping Attack on User Habitual Behaviors from Smart Home Wi-Fi Traffic. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 2023, 7. [Google Scholar] [CrossRef]
  44. Chi, H.; Fu, C.; Zeng, Q.; Du, X. Delay Wreaks Havoc on Your Smart Home: Delay-based Automation Interference Attacks. In Proceedings of the 2022 IEEE Symposium on Security and Privacy (SP), 2022; pp. 285–302. [Google Scholar] [CrossRef]
  45. Tariq, U.; Ahanger, T.A. Employing SAE-GRU deep learning for scalable botnet detection in smart city infrastructure. PeerJ Comput. Sci. 2025, 11, e2869. [Google Scholar] [CrossRef] [PubMed]
  46. Escolar, A.M.; Wang, Q.; Calero, J.M.A. Enhancing honeynet-based protection with network slicing for massive Pre-6G IoT Smart Cities deployments. J. Netw. Comput. Appl. 2024, 229, 103918. [Google Scholar] [CrossRef]
  47. Weng, X.; Beuran, R.; Tan, Y. SBCSE: Emulation platform for smart building control system security testing. Int. J. Crit. Infrastruct. Prot. 2026, 53, 100852. [Google Scholar] [CrossRef]
  48. Morales-Gonzalez, C.; Harper, M.; Yuan, B.; Fu, X. On Software Security of Building Automation Systems. In Proceedings of the 2025 International Conference on Computing, Networking and Communications (ICNC), 2025; pp. 382–386. [Google Scholar] [CrossRef]
  49. Balamurugan, S.; Granda, S.; Haile, S.; Petersen, A.; Wang, J.; Ling, J. A cybersecurity testbed for smart buildings. National Renewable Energy Laboratory (NREL), Technical report. Golden, CO (United States), 2023. [Google Scholar]
  50. Cash, M. On Vulnerabilities of Building Automation Systems. Thesis, 2024. [Google Scholar]
  51. Runge, I.M.; Akinci, B.; Bergés, M. Challenges in Cyber-Physical Attack Detection for Building Automation Systems. In Proceedings of the Proceedings of the 10th ACM International Conference on Systems for Energy-Efficient Buildings, Cities, and Transportation, New York, NY, USA, 2023; BuildSys ’23, pp. 236–239. [Google Scholar] [CrossRef]
  52. Li, G.; Yang, Z.; Fu, Y.; Ren, L.; O’Neill, Z.; Parikh, C. Development of a hardware-In-the-Loop (HIL) testbed for cyber-physical security in smart buildings. arXiv 2022, arXiv:2210.11234. [Google Scholar]
  53. Rondon, L.P.; Babun, L.; Aris, A.; Akkaya, K.; Uluagac, A.S. PoisonIvy: (In)secure Practices of Enterprise IoT Systems in Smart Buildings. In Proceedings of the Proceedings of the 7th ACM International Conference on Systems for Energy-Efficient Buildings, Cities, and Transportation, New York, NY, USA, 2020; BuildSys ’20, pp. 130–139. [Google Scholar] [CrossRef]
  54. Wu, F.; Qu, M. A cybersecurity framework for wireless-controlled smart buildings I: Two-position control. ASHRAE Trans. Accessed via ProQuest. 2020, 126, 412–420. [Google Scholar]
  55. Rahman, M.M.; Bouhafs, F.; Hoseini, S.A.; den Hartog, F. ARProof: A cross-protocol approach to detect and mitigate ARP-spoofing attacks in smart home networks. J. Netw. Comput. Appl. 2026, 246, 104396. [Google Scholar] [CrossRef]
  56. Kostage, K.; Peppers, S.; Gao, T.; Drefahl, P.; Vallar, G.; Guo, W.; Mazzola, L.; Qu, C. HiFINS: A Hierarchical Federated Learning-Based Interactive System for Smart Home Security. IEEE Access 2025, 13, 190471–190490. [Google Scholar] [CrossRef]
  57. Barman, P.; Chowdhury, R.; Chakraborty, T.; Goswami, A.; Ghosh, S.; Saha, B. SRF2T-ID: an implementation of ensemble learning-based IDS with wireless of things secure communication for smart residency environment. Neural Comput. Appl. 2025, 37, 12525–12564. [Google Scholar] [CrossRef]
  58. Javed, A.; Ehtsham, A.; Jawad, M.; Awais, M.N.; Qureshi, A.u.H.; Larijani, H. Implementation of Lightweight Machine Learning-Based Intrusion Detection System on IoT Devices of Smart Homes. Future Internet 2024, 16. [Google Scholar] [CrossRef]
  59. Karmous, N.; Ben Dhiab, Y.; Ould-Elhassen Aoueileyine, M.; Youssef, N.; Bouallegue, R.; Yazidi, A. Deep learning approaches for protecting IoT devices in smart homes from MitM attacks. Front. Comput. Sci. 2024, 6. [Google Scholar] [CrossRef]
  60. Yeboah-Ofori, A.; Hawsh, A. Evil Twin Attacks on Smart Home IoT Devices for Visually Impaired Users. In Proceedings of the 2023 IEEE International Smart Cities Conference (ISC2), 2023; pp. 1–7. [Google Scholar] [CrossRef]
  61. Jiang, C.; Fu, C.; Zhao, Z.; Du, X. Effective Anomaly Detection in Smart Home by Integrating Event Time Intervals. Procedia Computer Science The 13th International Conference on Emerging Ubiquitous Systems and Pervasive Networks (EUSPN) / The 12th International Conference on Current and Future Trends of Information and Communication Technologies in Healthcare (ICTH-2022) / Affiliated Workshops, 2022; 210, pp. 53–60. [Google Scholar] [CrossRef]
  62. He, F.; Tong, F.; Zhang, Y. A Bi-Layer Intrusion Detection Based on Device Behavior Profiling for Smart Home IoT. In Proceedings of the 2022 IEEE 19th International Conference on Mobile Ad Hoc and Smart Systems (MASS), 2022; pp. 373–379. [Google Scholar] [CrossRef]
  63. Dai, X.; Mao, J.; Li, J.; Lin, Q.; Liu, J. HomeGuardian: Detecting Anomaly Events in Smart Home Systems. Wirel. Commun. Mob. Comput. 2022, 2022, 8022033. [Google Scholar] [CrossRef]
  64. Liu, X.; Zeng, Q.; Du, X.; Valluru, S.L.; Fu, C.; Fu, X.; Luo, B. SniffMislead: Non-Intrusive Privacy Protection against Wireless Packet Sniffers in Smart Homes. In Proceedings of the Proceedings of the 24th International Symposium on Research in Attacks, Intrusions and Defenses, New York, NY, USA, 2021; RAID ’21, pp. 33–47. [Google Scholar] [CrossRef]
  65. Fu, C.; Zeng, Q.; Du, X. HAWatcher: Semantics-Aware Anomaly Detection for Appified Smart Homes. In Proceedings of the 30th USENIX Security Symposium (USENIX Security 21), 2021; USENIX Association; pp. 4223–4240. [Google Scholar]
  66. Akestoridis, D.G.; Harishankar, M.; Weber, M.; Tague, P. Zigator: analyzing the security of zigbee-enabled smart homes. In Proceedings of the Proceedings of the 13th ACM Conference on Security and Privacy in Wireless and Mobile Networks, New York, NY, USA, 2020; WiSec ’20, pp. 77–88. [Google Scholar] [CrossRef]
  67. Arif, S.; Khan, M.A.; Rehman, S.U.; Kabir, M.A.; Imran, M. Investigating Smart Home Security: Is Blockchain the Answer? IEEE Access 2020, 8, 117802–117816. [Google Scholar] [CrossRef]
  68. Benzel, T. The science of cyber security experimentation: the DETER project. In Proceedings of the Proceedings of the 27th Annual Computer Security Applications Conference, New York, NY, USA, 2011; ACSAC ’11, pp. 137–148. [Google Scholar] [CrossRef]
  69. Siaterlis, C.; Garcia, A.P.; Genge, B. On the Use of Emulab Testbeds for Scientifically Rigorous Experiments. IEEE Commun. Surv. Tutor. 2013, 15, 929–942. [Google Scholar] [CrossRef]
  70. Neto, E.C.P.; Dadkhah, S.; Ferreira, R.; Zohourian, A.; Lu, R.; Ghorbani, A.A. CICIoT2023: A Real-Time Dataset and Benchmark for Large-Scale Attacks in IoT Environment. Sensors 2023, 23. [Google Scholar] [CrossRef] [PubMed]
  71. CIC-IDS 2018 Dataset. 15 07 2026. Available online: https://www.kaggle.com/datasets/primus11/cic-ids-2018-dataset.
  72. CIC-DDoS 2019 Dataset. 15 07 2026. Available online: https://www.kaggle.com/datasets/dhoogla/cicddos2019.
  73. Koroniotis, N.; Moustafa, N.; Sitnikova, E.; Turnbull, B. Towards the development of realistic botnet dataset in the Internet of Things for network forensic analytics: Bot-IoT dataset. Future Gener. Comput. Syst. 2019, 100, 779–796. [Google Scholar] [CrossRef]
  74. Alsaedi, A.; Moustafa, N.; Tari, Z.; Mahmood, A.; Anwar, A. TON_IoT Telemetry Dataset: A New Generation Dataset of IoT and IIoT for Data-Driven Intrusion Detection Systems. IEEE Access 2020, 8, 165130–165150. [Google Scholar] [CrossRef]
  75. Garcia, S.; Parmisano, A.; Erquiaga, M.J. IoT-23: A labeled dataset with malicious and benign IoT network traffic. 2020. [Google Scholar] [CrossRef]
  76. Ullah, I.; Mahmoud, Q.H. A Scheme for Generating a Dataset for Anomalous Activity Detection in IoT Networks. In Proceedings of the Advances in Artificial Intelligence, Cham, 2020; pp. 508–520. [Google Scholar]
  77. Meidan, Y.; Bohadana, M.; Mathov, Y.; Mirsky, Y.; Shabtai, A.; Breitenbacher, D.; Elovici, Y. N-BaIoT—Network-Based Detection of IoT Botnet Attacks Using Deep Autoencoders. IEEE Pervasive Comput. 2018, 17, 12–22. [Google Scholar] [CrossRef]
  78. Takaronis, M.; Kollarou, A.; Kampourakis, V.; Gkioulos, V.; Katsikas, S. ICSSPulse: A Modular LLM-Assisted Platform for Industrial Control System Penetration Testing. In Proceedings of the ICT Systems Security and Privacy Protection, Cham, 2026; pp. 243–256. [Google Scholar] [CrossRef]
  79. Berg, T.H.; Amro, A.; Akbarzadeh, A.; Kavallieratos, G. From Words to Wires: Toward Rapid ICS Cyber-Range Construction Using LLMs. In Proceedings of the Computer Security. ESORICS 2025 International Workshops, Cham, 2026; pp. 462–481. [Google Scholar] [CrossRef]
Figure 1. A bird’s-eye view of the articles’ screening and selection process.
Figure 1. A bird’s-eye view of the articles’ screening and selection process.
Preprints 225564 g001
Figure 2. Temporal distribution of the studies included in the survey.
Figure 2. Temporal distribution of the studies included in the survey.
Preprints 225564 g002
Figure 3. Architectural taxonomy of the 28 included testbeds.
Figure 3. Architectural taxonomy of the 28 included testbeds.
Preprints 225564 g003
Figure 4. Frequency of the most prevalent tools and components across the 28 testbeds.
Figure 4. Frequency of the most prevalent tools and components across the 28 testbeds.
Preprints 225564 g004
Figure 5. Classification of attacks demonstrated on the testbeds of the 28 included studies. Badges denote the number of studies per attack type.
Figure 5. Classification of attacks demonstrated on the testbeds of the 28 included studies. Badges denote the number of studies per attack type.
Preprints 225564 g005
Figure 6. Classification of security controls devised by the 28 included studies. Badges denote the number of studies per control type.
Figure 6. Classification of security controls devised by the 28 included studies. Badges denote the number of studies per control type.
Preprints 225564 g006
Figure 7. Attack-control coverage across surveyed smart-environment testbeds.
Figure 7. Attack-control coverage across surveyed smart-environment testbeds.
Preprints 225564 g007
Figure 8. Conceptual relationship between testbed realism, attack-evaluation mode, and the strength of defensive evidence.
Figure 8. Conceptual relationship between testbed realism, attack-evaluation mode, and the strength of defensive evidence.
Preprints 225564 g008
Table 1. Comparison of the present review with the closest secondary studies. Y: explicit and substantive coverage; P: partial or secondary coverage; : not a central analytical dimension.
Table 1. Comparison of the present review with the closest secondary studies. Y: explicit and substantive coverage; P: partial or secondary coverage; : not a central analytical dimension.
Study Principal scope Testbed-centred Architecture/realism Tooling Attack coverage Control coverage Cyber–physical coupling Trust boundary Reproducibility Evidence maturity
Santana et al. [9] General IoT cybersecurity testbeds and testbed requirements Y Y P P Y
Cintuglu et al. [10] Cyber-physical smart-grid testbeds Y Y Y P P Y P
Minani et al. [11] General IoT testing objectives, methods, tools, and challenges P P Y P
Lonetti et al. [12] Model-based security testing for IoT P Y P P P
Graveto et al. [13] Security, safety, and privacy of BACS Y Y P P
Li et al. [14] Cyber-physical security and resilient control for BASs Y Y Y P
Panahi Rizi [15] Smart-city security, privacy, technologies, and solutions Y Y P P
Demertzi et al. [4] Threats, attacks, and countermeasures across smart-city domains Y Y P
Houichi et al. [16] Smart-city security synthesis and dataset-based AI IDS P Y Y P P
Joshi et al. [17] Bibliometric trends in smart-city and ITS cybersecurity P P
This work Cybersecurity testbeds for smart buildings, homes, and cities Y Y Y Y Y Y Y Y Y
Table 2. Search string groups and their logical combinations.
Table 2. Search string groups and their logical combinations.
Group 1: Domain Terms Group 2: Testbed/Resource Terms Group 3: Security Terms
smart building, building automation, smart home, home automation, smart city testbed
dataset
benchmark
HitL
emulation platform
simulation testbed
cybersecurity
intrusion detection
anomaly detection
attack*
IDS
Search logic: Group 1 AND Group 2 AND Group 3
Table 3. Inclusion and exclusion criteria for full-text eligibility.
Table 3. Inclusion and exclusion criteria for full-text eligibility.
Inclusion Criteria Exclusion Criteria
IC1: Designs, implements, or employs a cybersecurity testbed, emulation/simulation platform, or experimentally generated dataset for smart buildings, homes, or cities EC1: Addresses smart-environment cybersecurity only conceptually or at the policy level, with no experimental resource
IC2: Provides an explicit description of at least one of: implementation paradigm, protocols, attack scenarios, or detection mechanism EC2: Reuses only generic IT network datasets (e.g., KDD Cup 99, NSL-KDD, CICIDS) with no smart-environment component
IC3: Peer-reviewed venue or authoritative technical report (NREL/PNNL/NIST), published within the 2020-2026 window EC3: Targets a different CPS domain (automotive, generic ICS/SCADA, transmission-grid only) with no smart-environment instantiation
IC4: Reports sufficient detail to enable classification across the review’s analysis dimensions EC4: Published outside the 2020-2026 window, or a duplicate publication of an already-included work
Table 6. Qualitative coverage of demonstrated attack families by defensive-control function across the surveyed smart-environment testbeds.
Table 6. Qualitative coverage of demonstrated attack families by defensive-control function across the surveyed smart-environment testbeds.
Attack family Detection Preventive hardening Response / containment Recovery / safe state Interpretation of coverage
Availability (DoS/DDoS, RF jamming) Strong Minimal Partial None identified Frequently detected through learning-based classifiers and traffic analysis [42,45,56,57,58,62]; only selected studies evaluate active blocking or isolation [46,59].
Integrity and control (FDI, command injection, rule interference) Strong Minimal Minimal None identified Live and HitL studies demonstrate injection effects and detection [44,50,51,52,54,63,65]; pre-execution command enforcement and safe-state restoration remain largely unevaluated.
Network and communication (MitM, ARP spoofing, replay, Evil Twin) Partial Partial Partial None identified Concrete mechanisms include MQTTS encryption [47], ARP-spoofing mitigation [55], secure wireless links [57], SDN blocking [59], and network slicing or honeypot isolation [46], but coverage remains protocol-specific.
Malware and botnets Strong Minimal Minimal None identified Most work concerns recognition of malware, scanning, botnet, or DDoS traffic [42,45,46,56,57,62]; live infection is demonstrated only in selected cases [42,53].
Reconnaissance and enumeration Partial Minimal None identified None identified Scanning and enumeration are usually detectable traffic classes or precursors to other attacks [42,43,49,50,56,57], with little evaluation of controls that reduce discoverability.
Privacy and eavesdropping Minimal Minimal None identified None identified Traffic analysis exposes device states and user behaviour despite encrypted payloads [43,64,66]; decoy traffic and packet-pattern obfuscation provide an isolated preventive contribution [64].
Access and authentication compromise Minimal None identified None identified None identified Credential brute forcing and WPA-handshake capture are demonstrated in only two testbeds [52,60]; corresponding operational authentication controls are not evaluated end-to-end.
Software exploitation Minimal None identified None identified None identified Fuzzing identifies memory-safety defects in physical BACnet Secure Connect devices [48], but exploit mitigation, patch deployment, and runtime protection are not evaluated.
Table 7. Evidence maturity used to interpret attack-control coverage.
Table 7. Evidence maturity used to interpret attack-control coverage.
Evidence type Interpretation for defensive claims
Live attack and control evaluation Strongest evidence: an attack is executed against real components, the control operates within the testbed, and the resulting cyber or physical effect is observed; examples include live ARP mitigation and packet-obfuscation evaluation [55,64].
HitL evaluation Strong evidence for controller and cyber-physical behaviour, although the simulated plant or selected interfaces can omit deployment-specific effects [40,41,52].
Emulated attack and control evaluation Moderate evidence: enables repeatability and scale, but can omit device firmware, wireless effects, timing variation, and implementation-specific vulnerabilities [45,46,47,59].
Replayed or dataset-based evaluation Classifier evidence: supports detection performance under the represented traffic distribution, but does not independently demonstrate prevention, containment, or live attack resilience [55,56,57,62].
Recommendations only No operational defensive evidence: a mitigation is proposed but is not implemented and evaluated in the reported testbed [44,53,60,66].
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.