Preprint
Article

This version is not peer-reviewed.

Safety-Constrained Vehicle-to-Pedestrian Guidance for Visually Impaired Pedestrians: Multi-Camera Consensus Gating in Edge-Fog-Cloud ITS Architecture

Submitted:

15 August 2026

Posted:

18 August 2026

You are already at the latest version

Abstract
Driving assistance and Autonomous driving are among the fastest evolving domains of intelligent transportation systems (ITSs). However, visually impaired pedestrians (VIPs), remain weakly protected by the roadsides or vehicle perception systems, especially when the right of way must be communicated in an accessible and auditable manner. In this paper we describe a safety-constrained vehicle-to-pedestrian (V2P) architecture designed for VIPs crossing assistance. The system creates a time bounded interaction state between the pedestrians, the roadside infrastructure, and the vehicles at Fog, to avoid permissive instructions before all the safety stacks are respected in each level where there are formal time bounded strictures with appropriate tools to formally abide by these rules during each interaction between vehicles and VIPs. Our objective is not to design a new detector, tracker, or MLLM model. Instead, we combine edge descriptors, fog level multi-view descriptors, and multi-camera multiple objects tracking technique (MC-MOT) within a latency-constrained V2P loop. Our main empirical focus is the WildTrack multi-camera continuity, with a novel Multi-Camera Consensus Gating (MCCG) safety primitive criterion; we also conducted a JAAD experiment as secondary assistive-branch for feasibility check, indicating plausible crossing-intent estimates, sub-second restrictive alerts, and actionable grounded guidance. On WildTrack the reproduced Fog continuity branch reaches a MOTA = 88.8%, IDF1 = 91.7%, and HOTA = 65.5%. In the WildTrack settings with N = 7, MCCG with M = 4 blocks 75% of ID-switch considering a 5.0% permissive-eligibility cost. These findings support feasibility of the proposed ITS interaction loop, while leaving field deployment of V2X stack validation, and user studies for future work.
Keywords: 
;  ;  ;  ;  ;  

1. Introduction

Among all urban mobility road users, visually impaired people (VIPs) are among the most vulnerable. Modern assistive technologies, such as smart canes, wearable sensors (e.g. vest or goggles with sensors), ultrasonic devices, camera-based navigation aids, audio prompts, and haptic feedback, are available today to improve awareness of surroundings and confidence in unfamiliar environments [1,2,3]. However, existing assistive technologies generally operate only on the pedestrian side. These systems advise VIPs about the surrounding environment, but do not require a vehicle, an RSU, or a traffic control service to acknowledge the VIPs' priority.
Modern ITS deployments involve more vehicles, RSUs, cameras, and traffic control centers, and their developments predominantly remain vehicle-centric, through traffic optimization, obstacle detection, collision avoidance, without creating an accessible and assistive interaction channel for visually impaired pedestrians [4]. The critical part of modern urban mobility where VIPs can be temporarily hidden are street crossings. Busses, turning vehicles, railing structures, crowds, or street obstacles can temporarily hide VIPs from the perspective of the drivers or street equipment (e.g. cameras, sensors on the poles). Consequently, a single camera on a crosswalk can lose track of a visually impaired pedestrian if it is occluded. Similarly, a car might break without adequately indicating its intention. In addition, wearable assistive devices might produce a general alert that the VIP receives, without knowing whether the intent of the vehicle is specific and binding, with a smart-city governance loop fallback. This gap is critical as assistive technologies, and ITS infrastructures have evolved along separate trajectories, leaving pedestrians with obstacle and environment awareness devices, and vehicles with perception and roadside analytics tools to optimize traffic flow and avoid collisions only, without including VIPs into this loop.
In our paper we try to address this problem by defining a mutual recognition concept in which the protected pedestrian, the infrastructure, and the vehicle share a consistent and auditable understanding of presence, priority, and guidance. The key requirement is deliberately conservative: the system must not issue a permissive message such as "you may cross" unless the protected-pedestrian track is current, the conflicting vehicle has acknowledged the yield directive, and no uncertainty flag is raised. Our system does not replace conventional traffic rules or human supervision, it specifies how existing sensing, tracking, communication, and assistive-feedback modules could be orchestrated in a manner that VIPs receive an accessible confirmation instead of being treated only as a detected obstacle.

1.1. Contributions

In this work we investigated safety and guidance assistance provided to pedestrians. The contribution is not a new perception backbone, it is a transport-oriented interaction contract and evaluation protocol that binds sensing, communication, continuity, accessible feedback, and fallback logic into a measurable crossing service:
  • ITS Mutual Recognition for safe pedestrians: The vehicle's and pedestrian's interaction is viewed as a duration-limited relationship in the moment the vehicle can still potentially hit the pedestrian. It operates through cross-view continued pedestrian real-time observation, vehicle acknowledgment for giving the way, formal orchestration from the system.
  • Roadside delay-based negotiation (RDN): To address and prevent potentially dangerous permissive instructions an RDN is set up and is provided for interaction between the drivers and the city governance system to further improve communication and security. A conservative protocol is set to govern how information will be communicated with all relevant protagonists.
  • The current paper discusses the use of a 3D multi-camera system tracking, V2P signals, MLLM narrations for visual context and all required safety standards, without the requirement of sending raw video data, which raise privacy-related concerns, not to mention the large amount of processing power required for such operation.
  • A framework that is compelling with several performance measures based on specific datasets. We provide the performance results of our experiments on WildTrack dataset as the main empirical substrate because it directly tests the continuity requirement behind safety-constrained guidance, and JAAD dataset as a secondary feasibility check for the assistive interaction branch. HOTA is reported for the Fog continuity branch (65.5%), which is new related to the baseline studies.
  • Multi-Camera Consensus Gating (MCCG): in MCCG, an individual will not receive guidance until M out of N total number of Cameras successfully see the same person and communicate with the drivers of other cars for formal yielding and binding contract control. We propose MCCG and its sweep protocol as a reusable benchmark for safety-aware assistive multi-camera MOT evaluation.

1.2. Scope

Our paper is an architecture and system integration study that includes benchmarks using public datasets but is not intended to be a deployment-ready system. The WildTrack part investigates the multi-camera continuous monitoring criterion to fulfill the MCCG requirements and forms the bulk of this study. We have chosen to keep the JAAD part a secondary method to confirm the detection and warning of pedestrian intent and intent warning time, along with on-road assistance: the system will assist in the closed loop V2P cycle and reproduce a pedestrian crossing assisted by our system. Weather variations, the presence of failed street equipment or infrastructure, losing V2X packets due to interference or infrastructure faults, mixed traffic types, driver confidence in automated guidance systems and legally enforced traffic laws are examples of situations that have not been covered within the paper. The full cycle of road deployment, control with a real vehicle, wide-area communication infrastructure in a large community, regulatory compliance, and testing on VIP pedestrians are future work, and need a full orchestration mechanism with all the road traffic protagonists, including police for security cameras access, city infrastructure managements, and traffic control units.

3. Mutual Recognition as a Safety Protocol

3.1. Failure Modes and Conservative Handling

Guidance messages may influence pedestrian behavior in safety-critical situations; therefore, several worst-case scenarios must be considered. The system must implement conservative fallback decisions before issuing permissive instructions such as “you may cross”. A permissive instruction is generated only when three conditions are simultaneously satisfied:
(i)
The protected pedestrian path remains active above a predefined duration;
(ii)
The relevant vehicle or roadside unit has acknowledged the yield directive within the allowed reaction window;
(iii)
No uncertainty flag is active, including occlusion timeout, identity ambiguity, or communication loss.
We define this synchronized state as mutual recognition.
Let Pᵢ denote a protected pedestrian, Vⱼ a vehicle at sight, and Ω a crossing zone. At time t, the Edge tier emits an observation token:
oc,t = {bc,t, ec,t, τ(t), c}
where bc,t is a boundary box or location in a ground-plane, ec,t is a compact appearance descriptor, τ(t) is a time stamp, and c identifies the detection camera. The Fog tier fuses observations into a protected-pedestrian track belief:
Bi(t) = {x̂i(t), qi(t), ai(t), ui(t)}
where x̂i(t) is the estimated pedestrian state, qi(t) is the confidence in identity, ai(t) is the age of the last confidence observation and ui(t) is a set of uncertainty flags such as occlusion timeout, identity ambiguity, stale communication, or missing acknowledgment.
The Fog tier issues a yield directive token:
Yij(t) = {Pi, Vj, Ω, t, TTL, directive ID}
The vehicle must return an acknowledgment token Aij(t) before the pedestrian-facing device can deliver a permissive crossing message. The guidance function is therefore:
gi(t) = Cross if qi(t) ≥ τq, ai(t) ≤ Tocc, Aij(t) ≤ Tack, t - tY ≤ TTL, and ui(t) = ∅; otherwise, gi(t) = Stop/Wait.
Equation (1) is the core safety rule: uncertainty cannot produce permissive instruction. If the system is unsure, the pedestrian receives a restrictive alert such as “Stop” or “Wait, and the vehicle yield obligation remains active until the state is resolved or manually overridden by established traffic procedures.
Table 1 summarizes the failure-handling logic used by the proposed protocol. The table makes it explicit that permissive guidance is not a default output of the MLLM. It is a gated safety decision generated only after Fog-level state validation and vehicle acknowledgment.

3.2. Accessible V2P Intent Signaling

External human-machine interfaces, i.e. eHMIs and similar crosswalk displays projects are usually intended for pedestrians without impairment and become not viable for blind or impaired-vision pedestrians. For the VIPs the intention needs to be rendered using audio, haptic, vibrotactile and over phone/wearable-device. Our V2P signaling is proposed as a safety primitive of the user-interface detail in the proposed system. Our framework verifies the commitment that vehicle Vⱼ is yielding to protected pedestrian Pᵢ in crossing zone Ω, before the VIPs wearable renders a formal instruction. Previous work on communication of autonomous-vehicle intent to pedestrians and other vulnerable road-users also drives treatment such that explicit V2P signaling is perceived as a relevant safety interaction channel rather than a purely visual interface [17].
Figure 1 and Table 2 define linkages of accessible V2P signaling to various operational elements of the system in rest of the manuscript. in Figure 1 we illustrate the closed loop of mutual recognition; the Fog tier issues a time stamped yield directive followed by vehicle response or failure to respond, on the acknowledgment of vehicle response, on the pedestrian-facing channel either permissive or restrictive guidance is rendered, and finally the event is recorded for governance. In continuation, Table 2 in turn presents the loop in terms of architectural components, marking prior capabilities in use, and on the special integration role that each component assumes as prerequisites.
Together, Figure 1 and Table 2 show that accessible V2P intent signaling is not an optional interface layer, it is the visible endpoint of a constrained safety process, where perception and tracking create a protected-pedestrian state, Fog arbitration binds this state to a vehicle obligation, and the pedestrian receives guidance only after the obligation is recent under TTL bounds and formal acknowledgment. This interpretation also defines the boundary of our paper: the work integrates established detection, tracking, re-identification, V2P, and MLLM capabilities into a conservative interaction contract rather than claiming novelty in each underlying specific research field.
This safety primitive requires architectural support across the Edge, Fog, and Cloud tiers. Section 4 formalizes how these levels interact and how accountability, privacy, and latency constraints are embedded into the design.

4. Edge-Fog-Cloud Architecture

In this section, we present the building design and explain the reasoning behind the Edge-Fog-Cloud stack.
Figure 2 summarizes the architecture and its ethics-by-design feedback loop. The Edge tier performs on-device perception and publishes compact presence beacons (bbox, time, camID, low-dimensional appearance embeddings, no raw frames) over lightweight communication protocols (e.g. MQTT/V2X), while providing accessible audio/haptic feedback to the VIPs [18,19]. The Fog Roadside Units (RSU) fuse multi-camera observations, maintain identity continuity through occlusion, arbitrate right of way, and issue time-stamped yield directives to vehicles, with operational logging for audit [17], The Cloud hosts MLLM-based scene understanding for human-readable guidance, a model/policy hub, and governance services aligned with ISO/IEC 42001, IEEE 7000, and the NIST AI RMF ([20], p. 2023,21,22]). Ethics-by-design feedback closes the loop: uncertainty downgrades the guidance to “Stop/Wait” and all critical acts (yield issued/acknowledged, instruction delivered) are recorded for accountability.
Our architecture intentionally reuses strong modules from prior work, including lightweight edge detectors [18], MLLM scene understanding [19], multi-object and multi-camera tracking reviews [8], transformer-based tracking [9], graph association [10], and symmetry-driven re-identification descriptors [11]. The paper contribution is the constrained integration pattern: modules are linked through time-stamped tokens, conservative fallback rules, descriptor-only transfer, and governance logs so that a VIP crossing interaction can be treated as an auditable obligation rather than a transient detection event. Table 2 clarifies this boundary and distinguishes prior technical capabilities from the role they play in our architecture.

4.1. Edge Tier: Immediate Perception and Local Feedback

The Edge tier provides immediate perception and local safety reflexes, further ensuring accessible feedback. The system handles the above subjects while maintaining accountability hooks for audits and ethics-by-design requirements.
Lightweight deep learning models running on embedded hardware can perform real-time pedestrian detection and publish event present messages with latencies on the order of hundreds of milliseconds, using low bandwidth protocols such as MQTT rather than streaming raw videos [18]. This capability allows an Edge node to declare the presence of pedestrians fast enough for emergency braking or yielding decisions. Because the payload consists of compact descriptors (bounding box coordinates, timestamps, low dimensional embedding), bandwidth, and privacy cost are minimal [18,23]. In addition, impaired pedestrians receive explicit feedback about the V2P negotiations. This clears up the ambiguity regarding the orchestration of smart urban traffic with V2P [17].
Also, the Edge tier includes the visually impaired person's wearable device, smart cane or phone, the vehicle's onboard sensing and signaling system, and further any ultra-local roadside micro-node such as a smart pole with embedded camera and small processing unit. The system must work immediately: it should sense danger, announce it, and warn users in a time-bounded manner to react [18].
For the system requirements, accountability hooks are implemented in terms of performance tracking.
At the Edge, every safety action statement like “Vehicle 12 is yielding to you” is logged with basic details (time, pedestrian ID, vehicle ID, location) for further audit purposes if needed. The Edge does not store high-quality personal videos. In addition, it lacks the capability to archive detailed video content. Identities are processed forward using only small descriptors, not biometric data [10,11,23].

4.2. Fog Tier: Multi-Camera Continuity and Yield Arbitration

The Fog tier, located on an RSU or intersection controller, is responsible for spatial persistence. It combines observations across cameras, maintains the protected pedestrian track, and assigns yield obligations to approaching vehicles. This is where MC-MOT becomes safety-relevant: if pedestrian Pi disappears behind a bus, the Fog tier must preserve the identity hypothesis and keep the vehicle yield directive active until the system has sufficient evidence to resume permissive guidance or to remain restrictive.
The earlier broad discussion of distributed coordination is narrowed here to the safety-relevant rule in Eq. (1). The emphasis is not on introducing a new consensus algorithm, but on making delays operationally visible through token freshness, acknowledgment windows, and conservative fallback. This formulation is more directly related to VIP safety because the system can continue to require a vehicle yield obligation while withholding permissive guidance whenever the state is stale, ambiguous, or unacknowledged.
Modern MC-MOT systems for smart transport solve the association problem by connecting detections from different cameras and time periods to paths using only graph-based network methods [8,9,10]. These solvers use appearance, motion, and spatial cues to track the same protected pedestrian 'P' in different observations, even when the person is temporarily occluded. TrackFormer and similar transformer trackers handle tracking as sequence prediction instead of connecting objects later. This approach maintains object identity even in challenging situations [9]. the Fog tier can continue protecting the pedestrian 'P' even if 'P' goes behind a bus for two seconds and cannot be seen directly [8,9,10,11,14].
When the field-of-view changes from one area to another, and the cameras don't overlap, the handover happens through re-identification (Re-Id) methods, including symmetry-driven appearance descriptors (e.g. SDALF), to consistently determining that the new detection is still for person 'P', even without streaming facial biometrics [11]. This further maintains continuous protection across different areas, not just within the view of one camera.
From a control-theoretic perspective, distributed coordination under asynchronous interactions and communication delays has also been studied in the recent consensus-tracking literature. Recent work on asynchronous cooperation-competition networks and delayed substochastic consensus processes reinforces the importance of delay-aware agreement mechanisms when the safety-relevant state must remain consistent between distributed agents. Although these studies do not address VIP mobility directly, they support the design choice of handling right-of-way arbitration and acknowledgment logic in the Fog tier through explicitly time-bounded coordination rules [24,25].
Right-of-way disputes are resolved through arbitration processes. Proper enforcement mechanisms ensure compliance with established traffic regulations.
In the Fog tier's tracking of protected pedestrian 'P' in the crosswalk 'Ω’, it sends yield commands to relevant vehicles about giving way to Pedestrian 'P' immediately. This approach assigns clear right-of-way rather than using basic collision avoidance methods [17]. The vehicle must show the pedestrian that it will stop by using accessible V2P signals, such as audio/haptic confirmation messages, which are sent back to the protected person. This arbitration step makes the social contract work the same way; the pedestrian is not just detected, but acknowledged, protected, and informed.
The minimal fog tier, but binding logs events such as “Vehicle 12 yielded Pedestrian 'P' at 12:05:14 in crosswalk 'Ω'” as well as whether the confirmation message was delivered to 'P' [22], reduces privacy concerns by removing raw identity video, keeping accountability hooks for further review after incidents.

4.3. Cloud Tier: MLLM Narration, Policy, and Governance

The cloud tier handles three main topics: (i) Semantic interpretation and narration of the scene for the pedestrian; (ii) Distribution of the global policy to Fog and Edge tires; (iii) Control of the governance system and enforcement of rules in different layers.
Semantic narration occurs through multi-modal language models (MLLMs) that can fuse sensory context into accessible, human-centered guidance for users with visual impairments, in near real-time [19]. the cloud layer uses organized information from the fog layer regarding pedestrian safety and vehicle status to create clear instructions like “You may start crossing now; the yielding car is directly ahead of you and stopped.”. This step elevates the interaction from raw perception to accessible and assistive communication.
The Cloud tier works as a policy and model hub regarding data management. It distributes: (i) The system sends updated perception models to Edge devices, for example, to improve pedestrian detection systems that work well in dark conditions or during rain; (ii) Fog will send update to Association/tracking policies, stricter rules to allow walk in school areas, and longer waiting times for cane users at night will be sent to Fog; (iii) Fairness and retention constraints, such as the maximum time to store embedding data and conditions for audit records, applied to all levels ([20], p. 2023,21,22).
This central system sends rules to all traffic points to ensure consistent enforcement. The network follows standard mobility guidelines and privacy rules, rather than each intersection making its own decisions further.
Critically, the Cloud tier is where governance frameworks are embedded as operational controls. AI management and accountability standards such as ISO/IEC 42001 (AI management systems), IEEE 7000 (ethics-by-design for autonomous and intelligent systems), and the NIST AI Risk Management Framework (risk, bias, transparency, human oversight) are applied here as system requirements ([20], p. 2023,21,22) [26].
In practice, this means:
-
Audit of bias in detection and yielding decisions.
-
Documentation of how and when the “yield” directives were issued and enforced.
-
Traceability of any auditory or haptic assistive instruction provided to the pedestrian.
These governance hooks ensure accountability throughout the system stack, enabling reconstruction whenever auditing is conducted.
We have designed the full picture of our architecture, from the roles and artifacts Figure 1, Table 2, and the formal specifications of our architectural design Figure 2. In section 5, we quantify the operational budgets that make the mutual recognition loop viable in practice (latency, bandwidth/payloads, and safety-critical messaging) and specify the conservative fallback under uncertainty.

5. Real-Time Constraints: Latency, Bandwidth, and Safety-Critical Communication

At this point, we have explained what the system actually accomplishes (mutual recognition), how the respective calculations are carried out (different levels), and we will now examine whether this approach can function effectively under reasonable time/bandpass limitations while ensuring adequate guaranties for both the system's reliability and pedestrian safety. To address these requirements, we must consider at least three important factors: First, the full duration, from the moment a person is recognized to be intending to cross, up to the moment the car yields and the person receives an explicit OK to cross, must fall within the typical human range for reaction, specifically under 1 to 2 seconds. This time is sufficient both to alert pedestrians to yield and to let them cross the intersection ([21], page 2428). Second, communication constraints for bandwidth and privacy must be addressed. We limit our reliance on visual input not to disrupt our wireless data transmission capabilities (which include both V2X and backhaul systems). This is accomplished through minimal embedding and specific alert events which can allow us to perform the service citywide, even when we cannot pay the price of sending visual information of the pedestrians [9,11,18]; Third, determine the the escalation process when the connection has been temporarily lost (longer occlusion; failed car acknowledgment early response), the wearable device will issue a cautious alert (Never cross) and the entire system will log an incident to enable a subsequent study of the system's management practices. If there are problems with RSU warning messages or when other non-threatening pedestrians walk near, then the pedestrian must use other mechanisms to be able to cross [21,22].

5.1. Latency, Bandwidth, and Conservative Escalation

In the context of the architecture of Figure 1, four stages are necessary for an efficient latency operation.
Edge presence beacon: Lightweight pedestrian detection operates at frame times~150–300 ms. It publishes a presence event (bbox, timestamp, cam/agent ID, low dimensional appearance embedding) over MQTT/V2X. payloads are tens to few hundred bytes; transmission to Fog is essentially instantaneous on local links [18].
Fog fusion and arbitration: The Fog fuses multi-camera observations, maintains identity continuity through short occlusions and across non-overlapping views, and binds the obligation: “Vehicle 12 must yield to pedestrian 'P' in the crosswalk 'Ω' now”. This runs on the intersection rate time scale (tens of milliseconds for association/decision) [8,14].
Vehicle explicit intent compliance signaling: The vehicle receives the yield directive and performs braking/hold while exposing explicit intent through the eHMI/V2P medias. This step must remain sub-second to be meaningful in real scenarios.
Accessible confirmation for the pedestrian: A grounded message is generated at Fog (if high-compute GPUs are deployed) or Cloud: “The car ahead has stopped and is approaching you. Cross straight now”. The delivery to the mobile must feel immediate <1 s for “Stop/Do not stop”; <1–2 s for “You can cross now” [1,2,3,17].
The presence, arbitration, yielding, and confirmation loop should close in 1 to 2 s, to be consistent with the edge inference latencies reported and the RSU-grade decision times.

6. Experiments and Results

This section reports on benchmark-based component evidence for the proposed architecture. The main empirical contribution is the MCCG analysis of the "Wild Track" dataset, because it evaluates how multi-camera redundancy can gate permissive guidance under identity-switch risk. The JAAD branch is reported as a compact assistive-branch feasibility check for intent estimation, warning latency, and grounded guidance actionability. The results should not be read as a complete field trial, a validated V2X deployment, or proof that the entire architecture will generalize without additional roadside testing.

6.1. Datasets, Splits, Hardware, and Hyperparameters

Table 3 summarizes the experimental protocol. The JAAD branch is deliberately secondary and is used to check whether the pedestrian-facing direction branch can meet basic feasibility expectations about crossing intent, latency, and actionability using the public JAAD benchmark [12]. The WildTrack branch evaluates the multi-camera continuity requirement that cannot be tested with monocular clips, using the public WildTrack benchmark and its camera calibrations [13].

6.2. Assistive-Branch Feasibility Check on JAAD

Table 4 reports the results of the JAAD branch supporting the feasibility check. The table is retained to show that the proposed V2P loop has a plausible pedestrian-facing component: crossing intent can be estimated, restrictive messages can be delivered quickly, and grounded guidance can usually be expressed as actionable instruction.
Narration actionability was coded as the proportion of generated guidance messages that contained a specific state-based action or waiting instruction. Ambiguous or Fog-inconsistent messages are counted as non-actionable. The following sections therefore focus on WildTrack and MCCG, which directly evaluate the multi-camera continuity and identity-safety question.

6.3. WildTrack MC-MOT Positioning and Baselines

Table 5 reports the WildTrack reproduction used as the experimental substrate for the safety analysis, rather than as a claim of tracking superiority. The MOTA and IDF1 values for KSP-DO, EarlyBird, TrackTacular, and the attention-aware tracker are reported as contextual comparison values for the WildTrack dataset [14]. The Fog continuity branch is reproduced on a GPU cloud infrastructure and achieves MOTA 88.8%, IDF1 91.7%, and HOTA 65.5% under the present evaluation protocol. Baseline HOTA values remain non-comparable as they were not reported under baseline literature.
The numerical trend in Table 5 is informative for the ITS argument. The progression from KSP-DO to recent multi-view approaches shows that modern multi-view association can provide the continuity level needed by protected-pedestrian services, with MOTA increasing from 69.6% to 92.7% and IDF1 increasing from 73.2% to 96.1%. The reproduced Fog branch remains close to the EarlyBird comparison values and is used here mainly to generate realistic identity-switch cases. The main empirical claim is therefore not tracking superiority, but the safety-efficiency effect of MCCG in gating permissive guidance.
Table 5 representative WildTrack comparison and reproduced Fog continuity branch under the present evaluation protocol. The baseline MOTA and IDF1 values are reported from [14]. HOTA is computed using the higher-order tracking metric of Luiten et al. [15]. HOTA was not reported in a comparable study. Higher values are better.
The use of HOTA and explicit identity-failure reporting is consistent with recent trends in quantitative tracking validation for public transport spaces [16]. Moreover, we report DetA = 87.2% and AssA = 53.6%, which shows that the reproduced Fog continuity branch localizes pedestrians reliably but loses identity consistency more often than it loses detections

6.4. Multi-Camera Consensus Gating (MCCG)

The fact that the reproduced Fog continuity branch localizes pedestrians reliably but loses identity consistency more often than it loses detections is a very important finding of our work regarding safety-constrained V2P guidance: a localization error may keep the system conservative, whereas an identity switch during a crossing interaction can attach the yield obligation or permissive guidance to the wrong pedestrian. MCCG is introduced as the main empirical contribution of the paper because it turns the geometric redundancy of a multi-camera ITS deployment into an explicit safety gate.
gi(t) = Cross if qi(t) ≥ τq, ci(t) ≥ M, ai(t) ≤ Tocc, Aij(t) ≤ Tack, t - tY ≤ TTL, and ui(t) =;otherwise, gi(t) = Stop/Wait.
The camera support cᵢ(t) of a pedestrian detection is determined by projecting the estimated Fog-tier ground plane in camera i and accepting the observation if the projection falls inside the image with positive depth. MCCG does not re-detect every pedestrian but reuses the camera calibrations typically already performed in a multi-view system, thereby introducing no extra training overhead but serving as a solid safety benchmark for an assistive multi-camera MOT.
Table 6 shows a sweep of MCCG for M ∈ 1 to 7 on WildTrack. 'Permissive eligibility' is the fraction of tracks that satisfy the MCCG criteria, while 'overhead' is 1 minus this fraction. 'Switch-frames blocked' shows the number of 4 frames with an ID switch for which MCCG flagged the frame as 'uncertain' (i.e., where the best guess track had fewer than M cameras supporting it. At M=4, MCCG successfully blocks 3 of 4 frames with ID switches (75%), but only blocks 5% of all tracks, yielding 95.0% permissive eligibility. Increasing M to 6 yields complete coverage (4/4 = 100%), but decreases permissive eligibility to 62.4%, too high for crowded scenes. We therefore propose M=4 as the operating point for WildTrack with this specific tracker configuration. Because only 4 frames contained an ID switch, the 75% blocking rate is preliminary safety information: the 95% confidence interval for blocking 3/4 switch frames are broad (19.4%-99.4%), but the estimate of permissive eligibility at 95% is more stable as it is derived from 958 frames.

6.5. Operational Contract

Table 7 summarizes the operational contract of the mutual-recognition loop by linking each segment to an artifact, target, and fallback condition. The table is intended to make the proposed architecture auditable and traceable: each permissive or restrictive guidance message can be traced back to specific tokens, acknowledgments, uncertainty flags, and governance records.
Table 7 shows that the operational contract is organized around explicit artifacts rather than informal assumptions. Each stage produces a token, acknowledgment, utterance, or log entry that can be inspected after the event. The main design insight is that permissive guidance is not authorized by a single perception result or by unrestricted narration alone. It requires a current protected-pedestrian state, a new yield directive, vehicle recognition, and absence of active uncertainty. This contract also makes the system easier to test because each fallback condition can be reproduced in simulation or field trials.

7. Trustworthy AI Governance and Privacy for Inclusive Urban Mobility

In this section, we formalize how governance, accountability, and human protection are embedded in the Assistive Vehicular Intelligence architecture. The core claim of the present work is not only to solve the problem of real-time mutual recognition of the vehicle and a vulnerable pedestrian, but we also accomplish this in a way that is auditable, privacy-protected, complies with AI governance best practices codified by international standards, and is expressly biased in favor of pedestrian safety.
We ground this section in three complementary governance frameworks:
-
ISO/IEC 42001 Artificial intelligence management system: governs how AI systems should be run in terms of organizational accountability, risk management, and life cycle stages within which the AI systems are implemented. It treats AI safety and compliance as an ongoing management obligation, not a one-time certification ([20], p. 2023]).
-
IEEE 7000-2021: a standard that establishes a formal method for surfacing and resolving ethical issues during system design, including human well-being, unbalanced power, and foreseeable social disruption, especially for vulnerable stakeholders [26].
-
NIST AI Risk Management: defines a risk-oriented approach to trustworthiness, emphasizing transparency, accountability, explainability, and harm mitigation [22].

7.1. Governance Loop

In contrast with frameworks which deal with pedestrians as an obstacle avoidance subject, our architecture integrates governance and flow control rights. This paradigm is achieved through three elements:
Tracking obligation: When Fog issues a right-of-way percept (“Vehicle 12 must yield to Pedestrian 'P' in the crosswalk 'Ω'”), which is not just a percept event but an immediately enforceable obligation that is recorded with the same time stamp as the event issued, the vehicle identifier (Vehicle 12) and a pedestrian identification token (P).). This creates an accrual record: “Vehicle 12 was instructed to yield to 'P' at 't' in 'Ω'”. This aligns with ISO/IEC 42001 requirements for operational traceability of AI-based decisions, which stipulates that stakeholders must be able to re-construct what an AI system was instructed to do, under which conditions, and under which risk ([20], p. 2023).
Human-centered harm prevention as first-class behavior: Should the vehicle fail to comply within the allowable reaction window, the system proceeds with the following: (i) Immediately instructs the pedestrian “Do not cross/Stop,” with greater urgency than the earlier permissive message; (ii) Logs a non-compliance event for post-incident review. This is a direct example of the IEEE 7000 functional expectation requirement that systems must identify ethical failure modes and implement pathways that prioritize human well-being in the face of uncertainty [26]. In other words, non-compliance is not silently overlooked, it results in a safeguarding reaction from the vulnerable user and an auditable message to governance.
Risk-based fallback behavior: The conservative “safety block” reaction to ambiguous or degraded-visibility conditions, described above, is mandated and enforced. It is not a heuristic convenience. It is a direct embodiment of the NIST AI RMF principle, which states that trustworthy AI systems must recognize uncertainty, communicate about it, and mitigate the propagation of potential harm [22].
The edge wearable is never allowed to “hallucinate confidence”. It does not issue a “Cross now” instruction unless the system is in a verified safe state. If confidence decreases, it overrides any previous instructions, into the most conservative ones, and the downgrade is recorded [1,2,3,4].
Together, these three mechanisms mean that governance is in the loop, at the instant decisions are made, not only afterward.
Governance is specified as a runtime control loop. Figure 3 introduces this aspect before the detailed discussion of governance functions and privacy constraints. The figure shows how the same technical artifacts used for safety decisions, including yield directives, vehicle acknowledgments, wearable guidance, and fallback events, also become the audit evidence required for traceability, ethical mitigation, and risk fallback.
The governance loop has four runtime functions. First, it records whether a protected-pedestrian track was created, updated, lost, or recovered. Second, it records whether the vehicle yield directive was issued and acknowledged. Third, it records the exact class of pedestrian guidance provided: restrictive, permissive, or fallback template. Fourth, it stores enough context for incident reconstruction without creating a city-wide archive of identifiable video.
This design aligns with ISO/IEC 42001 for AI management systems, IEEE 7000 for ethics-by-design processes, and the NIST AI RMF for risk identification, mitigation, monitoring, and transparency. In practice, these standards are mapped to operational controls: threshold documentation, policy versioning, audit schemas, retention limits, failure-trigger logs, and human-review pathways.

7.2. Privacy and Data Minimization

The most common opposition to pedestrian safety of systems in a smart city is the possible use of such data in surveillance. Public tolerance is low for continuous facial tracking in shared space, especially for vulnerable populations. The proposed architecture addresses this at design time and runtime:
No continuous high-resolution face video exits the Edge: At the pedestrian-facing node there is on-device detection and local feature extraction. The only information sent upstream is a compact descriptor, a time 't', a camera ID, and an appearance embedding or structured descriptor for re-identification between cameras. This meets two constraints: (i) It preserves bandwidth; the upstream load for an event goes to a few hundred bytes, not megabits per second of video; (ii) Prevents the uncontrolled propagation of personally identifiable raw imagery.
This aligns with NIST AI RMF emphasis on minimization of sensitive information exposure and attack surface [22], and with ISO/IEC 42001 expectations around demonstrable data handling control in AI-enabled operations ([20], p. 2023).
Association and re-identification are appearance based, not biometric: The Fog does not match pedestrians across non-overlapping cameras using face crops. Instead, it uses structured appearance descriptors, such as SDALF-like signatures, which combine color layout, symmetry/asymmetry cues, and localized region features to maintain continuity of identity between views in crowded scenes [11]. With such settings, the Fog can confidently assert: "This is still Protected Pedestrian 'P'; keep yielding” as the person walks from Camera A to Camera B, without central streaming or biometric storage.
Retention is explicitly scoped to accountability: Our architecture does not propose indefinite city-wide logging. Logged events are tailored to compliance and incident reconstruction. These records satisfy the ISO/IEC 42001’s requirement for the auditability of AI-driven operational decisions ([20], p. 2023) and IEEE 7000’s expectation that vulnerable humans can contest unsafe outcomes through traceable system behavior [26].
Crucially, the log is for duties and instructions, not for full-resolution identity video. This is an important distinction for regulatory review and public acceptance.
In summary, the system treats privacy not as future compliance only, but as an architectural invariant to assure the public of the scope and finalities of such efforts.

8. Discussion, Limitations and Future Work

The paper positions the contribution narrowly and defensibly. It specifies how TrackFormer, graph neural association, SDALF descriptors, eHMIs, V2P communication, or MLLMs are used when they mediate a safety-critical interaction for a VIP. The main object of research is the protected-pedestrian mutual-recognition contract: a latency-sensitive, privacy-preserving, and auditable state in which vehicle yielding and pedestrian guidance are coupled.
Several limitations remain. First, the architecture has not yet been validated in full roadside deployment with live vehicles, real RSUs, diverse weather conditions, lighting changes, heterogeneous traffic, and mixed automated/human-driven vehicle behavior. Second, the V2P acknowledgment channel has been specified architecturally, but requires communication-stack testing under packet loss, congestion, interference, authentication failures, and delayed acknowledgments. Third, the JAAD and WildTrack evaluations support the selected branches of perception, timing, and continuity, but do not reproduce the full embodied experience of visually impaired pedestrians at a crossing. Fourth, the MLLM narration branch has been evaluated for actionability, but future work should compare MLLM-generated guidance against template-only guidance with VIP participants and institutional ethics approval. Fifth, MCCG has been validated in WildTrack with a fixed camera layout; its recommended M value may differ between scenes with different numbers of cameras, coverage patterns, or pedestrian densities, and the sweep protocol should be repeated for each new deployment context. Extending HOTA and MCCG reporting to all tracker baselines under a shared evaluation script and detection set remains an open benchmark task.
Future work will therefore focus on five tasks: (i) end-to-end pilot deployment at an instrumented crossing, (ii) broader MC-MOT benchmarking with shared detections, MOTA, IDF1, HOTA, and MCCG sweep reporting across tracker baselines, (iii) user studies comparing audio, haptic, and combined feedback for VIP comprehension, trust, cognitive load, and response time, (iv) formal verification or simulation-based stress testing of the delay-aware arbitration rule under bounded communication delays, and (v) validation of MCCG across scenes with different camera counts, densities, and occlusion patterns to establish a generalized M-selection guideline. These steps are necessary before the architecture can be described as deployment ready.

9. Conclusions

This paper presented a proposal for a bounded Edge-Fog-Cloud architecture for inclusive smart mobility centered on visually impaired pedestrians. The central argument is that pedestrian safety in AI-mediated traffic should not be reduced to object detection or isolated wearable warnings. Instead, cross-interaction should be treated as a mutual-recognition process in which the protected pedestrian, road infrastructure, and relevant vehicle share a time-bound state with respect to presence, priority, yield obligation, and accessible guidance. In this formulation, the visually impaired pedestrian is not only detected as a road obstacle but also a protected participant whose guidance depends on validated tracking continuity, fresh communication tokens, and explicit vehicle acknowledgment.
The contribution is therefore intentionally framed as a system-integration and architectural contribution rather than as a claim of a new standalone detector, tracker, re-identification model, V2P interface, or MLLM. The manuscript specifies how existing components can be organized into a safety-oriented interaction contract: edge nodes publish compact perception descriptors, the Fog tier maintains multi-camera continuity and performs delay-aware arbitration, the vehicle must acknowledge a yield directive before permissive guidance is issued, and the Cloud tier supports grounded narration, policy distribution, and governance. This organization makes the safety rule explicit: uncertainty, stale state, missing acknowledgment, identity ambiguity, or narration failure must lead to restrictive guidance rather than permissive crossing instructions.
The empirical results reported in the article support the feasibility of key components of this architecture without implying full deployment readiness. The JAAD-based branch is treated as a secondary feasibility check for pedestrian intent, guide timing, and narration actionability. The main empirical contribution is the WildTrack continuity protocol and MCCG analysis, which position the Fog branch in relation to representative multi-view tracking baselines, report a HOTA value (65.5%) under the reproduced evaluation protocol, and introduce Multi-Camera Consensus Gating as a proposed benchmark primitive for safety-aware assistive MC-MOT. These results help clarify the operational budgets and technical plausibility of the proposed loop, but they remain component-level evidence. Full validation still requires roadside pilots with live RSUs and vehicles, communication-stack testing under congestion and packet loss, uniform MC-MOT benchmarking with shared detections and HOTA reporting, and user studies with visually impaired participants under approved ethical protocols. Consequently, the paper should be read as a technically grounded step toward inclusive, auditable, and privacy-conscious smart mobility where city-scale deployment represents a natural future work.

References

  1. M. H. Abidi, A. Noor Siddiquee, H. Alkhalefah, et V. Srivastava, « A comprehensive review of navigation systems for visually impaired individuals », ;Heliyon;, vol. 10, n<sup>o</sup> 11, p. e31825, mai 2024, doi: 10.1016/j.heliyon.2024.e31825. vol. 10. [CrossRef] [PubMed]
  2. Santos, D. P. D.; Suzuki, A. H. G.; Medola, F. O.; Vaezipour, A. A Systematic Review of Wearable Devices for Orientation and Mobility of Adults With Visual Impairment and Blindness. IEEE Access 2021, vol. 9, 162306-162324. [Google Scholar] [CrossRef]
  3. Advancements in Smart Wearable Mobility Aids for Visual Impairments: A Bibliometric Narrative Review ». Consulté le: 29 octobre 2025. [En ligne]. Disponible sur: https://www.mdpi.com/1424-8220/24/24/7986?
  4. Waqar, A.; Alshehri, A. H.; Alanazi, F.; Alotaibi, S.; Almujibah, H. R. Evaluation of challenges to the adoption of intelligent transportation system for urban smart mobility. Res. Transp. Bus. Manag. 2023, vol. 51, 101060. [Google Scholar] [CrossRef]
  5. Rahman, M. A.; Mammeri, A.; Metari, S. « Intelligent Infrastructure for Enhancing Vulnerable Road User Safety using Machine Vision Technologies ». Int. J. Intell. Transp. Syst. Res. 2025, vol. 23(no 2), 1179-1196. [Google Scholar] [CrossRef]
  6. K. Watanabe, M. Nishi, et T. Ito, « Placement Design of Various Roadside Sensors for Stochastic Position Prediction of Traffic Participants on Community Roads », ;Int. J. Intell. Transp. Syst. Res.;, déc. 2025, doi: 10.1007/s13177-025-00574-w.
  7. J. Santos ;et al.;, « A Requirement-based Methodology for Selecting Simulation Solutions in C-ITS Safety Applications », ;Int. J. Intell. Transp. Syst. Res.;, vol. 23, n<sup>o</sup> 2, p. 1315-1330, août 2025, doi: 10.1007/s13177-025-00516-6. vol. 23, no 2. [CrossRef]
  8. L. Fei et B. Han, « Multi-Object Multi-Camera Tracking Based on Deep Learning for Intelligent Transportation: A Review », ;Sensors;, vol. 23, n<sup>o</sup> 8, p. 3852, avr. 2023, doi: 10.3390/s23083852.
  9. T. Meinhardt, A. Kirillov, L. Leal-Taixe, et C. Feichtenhofer, « TrackFormer: Multi-Object Tracking with Transformers », in ;2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR);, New Orleans, LA, USA: IEEE, juin 2022, p. 8834-8844. doi: 10.1109/CVPR52688.2022.00864. IEEE.
  10. G. Braso et L. Leal-Taixe, « Learning a Neural Solver for Multiple Object Tracking », in ;2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR);, Seattle, WA, USA: IEEE, juin 2020, p. 6246-6256. doi: 10.1109/CVPR42600.2020.00628. IEEE.
  11. M. Farenzena, L. Bazzani, A. Perina, V. Murino, et M. Cristani, « Person re-identification by symmetry-driven accumulation of local features », in ;2010 IEEE Computer Society Conference on Computer Vision and Pattern Recognition;, San Francisco, CA, USA: IEEE, juin 2010, p. 2360-2367. doi: 10.1109/cvpr.2010.5539926. San Francisco, CA, USA, juin 2010. [CrossRef]
  12. Kotseruba, A. Rasouli, et J. K. Tsotsos, « Joint Attention in Autonomous Driving (JAAD) », 22 avril 2020, ;arXiv;: arXiv:1609.04741. doi: 10.48550/arXiv.1609.04741.
  13. T. Chavdarova ;et al.;, « WILDTRACK: A Multi-camera HD Dataset for Dense Unscripted Pedestrian Detection », in ;2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition;, Salt Lake City, UT: IEEE, juin 2018, p. 5030-5039. doi: 10.1109/CVPR.2018.00528.
  14. R. Alturki, A. Hilton, et J.-Y. Guillemaut, « Attention-Aware Multi-View Pedestrian Tracking », 3 avril 2025, ;arXiv;: arXiv:2504.03047. doi: 10.48550/arXiv.2504.03047.
  15. J. Luiten ;et al.;, « HOTA: A Higher Order Metric for Evaluating Multi-object Tracking », ;Int. J. Comput. Vis.;, vol. 129, n<sup>o</sup> 2, p. 548-578, févr. 2021, doi: 10.1007/s11263-020-01375-2.
  16. R. Hishida, M. Iryo, R. Alhamdani, K. Mano, et Y. Yamaguchi, « Development of Methods for High-Density Crowd Measurement and Tracking in Railway Station Concourses », ;Int. J. Intell. Transp. Syst. Res.;, avr. 2026, doi: 10.1007/s13177-025-00610-9. [CrossRef]
  17. Habibovic ;et al.;, « Communicating Intent of Automated Vehicles to Pedestrians », ;Front. Psychol.;, vol. 9, août 2018, doi: 10.3389/fpsyg.2018.01336. [CrossRef] [PubMed]
  18. M. D. Alfikri et R. Kaliski, « Real-Time Pedestrian Detection on IoT Edge Devices: A Lightweight Deep Learning Approach », 24 septembre 2024, ;arXiv;: arXiv:2409.15740. doi: 10.48550/arXiv.2409.15740.
  19. M. Raif, H. Chaibi, N. Quadar, A. E. Rharras, et R. Saadane, « Edge-IoT and MLLMs for Education and Scene Understanding: Assisting Vision and Hearing-Impaired Individuals », ;IEEE Internet Things Mag.;, vol. 8, n<sup>o</sup> 4, p. 30-36, juill. 2025, doi: 10.1109/IOTM.001.2400202.
  20. « ISO/IEC 42001:2023 », ISO. Consulté le: 29 octobre 2025. [En ligne]. Disponible sur: https://www.iso.org/standard/42001.
  21. « 24748-7000-2022 - IEEE/ISO/IEC International Standard--Systems and software engineering--Life cycle management--Part 7000: Standard model process for addressing ethical concerns during system design ». Consulté le: 11 avril 2026. [En ligne]. Disponible sur: https://ieeexplore.ieee.org/document/9967807.
  22. E. Tabassi, « Artificial Intelligence Risk Management Framework (AI RMF 1.0) », National Institute of Standards and Technology (U.S.), Gaithersburg, MD, NIST AI 100-1, janv. 2023. doi: 10.6028/NIST.AI.100-1. National Institute of Standards and Technology (U.S.): Gaithersburg, MD.
  23. F. Zeng, B. Dong, Y. Zhang, T. Wang, X. Zhang, et Y. Wei, « MOTR: End-to-End Multiple-Object Tracking with Transformer », 19 juillet 2022, ;arXiv;: arXiv:2105.03247. doi: 10.48550/arXiv.2105.03247.
  24. W. Li ;et al.;, « Multiagent Consensus Tracking Control Over Asynchronous Cooperation–Competition Networks », ;IEEE Trans. Cybern.;, vol. 55, n<sup>o</sup> 9, p. 4347-4360, sept. 2025, doi: 10.1109/TCYB.2025.3583387. no 9.
  25. L. Shi, S. Yan, et W. Li, « Consensus and Products of Substochastic Matrices: Convergence Rate With Communication Delays », ;IEEE Trans. Syst. Man Cybern. Syst.;, vol. 55, n<sup>o</sup> 7, p. 4752-4761, juill. 2025, doi: 10.1109/TSMC.2025.3559667.
  26. « IEEE Standard Model Process for Addressing Ethical Concerns during System Design », ;IEEE Std 7000-2021;, p. 1-82, sept. 2021, doi: 10.1109/IEEESTD.2021.9536679.
Figure 1. Mutual-recognition loop for protected-pedestrian detection, vehicle-yield recognition, accessible feedback, and audit log.
Figure 1. Mutual-recognition loop for protected-pedestrian detection, vehicle-yield recognition, accessible feedback, and audit log.
Preprints 228553 g001
Figure 2. Edge-Fog-Cloud ITS architecture for protected-pedestrian continuity, yield arbitration, and grounded guidance.
Figure 2. Edge-Fog-Cloud ITS architecture for protected-pedestrian continuity, yield arbitration, and grounded guidance.
Preprints 228553 g002
Figure 3. Mutual-recognition governance loop.
Figure 3. Mutual-recognition governance loop.
Preprints 228553 g003
Table 1. Failure scenarios and conservative system responses.
Table 1. Failure scenarios and conservative system responses.
Failure scenario Trigger System response Logged evidence
Pedestrian occlusion No confident match for N frames or ai(t) > Tocc Maintain vehicle yield directive; downgrade wearable guidance to Stop/Wait Track ID, camera ID, occlusion start/end time
Cross-camera handover failure Descriptor or geometry match below threshold Suspend permissive guidance until re-identification succeeds Candidate IDs, confidence scores, handover window
Identity mismatch Appearance ambiguity or conflicting tracks Block permissive instruction and request re-verification Ambiguous tracks, rejected match, confidence threshold
Network congestion Stale Edge-Fog, Fog-vehicle, or Fog-wearable token Invalidate stale token; return restrictive guidance Token timestamp, TTL, missing link
Vehicle acknowledgment loss Yield directive not acknowledged within Tack Prohibit crossing instruction; keep vehicle obligation active Directive ID, vehicle ID, missing ACK
MLLM narration failure Hallucinated, delayed, or ungrounded message Use a verified template such as "Stop. Please wait." Prompt state, template ID, blocked output
Table 2. Evidence sources and the technical role of each architectural component.
Table 2. Evidence sources and the technical role of each architectural component.
Component Reused prior capability Integration role in the proposed architecture Current evidence / validation status
Edge perception Lightweight pedestrian detection and compact descriptors Publishes privacy-preserving presence tokens instead of raw video in ordinary operation JAAD latency and intent pipeline
Fog MC-MOT Multi-camera tracking, graph association, transformer continuity, re-ID Maintains protected-pedestrian continuity across occlusion and camera handover when confidence permits WildTrack continuity protocol
Delay-aware arbitration Time-stamped V2X and event messages Binds yield obligations and block permissive guidance when state is stale or uncertain Formal rule in Eq. (1); failure tests
Accessible V2P feedback eHMI, audio, haptic, and wearable signaling Converts vehicle intent into pedestrian-specific confirmation after Fog validation JAAD guidance-actionability evaluation
Cloud governance AI management, ethics-by-design, risk management Distributes policy, retention limits, audit requirements, and model updates Standard-aligned design specification
Table 3. Experimental protocol for the two validation branches.
Table 3. Experimental protocol for the two validation branches.
Item Secondary JAAD feasibility branch WildTrack continuity branch
Dataset role Monocular crossing-intention, warning delay, and grounded-guidance feasibility check Multi-camera pedestrian identity continuity under occlusion and handover
Split pedestrian track split into train/validation/test partitions; no pedestrian track is shared between partitions Official calibration and annotations; chronological train/test protocol with reduced-range ablation followed by full-range validation
Input RGB frames, pedestrian boxes, ego/scene context, crossing labels, and short text prompts Synchronized camera observations, calibration, detections, Compact Feature Descriptors, and ground-plane associations
Hardware Edge inference on embedded GPU-class hardware; Fog inference on a local GPU workstation; narration generated on a Fog GPU when available or Cloud GPU otherwise Fog workstation with GPU acceleration for association, descriptor matching, and TrackEval-style metric computation
Main hyperparameters Crossing-decision accuracy, Fog inference time, restrictive latency, permissive latency, narration actionability Occlusion timeout Tocc = 1.5 s; cosine descriptor threshold 0.62; ground-plane gate 1.0 m; Hungarian/graph association for candidate matching
Metrics Crossing-decision accuracy, Fog inference time, restrictive latency, permissive latency, narration actionability MOTA, IDF1, HOTA, ID switches, and handover failure rate
Table 4. Secondary feasibility check of the JAAD assistive branch. Values are mean ± standard deviation in evaluation folds or repeated timing runs.
Table 4. Secondary feasibility check of the JAAD assistive branch. Values are mean ± standard deviation in evaluation folds or repeated timing runs.
Metric Result Interpretation for mutual recognition
Crossing-decision accuracy 87.9% ± 0.4% Supports the feasibility of crossing-related state estimation before guidance generation
Fog-tier inference time 43 ms ± 2 ms Indicates that the guidance branch can leave the timing budget for V2P acknowledgment and wearable feedback
End-to-end restrictive alert latency 0.81 s ± 0.09 s Supports the feasibility of fast restrictive guidance under uncertainty
End-to-end permissive guidance latency 1.63 s ± 0.22 s Indicates that validated permissive guidance can remain within the intended human-response envelope
Narration actionability 91.2% ± 1.3% Shows that grounded guidance is usually specific enough for action after state validation
Table 5. Wildtrack continuity results.
Table 5. Wildtrack continuity results.
Methods Setting MOTA IDF1 HOTA Role in this paper
KSP-DO WildTrack 69.6 73.2 NR Baseline of the Lower-bound
EarlyBird WildTrack 89.5 92.3 NR Early-fusion baseline
TrackTacular WildTrack 91.8 95.3 NR Strong multi-view baseline
Attention-aware tracker [14] WildTrack 92.7 96.1 NR Recent review baseline
Our Fog continuity protocol WildTrack protocol 88.8 91.7 65.5 Reproducibility branch (torch 2.11, stock Colab); HOTA value under the reproduced protocol; MCCG validated (Section 6.4)
Table 6. MCCG Safety–Efficiency Trade-off.
Table 6. MCCG Safety–Efficiency Trade-off.
M Permissive-eligible Overhead Switch frames blocked Safety assessment
1 100.0% 0.0% 0/4 (0%) No safety gain
2 100.0% 0.0% 0/4 (0%) No safety gain
3 100.0% 0.0% 0/4 (0%) No safety gain
4 ← 95.0% 5.0% 3/4 (75%) Recommended operating point
5 94.2% 5.8% 3/4 (75%) Dominated by M=4
6 62.4% 37.6% 4/4 (100%) Too conservative for dense scenes
7 2.9% 97.1% 4/4 (100%) Operationally unusable
MCCG sweep in WildTrack for N = 7 cameras. The recommended operating point is M = 4. Permissive-eligible is the fraction of tracked observations that meet the M-camera support condition. Overhead is the complement of permissive-eligible. The blocked switch frames give the number of ID-switch-containing frames where the MCCG would withhold permissive guidance.
Table 7. Operational contract of the mutual-recognition loop.
Table 7. Operational contract of the mutual-recognition loop.
Segment Operational target Message artifact Safety contract
Edge-Fog presence Low-latency detection and compact event publication Presence token with box, timestamp, camera ID, descriptor Low confidence cannot authorize permissive guidance
Fog tracking and arbitration Maintain protected-pedestrian continuity and assign yield obligation Protected track state and yield directive token Timeout, ambiguity, or handover failure triggers Stop/Wait
Fog-vehicle directive Dispatch time-bounded yield instruction Directive token with pedestrian ID, zone, TTL, timestamp Missing ACK blocks permissive guidance
Vehicle-Fog ACK Confirm vehicle state within the acknowledgment window ACK token with directive ID, status, timestamp Explicit ACK required before "you may cross"
Fog/Cloud-wearable guidance Restrictive alerts under 1 s; permissive guidance in 1-2 s Grounded utterance or haptic code Any stale state, narration failure, or uncertainty triggers Stop/Wait
Governance log Preserve auditability without raw-video retention Decision/event token with participants, timestamps and outcome Records duties, ACKs, guidance, and fallback events
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.