Submitted:
03 October 2026
Posted:
07 October 2026
You are already at the latest version
Abstract
Screening site photographs against regulations requires deciding which provisions apply; fixed top-k retrieval ranks them by textual relevance rather than visible applicability, so widening it buys recall only by exposing inapplicable provisions. We frame applicability as set prediction under asymmetric miss/exposure cost and propose Risk-Controlled Applicability Set Retrieval (RCASR), which reweights a retrieval prior by evidence that each provision’s compiled subject and state conditions hold, from posteriors over 131 rule-agnostic visual atoms, and selects a set by width threshold or conformal risk control. On 346 photographs from 48 sites, with site-grouped cross-fitting, RCASR raised recorded-violation recall at a matched width of three from 0.515 (BM25) to 0.709, beat adaptive and ternary screens, and matched an agent retriever using two single-turn passes. A pre-specified width-two arm kept BM25’s width-three recall while cutting judged pairs by a third and false alarms per image from 1.28 to 1.00. Frozen before public scoring, RCASR raised ConstructionSite-10k violation detection for harness from 0.08 to 0.80 and for protective equipment from 0.43 to 0.67. Conformal risk control held α=0.10 at about ten candidates per image. The whole chain runs on a quadruped patrol robot’s onboard computer and has processed 2564 robot-recorded frames from construction sites.

Keywords:
construction safety
; regulation compliance
; vision–language models
; retrieval-augmented generation
; set-valued prediction
; conformal risk control
; site inspection
; false alarms
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.