Submitted:
11 September 2026
Posted:
14 September 2026
You are already at the latest version
Abstract
Visible–infrared detection improves perception in degraded illumination by combining visible texture with thermal contrast. Small objects remain difficult because their discriminative responses are weak, spatially sparse, and sensitive to cross-modal displacement. We present FGCFF-Net, a lightweight frequency-guided fusion detector for visible–infrared small object detection. The framework casts fusion as progressive evidence selection. A lightweight channel pruning module suppresses redundant dual-stream responses; frequency-guided alignment estimates residual spectral compensation; a cross-modal alignment and small-target framework predicts bidirectional spectral weights and deformable offsets; and a frequency-aware cross-modal aggregator conditions channel-spatial attention on frequency discrepancy. The training objective further couples dense detection loss with spectral consistency, offset smoothness, and small-target emphasis. On FLIR, M3FD, LLVIP, and KAIST, FGCFF-Netachieved 92.4% mAP50 on FLIR, 90.7% mAP50 on M3FD, 98.5% AP50 on LLVIP, and 5.72% log-average miss rate on KAIST, with 14.8M parameters and 68 FPS. These results indicate that frequency-aware alignment provides an efficient route for robust multispectral small-object perception.
Keywords:
visible–infrared detection
; small object detection
; cross-modal fusion
; frequency-aware attention
; lightweight neural networks
; multispectral perception
Copyright: This open access article is published under a Creative Commons CC BY 4.0 license, which permit the free download, distribution, and reuse, provided that the author and preprint are cited in any reuse.