Automated concrete crack detection increasingly employs multimodal sensing, although many existing systems assume that all sensors remain continuously available, reliable, and suitable for fusion. This study presents Reliability-Aware Selective Multimodal Fusion, RASMF, an image-first framework in which acoustic sensing is activated only when visual confidence is insufficient or the image is unavailable. The framework combines calibrated image, acoustic, and joint experts with an interpretable fusion model informed by predictive uncertainty, cross-modal disagreement, modality-quality descriptors, and availability indicators. An acoustic out-of-distribution safeguard rejects acquired signals whose characteristics depart from the development distribution. RASMF was evaluated on 4094 paired observations using five duplicate-aware outer folds and three random seeds under clean, degraded, and missing-modality conditions across multiple acquisition budgets. At the nominal 10% budget, the realised acoustic acquisition rate was 21.16%. Relative to the predefined missing-aware image-first baseline, RASMF reduced mean error from 0.9934% to 0.6754%, increased sensitivity from 96.6761% to 97.8209%, and reduced the Brier score from 0.0112 to 0.0071. These improvements were supported by paired statistical tests and two-way bootstrap intervals. Lower error was obtained in 12 of 16 degradation and modality-loss conditions, with the largest gains under substantial visual degradation. Moderate acoustic acquisition improved performance and calibration, whereas near-continuous acquisition increased operational cost without further benefit. RASMF therefore improved the balance between predictive reliability and conditional sensing relative to the designated baseline, although universal superiority over simpler fusion alternatives was not established.