EAAI Journal 2026 Journal Article
A weak-alignment multi-sensor fusion method for object detection using visible, thermal infrared, and radar data
- Xiaoyuan Liang
- Ruicong Zhi
- Lu Jia
Weak alignment across heterogeneous sensors remains a significant challenge for multi-modal object detection in real-world scenarios. Differences in temporal synchronization, spatial resolution, and sensor viewpoints among visible light cameras, thermal infrared sensors, and millimeter-wave radar often degrade the performance of methods that depend on accurate cross-modal alignment. To address this problem, we propose the Weak-Alignment Multi-Modal Visible Light–Thermal Infrared–Radar Network (WA-M 3 RTR), a weakly aligned multi-modal object detection framework that integrates visible light, thermal infrared, and radar data without requiring explicit geometric alignment. From a methodological perspective, the framework introduces three key components: (i) a trainable infrared edge-enhancement mechanism to emphasize small-object structures in thermal infrared imagery; (ii) a graph-based radar encoding module, termed Radar–Azimuth Graph Embedding (RA-GraphEmbed), that models range–azimuth maps via graph attention networks; and (iii) a semantics-guided cross-attention strategy that injects global thermal infrared and radar cues into visible light feature maps across multiple scales. From an application perspective, the proposed framework is designed for ground-to-air object detection under weak alignment conditions. We introduce the Multi-Modal Aerial Unaligned Detection Tri-Modal (MMAUD-Tri) dataset, a tri-modal dataset for aerial object detection with simulated spatial misalignments, and further evaluate the approach on the real-world Pohang Canal Object Detection and Tracking Dataset in Maritime Environments (PoLaRIS), which exhibits natural cross-modal misalignment. Experimental results on both datasets demonstrate that WA-M 3 RTR consistently outperforms visible light–only and dual-modality baselines.