Arrow Research search

Author name cluster

Song Wu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

14 papers
2 author rows

Possible papers

14

AAAI Conference 2026 Conference Paper

Views Attention Fusion of Granular-ball Fuzzy Representations Split for Improved Multi-view Clustering

  • Shuaiyu Liu
  • Song Wu
  • Jie Xu
  • Yazhou Ren
  • Yang Yang
  • Xiaorong Pu
  • Guoying Wang

Multi-View Clustering (MVC) is a pivotal multi-view learning paradigm widely adopted across various fields. Despite recent advances, existing methods primarily focus on enhancing the performance of fused multi-view representation, often neglecting the issue of Representation Degradation (RD) arising from discrepancies in the intrinsic quality of different views. To address the limitations, we propose a novel Granular-ball Fuzzy Split and Attention Fusion (GFSAF) learning, which leverages the nature of granular-ball to extract mutual and complementary representation separately. Meanwhile, the proposed method introduces an attention variant for fused representations to mitigate the RD issue. GFSAF mainly consists of two training stages: Split-Extract Stage and Views-Fusion Stage. Specifically, we design a novel Granular-ball Fuzzy Contrastive Learning to extract mutual representation, and introduce Noise Stripping Loss to reduce the influence of noise for complementary representation. Then, a novel multi-head Cross Views Attention is proposed to employ attention mechanism from multi-view perspectives for comprehensive fused representations. Experimental results on eight databases demonstrate that our GFSAF achieves superior performance compared to several state-of-the-art MVC methods.

JBHI Journal 2025 Journal Article

A Feature-Aware Approach to Acupoint Compatibility Prediction Using Residual Graph Attention Networks and Matrix Factorization

  • Ruiling Li
  • Ying Pan
  • Song Wu
  • Li Ma
  • Limei Peng

Compatibility among acupoints is a fundamental principle in acupuncture treatment within traditional Chinese medicine, playing a vital role in enhancing the effectiveness and scope of therapeutic interventions. With the increasing availability of acupuncture-related data, link prediction offers a data-driven approach that facilitates the evidence-based exploration and validation of acupoint compatibilities. However, existing link prediction methods often focus on mapping acupoints and their compatibility relationships into lower-dimensional spaces. These approaches can overlook essential acupoint features and make the predictions susceptible to noise interference. To address these challenges, we propose a novel acupoint compatibility prediction model based on a Feature-Aware Residual Graph Attention Network and Matrix Factorization (FRGATMF). Our model introduces a feature-aware connectivity fusion strategy that integrates acupoint attributes with structural information to enrich acupoint representations. Following this, a deep non-negative matrix factorization approach is employed to construct a denoised feature matrix. This matrix is processed through a residual graph attention network to derive comprehensive and effective node embeddings, which are crucial for accurate link prediction. Experimental results on the acupuncture dataset, along with three public datasets, demonstrate that FRGATMF significantly outperforms seven existing comparison models across various evaluation metrics. Additionally, link prediction can identify previously unconsidered or undocumented acupoint combinations that may offer better therapeutic results, thus expanding the range of treatment options and highlighting its potential in improving the prediction of acupoint compatibility relationships.

NeurIPS Conference 2025 Conference Paper

FreeControl: Efficient, Training-Free Structural Control via One-Step Attention Extraction

  • Jiang Lin
  • Xinyu Chen
  • Song Wu
  • Zhiqiu Zhang
  • Jizhi Zhang
  • Ye Wang
  • Qiang Tang
  • Qian Wang

Controlling the spatial and semantic structure of diffusion-generated images remains a challenge. Existing methods like ControlNet rely on handcrafted condition maps and retraining, limiting flexibility and generalization. Inversion-based approaches offer stronger alignment but incur high inference cost due to dual-path denoising. We present \textbf{FreeControl}, a training-free framework for semantic structural control in diffusion models. Unlike prior methods that extract attention across multiple timesteps, FreeControl performs \textit{one-step attention extraction} from a single, optimally chosen timestep and reuses it throughout denoising. This enables efficient structural guidance without inversion or retraining. To further improve quality and stability, we introduce \textit{Latent-Condition Decoupling (LCD)}: a principled separation of the timestep condition and the noised latent used in attention extraction. LCD provides finer control over attention quality and eliminates structural artifacts. FreeControl also supports compositional control via reference images assembled from multiple sources, enabling intuitive scene layout design and stronger prompt alignment. FreeControl introduces a new paradigm for test-time control—enabling structurally and semantically aligned, visually coherent generation directly from raw images, with the flexibility for intuitive compositional design and compatibility with modern diffusion models at ~5\% additional cost.

IROS Conference 2025 Conference Paper

MSPA-LIO: LiDAR-Inertial Odometry with Multi-Scale Plane Adjustment

  • Su Yan
  • Shuying Zhao
  • Yunzhou Zhang
  • Hengwang Ding
  • Wu Li
  • Sizhan Wang
  • Song Wu

Most current LiDAR-based odometry methods use point-to-local plane registration to constrain poses, ignoring the explicit plane structure in the environment. Due to noise interference and uneven distribution of point cloud, local planes are prone to tilt, resulting in registration errors. Therefore, we propose MSPA-LIO, a LiDAR-Inertial odometry with multi-scale plane adjustment, which uses geometric constraints and plane adjustment at both local voxel plane scale and large plane scale to improve odometry accuracy and enhance map consistency. In order to make full use of the planar structure in the environment, we propose an explicit large plane extraction method based on the voxel-based. We use large planes to correct the direction of the associated voxel planes, thereby overcoming the misregistration problem caused by local plane tilt. To further improve the odometry accuracy, we perform plane adjustments at the voxel plane scale and the large plane scale to make the pose and map more consistent. Experiments conducted on the VECtor Dataset and the Newer College Dataset demonstrate that our proposed algorithm outperforms four state-of-the-art algorithms.

NeurIPS Conference 2025 Conference Paper

Sim-LLM: Optimizing LLM Inference at the Edge through Inter-Task KV Reuse

  • Ruikun Luo
  • Changwei Gu
  • Qiang He
  • Feifei Chen
  • Song Wu
  • Hai Jin
  • Yun Yang

KV cache technology, by storing key-value pairs, helps reduce the computational overhead incurred by large language models (LLMs). It facilitates their deployment on resource-constrained edge computing nodes like edge servers. However, as the complexity and size of tasks increase, KV cache usage leads to substantial GPU memory consumption. Existing research has focused on mitigating KV cache memory usage through sequence length reduction, task-specific compression, and dynamic eviction policies. However, these methods are computationally expensive for resource-constrained edge computing nodes. To tackle this challenge, this paper presents Sim-LLM, a novel inference optimization mechanism that leverages task similarity to reduce KV cache memory consumption for LLMs. By caching KVs from processed tasks and reusing them for subsequent similar tasks during inference, Sim-LLM significantly reduces memory consumption while boosting system throughput and increasing maximum batch size, all with minimal accuracy degradation. Evaluated on both A40 and A100 GPUs, Sim-LLM achieves a system throughput improvement of up to 39. 40\% and a memory reduction of up to 34. 65%, compared to state-of-the-art approaches. Our source code is available at https: //github. com/CGCL-codes/SimLLM.

EAAI Journal 2024 Journal Article

Core-attributes enhanced generative adversarial networks for robust image enhancement

  • Shan Liu
  • Guoqiang Xiao
  • Michael S. Lew
  • Xinbo Gao
  • Song Wu

Automated image enhancement algorithms have a profound impact on human life today. To solve the problems of luminance, lack of detail information, and overall color tone bias of images taken by mobile devices, a novel framework of core-attributes enhanced generative adversarial network (CAE-GAN) is designed to improve these core attributes of enhanced images. The generator in CAE-GAN mainly consists of a luminance correction encoder (LCE) and a high-frequency supplementary decoder (HFSD). To target the adaptive luminance improvement for each location, the encoder based on LCE is designed by combining the extracted prior knowledge of luminance. Meanwhile, a decoder based on HFSD is proposed to fill in missing edge details during the image reconstruction process. In addition, a multi-scale statistical characteristics distinction branch (MSCDB) is proposed to correct the overall tone. Moreover, an upgrade adversarial loss function is designed to focus on the discrimination of both multi-scale and multi-perspective. The generator and discriminator are iteratively trained under the constraints of the total loss function, resulting in the generator that automatically improves the visualization of the images. Extensive experiments have shown that our CAE-GAN is capable of achieving excellent results in several evaluation metrics and subjective results. The source code of the proposed CAE-GAN is available at https: //github. com/SWU-CS-MediaLab/CAE-GAN.

IJCAI Conference 2024 Conference Paper

Cross-View Contrastive Fusion for Enhanced Molecular Property Prediction

  • Yan Zheng
  • Song Wu
  • Junyu Lin
  • Yazhou Ren
  • Jing He
  • Xiaorong Pu
  • Lifang He

Machine learning based molecular property prediction has been a hot topic in the field of computer aided drug discovery (CADD). However, current MPP methods face two prominent challenges: 1) single-view MPP methods do not sufficiently exploit the complementary information of molecular data across multiple views, generally producing suboptimal performance, and 2) most existing multi-view MPP methods ignore the disparities in data quality among different views, inadvertently introducing the risk of models being overshadowed by inferior views. To address the above challenges, we introduce a novel cross-view contrastive fusion for enhanced molecular property prediction method (MolFuse). First, we extract intricate molecular semantics and structures from both sequence and graph views to leverage the complementarity of multi-view data. Then, MolFuse employs two distinct graphs, the atomic graph and chemical bond graph, to enhance the representation of the molecular graph, allow us to integrate both the fundamental backbone attributes and the nuanced shape characteristics. Notably, we incorporate a dual learning mechanism to refine the initial feature representations, and global features are obtained by maximizing the coherence among diverse view-specific molecular representations for the downstream task. The overall learning processes are combined into a unified optimization problem for iterative training. Experiments on multiple benchmark datasets demonstrate the superiority of our MolFuse.

EAAI Journal 2024 Journal Article

Deep cross-modal hashing with multi-task latent space learning

  • Song Wu
  • Xiang Yuan
  • Guoqiang Xiao
  • Michael S. Lew
  • Xinbo Gao

Cross-modal Hashing (CMH) retrieval aims to mutually search data from heterogeneous modalities by projecting original modality data into a common hamming space, with the significant advantages of low storage and computing costs. However, CMH remains challenging for multi-label cross-modal datasets. Firstly, preserving content similarity would inevitably be deficient under the representation of short-length binary codes. Secondly, different semantics are treated independently, whereas their co-occurrences are neglected, reducing retrieval quality. Thirdly, the commonly used metric learning objective is ineffective in capturing similarity information at a fine-grained level, leading to the imprecise preservation of such information. Therefore, we propose a Deep Cross-Modal Hashing with Multi-Task Latent Space Learning (DMLSH) framework to tackle these bottlenecks. For a more thorough excavation of distinctive features with diverse characteristics underneath heterogeneous data, our DMLSH is designed to preserve three different types of knowledge. The first is the semantic relevance and co-occurrence with the integration of the attention module and the Long Short-Term Memory (LSTM) layer; The second is the highly precise pairwise correlation considering the quantification of semantic similarity with self-paced optimization; The last is the pairwise similarity information discovered by a self-supervised semantic network from a perspective of probabilistic knowledge transfer. Abundant knowledge from the latent spaces is seamlessly refined and fused into a common Hamming space by a hashing attention mechanism, facilitating the discrimination of hash codes and the elimination of modalities’ heterogeneity. Exhaustive experiments demonstrate the state-of-the-art performance of our proposed DMLSH on four mainstream cross-modal retrieval benchmarks.

NeurIPS Conference 2024 Conference Paper

E-Motion: Future Motion Simulation via Event Sequence Diffusion

  • Song Wu
  • Zhiyu Zhu
  • Junhui Hou
  • Guangming Shi
  • Jinjian Wu

Forecasting a typical object's future motion is a critical task for interpreting and interacting with dynamic environments in computer vision. Event-based sensors, which could capture changes in the scene with exceptional temporal granularity, may potentially offer a unique opportunity to predict future motion with a level of detail and precision previously unachievable. Inspired by that, we propose to integrate the strong learning capacity of the video diffusion model with the rich motion information of an event camera as a motion simulation framework. Specifically, we initially employ pre-trained stable video diffusion models to adapt the event sequence dataset. This process facilitates the transfer of extensive knowledge from RGB videos to an event-centric domain. Moreover, we introduce an alignment mechanism that utilizes reinforcement learning techniques to enhance the reverse generation trajectory of the diffusion model, ensuring improved performance and accuracy. Through extensive testing and validation, we demonstrate the effectiveness of our method in various complex scenarios, showcasing its potential to revolutionize motion flow prediction in computer vision applications such as autonomous vehicle guidance, robotic navigation, and interactive media. Our findings suggest a promising direction for future research in enhancing the interpretative power and predictive accuracy of computer vision systems. The source code ispublicly available at https: //github. com/p4r4mount/E-Motion.

JBHI Journal 2024 Journal Article

ECC-PolypDet: Enhanced CenterNet With Contrastive Learning for Automatic Polyp Detection

  • Yuncheng Jiang
  • Zixun Zhang
  • Yiwen Hu
  • Guanbin Li
  • Xiang Wan
  • Song Wu
  • Shuguang Cui
  • Silin Huang

Accurate polyp detection is critical for early colorectal cancer diagnosis. Although remarkable progress has been achieved in recent years, the complex colon environment and concealed polyps with unclear boundaries still pose severe challenges in this area. Existing methods either involve computationally expensive context aggregation or lack prior modeling of polyps, resulting in poor performance in challenging cases. In this paper, we propose the Enhanced CenterNet with Contrastive Learning (ECC-PolypDet), a two-stage training & end-to-end inference framework that leverages images and bounding box annotations to train a general model and fine-tune it based on the inference score to obtain a final robust model. Specifically, we conduct Box-assisted Contrastive Learning (BCL) during training to minimize the intra-class difference and maximize the inter-class difference between foreground polyps and backgrounds, enabling our model to capture concealed polyps. Moreover, to enhance the recognition of small polyps, we design the Semantic Flow-guided Feature Pyramid Network (SFFPN) to aggregate multi-scale features and the Heatmap Propagation (HP) module to boost the model's attention on polyp targets. In the fine-tuning stage, we introduce the IoU-guided Sample Re-weighting (ISR) mechanism to prioritize hard samples by adaptively adjusting the loss weight for each sample during fine-tuning. Extensive experiments on six large-scale colonoscopy datasets demonstrate the superiority of our model compared with previous state-of-the-art detectors.

IROS Conference 2024 Conference Paper

LA-LIO: Robust Localizability-Aware LiDAR-Inertial Odometry for Challenging Scenes

  • Junjie Huang
  • Yunzhou Zhang
  • Qingdong Xu
  • Song Wu
  • Jun Liu 0087
  • Guiyuan Wang
  • Wei Liu 0022

Modern robotic systems are increasingly deployed in complex and diverse environments, and reliable localization under challenging conditions becomes crucial for the safe and efficient operation of these systems. The odometry based on LiDAR is prone to system collapse caused by computational divergence under conditions of aggressive motion and information deficiency in spatial geometry. To enhance the robustness of systems in challenging scenes, this work proposes LA-LIO, robust localizability-aware LiDAR inertial odometry. It mainly consists of three parts. Firstly, this paper presents a LiDAR degeneration detection method that enables stable degeneration assessment. Secondly, a method for segmenting LiDAR point clouds is proposed to alleviate the issue of excessive distortion in point clouds under aggressive motion scenes. The last is an Errors State Kalman Filter (ESKF) method with adaptive weights to utilize the existing spatial information as much as possible to improve the stability of the system in degenerated scenarios. The proposed method is evaluated and compared in multiple experiments, demonstrating the performance and reliability improvements of this approach in challenging environments.

EAAI Journal 2024 Journal Article

Modality Blur and Batch Alignment Learning for Twin Noisy Labels-based Visible–infrared Person Re-identification

  • Song Wu
  • Shihao Shan
  • Guoqiang Xiao
  • Michael S. Lew
  • Xinbo Gao

The issue of Twin Noisy Labels, known as Noisy Annotations and Noisy Correspondences, increases the challenge in the engineering application of Visible–Infrared Person Re-identification (VI-ReID). This paper proposes an novel Modality Blur and Batch Alignment (MBBA) framework to address this issue in the practical application of VI-ReID. The MBBA consists of the Label Confidence Learning (LCL) module, the Modality Blur Learning (MBL) module, and the Batch Alignment Learning (BAL) module. The LCL utilizes the memorization effect of deep neural networks to estimate the confidence level of identity labels and rectify the noisy annotations and the noisy correspondences. The MBL uses the noiseless modality labels and the center loss to blur the boundary among modality data to make the discrimination of the learned latent feature space robust. Based on the designed alignment loss, the BAL employs the self-attention mechanism to align the significant prediction distributions among the cross-modal sample pairs from a batch-size perspective. The mean Average Precision (mAP) is improved by 4. 06% and 5. 34% compared to state-of-the-art methods on the 20% noisy RegDB dataset. And our MBBA exhibits contemporary state-of-the-art performance on Twin Noisy Labels based VI-ReID. Our MBBA is available at https: //github. com/SWU-CS-MediaLab/MBBA.

YNICL Journal 2023 Journal Article

Morphometric similarity network alterations in COVID-19 survivors correlate with behavioral features and transcriptional signatures

  • Jia Long
  • Jiao Li
  • Bing Xie
  • Zhuomin Jiao
  • Guoqiang Shen
  • Wei Liao
  • Xiaomin Song
  • Hongbo Le

OBJECTIVES: To explore the differences in the cortical morphometric similarity network (MSN) between COVID-19 survivors and healthy controls, and the correlation between these differences and behavioralfeatures and transcriptional signatures. MATERIALS & METHODS: 39 COVID-19 survivors and 39 age-, sex- and education years-matched healthy controls (HCs) were included. All participants underwent MRI and behavioral assessments (PCL-17, GAD-7, PHQ-9). MSN analysis was used to compute COVID-19 survivors vs. HCs differences across brain regions. Correlation analysis was used to determine the associations between regional MSN differences and behavioral assessments, and determine the spatial similarities between regional MSN differences and risk genes transcriptional activity. RESULTS: COVID-19 survivors exhibited decreased regional MSN in insula, precuneus, transverse temporal, entorhinal, para-hippocampal, rostral middle frontal and supramarginal cortices, and increased regional MSN in pars triangularis, lateral orbitofrontal, superior frontal, superior parietal, postcentral, and inferior temporal cortices. Regional MSN value of lateral orbitofrontal cortex was positively associated with GAD-7 and PHQ-9 scores, and rostral middle frontal was negatively related to PHQ-9 scores. The analysis of spatial similarities showed that seven risk genes (MFGE8, MOB2, NUP62, PMPCA, SDSL, TMEM178B, and ZBTB11) were related to regional MSN values. CONCLUSION: The MSN differences were associated with behavioral and transcriptional signatures, early psychological counseling or intervention may be required to COVID-19 survivors. Our study provided a new insight into understanding the altered coordination of structure in COVID-19 and may offer a new endophenotype to further investigate the brain substrate.

IJCAI Conference 2021 Conference Paper

PointLIE: Locally Invertible Embedding for Point Cloud Sampling and Recovery

  • Weibing Zhao
  • Xu Yan
  • Jiantao Gao
  • Ruimao Zhang
  • Jiayan Zhang
  • Zhen Li
  • Song Wu
  • Shuguang Cui

Point Cloud Sampling and Recovery (PCSR) is critical for massive real-time point cloud collection and processing since raw data usually requires large storage and computation. This paper addresses a fundamental problem in PCSR: How to downsample the dense point cloud with arbitrary scales while preserving the local topology of discarded points in a case-agnostic manner (i. e. , without additional storage for point relationships)? We propose a novel Locally Invertible Embedding (PointLIE) framework to unify the point cloud sampling and upsampling into one single framework through bi-directional learning. Specifically, PointLIE decouples the local geometric relationships between discarded points from the sampled points by progressively encoding the neighboring offsets to a latent variable. Once the latent variable is forced to obey a pre-defined distribution in the forward sampling path, the recovery can be achieved effectively through inverse operations. Taking the recover-pleasing sampled points and a latent embedding randomly drawn from the specified distribution as inputs, PointLIE can theoretically guarantee the fidelity of reconstruction and outperform state-of-the-arts quantitatively and qualitatively.

v2026.09.13