Arrow Research search

Author name cluster

Rong Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

36 papers
2 author rows

Possible papers

36

YNIMG Journal 2026 Journal Article

Biphasic adaptation of gBOLD-CSF coupling during sleep deprivation reflects compensatory enhancement and temporal disruption in glymphatic function

  • Dai Zhang
  • Rong Wang
  • Liqin Zhou
  • Ke Zhou
  • Zhentao Zuo
  • Guochen Sun

Sleep deprivation (SD) significantly impacts brain function, particularly through disruption of the glymphatic system, an essential mechanism for cerebral metabolic waste clearance dependent on cerebrospinal fluid (CSF) dynamics. Recent advances link CSF flow to global brain activity, measurable via global blood-oxygenation-level-dependent (gBOLD) signals. However, how gBOLD-CSF coupling changes during prolonged wakefulness remains unclear. Using resting-state functional magnetic resonance imaging (rs-fMRI), we investigated how 36-hour sleep deprivation affects gBOLD-CSF coupling in healthy participants. We observed a significant transient increase in gBOLD-CSF coupling strength as sleep deprivation progressed, peaking after approximately 30 h of wakefulness. Importantly, changes in coupling strength correlated quantitatively with heightened subjective sleep pressure but not with vigilance performance. Furthermore, SD induced a temporary phase shift in CSF signal timing relative to gBOLD, indicating disrupted temporal coordination. These results suggest that SD triggers both a transient enhancement and a temporal instability in neuro-fluid coupling, reflecting a biphasic modulation of brain-CSF coupling linked to glymphatic-related dynamics. Our findings reveal novel compensatory adjustments within the glymphatic system during prolonged wakefulness, advancing our understanding of the physiological underpinnings linking sleep loss, metabolic clearance, and brain function, with potential implications for cognitive health and neurodegenerative disease risk.

AAAI Conference 2026 Conference Paper

Directing Uncertainty-Aware Information Flow for Robust Diffusion Prediction

  • Weikang He
  • Yunpeng Xiao
  • Mengyang Huang
  • Xuemei Mou
  • Rong Wang
  • Qian Li

Information diffusion prediction is crucial for understanding social network dynamics, yet existing methods often neglect user participation uncertainty. This oversight typically stems from an implicit participation homogeneity assumption, which treats all observed interactions as equally reliable propagation signals, leading to fragile inferred topologies and uncertainty contamination. To address this, we propose SIEVE, a novel framework employing two synergistic strategies. First, robust node representations are learned via controllable uncertainty injection coupled with associated contrastive learning, mitigating topological fragility. Second, an uncertainty-aware directed graph aggregation mechanism is introduced, which dynamically constructs asymmetric aggregation topologies with adaptive weighting, thereby suppressing uncertainty contamination. Experiments on four public datasets demonstrate that SIEVE significantly outperforms state-of-the-art methods, offering valuable insights for designing robust information diffusion prediction models.

EAAI Journal 2026 Journal Article

Frequency-driven feature decoupling network for visible–infrared person re-identification

  • Xu Dong
  • Quange Tan
  • Rong Wang
  • Xiaowen Liu
  • Pingping Cao

Visible–infrared person re-identification (VI-ReID) remains a challenging task because of significant modality discrepancies and complex appearance variations between the visible and infrared domains. Existing approaches usually focus on holistic feature alignment or modality translation but often fail to separate the intrinsic differences between modality-specific features and identity-relevant features. To address this, we propose a novel frequency-driven feature decoupling network (FFD-Net) that uses frequency domain supervision to improve cross-modal matching. The design of FFD-Net is inspired by the observation that modality discrepancies are primarily encoded in low-frequency components, while appearance discrepancies are captured in high-frequency components. Thus, FFD-Net decomposes features into distinct frequency components using discrete wavelet transforms. From an information-theoretic perspective, this approach enables wavelet coefficients to preserve more mutual information with identity-related cues than alternative transforms, providing a solid theoretical foundation for frequency-driven feature decoupling. Additionally, FFD-Net incorporates a multi-scale aggregation module that enhances feature fusion. This module adaptively integrates complementary information from different frequency scales using attention-guided aggregation. To mitigate the loss of fine-grained details caused by spatial compression and frequency decomposition, an auxiliary learning strategy is introduced. This strategy reconstructs multi-level frequency features in a self-supervised manner, without introducing additional computational overhead during inference. Extensive experiments on the VI-ReID datasets demonstrate that FFD-Net significantly outperforms existing methods. From an artificial intelligence (AI) perspective, our work contributes by proposing a frequency-driven feature decoupling framework with a strong theoretical basis. From an engineering perspective, the method is applied to VI-ReID for intelligent surveillance, highlighting its practical value in real-world scenarios.

AAAI Conference 2026 Conference Paper

GCIB: Causal Intervention Guided Graph Information Bottleneck Framework

  • Hangyuan Du
  • Rong Wang
  • Lixin Cui
  • Gaoxia Jiang
  • Liang Bai
  • Wenjian Wang

Graph neural networks (GNNs) have demonstrated impressive performance in a broad spectrum of fields, but always suffer from the generalization problem when confronted with out-of-distribution (OOD) scenarios. Information bottleneck (IB) principle, which endeavors to learn the minimally sufficient representations for downstream tasks, has been shown to be a promising strategy in dealing with this problem. However, the IB-based methods do not inherently distinguish between causal and non-causal parts in the graph, leading to underperforming OOD generalization ability. In this paper, we develop the Graph Causal Information Bottleneck (GCIB) framework, a causal extension of the IB for graph data, which is capable of jointly compressing abundant information and capturing causal dependency from the input graph. Specifically, we endow graph IB with the ability of maintaining causal control by incorporating the underlying causal structure and introducing intervention operation. On this basis, we formulate the learning objective for GCIB and present its specific implementation. Graph representations learned by GCIB can effectively preserve causal information that fundamentally determines graph properties, resulting in outstanding OOD generalization ability. Extensive experiments on both synthetic and real-world datasets demonstrate the superiority of GCIB over state-of-the-art baselines.

AAAI Conference 2026 Conference Paper

Learning Intrinsic Hierarchy for Generalized Category Discovery

  • Yu Duan
  • Junzhi He
  • Zhanxuan Hu
  • Mengda Ji
  • Rong Wang
  • Quanxue Gao

Generalized Category Discovery (GCD) aims to classify unlabeled data by leveraging knowledge from labeled categories. While existing methods have achieved remarkable progress, they often treat images as flat feature sets, neglecting the intrinsic hierarchy: where key objects dominate meaning and backgrounds serve as context. For instance, in images of a dog either standing on grass or lying on a bed, the dog remains the central semantic element, whereas the background varies. Motivated by this, we propose LEArning Intrinsic Hierarchy (LEAH), a lightweight plug-and-play module designed to model hierarchical structure within images. LEAH consists of two components: a pruner that filters task-irrelevant tokens to extract key objects, and a constructor that embeds key objects and full images into hyperbolic space using adaptive entailment cones to capture compositional semantics. LEAH can be easily integrated into existing GCD frameworks with minimal modification. When applied to SimGCD, it achieves up to 13.2% accuracy improvement on fine-grained benchmarks, demonstrating its effectiveness in discovering subtle inter-class differences through hierarchical modeling.

JBHI Journal 2026 Journal Article

mRSubLoc: A Novel Multi-Label Learning Framework Integrating RNA Large Language Model for mRNA Subcellular Localization

  • Xiao Wang
  • Lixiang Yang
  • Rong Wang
  • Yongfeng Zhang

The subcellular localization of messenger RNA (mRNA) is essential for the regulation of gene expression and plays a pivotal role in targeted drug development. Although several computational models have been developed to predict mRNA localization, these approaches still face challenges in sequence representation and exhibit limited performance in handling multi-localization tasks. In this paper, we propose mRSubLoc, a novel multi-label deep learning framework for predicting mRNA subcellular localization. The model integrates the RNA large language model RNAErnie with one-hot encoding and Word2Vec embeddings to construct a comprehensive representation of mRNA sequences. A text convolutional neural network (TextCNN) is employed to capture local feature patterns, while a bidirectional long short-term memory network (BiLSTM) is used to capture long-range dependencies. These features are fused using a multi-head self-attention mechanism to effectively capture localization-specific characteristics. Finally, a multi-layer perceptron (MLP) explores complex dependencies among multiple localization sites, facilitating accurate mRNA subcellular localization prediction. Experimental results on a testing set demonstrate that mRSubLoc significantly outperforms state-of-the-art methods across multiple metrics, including Aiming (0. 7858), Coverage (0. 6212), Accuracy (0. 6161), Absolute-True (0. 3070), and Absolute-False (0. 1319). This study proposes a novel approach for predicting mRNA subcellular localization and provides new perspectives for advancing disease diagnosis and drug discovery in biomedical research.

YNIMG Journal 2026 Journal Article

Neural representations of emotional response inhibition reveal trait and state biomarkers in pediatric bipolar disorder

  • Jia Li
  • Rong Wang
  • Jianze Wu
  • Qian Xiao
  • Yuan Zhong

Pediatric bipolar disorder (PBD) is characterized by disrupted cognitive control, particularly in response inhibition under emotional interference. However, the neural underpinnings of these deficits, particularly how these impairments vary across emotional valence and whether they reflect trait markers or state alterations, remain unclear. While traditional univariate fMRI analyses reveal broad activation differences, they lack sensitivity to fine-grained neural patterns. This study aims to examine the neural representations of emotional response inhibition in PBD under valence-dependent interference using representational similarity analysis(RSA). We included manic (n = 15) and euthymic (n = 18) PBD patients, along with matched healthy controls (n = 17). Participants completed an emotional Go/NoGo task with happy, sad, and neutral faces during fMRI. Six contrast conditions were modeled to assess trait- and state-related effects. Whole-brain searchlight RSA (8 mm radius) was used to identify regions showing group differences in neural representational patterns. Results showed that emotional response inhibition engaged distributed neural systems, with distinct patterns across valence conditions. Compared to controls, PBD patients exhibited trait-related representational differences during happy inhibition, sad inhibition, and sad-specific inhibition, involving regions such as the precentral gyrus, middle frontal gyrus, and inferior parietal lobule. Manic patients showed state-related reductions in neural representations during sad-specific inhibition within frontal areas compared to euthymic patients. These findings indicate that emotional response inhibition deficits in PBD arise from both trait- and state-dependent abnormalities in neural representations. The study highlights the value of multivariate fMRI in uncovering clinically relevant biomarkers and provides a novel framework for developing phase-specific interventions.

AAAI Conference 2026 Conference Paper

Reliable-View 2D-3D Key-Part Aligned Transformer with Reinforced Masking for 3D Point Cloud Understanding

  • Xianglong Jin
  • Zheng Wang
  • Rong Wang
  • Feiping Nie

Self-supervised 3D point cloud understanding is crucial for scene understanding, where Masked Autoencoders (MAE) have achieved excellent performance in point cloud representation learning. However, existing MAE-style methods fail to consider spatial-semantic variations in masking strategies, and joint learning with multi-view images often overlooks view redundancy. To address these challenges, we propose an MAE framework enhanced with reliable multi-view 2D-3D Key-part alignment and Reinforced masking, named as KR-MAE. Our approach comprises three key innovations: Reinforced Masking (RM) strategically samples visible tokens based on semantic saliency to enhance reconstruction fidelity; Reliable Multi-View Selector (RVS) dynamically refines the most informative image subset by filtering occluded or low-texture views, mitigating detrimental redundancy; Reliable-view 2D-3D Key-part Aligned Transformer (KAT) establishes semantic-aligned correspondence between salient 3D point cloud parts and reliable multi-view 2D image patches, leveraging rich texture cues from 2D images to compensate for sparse geometry in point cloud. Extensive experiments on 3D classification and segmentation benchmarks demonstrate that KR-MAE achieves state-of-the-art performance, surpassing prior multi-modal methods.

AAAI Conference 2026 Conference Paper

S2-Boost: Synergistic Semantic Boosting for Coarse-to-Fine Ensemble Learning

  • Guanxiong He
  • Zheng Wang
  • Jie Wang
  • Liaoyuan Tang
  • Rong Wang
  • Feiping Nie

Neuroscientific evidence reveals that human visual recognition is not an instantaneous event but a hierarchical process, where the brain constructs a holistic perception by progressively integrating simple features like edges or texture into complex scenes. Ensemble learning successfully utilizes this principle, yet existing methods typically integrate models at the decision level, neglecting the rich, complementary information within the feature space itself and thus fundamentally limiting their potential. To address this, we introduce Synergistic Semantic Boosting (S2-Boosting), a framework that employs a self-supervised hierarchical semantic learning module to decompose an image into complementary, semantically meaningful parts autonomously. These parts guide a boosting procedure where a sequence of specialized learners, each focusing on a specific semantic partition, collaboratively corrects the ensemble's errors. We further present encouraging results on real-world image datasets, highlighting the intrinsic interpretability, paving the way for more robust and transparent models.

AAAI Conference 2026 Conference Paper

Towards Federated Clustering: A Client-wise Private Graph Aggregation Framework

  • Guanxiong He
  • Zheng Wang
  • Jie Wang
  • Liaoyuan Tang
  • Rong Wang
  • Feiping Nie

Federated clustering addresses the critical challenge of extracting patterns from decentralized, unlabeled data. However, it is hampered by the flaw that current approaches are forced to accept a compromise between performance and privacy: transmitting embedding representations risks sensitive data leakage, while sharing only abstract cluster prototypes leads to diminished model accuracy. To resolve this dilemma, we propose Structural Privacy-Preserving Federated Graph Clustering (SPP-FGC), a novel algorithm that innovatively leverages local structural graphs as the primary medium for privacy-preserving knowledge sharing, thus moving beyond the limitations of conventional techniques. Our framework operates on a clear client-server logic; on the client-side, each participant constructs a private structural graph that captures intrinsic data relationships, which the server then securely aggregates and aligns to form a comprehensive global graph from which a unified clustering structure is derived. The framework offers two distinct modes to suit different needs. SPP-FGC is designed as an efficient one-shot method that completes its task in a single communication round, ideal for rapid analysis. For more complex, unstructured data like images, SPP-FGC+ employs an iterative process where clients and the server collaboratively refine feature representations to achieve superior downstream performance. Extensive experiments demonstrate that our framework achieves state-of-the-art performance, improving clustering accuracy by up to 10% (NMI) over federated baselines while maintaining provable privacy guarantees.

IJCAI Conference 2025 Conference Paper

Capturing Individuality and Commonality Between Anchor Graphs for Multi-View Clustering

  • Zhoumin Lu
  • Yongbo Yu
  • Linru Ma
  • Feiping Nie
  • Rong Wang

The use of anchors often leads to better efficiency and scalability, making them highly favored. However, there is a challenge in anchor-based multi-view subspace learning. A unified anchor graph overly emphasize the commonality between views, failing to adequately capture the view-specific individuality. This has led some models to independently explore the individuality of each view before aligning and integrating them, often achieving better performance but making the process more cumbersome. Therefore, this paper proposes a new model, simultaneously capturing the individuality and commonality between anchor graphs for multi-view clustering. The model has three notable advantages: First, it allows view-specific anchor graphs to align in real-time with a common anchor graph as a reference, eliminating the need for post-alignment. Second, it enforces a cluster-wise structure among anchors and balances sample distribution among them, providing strong discriminative power. Lastly, it maintains linear complexity with respect to the numbers of samples and anchors, avoiding the significant time costs associated with their increase. Comprehensive experiments demonstrate the effectiveness and efficiency of our method compared to various state-of-the-art algorithms.

YNIMG Journal 2025 Journal Article

Exploring multidimensional brain mechanisms in robot-assisted surgical simulation

  • Haoxin Cui
  • Yujing Liang
  • Fankai Sun
  • Desheng Li
  • Xiangqing Wang
  • Rong Wang
  • Nan Zheng

The introduction of robotic-assisted surgical systems has revolutionized surgical procedures; however, current training programs often overlook the role of brain activity during surgery, making it challenging to detect cognitive differences between surgeons. To address this gap, this paper designed an experimental task closely resembling real surgical scenarios using a robotic surgical simulation system. The study introduced Principal Component Analysis (PCA) weights and Mahalanobis distance as metrics for identifying cognitive differences, with a focus on investigating the brain mechanisms underlying varying levels of surgical proficiency in terms of frequency domain, neural connectivity, and graph theory. Frequency domain analyses revealed that experienced surgeons exhibited greater activation in the alpha bands of the prefrontal cortex (Fp1, Fp2), occipital cortex (O1, O2), and midline parietal cortex (Pz) during task execution, compared to less experienced surgeons. Connectivity analysis indicated that high-level surgeons demonstrated superior neural efficiency, characterized by weaker localized activity but enhanced global integration of brain regions. Graph theoretical analyses further highlighted differences in network organization, with higher-level surgeons achieving a balanced interplay between local specialization and global integration of brain networks. Finally, classification and ablation experiments confirmed that the EEG features identified in this study effectively differentiate surgeons based on their operational expertise. These findings provide valuable insights into the underlying brain mechanisms involved in surgical proficiency and offer potential applications for supporting surgeon training and objective assessment of surgical skills. This research paves the way for the development of more targeted training programs for robotic surgery, ultimately enhancing the effectiveness of skill development and performance evaluation.

EAAI Journal 2025 Journal Article

Integrating multimodal biophysical features with hybrid deep learning for ribonucleic acid secondary structure prediction

  • Xiao Wang
  • Yongfeng Zhang
  • Lixiang Yang
  • Rong Wang

The secondary structure of ribonucleic acid (RNA) is pivotal for elucidating its functional roles. However, existing deep learning models predominantly rely on single-feature representations, which restricts their capacity to sufficiently capture the intricate information embedded in RNA sequences. To solve this problem, we propose a novel RNA secondary structure prediction method based on multimodal feature fusion and hybrid deep learning. By integrating multiple features of RNA are modeled to achieve the synergistic effect of the chemical microenvironment, local physical constraints and long-range interactions. We conceptualize RNA structure as an image and treat various RNA features as distinct channels of that image. To facilitate the transformation from RNA sequences to RNA images, we design a hybrid neural network architecture. Specifically, the sequence feature extraction module first extracts features directly from the RNA sequence. These features are then passed to the image feature extraction module, which focuses on capturing effective structural information from the multichannel RNA image, thereby enhancing the accuracy of RNA secondary-structure prediction. Experimental results indicate that our method achieves state-of-the-art performance compared with recent deep learning methods across several benchmark datasets and demonstrates potential in drug discovery applications.

AAAI Conference 2025 Conference Paper

Language Pre-training Guided Masking Representation Learning for Time Series Classification

  • Liaoyuan Tang
  • Zheng Wang
  • Jie Wang
  • Guanxiong He
  • Zhezheng Hao
  • Rong Wang
  • Feiping Nie

The representation learning of time series has a wide range of downstream tasks and applications in many practical scenarios. However, due to the complexity, spatiotemporality, and continuity of sequential stream data, compared with the representation learning of structural data such as images/videos, the time series self-supervised representation learning is even more challenging. Besides, the direct application of existing contrastive learning and masked autoencoder based approaches to time series representation learning encounters inherent theoretical limitations, such as ineffective augmentation and masking strategies. To this end, we propose a Language Pre-training guided Masking Representation Learning (LPMRL) for times series classification. Specifically, we first propose a novel language pre-training guided masking encoder for adaptively sampling semantic spatiotemporal patches via natural language descriptions and improving the discriminability of latent representations. Furthermore, we present the dual-information contrastive learning mechanism to explore both local and global information by meticulously designing high-quality hard negative samples of time series data samples. As a result, we also design various experiments, such as visualization of masking position and distribution and reconstruction error to verify the reasonability of proposed language guided masking technique. Last, we evaluate the performance of proposed representation learning via classification task conducted on 106 time series datasets, which demonstrates the effectiveness of proposed method.

ICLR Conference 2024 Conference Paper

Exploring the cloud of feature interaction scores in a Rashomon set

  • Sichao Li
  • Rong Wang
  • Quanling Deng
  • Amanda S. Barnard

Interactions among features are central to understanding the behavior of machine learning models. Recent research has made significant strides in detecting and quantifying feature interactions in single predictive models. However, we argue that the feature interactions extracted from a single pre-specified model may not be trustworthy since: *a well-trained predictive model may not preserve the true feature interactions and there exist multiple well-performing predictive models that differ in feature interaction strengths*. Thus, we recommend exploring feature interaction strengths in a model class of approximately equally accurate predictive models. In this work, we introduce the feature interaction score (FIS) in the context of a Rashomon set, representing a collection of models that achieve similar accuracy on a given task. We propose a general and practical algorithm to calculate the FIS in the model class. We demonstrate the properties of the FIS via synthetic data and draw connections to other areas of statistics. Additionally, we introduce a Halo plot for visualizing the feature interaction variance in high-dimensional space and a swarm plot for analyzing FIS in a Rashomon set. Experiments with recidivism prediction and image classification illustrate how feature interactions can vary dramatically in importance for similarly accurate predictive models. Our results suggest that the proposed FIS can provide valuable insights into the nature of feature interactions in machine learning models.

TMLR Journal 2024 Journal Article

Hyperspherical Prototype Node Clustering

  • Jitao Lu
  • Danyang Wu
  • Feiping Nie
  • Rong Wang
  • Xuelong Li

The general workflow of deep node clustering is to encode the nodes into node embeddings via graph neural networks and uncover clustering decisions from them, so clustering performance is heavily affected by the embeddings. However, existing works only consider preserving the semantics of the graph but ignore the inter-cluster separability of the nodes, so there's no guarantee that the embeddings can present a clear clustering structure. To remedy this deficiency, we propose Hyperspherical Prototype Node Clustering (HPNC), an end-to-end clustering paradigm that explicitly enhances the inter-cluster separability of learned node embeddings. Concretely, we constrain the embedding space to a unit-hypersphere, enabling us to scatter the cluster prototypes over the space with maximized pairwise distances. Then, we employ a graph autoencoder to map nodes onto the same hypersphere manifold. Consequently, cluster affinities can be directly retrieved from cosine similarities between node embeddings and prototypes. A clustering-oriented loss is imposed to sharpen the affinity distribution so that the learned node embeddings are encouraged to have small intra-cluster distances and large inter-cluster distances. Based on the proposed HPNC paradigm, we devise two schemes (HPNC-IM and HPNC-DEC) with distinct clustering backbones. Empirical results on popular benchmark datasets demonstrate the superiority of our method compared to other state-of-the-art clustering methods, and visualization results illustrate improved separability of the learned embeddings.

AAAI Conference 2024 Conference Paper

Multi-Class Support Vector Machine with Maximizing Minimum Margin

  • Feiping Nie
  • Zhezheng Hao
  • Rong Wang

Support Vector Machine (SVM) stands out as a prominent machine learning technique widely applied in practical pattern recognition tasks. It achieves binary classification by maximizing the "margin", which represents the minimum distance between instances and the decision boundary. Although many efforts have been dedicated to expanding SVM for multi-class case through strategies such as one versus one and one versus the rest, satisfactory solutions remain to be developed. In this paper, we propose a novel method for multi-class SVM that incorporates pairwise class loss considerations and maximizes the minimum margin. Adhering to this concept, we embrace a new formulation that imparts heightened flexibility to multi-class SVM. Furthermore, the correlations between the proposed method and multiple forms of multi-class SVM are analyzed. The proposed regularizer, akin to the concept of "margin", can serve as a seamless enhancement over the softmax in deep learning, providing guidance for network parameter learning. Empirical evaluations demonstrate the effectiveness and superiority of our proposed method over existing multi-classification methods. Complete version is available at https://arxiv.org/pdf/2312.06578.pdf. Code is available at https://github.com/zz-haooo/M3SVM.

IJCAI Conference 2024 Conference Paper

Perturbation Guiding Contrastive Representation Learning for Time Series Anomaly Detection

  • Liaoyuan Tang
  • Zheng Wang
  • Guanxiong He
  • Rong Wang
  • Feiping Nie

Time series anomaly detection is a critical task with applications in various domains. Due to annotation challenges, self-supervised methods have become the mainstream approach for time series anomaly detection in recent years. However, current contrastive methods categorize data perturbations into binary classes, normal or anomaly, which lack clarity on the specific impact of different perturbation methods. Inspired by the hypothesis that "the higher the probability of misclassifying perturbation types, the higher the probability of anomalies", we propose PCRTA, our approach firstly devises a perturbation classifier to learn the pseudo-labels of data perturbations. Furthermore, for addressing "class collapse issue" in contrastive learning, we propose a perturbation guiding positive and negative samples selection strategy by introducing learnable perturbation classification networks. Extensive experiments on six realworld datasets demonstrate the significant superiority of our model over thirteen state-of-the-art competitors, and obtains average 5. 14%, 8. 24% improvement in F1 score and AUC-PR, respectively.

NeurIPS Conference 2023 Conference Paper

DeepSimHO: Stable Pose Estimation for Hand-Object Interaction via Physics Simulation

  • Rong Wang
  • Wei Mao
  • Hongdong Li

This paper addresses the task of 3D pose estimation for a hand interacting with an object from a single image observation. When modeling hand-object interaction, previous works mainly exploit proximity cues, while overlooking the dynamical nature that the hand must stably grasp the object to counteract gravity and thus preventing the object from slipping or falling. These works fail to leverage dynamical constraints in the estimation and consequently often produce unstable results. Meanwhile, refining unstable configurations with physics-based reasoning remains challenging, both by the complexity of contact dynamics and by the lack of effective and efficient physics inference in the data-driven learning framework. To address both issues, we present DeepSimHO: a novel deep-learning pipeline that combines forward physics simulation and backward gradient approximation with a neural network. Specifically, for an initial hand-object pose estimated by a base network, we forward it to a physics simulator to evaluate its stability. However, due to non-smooth contact geometry and penetration, existing differentiable simulators can not provide reliable state gradient. To remedy this, we further introduce a deep network to learn the stability evaluation process from the simulator, while smoothly approximating its gradient and thus enabling effective back-propagation. Extensive experiments show that our method noticeably improves the stability of the estimation and achieves superior efficiency over test-time optimization. The code is available at https: //github. com/rongakowang/DeepSimHO.

AAAI Conference 2023 Conference Paper

Efficient Top-K Feature Selection Using Coordinate Descent Method

  • Lei Xu
  • Rong Wang
  • Feiping Nie
  • Xuelong Li

Sparse learning based feature selection has been widely investigated in recent years. In this study, we focus on the l2,0-norm based feature selection, which is effective for exact top-k feature selection but challenging to optimize. To solve the general l2,0-norm constrained problems, we novelly develop a parameter-free optimization framework based on the coordinate descend (CD) method, termed CD-LSR. Specifically, we devise a skillful conversion from the original problem to solving one continuous matrix and one discrete selection matrix. Then the nontrivial l2,0-norm constraint can be solved efficiently by solving the selection matrix with CD method. We impose the l2,0-norm on a vanilla least square regression (LSR) model for feature selection and optimize it with CD-LSR. Extensive experiments exhibit the efficiency of CD-LSR, as well as the discrimination ability of l2,0-norm to identify informative features. More importantly, the versatility of CD-LSR facilitates the applications of the l2,0-norm in more sophisticated models. Based on the competitive performance of l2,0-norm on the baseline LSR model, the satisfactory performance of its applications is reasonably expected. The source MATLAB code are available at: https://github.com/solerxl/Code_For_AAAI_2023.

NeurIPS Conference 2023 Conference Paper

Joint Feature and Differentiable $ k $-NN Graph Learning using Dirichlet Energy

  • Lei Xu
  • Lei Chen
  • Rong Wang
  • Feiping Nie
  • Xuelong Li

Feature selection (FS) plays an important role in machine learning, which extracts important features and accelerates the learning process. In this paper, we propose a deep FS method that simultaneously conducts feature selection and differentiable $ k $-NN graph learning based on the Dirichlet Energy. The Dirichlet Energy identifies important features by measuring their smoothness on the graph structure, and facilitates the learning of a new graph that reflects the inherent structure in new feature subspace. We employ Optimal Transport theory to address the non-differentiability issue of learning $ k $-NN graphs in neural networks, which theoretically makes our method applicable to other graph neural networks for dynamic graph learning. Furthermore, the proposed framework is interpretable, since all modules are designed algorithmically. We validate the effectiveness of our model with extensive experiments on both synthetic and real-world datasets.

IJCAI Conference 2022 Conference Paper

EMGC²F: Efficient Multi-view Graph Clustering with Comprehensive Fusion

  • Danyang Wu
  • Jitao Lu
  • Feiping Nie
  • Rong Wang
  • Yuan Yuan

This paper proposes an Efficient Multi-view Graph Clustering with Comprehensive Fusion (EMGC²F) model and a corresponding efficient optimization algorithm to address multi-view graph clustering tasks effectively and efficiently. Compared to existing works, our proposals have the following highlights: 1) EMGC²F directly finds a consistent cluster indicator matrix with a Super Nodes Similarity Minimization module from multiple views, which avoids time-consuming spectral decomposition in previous works. 2) EMGC²F comprehensively mines information from multiple views. More formally, it captures the consistency of multiple views via a Cross-view Nearest Neighbors Voting (CN²V) mechanism, meanwhile capturing the importance of multiple views via an adaptive weighted-learning mechanism. 3) EMGC²F is a parameter-free model and the time complexity of the proposed algorithm is far less than existing works, demonstrating the practicability. Empirical results on several benchmark datasets demonstrate that our proposals outperform SOTA competitors both in effectiveness and efficiency.

JBHI Journal 2022 Journal Article

Flexible Brain Transitions Between Hierarchical Network Segregation and Integration Associated With Cognitive Performance During a Multisource Interference Task

  • Rong Wang
  • Xiaoli Su
  • Zhao Chang
  • Pan Lin
  • Ying Wu

Cognition involves locally segregated and globally integrated processing. This process is hierarchically organized and linked to evidence from hierarchical modules in brain networks. However, researchers have not clearly determined how flexible transitions between these hierarchical processes are associated with cognitive performance. Here, we designed a multisource interference task (MSIT) and introduced the nested-spectral partition (NSP) method to detect hierarchical modules in brain functional networks. By defining hierarchical segregation and integration across multiple levels, we showed that the MSIT requires higher network segregation in the whole brain and most functional systems but generates higher integration in the control system. Meanwhile, brain networks have more flexible transitions between segregated and integrated configurations in the task state. Crucially, higher functional flexibility in the resting state, less flexibility in the task state and more efficient switching of the brain from resting to task states were associated with better task performance. Our hierarchical modular analysis was more effective at detecting alterations in functional organization and the phenotype of cognitive performance than graph-based network measures at a single level.

AAAI Conference 2022 Conference Paper

Transcribing Natural Languages for the Deaf via Neural Editing Programs

  • Dongxu Li
  • Chenchen Xu
  • Liu Liu
  • Yiran Zhong
  • Rong Wang
  • Lars Petersson
  • Hongdong Li

This work studies the task of glossification, of which the aim is to transcribe natural spoken language sentences for the Deaf (hard-of-hearing) community to ordered sign language glosses. Previous sequence-to-sequence language models trained with paired sentence-gloss data often fail to capture the rich connections between the two distinct languages, leading to unsatisfactory transcriptions. We observe that despite different grammars, glosses effectively simplify sentences for the ease of deaf communication, while sharing a large portion of vocabulary with sentences. This has motivated us to implement glossification by executing a collection of editing actions, e. g. word addition, deletion and copying, called editing programs, on their natural spoken language counterparts. Specifically, we design a new neural agent that learns to synthesize and execute editing programs, conditioned on sentence contexts and partial editing results. The agent is trained to imitate minimal editing programs, while exploring more widely the program space via policy gradients to optimize sequence-wise transcription quality. Results show that our approach outperforms previous glossification models by a large margin, improving the BLEU-4 score from 16. 45 to 18. 89 on RWTH-PHOENIX- WEATHER-2014T and from 18. 38 to 21. 30 on CSL-Daily.

TIST Journal 2021 Journal Article

Attentive Excitation and Aggregation for Bilingual Referring Image Segmentation

  • Qianli Zhou
  • Tianrui Hui
  • Rong Wang
  • Haimiao Hu
  • Si Liu

The goal of referring image segmentation is to identify the object matched with an input natural language expression. Previous methods only support English descriptions, whereas Chinese is also broadly used around the world, which limits the potential application of this task. Therefore, we propose to extend existing datasets with Chinese descriptions and preprocessing tools for training and evaluating bilingual referring segmentation models. In addition, previous methods also lack the ability to collaboratively learn channel-wise and spatial-wise cross-modal attention to well align visual and linguistic modalities. To tackle these limitations, we propose a Linguistic Excitation module to excite image channels guided by language information and a Linguistic Aggregation module to aggregate multimodal information based on image-language relationships. Since different levels of features from the visual backbone encode rich visual information, we also propose a Cross-Level Attentive Fusion module to fuse multilevel features gated by language information. Extensive experiments on four English and Chinese benchmarks show that our bilingual referring image segmentation model outperforms previous methods.

IJCAI Conference 2021 Conference Paper

Discrete Multiple Kernel k-means

  • Rong Wang
  • Jitao Lu
  • Yihang Lu
  • Feiping Nie
  • Xuelong Li

The multiple kernel k-means (MKKM) and its variants utilize complementary information from different kernels, achieving better performance than kernel k-means (KKM). However, the optimization procedures of previous works all comprise two stages, learning the continuous relaxed label matrix and obtaining the discrete one by extra discretization procedures. Such a two-stage strategy gives rise to a mismatched problem and severe information loss. To address this problem, we elaborate a novel Discrete Multiple Kernel k-means (DMKKM) model solved by an optimization algorithm that directly obtains the cluster indicator matrix without subsequent discretization procedures. Moreover, DMKKM can strictly measure the correlations among kernels, which is capable of enhancing kernel fusion by reducing redundancy and improving diversity. What’s more, DMKKM is parameter-free avoiding intractable hyperparameter tuning, which makes it feasible in practical applications. Extensive experiments illustrated the effectiveness and superiority of the proposed model.

IJCAI Conference 2021 Conference Paper

GSPL: A Succinct Kernel Model for Group-Sparse Projections Learning of Multiview Data

  • Danyang Wu
  • Jin Xu
  • Xia Dong
  • Meng Liao
  • Rong Wang
  • Feiping Nie
  • Xuelong Li

This paper explores a succinct kernel model for Group-Sparse Projections Learning (GSPL), to handle multiview feature selection task completely. Compared to previous works, our model has the following useful properties: 1) Strictness: GSPL innovatively learns group-sparse projections strictly on multiview data via ‘2; 0-norm constraint, which is different with previous works that encourage group-sparse projections softly. 2) Adaptivity: In GSPL model, when the total number of selected features is given, the numbers of selected features of different views can be determined adaptively, which avoids artificial settings. Besides, GSPL can capture the differences among multiple views adaptively, which handles the inconsistent problem among different views. 3) Succinctness: Except for the intrinsic parameters of projection-based feature selection task, GSPL does not bring extra parameters, which guarantees the applicability in practice. To solve the optimization problem involved in GSPL, a novel iterative algorithm is proposed with rigorously theoretical guarantees. Experimental results demonstrate the superb performance of GSPL on synthetic and real datasets.

IJCAI Conference 2020 Conference Paper

Discriminative Feature Selection via A Structured Sparse Subspace Learning Module

  • Zheng Wang
  • Feiping Nie
  • Lai Tian
  • Rong Wang
  • Xuelong Li

In this paper, we first propose a novel Structured Sparse Subspace Learning S^3L module to address the long-standing subspace sparsity issue. Elicited by proposed module, we design a new discriminative feature selection method, named Subspace Sparsity Discriminant Feature Selection S^2DFS which enables the following new functionalities: 1) Proposed S^2DFS method directly joints trace ratio objective and structured sparse subspace constraint via L2, 0-norm to learn a row-sparsity subspace, which improves the discriminability of model and overcomes the parameter-tuning trouble with comparison to the methods used L2, 1-norm regularization; 2) An alternative iterative optimization algorithm based on the proposed S^3L module is presented to explicitly solve the proposed problem with a closed-form solution and strict convergence proof. To our best knowledge, such objective function and solver are first proposed in this paper, which provides a new though for the development of feature selection methods. Extensive experiments conducted on several high-dimensional datasets demonstrate the discriminability of selected features via S^2DFS with comparison to several related SOTA feature selection methods. Source matlab code: https: //github. com/StevenWangNPU/L20-FS.

NeurIPS Conference 2020 Conference Paper

Efficient Clustering Based On A Unified View Of $K$-means And Ratio-cut

  • Shenfei Pei
  • Feiping Nie
  • Rong Wang
  • Xuelong Li

Spectral clustering and $k$-means, both as two major traditional clustering methods, are still attracting a lot of attention, although a variety of novel clustering algorithms have been proposed in recent years. Firstly, a unified framework of $k$-means and ratio-cut is revisited, and a novel and efficient clustering algorithm is then proposed based on this framework. The time and space complexity of our method are both linear with respect to the number of samples, and are independent of the number of clusters to construct, more importantly. These properties mean that it is easily scalable and applicable to large practical problems. Extensive experiments on 12 real-world benchmark and 8 facial datasets validate the advantages of the proposed algorithms compared to the state-of-the-art clustering algorithms. In particular, over 15x and 7x speed-up can be obtained with respect to $k$-means on the synthetic dataset of 1 million samples and the benchmark dataset (CelebA) of 200k samples, respectively [GitHub].

NeurIPS Conference 2020 Conference Paper

Learning Feature Sparse Principal Subspace

  • Lai Tian
  • Feiping Nie
  • Rong Wang
  • Xuelong Li

This paper presents new algorithms to solve the feature-sparsity constrained PCA problem (FSPCA), which performs feature selection and PCA simultaneously. Existing optimization methods for FSPCA require data distribution assumptions and are lack of global convergence guarantee. Though the general FSPCA problem is NP-hard, we show that, for a low-rank covariance, FSPCA can be solved globally (Algorithm 1). Then, we propose another strategy (Algorithm 2) to solve FSPCA for the general covariance by iteratively building a carefully designed proxy. We prove (data-dependent) approximation bound and convergence guarantees for the new algorithms. For the spectrum of covariance with exponential/Zipf's distribution, we provide exponential/posynomial approximation bound. Experimental results show the promising performance and efficiency of the new algorithms compared with the state-of-the-arts on both synthetic and real-world datasets.

IJCAI Conference 2020 Conference Paper

Semi-supervised Clustering via Pairwise Constrained Optimal Graph

  • Feiping Nie
  • Han Zhang
  • Rong Wang
  • Xuelong Li

In this paper, we present a technique of definitely addressing the pairwise constraints in the semi-supervised clustering. Our method contributes to formulating the cannot-link relations and propagating them over the affinity graph flexibly. The pairwise constrained instances are provably guaranteed to be in the same or different connected components of the graph. Combined with the Laplacian rank constraint, the proposed model learns a Pairwise Constrained structured Optimal Graph (PCOG), from which the specified c clusters supporting the known pairwise constraints are directly obtained. An efficient algorithm invoked by the label propagation is designed to solve the formulation. Additionally, we also provide a compact criterion to acquire the key pairwise constraints for prompting the semi-supervised graph clustering. Substantial experimental results show that the proposed method achieves the significant improvements by using a few prior pairwise constraints.

YNICL Journal 2020 Journal Article

Topological reorganization of brain functional networks in patients with mitochondrial encephalomyopathy with lactic acidosis and stroke‐like episodes

  • Rong Wang
  • Jie Lin
  • Chong Sun
  • Bin Hu
  • Xueling Liu
  • Daoying Geng
  • Yuxin Li
  • Liqin Yang

Mitochondrial encephalomyopathy with lactic acidosis and stroke-like episodes (MELAS) is a rare maternally inherited genetic disease; however, little is known about its underlying brain basis. Furthermore, the topological organization of brain functional network in MELAS has not been explored. Here, 45 patients with MELAS (22 at acute stage, 23 at chronic stage) and 22 normal controls were studied using resting- state functional magnetic resonance imaging and graph theory analysis approaches. Topological properties of brain functional networks including global and nodal metrics, rich club organization and modularity were analyzed. At the global level, MELAS patients exhibited reduced clustering coefficient, normalized clustering coefficient, normalized characteristic path length and local network efficiency compared with the controls. At the nodal level, several nodes with abnormal degree centrality and nodal efficiency were detected in MELAS patients, and the distribution of these nodes was partly consistent with the stroke-like lesions. For rich club organization, rich club nodes were reorganized and the connections among them were decreased in MELAS patients. Modularity analysis revealed that MELAS patents had altered intra- or inter-modular connections in default mode network, fronto-parietal network, sensorimotor network, occipital network and cerebellum network. Notably, the patients at acute stage showed more obvious changes in these topological properties than the patients at chronic stage. These findings indicated that MELAS patients, particularly those at acute stage, exhibited topological reorganization of the whole-brain functional network. This study may help us to understand the neuropathological mechanisms of MELAS.

JBHI Journal 2018 Journal Article

3-D Tracking for Augmented Reality Using Combined Region and Dense Cues in Endoscopic Surgery

  • Rong Wang
  • Mei Zhang
  • Xiangbing Meng
  • Zheng Geng
  • Fei-Yue Wang

An augmented reality (AR) technique has recently gained its popularity in minimally invasive surgery. Tracking is a crucial step to achieve precise AR. Besides optical tracking in traditional medical AR, visual tracking attracts a lot of attention due to its generality. Moreover, when the target organ's 3-D model can be obtained from preoperative images and under the model rigidity assumption, tracking is then converted into a problem of computing the six-degree-of-freedom pose of the 3-D model. In this paper, we introduce a robust tracking algorithm in our endoscopic AR system, where we combine the benefits of both region and dense cues in a unified framework. Each kind of cues alone may not be adequate for tracking in endoscopic surgery. However, they have complementary characteristics, with region cues being more robust to motion blur and fast motion, and dense cues being more accurate when motion is not large. We also propose an appearance model adaption method and an occlusion processing method to effectively handle occlusions. Experiments on both synthetic dataset and simulated surgical environment show the effectiveness and robustness of our proposed method. This work presents a novel tracking strategy in medical AR applications.

YNIMG Journal 2009 Journal Article

In vivo MRI of endogenous stem/progenitor cell migration from subventricular zone in normal and injured developing brains

  • Jian Yang
  • Jianxin Liu
  • Gang Niu
  • Kevin C. Chan
  • Rong Wang
  • Yong Liu
  • Ed X. Wu

Understanding the alterations of migratory activities of the endogenous neural stem/progenitor cells (NSPs) in injured developing brains is becoming increasingly imperative for curative reasons. In this study, 10-day-old neonatal rats with and without hypoxic–ischemic (HI) insult at postnatal day 7 were injected intraventricularly with micron-sized iron oxide particles (MPIOs), followed by serial high-resolution MRI at 7 T for 2 weeks. MRI findings were correlated to the histological analysis using iron staining and several immunohistochemical double staining. The results indicated that in normal and HI-injured brains the NSPs from the subventricular zone (SVZ) were labeled by MPIOs, and migrated as newly created cells (iron+/BrdU+), neuroblasts (iron+/nestin+), astrocytes or astrocytes-like progenitor cells (iron+/GFAP+), and mature neurons (iron+/NeuN+). In normal brains, the endogenous NSPs mainly exhibited a tangential pattern in both rostral and caudal directions. The NSP radial migratory pattern could be observed in some rats. In the HI-injured brains during the same developmental period, the NSPs mainly migrated towards the HI lesion sites. The tangential, rostrocaudal migrations could be observed but impaired. These findings suggest that the NSP migratory pathways in SVZ change in response to the HI insult, likely due to the self-repairing efforts known in the neonatal brains. The MRI approach demonstrated here is potentially applicable to the in vivo and longitudinal study of NSP cell activities in developing brains under normal and pathological conditions and in therapeutic interventions.

YNIMG Journal 2006 Journal Article

Transient blood pressure changes affect the functional magnetic resonance imaging detection of cerebral activation

  • Rong Wang
  • Tadeusz Foniok
  • Jaclyn I. Wamsteeker
  • Min Qiao
  • Boguslaw Tomanek
  • Rodrigo A. Vivanco
  • Ursula I. Tuor

Functional magnetic resonance imaging (fMRI) provides an indirect measure of cerebral activation that could be altered by factors directly affecting cerebral blood flow independent of changes in neuronal activation. Presently, we investigate how changes in blood pressure (BP) affect the activation detected with fMRI. fMRI scans were acquired in 33 rats under control conditions and following transient BP increases (norepinephrine, IV) or decreases (arfonad, IV) with and without electrical stimulation of the forepaw. Voxels correlating to either the stimulation or the change in BP time courses were identified. During transient hypertension, irrespective of forepaw stimulation, BP increases (i. e. , >10 mm Hg) produced a transient increase in the blood oxygen level-dependent (BOLD) intensity resulting in a significant numbers of voxels correlating to the BP time courses (P < 0. 05), and the number of these voxels increased as BP increased, becoming substantial at BP > 30 mm Hg. The activation patterns with BP increases and stimulation overlapped spatially resulting in an enhanced cerebral activation to simultaneous forepaw stimulation (P < 0. 05). BP decreases (>10 mm Hg) produced corresponding decreases in BOLD intensity, causing significant numbers of voxels correlating to the BP decreases (P < 0. 005), and these numbers increased as BP decreased (P < 0. 001). The BP decreases and stimulation time courses and responses were distinct, and hypotension did not affect the detection of the activation response to forepaw stimulation. The results indicate that substantial hypertension accompanying a stimulation paradigm produces a BOLD response that enhances the cerebral activation detected, whereas hypotension does not affect the detection of neuronal activation but does produce responses that could be interpreted as a ‘deactivation’.

v2026.09.13