Arrow Research search

Author name cluster

Wei Gao

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

70 papers
2 author rows

Possible papers

70

YNIMG Journal 2026 Journal Article

Associations between glymphatic dysfunction, white matter injury, and cognitive decline in Parkinson’s disease

  • Jun Yao
  • Yuting Xia
  • Song’an Shang
  • Ting Huang
  • Youyong Tian
  • Wei Gao
  • Yan Gu
  • Yu-Chen Chen

Parkinson's disease (PD) is a neurodegenerative disorder characterized by motor and nonmotor symptoms, with cognitive impairment significantly affecting patients' quality of life. This study aimed to investigate the relationships among glymphatic dysfunction, white matter (WM) injury and cognitive decline in PD patients. Seventy PD patients and 82 healthy controls (HCs) underwent clinical evaluations and magnetic resonance imaging (MRI) scans. Key metrics included the diffusion tensor imaging-along the perivascular space (DTI-ALPS) index, choroid plexus volume (CPV), white matter free water (WM-FW), and peak width of skeletonized mean diffusivity (PSMD). Statistical analyses included correlation analyses, mediation analysis and receiver operating characteristic (ROC) analyses. PD patients exhibited lower DTI-ALPS index (p < 0.001) and higher CPV (p < 0.001), WM-FW (p = 0.010), and PSMD (p < 0.001) compared to the HCs. The DTI-ALPS index was negatively correlated with WM-FW (r = -0.612, p < 0.001) and PSMD (r = -0.484, p < 0.001), whereas CPV was positively correlated with both (r = 0.613, p < 0.001; r = 0.540, p < 0.001). The DTI-ALPS index correlated positively with Montreal Cognitive Assessment (MoCA) score (ρ = 0.471, p < 0.001). CPV (ρ = -0.421, p = 0.002), WM-FW (ρ = -0.296, p = 0.029), and PSMD (ρ = -0.273, p = 0.044) correlated negatively with the MoCA score. Mediation analysis suggested that DTI-ALPS and CPV may be involved in the association between white matter injury and cognition, but the independent role of each individual indicator was not confirmed. Diagnostic performance evaluations indicated that the PSMD best predicted PD individually (AUC = 0.836), with the integrated four-biomarker model performing best (AUC = 0.855). These findings highlight the correlation between glymphatic function, WM integrity, and cognition in PD patients, supporting the use of these neuroimaging biomarkers for early diagnosis and monitoring of cognitive decline.

AAAI Conference 2026 Conference Paper

Correcting Quantization-Induced Gradient Mismatch in Neural Image Compression

  • Changhao Peng
  • Yuqi Ye
  • Wei Gao

In recent years, neural image compression methods have achieved impressive performance in image compression tasks, most of which are based on variational auto-encoder with hyper-prior and autoregressive Gaussian entropy model. We first demonstrate that the way these end-to-end approaches handle quantization during training leads to a mismatch between the gradients direction of entropy model parameters (i.e., mean and standard deviation) and the direction they should be optimized towards during inference, making neural network difficult to learn accurate estimates of entropy model parameters. To address this issue, we then propose a two-step improvement: in the first step, use straight-through estimator to align the forward propagation during training with inference, thereby correcting the gradients of standard deviation parameters; in the second step, utilize gradients transfer that we propose and MSE-guided gradients to manually compensate for the gradients of mean parameters lost due to straight-through estimator. Finally, we also propose to freeze the auto-encoder and hyper auto-encoder in pre-trained models provided by existing works, and fine-tune only the modules that predict the entropy model parameters, enabling efficient validation of proposed improvements. Experimental results show that our improvements bring appreciable performance gains to state-of-the-art neural image compression models in recent years. Meanwhile, our improvements require no modification to the structure of pre-trained models and only lightweight fine-tuning, which shows strong plug-and-play capability and practical utility.

AAAI Conference 2026 Conference Paper

MMBERT: Scaled Mixture-of-Experts Multimodal BERT for Robust Chinese Hate Speech Detection Under Cloaking Perturbations

  • Qiyao Xue
  • Yuchen Dou
  • Zheyuan Ryan Shi
  • Xiang Lorraine Li
  • Wei Gao

Hate speech detection on Chinese social networks presents distinct challenges, particularly due to the widespread use of cloaking techniques designed to evade conventional text-based detection systems. Although large language models (LLMs) have recently improved hate speech detection capabilities, the majority of existing work has concentrated on English datasets, with limited attention given to multimodal strategies in the Chinese context. In this study, we propose MMBERT, a novel BERT-based multimodal framework that integrates textual, speech, and visual modalities through a Mixture-of-Experts (MoE) architecture. To address the instability associated with directly integrating MoE into BERT-based models, we develop a progressive three-stage training paradigm. MMBERT incorporates modality-specific experts, a shared self-attention mechanism, and a router-based expert allocation strategy to enhance robustness against adversarial perturbations. Empirical results in several Chinese hate speech datasets show that MMBERT significantly surpasses fine-tuned BERT-based encoder models, fine-tuned LLMs, and LLMs utilizing in-context learning approaches.

AAAI Conference 2025 Conference Paper

AdaDPCC: Adaptive Rate Control and Rate-Distortion-Complexity Optimization for Dynamic Point Cloud Compression

  • Chenhao Zhang
  • Wei Gao

Dynamic point cloud compression (DPCC) is crucial in applications like autonomous driving and AR/VR. Current compression methods face challenges with complexity management and rate control. This paper introduces a novel dynamic coding framework that supports variable bitrate and computational complexities. Our approach includes a slimmable framework with multiple coding routes, allowing for efficient Rate-Distortion-Complexity Optimization (RDCO) within a single model. To address data sparsity in inter-frame prediction, we propose the coarse-to-fine motion estimation and compensation module that deconstructs geometric information while expanding the perceptive field. Additionally, we propose a precise rate control module that content-adaptively navigates point cloud frames through various coding routes to meet target bitrates. The experimental results demonstrate that our approach reduces the average BD-Rate by 5.81% and improves the BD-PSNR by 0.42 dB compared to the state-of-the-art method, while keeping the average bitrate error at 0.40%. Moreover, the average coding time is reduced by up to 44.6% compared to D-DPCC, underscoring its efficiency in real-time and bitrate-constrained DPCC scenarios.

EAAI Journal 2025 Journal Article

Assessment of hybrid kernel function in extreme support vector regression model for streamflow time series forecasting based on a bayesian estimator decomposition algorithm

  • Peng Shi
  • Lei Xu
  • Simin Qu
  • Hongshi Wu
  • Qiongfang Li
  • Yiqun Sun
  • Xiaoqiang Yang
  • Wei Gao

Diverse decomposition algorithms have been widely employed to streamflow time series forecasting. Their applications, however, are hindered by the plausible high accuracy in the overall decomposition-based framework. This paper firstly introduces a novel decomposition algorithm named Bayesian estimator of abrupt change, seasonality and trend (BEAST) into streamflow forecasting to alleviate the boundary effect. Practical samples are generated under the modified two-stage decomposition prediction (TSDP) framework. A hybrid kernel function, which benefits from two different standalone ones, is designed for kernel extreme support vector regression and the HKESVR model is trained on the samples using 10-fold cross-validation strategy. Comparative experiments are conducted on three monthly streamflow series from basins with diverse hydroclimatic conditions. The results in different lead times (1-, 3-, and 5-month-ahead) show that the BEAST algorithm imposes an average improvement of 5. 14% and 12. 25% for the root-mean-square error and Nash-Sutcliffe efficiency coefficient respectively on the standalone models and shares a comprehensive similar performance on the mean absolute percentage error. And the nonparametric test results reveal that the BEAST method shows a significant improvement on the comprehensive performance compared with a conventional decomposition method. By contrast, the differences between machine learning models are much smaller. The hybrid kernel function works well in some specific cases in which the standalone kernel function fails. The hybrid BEAST-HKESVR is reliable enough to rank the second place among the fifteen tested models. Finally, the effects of hyperparameters in the BEAST algorithm are discussed and relevant suggestions on them are provided.

EAAI Journal 2025 Journal Article

Enhanced neural-network-based iterative learning control considering iterative uncertainties for piezoelectric actuated micro-positioning platform

  • Miaolei Zhou
  • Yulong Sun
  • Xiuyu Zhang
  • Wei Gao
  • Chun-Yi Su

This research aims to developing a new enhanced data-driven sliding-mode iterative learning control (E-DDSILC) strategy for piezoelectric actuated micro-positioning (PAMP) platforms. For the first time, the analysis demonstrating that errors converge to 0 in E-DDSILC is successfully extended from strictly repetitive systems to systems with non-strictly repetitive initial conditions. This generalization expands the practical application range of E-DDSILC. Simultaneously, iterative uncertainties are considered, which are the major factor affecting the performance of iterative learning control. To address these uncertainties, a diagonal recurrent neural network is employed to fit and compensate for them within a dynamic linearization model, thereby further enhancing the tracking accuracy and practicability of E-DDSILC. Finally, Several experiments are performed on a PAMP platform to compare the developed E-DDSILC method with both classical DDSILC and traditional E-DDSILC schemes. Comparative experimental results prove the superiority of the developed controller.

IROS Conference 2025 Conference Paper

High-Precision Transformer-Based Visual Servoing for Humanoid Robots in Aligning Tiny Objects

  • Jialong Xue
  • Wei Gao
  • Yu Wang
  • Chao Ji
  • Dongdong Zhao
  • Shi Yan
  • Shiwu Zhang

High-precision tiny object alignment remains a common and critical challenge for humanoid robots in real world. To address this problem, this paper proposes a vision-based framework for precisely estimating and controlling the relative position between a handheld tool and a target object for humanoid robots, e. g. , a screwdriver tip and a screw head slot. By fusing images from the head and torso cameras on a robot with its head joint angles, the proposed Transformer-based visual servoing method can correct the handheld tool’s positional errors effectively, especially at a close distance. Experiments on M4-M8 screws demonstrate an average convergence error of 0. 8-1. 3 mm and a success rate of 93%-100%. Through comparative analysis, the results validate that this capability of high-precision tiny object alignment is enabled by the Distance Estimation Transformer architecture and the Multi-Perception-Head mechanism proposed in this paper.

EAAI Journal 2025 Journal Article

HMKRec: Optimize multi-user representation by hypergraph motifs for knowledge-aware recommendation

  • Di Wu
  • Mingjing Tang
  • Shu Zhang
  • Wei Gao

Knowledge graph-based recommender systems can explore users’ potential interests by learning user similarities, thereby further improving recommendation performance. However, existing methods focus only on the similarity between two users without considering the interaction patterns among multiple users, which overlook the influence of other users in user representation modeling. In this paper, we propose a novel framework using Hypergraph Motifs to optimize Multi-users representation for Recommendation (HMKRec). Specifically, HMKRec constructs a user–item hypergraph and maps it into a user–user adjacency graph. Then, it utilizes hypergraph motifs to model the interaction patterns of multiple users and reconstructs an implicit relationship network with weights and directions to explore high-order associations among multiple users. To learn the features of items and user relationships, we design a hierarchical graph convolution that integrates hypergraph convolutional networks and graph convolutional networks to obtain high-order representations of users. Finally, we propagate user preferences in the knowledge graph using the attention mechanism to obtain high-order representations of items for recommendation. Extensive experiments on three real-world datasets indicate that our method achieves at least a 1% performance improvement over the best-performing state-of-the-art baselines.

TMLR Journal 2025 Journal Article

LightTransfer: Your Long-Context LLM is Secretly a Hybrid Model with Effortless Adaptation

  • Xuan Zhang
  • Fengzhuo Zhang
  • Cunxiao Du
  • Chao Du
  • Tianyu Pang
  • Wei Gao
  • Min Lin

Scaling language models to handle longer contexts introduces substantial memory challenges due to the growing cost of key-value (KV) caches. Motivated by the efficiency gains of hybrid models and the broad availability of pretrained large transformer backbones, we explore transitioning transformer models into hybrid architectures for a more efficient generation. In this work, we propose \textsc{LightTransfer}, a lightweight method that transforms models such as LLaMA into hybrid variants. Our approach identifies \textit{lazy} layers---those focusing on recent or initial tokens---and replaces their full attention with streaming attention. This transformation can be performed without any training for long-context understanding tasks or with minimal fine-tuning for o1-like long reasoning generation tasks that require stronger reasoning capabilities. Experiments across diverse benchmarks and models (e.g., LLaMA, Mistral, QwQ-STILL) demonstrate that, even when half of the layers are identified as \textit{lazy}, \textsc{LightTransfer} achieves up to 2.17$\times$ throughput improvement with minimal performance loss ($<1.5\%$ on LongBench) and achieves 53.3\% on math benchmark AIME24 of advanced o1-like long reasoning model QwQ-STILL.

TIST Journal 2025 Journal Article

LLM-enhanced Multiple Instance Learning for Joint Rumor and Stance Detection with Social Context Information

  • Ruichao Yang
  • Jing Ma
  • Wei Gao
  • Hongzhan Lin

The proliferation of misinformation, such as rumors on social media, has drawn significant attention, prompting various expressions of stance among users. Although rumor detection and stance detection are distinct tasks, they can complement each other. Rumors can be identified by cross-referencing stances in related posts, and stances are influenced by the nature of the rumor. However, existing stance detection methods often require post-level stance annotations, which are costly to obtain. We propose a novel LLM-enhanced MIL approach to jointly predict post stance and claim class labels, supervised solely by claim labels, using an undirected microblog propagation model. Our weakly supervised approach relies only on bag-level labels of claim veracity, aligning with multi-instance learning (MIL) principles. To achieve this, we transform the multi-class problem into multiple MIL-based binary classification problems. We then employ a discriminative attention layer to aggregate the outputs from these classifiers into finer-grained classes. Experiments conducted on three rumor datasets and two stance datasets demonstrate the effectiveness of our approach, highlighting strong connections between rumor veracity and expressed stances in responding posts. Our method shows promising performance in joint rumor and stance detection compared to the state-of-the-art methods.

IJCAI Conference 2025 Conference Paper

On the Learning with Augmented Class via Forests

  • Fan Xu
  • Wuyang Chen
  • Wei Gao

Decision trees and forests have achieved successes in various real applications, most working with all testing classes known in training data. In this work, we focus on learning with augmented class via forests, where an augmented class may appear in testing data yet not in training data. We incorporate information of augmented class into trees' splitting, that is, augmented Gini impurity, a new splitting criterion is introduced to exploit some unlabeled data from testing distribution. We then develop the Learning with Augmented Class via Forests (short for LACForest) approach, which constructs shallow forests according to the augmented Gini impurity and then splits forests with pseudo-labeled augmented instances for better performance. We also develop deep neural forests via an optimization objective based on our augmented Gini impurity, which essentially utilizes the representation power of neural networks for forests. Theoretically, we present the convergence analysis for our augmented Gini impurity, and we finally conduct experiments to evaluate our approaches. The code is available at https: //github. com/nju-xuf/LACForest.

AAAI Conference 2025 Conference Paper

Point Cloud Semantic Segmentation with Sparse and Inhomogeneous Annotations

  • Zhiyi Pan
  • Nan Zhang
  • Wei Gao
  • Shan Liu
  • Ge Li

Utilizing uniformly distributed sparse annotations, weakly supervised learning alleviates the heavy reliance on fine-grained annotations in point cloud semantic segmentation tasks. However, few works discuss the inhomogeneity of sparse annotations, albeit it is common in real-world scenarios. Therefore, this work introduces the probability density function into the gradient sampling approximation method to qualitatively analyze the impact of annotation sparsity and inhomogeneity under weakly supervised learning. Based on our analysis, we propose an Adaptive Annotation Distribution Network (AADNet) capable of robust learning on arbitrarily distributed sparse annotations. Specifically, we propose a label-aware point cloud downsampling strategy to increase the proportion of annotations involved in the training stage. Furthermore, we design the multiplicative dynamic entropy as the gradient calibration function to mitigate the gradient bias caused by non-uniformly distributed sparse annotations and explicitly reduce the epistemic uncertainty. Without any prior restrictions and additional information, our proposed method achieves comprehensive performance improvements at multiple label rates and different annotation distributions.

JBHI Journal 2025 Journal Article

PointCHD: A Point Cloud Benchmark for Congenital Heart Disease Classification and Segmentation

  • Dinghao Yang
  • Wei Gao

Congenital heart disease (CHD) is one of the most common birth defects. Due to the lack of data and the difficulty of labeling, CHD datasets are scarce. Previous studies focused on CT and other medical image modalities, while point cloud is still unexplored. Point cloud can intuitively model organ shapes, which has obvious advantages in medical analysis and diagnosis assistance. However, the production of medical point cloud dataset is more complex than that of image dataset, and the 3D modeling of internal organs needs to be reconstructed after scanning by high-precision instruments. We propose PointCHD, the first point cloud dataset for CHD diagnosis, with a large number of high precision-annotated and wide-categorized data. PointCHD includes different types of three-dimensional data with varying degrees of distortion, and supports multiple analysis tasks, i. e. , classification, segmentation, reconstruction, etc. We also construct a benchmark on PointCHD with the goal of medical diagnosis, we design the analysis process and compare the performances of mainstream point cloud analysis methods. In view of the complex internal and external structures of heart point cloud, we propose a point cloud representation method based on manifold learning. By introducing normals to consider the surface continuity to construct a manifold learning method of adaptive projection plane, we can fully extract the structural features of heart, and achieve the best performance on each task of PointCHD benchmark. Finally, we summarize the existing problems of CHD point cloud analysis and prospects for potential future research directions.

IJCAI Conference 2025 Conference Paper

Stochasticity-aware No-Reference Point Cloud Quality Assessment

  • Songlin Fan
  • Wei Gao
  • Zhineng Chen
  • Ge Li
  • Guoqing Liu
  • Qicheng Wang

The evolution of point cloud processing algorithms necessitates an accurate assessment for their quality. Previous works consistently regard point cloud quality assessment (PCQA) as a MOS regression problem and devise a deterministic mapping, ignoring the stochasticity in generating MOS from subjective tests. This work presents the first probabilistic architecture for no-reference PCQA, motivated by the labeling process of existing datasets. The proposed method can model the quality judging stochasticity of subjects through a tailored conditional variational autoencoder (CVAE) and produces multiple intermediate quality ratings. These intermediate ratings simulate the judgments from different subjects and are then integrated into an accurate quality prediction, mimicking the generation process of a ground truth MOS. Specifically, our method incorporates a Prior Module, a Posterior Module, and a Quality Rating Generator, where the former two modules are introduced to model the judging stochasticity in subjective tests, while the latter is developed to generate diverse quality ratings. Extensive experiments indicate that our approach outperforms previous cutting-edge methods by a large margin and exhibits gratifying crossdataset robustness. Codes are available at https: //git. openi. org. cn/OpenPointCloud/nrpcqa.

AAAI Conference 2025 Conference Paper

Tackling Intertwined Data and Device Heterogeneities in Federated Learning with Unlimited Staleness

  • Haoming Wang
  • Wei Gao

Federated Learning (FL) can be affected by data and device heterogeneities, caused by clients' different local data distributions and latencies in uploading model updates (i.e., staleness). Traditional schemes consider these heterogeneities as two separate and independent aspects, but this assumption is unrealistic in practical FL scenarios where these heterogeneities are intertwined. In these cases, traditional FL schemes are ineffective, and a better approach is to convert a stale model update into a unstale one. In this paper, we present a new FL framework that ensures the accuracy and computational efficiency of this conversion, hence effectively tackling the intertwined heterogeneities that may cause unlimited staleness in model updates. Our basic idea is to estimate the distributions of clients' local training data from their uploaded stale model updates, and use these estimations to compute unstale client model updates. In this way, our approach does not require any auxiliary dataset nor the clients' local models to be fully trained, and does not incur any additional computation or communication overhead at client devices. We compared our approach with the existing FL strategies on mainstream datasets and models, and showed that our approach can improve the trained model accuracy by up to 25% and reduce the number of required training epochs by up to 35%.

AAAI Conference 2025 Conference Paper

UniPCGC: Towards Practical Point Cloud Geometry Compression via an Efficient Unified Approach

  • Kangli Wang
  • Wei Gao

Learning-based point cloud compression methods have made significant progress in terms of performance. However, these methods still encounter challenges including high complexity, limited compression modes, and a lack of support for variable rate, which restrict the practical application of these methods. In order to promote the development of practical point cloud compression, we propose an efficient unified point cloud geometry compression framework, dubbed as UniPCGC. It is a lightweight framework that supports lossy compression, lossless compression, variable rate and variable complexity. First, we introduce the Uneven 8-Stage Lossless Coder (UELC) in the lossless mode, which allocates more computational complexity to groups with higher coding difficulty, and merges groups with lower coding difficulty. Second, Variable Rate and Complexity Module (VRCM) is achieved in the lossy mode through joint adoption of a rate modulation module and dynamic sparse convolution. Finally, through the dynamic combination of UELC and VRCM, we achieve lossy compression, lossless compression, variable rate and complexity within a unified framework. Compared to the previous state-of-the-art method, our method achieves a compression ratio (CR) gain of 8.1% on lossless compression, and a Bjontegaard Delta Rate (BD-Rate) gain of 14.02% on lossy compression, while also supporting variable rate and variable complexity.

AAAI Conference 2025 Conference Paper

VE-Bench: Subjective-Aligned Benchmark Suite for Text-Driven Video Editing Quality Assessment

  • Shangkun Sun
  • Xiaoyu Liang
  • Songlin Fan
  • Wenxu Gao
  • Wei Gao

Text-driven video editing has recently experienced rapid development. Despite this, evaluating edited videos remains a considerable challenge. Current metrics tend to fail to align with human perceptions, and effective quantitative metrics for video editing are still notably absent. To address this, we introduce VE-Bench, a benchmark suite tailored to the assessment of text-driven video editing. This suite includes VE-Bench DB, a video quality assessment (VQA) database for video editing. VE-Bench DB encompasses a diverse set of source videos featuring various motions and subjects, along with multiple distinct editing prompts, editing results from 8 different models, and the corresponding Mean Opinion Scores (MOS) from 24 human annotators. Based on VE-Bench DB, we further propose VE-Bench QA, a quantitative human-aligned measurement for the text-driven video editing task. In addition to the aesthetic, distortion, and other visual quality indicators that traditional VQA methods emphasize, VE-Bench QA focuses on the text-video alignment and the relevance modeling between source and edited videos. It introduces a new assessment network for video editing that attains superior performance in alignment with human preferences.To the best of our knowledge, VE-Bench introduces the first quality assessment dataset for video editing and proposes an effective subjective-aligned quantitative metric for this domain. All models, data, and code will be publicly available to the community.

IROS Conference 2024 Conference Paper

Active Loop Closure for OSM-guided Robotic Mapping in Large-Scale Urban Environments

  • Wei Gao
  • Zezhou Sun
  • Mingle Zhao
  • ChengZhong Xu 0001
  • Hui Kong 0001

The autonomous mapping of large-scale urban scenes presents significant challenges for autonomous robots. To mitigate the challenges, global planning, such as utilizing prior GPS trajectories from OpenStreetMap (OSM), is often used to guide the autonomous navigation of robots for mapping. However, due to factors like complex terrain, unexpected body movement, and sensor noise, the uncertainty of the robot’s pose estimates inevitably increases over time, ultimately leading to the failure of robotic mapping. To address this issue, we propose a novel active loop closure procedure, enabling the robot to actively re-plan the previously planned GPS trajectory. The method can guide the robot to re-visit the previous places where the loop-closure detection can be performed to trigger the back-end optimization, effectively reducing errors and uncertainties in pose estimation. The proposed active loop closure mechanism is implemented and embedded into a real-time OSM-guided robot mapping framework. Empirical results on several large-scale outdoor scenarios demonstrate its effectiveness and promising performance.

NeurIPS Conference 2024 Conference Paper

Chain of Preference Optimization: Improving Chain-of-Thought Reasoning in LLMs

  • Xuan Zhang
  • Chao Du
  • Tianyu Pang
  • Qian Liu
  • Wei Gao
  • Min Lin

The recent development of chain-of-thought (CoT) decoding has enabled large language models (LLMs) to generate explicit logical reasoning paths for complex problem-solving. However, research indicates that these paths are not always deliberate and optimal. The tree-of-thought (ToT) method employs tree-searching to extensively explore the reasoning space and find better reasoning paths that CoT decoding might overlook. This deliberation, however, comes at the cost of significantly increased inference complexity. In this work, we demonstrate that fine-tuning LLMs leveraging the search tree constructed by ToT allows CoT to achieve similar or better performance, thereby avoiding the substantial inference burden. This is achieved through \emph{Chain of Preference Optimization} (CPO), where LLMs are fine-tuned to align each step of the CoT reasoning paths with those of ToT using the inherent preference information in the tree-search process. Extensive experimental results show that CPO significantly improves LLM performance in solving a variety of complex problems, including question answering, fact verification, and arithmetic reasoning, demonstrating its effectiveness. Our code is available at https: //github. com/sail-sg/CPO.

NeurIPS Conference 2024 Conference Paper

Distribution Guidance Network for Weakly Supervised Point Cloud Semantic Segmentation

  • Zhiyi Pan
  • Wei Gao
  • Shan Liu
  • Ge Li

Despite alleviating the dependence on dense annotations inherent to fully supervised methods, weakly supervised point cloud semantic segmentation suffers from inadequate supervision signals. In response to this challenge, we introduce a novel perspective that imparts auxiliary constraints by regulating the feature space under weak supervision. Our initial investigation identifies which distributions accurately characterize the feature space, subsequently leveraging this priori to guide the alignment of the weakly supervised embeddings. Specifically, we analyze the superiority of the mixture of von Mises-Fisher distributions (moVMF) among several common distribution candidates. Accordingly, we develop a Distribution Guidance Network (DGNet), which comprises a weakly supervised learning branch and a distribution alignment branch. Leveraging reliable clustering initialization derived from the weakly supervised learning branch, the distribution alignment branch alternately updates the parameters of the moVMF and the network, ensuring alignment with the moVMF-defined latent space. Extensive experiments validate the rationality and effectiveness of our distribution choice and network design. Consequently, DGNet achieves state-of-the-art performance under multiple datasets and various weakly supervised settings.

IROS Conference 2024 Conference Paper

Efficient Incremental Penetration Depth Estimation between Convex Geometries

  • Wei Gao

Penetration depth (PD) is essential for robotics due to its extensive applications in dynamic simulation, motion planning, haptic rendering, etc. The Expanding Polytope Algorithm (EPA) is the de facto standard for this problem, which estimates PD by expanding an inner polyhedral approximation of an implicit set. In this paper, we propose a novel optimization-based algorithm that incrementally estimates minimum penetration depth and its direction. One major advantage of our method is the capability to be warm-started by leveraging the spatial and temporal coherence. This coherence emerges naturally in many robotic applications (e. g. , the temporal coherence between adjacent simulation time knots). As a result, our algorithm achieves substantial speedup — we demonstrate it is 5-30x faster than EPA on several benchmarks. Moreover, our approach is built upon the same implicit geometry representation as EPA, which enables easy integration into existing software stacks. The code and supplemental document are available on: https://github.com/weigao95/mind-fcl.

AAAI Conference 2024 Conference Paper

End-to-End RGB-D Image Compression via Exploiting Channel-Modality Redundancy

  • Huiming Zheng
  • Wei Gao

As a kind of 3D data, RGB-D images have been extensively used in object tracking, 3D reconstruction, remote sensing mapping, and other tasks. In the realm of computer vision, the significance of RGB-D images is progressively growing. However, the existing learning-based image compression methods usually process RGB images and depth images separately, which cannot entirely exploit the redundant information between the modalities, limiting the further improvement of the Rate-Distortion performance. With the goal of overcoming the defect, in this paper, we propose a learning-based dual-branch RGB-D image compression framework. Compared with traditional RGB domain compression scheme, a YUV domain compression scheme is presented for spatial redundancy removal. In addition, Intra-Modality Attention (IMA) and Cross-Modality Attention (CMA) are introduced for modal redundancy removal. For the sake of benefiting from cross-modal prior information, Context Prediction Module (CPM) and Context Fusion Module (CFM) are raised in the conditional entropy model which makes the context probability prediction more accurate. The experimental results demonstrate our method outperforms existing image compression methods in two RGB-D image datasets. Compared with BPG, our proposed framework can achieve up to 15% bit rate saving for RGB images.

AAAI Conference 2024 Conference Paper

Fast Inter-frame Motion Prediction for Compressed Dynamic Point Cloud Attribute Enhancement

  • Wang Liu
  • Wei Gao
  • Xingming Mu

Recent years have witnessed the success of deep learning methods in quality enhancement of compressed point cloud. However, existing methods focus on geometry and attribute enhancement of single-frame point cloud. This paper proposes a novel compressed quality enhancement method for dynamic point cloud (DAE-MP). Specifically, we propose a fast inter-frame motion prediction module (IFMP) to explicitly estimate motion displacement and achieve inter-frame feature alignment. To maintain motion continuity between consecutive frames, we propose a motion consistency loss for supervised learning. Furthermore, a frequency component separation and fusion module is designed to extract rich frequency features adaptively. To the best of our knowledge, the proposed method is the first deep learning-based work to enhance the quality for compressed dynamic point cloud. Experimental results show that the proposed method can greatly improve the quality of compressed dynamic point cloud and provide a fast and efficient motion prediction plug-in for large-scale point cloud. For dynamic point cloud attribute with severely compressed artifact, our proposed DAE-MP method achieves up to 0.52dB (PSNR) performance gain. Moreover, the proposed IFMP module has a certain real-time processing ability for calculating the motion offset between dynamic point cloud frame.

AAAI Conference 2024 Conference Paper

Less Is More: Label Recommendation for Weakly Supervised Point Cloud Semantic Segmentation

  • Zhiyi Pan
  • Nan Zhang
  • Wei Gao
  • Shan Liu
  • Ge Li

Weak supervision has proven to be an effective strategy for reducing the burden of annotating semantic segmentation tasks in 3D space. However, unconstrained or heuristic weakly supervised annotation forms may lead to suboptimal label efficiency. To address this issue, we propose a novel label recommendation framework for weakly supervised point cloud semantic segmentation. Distinct from pre-training and active learning, the label recommendation framework consists of three stages: inductive bias learning, recommendations for points to be labeled, and point cloud semantic segmentation learning. In practice, we first introduce the point cloud upsampling task to induct inductive bias from structural information. During the recommendation stage, we present a cross-scene clustering strategy to generate centers of clustering as recommended points. Then we introduce a recommended point positions attention module LabelAttention to model the long-range dependency under sparse annotations. Additionally, we employ position encoding to enhance the spatial awareness of semantic features. Throughout the framework, the useful information obtained from inductive bias learning is propagated to subsequent semantic segmentation networks in the form of label positions. Experimental results demonstrate that our framework outperforms weakly supervised point cloud semantic segmentation methods and other methods for labeling efficiency on S3DIS and ScanNetV2, even at an extremely low label rate.

EAAI Journal 2024 Journal Article

Multi-degree-of-freedom unmanned aerial vehicle control combining a hybrid brain-computer interface and visual obstacle avoidance

  • Shanghong Xie
  • Wei Gao
  • Zhen Zeng
  • Qingfu Wu
  • Qian Huang
  • Nianming Ban
  • Qian Wu
  • Jiahui Pan

Objective The difficulty of unmanned aerial vehicle (UAV) control recently lies in multidirectional movement in 3-dimensional space, improving control accuracy and manipulation safety. To address these challenges, a UAV control system that incorporates a hybrid brain-computer interface (hBCI), gyroscope and visual obstacle avoidance based on monocular depth estimation is proposed. Approach. We propose an efficient steady-state visual evoked potential (SSVEP) classification network (CL-NET) featuring a one-dimensional convolutional neural network, a long short-term memory module and an attention module to identify the user's intention for UAV movement in the front, back, left and right directions. The take-off, landing and rising control of the UAV is realized by an electrooculogram (EOG) signal detection algorithm, a blink state detector. In addition, the UAV can fly in an oblique state and rotate according to the current head posture detected by a gyroscope. Furthermore, an improved monocular depth estimation network is employed to design the autonomous obstacle avoidance module of the UAV, ensuring the safety of the brain-controlled system in practice. Main results. The proposed CL-NET delivers an accuracy of 98. 67% on the public dataset and an accuracy of 97. 92% on the self-collected dataset, both of which surpass the performance of state-of-the-art models. Additionally, we set up a brain control group and a remote control group to conduct practical experiments in a realistic environment. In the experiments involving sixteen subjects, the proposed UAV control system reached an average information transfer rate (ITR) of 44. 09 bits/min, and the brain control group had a lower collision rate than the remote control group. Significance. The hybrid control method ensures that the multi-degree-of-freedom (multi-DOF) UAV control system maintains outstanding performance while ensuring good safety.

EAAI Journal 2024 Journal Article

Rank-based multimodal immune algorithm for many-objective optimization problems

  • Hainan Zhang
  • Jianhou Gan
  • Juxiang Zhou
  • Wei Gao

The immune algorithm (IA) is a prestigious heuristic algorithm based on a model of an artificial immune system, and the IA has shown promising results in the multi-objective optimization field. However, the algorithm’s low search ability in high-dimensional space and the clone assignment metric problem must be addressed. Thus, to solve these problems, we propose a rank-based multimodal immune algorithm (RMIA) for many-objective optimization problems. To alleviate the clone assignment metric problem, we design a novel vaccine selection mechanism, which is a rank-based clone selection method. We also propose a dynamic age-based elimination mechanism and a multimodal mutation strategy to address the poor searching ability of the IA in high-dimensional space, where the former is eliminated randomly via roulette in terms of the survival time and the advantages of antibodies in the population, and the latter adopts different mutation strategies based on the different states of antibodies. The proposed algorithm was evaluated and compared to multiple advanced multi-objective optimization immune algorithms (MOIAs) and many-objective optimization evolutionary algorithms (MaOEAs) to demonstrate its superiority. The code is available at https: //github. com/AizhEngHN/RMIA.

NeurIPS Conference 2024 Conference Paper

StreamFlow: Streamlined Multi-Frame Optical Flow Estimation for Video Sequences

  • Shangkun Sun
  • Jiaming Liu
  • Huaxia Li
  • Guoqing Liu
  • Thomas H. Li
  • Wei Gao

Prior multi-frame optical flow methods typically estimate flow repeatedly in a pair-wise manner, leading to significant computational redundancy. To mitigate this, we implement a Streamlined In-batch Multi-frame (SIM) pipeline, specifically tailored to video inputs to minimize redundant calculations. It enables the simultaneous prediction of successive unidirectional flows in a single forward pass, boosting processing speed by 44. 43% and reaching efficiencies on par with two-frame networks. Moreover, we investigate various spatiotemporal modeling methods for optical flow estimation within this pipeline. Notably, we propose a simple yet highly effective parameter-efficient Integrative spatiotemporal Coherence (ISC) modeling method, alongside a lightweight Global Temporal Regressor (GTR) to harness temporal cues. The proposed ISC and GTR bring powerful spatiotemporal modeling capabilities and significantly enhance accuracy, including in occluded areas, while adding modest computations to the SIM pipeline. Compared to the baseline, our approach, StreamFlow, achieves performance enhancements of 15. 45% and 11. 37% on the Sintel clean and final test sets respectively, with gains of 15. 53% and 10. 77% on occluded regions and only a 1. 11% rise in latency. Furthermore, StreamFlow exhibits state-of-the-art cross-dataset testing results on Sintel and KITTI, demonstrating its robust cross-domain generalization capabilities. The code is available here.

EAAI Journal 2023 Journal Article

Chaotic heterogeneous comprehensive learning PSO method for size and shape optimization of structures

  • Thu Huynh Van
  • Sawekchai Tangaramvong
  • Wei Gao

This paper proposes a novel chaotic heterogeneous comprehensive learning particle swarm optimization (CLPSO) method for the simultaneous size and shape design of structures. The heterogeneous CLPSO divides the particles into two explorative and exploitative subpopulations. The exploration performs the global searches for the set of best particles experienced solely within its own subpopulation, whilst the exploitation refines the deep searches learnt from the global best particle over an entire population. In essence, the proposed method maintains a good balance between the global explorative and local exploitative optimization schemes. The global searches within the explorative subpopulation are independent to the exploitative simulations even if the latter scheme prematurely converges to the local swarm position. For both subpopulations, the comprehensive learning approach constructs the particles through a cross-positioning process on the individual variable space and avoids the local optimal pitfall. Moreover, the chaotic logistic map within the exploitative optimization tests the global best particle through the set of diversely generated samples and hence enhances the local search ability. Various enriching techniques, including automatic adaptive (inertial weight and acceleration) parameters with dynamic space reduction, are incorporated to improve the likelihood of finding the optimal solution of practical-scale problems at modest computing efforts. The accuracy and robustness of the proposed method are illustrated through a number of planar and spatial truss design benchmarks subjected to the challenging nonconvex and/or nonsmooth programs.

EAAI Journal 2023 Journal Article

Non trust detection of decentralized federated learning based on historical gradient

  • Yikuan Chen
  • Li Liang
  • Wei Gao

As a paradigm of distributed machine learning, federated learning is widely used in various real scenarios due to its excellent privacy protection performance on preventing local data from being disclosed. However, the traditional federated learning has the defect that a third-party server aggregates the models of various users since it’s difficult to guarantee the reliability of the third party, and multicentre phenomena frequently appeared in various applications, such as social networks, banking and finance, medical health, etc. Users can’t be reassured in decentralization setting due to the mixture of malicious and untrustworthy ones among them. Although untrustworthy users are benign, they may be classified as the saboteurs because of poor efficiency performance in decentralized federated learning which is caused by missing or ambiguity of data. In this paper, we propose Decentralized Federated Learning Historical Gradient (DFedHG) approach to distinguish normal users, untrustworthy users and malicious users in the decentralized federated learning setting. Simultaneously, by means of DFedHG, malicious users are sub-divided into targetless attacks and targeted attacks, which is verified by adopting two types of data sets for confirmation. The experimental results show that the proposed approach achieves better performance compared with the conventional decentralized federated learning without untrustworthy users, and further present excellent differentiation of malicious users.

NeurIPS Conference 2023 Conference Paper

On the Exploration of Local Significant Differences For Two-Sample Test

  • Zhijian Zhou
  • Jie Ni
  • Jia-He Yao
  • Wei Gao

Recent years have witnessed increasing attentions on two-sample test with diverse real applications, while this work takes one more step on the exploration of local significant differences for two-sample test. We propose the ME$_\text{MaBiD}$, an effective test for two-sample testing, and the basic idea is to exploit local information by multiple Mahalanobis kernels and introduce bi-directional hypothesis for testing. On the exploration of local significant differences, we first partition the embedding space into several rectangle regions via a new splitting criterion, which is relevant to test power and data correlation. We then explore local significant differences based on our bi-directional masked $p$-value together with the ME$_\text{MaBiD}$ test. Theoretically, we present the asymptotic distribution and lower bounds of test power for our ME$_\text{MaBiD}$ test, and control the familywise error rate on the exploration of local significant differences. We finally conduct extensive experiments to validate the effectiveness of our proposed methods on two-sample test and the exploration of local significant differences.

NeurIPS Conference 2023 Conference Paper

On the Gini-impurity Preservation For Privacy Random Forests

  • XinRan Xie
  • Man-Jie Yuan
  • Xuetong Bai
  • Wei Gao
  • Zhi-Hua Zhou

Random forests have been one successful ensemble algorithms in machine learning. Various techniques have been utilized to preserve the privacy of random forests from anonymization, differential privacy, homomorphic encryption, etc. , whereas it rarely takes into account some crucial ingredients of learning algorithm. This work presents a new encryption to preserve data's Gini impurity, which plays a crucial role during the construction of random forests. Our basic idea is to modify the structure of binary search tree to store several examples in each node, and encrypt data features by incorporating label and order information. Theoretically, we prove that our scheme preserves the minimum Gini impurity in ciphertexts without decrypting, and present the security guarantee for encryption. For random forests, we encrypt data features based on our Gini-impurity-preserving scheme, and take the homomorphic encryption scheme CKKS to encrypt data labels due to their importance and privacy. We conduct extensive experiments to show the effectiveness, efficiency and security of our proposed method.

ICRA Conference 2022 Conference Paper

Deep Visual Navigation under Partial Observability

  • Bo Ai 0004
  • Wei Gao
  • Vinay
  • David Hsu

How can a robot navigate successfully in rich and diverse environments, indoors or outdoors, along office corridors or trails on the grassland, on the flat ground or the staircase? To this end, this work aims to address three challenges: (i) complex visual observations, (ii) partial observability of local visual sensing, and (iii) multimodal robot behaviors conditioned on both the local environment and the global navigation objective. We propose to train a neural network (NN) controller for local navigation via imitation learning. To tackle complex visual observations, we extract multi-scale spatial representations through CNNs. To tackle partial observability, we aggregate multi-scale spatial information over time and encode it in LSTMs. To learn multimodal behaviors, we use a separate memory module for each behavior mode. Importantly, we integrate the multiple neural network modules into a unified controller that achieves robust performance for visual navigation in complex, partially observable environments. We implemented the controller on the quadrupedal Spot robot and evaluated it on three challenging tasks: adversarial pedestrian avoidance, blind-spot obstacle avoidance, and elevator riding. The experiments show that the proposed NN architecture significantly improves navigation performance.

EAAI Journal 2022 Journal Article

Life prediction of underground structure by sulfate corrosion using Harris hawks optimizing genetic programming

  • Yuan Xie
  • Wei Gao
  • Yiwei Wang
  • Xin Chen
  • Shuangshuang Ge
  • Sen Wang

A corrosive sulfate environment can cause strong deterioration and destruction of reinforced concrete (RC) underground structures and seriously reduce their service life. Thus, it is very important to predict the service life of RC underground structures in corrosive sulfate environments. However, the service life of underground structures is affected by numerous complicated engineering and environmental factors and cannot be determined by traditional theoretical and experimental investigations. Therefore, to solve this problem, a new data-driven method based on Harris hawks optimizing genetic programming (HHO-GP) is proposed. In this new method, to improve the traditional genetic programming (GP), a new global optimization algorithm called Harris hawks optimization (HHO) is adopted to optimize its main controlling parameters. Based on 25 groups of real engineering data, the life prediction model of underground structures in corrosive sulfate environments with 12 main engineering and environmental influence factors is established by the HHO-GP method. The results show that the average relative training error (5. 5%) and predicting error (6. 3%) of the new prediction model are small. Therefore, the proposed HHO-GP method can construct a suitable life prediction model based on only real engineering data, regardless of how many complicated influencing factors are considered. Moreover, our data-driven life prediction model is described by one explicit polynomial function based on 12 influencing factors. Thus, it can be applied in real engineering simply and easily. Finally, the influence of the main controlling parameters of the HHO-GP on its accuracy and efficiency is analyzed. The results reveal that considering the computing accuracy and efficiency and the model completeness, the small population size and maximum iterations of HHO are suitable, whose recommended values are all 15. The population size and maximum number of iterations of GP have little influence on the prediction accuracy. Their recommended values all can be 50.

AAAI Conference 2022 Conference Paper

OctAttention: Octree-Based Large-Scale Contexts Model for Point Cloud Compression

  • Chunyang Fu
  • Ge Li
  • Rui Song
  • Wei Gao
  • Shan Liu

In point cloud compression, sufficient contexts are significant for modeling the point cloud distribution. However, the contexts gathered by the previous voxel-based methods decrease when handling sparse point clouds. To address this problem, we propose a multiple-contexts deep learning framework called OctAttention employing the octree structure, a memory-efficient representation for point clouds. Our approach encodes octree symbol sequences in a lossless way by gathering the information of sibling and ancestor nodes. Expressly, we first represent point clouds with octree to reduce spatial redundancy, which is robust for point clouds with different resolutions. We then design a conditional entropy model with a large receptive field that models the sibling and ancestor contexts to exploit the strong dependency among the neighboring nodes and employ an attention mechanism to emphasize the correlated nodes in the context. Furthermore, we introduce a mask operation during training and testing to make a trade-off between encoding time and performance. Compared to the previous state-of-the-art works, our approach obtains a 10%-35% BD-Rate gain on the LiDAR benchmark (e. g. SemanticKITTI) and object point cloud dataset (e. g. MPEG 8i, MVUB), and saves 95% coding time compared to the voxel-based baseline. The code is available at https: //github. com/zb12138/OctAttention.

IJCAI Conference 2022 Conference Paper

On the Optimization of Margin Distribution

  • Meng-Zhang Qian
  • Zheng Ai
  • Teng Zhang
  • Wei Gao

Margin has played an important role on the design and analysis of learning algorithms during the past years, mostly working with the maximization of the minimum margin. Recent years have witnessed the increasing empirical studies on the optimization of margin distribution according to different statistics such as medium margin, average margin, margin variance, etc. , whereas there is a relative paucity of theoretical understanding. In this work, we take one step on this direction by providing a new generalization error bound, which is heavily relevant to margin distribution by incorporating ingredients such as average margin and semi-variance, a new margin statistics for the characterization of margin distribution. Inspired by the theoretical findings, we propose the MSVMAv, an efficient approach to achieve better performance by optimizing margin distribution in terms of its empirical average margin and semi-variance. We finally conduct extensive experiments to show the superiority of the proposed MSVMAv approach.

NeurIPS Conference 2022 Conference Paper

Receding Horizon Inverse Reinforcement Learning

  • Yiqing Xu
  • Wei Gao
  • David Hsu

Inverse reinforcement learning (IRL) seeks to infer a cost function that explains the underlying goals and preferences of expert demonstrations. This paper presents Receding Horizon Inverse Reinforcement Learning (RHIRL), a new IRL algorithm for high-dimensional, noisy, continuous systems with black-box dynamic models. RHIRL addresses two key challenges of IRL: scalability and robustness. To handle high-dimensional continuous systems, RHIRL matches the induced optimal trajectories with expert demonstrations locally in a receding horizon manner and stitches'' together the local solutions to learn the cost; it thereby avoids the curse of dimensionality''. This contrasts sharply with earlier algorithms that match with expert demonstrations globally over the entire high-dimensional state space. To be robust against imperfect expert demonstrations and control noise, RHIRL learns a state-dependent cost function ``disentangled'' from system dynamics under mild conditions. Experiments on benchmark tasks show that RHIRL outperforms several leading IRL algorithms in most instances. We also prove that the cumulative error of RHIRL grows linearly with the task duration.

AIJ Journal 2022 Journal Article

Towards convergence rate analysis of random forests for classification

  • Wei Gao
  • Fan Xu
  • Zhi-Hua Zhou

Random forests have been one of the successful ensemble algorithms in machine learning, and the basic idea is to construct a large number of random trees individually and make predictions based on an average of their predictions. The great successes have attracted much attention on theoretical understandings of random forests, mostly focusing on regression problems. This work takes one step towards the convergence rates of random forests for classification. We present the first finite-sample rate O ( n − 1 / ( 8 d + 2 ) ) on the convergence of purely random forests for binary classification, which can be improved to be of O ( n − 1 / ( 3. 87 d + 2 ) ) by considering the midpoint splitting mechanism. We introduce another variant of random forests, which follows Breiman's original random forests but with different mechanisms on splitting dimensions and positions. We present the convergence rate O ( n − 1 / ( d + 2 ) ( ln ⁡ n ) 1 / ( d + 2 ) ) for the variant of random forests, which reaches the minimax rate, except for a factor ( ln ⁡ n ) 1 / ( d + 2 ), of the optimal plug-in classifier under the L-Lipschitz assumption. We achieve the tighter convergence rate O ( ln ⁡ n / n ) under some assumptions over structural data. This work also takes one step towards the convergence rate of random forests for multi-class learning, and presents the same convergence rates of random forests for multi-class learning as that of binary classification, yet with different constants. We finally provide empirical studies to support the theoretical analysis.

AAAI Conference 2021 Conference Paper

Evidence Aware Neural Pornographic Text Identification for Child Protection

  • Kaisong Song
  • Yangyang Kang
  • Wei Gao
  • Zhe Gao
  • Changlong Sun
  • Xiaozhong Liu

Identifying pornographic text online is practically useful to protect children from access to such adult content. However, some authors may intentionally avoid using sensitive words in their pornographic texts to take advantage of the lack of human audits. Without prior knowledge guidance, real semantics of such pornographic text is difficult to understand by existing methods due to its high context-sensitivity and heavy usage of figurative language, which brings huge challenges to the porn detection systems used in social media platforms. In this paper, we approach to the problem as a document-level porn identification task by locating and integrating sentencelevel evidence and propose a novel Evidence-Aware Neural Porn Classification (eNPC) model. Specifically, we first propose a basic model which locates porn indicative sentences in the document with a multiple instance learning model, and then aggregate the sentence-level evidence to induce document label with self-attention mechanism. Moreover, we consider label dependencies within local context. Finally, we further enhance the sentence representation with prior knowledge produced by an automatic porn lexicon construction strategy. Extensive experimental results show that our model exhibits consistent superiority over competitors on two realworld Chinese novel datasets and an English story dataset.

YNIMG Journal 2021 Journal Article

Functional coupling of the orbitofrontal cortex and the basolateral amygdala mediates the association between spontaneous reappraisal and emotional response

  • Wei Gao
  • Bharat Biswal
  • ShengDong Chen
  • Xinran Wu
  • Jiajin Yuan

Emotional regulation is known to be associated with activity in the amygdala. The amygdala is an emotion-generative region that comprises of structurally and functionally distinct nuclei. However, little is known about the contributions of different frontal-amygdala sub-region pathways to emotion regulation. Here, we investigated how functional couplings between frontal regions and amygdala sub-regions are involved in different spontaneous emotion regulation processes by using an individual-difference approach and a generalized psycho-physiological interaction (gPPI) approach. Specifically, 50 healthy participants reported their dispositional use of spontaneous cognitive reappraisal and expressive suppression in daily life and their actual use of these two strategies during the performance of an emotional-picture watching task. Results showed that functional coupling between the orbitofrontal cortex (OFC) and the basolateral amygdala (BLA) was associated with higher scores of both dispositional and actual uses of reappraisal. Similarly, functional coupling between the dorsolateral prefrontal cortex (dlPFC) and the centromedial amygdala (CMA) was associated with higher scores of both dispositional and actual uses of suppression. Mediation analyses indicated that functional coupling of the right OFC-BLA partially mediated the association between reappraisal and emotional response, irrespective of whether reappraisal was measured by dispositional use (indirect effect(SE)=-0. 2021 (0. 0811), 95%CI(BC)= [-0. 3851, -0. 0655]) or actual use (indirect effect(SE)=-0. 1951 (0. 0796), 95%CI(BC)= [-0. 3654, -0. 0518])). These findings suggest that spontaneous reappraisal and suppression involve distinct frontal- amygdala functional couplings, and the modulation of BLA activity from OFC may be necessary for changing emotional response during spontaneous reappraisal.

AIJ Journal 2021 Journal Article

On the noise estimation statistics

  • Wei Gao
  • Teng Zhang
  • Bin-Bin Yang
  • Zhi-Hua Zhou

Learning with noisy labels has attracted much attention during the past few decades. A fundamental problem is how to estimate noise proportions from corrupted data. Previous studies on this issue resort to the estimations of class distributions, conditional distributions, or the kernel embedding of distributions. In this paper, we present another simple and effective approach for noise estimation. The basic idea is to utilize the first- and second-order statistics of observed data, and the positive semi-definiteness of covariance matrices. Then, an upper bound on noise estimation is provided without additional assumptions over data distribution. Based on this idea and using the locality property of random noise, we develop the Noise Estimation Statistics with Clusters (NESC) method, which firstly clusters the corrupted data by k-means algorithm, and then makes noise estimation from clusters based on the first- and second-order statistics. We present the existence, uniqueness and convergence analysis of our noise estimation, and empirical studies verify the effectiveness of the NESC method.

TIST Journal 2020 Journal Article

An Attention-based Rumor Detection Model with Tree-structured Recursive Neural Networks

  • Jing Ma
  • Wei Gao
  • Shafiq Joty
  • Kam-Fai Wong

Rumor spread in social media severely jeopardizes the credibility of online content. Thus, automatic debunking of rumors is of great importance to keep social media a healthy environment. While facing a dubious claim, people often dispute its truthfulness sporadically in their posts containing various cues, which can form useful evidence with long-distance dependencies. In this work, we propose to learn discriminative features from microblog posts by following their non-sequential propagation structure and generate more powerful representations for identifying rumors. For modeling non-sequential structure, we first represent the diffusion of microblog posts with propagation trees, which provide valuable clues on how a claim in the original post is transmitted and developed over time. We then present a bottom-up and a top-down tree-structured models based on Recursive Neural Networks (RvNN) for rumor representation learning and classification, which naturally conform to the message propagation process in microblogs. To enhance the rumor representation learning, we reveal that effective rumor detection is highly related to finding evidential posts, e.g., the posts expressing specific attitude towards the veracity of a claim, as an extension of the previous RvNN-based detection models that treat every post equally. For this reason, we design discriminative attention mechanisms for the RvNN-based models to selectively attend on the subset of evidential posts during the bottom-up/top-down recursive composition. Experimental results on four datasets collected from real-world microblog platforms confirm that (1) our RvNN-based models achieve much better rumor detection and classification performance than state-of-the-art approaches; (2) the attention mechanisms for focusing on evidential posts can further improve the performance of our RvNN-based method; and (3) our approach possesses superior capacity on detecting rumors at a very early stage.

AAAI Conference 2020 Conference Paper

AUC Optimization with a Reject Option

  • Song-Qing Shen
  • Bin-Bin Yang
  • Wei Gao

Making an erroneous decision may cause serious results in diverse mission-critical tasks such as medical diagnosis and bioinformatics. Previous work focuses on classification with a reject option, i. e. , abstain rather than classify an instance of low confidence. Most mission-critical tasks are always accompanied with class imbalance and cost sensitivity, where AUC has been shown a preferable measure than accuracy in classification. In this work, we propose the framework of AUC optimization with a reject option, and the basic idea is to withhold the decision of ranking a pair of positive and negative instances with a lower cost, rather than mis-ranking. We obtain the Bayes optimal solution for ranking, and learn the reject function and score function for ranking, simultaneously. An online algorithm has been developed for AUC optimization with a reject option, by considering the convex relaxation and plug-in rule. We verify, both theoretically and empirically, the effectiveness of the proposed algorithm.

NeurIPS Conference 2020 Conference Paper

Towards Convergence Rate Analysis of Random Forests for Classification

  • Wei Gao
  • Zhi-Hua Zhou

Random forests have been one of the successful ensemble algorithms in machine learning. The basic idea is to construct a large number of random trees individually and make prediction based on an average of their predictions. The great successes have attracted much attention on the consistency of random forests, mostly focusing on regression. This work takes one step towards convergence rates of random forests for classification. We present the first finite-sample rate O(n^{-1/(8d+2)}) on the convergence of pure random forests for classification, which can be improved to be of O(n^{-1/(3. 87d+2)}) by considering the midpoint splitting mechanism. We introduce another variant of random forests, which follow Breiman's original random forests but with different mechanisms on splitting dimensions and positions. We get a convergence rate O(n^{-{1}/(d+2)}(\ln n)^{{1}/(d+2)}) for the variant of random forests, which reaches the minimax rate, except for a factor (\ln n)^{{1}/(d+2)}, of the optimal plug-in classifier under the L-Lipschitz assumption. We achieve tighter convergence rate O(\sqrt{\ln n/n}) under proper assumptions over structural data.

YNIMG Journal 2019 Journal Article

A review on neuroimaging studies of genetic and environmental influences on early brain development

  • Wei Gao
  • Karen Grewen
  • Rebecca C. Knickmeyer
  • Anqi Qiu
  • Andrew Salzwedel
  • Weili Lin
  • John H. Gilmore

The past decades witnessed a surge of interest in neuroimaging study of normal and abnormal early brain development. Structural and functional studies of normal early brain development revealed massive structural maturation as well as sequential, coordinated, and hierarchical emergence of functional networks during the infancy period, providing a great foundation for the investigation of abnormal early brain development mechanisms. Indeed, studies of altered brain development associated with either genetic or environmental risks emerged and thrived. In this paper, we will review selected studies of genetic and environmental risks that have been relatively more extensively investigated-familial risks, candidate risk genes, and genome-wide association studies (GWAS) on the genetic side; maternal mood disorders and prenatal drug exposures on the environmental side. Emerging studies on environment-gene interactions will also be reviewed. Our goal was not to perform an exhaustive review of all studies in the field but to leverage some representative ones to summarize the current state, point out potential limitations, and elicit discussions on important future directions.

IJCAI Conference 2019 Conference Paper

Cold-Start Aware Deep Memory Network for Multi-Entity Aspect-Based Sentiment Analysis

  • Kaisong Song
  • Wei Gao
  • Lujun Zhao
  • Jun Lin
  • Changlong Sun
  • Xiaozhong Liu

Various types of target information have been considered in aspect-based sentiment analysis, such as entities and aspects. Existing research has realized the importance of targets and developed methods with the goal of precisely modeling their contexts via generating target-specific representations. However, all these methods ignore that these representations cannot be learned well due to the lack of sufficient human-annotated target-related reviews, which leads to the data sparsity challenge, a. k. a. cold-start problem here. In this paper, we focus on a more general multiple entity aspect-based sentiment analysis (ME-ABSA) task which aims at identifying the sentiment polarity of different aspects of multiple entities in their context. Faced with severe cold-start scenario, we develop a novel and extensible deep memory network framework with cold-start aware computational layers which use frequency-guided attention mechanism to accentuate on the most related targets, and then compose their representations into a complementary vector for enhancing the representations of cold-start entities and aspects. To verify the effectiveness of the framework, we instantiate it with a concrete context encoding method and then apply the model to the ME-ABSA task. Experimental results conducted on two public datasets demonstrate that the proposed approach outperforms state-of-the-art baselines on ME-ABSA task.

YNIMG Journal 2019 Journal Article

The UNC/UMN Baby Connectome Project (BCP): An overview of the study design and protocol development

  • Brittany R. Howell
  • Martin A. Styner
  • Wei Gao
  • Pew-Thian Yap
  • Li Wang
  • Kristine Baluyot
  • Essa Yacoub
  • Geng Chen

The human brain undergoes extensive and dynamic growth during the first years of life. The UNC/UMN Baby Connectome Project (BCP), one of the Lifespan Connectome Projects funded by NIH, is an ongoing study jointly conducted by investigators at the University of North Carolina at Chapel Hill and the University of Minnesota. The primary objective of the BCP is to characterize brain and behavioral development in typically developing infants across the first 5 years of life. The ultimate goals are to chart emerging patterns of structural and functional connectivity during this period, map brain-behavior associations, and establish a foundation from which to further explore trajectories of health and disease. To accomplish these goals, we are combining state of the art MRI acquisition and analysis techniques, including high-resolution structural MRI (T1-and T2-weighted images), diffusion imaging (dMRI), and resting state functional connectivity MRI (rfMRI). While the overall design of the BCP largely is built on the protocol developed by the Lifespan Human Connectome Project (HCP), given the unique age range of the BCP cohort, additional optimization of imaging parameters and consideration of an age appropriate battery of behavioral assessments were needed. Here we provide the overall study protocol, including approaches for subject recruitment, strategies for imaging typically developing children 0–5 years of age without sedation, imaging protocol and optimization, a description of the battery of behavioral assessments, and QA/QC procedures. Combining HCP inspired neuroimaging data with well-established behavioral assessments during this time period will yield an invaluable resource for the scientific community.

AAAI Conference 2019 Conference Paper

Weighted Oblique Decision Trees

  • Bin-Bin Yang
  • Song-Qing Shen
  • Wei Gao

Decision trees have attracted much attention during the past decades. Previous decision trees include axis-parallel and oblique decision trees; both of them try to find the best splits via exhaustive search or heuristic algorithms in each iteration. Oblique decision trees generally simplify tree structure and take better performance, but are always accompanied with higher computation, as well as the initialization with the best axis-parallel splits. This work presents the Weighted Oblique Decision Tree (WODT) based on continuous optimization with random initialization. We consider different weights of each instance for child nodes at all internal nodes, and then obtain a split by optimizing the continuous and differentiable objective function of weighted information entropy. Extensive experiments show the effectiveness of the proposed algorithm.

ICRA Conference 2018 Conference Paper

Dynamic Actuator Selection and Robust State-Feedback Control of Networked Soft Actuators

  • Nafiseh Ebrahimi
  • Sebastian Nugroho
  • Ahmad F. Taha
  • Nikolaos Gatsis
  • Wei Gao
  • Amir Jafari

The design of robots that are light, soft, powerful is a grand challenge. Since they can easily adapt to dynamic environments, soft robotic systems have the potential of changing the status-quo of bulky robotics. A crucial component of soft robotics is a soft actuator that is activated by external stimuli to generate desired motions. Unfortunately, there is a lack of powerful soft actuators that operate through lightweight power sources. To that end, we recently designed a highly scalable, flexible, biocompatible Electromagnetic Soft Actuator (ESA). With ESAs, artificial muscles can be designed by integrating a network of ESAs. The main research gap addressed in this work is in the absence of system-theoretic understanding of the impact of the realtime control and actuator selection algorithms on the performance of networked soft-body actuators and ESAs. The objective of this paper is to establish a framework that guides the analysis and robust control of networked ESAs. A novel ESA is described, and a configuration of soft actuator matrix to resemble artificial muscle fiber is presented. A mathematical model which depicts the physical network is derived, considering the disturbances due to external forces and linearization errors as an integral part of this model. Then, a robust control and minimal actuator selection problem with logistic constraints and control input bounds is formulated, and tractable computational routines are proposed with numerical case studies.

IJCAI Conference 2018 Conference Paper

Tri-net for Semi-Supervised Deep Learning

  • Dong-Dong Chen
  • Wei Wang
  • Wei Gao
  • Zhi-Hua Zhou

Deep neural networks have witnessed great successes in various real applications, but it requires a large number of labeled data for training. In this paper, we propose tri-net, a deep neural network which is able to use massive unlabeled data to help learning with limited labeled data. We consider model initialization, diversity augmentation and pseudo-label editing simultaneously. In our work, we utilize output smearing to initialize modules, use fine-tuning on labeled data to augment diversity and eliminate unstable pseudo-labels to alleviate the influence of suspicious pseudo-labeled data. Experiments show that our method achieves the best performance in comparison with state-of-the-art semi-supervised deep learning methods. In particular, it achieves 8. 30% error rate on CIFAR-10 by using only 4000 labeled examples.

NeurIPS Conference 2018 Conference Paper

Unorganized Malicious Attacks Detection

  • Ming Pang
  • Wei Gao
  • Min Tao
  • Zhi-Hua Zhou

Recommender systems have attracted much attention during the past decade. Many attack detection algorithms have been developed for better recommendations, mostly focusing on shilling attacks, where an attack organizer produces a large number of user profiles by the same strategy to promote or demote an item. This work considers another different attack style: unorganized malicious attacks, where attackers individually utilize a small number of user profiles to attack different items without organizer. This attack style occurs in many real applications, yet relevant study remains open. We formulate the unorganized malicious attacks detection as a matrix completion problem, and propose the Unorganized Malicious Attacks detection (UMA) algorithm, based on the alternating splitting augmented Lagrangian method. We verify, both theoretically and empirically, the effectiveness of the proposed approach.

IJCAI Conference 2017 Conference Paper

Efficient Label Contamination Attacks Against Black-Box Learning Models

  • Mengchen Zhao
  • Bo An
  • Wei Gao
  • Teng Zhang

Label contamination attack (LCA) is an important type of data poisoning attack where an attacker manipulates the labels of training data to make the learned model beneficial to him. Existing work on LCA assumes that the attacker has full knowledge of the victim learning model, whereas the victim model is usually a black-box to the attacker. In this paper, we develop a Projected Gradient Ascent (PGA) algorithm to compute LCAs on a family of empirical risk minimizations and show that an attack on one victim model can also be effective on other victim models. This makes it possible that the attacker designs an attack against a substitute model and transfers it to a black-box victim model. Based on the observation of the transferability, we develop a defense algorithm to identify the data points that are most likely to be attacked. Empirical studies show that PGA significantly outperforms existing baselines and linear learning models are better substitute models than nonlinear ones.

YNIMG Journal 2017 Journal Article

Functional circuit mapping of striatal output nuclei using simultaneous deep brain stimulation and fMRI

  • Nathalie Van Den Berge
  • Daniel L. Albaugh
  • Andrew Salzwedel
  • Christian Vanhove
  • Roel Van Holen
  • Wei Gao
  • Garret D. Stuber
  • Yen-Yu Ian Shih

The substantia nigra pars reticulata (SNr) and external globus pallidus (GPe) constitute the two major output targets of the rodent striatum. Both the SNr and GPe converge upon thalamic relay nuclei (directly or indirectly, respectively), and are traditionally modeled as functionally antagonistic relay inputs. However, recent anatomical and functional studies have identified unanticipated circuit connectivity in both the SNr and GPe, demonstrating their potential as far more than relay nuclei. In the present study, we employed simultaneous deep brain stimulation and functional magnetic resonance imaging (DBS-fMRI) with cerebral blood volume (CBV) measurements to functionally and unbiasedly map the circuit- and network level connectivity of the SNr and GPe. Sprague-Dawley rats were implanted with a custom-made MR-compatible stimulating electrode in the right SNr (n=6) or GPe (n=7). SNr- and GPe-DBS, conducted across a wide range of stimulation frequencies, revealed a number of surprising evoked responses, including unexpected CBV decreases within the striatum during DBS at either target, as well as GPe-DBS-evoked positive modulation of frontal cortex. Functional connectivity MRI revealed global modulation of neural networks during DBS at either target, sensitive to stimulation frequency and readily reversed following cessation of stimulation. This work thus contributes to a growing literature demonstrating extensive and unanticipated functional connectivity among basal ganglia nuclei.

IJCAI Conference 2017 Conference Paper

Recommendation vs Sentiment Analysis: A Text-Driven Latent Factor Model for Rating Prediction with Cold-Start Awareness

  • Kaisong Song
  • Wei Gao
  • Shi Feng
  • Daling Wang
  • Kam-Fai Wong
  • Chengqi Zhang

Review rating prediction is an important research topic. The problem was approached from either the perspective of recommender systems (RS) or that of sentiment analysis (SA). Recent SA research using deep neural networks (DNNs) has realized the importance of user and product interaction for better interpreting the sentiment of reviews. However, the complexity of DNN models in terms of the scale of parameters is very high, and the performance is not always satisfying especially when user-product interaction is sparse. In this paper, we propose a simple, extensible RS-based model, called Text-driven Latent Factor Model (TLFM), to capture the semantics of reviews, user preferences and product characteristics by jointly optimizing two components, a user-specific LFM and a product-specific LFM, each of which decomposes text into a specific low-dimension representation. Furthermore, we address the cold-start issue by developing a novel Pairwise Rating Comparison strategy (PRC), which utilizes the difference between ratings on common user/product as supplementary information to calibrate parameter estimation. Experiments conducted on IMDB and Yelp datasets validate the advantage of our approach over state-of-the-art baseline methods.

IJCAI Conference 2016 Conference Paper

Detecting Rumors from Microblogs with Recurrent Neural Networks

  • Jing Ma
  • Wei Gao
  • Prasenjit Mitra
  • Sejeong Kwon
  • Bernard J. Jansen
  • Kam-Fai Wong
  • Meeyoung Cha

Microblogging platforms are an ideal place for spreading rumors and automatically debunking rumors is a crucial problem. To detect rumors, existing approaches have relied on hand-crafted features for employing machine learning algorithms that require daunting manual effort. Upon facing a dubious claim, people dispute its truthfulness by posting various cues over time, which generates long-distance dependencies of evidence. This paper presents a novel method that learns continuous representations of microblog events for identifying rumors. The proposed model is based on recurrent neural networks (RNN) for learning the hidden representations that capture the variation of contextual information of relevant posts over time. Experimental results on datasets from two real-world microblog platforms demonstrate that (1) the RNN method outperforms state-of-the-art rumor detection models that use hand-crafted features; (2) performance of the RNN-based algorithm is further improved via sophisticated recurrent units and extra hidden layers; (3) RNN-based method detects rumors more quickly and accurately than existing techniques, including the leading online rumor debunking services.

AIJ Journal 2016 Journal Article

One-pass AUC optimization

  • Wei Gao
  • Lu Wang
  • Rong Jin
  • Shenghuo Zhu
  • Zhi-Hua Zhou

AUC is an important performance measure that has been used in diverse tasks, such as class-imbalanced learning, cost-sensitive learning, learning to rank, etc. In this work, we focus on one-pass AUC optimization that requires going through training data only once without having to store the entire training dataset. Conventional online learning algorithms cannot be applied directly to one-pass AUC optimization because AUC is measured by a sum of losses defined over pairs of instances from different classes. We develop a regression-based algorithm which only needs to maintain the first and second-order statistics of training data in memory, resulting in a storage requirement independent of the number of training data. To efficiently handle high-dimensional data, we develop two deterministic algorithms that approximate the covariance matrices. We verify, both theoretically and empirically, the effectiveness of the proposed algorithms.

YNIMG Journal 2016 Journal Article

Resting state network topology of the ferret brain

  • Zhe Charles Zhou
  • Andrew P. Salzwedel
  • Susanne Radtke-Schuller
  • Yuhui Li
  • Kristin K. Sellers
  • John H. Gilmore
  • Yen-Yu Ian Shih
  • Flavio Fröhlich

Resting state functional magnetic resonance imaging (rsfMRI) has emerged as a versatile tool for non-invasive measurement of functional connectivity patterns in the brain. RsfMRI brain dynamics in rodents, non-human primates, and humans share similar properties; however, little is known about the resting state functional connectivity patterns in the ferret, an animal model with high potential for developmental and cognitive translational study. To address this knowledge-gap, we performed rsfMRI on anesthetized ferrets using a 9. 4T MRI scanner, and subsequently performed group-level independent component analysis (gICA) to identify functionally connected brain networks. Group-level ICA analysis revealed distributed sensory, motor, and higher-order networks in the ferret brain. Subsequent connectivity analysis showed interconnected higher-order networks that constituted a putative default mode network (DMN), a network that exhibits altered connectivity in neuropsychiatric disorders. Finally, we assessed ferret brain topological efficiency using graph theory analysis and found that the ferret brain exhibits small-world properties. Overall, these results provide additional evidence for pan-species resting-state networks, further supporting ferret-based studies of sensory and cognitive function.

AAAI Conference 2016 Conference Paper

Risk Minimization in the Presence of Label Noise

  • Wei Gao
  • Lu Wang
  • Yu-Feng Li
  • Zhi-Hua Zhou

Matrix concentration inequalities have attracted much attention in diverse applications such as linear algebra, statistical estimation, combinatorial optimization, etc. In this paper, we present new Bernstein concentration inequalities depending only on the first moments of random matrices, whereas previous Bernstein inequalities are heavily relevant to the first and second moments. Based on those results, we analyze the empirical risk minimization in the presence of label noise. We find that many popular losses used in risk minimization can be decomposed into two parts, where the first part won’t be affected and only the second part will be affected by noisy labels. We show that the influence of noisy labels on the second part can be reduced by our proposed LICS (Labeled Instance Centroid Smoothing) approach. The effectiveness of the LICS algorithm is justified both theoretically and empirically.

YNIMG Journal 2016 Journal Article

Use of a steady-state baseline to address evoked vs. oscillation models of visual evoked potential origin

  • Minpeng Xu
  • Yihong Jia
  • Hongzhi Qi
  • Yong Hu
  • Feng He
  • Xin Zhao
  • Peng Zhou
  • Lixin Zhang

There has been a long debate about the neural mechanism of event-related potentials (ERPs). Previously, no evidence or method was apparent to validate the two competing models, the evoked model and the oscillation model. One argument is whether the pre-stimulus brain oscillation could influence the following ERP. This study carried out an innovative visual oddball task experiment to investigate the dynamic process of visual evoked potentials. A period of stable oscillations of specified dominant frequencies and initial phases, i. e. the steady-state baseline, would be induced before responses to transient stimuli of different contrasts, which could overcome the artifact problem caused by the ‘sorting’ method. The result first revealed a ‘three-period-transition’ for the generation of visual evoked potentials by an objective decomposition. The ERP almost retained the preceding oscillation during the first period, provided an unstable negative potential in the second period, and generated the N1 component in the third period. The cross term analysis showed that the evoked model couldn't be the whole explanation for the ERP generation. Furthermore, the component analysis revealed that the N1 latency was sensitive to the initial phase under the low stimulus contrast (supporting the oscillation model) but not under the high stimulus contrast (supporting the evoked model). It demonstrated that the external stimulus contrast is a significant factor deciding the explicit model for ERPs. Our method and preliminary results may help reconcile the previous, seemly contradictory findings on the ERP mechanism.

IJCAI Conference 2015 Conference Paper

On the Consistency of AUC Pairwise Optimization

  • Wei Gao
  • Zhi-Hua Zhou

AUC (Area Under ROC Curve) has been an important criterion widely used in diverse learning tasks. To optimize AUC, many learning approaches have been developed, most working with pairwise surrogate losses. Thus, it is important to study the AUC consistency based on minimizing pairwise surrogate losses. In this paper, we introduce the generalized calibration for AUC optimization, and prove that it is a necessary condition for AUC consistency. We then provide a sufficient condition for AUC consistency, and show its usefulness in studying the consistency of various surrogate losses, as well as the invention of new consistent losses. We further derive regret bounds for exponential and logistic losses, and present regret bounds for more general surrogate losses in the realizable setting. Finally, we prove regret bounds that disclose the equivalence between the pairwise exponential loss of AUC and univariate exponential loss of accuracy.

IJCAI Conference 2015 Conference Paper

Personalized Sentiment Classification Based on Latent Individuality of Microblog Users

  • Kaisong Song
  • Shi Feng
  • Wei Gao
  • Daling Wang
  • Ge Yu
  • Kam-Fai Wong

Sentiment expression in microblog posts often reflects user’s specific individuality due to different language habit, personal character, opinion bias and so on. Existing sentiment classification algorithms largely ignore such latent personal distinctions among different microblog users. Meanwhile, sentiment data of microblogs are sparse for individual users, making it infeasible to learn effective personalized classifier. In this paper, we propose a novel, extensible personalized sentiment classification method based on a variant of latent factor model to capture personal sentiment variations by mapping users and posts into a low-dimensional factor space. We alleviate the sparsity of personal texts by decomposing the posts into words which are further represented by the weighted sentiment and topic units based on a set of syntactic units of words obtained from dependency parsing results. To strengthen the representation of users, we leverage users following relation to consolidate the individuality of a user fused from other users with similar interests. Results on real-world microblog datasets confirm that our method outperforms stateof-the-art baseline algorithms with large margins.

AAAI Conference 2014 Conference Paper

Fast Multi-Instance Multi-Label Learning

  • Sheng-Jun Huang
  • Wei Gao
  • Zhi-Hua Zhou

In multi-instance multi-label learning (MIML), one object is represented by multiple instances and simultaneously associated with multiple labels. Existing MIML approaches have been found useful in many applications; however, most of them can only handle moderatesized data. To efficiently handle large data sets, we propose the MIMLfast approach, which first constructs a low-dimensional subspace shared by all labels, and then trains label specific linear models to optimize approximated ranking loss via stochastic gradient descent. Although the MIML problem is complicated, MIMLfast is able to achieve excellent performance by exploiting label relations with shared space and discovering sub-concepts for complicated labels. Experiments show that the performance of MIMLfast is highly competitive to state-of-the-art techniques, whereas its time cost is much less; particularly, on a data set with 30K bags and 270K instances, where none of existing approaches can return results in 24 hours, MIMLfast takes only 12 minutes. Moreover, our approach is able to identify the most representative instance for each label, and thus providing a chance to understand the relation between input patterns and output semantics.

IROS Conference 2014 Conference Paper

HexaMorph: A reconfigurable and foldable hexapod robot inspired by origami

  • Wei Gao
  • Ke Huo
  • Jasjeet Singh Seehra
  • Karthik Ramani
  • Raymond J. Cipra

Origami affords the creation of diverse 3D objects through explicit folding processes from 2D sheets of material. Originally as a paper craft from 17th century AD, origami designs reveal the rudimentary characteristics of sheet folding: it is lightweight, inexpensive, compact and combinatorial. In this paper, we present “HexaMorph”, a novel starfish-like hexapod robot designed for modularity, foldability and reconfigurability. Our folding scheme encompasses periodic foldable tetrahedral units, called “Basic Structural Units” (BSU), for constructing a family of closed-loop spatial mechanisms and robotic forms. The proposed hexapod robot is fabricated using single sheets of cardboard. The electronic and battery components for actuation are allowed to be preassembled on the flattened crease-cut pattern and enclosed inside when the tetrahedral modules are folded. The self-deploying characteristic and the mobility of the robot are investigated, and we discuss the motion planning and control strategies for its squirming locomotion. Our design and folding paradigm provides a novel approach for building reconfigurable robots using a range of lightweight foldable sheets.

TIST Journal 2013 Journal Article

Dynamic joint sentiment-topic model

  • Yulan He
  • Chenghua Lin
  • Wei Gao
  • Kam-Fai Wong

Social media data are produced continuously by a large and uncontrolled number of users. The dynamic nature of such data requires the sentiment and topic analysis model to be also dynamically updated, capturing the most recent language use of sentiments and topics in text. We propose a dynamic Joint Sentiment-Topic model (dJST) which allows the detection and tracking of views of current and recurrent interests and shifts in topic and sentiment. Both topic and sentiment dynamics are captured by assuming that the current sentiment-topic-specific word distributions are generated according to the word distributions at previous epochs. We study three different ways of accounting for such dependency information: (1) sliding window where the current sentiment-topic word distributions are dependent on the previous sentiment-topic-specific word distributions in the last S epochs; (2) skip model where history sentiment topic word distributions are considered by skipping some epochs in between; and (3) multiscale model where previous long- and short- timescale distributions are taken into consideration. We derive efficient online inference procedures to sequentially update the model with newly arrived data and show the effectiveness of our proposed model on the Mozilla add-on reviews crawled between 2007 and 2011.

IJCAI Conference 2013 Conference Paper

Multi-View Discriminant Transfer Learning

  • Pei Yang
  • Wei Gao

We study to incorporate multiple views of data in a perceptive transfer learning framework and propose a Multi-view Discriminant Transfer (MDT) learning approach for domain adaptation. The main idea is to find the optimal discriminant weight vectors for each view such that the correlation between the two-view projected data is maximized, while both the domain discrepancy and the view disagreement are minimized simultaneously. Furthermore, we analyze MDT theoretically from discriminant analysis perspective to explain the condition and reason, under which the proposed method is not applicable. The analytical results allow us to investigate whether there exist within-view and/or betweenview conflicts, and thus provides a deep insight into whether the transfer learning algorithm work properly or not in the view-based problems and the combined learning problem. Experiments show that MDT significantly outperforms the state-of-the-art baselines including some typical multi-view learning approaches in single- or cross-domain.

AIJ Journal 2013 Journal Article

On the consistency of multi-label learning

  • Wei Gao
  • Zhi-Hua Zhou

Multi-label learning has attracted much attention during the past few years. Many multi-label approaches have been developed, mostly working with surrogate loss functions because multi-label loss functions are usually difficult to optimize directly owing to their non-convexity and discontinuity. These approaches are effective empirically, however, little effort has been devoted to the understanding of their consistency, i. e. , the convergence of the risk of learned functions to the Bayes risk. In this paper, we present a theoretical analysis on this important issue. We first prove a necessary and sufficient condition for the consistency of multi-label learning based on surrogate loss functions. Then, we study the consistency of two well-known multi-label loss functions, i. e. , ranking loss and hamming loss. For ranking loss, our results disclose that, surprisingly, none of convex surrogate loss is consistent; we present the partial ranking loss, with which some surrogate losses are proven to be consistent. We also discuss on the consistency of univariate surrogate losses. For hamming loss, we show that two multi-label learning methods, i. e. , one-vs-all and pairwise comparison, which can be regarded as direct extensions from multi-class learning, are inconsistent in general cases yet consistent under the dominating setting, and similar results also hold for some recent multi-label approaches that are variations of one-vs-all. In addition, we discuss on the consistency of learning approaches that address multi-label learning by decomposing into a set of binary classification problems.

AIJ Journal 2013 Journal Article

On the doubt about margin explanation of boosting

  • Wei Gao
  • Zhi-Hua Zhou

Margin theory provides one of the most popular explanations to the success of AdaBoost, where the central point lies in the recognition that margin is the key for characterizing the performance of AdaBoost. This theory has been very influential, e. g. , it has been used to argue that AdaBoost usually does not overfit since it tends to enlarge the margin even after the training error reaches zero. Previously the minimum margin bound was established for AdaBoost, however, Breiman (1999) [9] pointed out that maximizing the minimum margin does not necessarily lead to a better generalization. Later, Reyzin and Schapire (2006) [37] emphasized that the margin distribution rather than minimum margin is crucial to the performance of AdaBoost. In this paper, we first present the kth margin bound and further study on its relationship to previous work such as the minimum margin bound and Emargin bound. Then, we improve the previous empirical Bernstein bounds (Audibert et al. 2009; Maurer and Pontil, 2009) [2, 30], and based on such findings, we defend the margin-based explanation against Breimanʼs doubts by proving a new generalization error bound that considers exactly the same factors as Schapire et al. (1998) [39] but is sharper than Breimanʼs (1999) [9] minimum margin bound. By incorporating factors such as average margin and variance, we present a generalization error bound that is heavily related to the whole margin distribution. We also provide margin distribution bounds for generalization error of voting classifiers in finite VC-dimension space.

IJCAI Conference 2013 Conference Paper

Uniform Convergence, Stability and Learnability for Ranking Problems

  • Wei Gao
  • Zhi-Hua Zhou

Most studies were devoted to the design of efficient algorithms and the evaluation and application on diverse ranking problems, whereas few work has been paid to the theoretical studies on ranking learnability. In this paper, we study the relation between uniform convergence, stability and learnability of ranking. In contrast to supervised learning where the learnability is equivalent to uniform convergence, we show that the ranking uniform convergence is sufficient but not necessary for ranking learnability with AERM, and we further present a sufficient condition for ranking uniform convergence with respect to bipartite ranking loss. Considering the ranking uniform convergence being unnecessary for ranking learnability, we prove that the ranking average stability is a necessary and sufficient condition for ranking learnability.

YNIMG Journal 2012 Journal Article

Altered structural connectivity in neonates at genetic risk for schizophrenia: A combined study using morphological and white matter networks

  • Feng Shi
  • Pew-Thian Yap
  • Wei Gao
  • Weili Lin
  • John H. Gilmore
  • Dinggang Shen

Recently, an increasing body of evidence suggests that developmental abnormalities related to schizophrenia may occur as early as the neonatal stage. Impairments of brain gray matter and wiring problems of axonal fibers are commonly suspected to be responsible for the disconnection hypothesis in schizophrenia adults, but significantly less is known in neonates. In this study, we investigated 26 neonates who were at genetic risk for schizophrenia and 26 demographically matched healthy neonates using both morphological and white matter networks to examine possible brain connectivity abnormalities. The results showed that both populations exhibited small-world network topology. Morphological network analysis indicated that the brain structural associations of the high-risk neonates tended to have globally lower efficiency, longer connection distance, and less number of hub nodes and edges with relatively higher betweenness. Subgroup analysis showed that male neonates were significantly disease-affected, while the female neonates were not. White matter network analysis, however, showed that the fiber networks were globally unaffected, although several subcortical–cortical connections had significantly less number of fibers in high-risk neonates. This study provides new lines of evidence in support of the disconnection hypothesis, reinforcing the notion that the genetic risk of schizophrenia induces alterations in both gray matter structural associations and white matter connectivity.

YNIMG Journal 2011 Journal Article

Sex differences in grey matter atrophy patterns among AD and aMCI patients: Results from ADNI

  • Martha Skup
  • Hongtu Zhu
  • Yaping Wang
  • Kelly S. Giovanello
  • Ja-an Lin
  • Dinggang Shen
  • Feng Shi
  • Wei Gao

We used longitudinal magnetic resonance imaging (MRI) data to determine whether there are any gender differences in grey matter atrophy patterns over time in 197 individuals with probable Alzheimer's disease (AD) and 266 with amnestic mild cognitive impairment (aMCI), compared with 224 healthy controls participating in the Alzheimer's Disease Neuroimaging Initiative (ADNI). While previous research has differentiated probable AD and aMCI groups from controls in brain atrophy, it is unclear whether and how sex plays a role in patterns of change over time. Using regional volumetric maps, we fit longitudinal models to the grey matter data collected at repeated occasions, seeking differences in patterns of volume change over time by sex and diagnostic group in a voxel-wise analysis. Additionally, using a region-of-interest approach, we fit longitudinal models to the global volumetric data of predetermined brain regions to determine whether this more conventional approach is sufficient for determining sex and group differences in atrophy. Our longitudinal analyses revealed that, of the various grey matter regions investigated, males and females in the AD group and the aMCI group showed different patterns of decline over time compared to controls in the bilateral precuneus, bilateral caudate nucleus, right entorhinal gyrus, bilateral thalamus, bilateral middle temporal gyrus, left insula, and right amygdala. As one of the first investigation to model more than two time points of structural MRI data over time, our findings add insight into how AD and aMCI males and females differ from controls and from each other over time.

YNIMG Journal 2009 Journal Article

A unified optimization approach for diffusion tensor imaging technique

  • Wei Gao
  • Hongtu Zhu
  • Weili Lin

An optimization approach for diffusion tensor imaging (DTI) technique is proposed, aiming to improve the estimates of tensors, fractional anisotropy (FA), and fiber directions. With the simulated annealing algorithm, the proposed approach simultaneously optimizes imaging parameters (gradient duration/separation, read-out time, and TE), b-values, and diffusion gradient directions either with or without incorporating prior knowledge of tensor fields. In addition, the method through which tensors are estimated, least squares in our study, was also considered in the optimization procedures. Monte-Carlo simulations were performed for three different scenarios of prior fiber distributions including fibers orientated in 1 (CONE1) and 3 (CONE3) cone areas (50 tensors orderly oriented within a diverging angle of 20° in each cone) and a uniform fiber distribution (UNIF). In addition, three imaging acquisition schemes together with different signal-to-noise ratios were tested, including M/N=1/6, 2/12, and 5/30 for each prior fiber distribution where M and N were the number of b=0 and b>0 images, respectively. Our results show that the optimal b-value ranges between 0. 7 and 1. 0×109 s/m2 for UNIF. However, the optimal b-value ranges become both higher and wider for CONE1 and CONE3 than that of UNIF. In addition, the biases and standard deviations (SD) of tensors, and SD of FA are substantially reduced and the accuracy of fiber directional estimates is improved using the proposed approach particularly in CONE1 when compared with the conventional approaches. Together, the proposed unified optimization approach may offer a direct and simultaneous means to optimize DTI experiments.

v2026.09.13