Arrow Research search

Author name cluster

Chen Cao

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

9 papers
1 author row

Possible papers

9

AAAI Conference 2025 Conference Paper

Advancing Audio-Based Text Generation with Imbalance Preference Optimization

  • Zhenghao Zhou
  • Yongjie Liu
  • Chen Cao

Human feedback in generative systems is a highly active frontier of research that aims to improve the quality of generated content and align it with subjective preferences. Existing efforts predominantly focus on text-only large language models (LLMs) or text-based image generation, while cross-modal generation between audio and text remains largely unexplored. Moreover, there is currently no open-source preference dataset to support the deployment of alignment algorithms in this domain. In this work, we take audio speech translation (AST) and audio captioning (AAC) tasks as examples to explore how to enhance the performance of mainstream audio-based text generation models with limited human annotation. Specifically, we propose an novel framework named IPO that includes a model adversarial sampling concept--human annotators act as referees to determine model outcomes, using these results as pseudo-labels for the corresponding beam search hypotheses. Given these imbalance win-loss results, IPO effectively enable the two models to update interactively to win the next round of adversarial sampling. We conduct both subjective and objective evaluations to demonstrate the alignment benefits of IPO and its enhancement on model perception and generation capacities. On both AAC and AST, a few hundreds of annotations significantly enhance the weak model, and the strong model can also be encouraged to achieve new state-of-the-art results in terms of different objective metrics. Additionally, we show the extensibility of IPO by applying it to the reverse task of text-to-speech generation, improving the robustness of system on unseen reference speaker.

JBHI Journal 2025 Journal Article

Medical Graph Diffusion: Hybrid Graph Diffusion With Heterogeneous Graph Convolutional Networks for Medical Text Classification

  • Guishen Wang
  • Shengnan Li
  • Keshuang Liu
  • Xiaowen Hu
  • Chen Cao

Text classification is a critical task for understanding the knowledge behind text, especially in medical text. In this paper, we propose a medical graph diffusion model, named the MGD model, for the medical text classification task. To model more structural relationships within a document, our MGD model constructs a text heterogeneous graph to represent word-level, sentence-level, and word-sentence-level structural relationships. To overcome the limitation of only considering direct neighbors, a graph diffusion convolution is employed to reconstruct the text heterogeneous graph. Subsequently, a heterogeneous graph convolutional network and a multilayer perceptron are used to complete the medical text classification task. To evaluate the performance of our MGD model, various text classification benchmarks, including long text standard benchmarks, short text standard benchmarks, and medical text benchmarks, are used to comprehensively assess the effectiveness and robustness of our MGD model. Compared with other representative baselines, it achieved notable improvements in both Accuracy and F1 score evaluation metrics. Ablation experiment results further demonstrated that the construction of heterogeneous graphs and the use of diffusion graph convolutional networks significantly impact the performance of our MGD model.

JBHI Journal 2025 Journal Article

MMDDI-SSE: A Novel Multi-Modal Feature Fusion Model With Static Subgraph Embedding for Drug-Drug Interaction Event Prediction

  • Guishen Wang
  • Honghan Chen
  • Handan Wang
  • Hairong Gao
  • Xiaowen Hu
  • Chen Cao

Artificial intelligence techniques play a pivotal role in the accurate identification of drug-drug interaction (DDI) events, thereby informing clinical decisions and treatment regimens. While existing DDI prediction models have made significant progress by leveraging sequence features such as chemical substructures, targets, and enzymes, they often face limitations in integrating and effectively utilizing multi-modal drug representations. To address these limitations, this study proposes a novel multi-modal feature fusion model for DDI event prediction: MMDDI-SSE. Our approach integrates drug sequence modality with DDI graph representations through a novel architecture that employs static subgraph generation to capture structural properties. The model utilizes a graph autoencoder architecture to learn both local and global topological features from these subgraphs, while simultaneously processing diverse sequence-based characteristics including semantically enhanced pharmacodynamic features, chemical substructures, target proteins, and enzyme information. Through comprehensive evaluation on two distinct datasets, MMDDI-SSE demonstrates superior predictive performance compared to state-of-the-art baselines. Ablation studies further validate the effectiveness of each architectural component in enhancing DDI prediction accuracy.

JBHI Journal 2024 Journal Article

EHR-HGCN: An Enhanced Hybrid Approach for Text Classification Using Heterogeneous Graph Convolutional Networks in Electronic Health Records

  • Guishen Wang
  • Xiaoxue Lou
  • Fang Guo
  • Devin Kwok
  • Chen Cao

Text classification is a central part of natural language processing, with important applications in understanding the knowledge behind biomedical texts including electronic health records (EHR). In this article, we propose a novel heterogeneous graph convolutional network method for classifying EHR texts. Our method, called EHR-HGCN, is able to combine context-sensitive word and sentence embeddings with structural sentence-level and word-level relation information to perform text classification. EHR-HGCN reframes EHR text classification as a graph classification task to better capture structural information about the document using a heterogeneous graph. To mine contextual information from a document, EHR-HGCN first applies a bidirectional recurrent neural network (BiRNN) on word embeddings obtained via Global Vectors for word representation (GloVe) to obtain context-sensitive word-level and sentence-level embeddings. To mine structural relationships from the document, EHR-HGCN then constructs a heterogeneous graph over the word and sentence embeddings, where sentence-word and word-word relationships are represented by graph edges. Finally, a heterogeneous graph convolutional neural network is used to classify documents by their graph representation. We evaluate EHR-HGCN on a variety of standard text classification benchmarks and find that EHR-HGCN has higher accuracy and F1-score than other representative machine learning and deep learning methods. We also apply EHR-HGCN to the MedLit benchmark and find it performs with high accuracy and F1-score on the task of section classification in EHR texts. Our ablation experiments show that the heterogeneous graph construction and heterogeneous graph convolutional network are critical to the performance of EHR-HGCN.

TIST Journal 2023 Journal Article

What Your Next Check-in Might Look Like: Next Check-in Behavior Prediction

  • Heli Sun
  • Chen Cao
  • Xuguang Chu
  • Tingting Hu
  • Junzhi Lu
  • Liang He
  • Zhi Wang
  • Hui He

In recent years, the next-POI recommendation has become a trending research topic in the field of trajectory data mining. For protection of user privacy, users’ complete GPS trajectories are difficult to obtain. The check-in information posted by users on social networks has become an important data source for Spatio-temporal Trajectory research. However, state-of-the-art methods neglect the social meaning and the information dissemination function of check-in behavior. The social meaning is an important reason why users are willing to post check-in on social networks, and the information dissemination function means, users can affect each other’s behavior by check-ins. The above characteristics of the check-in behavior make it different from the visiting behavior. We consider a new problem of predicting the next check-in behavior including the check-in time, the POI (point-of-interest) where the check-in is located, functional semantics of the POI, and so on. To solve the proposed problem, we build a multi-task learning model called DPMTM, and a pre-training module is designed to extract dynamic social semantics of check-in behaviors. Our results show that the DPMTM model works well in the check-in behavior problem.

AIIM Journal 2022 Journal Article

Brain gray matter nuclei segmentation on quantitative susceptibility mapping using dual-branch convolutional neural network

  • Chao Chai
  • Pengchong Qiao
  • Bin Zhao
  • Huiying Wang
  • Guohua Liu
  • Hong Wu
  • Wen Shen
  • Chen Cao

Abnormal iron accumulation in the brain subcortical nuclei has been reported to be correlated to various neurodegenerative diseases, which can be measured through the magnetic susceptibility from the quantitative susceptibility mapping (QSM). To quantitatively measure the magnetic susceptibility, the nuclei should be accurately segmented, which is a tedious task for clinicians. In this paper, we proposed a dual-branch residual-structured U-Net (DB-ResUNet) based on 3D convolutional neural network (CNN) to automatically segment such brain gray matter nuclei. Due to memory limit, 3D-CNN-based methods typically adopted image patches, instead of the whole volumetric image, which, however, ignored the spatial contextual information of the neighboring patches, and therefore led to the accuracy loss. To better tradeoff segmentation accuracy and the memory efficiency, the proposed DB-ResUNet incorporated patches with different resolutions. By jointly using QSM and 3D T1 weighted imaging (T1WI) as inputs, the proposed method was able to achieve better segmentation accuracy over its single-branch counterpart, as well as the conventional atlas-based method and the classical 3D CNN structures. The susceptibility values and the volumes were also measured, which indicated that the measurements from the proposed DB-ResUNet was able to present high correlation with values from the manually annotated regions of interest.

NeurIPS Conference 2020 Conference Paper

Fully Convolutional Mesh Autoencoder using Efficient Spatially Varying Kernels

  • Yi Zhou
  • Chenglei Wu
  • Zimo Li
  • Chen Cao
  • Yuting Ye
  • Jason Saragih
  • Hao Li
  • Yaser Sheikh

Learning latent representations of registered meshes is useful for many 3D tasks. Techniques have recently shifted to neural mesh autoencoders. Although they demonstrate higher precision than traditional methods, they remain unable to capture fine-grained deformations. Furthermore, these methods can only be applied to a template-specific surface mesh, and is not applicable to more general meshes, like tetrahedrons and non-manifold meshes. While more general graph convolution methods can be employed, they lack performance in reconstruction precision and require higher memory usage. In this paper, we propose a non-template-specific fully convolutional mesh autoencoder for arbitrary registered mesh data. It is enabled by our novel convolution and (un)pooling operators learned with globally shared weights and locally varying coefficients which can efficiently capture the spatially varying contents presented by irregular mesh connections. Our model outperforms state-of-the-art methods on reconstruction accuracy. In addition, the latent codes of our network are fully localized thanks to the fully convolutional structure, and thus have much higher interpolation capability than many traditional 3D mesh generation models.

YNICL Journal 2019 Journal Article

Differential involvement of rubral branches in chronic capsular and pontine stroke

  • Jun Guo
  • Jingchun Liu
  • Caihong Wang
  • Chen Cao
  • Lejun Fu
  • Tong Han
  • Jingliang Cheng
  • Chunshui Yu

Background and Purpose Early studies have indicated that the cortico-rubro-spinal tracts play important roles in motor dysfunction after stroke. However, the differential involvement of the rubral branches in capsular and pontine stroke, and their associations with the motor impairment are still unknown. Methods The present study recruited 144 chronic stroke patients and 91 normal controls (NC) from three hospitals, including 102 cases with capsular stroke (CS) and 42 cases with pontine stroke (PS). The rubral branches, including bilateral corticorubral tracts (CRT), dentatorubral tracts (DRT), and rubrospinal tracts (RST), and the cortico-spinal tract (CST) were reconstructed based on the dataset of the Human Connectome Project. Group differences in diffusion scalars of each rubral branch were compared, and the associations between the diffusion measures of rubral branches and the Fugl-Meyer assessment (FMA) scores were tested. Results The bilateral CRT of the CS cases showed significantly lower factional anisotropy (FA) than in the NC. The bilateral DRT of the PS cases had lower FA than in the NC. Both CS and PS cases had significantly lower FA of the bilateral RST than the NC. Besides, the stroke patients demonstrated significantly lower FA in bilateral CSTs than the NC. Partial correlation analysis identified significantly positive correlations between the FA of the ipsilesional and CRT and the FMA scores in the CS group, and significantly positive correlations between the FA of the RST bilaterally and the FMA scores in the CS and PS groups. Furthermore, the association between RST integrity and FMA scores still survived after controlling for the effect of the CST. Finally, multiple regression modelling found that rubral tract FA explained 39. 2% of the variance in FMA scores for CS patients, and 48. 8% of the variance in FMA scores for PS patients. Conclusions The bilateral rubral branches were differentially involved in the chronic capsular and pontine stroke, and the impairment severity of each rubral branch was dependent on lesion locations. The integrity of the rubral branches is related to motor impairment in both the chronic capsular and pontine stroke.

YNICL Journal 2017 Journal Article

Gray matter volume changes in chronic subcortical stroke: A cross-sectional study

  • Qingqing Diao
  • Jingchun Liu
  • Caihong Wang
  • Chen Cao
  • Jun Guo
  • Tong Han
  • Jingliang Cheng
  • Xuejun Zhang

This study aimed to investigate the effects of lesion side and degree of motor recovery on gray matter volume (GMV) difference relative to healthy controls in right-handed subcortical stroke. Structural MRI data were collected in 97 patients with chronic subcortical ischemic stroke and 79 healthy controls. Voxel-wise GMV analysis was used to investigate the effects of lesion side and degree of motor recovery on GMV difference in right-handed chronic subcortical stroke patients. Compared with healthy controls, right-lesion patients demonstrated GMV increase (P <0. 05, voxel-wise false discovery rate correction) in the bilateral paracentral lobule (PCL) and supplementary motor area (SMA) and the right middle occipital gyrus (MOG); while left-lesion patients did not exhibit GMV difference under the same threshold. Patients with complete and partial motor recovery showed similar degree of GMV increase in right-lesion patients. However, the motor recovery was correlated with the GMV increase in the bilateral SMA in right-lesion patients. These findings suggest that there exists a lesion-side effect on GMV difference relative to healthy controls in right-handed patients with chronic subcortical stroke. The GMV increase in the SMA may facilitate motor recovery in subcortical stroke patients.

v2026.09.13