Arrow Research search

Author name cluster

Hu Chen

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

5 papers
1 author row

Possible papers

5

JBHI Journal 2026 Journal Article

DEW-Net: A W-Shaped Dual-Encoder Network with Attention Fusion Mechanisms for Pathological H&E Image Segmentation

  • Fuhan Meng
  • Xixiang Deng
  • Yingbo Qu
  • Chentao Li
  • Qiang Ma
  • Yusong Mao
  • Xinwei Zhang
  • Pan Huang

Segmenting programmed cell death-ligand 1 (PD-L1) expression regions in lung squamous cell carcinoma from pathological H&E images represents a challenging pixel-level prediction task, attributed to the morphological heterogeneity and size discrepancies of expression areas. Although hybrid architectures of CNN and Transformer can extract local features and capture long-range dependencies, they inadequately address information interaction and redundant information elimination during the fusion process, adversely impacting PD-L1 segmentation accuracy. To address this, we propose a W-shaped dual-encoder network (DEW-Net) with novel attention fusion mechanisms. First, a CNN encoder and a Swin Transformer encoder are connected in parallel to extract multi-layer local and global features from pathological images, respectively. Second, a Cross-Attention Fusion (CAF) module is proposed to strengthen information interaction and semantic feature fusion. Additionally, a Channel Attention (CA) is introduced in skip connections to enhance the channel-wise information of shallow features, while a Bilateral-voting Position Attention (BPA) module is further proposed to eliminate positional noise in same-scale shallow features and reinforce position-wise information. We conducted extensive experiments on four datasets. On the PD-L1 segmentation dataset, DEW-Net achieved superior performance, with DSC and IoU reaching 79. 93% and 71. 27%, respectively. These results demonstrate its strong performance and generalization capability compared to other state-of-the-art (SOTA) methods.

JBHI Journal 2026 Journal Article

Uncertainty-Aware Cross-Modal Retrieval for Medical Report Generation

  • Nan Zhou
  • Meng Liu
  • Linchao He
  • Mengting Luo
  • Yidi Chen
  • Yi Zhang
  • Ke Zou
  • Hu Chen

Automatic medical report generation (MRG) has advanced significantly with retrieval-augmented strategies. However, existing methods face two persistent challenges: 1) a largely reliance on single-modal retrieval, which limits multimodal semantic capture and cross-modal alignment; and 2) a lack of reliable information control, leading to irrelevant noisy content and potential hallucinations. To address these limitations, we propose Uncertainty-aware Cross-modal Alignment and Refinement, named U-CAR, a unified framework that enhances both semantic integration and retrieval reliability. First, a cross-modal alignment module explicitly learns fine-grained correspondences between visual and textual representations, ensuring consistent semantics across modalities. This alignment guides the construction of dual-path retrieval-aware memory banks, with one in the visual domain and one in the textual domain, enabling retrieval to capture complementary cues from both modalities. Second, we design a cross-modal retrieval-augmented generation strategy that jointly attends to the retrieved visual and textual context, thereby enriching semantic coverage and reinforcing the integration of multi-modal evidence in the generated reports. In parallel, we introduce an uncertainty-aware refinement mechanism that quantifies generation confidence to adaptively determine the necessity of retrieval. Experiments on the IU X-Ray and MIMIC-CXR datasets demonstrate that U-CAR outperforms the current state-of-the-art methods, achieving a 9% improvement in CIDEr on IU X-Ray. and a 4% gain in BLEU-4 on MIMIC-CXR. These results underscore U-CAR's effectiveness in generating accurate, coherent, and clinically relevant medical reports. Codes are available in https://github.com/Zhounan1222/U-CAR/tree/main.

JBHI Journal 2024 Journal Article

LA-ViT: A Network With Transformers Constrained by Learned-Parameter-Free Attention for Interpretable Grading in a New Laryngeal Histopathology Image Dataset

  • Pan Huang
  • Hualiang Xiao
  • Peng He
  • Chentao Li
  • Xiaodong Guo
  • Sukun Tian
  • Peng Feng
  • Hu Chen

Grading laryngeal squamous cell carcinoma (LSCC) based on histopathological images is a clinically significant yet challenging task. However, more low-effect background semantic information appeared in the feature maps, feature channels, and class activation maps, which caused a serious impact on the accuracy and interpretability of LSCC grading. While the traditional transformer block makes extensive use of parameter attention, the model overlearns the low-effect background semantic information, resulting in ineffectively reducing the proportion of background semantics. Therefore, we propose an end-to-end network with transformers constrained by learned-parameter-free attention (LA-ViT), which improve the ability to learn high-effect target semantic information and reduce the proportion of background semantics. Firstly, according to generalized linear model and probabilistic, we demonstrate that learned-parameter-free attention (LA) has a stronger ability to learn highly effective target semantic information than parameter attention. Secondly, the first-type LA transformer block of LA-ViT utilizes the feature map position subspace to realize the query. Then, it uses the feature channel subspace to realize the key, and adopts the average convergence to obtain a value. And those construct the LA mechanism. Thus, it reduces the proportion of background semantics in the feature maps and feature channels. Thirdly, the second-type LA transformer block of LA-ViT uses the model probability matrix information and decision level weight information to realize key and query, respectively. And those realize the LA mechanism. So, it reduces the proportion of background semantics in class activation maps. Finally, we build a new complex semantic LSCC pathology image dataset to address the problem, which is less research on LSCC grading models because of lacking clinically meaningful datasets. After extensive experiments, the whole metrics of LA-ViT outperform those of other state-of-the-art methods, and the visualization maps match better with the regions of interest in the pathologists' decision-making. Moreover, the experimental results conducted on a public LSCC pathology image dataset show that LA-ViT has superior generalization performance to that of other state-of-the-art methods.

AAAI Conference 2023 Conference Paper

Low-Resource Personal Attribute Prediction from Conversations

  • Yinan Liu
  • Hu Chen
  • Wei Shen
  • Jiaoyan Chen

Personal knowledge bases (PKBs) are crucial for a broad range of applications such as personalized recommendation and Web-based chatbots. A critical challenge to build PKBs is extracting personal attribute knowledge from users' conversation data. Given some users of a conversational system, a personal attribute and these users' utterances, our goal is to predict the ranking of the given personal attribute values for each user. Previous studies often rely on a relative number of resources such as labeled utterances and external data, yet the attribute knowledge embedded in unlabeled utterances is underutilized and their performance of predicting some difficult personal attributes is still unsatisfactory. In addition, it is found that some text classification methods could be employed to resolve this task directly. However, they also perform not well over those difficult personal attributes. In this paper, we propose a novel framework PEARL to predict personal attributes from conversations by leveraging the abundant personal attribute knowledge from utterances under a low-resource setting in which no labeled utterances or external data are utilized. PEARL combines the biterm semantic information with the word co-occurrence information seamlessly via employing the updated prior attribute knowledge to refine the biterm topic model's Gibbs sampling process in an iterative manner. The extensive experimental results show that PEARL outperforms all the baseline methods not only on the task of personal attribute prediction from conversations over two data sets, but also on the more general weakly supervised text classification task over one data set.

JBHI Journal 2023 Journal Article

SemiMAR: Semi-Supervised Learning for CT Metal Artifact Reduction

  • Tao Wang
  • Hui Yu
  • Zhiwen Wang
  • Hu Chen
  • Yan Liu
  • Jingfeng Lu
  • Yi Zhang

Metal artifacts lead to CT imaging quality degradation. With the success of deep learning (DL) in medical imaging, a number of DL-based supervised methods have been developed for metal artifact reduction (MAR). Nonetheless, fully-supervised MAR methods based on simulated data do not perform well on clinical data due to the domain gap. Although this problem can be avoided in an unsupervised way to a certain degree, severe artifacts cannot be well suppressed in clinical practice. Recently, semi-supervised metal artifact reduction (MAR) methods have gained wide attention due to their ability in narrowing the domain gap and improving MAR performance in clinical data. However, these methods typically require large model sizes, posing challenges for optimization. To address this issue, we propose a novel semi-supervised MAR framework. In our framework, only the artifact-free parts are learned, and the artifacts are inferred by subtracting these clean parts from the metal-corrupted CT images. Our approach leverages a single generator to execute all complex transformations, thereby reducing the model's scale and preventing overlap between clean part and artifacts. To recover more tissue details, we distill the knowledge from the advanced dual-domain MAR network into our model in both image domain and latent feature space. The latent space constraint is achieved via contrastive learning. We also evaluate the impact of different generator architectures by investigating several mainstream deep learning-based MAR backbones. Our experiments demonstrate that the proposed method competes favorably with several state-of-the-art semi-supervised MAR techniques in both qualitative and quantitative aspects.

v2026.09.13