Arrow Research search

Author name cluster

Pengcheng Shi

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

8 papers
2 author rows

Possible papers

8

JBHI Journal 2025 Journal Article

Ensemble Feature Selection for Microarray Data Classification

  • Xiaojian Ding
  • Pengcheng Shi
  • Xin Wang
  • Kaixiang Wang

Microarray data classification is challenged by high dimensionality and small sample sizes, causing feature selection instability. Traditional ensemble feature selection methods struggle to balance diversity and quality effectively. We propose a novel Ensemble Feature Selection Method (EFSM) that introduces a feature mapping diversity metric to generate a robust candidate pool. EFSM first generates a diverse candidate pool of feature selectors by leveraging randomized neural networks to create multiple non-linear feature mappings (views) of the original data. Its core innovation is an ensemble pruning technique formulated as an optimization problem that jointly maximizes both the predictive accuracy of individual selectors and their pairwise diversity. We simplify this NP-hard problem by converting it into a Semi-Definite Programming (SDP) problem and deriving a novel bound for efficient solution. Finally, the rankings from the pruned ensemble are aggregated using the Borda count method. Extensive experiments on 15 biological datasets demonstrate that EFSM outperforms nine state-of-the-art feature selection methods across popular classifiers, achieving superior and stable performance for high-dimensional data analysis.

IROS Conference 2024 Conference Paper

Hyperbolic Image-and-Pointcloud Contrastive Learning for 3D Classification

  • Naiwen Hu
  • Haozhe Cheng
  • Yifan Xie
  • Pengcheng Shi
  • Jihua Zhu

3D contrastive representation learning has exhibited remarkable efficacy across various downstream tasks. However, existing contrastive learning paradigms based on cosine similarity fail to deeply explore the potential intra-modal hierarchical and cross-modal semantic correlations about multi-modal data in Euclidean space. In response, we seek solutions in hyperbolic space and propose a hyperbolic image-and-pointcloud contrastive learning method (HyperIPC). For the intra-modal branch, we rely on the intrinsic geometric structure to explore the hyperbolic embedding representation of point cloud to capture invariant features. For the cross-modal branch, we leverage images to guide the point cloud in establishing strong semantic hierarchical correlations. Empirical experiments underscore the outstanding classification performance of HyperIPC. Notably, HyperIPC enhances object classification results by 2. 8% and few-shot classification outcomes by 5. 9% on ScanObjectNN compared to the baseline. Furthermore, ablation studies and confirmatory testing validate the rationality of HyperIPC’s parameter settings and the effectiveness of its submodules.

NeurIPS Conference 2024 Conference Paper

Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?

  • Pedro R. Bassi
  • Wenxuan Li
  • Yucheng Tang
  • Fabian Isensee
  • Zifu Wang
  • Jieneng Chen
  • Yu-Cheng Chou
  • Saikat Roy

How can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified metrics, unfair comparisons, and short-term outcome pressure. As a consequence, good performance on standard benchmarks does not guarantee success in real-world scenarios. To address these problems, we present Touchstone, a large-scale collaborative segmentation benchmark of 9 types of abdominal organs. This benchmark is based on 5, 195 training CT scans from 76 hospitals around the world and 5, 903 testing CT scans from 11 additional hospitals. This diverse test set enhances the statistical significance of benchmark results and rigorously evaluates AI algorithms across various out-of-distribution scenarios. We invited 14 inventors of 19 AI algorithms to train their algorithms, while our team, as a third party, independently evaluated these algorithms on three test sets. In addition, we also evaluated pre-existing AI frameworks---which, differing from algorithms, are more flexible and can support different algorithms—including MONAI from NVIDIA, nnU-Net from DKFZ, and numerous other open-source frameworks. We are committed to expanding this benchmark to encourage more innovation of AI algorithms for the medical domain.

AAAI Conference 2021 Conference Paper

A Continual Learning Framework for Uncertainty-Aware Interactive Image Segmentation

  • Ervine Zheng
  • Qi Yu
  • Rui Li
  • Pengcheng Shi
  • Anne Haake

Deep learning models have achieved state-of-the-art performance in semantic image segmentation, but the results provided by fully automatic algorithms are not always guaranteed satisfactory to users. Interactive segmentation offers a solution by accepting user annotations on selective areas of the images to refine the segmentation results. However, most existing models only focus on correcting the current image’s misclassified pixels, with no knowledge carried over to other images. In this work, we formulate interactive image segmentation as a continual learning problem and propose a framework to effectively learn from user annotations, aiming to improve the segmentation on both the current image and unseen images in future tasks while avoiding deteriorated performance on previously-seen images. It employs a probabilistic mask to control the neural network’s kernel activation and extract the most suitable features for segmenting images in each task. We also apply a task-aware embedding to automatically infer the optimal kernel activation for initial segmentation and subsequent refinement. Interactions with users are guided through multi-source uncertainty estimation so that users can focus on the most important areas to minimize the overall manual annotation effort. Experiments are performed on both medical and natural image datasets to illustrate the proposed framework’s effectiveness on basic segmentation performance, forward knowledge transfer, and backward knowledge transfer.

JBHI Journal 2021 Journal Article

Evaluating Technology-Mediated Collaborative Workflows for Telehealth

  • Christopher Bondy
  • Linlin Chen
  • Pamela Grover
  • Vicki Hanson
  • Rui Li
  • Pengcheng Shi

Goals: This paper discusses the need for a predictable method to evaluate gains and gaps of collaborative technology-mediated workflows and introduces an evaluation framework to address this need. Methods: The Collaborative Space – Analysis Framework (CS-AF), introduced in this research, is a cross-disciplinary evaluation method designed to evaluate technology-mediated collaborative workflows. The 5-step CS-AF meta-process includes: (1) current-state workflow definition, (2) current-state (baseline) workflow assessment, (3) technology-mediated workflow development and deployment, (4) technology-mediated workflow assessment, (5) analysis and conclusions. For this research, a comprehensive, empirical study of hypertension exam workflow for telehealth was conducted using the CS-AF approach. Results: The CS-AF systemized approach reveals critical cross-disciplinary evaluation data concerning gains and gaps of collaborative workflows when technology-mediated enhancements are characterized and compared with a baseline workflow for the goal of continuous workflow improvement. Conclusion: The CS-AF is an effective meta-analysis process that can be adapted for use in multiple domains.

NeurIPS Conference 2020 Conference Paper

Dynamic Fusion of Eye Movement Data and Verbal Narrations in Knowledge-rich Domains

  • Ervine Zheng
  • Qi Yu
  • Rui Li
  • Pengcheng Shi
  • Anne Haake

We propose to jointly analyze experts' eye movements and verbal narrations to discover important and interpretable knowledge patterns to better understand their decision-making processes. The discovered patterns can further enhance data-driven statistical models by fusing experts' domain knowledge to support complex human-machine collaborative decision-making. Our key contribution is a novel dynamic Bayesian nonparametric model that assigns latent knowledge patterns into key phases involved in complex decision-making. Each phase is characterized by a unique distribution of word topics discovered from verbal narrations and their dynamic interactions with eye movement patterns, indicating experts' special perceptual behavior within a given decision-making stage. A new split-merge-switch sampler is developed to efficiently explore the posterior state space with an improved mixing rate. Case studies on diagnostic error prediction and disease morphology categorization help demonstrate the effectiveness of the proposed model and discovered knowledge patterns.

JBHI Journal 2015 Journal Article

Representing Variability and Transmural Differences in a Model of Human Heart Failure

  • Mohamed M. Elshrif
  • Pengcheng Shi
  • Elizabeth M. Cherry

During heart failure (HF) at the cellular level, the electrophysiological properties of single myocytes get remodeled, which can trigger the occurrence of ventricular arrhythmias that could be manifested in many forms such as early afterdepolarizations (EADs) and alternans (ALTs). In this paper, based on experimentally observed human HF data, specific ionic and exchanger current strengths are modified from a recently developed human ventricular cell model: the O’Hara–Virág–Varró–Rudy (OVVR) model. A new transmural HF-OVVR model is developed that incorporates HF changes and variability of the observed remodeling. This new heterogeneous HF-OVVR model is able to replicate many of the failing action potential (AP) properties and the dynamics of both $\bf {[Ca^{2+}]_i}$ and $\bf {[Na^{+}]_i}$ in accordance with experimental data. Moreover, it is able to generate EADs for different cell types and exhibits ALTs at modest pacing rate for transmural cell types. We have assessed the HF-OVVR model through the examination of the AP duration and the major ionic currents’ rate dependence in single myocytes. The evaluation of the model comes from utilizing the steady-state (S-S) and S1-S2 restitution curves and from probing the accommodation of the HF-OVVR model to an abrupt change in cycle length. In addition, we have investigated the effect of chosen currents on the AP properties, such as blocking the slow sodium current to shorten the AP duration and suppress the EADs, and have found good agreement with experimental observations. This study should help elucidate arrhythmogenic mechanisms at the cellular level and predict unseen properties under HF conditions. In addition, this AP cell model might be useful for modeling and simulating HF at the tissue and organ levels.

AIIM Journal 2014 Journal Article

From spoken narratives to domain knowledge: Mining linguistic data for medical image understanding

  • Xuan Guo
  • Qi Yu
  • Cecilia Ovesdotter Alm
  • Cara Calvelli
  • Jeff B. Pelz
  • Pengcheng Shi
  • Anne R. Haake

Objectives Extracting useful visual clues from medical images allowing accurate diagnoses requires physicians’ domain knowledge acquired through years of systematic study and clinical training. This is especially true in the dermatology domain, a medical specialty that requires physicians to have image inspection experience. Automating or at least aiding such efforts requires understanding physicians’ reasoning processes and their use of domain knowledge. Mining physicians’ references to medical concepts in narratives during image-based diagnosis of a disease is an interesting research topic that can help reveal experts’ reasoning processes. It can also be a useful resource to assist with design of information technologies for image use and for image case-based medical education systems. Methods and materials We collected data for analyzing physicians’ diagnostic reasoning processes by conducting an experiment that recorded their spoken descriptions during inspection of dermatology images. In this paper we focus on the benefit of physicians’ spoken descriptions and provide a general workflow for mining medical domain knowledge based on linguistic data from these narratives. The challenge of a medical image case can influence the accuracy of the diagnosis as well as how physicians pursue the diagnostic process. Accordingly, we define two lexical metrics for physicians’ narratives—lexical consensus score and top N relatedness score—and evaluate their usefulness by assessing the diagnostic challenge levels of corresponding medical images. We also report on clustering medical images based on anchor concepts obtained from physicians’ medical term usage. These analyses are based on physicians’ spoken narratives that have been preprocessed by incorporating the Unified Medical Language System for detecting medical concepts. Results The image rankings based on lexical consensus score and on top 1 relatedness score are well correlated with those based on challenge levels (Spearman correlation >0. 5 and Kendall correlation >0. 4). Clustering results are largely improved based on our anchor concept method (accuracy >70% and mutual information >80%). Conclusions Physicians’ spoken narratives are valuable for the purpose of mining the domain knowledge that physicians use in medical image inspections. We also show that the semantic metrics introduced in the paper can be successfully applied to medical image understanding and allow discussion of additional uses of these metrics.

v2026.09.13