Arrow Research search

Author name cluster

Ao Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

23 papers
2 author rows

Possible papers

23

EAAI Journal 2026 Journal Article

An interpretable Transformer–LSTM denoising autoencoder for semi-supervised fault diagnosis in chemical processes

  • Lijie Guo
  • Jiaqi Shi
  • Jianxin Kang
  • Ao Li

To tackle key challenges in chemical process fault diagnosis, such as limited labeled data, complex feature extraction, and low interpretability, we propose an interpretable Transformer and Long Short-Term Memory (LSTM) denoising autoencoder (TrLAe) model for semi-supervised fault diagnosis. The encoder combines Transformer and dual-branch LSTMs to capture both global multivariate dependencies and local temporal dynamics, enabling complementary feature fusion. The decoder reconstructs high-dimensional latent features through a series of Transformer–LSTM modules. By leveraging the autoencoder's reconstruction learning, the model effectively uses abundant unlabeled data to learn latent feature distributions in an unsupervised manner, producing robust representations even with limited labeled samples. A denoising mechanism is employed to enhance generalization, and spatiotemporal attention is incorporated to improve interpretability and identify fault-related variables. To support engineering decision-making, these identified fault-related variables are then mapped to the deviation–cause–consequence framework in hazard and operability (HAZOP) analysis, creating a knowledge-driven, interpretable approach to fault propagation. Experiments on the Tennessee Eastman process show that the TrLAe model achieves high diagnostic accuracy and strong generalization, highlighting its potential to enhance the reliability and safety of modern chemical processes.

EAAI Journal 2026 Journal Article

Cross-stain knowledge distillation for low-cost lung cancer programmed death ligand-1 assessment with multi-granularity multiple instance learning

  • Yi Shi
  • Chong Ge
  • Fang Zhao
  • Anli Zhang
  • Ao Li
  • Haibo Wu
  • Minghui Wang

Accurately assessing programmed death ligand-1 (PD-L1) status, recognizing patients potentially responsive to immunotherapies. Since immunohistochemistry (IHC) staining is gold standard in identifying molecular biomarker but often expensive and unattainable, routine hematoxylin and eosin (H&E) staining offers a low-cost alternative. However, H&E images primarily reveal morphological knowledge and inherently lack PD-L1-related molecular information, resulting in a severe mono-stain knowledge limitation. Additionally, most existing approaches analyze gigapixel whole-slide images at only a single magnification, which fails to unravel complex pathological information across granularities, leading to a significant uni-granularity information limitation. Therefore, we propose an innovative cross-stain knowledge distillation with multi-granularity framework, namely CroSMuG. First, to alleviate uni-granularity information limitation, a new multi-granularity multiple instance learning framework is introduced. This is based on macro-micro dual branches, which comprises a macro-branch and a micro-branch to extract global and local pathological information. Furthermore, we develop a novel cross-stain knowledge distillation strategy featuring triple-united distillation loss. Specifically, this approach introduces globality-, locality- and task-aware knowledge distillation, enabling the H&E-based predictive network as a student to learn crucial molecular knowledge from an IHC teacher network. Extensive experiments are conducted on diverse real-world datasets from multiple medical centers, and CroSMuG achieves the superior performance with area under the curve (AUC) of 83. 6 % on internal dataset and 81. 2 % on external dataset. These results highlight the generalizability of CroSMuG for accurate H&E-based PD-L1 assessment, offering the potential for practical applications in lung cancer immunotherapy decision-making in clinical practices.

AAAI Conference 2026 Conference Paper

Enhancing Kernel Power $K$-means: Scalable and Robust Clustering with Random Fourier Features and Possibilistic Method

  • Yixi Chen
  • Weixuan Liang
  • Tianrui Liu
  • Jun-Jie Huang
  • Ao Li
  • Xueling Zhu
  • Xinwang Liu

Kernel power k-means (KPKM) leverages a family of means to mitigate local minima issues in kernel k-means. However, KPKM faces two key limitations: (1) the computational burden of the full kernel matrix restricts its use on extensive data, and (2) the lack of authentic centroid-sample assignment learning reduces its noise robustness. To overcome these challenges, we propose RFF-KPKM, introducing the first approximation theory for applying random Fourier features (RFF) to KPKM. RFF-KPKM employs RFF to generate efficient, low dimensional feature maps, bypassing the need for the whole kernel matrix. Crucially, we are the first to establish strong theoretical guarantees for this combination: (1) an excess risk bound of O( k^3/n), (2) strong consistency with membership values, and (3) a (1 + ε) relative error bound achievable using the RFF of dimension poly(ε^{−1} logk). Furthermore, to improve robustness and the ability to learn multiple kernels, we propose IP-RFF-MKPKM, an improved possibilistic RFF-based multiple kernel power k-means. IP-RFF-MKPKM ensures the scalability of MKPKM via RFF and refines cluster assignments by combining the merits of the possibilistic and fuzzy membership. Experiments on large-scale datasets demonstrate the superior efficiency and clustering accuracy of the proposed methods compared to the state-of-the-art alternatives.

JBHI Journal 2026 Journal Article

MTS-LOF: Medical Time-Series Representation Learning via Occlusion-Invariant Features

  • Huayu Li
  • Ana S. Carreon-Rascon
  • Xiwen Chen
  • Geng Yuan
  • Ao Li

Medical time series data are indispensable in healthcare, providing critical insights for disease diagnosis, treatment planning, and patient management. The exponential growth in data complexity, driven by advanced sensor technologies, has presented challenges related to data labeling. Self-supervised learning (SSL) has emerged as a transformative approach to address these challenges, eliminating the need for extensive human annotation. In this study, we introduce a novel framework for Medical Time Series Representation Learning, known as MTS-LOF. MTS-LOF leverages the strengths of Joint-Embedding SSL and Masked Autoencoder (MAE) methods, offering a unique approach to representation learning for medical time series data. By combining these techniques, MTS-LOF enhances the potential of healthcare applications by providing more sophisticated, context-rich representations. Additionally, MTS-LOF employs a multi-masking strategy to facilitate occlusion-invariant feature learning. This approach allows the model to create multiple views of the data by masking portions of it. By minimizing the discrepancy between the representations of these masked patches and the fully visible patches, MTS-LOF learns to capture rich contextual information within medical time series datasets. The results of experiments conducted on diverse medical time series datasets demonstrate the superiority of MTS-LOF over other methods. These findings hold promise for significantly enhancing healthcare applications by improving representation learning. Furthermore, our work delves into the integration of Joint-Embedding SSL and MAE techniques, shedding light on the intricate interplay between temporal and structural dependencies in healthcare data. This understanding is crucial, as it allows us to grasp the complexities of healthcare data analysis.

AAAI Conference 2026 Conference Paper

Sample-specific Modality Diagnosis and Cross-modal Enhancement for Incomplete Multimodal Representations

  • Junsong Chen
  • Jiyuan Liu
  • Suyuan Liu
  • Wei Zhang
  • Ao Li
  • En Zhu
  • Xinwang Liu

In multimodal sentiment analysis, modality missingness and quality degradation are common. Existing methods often rely on batch-level modality generation, generation but neglect sample-level missingness, hence their flexibility is limited severely in real-world scenarios. To address this, Sample-specific Modality Diagnosis and Cross-modal Enhancement for Incomplete Multimodal Representations (SMCIR) is proposed. Specifically, The Dynamic Multi-feature Fusion Detector (DMFD) is presented, which detects missingness and severity at the sample-level using indicators such as information entropy, modality similarity, and mutual information. Unlike batch-based methods, the DMFD provides fine-grained detection and adaptive responses, improving sensitivity to modality disturbances. Meanwhile, the Context-aware Modality Completion Generator (CMCG) is developed to restore missing modalities through context-guided reconstruction using multiscale feature fusion and cross-modal attention. In this way, the proposed CMCG method can avoid redundancy and inconsistency, enhancing the consistency and discriminativity of the fused representation. In CMCG, the text modality serves as a stable guide to improve context consistency. Experiments on the CMU-MOSI and CMU-MOSEI datasets show that SMCIR outperforms existing full-modal and non-recovery-based methods, well validating its efficacy and superiority in multimodal learning.

ICLR Conference 2025 Conference Paper

Agent-Oriented Planning in Multi-Agent Systems

  • Ao Li
  • Yuexiang Xie
  • Songze Li
  • Fugee Tsung
  • Bolin Ding
  • Yaliang Li

Through the collaboration of multiple LLM-empowered agents possessing diverse expertise and tools, multi-agent systems achieve impressive progress in solving real-world problems. Given the user queries, the meta-agents, serving as the brain within multi-agent systems, are required to decompose the queries into multiple sub-tasks that can be allocated to suitable agents capable of solving them, so-called agent-oriented planning. In this study, we identify three critical design principles of agent-oriented planning, including solvability, completeness, and non-redundancy, to ensure that each sub-task can be effectively resolved, resulting in satisfactory responses to user queries. These principles further inspire us to propose AOP, a novel framework for agent-oriented planning in multi-agent systems, leveraging a fast task decomposition and allocation process followed by an effective and efficient evaluation via a reward model. According to the evaluation results, the meta-agent is also responsible for promptly making necessary adjustments to sub-tasks and scheduling. Besides, we integrate a feedback loop into AOP to further enhance the effectiveness and robustness of such a problem-solving process. Extensive experiments demonstrate the advancement of AOP in solving real-world problems compared to both single-agent systems and existing planning strategies for multi-agent systems. The source code is available at https://github.com/lalaliat/Agent-Oriented-Planning

JBHI Journal 2025 Journal Article

Agnostic-Specific Modality Learning for Cancer Survival Prediction From Multiple Data

  • Honglei Liu
  • Yi Shi
  • Ying Xu
  • Ao Li
  • Minghui Wang

Cancer is a pressing public health problem and one of the main causes of mortality worldwide. The development of advanced computational methods for predicting cancer survival is pivotal in aiding clinicians to formulate effective treatment strategies and improve patient quality of life. Recent advances in survival prediction methods show that integrating diverse information from various cancer-related data, such as pathological images and genomics, is crucial for improving prediction accuracy. Despite promising results of existing approaches, there are great challenges of modality gap and semantic redundancy presented in multiple cancer data, which could hinder the comprehensive integration and pose substantial obstacles to further enhancing cancer survival prediction. In this study, we propose a novel agnostic-specific modality learning (ASML) framework for accurate cancer survival prediction. To bridge the modality gap and provide a comprehensive view of distinct data modalities, we employ an agnostic-specific learning strategy to learn the commonality across modalities and the uniqueness of each modality. Moreover, a cross-modal fusion network is exerted to integrate multimodal information by modeling modality correlations and diminish semantic redundancy in a divide-and-conquer manner. Extensive experiment results on three TCGA datasets demonstrate that ASML reaches better performance than other existing cancer survival prediction methods for multiple data.

EAAI Journal 2025 Journal Article

Deep spectral clustering network for incomplete multi-view clustering

  • Ao Li
  • Sanlin Mei
  • Cong Feng
  • Tianyu Gao
  • Hai Huang

Spectral clustering is a helpful technique for clustering non-convex data, which extends the clustering to data with multiple partial views. Still, it has a higher running time thanks to cubic time complexity, while extending spectral embeddings to unseen samples is non-trivial and prevents model reuse. In light of this, we propose a Deep Spectral Clustering Network for Incomplete Multi-view Clustering (DSCN-IMC). At its core, DSCN-IMC jointly learns the approximate spectral embeddings and the soft cluster assignments using a deep neural network in an unsupervised and end-to-end fashion. Specifically, we integrate the graph convolutional layers and manifold loss into our network to achieve robustness against incomplete multi-view data. After that, an orthnormalization layer is exploited to fulfill the orthogonal constraint of spectral embeddings, thereby achieving the lower running time of spectral clustering. Building upon this, a self-supervised clustering module is introduced to obtain cluster assignments of unseen samples and facilitates model reuse. Our method’s efficacy is demonstrated through experiments conducted on nine datasets, wherein we compare its performance against nine state-of-the-art baselines.

JBHI Journal 2025 Journal Article

DRLSurv: Disentangled Representation Learning for Cancer Survival Prediction by Mining Multimodal Consistency and Complementarity

  • Ying Xu
  • Yi Shi
  • Honglei Liu
  • Ao Li
  • Anli Zhang
  • Minghui Wang

Accurate cancer survival prediction is crucial in devising optimal treatment plans and offering individualized care to improve clinical outcomes. Recent researches confirm that integrating heterogenous cancer data such as histopathological images and genomic data, can enhance our understanding of cancer progression and provides a multimodal perspective on patient survival chances. However, existing methods often over-look the fundamental aspects of multimodal data, i. e. , consistency and complementarity, which in consequence significantly hinder advancements in cancer survival prediction. To address this issue, we represent DRLSurv, a novel multimodal deep learning method that leverages disentangled representation learning for precise cancer survival prediction. Through dedicated deep encoding networks, DRLSurv decomposes each modality into modality-invariant and modality-specific representations, which are mapped to common and unique feature subspaces for simultaneously mining the distinct aspects of cancer multimodal data. Moreover, our method innovatively introduces a subspace-based proximity contrastive loss and re-disentanglement loss, thus ensuring the successful decomposition of consistent and complementary information while maintaining the multimodal fidelity during the learning of disentangled representations. Both quantitative analyses and visual assessments on different datasets validate the superiority of DRLSurv over existing survival prediction approaches, demonstrating its powerful capability to exploit enriched survival-related information from cancer multimodal data. Therefore, DRLSurv not only offers a unified and comprehensive deep learning framework for advancing multimodal survival predictions, but also provides valuable insights for cancer prognosis and survival analysis.

IROS Conference 2025 Conference Paper

Exploratory Movement Strategies for Texture Discrimination with a Neuromorphic Tactile Sensor

  • Xingchen Xu
  • Ao Li
  • Benjamin Ward-Cherrier

We propose a neuromorphic tactile sensing frame-work for robotic texture classification that is inspired by human exploratory strategies. Our system utilizes the NeuroTac sensor to capture neuromorphic tactile data during a series of exploratory motions. We first tested six distinct motions for texture classification under fixed environment: sliding, rotating, tapping, as well as the combined motions: sliding+rotating, tapping+rotating, and tapping+sliding. We chose sliding and sliding+rotating as the best motions based on final accuracy and the sample timing length needed to reach converged accuracy. In the second experiment designed to simulate complex real-world conditions, these two motions were further evaluated under varying contact depth and speeds. Under these conditions, our framework attained the highest accuracy of 87. 33% with sliding+rotating while maintaining an extremely low power consumption of only 8. 04 mW. These results suggest that the sliding+rotating motion is the optimal exploratory strategy for neuromorphic tactile sensing deployment in texture classification tasks and holds significant promise for enhancing robotic environmental interaction.

ICLR Conference 2025 Conference Paper

InstaRevive: One-Step Image Enhancement via Dynamic Score Matching

  • Yixuan Zhu
  • Haolin Wang 0006
  • Ao Li
  • Wenliang Zhao
  • Yansong Tang
  • Jingxuan Niu
  • Lei Chen 0069
  • Jie Zhou 0001

Image enhancement finds wide-ranging applications in real-world scenarios due to complex environments and the inherent limitations of imaging devices. Recent diffusion-based methods yield promising outcomes but necessitate prolonged and computationally intensive iterative sampling. In response, we propose InstaRevive, a straightforward yet powerful image enhancement framework that employs score-based diffusion distillation to harness potent generative capability and minimize the sampling steps. To fully exploit the potential of the pre-trained diffusion model, we devise a practical and effective diffusion distillation pipeline using dynamic noise control to address inaccuracies in updating direction during score matching. Our noise control strategy enables a dynamic diffusing scope, facilitating precise learning of denoising trajectories within the diffusion model and ensuring accurate distribution matching gradients during training. Additionally, to enrich guidance for the generative power, we incorporate textual prompts via image captioning as auxiliary conditions, fostering further exploration of the diffusion model. Extensive experiments substantiate the efficacy of our framework across a diverse array of challenging tasks and datasets, unveiling the compelling efficacy and efficiency of InstaRevive in delivering high-quality and visually appealing results.

JBHI Journal 2025 Journal Article

MIF: Multi-Shot Interactive Fusion Model for Cancer Survival Prediction Using Pathological Image and Genomic Data

  • Yi Shi
  • Minghui Wang
  • Honglei Liu
  • Fang Zhao
  • Ao Li
  • Xun Chen

Accurate cancer survival prediction is crucial for oncologists to determine therapeutic plan, which directly influences the treatment efficacy and survival outcome of patient. Recently, multimodal fusion-based prognostic methods have demonstrated effectiveness for survival prediction by fusing diverse cancer-related data from different medical modalities, e. g. , pathological images and genomic data. However, these works still face significant challenges. First, most approaches attempt multimodal fusion by simple one-shot fusion strategy, which is insufficient to explore complex interactions underlying in highly disparate multimodal data. Second, current methods for investigating multimodal interactions face the capability-efficiency dilemma, which is the difficult balance between powerful modeling capability and applicable computational efficiency, thus impeding effective multimodal fusion. In this study, to encounter these challenges, we propose an innovative multi-shot interactive fusion method named MIF for precise survival prediction by utilizing pathological and genomic data. Particularly, a novel multi-shot fusion framework is introduced to promote multimodal fusion by decomposing it into successive fusing stages, thus delicately integrating modalities in a progressive way. Moreover, to address the capacity-efficiency dilemma, various affinity-based interactive modules are introduced to synergize the multi-shot framework. Specifically, by harnessing comprehensive affinity information as guidance for mining interactions, the proposed interactive modules can efficiently generate low-dimensional discriminative multimodal representations. Extensive experiments on different cancer datasets unravel that our method not only successfully achieves state-of-the-art performance by performing effective multimodal fusion, but also possesses high computational efficiency compared to existing survival prediction methods.

EAAI Journal 2025 Journal Article

Multimodal emotion recognition by fusing complementary patterns from central to peripheral neurophysiological signals across feature domains

  • Zhuang Ma
  • Ao Li
  • Jiehao Tang
  • Jianhua Zhang
  • Zhong Yin

The implementation and application of artificial intelligence are propelling various advanced affective computing frameworks. Automatic recognition of emotions using multimodal physiological signals enhances the efficiency of systems such as health-care applications, pilot cognitive state monitoring, and passive brain-computer interfaces. However, challenges remain in capturing topological-frequency patterns in diverse electroencephalogram (EEG) electrode layouts, uncovering coupling dynamics across adjacent peripheral modalities, and integrating complementary affective patterns from the feature to modality levels. To address these issues, we propose a central-to-peripheral complementary integration network that employs hybrid encoders to extract and integrate affective patterns from EEG and peripheral signals. For EEG, the model unifies features from different channels into a single map to extract local-to-global representations, while for peripheral signals, adjacent cross-modal information is embedded into global affective patterns. These abstractions are systematically aggregated for emotion recognition by aligning affective relevance across domains and modalities within the central and peripheral nervous systems. The proposed model was evaluated on four publicly available multimodal databases using a leave-one-subject-out cross-validation approach. On the Database for Emotion Analysis using Physiological signal (DEAP), the binary recognition accuracy for valence and arousal scales was 75. 00% and 77. 33%, respectively. On the Human-Computer Interaction (HCI) database, the corresponding binary accuracies were 78. 78% and 75. 38%. For the SJTU Emotion EEG Datasets IV and V (SEED-IV and SEED-V), the four-class and five-class accuracies were 71. 94% and 84. 83%, respectively. These results validate the robustness and remarkable generalization capability of the proposed method.

AIIM Journal 2025 Journal Article

SCLResNet and DSAF: A self-supervised contrastive learning and deep self-attention fusion-based multimodal network for predicting central lymph node metastasis in papillary thyroid carcinoma

  • Shidi Miao
  • Yuyang Jiang
  • Wenjuan Huang
  • Yuxin Jiang
  • Mengzhuo Sun
  • Mingxuan Wang
  • Hongzhuo Qi
  • Ao Li

Accurate prediction of central lymph node metastasis (CLNM) in papillary thyroid carcinoma (PTC) is crucial to avoid unnecessary invasive procedures, yet existing models often fall short. We constructed the SCLResNet101 model based on a contrastive learning framework to extract network features of tumor ultrasound (US). SeResnet101 was used to extract network features of peri-vascular adipose tissue (PVAT) from the computed tomography (CT) of C6 (the arterial and venous layers beneath the thyroid). Univariate and multivariate analyses were performed using binary logistic regression to select clinical features. Finally, we constructed a Deep Self-Attention Fusion (DSAF) network to integrate features from these three modalities for CLNM prediction. Univariate and multivariate analyses revealed that Gender, Age, Size of US, and Extrathyroidal Extension (ETE) were independent risk factors for CLNM. In the internal test cohort (I-T), the area under the curve (AUC) of model was 0. 863 (95 % CI: 0. 779–0. 932). In the external test cohort (E-T), the AUC was 0. 839 (95 % CI: 0. 755–0. 905). Compared to all radiologists, the model significantly reduced both false-positive and false-negative rates in both the I-T and E-T. This study incorporates PVAT, which significantly enhances the performance of the multimodal deep learning model and may assist surgeons in making more informed and precise surgical decisions in the treatment of PTC.

JBHI Journal 2025 Journal Article

ULNC: Universal Language-Guided Nuclei Classification Via Instance-Aware Prompting

  • Kai Fan
  • Aiqiu Wu
  • Binbin Zheng
  • Anli Zhang
  • Ao Li
  • Minghui Wang

Accurate classification of nuclei is a crucial step in advancing pathology image analysis for disease diagnosis and treatment. Recently, prompt learning has shown great promise in universal nuclei classification for multiple datasets, as it can harvest the common knowledge by representing nuclei semantics across different sources. However, remarkable intra- and cross-dataset variability in nuclei categories and their complicated semantic relationships with numerous nuclei are often observed in distinct image instances, leading to two key challenges of instance variability and semantics ambiguity. In this paper, we propose a universal language-guided nuclei classification framework (ULNC) that leverages prompt-based language supervision to overcome these obstacles at the instance level. To address the instance variability issue, we introduce an innovative prompt learning approach that fully exploits unique contextual information of each image instance and generates instance-aware text embeddings highly adaptable to the categorical semantics of varied data sources. Additionally, to tackle the problem of semantics ambiguity, we employ a local vision-language matching loss that explicitly reinforces semantic connections between localized image regions and text prompts for nuclei categories, thus promoting the universal model's ability to learn discriminative image features for generalized nuclei classification. Extensive experiments conducted on several public datasets demonstrate that ULNC outperforms state-of-the-art methods in both accuracy and generalization to unseen domains, highlighting its potential for robust nuclei classification across diverse datasets.

EAAI Journal 2024 Journal Article

Acoustic tomography temperature reconstruction based on improved sparse reconstruction model and multi-scale feature fusion network

  • Xianghu Dong
  • Lifeng Zhang
  • Lifeng Qian
  • Chuanbao Wu
  • Zhihao Tang
  • Ao Li

Acoustic tomography is a widely used non-contact method for visualizing temperature distribution. A temperature distribution reconstruction algorithm based on an improved sparse reconstruction model and multi-scale feature fusion network is proposed. First, the acoustic temperature measurement sparse reconstruction model is improved by combining the error function (ERF) and the iterative reweighting algorithm, and the alternating direction method of multipliers algorithm (ADMM) is used to solve the model to obtain the initial temperature distribution. Then the feature extraction network is constructed to extract multi-scale features of acoustic time of flight (TOF) as prior information. Finally, the feature fusion reconstruction network is constructed to fuse and reconstruct the initial temperature distribution and multi-scale features to obtain a high-precision temperature distribution. Simulation and experimental tests were conducted respectively, and compared with other algorithms. The results show that the average relative error and root mean square error of the simulated temperature distribution reconstruction are 0. 073% and 0. 1% respectively, the average reconstruction error of the temperature points set in the experimental test is 0. 38%, and the reconstruction errors are lower than other algorithms. The proposed method effectively utilizes prior information to correct sparse reconstruction temperature distribution reconstruction results, significantly improving the quality of temperature distribution reconstruction.

IJCAI Conference 2024 Conference Paper

Benchmarking Fish Dataset and Evaluation Metric in Keypoint Detection - Towards Precise Fish Morphological Assessment in Aquaculture Breeding

  • Weizhen Liu
  • Jiayu Tan
  • Guangyu Lan
  • Ao Li
  • Dongye Li
  • Le Zhao
  • Xiaohui Yuan
  • Nanqing Dong

Accurate phenotypic analysis in aquaculture breeding necessitates the quantification of subtle morphological phenotypes. Existing datasets suffer from limitations such as small scale, limited species coverage, and inadequate annotation of keypoints for measuring refined and complex morphological phenotypes of fish body parts. To address this gap, we introduce FishPhenoKey, a comprehensive dataset comprising 23, 331 high-resolution images spanning six fish species. Notably, FishPhenoKey includes 22 phenotype-oriented annotations, enabling the capture of intricate morphological phenotypes. Motivated by the nuanced evaluation of these subtle morphologies, we also propose a new evaluation metric, Percentage of Measured Phenotypes (PMP). It is designed to assess the accuracy of individual keypoint positions and is highly sensitive to the phenotype measured using the corresponding keypoints. To enhance keypoint detection accuracy, we further propose a novel loss, Anatomically-Calibrated Regularization (ACR), that can be integrated into keypoint detection models, leveraging biological insights to refine keypoint localization. Our contributions set a new benchmark in fish phenotype analysis, addressing the challenges of precise morphological quantification and opening new avenues for research in sustainable aquaculture and genetic studies. Our dataset and code are available at https: //github. com/WeizhenLiuBioinform/FishPhenotype-Detect.

JBHI Journal 2024 Journal Article

Cross-Domain Nuclei Detection in Histopathology Images Using Graph-Based Nuclei Feature Alignment

  • Zhi Wang
  • Kai Fan
  • Xiaoya Zhu
  • Honglei Liu
  • Gang Meng
  • Minghui Wang
  • Ao Li

As powerful tools deep neural networks have been successfully adopted for nuclei detection in histopathology images, whereas require the same probability distribution between training and testing data. However, domain shift among histopathology images widely exists in real-world applications and severely deteriorates the detection performance of deep neural networks. Despite encouraging results of existing domain adaptation methods, there remain challenges for cross-domain nuclei detection task. First, in view of the tiny size of nuclei, it is actually very difficult to obtain sufficient nuclei features, thus leading to a negative influence for feature alignment. Second, due to unavailable annotations in target domain, some extracted features contain background pixels and are thereby indiscriminative, which can largely confuse the alignment procedure. To address these challenges, in this article, we propose an end-to-end graph-based nuclei feature alignment (GNFA) method for boosting cross-domain nuclei detection. Concretely, sufficient nuclei features are generated from nuclei graph convolutional network (NGCN) by aggregating information of adjacent nuclei upon construction of nuclei graph for successful alignment. In addition, importance learning module (ILM) is designed to further select discriminative nuclei features for mitigating negative influence of background pixels in target domain during alignment. By utilizing sufficient and discriminative node features generated from GNFA, our method can successfully perform feature alignment and effectively alleviate domain shift problem for nuclei detection. Extensive experiments of multiple adaptation scenarios reveal that our method achieves state-of-the-art performance in cross-domain nuclei detection compared with existing domain adaptation methods.

JBHI Journal 2024 Journal Article

DeScoD-ECG: Deep Score-Based Diffusion Model for ECG Baseline Wander and Noise Removal

  • Huayu Li
  • Gregory Ditzler
  • Janet Roveda
  • Ao Li

Objective: Electrocardiogram (ECG) signals commonly suffer noise interference, such as baseline wander. High-quality and high-fidelity reconstruction of the ECG signals is of great significance to diagnosing cardiovascular diseases. Therefore, this paper proposes a novel ECG baseline wander and noise removal technology. Methods: We extended the diffusion model in a conditional manner that was specific to the ECG signals, namely the Deep Score-Based Diffusion model for Electrocardiogram baseline wander and noise removal (DeScoD-ECG). Moreover, we deployed a multi-shots averaging strategy that improved signal reconstructions. We conducted the experiments on the QT Database and the MIT-BIH Noise Stress Test Database to verify the feasibility of the proposed method. Baseline methods are adopted for comparison, including traditional digital filter-based and deep learning-based methods. Results: The quantities evaluation results show that the proposed method obtained outstanding performance on four distance-based similarity metrics with at least 20% overall improvement compared with the best baseline method. Conclusion: This paper demonstrates the state-of-the-art performance of the DeScoD-ECG for ECG baseline wander and noise removal, which has better approximations of the true data distribution and higher stability under extreme noise corruptions. Significance: This study is one of the first to extend the conditional diffusion-based generative model for ECG noise removal, and the DeScoD-ECG has the potential to be widely used in biomedical applications.

IJCAI Conference 2024 Conference Paper

Revealing Hierarchical Structure of Leaf Venations in Plant Science via Label-Efficient Segmentation: Dataset and Method

  • Weizhen Liu
  • Ao Li
  • Ze Wu
  • Yue Li
  • Baobin Ge
  • Guangyu Lan
  • Shilin Chen
  • Minghe Li

Hierarchical leaf vein segmentation is a crucial but under-explored task in agricultural sciences, where analysis of the hierarchical structure of plant leaf venation can contribute to plant breeding. While current segmentation techniques rely on data-driven models, there is no publicly available dataset specifically designed for hierarchical leaf vein segmentation. To address this gap, we introduce the HierArchical Leaf Vein Segmentation (HALVS) dataset, the first public hierarchical leaf vein segmentation dataset. HALVS comprises 5, 057 real-scanned high-resolution leaf images collected from three plant species: soybean, sweet cherry, and London planetree. It also includes human-annotated ground truth for three orders of leaf veins, with a total labeling effort of 83. 8 person-days. Based on HALVS, we further develop a label-efficient learning paradigm that leverages partial label information, i. e. missing annotations for tertiary veins. Empirical studies are performed on HALVS, revealing new observations, challenges, and research directions on leaf vein segmentation. Our dataset and code are available at https: //github. com/WeizhenLiuBioinform/ HALVS-Hierarchical-Vein-Segment.

AIIM Journal 2022 Journal Article

Global and local attentional feature alignment for domain adaptive nuclei detection in histopathology images

  • Zhi Wang
  • Xiaoya Zhu
  • Ao Li
  • Yuan Wang
  • Gang Meng
  • Minghui Wang

Automated nuclei detection is crucial prerequisites for a number of histopathology related image analysis such as cancer diagnosis. Although existing deep learning based nuclei detection methods have achieved promising results, they cannot effectively deal with domain shift problem caused by different staining procedures and organ specific nuclear morphology. To handle this problem, in this paper a novel adversarial feature alignment method is proposed for domain adaptive nuclei detection, which includes both global alignment and local attentional alignment components to transfer the knowledge from source domain to target domain. Specifically, in local attentional alignment component, by using nuclei locations as guidance we extract local features and perform adversarial alignment. Furthermore, to address the issue that these local features from nuclei regions often contain insufficient information because of the small size of nuclei, we introduce an efficient location-aware self-attention (LocSA) module to refine local features by utilizing cues from all nuclei for obtaining discriminative features to perform successful feature alignment. Extensive experimental results are provided on two adaptation scenarios and our method demonstrates favorable performance against existing domain adaptation methods, which highlights the effectiveness of the proposed method for domain adaptive nuclei detection.

JBHI Journal 2022 Journal Article

Reconstruction-Assisted Feature Encoding Network for Histologic Subtype Classification of Non-Small Cell Lung Cancer

  • Haichun Li
  • Qilong Song
  • Dongqi Gui
  • Minghui Wang
  • Xuhong Min
  • Ao Li

Accurate histological subtype classification between adenocarcinoma (ADC) and squamous cell carcinoma (SCC) using computed tomography (CT) images is of great importance to assist clinicians in determining treatment and therapy plans for non-small cell lung cancer (NSCLC) patients. Although current deep learning approaches have achieved promising progress in this field, they are often difficult to capture efficient tumor representations due to inadequate training data, and in consequence show limited performance. In this study, we propose a novel and effective reconstruction-assisted feature encoding network (RAFENet) for histological subtype classification by leveraging an auxiliary image reconstruction task to enable extra guidance and regularization for enhanced tumor feature representations. Different from existing reconstruction-assisted methods that directly use generalizable features obtained from shared encoder for primary task, a dedicated task-aware encoding module is utilized in RAFENet to perform refinement of generalizable features. Specifically, a cascade of cross-level non-local blocks are introduced to progressively refine generalizable features at different levels with the aid of lower-level task-specific information, which can successfully learn multi-level task-specific features tailored to histological subtype classification. Moreover, in addition to widely adopted pixel-wise reconstruction loss, we introduce a powerful semantic consistency loss function to explicitly supervise the training of RAFENet, which combines both feature consistency loss and prediction consistency loss to ensure semantic invariance during image reconstruction. Extensive experimental results show that RAFENet effectively addresses the difficult issues that cannot be resolved by existing reconstruction-based methods and consistently outperforms other state-of-the-art methods on both public and in-house NSCLC datasets. Supplementary material is available at https://github.com/lhch1994/Rafenet_sup_material.

JBHI Journal 2020 Journal Article

A Novel MKL Method for GBM Prognosis Prediction by Integrating Histopathological Image and Multi-Omics Data

  • Ya Zhang
  • Ao Li
  • Jie He
  • Minghui Wang

Glioblastoma multiforme (GBM) is one of the most malignant brain tumors with very short prognosis expectation. To improve patients’ clinical treatment and their life quality after surgery, researches have developed tremendous in silico models and tools for predicting GBM prognosis based on molecular datasets and have earned great success. However, pathology still plays the most critical role in cancer diagnosis and prognosis in the clinic at present. Recent advancement of storing and processing histopathological images has drawn attention of researchers. Models based on histopathological images are developed, which show great potential for computer-aided pathological diagnoses. But models based on both molecular and histopathological images that could predict GBM prognosis with high accuracy are not present yet. In our previous research, we used the simple MKL method to integrate multi-omics data to improve GBM prognosis prediction successfully. In this paper, we have developed a novel multiple kernel learning (MKL) method, named histopathological integrating multiple kernel learning (HI-MKL), that could integrate both histopathological images and multi-omics data efficiently. By using datasets from The Cancer Genome Atlas project, we have built a system that could predict the GBM prognosis with high accuracy. Our research shows that HI-MKL is an accurate, robust, and generalized MKL method, which performs well in a GBM prognosis task.

v2026.09.13