Arrow Research search

Author name cluster

Fei Ma

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

19 papers
1 author row

Possible papers

19

AAAI Conference 2026 Conference Paper

A Principle-Driven Adaptive Policy for Group Cognitive Stimulation Dialogue for Elderly with Cognitive Impairment

  • Jiyue Jiang
  • Yanyu Chen
  • Pengan CHEN
  • Kai Liu
  • Jingqi Zhou
  • Zheyong Zhu
  • He Hu
  • Fei Ma

Cognitive impairment is becoming a major public health challenge. Cognitive Stimulation Therapy (CST) is an effective intervention for cognitive impairment, but traditional methods are difficult to scale, and existing digital systems struggle with group dialogues and cognitive stimulation principles. While Large Language Models (LLMs) are powerful, their application in this context faces key challenges: cognitive stimulation dialogue paradigms, a lack of therapeutic reasoning, and static-only user modeling. To address these issues, we propose a principle-driven adaptive policy actualized through a Group Cognitive Stimulation Dialogue (GCSD) system. We first construct a dataset with over 500 hours of real-world CST conversations and 10,000+ simulated dialogues generated via our Principle-Guided Scenario Simulation strategy. Our GCSD system then integrates four core modules to overcome LLM limitations: (i) a multi-speaker context controller to resolve role confusion; (ii) dynamic participant cognitive state modeling for personalized interaction; (iii) a cognitive stimulation-focused attention loss to instill cognitive stimulation reasoning; and (iv) a multi-dimensional reward strategy to enhance response value. Experimental results demonstrate that GCSD significantly outperforms baseline models across various evaluation metrics. Future work will focus on long-term clinical validation to bridge the gap between computational performance and clinical efficacy.

AAAI Conference 2026 Conference Paper

D-GARA: A Dynamic Benchmarking Framework for GUI Agent Robustness in Real-World Anomalies

  • Sen Chen
  • Tong Zhao
  • Yi Bin
  • Fei Ma
  • Wenqi Shao
  • Zheng Wang

Developing intelligent agents capable of operating a wide range of Graphical User Interfaces (GUIs) with human-level proficiency is a key milestone on the path toward Artificial General Intelligence. While most existing datasets and benchmarks for training and evaluating GUI agents are static and idealized, failing to reflect the complexity and unpredictability of real-world environments, particularly the presence of anomalies. To bridge this research gap, we propose D-GARA, a dynamic benchmarking framework, to evaluate Android GUI agent robustness in real-world anomalies. D-GARA introduces a diverse set of real-world anomalies that GUI agents commonly face in practice, including interruptions such as permission dialogs, battery warnings, and update prompts. Based on D-GARA framework, we construct and annotate a benchmark featuring commonly used Android applications with embedded anomalies to support broader community research. Comprehensive experiments and results demonstrate substantial performance degradation in state-of-the-art GUI agents when exposed to anomaly-rich environments, highlighting the need for robustness-aware learning. D-GARA is modular and extensible, supporting the seamless integration of new tasks, anomaly types, and interaction scenarios to meet specific evaluation goals.

EAAI Journal 2026 Journal Article

For automated cell culture systems: A high-speed style-transfer and prior-guided network for high-precision cell monitoring in bright-field microscopy

  • Jing Meng
  • Xiaolin Zhang
  • Jue Hou
  • Sumin Qi
  • Fei Ma
  • Zenan Wang
  • Dengwang Li

Bright-field (BF) microscopy is commonly used for cell growth monitoring in spatially constrained automated culture systems, but its low image contrast and diverse cell morphologies limit imaging clarity and the accuracy of confluence analysis. To address this, this study proposes a lightweight BF to phase contrast (PC) style transfer network (LW-P2P), which achieves fast and robust contrast enhancement through a brightness centralization method and a multi-criterion optimization strategy. Building on this, a MA-TPP segmentation network is developed, it incorporates a texture prior prompt (TPP) and a mixed attention mechanism (MA) to significantly enhance the recognition of low-contrast and halo-affected regions and enable fully automated confluence calculation. Evaluated on four datasets spanning five cell lines under varied densities and defocus conditions, LW-P2P improves signal-to-noise ratio by 21. 02% and achieves 8 × faster inference than state-of-the-art methods. MA-TPP Net achieves state-of-the-art segmentation on three diverse datasets covering six cell types, reducing confluence error to 0. 95%—a 26. 36% improvement over prior best methods. This study establishes an efficient and intelligent imaging-analysis framework for automated cell culture monitoring, with promising potential in regenerative medicine and related fields.

EAAI Journal 2026 Journal Article

Restoring neural radiance fields performance under adverse weather conditions

  • Ying He
  • Gan Chen
  • F. Richard Yu
  • Ming Li
  • Fei Ma
  • Guang Zhou

Neural Radiance Fields (NeRFs) have emerged as a powerful paradigm for modeling complex, photorealistic three-dimensional environments, garnering significant attention in the field of scene-based robotic localization. Regrettably, environments encountered in robotic applications are frequently susceptible to adverse weather conditions (e. g. , rain, snow, fog). Under such conditions, the inherent quality degradation introduced by existing image restoration algorithms severely disrupts the spatial consistency reconstruction of NeRFs. To address this challenge, this paper proposes a novel methodology for three-dimensional Scene Reconstruction under adverse weather, termed WeatherNeRF, which seamlessly integrates an image restoration algorithm with the neural radiance field framework. The proposed method effectively leverages the image restoration algorithm to process input image sequences affected by inclement weather. To mitigate the introduction of invalid artifacts by these processed sequences during scene reconstruction, we employ two regularization functions specifically designed to enhance scene compactness. Furthermore, to bolster scene consistency and facilitate effective scene restoration, we incorporate two-dimensional prior knowledge extracted from an image restoration model—WeatherDiffusion—during the reconstruction process, utilizing Score Distillation Sampling (SDS). Comprehensive experimental evaluations demonstrate that the proposed WeatherNeRF framework effectively restores neural radiance fields in everyday scenes degraded by adverse weather conditions and is capable of synthesizing high-fidelity novel view images. The code and data are publicly available at https: //github. com/C2022G/WeatherNeRF.

EAAI Journal 2026 Journal Article

Transformer-based explicit model predictive control with variable prediction horizon

  • Sichao Wu
  • Jiang Wu
  • Xingyu Cao
  • Fawang Zhang
  • Guangyuan Yu
  • Junjie Zhao
  • Yue Qu
  • Fei Ma

Traditional online Model Predictive Control (MPC) methods often suffer from excessive computational complexity, limiting their practical deployment. Explicit MPC mitigates online computational load by pre-computing control policies offline, however, existing explicit MPC methods typically rely on simplified system dynamics and cost functions, restricting their accuracy for complex systems. This paper proposes a novel Transformer-based explicit MPC algorithm (TransMPC) capable of generating highly accurate control sequences in real-time for complex dynamic systems. The Transformer is a deep learning architecture utilizing self-attention mechanisms to process sequential data. Specifically, we formulate the MPC policy as an encoder-only Transformer leveraging bidirectional self-attention, enabling simultaneous inference of entire control sequences in a single forward pass. This design inherently accommodates variable prediction horizons while ensuring low inference latency. Furthermore, we introduce a direct policy optimization framework that alternates between sampling and learning phases. Unlike imitation-based approaches dependent on precomputed optimal trajectories, TransMPC directly optimizes the true finite-horizon cost via automatic differentiation. Random horizon sampling combined with a replay buffer provides independent and identically distributed (i. i. d.) training samples, ensuring robust generalization across varying states and horizon lengths. Extensive simulations and real-world multi-platform experiments demonstrate that TransMPC achieves up to 81. 97%–516. 74% faster inference than recurrent-based methods, while maintaining high accuracy in both tracking and manipulation tasks. Specifically, in the mobile robot trajectory tracking task, TransMPC achieves lateral tracking errors as low as 0. 008 m, along with millimeter-level positioning and sub-degree orientation accuracy on a 7-degrees of freedom drill arm.

IJCAI Conference 2025 Conference Paper

Active Multimodal Distillation for Few-shot Action Recognition

  • Weijia Feng
  • Yichen Zhu
  • Ruojia Zhang
  • Chenyang Wang
  • Fei Ma
  • Xiaobao Wang
  • Xiaobai Li

Owing to its rapid progress and broad application prospects, few-shot action recognition has attracted considerable interest. However, current methods are predominantly based on limited single-modal data, which does not fully exploit the potential of multimodal information. This paper presents a novel framework that actively identifies reliable modalities for each sample using task-specific contextual cues, thus significantly improving recognition performance. Our framework integrates an Active Sample Inference (ASI) module, which utilizes active inference to predict reliable modalities based on posterior distributions and subsequently organizes them accordingly. Unlike reinforcement learning, active inference replaces rewards with evidence-based preferences, making more stable predictions. Additionally, we introduce an active mutual distillation module that enhances the representation learning of less reliable modalities by transferring knowledge from more reliable ones. Adaptive multimodal inference is employed during the meta-test to assign higher weights to reliable modalities. Extensive experiments across multiple benchmarks demonstrate that our method significantly outperforms existing approaches.

EAAI Journal 2025 Journal Article

Adaptive Patch Contrast for Weakly Supervised Semantic Segmentation

  • Wangyu Wu
  • Tianhong Dai
  • Zhenhong Chen
  • Xiaowei Huang
  • Jimin Xiao
  • Fei Ma
  • Renrong Ouyang

Weakly Supervised Semantic Segmentation (WSSS), using only image-level labels, has garnered significant attention due to its cost-effectiveness. Typically, the framework involves using image-level labels as training data to generate pixel-level pseudo-labels with refinements. Recently, methods based on Vision Transformers (ViT) have demonstrated superior capabilities in generating reliable pseudo-labels, particularly in recognizing complete object regions. However, current ViT-based approaches have some limitations in the use of patch embeddings, being prone to being dominated by certain abnormal patches, as well as many multi-stage methods being time-consuming and lengthy in training, thus lacking efficiency. Therefore, in this paper, we introduce a novel ViT-based WSSS method named Adaptive Patch Contrast (APC) that significantly enhances patch embedding learning for improved segmentation effectiveness. APC utilizes an Adaptive-K Pooling (AKP) layer to address the limitations of previous max pooling selection methods. Additionally, we propose a Patch Contrastive Learning (PCL) to enhance patch embeddings, thereby further improving the final results. We developed an end-to-end single-stage framework without CAM, which improved training efficiency. Experimental results demonstrate that our method performs exceptionally well on public datasets, outperforming other state-of-the-art WSSS methods with a shorter training time.

EAAI Journal 2025 Journal Article

Dynamic momentum contrastive learning network for diabetic retinopathy grading

  • Yanfei Guo
  • Chenglong Yang
  • Hangli Du
  • Yuanke Zhang
  • Fei Ma
  • Shasha Yuan

Diabetic retinopathy is the leading cause of blindness among the global working population. Automated and accurate grading of diabetic retinopathy is crucial for the diagnosis and treatment of retinal diseases. However, challenges arise in grading due to class imbalance, inter-class similarity, intra-class variability, and the small scale of lesions in different stages of diabetic retinopathy. To address the issues of class imbalance and small lesion scales, the dynamic momentum contrastive learning network (DMCLNet) for diabetic retinopathy grading is proposed in this paper. Firstly, a new strategy is introduced for constructing positive and negative samples to enhance individual differences. Then, an encoder and a momentum encoder are used to extract features from the main image, the positive and negative samples, respectively. The dynamic balancing strategy is presented to update the multi-class feature queue and calculate the similarity matrix between the main image and each positive and negative sample for contrastive learning. A dual dimensional loss function guides the training of the proposed model to fully capture the subtle differences between images of different classes, which improves inter-class discrimination ability. Finally, channel attention and spatial attention mechanisms are applied to enhance disease-related fine-grained features and suppress irrelevant redundant information, which accurately identifies small-scale lesions and improves the precision of severity grading. Extensive comparative experiments and ablation studies on three public datasets demonstrate that DMCLNet outperforms other state-of-the-art methods. It can achieve superior diabetic retinopathy grading performance across multiple datasets without relying on pixel-level lesion annotation, with high accuracy and generalization.

IJCAI Conference 2025 Conference Paper

Inter3D: A Benchmark and Strong Baseline for Human-Interactive 3D Object Reconstruction

  • Gan Chen
  • Ying He
  • Mulin Yu
  • F. Richard Yu
  • Gang Xu
  • Fei Ma
  • Ming Li
  • Guang Zhou

Recent advancements in implicit 3D reconstruction methods, e. g. , neural rendering fields and Gaussian splatting, have primarily focused on novel view synthesis of static or dynamic objects with continuous motion states. However, these approaches struggle to efficiently model a human-interactive object with n movable parts, requiring 2^n separate models to represent all discrete states. To overcome this limitation, we propose Inter3D, a new benchmark and approach for novel state synthesis of human-interactive objects. We introduce a self-collected dataset featuring commonly encountered interactive objects and a new evaluation pipeline, where only individual part states are observed during training, while part combination states remain unseen. We also propose a strong baseline approach that leverages Space Discrepancy Tensors to efficiently modelling all states of an object. To alleviate the impractical constraints on camera trajectories across training states, we propose a Mutual State Regularization mechanism to enhance the spatial density consistency of movable parts. In addition, we explore two occupancy grid sampling strategies to facilitate training efficiency. We conduct extensive experiments on the proposed benchmark, showcasing the challenges of the task and the superiority of our approach. The code and data are publicly available at https: //github. com/Inter3D-ui/Inter3D.

AAAI Conference 2025 Conference Paper

PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head Synthesis

  • Yifan Xie
  • Tao Feng
  • Xin Zhang
  • Xiangyang Luo
  • Zixuan Guo
  • Weijiang Yu
  • Heng Chang
  • Fei Ma

Talking head synthesis with arbitrary speech audio is a crucial challenge in the field of digital humans. Recently, methods based on radiance fields have received increasing attention due to their ability to synthesize high-fidelity and identity-consistent talking heads from just a few minutes of training video. However, due to the limited scale of the training data, these methods often exhibit poor performance in audio-lip synchronization and visual quality. In this paper, we propose a novel 3D Gaussian-based method called PointTalk, which constructs a static 3D Gaussian field of the head and deforms it in sync with the audio. It also incorporates an audio-driven dynamic lip point cloud as a critical component of the conditional information, thereby facilitating the effective synthesis of talking heads. Specifically, the initial step involves generating the corresponding lip point cloud from the audio signal and capturing its topological structure. The design of the dynamic difference encoder aims to capture the subtle nuances inherent in dynamic lip movements more effectively. Furthermore, we integrate the audio-point enhancement module, which not only ensures the synchronization of the audio signal with the corresponding lip point cloud within the feature space, but also facilitates a deeper understanding of the interrelations among cross-modal conditional features. Extensive experiments demonstrate that our method achieves superior high-fidelity and audio-lip synchronization in talking head synthesis compared to previous methods.

AAAI Conference 2025 Conference Paper

ReMask-Animate: Refined Character Image Animation Using Mask-Guided Adapters

  • Xunzhi Xiang
  • Haiwei Xue
  • Zonghong Dai
  • Di Wang
  • Minglei Li
  • Ye Yue
  • Fei Ma
  • Weijiang Yu

Pose-controlled human video generation is of significant interest and finds extensive applications in areas such as automated advertising and content creation on social media platforms. While existing methods employing pose sequences and reference images for human image animation have exhibited notable performance, they tend to encounter issues such as specific region blurring, background sharpening, and decreased identity consistency. In this paper, we introduce ReMask-Animate, which utilizes masks as additional priors to guide the model's local visual attention to specific areas, thereby alleviating feature confusion between different regions of the image. Three distinct mask-guided adapters are designed for cross-condition regional fusion of hand and face pose features, mitigating feature confusion between the foreground and background, and enhancing the visual consistency of character identity. Moreover, these lightweight adapters introduce minimal computational overhead and can be seamlessly integrated into specific layers of the backbone architecture. Extensive experiments show that our method outperforms state-of-the-art methods on five metrics in public datasets. Additionally, qualitative evaluations highlight a significant improvement in the quality of generated videos, demonstrating our approach's superiority.

NeurIPS Conference 2025 Conference Paper

RoMa: A Robust Model Watermarking Scheme for Protecting IP in Diffusion Models

  • Yingsha Xie
  • Rui Min
  • Zeyu Qin
  • Fei Ma
  • Li Shen
  • Fei Yu
  • Xiaochun Cao

Preserving intellectual property (IP) within a pre-trained diffusion model is critical for protecting the model's copyright and preventing unauthorized model deployment. In this regard, model watermarking is a common practice for IP protection that embeds traceable information within models and allows for further verification. Nevertheless, existing watermarking schemes often face challenges due to their vulnerability to fine-tuning, limiting their practical application in general pre-training and fine-tuning paradigms. Inspired by using mode connectivity to analyze model performance between a pair of connected models, we investigate watermark vulnerability by leveraging Linear Mode Connectivity (LMC) as a proxy to analyze the fine-tuning dynamics of watermark performance. Our results show that existing watermarked models tend to converge to sharp minima in the loss landscape, thus making them vulnerable to fine-tuning. To tackle this challenge, we propose RoMa, a Ro bust M odel w a termarking scheme that improves the robustness of watermarks against fine-tuning. Specifically, RoMa decomposes watermarking into two components, including Embedding Functionality, which preserves reliable watermark detection capability, and Path-specific Smoothness, which enhances the smoothness along the watermark-connected path to improve robustness. Extensive experiments on benchmark datasets MS-COCO-2017 and CUB-200-2011 demonstrate that RoMa significantly improves watermark robustness against fine-tuning while maintaining generation quality, outperforming baselines. The code is available at https: //github. com/xiekks/RoMa.

AAAI Conference 2025 Conference Paper

Subgraph Invariant Learning Towards Large-Scale Graph Node Classification

  • Leilei Wang
  • Si Shi
  • Fei Ma
  • Fei Richard Yu
  • Pengteng Li
  • Ying Tiffany He

Graph Neural Networks (GNNs) have shown efficacy in graph node classification, but face computational challenges on large-scale graphs. Although existing graph reduction methods address these issues, they still require high computational resources and fail to prioritize robust performance on out-of-distribution data. To tackle these challenges, we introduce the subgraph invariant learning paradigm, inspired by the small-world phenomenon. This approach enables models trained on specific subgraphs to generalize across diverse subgraphs, reducing computational demands, and enhancing scalability. To promote generalization, we maximize the invariance log-likelihood by deriving a theoretical lower bound of it and formulating the InVar loss. This loss minimizes the discrepancy between node representations and their corresponding invariance representations while maximizing the entropy of the node representation. In response to InVar loss, we propose the Invariance Facilitation Model (IFM), comprising the Invariance Representation Encoder (IRE) and Node Representation Encoder (NRE). IRE, capturing the invariance representations, utilizes Invariance ATTention (InvarATT) to compress long-range dependencies, while NRE learns the node representation, by integrating invariance representations via Telematic ATTention (TeleATT) and exchanging local information within each subgraph through GNNs. Evaluations on four large-scale graph datasets demonstrate the effectiveness, computational efficiency, and interpretability of IFM for large-scale graph node classification.

NeurIPS Conference 2025 Conference Paper

Universal Visuo-Tactile Video Understanding for Embodied Interaction

  • Yifan Xie
  • Mingyang Li
  • Shoujie Li
  • Xingting Li
  • Guangyu Chen
  • Fei Ma
  • Fei Yu
  • Wenbo Ding

Tactile perception is essential for embodied agents to understand the physical attributes of objects that cannot be determined through visual inspection alone. While existing methods have made progress in visual and language modalities for physical understanding, they fail to effectively incorporate tactile information that provides crucial haptic feedback for real-world interaction. In this paper, we present VTV-LLM, the first multi-modal large language model that enables universal Visuo-Tactile Video (VTV) understanding, bridging the gap between tactile perception and natural language. To address the challenges of cross-sensor and cross-modal integration, we contribute VTV150K, a comprehensive dataset comprising 150, 000 video frames from 100 diverse objects captured across three different tactile sensors (GelSight Mini, DIGIT, and Tac3D), annotated with four fundamental tactile attributes (hardness, protrusion, elasticity, and friction). We develop a novel three-stage training paradigm that includes VTV enhancement for robust visuo-tactile representation, VTV-text alignment for cross-modal correspondence, and text prompt finetuning for natural language generation. Our framework enables sophisticated tactile reasoning capabilities including feature assessment, comparative analysis, and scenario-based decision-making. Extensive experimental evaluations demonstrate that VTV-LLM achieves superior performance in tactile reasoning tasks, establishing a foundation for more intuitive human-machine interaction in tactile domains.

IJCAI Conference 2025 Conference Paper

VideoHumanMIB: Unlocking Appearance Decoupling for Video Human Motion In-betweening

  • Haiwei Xue
  • Zhensong Zhang
  • Minglei Li
  • Zonghong Dai
  • Fei Yu
  • Fei Ma
  • Zhiyong Wu

We propose VideoHumanMIB, a novel framework for Video Human Motion In-betweening that enables seamless transitions between different motion video clips, facilitating the generation of longer and more natural digital human videos. While existing video frame interpolation methods work well for similar motions in adjacent frames, they often struggle with complex human movements, resulting in artifacts and unrealistic transitions. To address these challenges, we introduce a two-stage approach: First, we design an Appearance Reconstruction AutoEncoder to decouple appearance and motion information, extracting robust appearance-invariant features. Second, we develop an enhanced diffusion pretrained network that leverages both motion optical flow and human pose as guidance conditions, enabling the model to learn comprehensive latent distributions of possible motions. Rather than operating directly in pixel space, our model works in a learned latent space, allowing it to better capture the underlying motion dynamics. The framework is optimized with a dual-frame constraint loss and a motion flow loss to ensure temporal consistency and natural movement transitions. Extensive experiments demonstrate that our approach generates highly realistic transition sequences that significantly outperform existing methods, particularly in challenging scenarios with large motion variations. The proposed VideoHumanMIB establishes a new baseline for human motion synthesis and enables more natural and controllable digital human animation.

EAAI Journal 2023 Journal Article

Deep ensemble learning for high-dimensional subsurface fluid flow modeling

  • Abouzar Choubineh
  • Jie Chen
  • David A. Wood
  • Frans Coenen
  • Fei Ma

The accuracy of Deep Learning (DL) algorithms can be improved by combining several deep learners into an ensemble. This avoids the continuous endeavor required to adjust the architecture of individual networks or the nature of the propagation. This study investigates prediction improvements possible using Deep Ensemble Learning (DEL) to determine four distinct multiscale basis functions in the mixed Generalized Multiscale Finite Element Method (GMsFEM), involving the permeability field as the only input. 376, 250 samples were initially generated, filtered down to 367, 811 after data pre-processing. A standard Convolutional Neural Network (CNN) named SkiplessCNN and three skip connection-based CNNs named FirstSkipCNN, MidSkipCNN, and DualSkipCNN were developed for the base learners. For each basis function, these four CNNs were combined into an ensemble model using linear regression and ridge regression, separately, as part of the stacking technique. A comparison of the coefficient of determination (R 2 ) and Mean Squared Error (MSE) confirms the effectiveness of all three skip connections in enhancing the performance of the standard CNN, with DualSkip being the most effective among them. Additionally, as evaluated on the testing subset, the combined models meaningfully outperform the individual models for all basis functions. The case that applies linear regression delivers R 2 ranging from 0. 8456 to 0. 9191 and MSE ranging from 0. 0092 to 0. 0369. The ridge regression case achieves marginally better predictions with R 2 ranging from 0. 8539 to 0. 922, and MSE ranging from 0. 009 to 0. 0349 because its solution involves more evenly distributed weights.

IJCAI Conference 2021 Conference Paper

Pairwise Half-graph Discrimination: A Simple Graph-level Self-supervised Strategy for Pre-training Graph Neural Networks

  • Pengyong Li
  • Jun Wang
  • Ziliang Li
  • Yixuan Qiao
  • Xianggen Liu
  • Fei Ma
  • Peng Gao
  • Sen Song

Self-supervised learning has gradually emerged as a powerful technique for graph representation learning. However, transferable, generalizable, and robust representation learning on graph data still remains a challenge for pre-training graph neural networks. In this paper, we propose a simple and effective self-supervised pre-training strategy, named Pairwise Half-graph Discrimination (PHD), that explicitly pre-trains a graph neural network at graph-level. PHD is designed as a simple binary classification task to discriminate whether two half-graphs come from the same source. Experiments demonstrate that the PHD is an effective pre-training strategy that offers comparable or superior performance on 13 graph classification tasks compared with state-of-the-art strategies, and achieves notable improvements when combined with node-level strategies. Moreover, the visualization of learned representation revealed that PHD strategy indeed empowers the model to learn graph-level knowledge like the molecular scaffold. These results have established PHD as a powerful and effective self-supervised learning strategy in graph-level representation learning.

JBHI Journal 2020 Journal Article

Length-of-Stay Prediction for Pediatric Patients With Respiratory Diseases Using Decision Tree Methods

  • Fei Ma
  • Limin Yu
  • Lishan Ye
  • David D. Yao
  • Weifen Zhuang

Accurate prediction of a patient's length-of-stay (LOS) in the hospital enables an efficient and effective management of hospital beds. This paper studies LOS prediction for pediatric patients with respiratory diseases using three decision tree methods: Bagging, Adaboost, and Random forest. A data set of 11, 206 records retrieved from the hospital information system is used for analysis after preprocessing and transformation through a computation and an expansion method. Two tests, namely bisection test and periodic test, are designed to assess the performance of the prediction methods. Bagging shows the best result on the bisection test (0. 296 RMSE, 0. 831 R 2, and 0. 723 Acc ± 1) for the testing set of the whole data test. The performances of the three methods are similar on the periodic test, whereas Adaboost performs slightly better than the other two methods. Results indicate that the three methods are all effective for the LOS prediction. This study also investigates the importance of different data fields to the LOS prediction, and finds that hospital treatment-related data fields contribute more to the LOS prediction than other categories of fields.

TCS Journal 2018 Journal Article

An iteration method for computing the total number of spanning trees and its applications in graph theory

  • Fei Ma
  • Bing Yao

Calculating and analyzing the number of spanning trees of graphs (network models) is an important and interesting research project in wide variety of fields, such as mathematics, physics, theoretical computer science, chemistry and so on. The number of spanning trees of graphs (models) displays amounts of information on its structural features and also on some relevant dynamical properties, in particular network security, random walks and percolation. In this paper, firstly, due to lots of graphs (models) are built on the basis of various simple and small elements (components), we provide primarily some helpful network-operation, such as link-operation and merge-operation, to generate more realistic and complicated graphs (models). Secondly, considering reliability of fault-tolerance to random faults and of intrusion-tolerance to selectively remove attacks, synchronization capability and diffusion properties of networks, we present an iterative method (algorithm) for computing the total number of spanning trees. As a pellucid example, we apply our method to tree space and cycle space, notice that it is proved to be indeed a better tool. In order to reflect more widely practical meanings, we study its applications in graph theory, including ladder-graph with zero clustering coefficient, wheel-graph having nonzero clustering coefficient as constituent ingredients of maximal planar graphs. In the rest of this paper, we make a brief summary that the method described by us can be designed a program (algorithm) for obtaining easily the exact number of spanning trees of some models.

v2026.09.13