Arrow Research search

Author name cluster

Dong Ni

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

20 papers
1 author row

Possible papers

20

JBHI Journal 2026 Journal Article

OnUVS: An Online Motion Transfer Framework with Content-Texture Decoupling for High-Fidelity Ultrasound Video Synthesis

  • Han Zhou
  • Rusi Chen
  • Xin Yang
  • Ao Chang
  • Junxuan Yu
  • Yuhao Huang
  • Ruobing Huang
  • Xinrui Zhou

Ultrasound (US) imaging plays a crucial role in diagnosing heart and pelvic diseases, where sonographers tend to evaluate dynamic motion and structure. However, the scarcity of US videos for rare cases limits training opportunities for novice sonographers and deep learning models, hindering detection rates and clinical di agnostic applications. US video synthesis is a promising solution to this issue. Nevertheless, accurately imitating the intricate motion of the anatomy while preserving image f idelity presents a significant challenge. In this work, we propose OnUVS, a novel online feature-decoupling frame work for high-fidelity US video synthesis. First, to simulate realistic motion, we incorporate keypoints into anatomical learning through a weakly supervised training approach, which enhances motion representation and minimizes the need for fully annotated data. Second, we implement a dual decoder generator that effectively balances content and textural features of generated frames, significantly enhancing the image fidelity of US videos. Third, a multi-scale discriminator further refines the sharpness and fine details, ensuring high-fidelity video synthesis. Fourth, an online learning strategy is designed to smooth coherence between frames by constraining the keypoint trajectories during inference. Validation on echocardiographic and pelvic floor US datasets demonstrates that OnUVS outperforms existing methods, achieving a 22. 08% improvement in motion consistency (FVD) and 25. 04% in image fidelity (FID). To facilitate reproducibility, we publicly release the code of OnUVSat: https://github.com/LucyChen159/OnUVS.

AAAI Conference 2025 Conference Paper

AeroGTO: An Efficient Graph-Transformer Operator for Learning Large-Scale Aerodynamics of 3D Vehicle Geometries

  • Pengwei Liu
  • Pengkai Wang
  • Xingyu Ren
  • Hangjie Yuan
  • Zhongkai Hao
  • Chao Xu
  • Shengze Cai
  • Dong Ni

Obtaining high-precision aerodynamics in the automotive industry relies on large-scale simulations with computational fluid dynamics, which are generally time-consuming and computationally expensive. Recent advances in operator learning for partial differential equations offer promising improvements in terms of efficiency. However, capturing intricate physical correlations from extensive and varying geometries while balancing large-scale discretization and computational costs remains a significant challenge. To address these issues, we propose **AeroGTO**, an efficient graph-transformer operator designed specifically for learning large-scale aerodynamics in engineering applications. AeroGTO combines local feature extraction through message passing and global correlation capturing via projection-inspired attention, employing a frequency-enhanced graph neural network augmented with k-nearest neighbors to handle three-dimensional (3D) irregular geometries. Moreover, the transformer architecture adeptly manages multi-level dependencies with only linear complexity concerning the number of mesh points, enabling fast inference of the model. Given a car's 3D mesh, AeroGTO accurately predicts surface pressure and estimates drag. In comparisons with five advanced models, AeroGTO is extensively tested on two industry-standard benchmarks, Ahmed-Body and DrivAerNet, achieving a 7.36% improvement in surface pressure prediction and a 10.71% boost in drag coefficient estimation, with fewer FLOPs and only 1% of the parameters used by the prior leading method.

NeurIPS Conference 2025 Conference Paper

Uncertainty-Informed Meta Pseudo Labeling for Surrogate Modeling with Limited Labeled Data

  • Xingyu Ren
  • Pengwei Liu
  • Pengkai Wang
  • Guanyu Chen
  • Qinxin Wu
  • Dong Ni

Deep neural networks, particularly neural operators, provide an efficient alternative to costly simulations in surrogate modeling. However, their performance is often constrained by the need for large-scale labeled datasets, which are costly and challenging to acquire in many scientific domains. Semi-supervised learning reduces label reliance by leveraging unlabeled data yet remains vulnerable to noisy pseudo-labels that mislead training and undermine robustness. To address these challenges, we propose a novel framework, Uncertainty-Informed Meta Pseudo Labeling (UMPL). The core mechenism is to refine pseudo-label quality through uncertainty-informed feedback signals. Specifically, the teacher model generates pseudo labels via epistemic uncertainty, while the student model learns from these labels and provides feedback based on aleatoric uncertainty. This interplay forms a meta-learning loop where enhanced generalization and improved pseudo-label quality reinforce each other, enabling the student model to achieve more stable uncertainty estimation and leading to more robust training. Notably, This framework is model-agnostic and can be seamlessly integrated into various neural architectures, facilitating effective exploitation of unlabeled data to enhance generalization in distribution shifts and out-of-distribution scenarios. Extensive evaluations of four models across seven tasks covering steady state and transient prediction problems demonstrate that UMPL consistently outperforms the best existing semi-supervised regression methods. When using only 10% of the fully supervised training data, UMPL achieves a 14. 18% improvement, highlighting its strong effectiveness under limited supervision. Our codes are available at https: //github. com/small-dumpling/UMPL.

EAAI Journal 2024 Journal Article

A method for real-time optimal heliostat aiming strategy generation via deep learning

  • Sipei Wu
  • Dong Ni

Optimal aiming strategies are essential for efficient solar power tower technology operation. However, the high calculation complexity makes it difficult for existing optimization methods to solve the optimization problem in real-time directly. This work proposes a real-time optimal heliostat aiming strategy generation method via deep learning. First, a two-stage learning scheme where the neural network models are trained by genetic algorithm (GA) benchmark solutions to produce an optimal aiming strategy is presented. Then, an end-to-end model without needing GA solutions for training is developed and discussed. Furthermore, a robust end-to-end training method using randomly sampled flux maps is also proposed. The proposed models demonstrated comparable performance as GA with two orders of magnitude less computation time through case studies. Among the proposed models, the end-to-end model shows significantly better generalization ability than the pure data-driven two-stage model on the test set. A robust end-to-end model with data enhancement has better robustness on unseen flux maps.

IS Journal 2023 Journal Article

Spatial Anomaly Detection in Hyperspectral Imaging Using Optical Neural Networks

  • Lingfeng Liu
  • Dong Ni
  • Liankui Dai

Hyperspectral imaging (HSI) is a widely used technology, yet hard to implement in real-time anomaly detection due to its extensive data flow volume. An autoencoder structured hybrid optical–electrical neural network method is proposed in this work that realizes feature exaction and anomaly detection during the hyperspectral data acquisition process to address such issues. In the proposed method, a digital micromirror device functions as the core optical processor to extract low-level features from the hyperspectral data flow. Weight binarization and a conditional subnetwork are utilized to suit optical computation. Pretraining by artificial data is implemented to ease the training data burden. Case studies on defect detection and foreign object detection have demonstrated that the proposed method can significantly reduce the sampling time by orders of magnitude without loss of detection accuracy.

AAAI Conference 2022 Conference Paper

Detecting Human-Object Interactions with Object-Guided Cross-Modal Calibrated Semantics

  • Hangjie Yuan
  • Mang Wang
  • Dong Ni
  • Liangpeng Xu

Human-Object Interaction (HOI) detection is an essential task to understand human-centric images from a fine-grained perspective. Although end-to-end HOI detection models thrive, their paradigm of parallel human/object detection and verb class prediction loses two-stage methods’ merit: objectguided hierarchy. The object in one HOI triplet gives direct clues to the verb to be predicted. In this paper, we aim to boost end-to-end models with object-guided statistical priors. Specifically, We propose to utilize a Verb Semantic Model (VSM) and use semantic aggregation to profit from this object-guided hierarchy. Similarity KL (SKL) loss is proposed to optimize VSM to align with the HOI dataset’s priors. To overcome the static semantic embedding problem, we propose to generate cross-modality-aware visual and semantic features by Cross-Modal Calibration (CMC). The above modules combined composes Object-guided Cross-modal Calibration Network (OCN). Experiments conducted on two popular HOI detection benchmarks demonstrate the significance of incorporating the statistical prior knowledge and produce state-of-the-art performances. More detailed analysis indicates proposed modules serve as a stronger verb predictor and a more superior method of utilizing prior knowledge. The codes are available at https: //github. com/JacobYuan7/OCN- HOI-Benchmark.

JBHI Journal 2022 Journal Article

Improved Segmentation of Echocardiography With Orientation-Congruency of Optical Flow and Motion-Enhanced Segmentation

  • Wufeng Xue
  • Heng Cao
  • Junqiang Ma
  • Ti Bai
  • Tianfu Wang
  • Dong Ni

Quantification of left ventricular (LV) ejection fraction (EF) from echocardiography depends upon the identification of endocardium boundaries as well as the calculation of end-diastolic (ED) and end-systolic (ES) LV volumes. It's critical to segment the LV cavity for precise calculation of EF from echocardiography. Most of the existing echocardiography segmentation approaches either only segment ES and ED frames without leveraging the motion information, or the motion information is only utilized as an auxiliary task. To address the above drawbacks, in this work, we propose a novel echocardiography segmentation method which can effectively utilize the underlying motion information by accurately predicting optical flow (OF) fields. First, we devised a feature extractor shared by the segmentation and the optical flow sub-tasks for efficient information exchange. Then, we proposed a new orientation congruency constraint for the OF estimation sub-task by promoting the congruency of optical flow orientation between successive frames. Finally, we design a motion-enhanced segmentation module for the final segmentation. Experimental results show that the proposed method achieved state-of-the-art performance for EF estimation, with a Pearson correlation coefficient of 0. 893 and a Mean Absolute Error of 5. 20% when validated with echo sequences of 450 patients.

JBHI Journal 2022 Journal Article

Joint Landmark and Structure Learning for Automatic Evaluation of Developmental Dysplasia of the Hip

  • Xindi Hu
  • Limin Wang
  • Xin Yang
  • Xu Zhou
  • Wufeng Xue
  • Yan Cao
  • Shengfeng Liu
  • Yuhao Huang

The ultrasound (US) screening of the infant hip is vital for the early diagnosis of developmental dysplasia of the hip (DDH). The US diagnosis of DDH refers to measuring alpha and beta angles that quantify hip joint development. These two angles are calculated from key anatomical landmarks and structures of the hip. However, this measurement process is not trivial for sonographers and usually requires a thorough understanding of complex anatomical structures. In this study, we propose a multi-task framework to learn the relationships among landmarks and structures jointly and automatically evaluate DDH. Our multi-task networks are equipped with three novel modules. Firstly, we adopt Mask R-CNN as the basic framework to detect and segment key anatomical structures and add one landmark detection branch to form a new multi-task framework. Secondly, we propose a novel shape similarity loss to refine the incomplete anatomical structure prediction robustly and accurately. Thirdly, we further incorporate the landmark-structure consistent prior to ensure the consistency of the bony rim estimated from the segmented structure and the detected landmark. In our experiments, 1231 US images of the infant hip from 632 patients are collected, of which 247 images from 126 patients are tested. The average errors in alpha and beta angles are 2. 221 ${}^{\circ }$ and 2. 899 ${}^{\circ }$. About 93% and 85% estimates of alpha and beta angles have errors less than 5 degrees, respectively. Experimental results demonstrate that the proposed method can accurately and robustly realize the automatic evaluation of DDH, showing great potential for clinical application.

JBHI Journal 2022 Journal Article

Regional Cardiac Motion Scoring With Multi-Scale Motion-Based Spatial Attention

  • Wufeng Xue
  • Zejian Chen
  • Tianfu Wang
  • Shuo Li
  • Dong Ni

Regional cardiac motion scoring aims to classify the motion status of each myocardium segment into one of the four categories (normal, hypokinetic, akinetic, and dyskinetic) from multiple short-axis MR sequences. It is essential for prognosis and early diagnosis for various cardiac diseases. However, the complex motion procedure of the myocardium and the invisible pattern differences pose great challenges, leading to low performance for automatic methods. Most existing works mitigate the task by differentiating the normal motion patterns from the abnormal ones, without fine-grained motion scoring. We propose an effective method for the task of cardiac motion scoring by connecting a bottom-up and another top-down branch with a novel motion-based spatial attention module in multi-scale space. Specifically, we use the convolution blocks for low-level feature extraction that acts as a bottom-up mechanism, and the task of optical flow for explicit motion extraction that acts as a top-down mechanism for high-level allocation of spatial attention. To this end, a newly designed Multi-scale Motion-based Spatial Attention (MMSA) module is used as the pivot connecting the bottom-up part and the top-down part, and adaptively weight the low-level features according to the motion information. Experimental results on a newly constructed dataset of 1440 myocardium segments from 90 subjects demonstrate that the proposed MMSA can accurately analyze the regional myocardium motion, with accuracies of 79. 3% for 4-way motion scoring, 89. 0% for abnormality detection, and correlation of 0. 943 for estimation of motion score index. This work has great potential for practical assessmentof cardiac motion function.

NeurIPS Conference 2022 Conference Paper

RLIP: Relational Language-Image Pre-training for Human-Object Interaction Detection

  • Hangjie Yuan
  • Jianwen Jiang
  • Samuel Albanie
  • Tao Feng
  • Ziyuan Huang
  • Dong Ni
  • Mingqian Tang

The task of Human-Object Interaction (HOI) detection targets fine-grained visual parsing of humans interacting with their environment, enabling a broad range of applications. Prior work has demonstrated the benefits of effective architecture design and integration of relevant cues for more accurate HOI detection. However, the design of an appropriate pre-training strategy for this task remains underexplored by existing approaches. To address this gap, we propose $\textit{Relational Language-Image Pre-training}$ (RLIP), a strategy for contrastive pre-training that leverages both entity and relation descriptions. To make effective use of such pre-training, we make three technical contributions: (1) a new $\textbf{Par}$allel entity detection and $\textbf{Se}$quential relation inference (ParSe) architecture that enables the use of both entity and relation descriptions during holistically optimized pre-training; (2) a synthetic data generation framework, Label Sequence Extension, that expands the scale of language data available within each minibatch; (3) ambiguity-suppression mechanisms, Relation Quality Labels and Relation Pseudo-Labels, to mitigate the influence of ambiguous/noisy samples in the pre-training data. Through extensive experiments, we demonstrate the benefits of these contributions, collectively termed RLIP-ParSe, for improved zero-shot, few-shot and fine-tuning HOI detection performance as well as increased robustness to learning from noisy annotations. Code will be available at https: //github. com/JacobYuan7/RLIP.

JBHI Journal 2021 Journal Article

Learn Fine-Grained Adaptive Loss for Multiple Anatomical Landmark Detection in Medical Images

  • Guang-Quan Zhou
  • Juzheng Miao
  • Xin Yang
  • Rui Li
  • En-Ze Huo
  • Wenlong Shi
  • Yuhao Huang
  • Jikuan Qian

Automatic and accurate detection of anatomical landmarks is an essential operation in medical image analysis with a multitude of applications. Recent deep learning methods have improved results by directly encoding the appearance of the captured anatomy with the likelihood maps (i. e. , heatmaps). However, most current solutions overlook another essence of heatmap regression, the objective metric for regressing target heatmaps and rely on hand-crafted heuristics to set the target precision, thus being usually cumbersome and task-specific. In this paper, we propose a novel learning-to-learn framework for landmark detection to optimize the neural network and the target precision simultaneously. The pivot of this work is to leverage the reinforcement learning (RL) framework to search objective metrics for regressing multiple heatmaps dynamically during the training process, thus avoiding setting problem-specific target precision. We also introduce an early-stop strategy for active termination of the RL agent's interaction that adapts the optimal precision for separate targets considering exploration-exploitation tradeoffs. This approach shows better stability in training and improved localization accuracy in inference. Extensive experimental results on two different applications of landmark localization: 1) our in-house prenatal ultrasound (US) dataset and 2) the publicly available dataset of cephalometric X-Ray landmark detection, demonstrate the effectiveness of our proposed method. Our proposed framework is general and shows the potential to improve the efficiency of anatomical landmark detection.

AAAI Conference 2021 Conference Paper

Learning Visual Context for Group Activity Recognition

  • Hangjie Yuan
  • Dong Ni

Group activity recognition aims to recognize an overall activity in a multi-person scene. Previous methods strive to reason on individual features. However, they under-explore the person-specific contextual information, which is significant and informative in computer vision tasks. In this paper, we propose a new reasoning paradigm to incorporate global contextual information. Specifically, we propose two modules to bridge the gap between group activity and visual context. The first is Transformer based Context Encoding (TCE) module, which enhances individual representation by encoding global contextual information to individual features and refining the aggregated information. The second is Spatial-Temporal Bilinear Pooling (STBiP) module. It firstly further explores pairwise relationships for the context encoded individual representation, then generates semantic representations via gated message passing on a constructed spatial-temporal graph. On their basis, we further design a two-branch model that integrates the designed modules into a pipeline. Systematic experiments demonstrate each module’s effectiveness on either branch. Visualizations indicate that visual contextual cues can be aggregated globally by TCE. Moreover, our method achieves state-of-the-art results on two widely used benchmarks using only RGB images as input and 2D backbones.

JBHI Journal 2020 Journal Article

CR-Unet: A Composite Network for Ovary and Follicle Segmentation in Ultrasound Images

  • Haoming Li
  • Jinghui Fang
  • Shengfeng Liu
  • Xiaowen Liang
  • Xin Yang
  • Zixin Mai
  • Manh The Van
  • Tianfu Wang

Transvaginal ultrasound (TVUS) is widely used in infertility treatment. The size and shape of the ovary and follicles must be measured manually for assessing their physiological status by sonographers. However, this process is extremely time-consuming and operator-dependent. In this study, we propose a novel composite network, namely CR-Unet, to simultaneously segment the ovary and follicles in TVUS. The CR-Unet incorporates the spatial recurrent neural network (RNN) into a plain U-Net. It can effectively learn multi-scale and long-range spatial contexts to combat the challenges of this task, such as the poor image quality, low contrast, boundary ambiguity, and complex anatomy shapes. We further adopt deep supervision strategy to make model training more effective and efficient. In addition, self-supervision is employed to iteratively refine the segmentation results. Experiments on 3204 TVUS images from 219 patients demonstrate the proposed method achieved the best segmentation performance compared to other state-of-the-art methods for both the ovary and follicles, with a Dice Similarity Coefficient (DSC) of 0. 912 and 0. 858, respectively.

JBHI Journal 2019 Journal Article

Dense Deconvolutional Network for Skin Lesion Segmentation

  • Hang Li
  • Xinzi He
  • Feng Zhou
  • Zhen Yu
  • Dong Ni
  • Siping Chen
  • Tianfu Wang
  • Baiying Lei

Automatic delineation of skin lesion contours from dermoscopy images is a basic step in the process of diagnosis and treatment of skin lesions. However, it is a challenging task due to the high variation of appearances and sizes of skin lesions. In order to deal with such challenges, we propose a new dense deconvolutional network (DDN) for skin lesion segmentation based on residual learning. Specifically, the proposed network consists of dense deconvolutional layers (DDLs), chained residual pooling (CRP), and hierarchical supervision (HS). First, unlike traditional deconvolutional layers, DDLs are adopted to maintain the dimensions of the input and output images unchanged. The DDNs are trained in an end-to-end manner without the need of prior knowledge or complicated postprocessing procedures. Second, the CRP aims to capture rich contextual background information and to fuse multilevel features. By combining the local and global contextual information via multilevel feature fusion, the high-resolution prediction output is obtained. Third, HS is added to serve as an auxiliary loss and to refine the prediction mask. Extensive experiments based on the public ISBI 2016 and 2017 skin lesion challenge datasets demonstrate the superior segmentation results of our proposed method over the state-of-the-art methods.

JBHI Journal 2019 Journal Article

Neuroimaging Retrieval via Adaptive Ensemble Manifold Learning for Brain Disease Diagnosis

  • Baiying Lei
  • Peng Yang
  • Yinan Zhuo
  • Feng Zhou
  • Dong Ni
  • Siping Chen
  • Xiaohua Xiao
  • Tianfu Wang

Alzheimer's disease (AD) is a neurodegenerative and non-curable disease, with serious cognitive impairment, such as dementia. Clinically, it is critical to study the disease with multi-source data in order to capture a global picture of it. In this respect, an adaptive ensemble manifold learning (AEML) algorithm is proposed to retrieve multi-source neuroimaging data. Specifically, an objective function based on manifold learning is formulated to impose geometrical constraints by similarity learning. The complementary characteristics of various sources of brain disease data for disorder discovery are investigated by tuning weights from ensemble learning. In addition, a generalized norm is explicitly explored for adaptive sparseness degree control. The proposed AEML algorithm is evaluated by the public AD neuroimaging initiative database. Results obtained from the extensive experiments demonstrate that our algorithm outperforms the traditional methods.

JBHI Journal 2018 Journal Article

A Deep Convolutional Neural Network-Based Framework for Automatic Fetal Facial Standard Plane Recognition

  • Zhen Yu
  • Ee-Leng Tan
  • Dong Ni
  • Jing Qin
  • Siping Chen
  • Shengli Li
  • Baiying Lei
  • Tianfu Wang

Ultrasound imaging has become a prevalent examination method in prenatal diagnosis. Accurate acquisition of fetal facial standard plane (FFSP) is the most important precondition for subsequent diagnosis and measurement. In the past few years, considerable effort has been devoted to FFSP recognition using various hand-crafted features, but the recognition performance is still unsatisfactory due to the high intraclass variation of FFSPs and the high degree of visual similarity between FFSPs and other non-FFSPs. To improve the recognition performance, we propose a method to automatically recognize FFSP via a deep convolutional neural network (DCNN) architecture. The proposed DCNN consists of 16 convolutional layers with small 3 × 3 size kernels and three fully connected layers. A global average pooling is adopted in the last pooling layer to significantly reduce network parameters, which alleviates the overfitting problems and improves the performance under limited training data. Both the transfer learning strategy and a data augmentation technique tailored for FFSP are implemented to further boost the recognition performance. Extensive experiments demonstrate the advantage of our proposed method over traditional approaches and the effectiveness of DCNN to recognize FFSP for clinical diagnosis.

JBHI Journal 2018 Journal Article

Automatic Fetal Head Circumference Measurement in Ultrasound Using Random Forest and Fast Ellipse Fitting

  • Jing Li
  • Yi Wang
  • Baiying Lei
  • Jie-Zhi Cheng
  • Jing Qin
  • Tianfu Wang
  • Shengli Li
  • Dong Ni

Head circumference (HC) is one of the most important biometrics in assessing fetal growth during prenatal ultrasound examinations. However, the manual measurement of this biometric by doctors often requires substantial experience. We developed a learning-based framework that used prior knowledge and employed a fast ellipse fitting method (ElliFit) to measure HC automatically. We first integrated the prior knowledge about the gestational age and ultrasound scanning depth into a random forest classifier to localize the fetal head. We further used phase symmetry to detect the center line of the fetal skull and employed ElliFit to fit the HC ellipse for measurement. The experimental results from 145 HC images showed that our method had an average measurement error of 1. 7 mm and outperformed traditional methods. The experimental results demonstrated that our method shows great promise for applications in clinical practice.

AAAI Conference 2017 Conference Paper

Fine-Grained Recurrent Neural Networks for Automatic Prostate Segmentation in Ultrasound Images

  • Xin Yang
  • Lequan Yu
  • Lingyun Wu
  • Yi Wang
  • Dong Ni
  • Jing Qin
  • Pheng-Ann Heng

Boundary incompleteness raises great challenges to automatic prostate segmentation in ultrasound images. Shape prior can provide strong guidance in estimating the missing boundary, but traditional shape models often suffer from hand-crafted descriptors and local information loss in the fitting procedure. In this paper, we attempt to address those issues with a novel framework. The proposed framework can seamlessly integrate feature extraction and shape prior exploring, and estimate the complete boundary with a sequential manner. Our framework is composed of three key modules. Firstly, we serialize the static 2D prostate ultrasound images into dynamic sequences and then predict prostate shapes by sequentially exploring shape priors. Intuitively, we propose to learn the shape prior with the biologically plausible Recurrent Neural Networks (RNNs). This module is corroborated to be effective in dealing with the boundary incompleteness. Secondly, to alleviate the bias caused by different serialization manners, we propose a multi-view fusion strategy to merge shape predictions obtained from different perspectives. Thirdly, we further implant the RNN core into a multiscale Auto-Context scheme to successively refine the details of the shape prediction map. With extensive validation on challenging prostate ultrasound images, our framework bridges severe boundary incompleteness and achieves the best performance in prostate boundary delineation when compared with several advanced methods. Additionally, our approach is general and can be extended to other medical image segmentation tasks, where boundary incompleteness is one of the main challenges.

JBHI Journal 2017 Journal Article

Segmentation, Splitting, and Classification of Overlapping Bacteria in Microscope Images for Automatic Bacterial Vaginosis Diagnosis

  • Youyi Song
  • Liang He
  • Feng Zhou
  • Siping Chen
  • Dong Ni
  • Baiying Lei
  • Tianfu Wang

Quantitative analysis of bacterial morphotypes in the microscope images plays a vital role in diagnosis of bacterial vaginosis (BV) based on the Nugent score criterion. However, there are two main challenges for this task: 1) It is quite difficult to identify the bacterial regions due to various appearance, faint boundaries, heterogeneous shapes, low contrast with the background, and small bacteria sizes with regards to the image. 2) There are numerous bacteria overlapping each other, which hinder us to conduct accurate analysis on individual bacterium. To overcome these challenges, we propose an automatic method in this paper to diagnose BV by quantitative analysis of bacterial morphotypes, which consists of a three-step approach, i. e. , bacteria regions segmentation, overlapping bacteria splitting, and bacterial morphotypes classification. Specifically, we first segment the bacteria regions via saliency cut, which simultaneously evaluates the global contrast and spatial weighted coherence. And then Markov random field model is applied for high-quality unsupervised segmentation of small object. We then decompose overlapping bacteria clumps into markers, and associate a pixel with markers to identify evidence for eventual individual bacterium splitting. Next, we extract morphotype features from each bacterium to learn the descriptors and to characterize the types of bacteria using an Adaptive Boosting machine learning framework. Finally, BV diagnosis is implemented based on the Nugent score criterion. Experiments demonstrate that our proposed method achieves high accuracy and efficiency in computation for BV diagnosis.

JBHI Journal 2015 Journal Article

Standard Plane Localization in Fetal Ultrasound via Domain Transferred Deep Neural Networks

  • Hao Chen
  • Dong Ni
  • Jing Qin
  • Shengli Li
  • Xin Yang
  • Tianfu Wang
  • Pheng Ann Heng

Automatic localization of the standard plane containing complicated anatomical structures in ultrasound (US) videos remains a challenging problem. In this paper, we present a learning-based approach to locate the fetal abdominal standard plane (FASP) in US videos by constructing a domain transferred deep convolutional neural network (CNN). Compared with previous works based on low-level features, our approach is able to represent the complicated appearance of the FASP and hence achieve better classification performance. More importantly, in order to reduce the overfitting problem caused by the small amount of training samples, we propose a transfer learning strategy, which transfers the knowledge in the low layers of a base CNN trained from a large database of natural images to our task-specific CNN. Extensive experiments demonstrate that our approach outperforms the state-of-the-art method for the FASP localization as well as the CNN only trained on the limited US training samples. The proposed approach can be easily extended to other similar medical image computing problems, which often suffer from the insufficient training samples when exploiting the deep CNN to represent high-level features.

v2026.09.13