Arrow Research search

Author name cluster

Fan Bai

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

5 papers
1 author row

Possible papers

5

NeurIPS Conference 2024 Conference Paper

SegVol: Universal and Interactive Volumetric Medical Image Segmentation

  • Yuxin Du
  • Fan Bai
  • Tiejun Huang
  • Bo Zhao

Precise image segmentation provides clinical study with instructive information. Despite the remarkable progress achieved in medical image segmentation, there is still an absence of a 3D foundation segmentation model that can segment a wide range of anatomical categories with easy user interaction. In this paper, we propose a 3D foundation segmentation model, named SegVol, supporting universal and interactive volumetric medical image segmentation. By scaling up training data to 90K unlabeled Computed Tomography (CT) volumes and 6K labeled CT volumes, this foundation model supports the segmentation of over 200 anatomical categories using semantic and spatial prompts. To facilitate efficient and precise inference on volumetric images, we design a zoom-out-zoom-in mechanism. Extensive experiments on 22 anatomical segmentation tasks verify that SegVol outperforms the competitors in 19 tasks, with improvements up to 37. 24\% compared to the runner-up methods. We demonstrate the effectiveness and importance of specific designs by ablation study. We expect this foundation model can promote the development of volumetric medical image analysis. The model and code are publicly available at https: //github. com/BAAI-DCAI/SegVol.

NeurIPS Conference 2024 Conference Paper

Touchstone Benchmark: Are We on the Right Way for Evaluating AI Algorithms for Medical Segmentation?

  • Pedro R. Bassi
  • Wenxuan Li
  • Yucheng Tang
  • Fabian Isensee
  • Zifu Wang
  • Jieneng Chen
  • Yu-Cheng Chou
  • Saikat Roy

How can we test AI performance? This question seems trivial, but it isn't. Standard benchmarks often have problems such as in-distribution and small-size test sets, oversimplified metrics, unfair comparisons, and short-term outcome pressure. As a consequence, good performance on standard benchmarks does not guarantee success in real-world scenarios. To address these problems, we present Touchstone, a large-scale collaborative segmentation benchmark of 9 types of abdominal organs. This benchmark is based on 5, 195 training CT scans from 76 hospitals around the world and 5, 903 testing CT scans from 11 additional hospitals. This diverse test set enhances the statistical significance of benchmark results and rigorously evaluates AI algorithms across various out-of-distribution scenarios. We invited 14 inventors of 19 AI algorithms to train their algorithms, while our team, as a third party, independently evaluated these algorithms on three test sets. In addition, we also evaluated pre-existing AI frameworks---which, differing from algorithms, are more flexible and can support different algorithms—including MONAI from NVIDIA, nnU-Net from DKFZ, and numerous other open-source frameworks. We are committed to expanding this benchmark to encourage more innovation of AI algorithms for the medical domain.

IJCAI Conference 2022 Conference Paper

C3-STISR: Scene Text Image Super-resolution with Triple Clues

  • Minyi Zhao
  • Miao Wang
  • Fan Bai
  • Bingjia Li
  • Jie Wang
  • Shuigeng Zhou

Scene text image super-resolution (STISR) has been regarded as an important pre-processing task for text recognition from low-resolution scene text images. Most recent approaches use the recognizer's feedback as clues to guide super-resolution. However, directly using recognition clue has two problems: 1) Compatibility. It is in the form of probability distribution, has an obvious modal gap with STISR - a pixel-level task; 2) Inaccuracy. it usually contains wrong information, thus will mislead the main task and degrade super-resolution performance. In this paper, we present a novel method C3-STISR that jointly exploits the recognizer's feedback, visual and linguistical information as clues to guide super-resolution. Here, visual clue is from the images of texts predicted by the recognizer, which is informative and more compatible with the STISR task; while linguistical clue is generated by a pre-trained character-level language model, which is able to correct the predicted texts. We design effective extraction and fusion mechanisms for the triple cross-modal clues to generate a comprehensive and unified guidance for super-resolution. Extensive experiments on TextZoom show that C3-STISR outperforms the SOTA methods in fidelity and recognition performance. Code is available in https: //github. com/zhaominyiz/C3-STISR.

AAAI Conference 2021 Conference Paper

GIF Thumbnails: Attract More Clicks to Your Videos

  • Yi Xu
  • Fan Bai
  • Yingxuan Shi
  • Qiuyu Chen
  • Longwen Gao
  • Kai Tian
  • Shuigeng Zhou
  • Huyang Sun

With the rapid increase of mobile devices and online media, more and more people prefer posting/viewing videos online. Generally, these videos are presented on video streaming sites with image thumbnails and text titles. While facing huge amounts of videos, a viewer clicks through a certain video with high probability because of its eye-catching thumbnail. However, current video thumbnails are created manually, which is time-consuming and quality-unguaranteed. And static image thumbnails contain very limited information of the corresponding videos, which prevents users from successfully clicking what they really want to view. In this paper, we address a novel problem, namely GIF thumbnail generation, which aims to automatically generate GIF thumbnails for videos and consequently boost their Click- Through-Rate (CTR). Here, a GIF thumbnail is an animated GIF file consisting of multiple segments from the video, containing more information of the target video than a static image thumbnail. To support this study, we build the first GIF thumbnails benchmark dataset that consists of 1070 videos covering a total duration of 69. 1 hours, and 5394 corresponding manually-annotated GIFs. To solve this problem, we propose a learning-based automatic GIF thumbnail generation model, which is called Generative Variational Dual- Encoder (GEVADEN). As not relying on any user interaction information (e. g. time-sync comments and real-time view counts), this model is applicable to newly-uploaded/rarelyviewed videos. Experiments on our built dataset show that GEVADEN significantly outperforms several baselines, including video-summarization and highlight-detection based ones. Furthermore, we develop a pilot application of the proposed model on an online video platform with 9814 videos covering 1231 hours, which shows that our model achieves a 37. 5% CTR improvement over traditional image thumbnails. This further validates the effectiveness of the proposed model and the promising application prospect of GIF thumbnails.

YNIMG Journal 2015 Journal Article

Rhesus monkey brain development during late infancy and the effect of phencyclidine: A longitudinal MRI and DTI study

  • Cirong Liu
  • Xiaoguang Tian
  • Huilang Liu
  • Yin Mo
  • Fan Bai
  • Xudong Zhao
  • Yuanye Ma
  • Jianhong Wang

Early brain development is a complex and rapid process, the disturbance of which may cause the onset of brain disorders. Based on longitudinal imaging data acquired from 6 to 16months postnatal, we describe a systematic trajectory of monkey brain development during late infancy, and demonstrate the influence of phencyclidine (PCP) on this trajectory. Although the general developmental trajectory of the monkey brain was close to that of the human brain, the development in monkeys was faster and regionally specific. Gray matter volume began to decrease during late infancy in monkeys, much earlier than in humans in whom it occurs in adolescence. Additionally, the decrease of gray matter volume in higher-order association regions (the frontal, parietal and temporal lobes) occurred later than in regions for primary functions (the occipital lobe and cerebellum). White matter volume displayed an increasing trend in most brain regions, but not in the occipital lobe, which had a stable volume. In addition, based on diffusion tensor imaging, we found an increase in fractional anisotropy and a decrease in diffusivity, which may be associated with myelination and axonal changes in white matter tracts. Meanwhile, we tested the influence of 14-day PCP treatment on the developmental trajectories. Such treatment tended to accelerated brain maturation during late infancy, although not statistically significant. These findings provide comparative information for the understanding of primate brain maturation and neurodevelopmental disorders.

v2026.09.13