Arrow Research search

Author name cluster

Ying Chen

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

38 papers
2 author rows

Possible papers

38

EAAI Journal 2026 Journal Article

Clinician-informed offline reinforcement learning for vasopressor administration optimization in shock management

  • Feier Qiu
  • Ying Chen
  • Xiuxian Wang
  • Na Geng
  • Zhitao Yang

Vasopressor treatment strategies are essential for managing shock patients, yet determining optimal type, dosage, and timing of vasopressors remains challenging given variable clinician expertise and patient conditions. Applying existing reinforcement learning (RL) algorithms in treatment decision making risks Q-value overestimation and ignores the gap between artificial intelligence (AI)-driven recommendations and established clinical practices. This work introduces an offline RL algorithm called Safe Conservative Q-learning (SafeCQL), integrating conservative regularization and clinician-informed safety constraints. Patient trajectories from a large real-world intensive care database are modeled as a Markov decision process (MDP) to optimize vasopressor administration. Off-policy evaluations with model-based and model-free estimation methods demonstrate the superior performance of SafeCQL over clinician policies and existing RL models in improving survival rates and policy robustness. Findings show that SafeCQL improves the survival rate by 5. 1% and raises expected returns from 45. 78 to 69. 15 (model-based) and 44. 33 to 61. 24 (model-free). SafeCQL can derive a treatment policy that aligns closely with clinician preferences while surpassing clinician expertise in decision quality. This work offers a deployable solution for personalized vasopressor management in dynamic critical care environments.

EAAI Journal 2026 Journal Article

Decarbonization responsive scheduling interactions in cogeneration units based sustainable city energy ecosystem considering carbon capture and thermal storage device with digital social welfare

  • Jingxuan Dong
  • Ying Chen

Enhancing the flexibility and decarbonization of integrated electricity and heat systems embodies a significant challenge in the ongoing development of sustainable urban energy ecosystems. Conventional combined heat and power (CHP) systems have been plagued with substantially rigid thermal and electric couplings, limited responses to renewable variable energy source fluctuations, and a lack of coordination with carbon capture processes. This study proposes a new decision-theoretic game-based dual-stage dispatch model which leverages a Directional Flow Thermal Storage Module (DFTSM) integrated with carbon capture coordination in a digital optimization framework. This framework exploits a simulated day-ahead and real-time scheduling combination based on a Markov Decision Process (MDP) that enhances the flexibility of energy allocation under uncertainty and operational efficiency while coordinating thermal and electric demand. The DFTSM, designed out of a multi-stage Tesla valve, allows flexible bidirectional heat transfer and increased peak-load regulation ability. The game-based model adopts cooperative interaction between system agents and welfare-based decision-making. The results confirm the MDP-based decision-theoretic game framework delivers and operates with; 7. 8% less total operating cost, 48% less carbon, and 35% more renewable benefit than current systems. Ultimately, this work demonstrates the structure of our model can bridge economic performance, environmental responsibility, and digital intelligence in low-carbon cogeneration dispatch decision-making. The model creates an industry-first integration of the dynamics of physical storage, carbon-aware optimization, and digital welfare analytics, which offers a scalable opportunity to provide resilient, decarbonized, and socially responsible energy within urban communities.

EAAI Journal 2026 Journal Article

Enhancing small object detection in low-altitude remote sensing via high-resolution feature extraction and multi-scale fusion

  • Xinyuan Le
  • Ying Chen
  • Wei Zeng
  • Xiang Ao
  • Huiling Chen
  • Jingyan Xie

To address insufficient feature representation, redundancy in multi-scale fusion, and weak directional perception in low-altitude remote sensing images, this paper proposes a detection model based on convolutional neural networks and attention mechanisms, named Multi-Scale Fusion You Only Look Once (MSF-YOLO), with enhanced feature extraction and multi-scale fusion. First, the detection head hierarchy is reconstructed by adding a P2 tiny object head and removing the large object head, enhancing high-resolution feature utilization. Second, the Improved Selective Contour Aggregation (ISBA) module is designed to construct the Improved Selective Contour Aggregation Network (ISBANet), which dynamically adjusts fusion weights and performs consistency correction. Third, Enhanced Spatial and Directional Convolution (ESDConv) is proposed, improving small-object feature representation via spatial slicing-channel concatenation and multi-direction convolution while avoiding detail loss from traditional downsampling. Experimental results on the Vision Meets Drone Object Detection in Image Challenge (2019) (VisDrone-DET2019) dataset show that MSF-YOLO-n achieves detection accuracy close to You Only Look Once version 8x (YOLOv8x) with only 4% of its parameters and 18. 4% of its computational cost. Compared to baseline You Only Look Once version 8n (YOLOv8n), MSF-YOLO-n reduces parameters by 10% while increasing mean average precision at IoU thresholds 50–95% (mAP50-95) by 9. 3%, albeit with an expected increase in floating-point operations (FLOPs). Verification on the Unmanned Aerial Vehicle Benchmark Object Detection and Tracking (UAVDT) dataset confirms the model's generalization ability, demonstrating the effectiveness of the proposed artificial intelligence method for low-altitude remote sensing small-object detection.

AAAI Conference 2026 Conference Paper

GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and a Comprehensive Multimodal Dataset Towards General Medical AI

  • Tianbin Li
  • Yanzhou Su
  • Wei Li
  • Bin Fu
  • Zhe Chen
  • Ziyan Huang
  • Guoan Wang
  • Chenglong Ma

Despite significant advancements in general AI, its effectiveness in the medical domain is limited by the lack of specialized medical knowledge. To address this, we formulate GMAI-VL-5.5M, a multimodal medical dataset created by converting hundreds of specialized medical datasets with various annotations into high-quality image-text pairs. This dataset offers comprehensive task coverage, diverse modalities, and rich image-text data. Building upon this dataset, we develop GMAI-VL, a 7B-parameter general medical vision-language model, with a three-stage training strategy that enhances the integration of visual and textual information. This approach significantly improves the model's ability to process multimodal data, supporting accurate diagnoses and clinical decision-making. Experiments show that GMAI-VL achieves state-of-the-art performance across various multimodal medical tasks, including visual question answering and medical image diagnosis.

TAAS Journal 2026 Journal Article

Joint Resource Allocation and Task Slicing for Mobile Multimedia Computing in Edge-based Autonomous Systems

  • Jiwei Huang
  • Yajing Leng
  • Jiarong Bao
  • Songyuan Li
  • Ying Chen

Mobile multimedia applications such as real-time video processing, augmented reality, and mobile gaming have raised high requirements for low latency and high efficiency. Edge-based autonomous systems have become a key technology for processing these application tasks. This article focuses on joint resource allocation and task slicing for mobile multimedia computing in edge-based autonomous systems. We propose an efficient resource allocation and task slicing strategy, aiming at the optimization of the overall utility of both edge servers and mobile devices simultaneously. We transform the resource allocation problem into resource pricing and purchasing behaviors. We present a Stackelberg game model and prove theorems for the existence of equilibrium and optimality. Based on the theorems, we design an algorithm namely G-RPTSS for resource purchasing and computation task slicing. Then, we employ Deep Reinforcement Learning (DRL) techniques in resource pricing and propose the DRL-ESRP algorithm which is capable of adaptively responding to dynamic computational scenarios in edge-based autonomous systems. Our scheme leverages the DRL technique for autonomous learning and policy adjustment. Simulation experiments, based on real-world scenario data, demonstrate the superior of our approach in learning efficiency and performance advantages to existing both non-DRL and other DRL algorithms.

JBHI Journal 2026 Journal Article

Point-Supervised Coronary Semantic Segmentation in X-Ray Angiographic Images

  • Ying Chen
  • Danni Ai
  • Jianyu Du
  • Yuanyuan Wang
  • Tianyu Fu
  • Deqiang Xiao
  • Yucong Lin
  • Long Shao

Coronary semantic segmentation in X-ray angiography is essential for computer-aided diagnosis and treatment planning of coronary artery disease (CAD). Despite its importance, this task remains highly challenging due to the complex and interconnected vascular topology, as well as the similar visual characteristics among different branches, making dense pixel-level manual annotation difficult and labor-intensive. To alleviate this burden, we propose a point-supervised coronary semantic segmentation framework that significantly reduces annotation effort without compromising segmentation accuracy. The primary challenge of point label based supervision lies in the model's tendency to overfit sparse point labels, leading to limited generalization to pixel-level predictions. To enrich the supervision signals and stabilize the training process with the sparse point labels, we propose an adaptive foreground mask generation module and a region regularization strategy to ensure accurate semantic guidance while maximizing meaningful coverage of the vascular structures. To enhance coronary topology perception and branch differentiation, we propose a multi-task learning framework that jointly performs keypoint detection and coronary semantic segmentation through a shared feature extraction encoder and two task-specific decoders. The experimental results demonstrate that our point-supervised model achieves performance comparable to fully supervised model, and outperforms the existing state-of-the-art point-supervised semantic segmentation methods.

AAAI Conference 2026 Conference Paper

S2-UniSeg: Fast Universal Agglomerative Pooling for Scalable Segment Anything Without Supervision

  • Huihui Xu
  • Jin Ye
  • Hongqiu Wang
  • Changkai Ji
  • Jiashi Lin
  • Ming Hu
  • Ziyan Huang
  • Ying Chen

Recent self-supervised image segmentation models have achieved promising performance on semantic segmentation and class-agnostic instance segmentation. However, their pretraining schedule is multi-stage, requiring a time-consuming pseudo-masks generation process between each training epoch. This time-consuming offline process not only makes it difficult to scale with training dataset size, but also leads to sub-optimal solutions due to its discontinuous optimization routine. To solve these, we first present a novel pseudo-mask algorithm, Fast Universal Agglomerative Pooling (UniAP). Each layer of UniAP can identify groups of similar nodes in parallel, allowing to generate both semantic-level and instance-level and multi-granular pseudo-masks within ens of milliseconds for one image. Based on the fast UniAP, we propose the Scalable Self-Supervised Universal Segmentation (S2-UniSeg), which employs a student and a momentum teacher for continuous pretraining. A novel segmentation-oriented pretext task, Query-wise Self-Distillation (QuerySD), is proposed to pretrain S2-UniSeg to learn the local-to-global correspondences. Under the same setting, S2-UniSeg outperforms the SOTA UnSAM model, achieving notable improvements of AP+6.9 on COCO, AR+11.1 on UVO, PixelAcc+4.5 on COCOStuff-27, RQ+8.0 on Cityscapes. After scaling up to a larger 2M-image subset of SA-1B, S2-UniSeg further achieves performance gains on all four benchmarks.

EAAI Journal 2026 Journal Article

Self-supervised multi-level trajectory representation model for field-road trajectory segmentation

  • Xiaoqiang Zhang
  • Qianrun Wei
  • Bingbing Hu
  • Liwei Pan
  • Caicong Wu
  • Claus Aage Grøn Sørensen
  • Ying Chen
  • Kun Zhou

Classifying each point in global navigation satellite system positioning trajectories as either in-field or on-road is pivotal for analyzing the operational performance of agricultural vehicles. This paper introduces a field-road trajectory segmentation method which significantly enhances segmentation robustness through a self-supervised learning approach. The method involves pre-training a trajectory representation model with self-supervised learning, which is subsequently fine-tuned for trajectory segmentation applications. Our model employs dual encoders: point-level and trajectory-level to capture essential multi-level spatio-temporal features for accurate trajectory segmentation. The point-level encoder focuses on extracting detailed features for individual points and performing point-density classification, while the trajectory-level encoder enriches these features by integrating trajectory similarity computations. Meanwhile, utilizing a combination of Convolutional Neural Networks and Transformer networks, the model adeptly handles both temporal and spatial dependencies in trajectory data, crucial for dynamic adaptation to various trajectories. The accuracy of our method achieves 93. 95% and 89. 32% on two manually labeled datasets, respectively, and experiments on the raw trajectory dataset demonstrate that a pre-training trajectory representation model can effectively capture the trajectory characteristic. Extensive validation confirms the superior efficacy of the proposed method and its potential impact on the evaluation efficiency of practical agricultural operations. The source code is available at the following address: https: //github. com/peanut2code/PreTR-TS.

AAAI Conference 2026 Conference Paper

Trainable EEG Interpolation and Structure-Sharing Dual-Path Encoders for Brain-Assisted Target Speaker Extraction

  • Zhao Lv
  • Haoran Zhou
  • Ying Chen
  • Youdian Gao
  • Xinhui Li
  • Ruibo Fu
  • Cunhang Fan

Brain-assisted target speaker extraction (TSE) isolates a target speaker's voice from a mixture by leveraging task-specific representations in Electroencephalogram (EEG) signals. However, existing methods rely on fixed interpolation for EEG-audio alignment, introducing redundant computations. They also employ single-path encoders that extract only target-relevant features while neglecting complementary, irrelevant ones, limiting discriminability. To address these limitations, this paper proposes a Trainable EEG Interpolation and Structure-sharing Dual-path Encoders network (TIDENet). The proposed Trainable EEG Interpolation (TEI) uses a neural network module to leverage cross-sample EEG information during resampling by parameters updating, thereby overcoming the limitations of fixed interpolation. The Structure-sharing Dual-path Encoders (SSDPE) extend existing speech and EEG encoders by introducing dual paths that separately process features relevant and irrelevant to the target speaker and incorporates interactive fusion between them, which enhances the encoder's ability to capture task-relevant information. Experimental results on public datasets demonstrate that TIDENet achieves relative improvements of up to 20.47%, 22.22%, 2.91%, 6.20%, and 15.84% in signal-to-distortion ratio (SDR), scale-invariant SDR (SI-SDR), short-time objective intelligibility (STOI), extended STOI (ESTOI), and perceptual evaluation of speech quality (PESQ), respectively, compared to the state-of-the-art. These significant gains validate the effectiveness of the proposed TEI method and SSDPE architecture.

EAAI Journal 2026 Journal Article

Vibration-induced deformation prediction and multi-staged parameter optimization of coarse-grained soils: Trade-off between energy and time

  • Ying Chen
  • Qun Qi
  • Zhihong Nie

The relationship between compaction deformation and vibration parameters in coarse-grained soils, a typical geotechnical material, is highly nonlinear, which limits the prediction of vibration-induced deformation and thus restricts the optimization of vibration parameters. Since the compaction deformation is dynamic and corresponds to different optimal vibration parameters, the parameter optimization should be multi-staged. Vibration parameters determine the energy and time costs, so multi-objective optimization, trading off vibration time and vibration energy, is the key to determining the optimal vibration parameters. Therefore, an artificial intelligence (AI)-based method was proposed for predicting compaction deformation and optimizing vibration parameters, where the former was predicted by the BO-FCNN (Bayesian optimized fully connected neural network) algorithm, and the latter was optimized by the proposed MC-NSGA-II (multi-stage and multi-objective) algorithm through adjusting the weights of time and energy. Through vibration compaction tests, the effectiveness of the method was proven. The BO-FCNN algorithm exhibits high accuracy in predicting compaction deformation, and the MC-NSGA-II algorithm shows a significant Pareto trade-off relationship between vibration energy and vibration time. In time-priority optimization, vibration time was reduced by 16. 9 % using high excitation force and frequency, while the energy was higher. In energy-priority optimization, vibration energy was reduced by 14. 2 % through alternating between low and high excitation forces, while the time was higher. The optimized scheme outperforms the traditional one, inducing time and energy optimization by 46. 2 % and 53. 2 % respectively. This study provides AI-based insights for improving the compaction efficiency and reducing the energy consumption in subgrade construction engineering.

EAAI Journal 2025 Journal Article

A novel convolutional neural network with global perception for bearing fault diagnosis

  • Xianguo Li
  • Ying Chen
  • Yi Liu

Bearings are key support components in rotating machinery, and their stability is crucial to the reliability of the entire mechanical system. To address the limitations of existing Transformer architectures in edge-side optimization and convolutional neural networks in global feature extraction, especially the resulting poor real-time performance and low accuracy in bearing fault diagnosis based on acoustic signals, this paper proposes a novel global-aware convolutional neural network based on residual masking and position-aware strategies (ParC-ReSMNet). Firstly, the network is based on the residual network (ResNet-18) and designed with a residual mask block combined with an improved spatial pyramid mask attention mechanism, which effectively removes redundant spatial information and focuses on critical fault features, thereby enhancing the robustness and generalization of the network. Secondly, a position-aware circular module is introduced to replace specific residual blocks in the original network, achieving an effective fusion of positional and global information, thereby augmenting the modeling capability of the convolutional neural network. Experiments are conducted on a self-made belt conveyor idler dataset and the Detection and Classification of Acoustic Scenes and Events (DCASE) 2023 Task2 bearing dataset, with results showing that ParC-ReSMNet achieves 95. 49% and 96. 67% accuracy, respectively. Compared to seven state-of-the-art models, it has the highest precision and recall, along with good real-time performance, which demonstrates great application value for fault monitoring of belt conveyors used in coal mines, power plants, ports, and other rotating machinery. The code library is available at: https: //github. com/xgli411/Parc-ResMNet.

JBHI Journal 2025 Journal Article

BioMTAN: A Biological Knowledge-guided Multi-task Attention Network for Co-enhanced Cancer Diagnosis and Prognosis

  • Ying Chen
  • Jiajing Xie
  • Yuxiang Lin
  • Yuhang Song
  • Wenxian Yang
  • Rongshan Yu

With the advancement of precision medicine, gene expression data have become a crucial tool in both cancer diagnosis and prognosis for different cancer types. The incorporation of biological pathways as prior knowledge has gained increasing interest in tackling the difficulties of high dimensionality and noisy information within gene expression data. However, most existing approaches guided by biological pathways ignore the intrinsic link between diagnostic and prognostic tasks in cancer research. They fail to capitalize on the potential of leveraging shared biological information from both tasks to enhance gene pathway representations. To this end, we introduce the Biological Knowledge-guided Multi-task Attention Network (BioMTAN), a novel multi-task learning framework designed for simultaneous prediction of molecular subtypes and survival risk. Specifically, we compile tailored knowledge collections that comprise multiple pathways for the two tasks, model them as unique subgraphs and use a multi-level information fusion strategy to provide a wealth of biological insights. Moreover, we develop a Multi-task Attention Module, which extracts essential global information functioning as the key and value by interacting with biological pathways from different collections, and utilizes task-specific local information as the query, efficiently decoding task-awareness feature for each task and facilitating communication across tasks within cancer diagnosis and prognosis. Extensive validation on the public The Cancer Genome Atlas (TCGA) datasets confirms the enhanced performance of BioMTAN and highlights the significant pathways in each task, underscoring its potential as an instrumental asset in precision oncology.

YNIMG Journal 2025 Journal Article

Disrupted structural connectivity-gray matter covariance coupling and associated cytoarchitectural and transcriptomic profiles in attention-deficit/hyperactivity disorder

  • Yajing Long
  • Nanfang Pan
  • Song Wang
  • Kun Qin
  • Qiuxing Chen
  • Clara S. Vetter
  • Manpreet K. Singh
  • Alex Fornito

BACKGROUND: Attention-deficit/hyperactivity disorder (ADHD) has been associated with disrupted axonal connectivity (termed structural connectivity, SC) and altered interregional coupling of gray matter morphometry (termed gray matter covariance, GMC). However, the relationship between SC and GMC in ADHD remains understudied. METHODS: We investigated this relationship by quantifying the coupling between SC and GMC using neuroimaging data from 109 children with ADHD (aged 10.8 ± 2.3) and 105 typically developing controls (aged 11.2 ± 2.4) comparable for age and sex. Publicly accessible cytoarchitectural and transcriptomic datasets were employed to characterize the cellular and molecular correlates of ADHD-related SC-GMC coupling differences, and a machine learning pipeline was used to investigate its potential in classifying children with ADHD. RESULTS: Children with ADHD showed aberrant SC-GMC coupling patterns in the right putamen, left hippocampus, and ventral attention network compared to controls. Their abnormal SC-GMC coupling patterns were correlated with sensory-fugal gradient of cytohistological variation and spatially associated with gene expression enriched for neurodevelopment-related biological pathways, including neuron projection development. The classification model based on SC-GMC couplings achieved an area under the receiver operating characteristic curve (AUC) value of 0.67. CONCLUSIONS: Our findings provide novel insights into atypical couplings between brain gray and white matter structural connectomes in ADHD, their histological and transcriptional correlates, and prospects of using these data to expand clinical phenotyping.

EAAI Journal 2025 Journal Article

Fully logits guided distillation with intermediate decision learning for deep model compression

  • Yiqin Wang
  • Ying Chen

Knowledge distillation, as a model compression technique, has been widely applied in artificial intelligence to improve the efficiency of deep learning models, especially in resource-constrained environments. Considering that logit contains more decision information compared to the intermediate feature maps, fully logits guided distillation is proposed, which allows student networks to have better access to the guidance from both the intermediate and decision levels of the teacher’s network. Intermediate feature logicalization is designed, which perform logit transformations on the intermediate feature maps to obtain intermediate decision information. A logit matrisation strategy is proposed, which aim to capture inter-class information of the logits. Furthermore, cross layer distillation is presented in order to enable the final logit of the teacher to provide guidance to the intermediate layers of the student. The proposed mechanism can be embedded into State-of-the-art distillation frameworks to further improve the accuracy. Experiments conducted on the CIFAR-10, CIFAR-100, and ImageNet datasets demonstrate the effectiveness of the proposed method. Image classification accuracy was used as the evaluation metric, and the results show that the proposed method improves accuracy by an average of 1. 84%, with a best improvement of 3. 19% over the baseline models. Code is available at https: //github. com/YiqinWang-JN/FLGD.

AAAI Conference 2025 Conference Paper

Imitate Before Detect: Aligning Machine Stylistic Preference for Machine-Revised Text Detection

  • Jiaqi Chen
  • Xiaoye Zhu
  • Tianyang Liu
  • Ying Chen
  • Chen Xinhui
  • Yiwen Yuan
  • Chak Tou Leong
  • Zuchao Li

Large Language Models (LLMs) have revolutionized text generation, making detecting machine-generated text increasingly challenging. Although past methods have achieved good performance on detecting pure machine-generated text, those detectors have poor performance on distinguishing machine-revised text (rewriting, expansion, and polishing), which can have only minor changes from its original human prompt. As the content of text may originate from human prompts, detecting machine-revised text often involves identifying distinctive machine styles, e.g., worded favored by LLMs. However, existing methods struggle to detect machine-style phrasing hidden within the content contributed by humans. We propose the “Imitate Before Detect” (ImBD) approach, which first imitates the machine-style token distribution, and then compares the distribution of the text to be tested with the machine-style distribution to determine whether the text has been machine-revised. To this end, we introduce Style Preference Optimization (SPO), which aligns a scoring LLM model to the preference of text styles generated by machines. The aligned scoring model is then used to calculate the style-conditional probability curvature (Style-CPC), quantifying the log probability difference between the original and conditionally sampled texts for effective detection. We conduct extensive comparisons across various scenarios, encompassing text revisions by six LLMs, four distinct text domains, and three machine revision types. Compared to existing state-of-the-art methods, our method yields a 13% increase in AUC for detecting text revised by open-source LLMs, and improves performance by 5% and 19% for detecting GPT-3.5 and GPT-4o revised text, respectively. Notably, our method surpasses the commercially trained GPT-Zero with just 1,000 samples and five minutes of SPO, demonstrating its efficiency and effectiveness.

IJCAI Conference 2025 Conference Paper

ListenNet: A Lightweight Spatio-Temporal Enhancement Nested Network for Auditory Attention Detection

  • Cunhang Fan
  • Xiaoke Yang
  • Hongyu Zhang
  • Ying Chen
  • Lu Li
  • Jian Zhou
  • Zhao Lv

Auditory attention detection (AAD) aims to identify the direction of the attended speaker in multi-speaker environments from brain signals, such as Electroencephalography (EEG) signals. However, existing EEG-based AAD methods overlook the spatio-temporal dependencies of EEG signals, limiting their decoding and generalization abilities. To address these issues, this paper proposes a Lightweight Spatio-Temporal Enhancement Nested Network (ListenNet) for AAD. The ListenNet has three key components: Spatio-temporal Dependency Encoder (STDE), Multi-scale Temporal Enhancement (MSTE), and Cross-Nested Attention (CNA). The STDE reconstructs dependencies between consecutive time windows across channels, improving the robustness of dynamic pattern extraction. The MSTE captures temporal features at multiple scales to represent both fine-grained and long-range temporal patterns. In addition, the CNA integrates hierarchical features more effectively through novel dynamic attention mechanisms to capture deep spatio-temporal correlations. Experimental results on three public datasets demonstrate the superiority of ListenNet over state-of-the-art methods in both subject-dependent and challenging subject-independent settings, while reducing the trainable parameter count by approximately 7 times. Code is available at: https: //github. com/fchest/ListenNet.

IJCAI Conference 2025 Conference Paper

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction

  • Cunhang Fan
  • Ying Chen
  • Jian Zhou
  • Zexu Pan
  • Jingjing Zhang
  • Youdian Gao
  • Xiaoke Yang
  • Zhengqi Wen

The brain-assisted target speaker extraction (TSE) aims to extract the attended speech from mixed speech by utilizing the brain neural activities, for example Electroencephalography (EEG). However, existing models overlook the issue of temporal misalignment between speech and EEG modalities, which hampers TSE performance. In addition, the speech encoder in current models typically uses basic temporal operations (e. g. , one-dimensional convolution), which are unable to effectively extract target speaker information. To address these issues, this paper proposes a multi-scale and multi-modal alignment network (M3ANet) for brain-assisted TSE. Specifically, to eliminate the temporal inconsistency between EEG and speech modalities, the modal alignment module that uses a contrastive learning strategy is applied to align the temporal features of both modalities. Additionally, to fully extract speech information, multi-scale convolutions with GroupMamba modules are used as the speech encoder, which scans speech features at each scale from different directions, enabling the model to capture deep sequence information. Experimental results on three publicly available datasets show that the proposed model outperforms current state-of-the-art methods across various evaluation metrics, highlighting the effectiveness of our proposed method. The source code is available at: https: //github. com/fchest/M3ANet.

JBHI Journal 2025 Journal Article

SleepHybridNet: A Lightweight Hybrid CNN-Transformer Model for Enhanced N1 Sleep Staging From Single-Channel EEG

  • Hao Zhou
  • Mengxiang Su
  • Jeng-Shyang Pan
  • Chenglong Dai
  • Ying Chen
  • Shu-Chuan Chu

This study introduces SleepHybridNet, a lightweight hybrid CNN-Transformer model designed to enhance the classification of non-rapid eye movement stage 1 (N1) sleep using single-channel electroencephalogram (EEG) signals. Accurate identification of the N1 stage is of critical importance in both sleep neuroscience and clinical practice. However, due to the ambiguous features during N1 stage, current deep learning models still struggle to achieve satisfactory performance. To address these challenges, SleepHybridNet integrates multi-scale feature fusion and sequence modeling through a novel architecture. It consists of a Multi-Scale Convolutional Neural Network (MSCNN) module, a Transformer encoder, a spectral feature extraction unit, and a multi-task classifier. Experimental results based on the publicly available Sleep-EDF Expanded dataset demonstrate that SleepHybridNet outperforms existing methods in both classification accuracy and generalization capability. Specifically, the model achieves an overall accuracy of 88. 2% and an F1-score of 0. 633 for the N1 stage, showing superior performance particularly in underrepresented classes such as N1 and N3 stages. With only 5. 1 M parameters, the lightweight design of the model can enable practical deployment in clinical settings, bridging the gap between high-performance deep learning algorithms and practical applicability in sleep medicine. Future work may explore the integration of multimodal data from wearable sensors to further expand its use in diverse application scenarios.

ICRA Conference 2024 Conference Paper

3D Object Detection with VI-SLAM Point Clouds: The Impact of Object and Environment Characteristics on Model Performance

  • Lin Duan
  • Tim Scargill
  • Ying Chen
  • Maria Gorlatova

3D object detection (OD) is a crucial element in scene understanding. However, most existing 3D OD models have been tailored to work with light detection and ranging (LiDAR) and RGB-D point cloud data, leaving their performance on commonly available visual-inertial simultaneous localization and mapping (VI-SLAM) point clouds unexamined. In this paper, we create and release two datasets: VIP500, 4772 VI-SLAM point clouds covering 500 different object and environment configurations, and VIP500-D, an accompanying set of 20 RGB-D point clouds for the object classes and shapes in VIP500. We then use these datasets to quantify the differences between VI-SLAM point clouds and dense RGB-D point clouds, as well as the discrepancies between VI-SLAM point clouds generated with different object and environment characteristics. Finally, we evaluate the performance of three leading OD models on the diverse data in our VIP500 dataset, revealing the promise of OD models trained on VI-SLAM data; we examine the extent to which both object and environment characteristics impact performance, along with the underlying causes.

AAAI Conference 2024 Conference Paper

Unsupervised Continual Anomaly Detection with Contrastively-Learned Prompt

  • Jiaqi Liu
  • Kai Wu
  • Qiang Nie
  • Ying Chen
  • Bin-Bin Gao
  • Yong Liu
  • Jinbao Wang
  • Chengjie Wang

Unsupervised Anomaly Detection (UAD) with incremental training is crucial in industrial manufacturing, as unpredictable defects make obtaining sufficient labeled data infeasible. However, continual learning methods primarily rely on supervised annotations, while the application in UAD is limited due to the absence of supervision. Current UAD methods train separate models for different classes sequentially, leading to catastrophic forgetting and a heavy computational burden. To address this issue, we introduce a novel Unsupervised Continual Anomaly Detection framework called UCAD, which equips the UAD with continual learning capability through contrastively-learned prompts. In the proposed UCAD, we design a Continual Prompting Module (CPM) by utilizing a concise key-prompt-knowledge memory bank to guide task-invariant 'anomaly' model predictions using task-specific 'normal' knowledge. Moreover, Structure-based Contrastive Learning (SCL) is designed with the Segment Anything Model (SAM) to improve prompt learning and anomaly segmentation results. Specifically, by treating SAM's masks as structure, we draw features within the same mask closer and push others apart for general feature representations. We conduct comprehensive experiments and set the benchmark on unsupervised continual anomaly detection and segmentation, demonstrating that our method is significantly better than anomaly detection methods, even with rehearsal training. The code will be available at https://github.com/shirowalker/UCAD.

TIST Journal 2023 Journal Article

Attention-guided Adversarial Attack for Video Object Segmentation

  • Rui Yao
  • Ying Chen
  • Yong Zhou
  • Fuyuan Hu
  • Jiaqi Zhao
  • Bing Liu
  • Zhiwen Shao

Video Object Segmentation (VOS) methods have made many breakthroughs with the help of the continuous development and advancement of deep learning. However, the deep learning model is vulnerable to malicious adversarial attacks, which mislead the model to make wrong decisions by adding adversarial perturbation that humans cannot perceive to the input image. Threats to deep learning models remind us that video object segmentation methods are also vulnerable to attacks, thereby threatening their security. Therefore, we study adversarial attacks on the VOS task to better identify the vulnerabilities of the VOS method, which in turn provides an opportunity to improve its robustness. In this paper, we propose an attention-guided adversarial attack method, which uses spatial attention blocks to capture features with global dependencies to construct correlations between consecutive video frames, and performs multipath aggregation to effectively integrate spatial-temporal perturbation, thereby guiding the deconvolution network to generate adversarial examples with strong attack capability. Specifically, the class loss function is designed to enable the deconvolution network to better activate noise in other regions and suppress the activation related to the object class based on the enhanced feature map of the object class. At the same time, attentional feature loss is designed to enhance the transferability against attack. The experimental results on the DAVIS dataset show that the proposed attention-guided adversarial attack method can significantly reduce the segmentation accuracy of OSVOS, and the J & F mean on DAVIS 2016 can reach 73.6% drop rate. The generated adversarial examples are also highly transferable to other video object segmentation models.

AAAI Conference 2023 Conference Paper

Copyright-Certified Distillation Dataset: Distilling One Million Coins into One Bitcoin with Your Private Key

  • Tengjun Liu
  • Ying Chen
  • Wanxuan Gu

The rapid development of neural network dataset distillation in recent years has provided new ideas in many areas such as continuous learning, neural network architecture search and privacy preservation. Dataset distillation is a very effective method to distill large training datasets into small data, thus ensuring that the test accuracy of models trained on their synthesized small datasets matches that of models trained on the full dataset. Thus, dataset distillation itself is commercially valuable, not only for reducing training costs, but also for compressing storage costs and significantly reducing the training costs of deep learning. However, copyright protection for dataset distillation has not been proposed yet, so we propose the first method to protect intellectual property by embedding watermarks in the dataset distillation process. Our approach not only popularizes the dataset distillation technique, but also authenticates the ownership of the distilled dataset by the models trained on that distilled dataset.

NeurIPS Conference 2023 Conference Paper

Geometric Analysis of Matrix Sensing over Graphs

  • Haixiang Zhang
  • Ying Chen
  • Javad Lavaei

In this work, we consider the problem of matrix sensing over graphs (MSoG). As a general case of matrix completion and matrix sensing problems, the MSoG problem has not been analyzed in the literature and the existing results cannot be directly applied to the MSoG problem. This work provides the first theoretical results on the optimization landscape of the MSoG problem. More specifically, we propose a new condition, named the $\Omega$-RIP condition, to characterize the optimization complexity of the problem. In addition, with an improved regularizer of the incoherence, we prove that the strict saddle property holds for the MSoG problem with high probability under the incoherence condition and the $\Omega$-RIP condition, which guarantees the polynomial-time global convergence of saddle-avoiding methods. Compared with state-of-the-art results, the bounds in this work are tight up to a constant. Besides the theoretical guarantees, we numerically illustrate the close relation between the $\Omega$-RIP condition and the optimization complexity.

ICLR Conference 2023 Conference Paper

TextShield: Beyond Successfully Detecting Adversarial Sentences in text classification

  • Lingfeng Shen
  • Ze Zhang
  • Haiyun Jiang
  • Ying Chen

Adversarial attack serves as a major challenge for neural network models in NLP, which precludes the model's deployment in safety-critical applications. A recent line of work, detection-based defense, aims to distinguish adversarial sentences from benign ones. However, {the core limitation of previous detection methods is being incapable of giving correct predictions on adversarial sentences unlike defense methods from other paradigms.} To solve this issue, this paper proposes TextShield: (1) we discover a link between text attack and saliency information, and then we propose a saliency-based detector, which can effectively detect whether an input sentence is adversarial or not. (2) We design a saliency-based corrector, which converts the detected adversary sentences to benign ones. By combining the saliency-based detector and corrector, TextShield extends the detection-only paradigm to a detection-correction paradigm, thus filling the gap in the existing detection-based defense. Comprehensive experiments show that (a) TextShield consistently achieves higher or comparable performance than state-of-the-art defense methods across various attacks on different benchmarks. (b) our saliency-based detector outperforms existing detectors for detecting adversarial sentences.

YNICL Journal 2022 Journal Article

Changes in brain connectivity linked to multisensory processing of pain modulation in migraine with acupuncture treatment

  • Lu Liu
  • Tian-Li Lyu
  • Ming-Yang Fu
  • Lin-Peng Wang
  • Ying Chen
  • Jia-Hui Hong
  • Qiu-Yi Chen
  • Yu-Pu Zhu

Migraine without aura (MWoA) is a major neurological disorder with unsatisfactory adherence to current medications. Acupuncture has emerged as a promising method for treating MWoA. However, the brain mechanism underlying acupuncture is yet unclear. The present study aimed to examine the effects of acupuncture in regulating brain connectivity of the key regions in pain modulation. In this study, MWoA patients were recruited and randomly assigned to 4 weeks of real or sham acupuncture. Resting-state functional magnetic resonance imaging (fMRI) data were collected before and after the treatment. A modern neuroimaging literature meta-analysis of 515 fMRI studies was conducted to identify pain modulation-related key regions as regions of interest (ROIs). Seed-to-voxel resting state-functional connectivity (rsFC) method and repeated-measures two-way analysis of variance were conducted to determine the interaction effects between the two groups and time (baseline and post-treatment). The changes in rsFC were evaluated between baseline and post-treatment in real and sham acupuncture groups, respectively. Clinical data at baseline and post-treatment were also recorded in order to determine between-group differences in clinical outcomes as well as correlations between rsFC changes and clinical effects. 40 subjects were involved in the final analysis. The current study demonstrated significant improvement in real acupuncture vs sham acupuncture on headache severity (monthly migraine days), headache impact (6-item Headache Impact Test), and health-related quality of life (Migraine-Specific Quality of Life Questionnaire). Five pain modulation-related key regions, including the right amygdala (AMYG), left insula (INS), left medial orbital superior frontal gyrus (PFCventmed), left middle occipital gyrus (MOG), and right middle cingulate cortex (MCC), were selected based on the meta-analysis on brain imaging studies. This study found that 1) after acupuncture treatment, migraine patients of the real acupuncture group showed significantly enhanced connectivity in the right AMYG/MCC-left MTG and the right MCC-right superior temporal gyrus (STG) compared to that of the sham acupuncture group; 2) negative correlations were established between clinical effects and increased rsFC in the right AMYG/MCC-left MTG; 3) baseline right AMYG-left MTG rsFC predicts monthly migraine days reduction after treatment. The current results suggested that acupuncture may concurrently regulate the rsFC of two pain modulation regions in the AMYG and MCC. MTG and STG may be the key nodes linked to multisensory processing of pain modulation in migraine with acupuncture treatment. These findings highlighted the potential of acupuncture for migraine management and the mechanisms underlying the modulation effects.

AAAI Conference 2022 Conference Paper

Guide Local Feature Matching by Overlap Estimation

  • Ying Chen
  • Dihe Huang
  • Shang Xu
  • Jianlin Liu
  • Yong Liu

Local image feature matching under large appearance, viewpoint, and distance changes is challenging yet important. Conventional methods detect and match tentative local features across the whole images, with heuristic consistency checks to guarantee reliable matches. In this paper, we introduce a novel Overlap Estimation method conditioned on image pairs with TRansformer, named OETR, to constrain local feature matching in the commonly visible region. OETR performs overlap estimation in a two step process of feature correlation and then overlap regression. As a preprocessing module, OETR can be plugged into any existing local feature detection and matching pipeline, to mitigate potential view angle or scale variance. Intensive experiments show that OETR can boost state of the art local feature matching performance substantially, especially for image pairs with small shared regions. The code will be publicly available at https: //github. com/AbyssGaze/OETR.

AAAI Conference 2022 Conference Paper

KATG: Keyword-Bias-Aware Adversarial Text Generation for Text Classification

  • Lingfeng Shen
  • Shoushan Li
  • Ying Chen

Recent work has shown that current text classification models are vulnerable to a small adversarial perturbation on inputs, and adversarial training that re-trains the models with the support of adversarial examples is the most popular way to alleviate the impact of the perturbation. However, current adversarial training methods have two principal problems: a drop in model’s generalization and ineffective defending against other text attacks. In this paper, we propose a Keywordbias-aware Adversarial Text Generation model (KATG) that implicitly generates adversarial sentences using a generatordiscriminator structure. Instead of using a benign sentence to generate an adversarial sentence, the KATG model utilizes extra multiple benign sentences (namely prior sentences) to guide adversarial sentence generation. Furthermore, to cover more perturbations used in existing attacks, a keyword-biasbased sampling is proposed to select sentences containing biased words as prior sentences. Besides, to effectively utilize prior sentences, a generative flow mechanism is proposed to construct a latent semantic space for learning a latent representation of the prior sentences. Experiments demonstrate that adversarial sentences generated by our KATG model can strengthen the generalization and the robustness of text classification models. Benign Sentence Sixthreezero is good, I’ve used it for a long time, only changed because I got tired of the same old bike. (Pos) Prior Sentences S1: Blackberry may work on the systems, but I’m not willing to take that chance on a new expensive phone. (Neg) S2: Iphone4s is in ok previously used condition as stated. But I was disappointed I couldn’t activate the phone upon arrival. (Neg) Adv. Sentence Amazing Iphone4s, used it for so long, only changed because I got tired of the old expensive Blackberry. (Pos) Table 1: Benign sentence, prior sentences and adversarial sentence used in our KATG model. *the corresponding author Copyright © 2022, Association for the Advancement of Artificial Intelligence (www. aaai. org). All rights reserved.

IJCAI Conference 2021 Conference Paper

KDExplainer: A Task-oriented Attention Model for Explaining Knowledge Distillation

  • Mengqi Xue
  • Jie Song
  • Xinchao Wang
  • Ying Chen
  • Xingen Wang
  • Mingli Song

Knowledge distillation (KD) has recently emerged as an efficacious scheme for learning compact deep neural networks (DNNs). Despite the promising results achieved, the rationale that interprets the behavior of KD has yet remained largely understudied. In this paper, we introduce a novel task-oriented attention model, termed as KDExplainer, to shed light on the working mechanism underlying the vanilla KD. At the heart of KDExplainer is a Hierarchical Mixture of Experts (HME), in which a multi-class classification is reformulated as a multi-task binary one. Through distilling knowledge from a free-form pre-trained DNN to KDExplainer, we observe that KD implicitly modulates the knowledge conflicts between different subtasks, and in reality has much more to offer than label smoothing. Based on such findings, we further introduce a portable tool, dubbed as virtual attention module (VAM), that can be seamlessly integrated with various DNNs to enhance their performance under KD. Experimental results demonstrate that with a negligible additional cost, student models equipped with VAM consistently outperform their non-VAM counterparts across different benchmarks. Furthermore, when combined with other KD methods, VAM remains competent in promoting results, even though it is only motivated by vanilla KD. The code is available at https: // github. com/zju-vipa/KDExplainer.

AAAI Conference 2021 Conference Paper

MANGO: A Mask Attention Guided One-Stage Scene Text Spotter

  • Liang Qiao
  • Ying Chen
  • Zhanzhan Cheng
  • Yunlu Xu
  • Yi Niu
  • Shiliang Pu
  • Fei Wu

Recently end-to-end scene text spotting has become a popular research topic due to its advantages of global optimization and high maintainability in real applications. Most methods attempt to develop various region of interest (RoI) operations to concatenate the detection part and the sequence recognition part into a two-stage text spotting framework. However, in such framework, the recognition part is highly sensitive to the detected results (e. g. , the compactness of text contours). To address this problem, in this paper, we propose a novel Mask AttentioN Guided One-stage text spotting framework named MANGO, in which character sequences can be directly recognized without RoI operation. Concretely, a positionaware mask attention module is developed to generate attention weights on each text instance and its characters. It allows different text instances in an image to be allocated on different feature map channels which are further grouped as a batch of instance features. Finally, a lightweight sequence decoder is applied to generate the character sequences. It is worth noting that MANGO inherently adapts to arbitraryshaped text spotting and can be trained end-to-end with only coarse position information (e. g. , rectangular bounding box) and text annotations. Experimental results show that the proposed method achieves competitive and even new state-ofthe-art performance on both regular and irregular text spotting benchmarks, i. e. , ICDAR 2013, ICDAR 2015, Total-Text, and SCUT-CTW1500.

NeurIPS Conference 2021 Conference Paper

Refining Language Models with Compositional Explanations

  • Huihan Yao
  • Ying Chen
  • Qinyuan Ye
  • Xisen Jin
  • Xiang Ren

Pre-trained language models have been successful on text classification tasks, but are prone to learning spurious correlations from biased datasets, and are thus vulnerable when making inferences in a new domain. Prior work reveals such spurious patterns via post-hoc explanation algorithms which compute the importance of input features. Further, the model is regularized to align the importance scores with human knowledge, so that the unintended model behaviors are eliminated. However, such a regularization technique lacks flexibility and coverage, since only importance scores towards a pre-defined list of features are adjusted, while more complex human knowledge such as feature interaction and pattern generalization can hardly be incorporated. In this work, we propose to refine a learned language model for a target domain by collecting human-provided compositional explanations regarding observed biases. By parsing these explanations into executable logic rules, the human-specified refinement advice from a small set of explanations can be generalized to more training examples. We additionally introduce a regularization term allowing adjustments for both importance and interaction of features to better rectify model behavior. We demonstrate the effectiveness of the proposed approach on two text classification tasks by showing improved performance in target domain as well as improved model fairness after refinement.

AAAI Conference 2021 Conference Paper

Temporal-Coded Deep Spiking Neural Network with Easy Training and Robust Performance

  • Shibo Zhou
  • Xiaohua Li
  • Ying Chen
  • Sanjeev T. Chandrasekaran
  • Arindam Sanyal

Spiking neural network (SNN) is promising but the development has fallen far behind conventional deep neural networks (DNNs) because of difficult training. To resolve the training problem, we analyze the closed-form input-output response of spiking neurons and use the response expression to build abstract SNN models for training. This avoids calculating membrane potential during training and makes the direct training of SNN as efficient as DNN. We show that the nonleaky integrate-and-fire neuron with single-spike temporalcoding is the best choice for direct-train deep SNNs. We develop an energy-efficient phase-domain signal processing circuit for the neuron and propose a direct-train deep SNN framework. Thanks to easy training, we train deep SNNs under weight quantizations to study their robustness over low-cost neuromorphic hardware. Experiments show that our direct-train deep SNNs have the highest CIFAR-10 classification accuracy among SNNs, achieve ImageNet classification accuracy within 1% of the DNN of equivalent architecture, and are robust to weight quantization and noise perturbation.

AIIM Journal 2020 Journal Article

ADHD classification by dual subspace learning using resting-state functional connectivity

  • Ying Chen
  • Yibin Tang
  • Chun Wang
  • Xiaofeng Liu
  • Li Zhao
  • Zhishun Wang

As one of the most common neurobehavioral diseases in school-age children, Attention Deficit Hyperactivity Disorder (ADHD) has been increasingly studied in recent years. But it is still a challenge problem to accurately identify ADHD patients from healthy persons. To address this issue, we propose a dual subspace classification algorithm by using individual resting-state Functional Connectivity (FC). In detail, two subspaces respectively containing ADHD and healthy control features, called as dual subspaces, are learned with several subspace measures, wherein a modified graph embedding measure is employed to enhance the intra-class relationship of these features. Therefore, given a subject (used as test data) with its FCs, the basic classification principle is to compare its projected component energy of FCs on each subspace and then predict the ADHD or control label according to the subspace with larger energy. However, this principle in practice works with low efficiency, since the dual subspaces are unstably obtained from ADHD databases of small size. Thereby, we present an ADHD classification framework by a binary hypothesis testing of test data. Here, the FCs of test data with its ADHD or control label hypothesis are employed in the discriminative FC selection of training data to promote the stability of dual subspaces. For each hypothesis, the dual subspaces are learned from the selected FCs of training data. The total projected energy of these FCs is also calculated on the subspaces. Sequentially, the energy comparison is carried out under the binary hypotheses. The ADHD or control label is finally predicted for test data with the hypothesis of larger total energy. In the experiments on ADHD-200 dataset, our method achieves a significant classification performance compared with several state-of-the-art machine learning and deep learning methods, where our accuracy is about 90 % for most of ADHD databases in the leave-one-out cross-validation test.

JAIR Journal 2019 Journal Article

DSTL: Solution to Limitation of Small Corpus in Speech Emotion Recognition

  • Ying Chen
  • Zhongzhe Xiao
  • Xiaojun Zhang
  • Zhi Tao

Traditional machine learning methods share a common hypothesis: training and testing datasets must be in a common feature space with the same distribution. However, in reality, the labeled target data may be rare, so that target space does not share the same feature space or distribution as an available training set (source domain). To address the mismatch of domains, we propose a Dual-Subspace Transfer Learning (DSTL) framework that considers both the common and specific information of the two domains. In DSTL, a latent common subspace is first learned to preserve the data properties and reduce the discrepancy of domains. Then, we propose a mapping strategy to transfer the sourcespecific information to the target subspace. The integration of the domain-common and specific information constructs the proposed DSTL framework. In comparison to the stateart-of works, the main contribution of our work is that the DSTL framework not only considers the commonalities, but also exploits the specific information. Experiments on three emotional speech corpora verify the effectiveness of our approach. The results show that the methods which include both domain-common and specific information perform better than the baseline methods which only exploit the domain commonalities.

AAAI Conference 2019 Conference Paper

M2Det: A Single-Shot Object Detector Based on Multi-Level Feature Pyramid Network

  • Qijie Zhao
  • Tao Sheng
  • Yongtao Wang
  • Zhi Tang
  • Ying Chen
  • Ling Cai
  • Haibin Ling

Feature pyramids are widely exploited by both the state-ofthe-art one-stage object detectors (e. g. , DSSD, RetinaNet, RefineDet) and the two-stage object detectors (e. g. , Mask R- CNN, DetNet) to alleviate the problem arising from scale variation across object instances. Although these object detectors with feature pyramids achieve encouraging results, they have some limitations due to that they only simply construct the feature pyramid according to the inherent multiscale, pyramidal architecture of the backbones which are originally designed for object classification task. Newly, in this work, we present Multi-Level Feature Pyramid Network (MLFPN) to construct more effective feature pyramids for detecting objects of different scales. First, we fuse multi-level features (i. e. multiple layers) extracted by backbone as the base feature. Second, we feed the base feature into a block of alternating joint Thinned U-shape Modules and Feature Fusion Modules and exploit the decoder layers of each Ushape module as the features for detecting objects. Finally, we gather up the decoder layers with equivalent scales (sizes) to construct a feature pyramid for object detection, in which every feature map consists of the layers (features) from multiple levels. To evaluate the effectiveness of the proposed MLFPN, we design and train a powerful end-to-end one-stage object detector we call M2Det by integrating it into the architecture of SSD, and achieve better detection performance than state-of-the-art one-stage detectors. Specifically, on MS- COCO benchmark, M2Det achieves AP of 41. 0 at speed of 11. 8 FPS with single-scale inference strategy and AP of 44. 2 with multi-scale inference strategy, which are the new stateof-the-art results among one-stage detectors. The code will be made available on https: //github. com/qijiezhao/M2Det.

YNICL Journal 2018 Journal Article

Disrupted grey matter network morphology in pediatric posttraumatic stress disorder

  • Running Niu
  • Du Lei
  • Fuqin Chen
  • Ying Chen
  • Xueling Suo
  • Lingjiang Li
  • Su Lui
  • Xiaoqi Huang

Introduction: Disrupted topological organization of brain functional networks has been widely observed in posttraumatic stress disorder (PTSD). However, the topological organization of the brain grey matter (GM) network has not yet been investigated in pediatric PTSD who was more vulnerable to develop PTSD when exposed to stress. Materials and methods: Twenty two pediatric PTSD patients and 22 matched trauma-exposed controls who survived a massive earthquake (8.0 magnitude on Richter scale) in Sichuan Province of western China in 2008 underwent structural brain imaging with MRI 8-15 months after the earthquake. Brain networks were constructed based on the morphological similarity of GM across regions, and analyzed using graph theory approaches. Nonparametric permutation testing was performed to assess group differences in each topological metric. Results: Compared with controls, brain networks of PTSD patients were characterized by decreased characteristic path length (P = 0.0060) and increased clustering coefficient (P = 0.0227), global efficiency (P = 0.0085) and local efficiency (P = 0.0024). Locally, patients with PTSD exhibited increased centrality in nodes of the default-mode (DMN), central executive (CEN) and salience networks (SN), involving medial prefrontal (mPFC), parietal, anterior cingulate (ACC), occipital and olfactory cortex and hippocampus. Conclusions: Our analyses of topological brain networks in children with PTSD indicate a significantly more segregated and integrated organization. The associations and disassociations between these grey matter findings and white matter (WM) and functional changes previously reported in this sample may be important for diagnostic purposes and understanding the brain maturational effects of pediatric PTSD.

YNICL Journal 2018 Journal Article

Volume alteration of hippocampal subfields in first-episode antipsychotic-naïve schizophrenia patients before and after acute antipsychotic treatment

  • Wenbin Li
  • Kaiming Li
  • Pujun Guan
  • Ying Chen
  • Yuan Xiao
  • Su Lui
  • John A. Sweeney
  • Qiyong Gong

The nature of hippocampal changes in schizophrenia before first treatment, and whether hippocampal subfields are affected by antipsychotic treatment are important questions for schizophrenia research. Forty-one first-episode antipsychotic-naïve acutely ill schizophrenia inpatients had MRI scans before and six weeks after antipsychotic treatment. Thirty-nine matched healthy controls were also scanned, twenty-two of which were scanned a second time six weeks later. Volumes of hippocampal subfields were measured via FreeSurfer v6.0 using a longitudinal analysis pipeline. Before treatment, schizophrenia patients had no significant changes in total hippocampal volume but exhibited significantly greater subfield volumes than controls in bilateral molecular layers of the hippocampus (ML), bilateral granular cell layers of the dentate gyrus (GC-DG), and bilateral cornu ammonis area 4 (CA4). After six weeks of antipsychotic treatment, patients showed volume reductions compared with pretreatment scans in total hippocampus bilaterally, with subfield volume reduction noted in previously enlarged subfields (i.e., bilateral ML, GC-DG and CA4) and in bilateral hippocampal tails, left CA1, CA3, and fimbria. Subfields with volume increases before treatment were reduced to the level of healthy controls (bilateral ML and GC-DG) or near to it (bilateral CA4) after treatment. These results indicate subfield-specific hippocampal hypertrophy prior to treatment, and that these abnormalities were reduced after acute antipsychotic therapy in a dose-related manner together with volume reductions in other areas that were not hypertrophic before treatment.

IS Journal 2015 Journal Article

A Study on a Cabled Seafloor Observatory

  • Fengzhong Qu
  • Zhenduo Wang
  • Hong Song
  • Ying Chen
  • Liuqing Yang

This article discusses the development of Zhejiang University's Zhairuoshan Island Experimental Research Observatory (ZERO). The authors discuss its background, network structure, components, sea trials, and future plans. The authors predict that ZERO will have an important influence as a collaborative center for scientists, engineers, and the public.

AAAI Conference 2011 Conference Paper

Cross Media Entity Extraction and Linkage for Chemical Documents

  • Su Yan
  • Scott Spangler
  • Ying Chen

Text and images are two major sources of information in scientific literature. Information from these two media typically reinforce and complement each other, thus simplifying the process for human to extract and comprehend information. However, machines cannot create the links or have the semantic understanding between images and text. We propose to integrate text analysis and image processing techniques to bridge the gap between the two media, and discover knowledge from the combined information sources, which would be otherwise lost by traditional single-media based mining systems. The focus is on the chemical entity extraction task because images are well known to add value to the textual content in chemical literature. Annotation of US chemical patent documents demonstrates the effectiveness of our proposal.

v2026.09.13