Arrow Research search

Author name cluster

Yong Zhou

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

25 papers
2 author rows

Possible papers

25

AAAI Conference 2026 Conference Paper

Causal Decoupling Domain Generalization for Remote Sensing Change Detection

  • Jiaqi Zhao
  • Jianpeng Xie
  • Yong Zhou
  • Wen-Liang Du
  • Hancheng Zhu
  • Rui Yao

While current state-of-the-art Remote Sensing Change Detection (RSCD) methods can achieve impressive results on individual datasets, they become unreliable in unseen environments and imaging conditions, with performance metrics declining by as much as 60% to 80%. Simultaneously, variable environments and complex imaging conditions are the main characteristics of remote sensing data, calling for generalizable RSCD methods. To address this issue, we propose a novel RSCD method capable of domain generalization—CDDGNet. This method is based on causal decoupling theory, which progressively decouples invariant change features from variable domain features to extract generalizable characteristics. This enables a network trained on a single domain to accurately identify change regions in other domains. Specifically, firstly, the Causal Feature Adaptation Module is proposed to preliminarily decouple and simplify feature information during the encoding process by using wavelet transformation and feature energy spectralization methods. Secondly, the Causal Feature Fusion Module is presented to fully decouple features and aggregate significant change features during the decoding process through frequency domain processing and feature re-attention mechanisms. Thirdly, the Decoupling Effect Loss Function is proposed to optimize the process by evaluating the effectiveness of causal decoupling. Extensive experiments have shown that our model significantly outperforms existing methods across multiple groups of generalization tasks with varying levels of difficulty.

AAAI Conference 2026 Conference Paper

CLIPDet3D: Vision-Language Collaborative Distillation for 3D Object Detection

  • Jiaqi Zhao
  • Huanfeng Hu
  • Yong Zhou
  • Wen-Liang Du
  • Kunyang Sun
  • Rui Yao
  • Qigong Sun

Multi-view 3D object detection plays a vital role in autonomous driving systems due to its ability to perceive complex scenes accurately. However, real-world driving data often exhibits a long-tailed distribution, causing significant drops in detection accuracy for rare categories in existing methods. To mitigate this issue, we propose CLIPDet3D, a novel vision-language collaborative framework for multi-view 3D object detection. First, to tackle the difficulty of capturing the semantic information of rare categories, a Vision-Language Collaborative Learning strategy is proposed to incorporate class-level semantic priors from CLIP. Second, a Depth Feature Contrastive Distillation module is designed to overcome the large depth estimation error for rare categories by aligning depth features between a teacher and a student network. Furthermore, to alleviate the difficulty in focusing on regions of rare categories, a Dual-Stream Prompt Attention mechanism is devised to inject learnable prompts and compute attention along both horizontal and vertical BEV directions. Evaluations on the nuScenes dataset demonstrate that CLIPDet3D achieves state-of-the-art accuracy while maintaining efficient inference.

JMLR Journal 2026 Journal Article

Communication-efficient Distributed Statistical Inference for Massive Data with Heterogeneous Auxiliary Information

  • Miaomiao Yu
  • Zhongfeng Jiang
  • Jiaxuan Li
  • Yong Zhou

Heterogeneous auxiliary information commonly arises in big data due to diverse study settings and privacy constraints. Excluding such indirect evidence often results in a substantial loss of statistical inference efficiency. This article proposes a novel framework for integrating a mixture of individual-level data and multiple external heterogeneous summary statistics by multiplying likelihood functions and confidence densities. Theoretically, we show that the proposed method possesses desirable properties and can achieve statistical efficiency comparable to that of the individual participant data (IPD) estimator, which uses all available individual-level data. Furthermore, we develop a communication-efficient distributed inference procedure for massive datasets with heterogeneous auxiliary information. We demonstrate that the proposed iterative algorithm achieves linear convergence under general conditions or generalized linear models. Finally, extensive simulations and real data applications are conducted to illustrate the performance of the proposed methods. [abs] [ pdf ][ bib ] &copy JMLR 2026. ( edit, beta )

AAAI Conference 2026 Conference Paper

DTTNet: Improving Video Shadow Detection via Dark-Aware Guidance and Tokenized Temporal Modeling

  • Zhicheng Li
  • Kunyang Sun
  • Rui Yao
  • Hancheng Zhu
  • Fuyuan Hu
  • Jiaqi Zhao
  • Zhiwen Shao
  • Yong Zhou

Video shadow detection confronts two entwined difficulties: distinguishing shadows from complex backgrounds and modeling dynamic shadow deformations under varying illumination. To address shadow-background ambiguity, we leverage linguistic priors through the proposed Vision-language Match Module (VMM) and a Dark-aware Semantic Block (DSB), extracting text-guided features to explicitly differentiate shadows from dark objects. Furthermore, we introduce adaptive mask reweighting to downweight penumbra regions during training and apply edge masks at the final decoder stage for better supervision. For temporal modeling of variable shadow shapes, we propose a Tokenized Temporal Block (TTB) that decouples spatiotemporal learning. TTB summarizes cross-frame shadow semantics into learnable temporal tokens, enabling efficient sequence encoding with minimal computation overhead. Comprehensive Experiments on multiple benchmark datasets demonstrate state-of-the-art accuracy and real-time inference efficiency.

TCS Journal 2026 Journal Article

Open shop scheduling problem with a flexible maintenance period: Revisited

  • Yuan Yuan
  • Xin Han
  • Xingwu Liu
  • Yong Zhou
  • Hao Lu

This paper addresses the two-machine open shop scheduling problem where one flexible maintenance period is imposed on the second machine. The maintenance period must start within a given time window and has a fixed duration. The objective is to minimize the makespan. We demonstrate that a ( 1 + ϵ ) ρ -approximation algorithm can be constructed for the studied problem if there exists a ρ-approximation algorithm for the version with a fixed maintenance period. As a consequence, by applying the polynomial-time approximation scheme (PTAS) for the fixed maintenance period, we derive a PTAS for the problem under consideration, thereby solving an open question in the literature. Furthermore, we propose a 4/3-approximation algorithm with O(n) time complexity, which outperforms the previous 3/2-approximation algorithm presented in the literature.

AAAI Conference 2026 Conference Paper

Unified Representation Causal Prompt Distillation for Re-Inference-Free Lifelong Person Re-Identification

  • Jiaqi Zhao
  • Jie Luo
  • Yong Zhou
  • Wen-Liang Du
  • Xixi Li
  • Rui Yao

Lifelong person re-identification (LReID) aims to retrieve the target person from sequentially collected data. Due to significant domain gaps between datasets and the continuous increase of training data from different scenarios, weak inter-domain generalization and catastrophic forgetting issues have remained major challenges for LReID. To tackle these issues, a novel LReID method called Unified Representation Causal Prompt Distillation (URCPD) is proposed. Specifically, to reduce domain gaps among different scene datasets and improve model inter-domain generalization capability, a Feature Decoupling Style Transfer module (FDST) is proposed to map new features into a unified feature space. Furthermore, to reduce the accumulated forgetting of old knowledge during the training stage, a Causal Prompt Distillation module (CPD) is introduced. This module eliminates the re-inference process for distillation and embeds memory prompts to combat catastrophic forgetting. Extensive experiments on five classic LReID seen datasets and seven unseen datasets demonstrate that our method significantly outperforms state-of-the-art methods.

IJCAI Conference 2025 Conference Paper

Beyond Individual and Point: Next POI Recommendation via Region-aware Dynamic Hypergraph with Dual-level Modeling

  • Xixi Li
  • Zhuo Gu
  • Rui Yao
  • Yong Zhou
  • Hancheng Zhu
  • Jiaqi Zhao
  • Wen-Liang Du

Next POI recommendation contributes to the prosperity of various intelligent location-based services. Existing studies focus on exploring sequential patterns and POI interactions using sequential and graph-based methods to enhance recommendation performance. However, they don't effectively exploit geographical information. In addition, methods that focus on modeling mobility patterns using individual limited data may suffer from data sparsity and the information cocoons problem. Moreover, most graph structures focus on adjacent nodes, failing to capture potential high-order associations among POIs. To address these challenges, we propose the Region-aware dynamic Hypergraph learning method with Dual-level interaction Modeling (ReHDM), which exploits users' dynamic mobility beyond individual and point. Specifically, ReHDM utilizes regional encoding to mine the potential spatial relationships among POIs with coarse-grained geographical information. By incorporating POI-level and trajectory-level associations within a hypergraph convolutional network, ReHDM comprehensively captures cross-user collaborative information. Furthermore, ReHDM captures not only dependencies among POIs within each trajectory for a single user, but also the high-order collaborative information across individual user trajectories and associated users' trajectories. Experimental results on three public datasets demonstrate the superiority of ReHDM to the state-of-the-art.

IJCAI Conference 2025 Conference Paper

Counterfactual Knowledge Maintenance for Unsupervised Domain Adaptation

  • Yao Li
  • Yong Zhou
  • Jiaqi Zhao
  • Wen-Liang Du
  • Rui Yao
  • Bing Liu

Traditional unsupervised domain adaptation (UDA) struggles to extract rich semantics due to backbone limitations. Recent large-scale pre-trained visual-language models (VLMs) have shown strong zero-shot learning capabilities in UDA tasks. However, directly using VLMs results in a mixture of semantic and domain-specific information, complicating knowledge transfer. Complex scenes with subtle semantic differences are prone to misclassification, which in turn can result in the loss of features that are crucial for distinguishing between classes. To address these challenges, we propose a novel counterfactual knowledge maintenance UDA framework. Specifically, we employ counterfactual disentanglement to separate the representation of semantic information from domain features, thereby reducing domain bias. Furthermore, to clarify ambiguous visual information specific to classes, we maintain the discriminative knowledge of both visual and textual information. This approach synergistically leverages multimodal information to preserve modality-specific distinguishable features. We conducted extensive experimental evaluations on several public datasets to demonstrate the effectiveness of our method. The source code is available at https: //github. com/LiYaolab/CMKUDA

JBHI Journal 2025 Journal Article

Cross-Domain Human Activity Recognition via Domain Adaptation and Fused Attention

  • Tianyun Zhu
  • Yilin Dong
  • Yong Zhou
  • Changming Zhu
  • Lei Cao

In recent years, the utilization of wearable sensors for Human Activity Recognition (HAR) has garnered significant interest in the fields of medical health monitoring and sports management. However, HAR often suffer the poor generalization from the insufficient labeled data for complex activities. To address this issue, the novel Transfer Component Analysis-Bidirectional Long Short-Term Memory network (TCA-BiLSTM) with the fused attention mechanism is presented in this paper. Specifically, TCA-BiLSTM first leverages the Maximum Mean Difference (MMD) within the Reproducing Kernel Hilbert Space (RKHS) to learn transfer components for sensor-based HAR. These derived transfer components align the data collected from sensors deployed on different body parts, facilitating the mapping of cross-domain HAR data. Then, the two-layer BiLSTM with the novel fused attention mechanism is given to classify the unseen activities, which aims to capture the multi-granularity activity information after the TCA-based domain adaptation. To evaluate the effectiveness of TCA-BiLSTM, a series of experiments were conducted using the DSADS and PAMAP2 datasets. The results demonstrate that TCA-BiLSTM outperforms the state-of-art methods such as DSAN and FNet, achieving performance improvements of 6. 1% and 2. 5%, respectively.

IJCAI Conference 2025 Conference Paper

GSDet: Gaussian Splatting for Oriented Object Detection

  • Zeyu Ding
  • Jiaqi Zhao
  • Yong Zhou
  • Wen-Liang Du
  • Hancheng Zhu
  • Rui Yao

Oriented object detection has advanced with the development of convolutional neural networks (CNNs) and transformers. However, modern detectors still rely on predefined object candidates, such as anchors in CNN-based methods or queries in transformer-based methods, which struggle to capture spatial information effectively. To address the limitations, we propose GSDet, a novel framework that formulates oriented object detection as Gaussian splatting. Specifically, our approach performs detection within a 3D feature space constructed from image features, where 3D Gaussians are employed to represent oriented objects. These 3D Gaussians are projected onto the image plane to form 2D Gaussians, which are then transformed into oriented boxes. Furthermore, we optimize the mean, anisotropic covariance, and confidence scores of these randomly initialized 3D Gaussians, using a decoder that incorporates 3D Gaussian sampling. Moreover, our method exhibits flexibility, enabling adaptive control and a dynamic number of Gaussians during inference. Experiments on 3 datasets indicate that GSDet achieves AP50 gains of 0. 7% on DIOR-R, 0. 3% on DOTA-v1. 0, and 0. 55% on DOTA-v1. 5 when evaluated with adaptive control and outperforms mainstream detectors.

IROS Conference 2025 Conference Paper

Learning to Exploit Leg Odometry Enables Terrain-Aware Quadrupedal Locomotion

  • Yong Zhou
  • Jiawei Jiang
  • Bo Du
  • Zengmao Wang

The geometry of terrain is crucial for developing terrain-aware locomotion policies. Recent advancements in quadrupedal locomotion based on learning rely on depth information obtained from LiDARs and depth cameras. Despite the capabilities of these locomotion policies on terrains, they pose challenges in processing high-dimensional data in real time with onboard hardware. In this study, we develop a lightweight framework that utilizes only the intrinsic sensors of a quadrupedal robot to facilitate terrain-aware locomotion. We introduce a learning-based leg odometry, integrated with a locomotion policy trained through reinforcement learning. Utilizing blind localization from leg odometry alongside a pre-constructed height map enables the robot to navigate steps and stairs without incident. We assess the efficacy of our framework through simulations, where our results indicate that the robot achieves up to a 17% improvement in successful traversal rates and requires fewer point samples. By compensating for slippage during locomotion, our learning-based leg odometry surpasses traditional inertialleg odometry. Lastly, we validate the practical applicability of our models on a real robot, confirming their effectiveness in real-world settings.

IJCAI Conference 2025 Conference Paper

Modality-Guided Dynamic Graph Fusion and Temporal Diffusion for Self-Supervised RGB-T Tracking

  • Shenglan Li
  • Rui Yao
  • Yong Zhou
  • Hancheng Zhu
  • Kunyang Sun
  • Bing Liu
  • Zhiwen Shao
  • Jiaqi Zhao

To reduce the reliance on large-scale annotations, self-supervised RGB-T tracking approaches have garnered significant attention. However, the omission of the object region by erroneous pseudo-label or the introduction of background noise affects the efficiency of modality fusion, while pseudo-label noise triggered by similar object noise can further affect the tracking performance. In this paper, we propose GDSTrack, a novel approach that introduces dynamic graph fusion and temporal diffusion to address the above challenges in self-supervised RGB-T tracking. GDSTrack dynamically fuses the modalities of neighboring frames, treats them as distractor noise, and leverages the denoising capability of a generative model. Specifically, by constructing an adjacency matrix via an Adjacency Matrix Generator (AMG), the proposed Modality-guided Dynamic Graph Fusion (MDGF) module uses a dynamic adjacency matrix to guide graph attention, focusing on and fusing the object’s coherent regions. Temporal Graph-Informed Diffusion (TGID) models MDGF features from neighboring frames as interference, and thus improving robustness against similar-object noise. Extensive experiments conducted on four public RGB-T tracking datasets demonstrate that GDSTrack outperforms the existing state-of-the-art methods. The source code is available at https: //github. com/LiShenglana/GDSTrack.

EAAI Journal 2025 Journal Article

Prebuilt spatiotemporal index: An exploration of efficient real-time data storage in intelligent transportation systems

  • Yiran Shao
  • Kangshuai Zhang
  • Yong Zhou
  • Zhenwu Chen
  • Yang Yang
  • Lei Peng

With the widespread deployment of Internet of Things (IoT) devices in Intelligent Transportation Systems (ITS), real-time spatiotemporal data has grown rapidly. Hybrid data-bases have emerged as the mainstream solution for managing such data. The LSM R*-tree, which integrates Log-Structured Merge-trees (LSM-trees) for time-series ingestion and R*-trees for spatial queries, is now a widely used spatiotemporal index. However, storage schemes based on LSM R*-tree structures involve real-time index construction, leading to significant overhead and write latency. To address these challenges, this paper proposes a prebuilt index-based data storage workflow that shifts index construction ahead of data writing. This approach allows the database to directly apply the prebuilt index at write time, thereby minimizing real-time construction costs and enhancing write performance. To support index prebuild, we propose a lightweight embedding scheme, Index2Vec, specifically designed for R*-tree structures. Based on this, we extend the Transformer architecture and develop the R*-tree Prediction Network for ITS (RTPN4ITS), which achieves efficient inference on resource-constrained edge devices. Experimental results show that the prebuilt R*-tree index improves query efficiency by up to 90% over time-index-only schemes and enhances write performance by nearly 50% compared to real-time indexing. The proposed RTPN4ITS model achieves robust accuracy across varying traffic densities, reaching 85% accuracy in dense conditions. Moreover, the Index2Vec embedding enhances the Transformer’s structural awareness. In summary, this paper proposes an efficient prebuilt indexing strategy and lightweight embedding-based model for real-time spatiotemporal data management in ITS.

AAAI Conference 2025 Conference Paper

Structured IB: Improving Information Bottleneck with Structured Feature Learning

  • Hanzhe Yang
  • Youlong Wu
  • Dingzhu Wen
  • Yong Zhou
  • Yuanming Shi

The Information Bottleneck (IB) principle has emerged as a promising approach for enhancing the generalization, robustness, and interpretability of deep neural networks, demonstrating efficacy across image segmentation, document clustering, and semantic communication. Among IB implementations, the IB Lagrangian method, employing Lagrangian multipliers, is widely adopted. While numerous methods for the optimizations of IB Lagrangian based on variational bounds and neural estimators are feasible, their performance is highly dependent on the quality of their design, which is inherently prone to errors. To address this limitation, we introduce Structured IB, a framework for investigating potential structured features. By incorporating auxiliary encoders to extract missing informative features, we generate more informative representations. Our experiments demonstrate superior prediction accuracy and task-relevant information preservation compared to the original IB Lagrangian method, even with reduced network size.

TCS Journal 2024 Journal Article

Flow shop scheduling problems with transportation constraints revisited

  • Yan Lan
  • Yuan Yuan
  • Yinling Wang
  • Xin Han
  • Yong Zhou

This paper investigates two flow shop scheduling problems with two machines A and B, and a single transporter V. In the first problem, the transporter V is located at machine A initially, each job has to be processed first on A, then transported to B for further processing. While in the second problem, the transporter V is located at machine B initially, each job needs to be processed first on A, then on B, finally transported to the destination. In both problems, the transporter V can carry up to c (where c is a constant and c ≥ 1 ) jobs at a time and the objective is to minimize the makespan. For the former problem, the best approximation algorithm guarantees a worst-case ratio bound of ( 5 3 + ε ), while for the latter one, the best approximation algorithm provides a worst-case ratio bound of 2. For each problem, we design a polynomial-time approximation scheme (PTAS).

AAAI Conference 2024 Conference Paper

LRANet: Towards Accurate and Efficient Scene Text Detection with Low-Rank Approximation Network

  • Yuchen Su
  • Zhineng Chen
  • Zhiwen Shao
  • Yuning Du
  • Zhilong Ji
  • Jinfeng Bai
  • Yong Zhou
  • Yu-Gang Jiang

Recently, regression-based methods, which predict parameterized text shapes for text localization, have gained popularity in scene text detection. However, the existing parameterized text shape methods still have limitations in modeling arbitrary-shaped texts due to ignoring the utilization of text-specific shape information. Moreover, the time consumption of the entire pipeline has been largely overlooked, leading to a suboptimal overall inference speed. To address these issues, we first propose a novel parameterized text shape method based on low-rank approximation. Unlike other shape representation methods that employ data-irrelevant parameterization, our approach utilizes singular value decomposition and reconstructs the text shape using a few eigenvectors learned from labeled text contours. By exploring the shape correlation among different text contours, our method achieves consistency, compactness, simplicity, and robustness in shape representation. Next, we propose a dual assignment scheme for speed acceleration. It adopts a sparse assignment branch to accelerate the inference speed, and meanwhile, provides ample supervised signals for training through a dense assignment branch. Building upon these designs, we implement an accurate and efficient arbitrary-shaped text detector named LRANet. Extensive experiments are conducted on several challenging benchmarks, demonstrating the superior accuracy and efficiency of LRANet compared to state-of-the-art methods. Code is available at: https://github.com/ychensu/LRANet.git

JMLR Journal 2024 Journal Article

Semi-supervised Inference for Block-wise Missing Data without Imputation

  • Shanshan Song
  • Yuanyuan Lin
  • Yong Zhou

We consider statistical inference for single or low-dimensional parameters in a high-dimensional linear model under a semi-supervised setting, wherein the data are a combination of a labelled block-wise missing data set of a relatively small size and a large unlabelled data set. The proposed method utilises both labelled and unlabelled data without any imputation or removal of the missing observations. The asymptotic properties of the estimator are established under regularity conditions. Hypothesis testing for low-dimensional coefficients are also studied. Extensive simulations are conducted to examine the theoretical results. The method is evaluated on the Alzheimer’s Disease Neuroimaging Initiative data. [abs] [ pdf ][ bib ] &copy JMLR 2024. ( edit, beta )

TIST Journal 2023 Journal Article

Attention-guided Adversarial Attack for Video Object Segmentation

  • Rui Yao
  • Ying Chen
  • Yong Zhou
  • Fuyuan Hu
  • Jiaqi Zhao
  • Bing Liu
  • Zhiwen Shao

Video Object Segmentation (VOS) methods have made many breakthroughs with the help of the continuous development and advancement of deep learning. However, the deep learning model is vulnerable to malicious adversarial attacks, which mislead the model to make wrong decisions by adding adversarial perturbation that humans cannot perceive to the input image. Threats to deep learning models remind us that video object segmentation methods are also vulnerable to attacks, thereby threatening their security. Therefore, we study adversarial attacks on the VOS task to better identify the vulnerabilities of the VOS method, which in turn provides an opportunity to improve its robustness. In this paper, we propose an attention-guided adversarial attack method, which uses spatial attention blocks to capture features with global dependencies to construct correlations between consecutive video frames, and performs multipath aggregation to effectively integrate spatial-temporal perturbation, thereby guiding the deconvolution network to generate adversarial examples with strong attack capability. Specifically, the class loss function is designed to enable the deconvolution network to better activate noise in other regions and suppress the activation related to the object class based on the enhanced feature map of the object class. At the same time, attentional feature loss is designed to enhance the transferability against attack. The experimental results on the DAVIS dataset show that the proposed attention-guided adversarial attack method can significantly reduce the segmentation accuracy of OSVOS, and the J & F mean on DAVIS 2016 can reach 73.6% drop rate. The generated adversarial examples are also highly transferable to other video object segmentation models.

JMLR Journal 2023 Journal Article

Distributed Algorithms for U-statistics-based Empirical Risk Minimization

  • Lanjue Chen
  • Alan T.K. Wan
  • Shuyi Zhang
  • Yong Zhou

Empirical risk minimization, where the underlying loss function depends on a pair of data points, covers a wide range of application areas in statistics including pairwise ranking and survival analysis. The common empirical risk estimator obtained by averaging values of a loss function over all possible pairs of observations is essentially a U-statistic. One well-known problem with minimizing U-statistic type empirical risks, is that the computational complexity of U-statistics increases quadratically with the sample size. When faced with big data, this poses computational challenges as the colossal number of observation pairs virtually prohibits centralized computing to be performed on a single machine. This paper addresses this problem by developing two computationally and statistically efficient methods based on the divide-and-conquer strategy on a decentralized computing system, whereby the data are distributed among machines to perform the tasks. One of these methods is based on a surrogate of the empirical risk, while the other method extends the one-step updating scheme in classical M-estimation to the case of pairwise loss. We show that the proposed estimators are as asymptotically efficient as the benchmark global U-estimator obtained under centralized computing. As well, we introduce two distributed iterative algorithms to facilitate the implementation of the proposed methods, and conduct extensive numerical experiments to demonstrate their merit. [abs] [ pdf ][ bib ] &copy JMLR 2023. ( edit, beta )

EAAI Journal 2022 Journal Article

Edge-aware and spectral–spatial information aggregation network for multispectral image semantic segmentation

  • Di Zhang
  • Jiaqi Zhao
  • Jingyang Chen
  • Yong Zhou
  • Boyu Shi
  • Rui Yao

Semantic segmentation is a fundamental task in the field of remote sensing image intelligent interpretation and computer vision. Multispectral remote sensing images have attracted more and more researchers’ attention because they can accurately describe different types of reflection spectra. However, inaccurate multispectral feature description leads to edge semantic ambiguity and misclassification of small objects. In this article, we propose a novel network named edge-aware and spectral–spatial information aggregation net (ESSANet) to capture both high-level semantic features and low-level edge details for semantic segmentation of remote sensing images. Specifically, on the one hand, in order to improve the representation ability of discriminant features, we design a two-stream spectral–spatial feature extraction network via 3D hybrid convolution and multi-level aggregation network. On the other hand, in order to eliminate the effect of edge semantic ambiguity, we develop a siamese edge-aware structure and multi-stage edge loss function. Experimental results show that our method achieved 3. 5% and 4. 09% mean intersection over union (mIoU) score improvements and 2. 59% and 3. 32% Kappa score improvements compared with the competitive baseline algorithm on the SEN12MS and US3D datasets, respectively. In addition, the method proposed in this paper also achieves a better trade-off between speed and accuracy.

JBHI Journal 2021 Journal Article

Attention-RefNet: Interactive Attention Refinement Network for Infected Area Segmentation of COVID-19

  • Titinunt Kitrungrotsakul
  • Qingqing Chen
  • Huitao Wu
  • Yutaro Iwamoto
  • Hongjie Hu
  • Wenchao Zhu
  • Chao Chen
  • Fangyi Xu

COVID-19 pneumonia is a disease that causes an existential health crisis in many people by directly affecting and damaging lung cells. The segmentation of infected areas from computed tomography (CT) images can be used to assist and provide useful information for COVID-19 diagnosis. Although several deep learning-based segmentation methods have been proposed for COVID-19 segmentation and have achieved state-of-the-art results, the segmentation accuracy is still not high enough (approximately 85%) due to the variations of COVID-19 infected areas (such as shape and size variations) and the similarities between COVID-19 and non-COVID-infected areas. To improve the segmentation accuracy of COVID-19 infected areas, we propose an interactive attention refinement network (Attention RefNet). The interactive attention refinement network can be connected with any segmentation network and trained with the segmentation network in an end-to-end fashion. We propose a skip connection attention module to improve the important features in both segmentation and refinement networks and a seed point module to enhance the important seeds (positions) for interactive refinement. The effectiveness of the proposed method was demonstrated on public datasets (COVID-19CTSeg and MICCAI) and our private multicenter dataset. The segmentation accuracy was improved to more than 90%. We also confirmed the generalizability of the proposed network on our multicenter dataset. The proposed method can still achieve high segmentation accuracy.

EAAI Journal 2021 Journal Article

Cluster-based fine-to-coarse superpixel segmentation

  • Xiangjun Li
  • Yong Zhou
  • Xinping Zhang
  • Su Xu
  • Peng Yu

As an image preprocessing technology, superpixel segmentation has become an important tool in the field of computer vision. How to obtain a more accurate, faster, and easier-to-apply superpixel segmentation algorithm is a problem faced by researchers. In this paper, a cluster-based fine-to-coarse superpixel segmentation (FCSS) algorithm is proposed. By introducing color thresholds and depth thresholds with practical physical meanings as algorithm parameters, high-quality segmentation with fewer superpixels is achieved. It not only reduces the complexity of the upper application, but also provides an easy to understand interface. Superpixel segmentation methods often cannot achieve high-quality segmentation through a set of parameters. Experimental results show that FCSS can achieve finer segmentation by setting different parameters, and the segmentation results are superior to other algorithms. When the number of superpixels is 100, the segmentation performance of FCSS is better than that of existing state-of-the-art methods.

TIST Journal 2021 Journal Article

Multi-Stage Fusion and Multi-Source Attention Network for Multi-Modal Remote Sensing Image Segmentation

  • Jiaqi Zhao
  • Yong Zhou
  • Boyu Shi
  • Jingsong Yang
  • Di Zhang
  • Rui Yao

With the rapid development of sensor technology, lots of remote sensing data have been collected. It effectively obtains good semantic segmentation performance by extracting feature maps based on multi-modal remote sensing images since extra modal data provides more information. How to make full use of multi-model remote sensing data for semantic segmentation is challenging. Toward this end, we propose a new network called Multi-Stage Fusion and Multi-Source Attention Network ((MS) 2 -Net) for multi-modal remote sensing data segmentation. The multi-stage fusion module fuses complementary information after calibrating the deviation information by filtering the noise from the multi-modal data. Besides, similar feature points are aggregated by the proposed multi-source attention for enhancing the discriminability of features with different modalities. The proposed model is evaluated on publicly available multi-modal remote sensing data sets, and results demonstrate the effectiveness of the proposed method.

TIST Journal 2020 Journal Article

Video Object Segmentation and Tracking

  • Rui Yao
  • Guosheng Lin
  • Shixiong Xia
  • Jiaqi Zhao
  • Yong Zhou

Object segmentation and object tracking are fundamental research areas in the computer vision community. These two topics are difficult to handle some common challenges, such as occlusion, deformation, motion blur, scale variation, and more. The former contains heterogeneous object, interacting object, edge ambiguity, and shape complexity; the latter suffers from difficulties in handling fast motion, out-of-view, and real-time processing. Combining the two problems of Video Object Segmentation and Tracking (VOST) can overcome their respective difficulties and improve their performance. VOST can be widely applied to many practical applications such as video summarization, high definition video compression, human computer interaction, and autonomous vehicles. This survey aims to provide a comprehensive review of the state-of-the-art VOST methods, classify these methods into different categories, and identify new trends. First, we broadly categorize VOST methods into Video Object Segmentation (VOS) and Segmentation-based Object Tracking (SOT). Each category is further classified into various types based on the segmentation and tracking mechanism. Moreover, we present some representative VOS and SOT methods of each time node. Second, we provide a detailed discussion and overview of the technical characteristics of the different methods. Third, we summarize the characteristics of the related video dataset and provide a variety of evaluation metrics. Finally, we point out a set of interesting future works and draw our own conclusions.

TCS Journal 2003 Journal Article

Normal conditions for inference relations and injective models

  • Zhaohui Zhu
  • Xi'an Xiao
  • Yong Zhou
  • Wujia Zhu

Although fruitful representation results induced by some kinds of injective models, e. g. , filtered, ranked and quasi-linear injective models, etc. , have been established in the literature, it is still an open problem to characterize the family of all injective inference relations in terms of rules. The type of postulates appearing in recent literature seems to be unable to characterize this family. This brings up an interesting theoretical problem: What kind of injective inference relations may be characterized by existent types of postulates? This paper makes an initial step to answer this question. To this end, a notion of a normal condition is introduced, which subsumes all Horn and non-Horn conditions presented in the literature. We obtain some results on injective models generating inferences characterized by normal conditions, and show that these injective models must be specific standard models. Moreover, for any set of injective models determined only by a structural property of preferential orders, if the family of inference relations induced by it can be characterized by normal conditions, then it must be a subset of filtered models in this circumstance. Thus, its associated inference relations satisfy the non-Horn rule disjunctive rationality.

v2026.09.13