Arrow Research search

Author name cluster

Xiaolin Zhang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

14 papers
2 author rows

Possible papers

14

AAAI Conference 2026 Short Paper

AEFGL: Reverse Auction and Value Evaluation-Based Federated Graph Learning Incentive Mechanism (Student Abstract)

  • Xin Chang
  • Lixin Liu
  • Jingyu Wang
  • Jinling Yu
  • Xiaolin Zhang

Federated Graph Learning enables multiple clients to collaboratively train graph models while protecting local private data. However, most studies have assumed that all clients contribute data voluntarily and actively. Without reasonable incentives, clients are often reluctant to contribute personal data for model training. Furthermore, the budget for incentives is limited, and if clients with low-quality graph data are incentivized to participate in training, it will negatively impact the training performance of all parties in the system. To address this, we propose AEFGL, a Reverse Auction and Value Evaluation-Based Incentive Mechanism for Federated Graph Learning. First, we design a reverse auction mechanism combining graph structural attribute motifs with client production value. Then, we propose a method for evaluating client production value based on the comparison of the client's expected reward and actual value. This mechanism can incentivize clients with high-quality graph data to participate in training within budget constraints, thereby improving the model quality. Experimental results validate the superiority of the AEFGL mechanism and the economic properties it satisfies.

EAAI Journal 2026 Journal Article

For automated cell culture systems: A high-speed style-transfer and prior-guided network for high-precision cell monitoring in bright-field microscopy

  • Jing Meng
  • Xiaolin Zhang
  • Jue Hou
  • Sumin Qi
  • Fei Ma
  • Zenan Wang
  • Dengwang Li

Bright-field (BF) microscopy is commonly used for cell growth monitoring in spatially constrained automated culture systems, but its low image contrast and diverse cell morphologies limit imaging clarity and the accuracy of confluence analysis. To address this, this study proposes a lightweight BF to phase contrast (PC) style transfer network (LW-P2P), which achieves fast and robust contrast enhancement through a brightness centralization method and a multi-criterion optimization strategy. Building on this, a MA-TPP segmentation network is developed, it incorporates a texture prior prompt (TPP) and a mixed attention mechanism (MA) to significantly enhance the recognition of low-contrast and halo-affected regions and enable fully automated confluence calculation. Evaluated on four datasets spanning five cell lines under varied densities and defocus conditions, LW-P2P improves signal-to-noise ratio by 21. 02% and achieves 8 × faster inference than state-of-the-art methods. MA-TPP Net achieves state-of-the-art segmentation on three diverse datasets covering six cell types, reducing confluence error to 0. 95%—a 26. 36% improvement over prior best methods. This study establishes an efficient and intelligent imaging-analysis framework for automated cell culture monitoring, with promising potential in regenerative medicine and related fields.

EAAI Journal 2024 Journal Article

A novel CT image segmentation model

  • Jingdong Yang
  • Han Wang
  • Wei Liu
  • Xianyou Zheng
  • Xiaolin Zhang
  • Shaoqing Yu

Convolutional neural networks (CNNs) can be used for clinical medical image segmentation to improve detection efficiency and accuracy. However, because of the fixed size of convolution kernel and receptive field for existing CNN model, some associated pixel features are ignored and segmentation performance is impaired. Therefore, we propose an effective image segmentation model, COPANet, which uses RepVGG as the backbone and replaces the standard convolution with dilated convolution of multiple kernels to increase the receptive field size, so that COPANet can make full use of pixel correlations and acquire contexture semantic information. We also build skip-connections between global and local features, where a parallel attention (PANet) is employed to extract important location information in the downsampling. PANet is a fused mechanism integrating channel attention with spatial attention in parallel, which can fully extract more global information. We also apply weighted combined loss to reduce the effect of class imbalance of foreground and background pixels on segmentation performance and speed up convergence. In addition, we conduct experiments on 480 cases of clinical CT sinus from Shanghai Tongji Hospital and 236 cases of CT patella fracture from Shanghai Sixth People's Hospital. The evaluation indexes after 5-Fold cross-validation are as follows: Precision is 92. 76% and 96. 55%, Recall is 88. 25% and 57. 58%, Specificity is 99. 85% and 96. 55%, IOU is 91. 08% and 77. 38%, Hausdorff Distance is 2. 8135 and 11. 1879, respectively. Compared with the state-of-the-art models, COPANet has higher segmentation accuracy and better generalization performance, which can assist the clinical diagnosis.

ICRA Conference 2024 Conference Paper

BEE-Net: Bridging Semantic and Instance with Gated Encoding and Edge Constraint for Efficient Panoptic Segmentation

  • Xinyang Huang
  • Guanghui Zhang
  • Dongchen Zhu
  • Yunpeng Sun
  • Wenjun Shi
  • Gang Ye
  • Yang Xiao
  • Lei Wang 0202

Panoptic segmentation is a challenging perception task, which can help robots to comprehensively perceive the surrounding environment. In the task, we notice that semantic, instance, and panoptic have rich relations, however, which are rarely explored. In this work, we propose a novel panoptic, instance, and semantic bridged network to delve into the reciprocal relation. To make semantic and instance benefit from each other, we design a novel Gated Encoding (GE) module, incorporating complementary cues between semantic and instance heads through the gated mechanism. In addition, a novel edge-aware consistency constraint among edges of each task is presented, which exhaustedly exploits geometric constraints, to boost the segmentation quality of challenging edges. Experimental results on the Cityscapes and MS-COCO datasets demonstrate that our approach achieves state-of-the-art performance in an efficient CNN-based paradigm, attaining a balance between accuracy and efficiency.

ICRA Conference 2024 Conference Paper

CVFormer: Learning Circum-View Representation and Consistency for Vision-Based Occupancy Prediction via Transformers

  • Zhengqi Bai
  • Wenjun Shi
  • Dongchen Zhu
  • Hanlong Kang
  • Guanghui Zhang
  • Gang Ye
  • Yang Xiao
  • Lei Wang 0202

With the increasing demands for perception accuracy in autonomous driving, there is a growing focus on fine-grained 3D semantic occupancy prediction. Effectively representing detailed three-dimensional scenes has become a significant challenge in the development of this task. In this paper, we present a novel transformer-based framework named CVFormer, which leverages two-dimensional circum-views from the ego to excavate three-dimensional features of the surrounding environment. Circum-views provide a novel solution for effectively addressing the representation of dense and fine-grained scenes. Specifically, a multi-attention module CTMA is designed for fusing temporal features from circum-views to fully exploit the spatiotemporal correlations between frames and capture more comprehensive clues. Furthermore, a novel 2D projection constraint is established by observing objects from different perspective directions, and multiple 3D constraints based on object invariance and semantic consistency are also conducted for supervising the network, which enhances its performance of understanding the scene. Experimental results on nuScenes dataset demonstrate that the proposed CVFormer obviously outperforms existing methods for occupancy prediction.

ICRA Conference 2023 Conference Paper

Fast Extrinsic Calibration for Multiple Inertial Measurement Units in Visual-Inertial System

  • Youwei Yu
  • Yanqing Liu
  • Fengjie Fu
  • Sihan He
  • Dongchen Zhu
  • Lei Wang 0202
  • Xiaolin Zhang
  • Jiamao Li

In this paper, we propose a fast extrinsic calibration method for fusing multiple inertial measurement units (MIMU) to improve visual-inertial odometry (VIO) localization accuracy. Currently, data fusion algorithms for MIMU highly depend on the number of inertial sensors. Based on the assumption that extrinsic parameters between inertial sensors are perfectly calibrated, the fusion algorithm provides better localization accuracy with more IMUs, while neglecting the effect of extrinsic calibration error. Our method builds two non-linear least-squares problems to estimate the MIMU relative position and orientation separately, independent of external sensors and inertial noises online estimation. Then we give the general form of the virtual IMU (VIMU) method and propose its propagation on manifold. We perform our method on datasets, our self-made sensor board, and board with different IMUs, validating the superiority of our method over competing methods concerning speed, accuracy, and robustness. In the simulation experiment, we show that only fusing two IMUs with our calibration method to predict motion can rival nine IMUs. Real-world experiments demonstrate better localization accuracy of the VIO integrated with our calibration method and VIMU propagation on manifold.

IROS Conference 2023 Conference Paper

FeatDANet: Feature-level Domain Adaptation Network for Semantic Segmentation

  • Jiao Li
  • Wenjun Shi
  • Dongchen Zhu
  • Guanghui Zhang
  • Xiaolin Zhang
  • Jiamao Li

Unsupervised domain adaptation (UDA) is proposed to better adapt the network trained on labeled synthetic data to unlabeled real-world data for addressing the annotation cost. However, most of these methods pay more attention to domain distributions in input and output stages while ignoring the important differences in semantic expressions and local details in middle feature stages. Therefore, a novel UDA network named FeatDANet is presented to align feature-level domain distributions at each encoder layer. Specifically, two attention-based modules abbreviated as IFAM and DFLM are designed and implemented by mixing queries and keys between domains for advisable domain adaptation. The former realizes Inter-domain Features Alignment by transferring feature style, and the latter achieves Domain-invariant Features Learning robustly for the domain shift. Furthermore, FeatDANet is constructed as a self-training network with three weight-sharing branches, and an improved pseudo-labels learning strategy is suggested by identifying more confident pseudolabels and maximizing the use of pseudo-labels. It increases the participation of unlabeled data and also ensures stability in training. Extensive experiments show that FeatDANet achieves state-of-the-art performances on the tasks of GTA→Cityscapes and Synthia→Cityscapes.

IJCAI Conference 2023 Conference Paper

G2Pxy: Generative Open-Set Node Classification on Graphs with Proxy Unknowns

  • Qin Zhang
  • Zelin Shi
  • Xiaolin Zhang
  • Xiaojun Chen
  • Philippe Fournier-Viger
  • Shirui Pan

Node classification is the task of predicting the labels of unlabeled nodes in a graph. State-of-the-art methods based on graph neural networks achieve excellent performance when all labels are available during training. But in real-life, models are of ten applied on data with new classes, which can lead to massive misclassification and thus significantly degrade performance. Hence, developing open-set classification methods is crucial to determine if a given sample belongs to a known class. Existing methods for open-set node classification generally use transductive learning with part or all of the features of real unseen class nodes to help with open-set classification. In this paper, we propose a novel generative open-set node classification method, i. e. , G2Pxy, which follows a stricter inductive learning setting where no information about unknown classes is available during training and validation. Two kinds of proxy unknown nodes, inter-class unknown proxies and external unknown proxies are generated via mixup to efficiently anticipate the distribution of novel classes. Using the generated proxies, a closed-set classifier can be transformed into an open-set one, by augmenting it with an extra proxy classifier. Under the constraints of both cross entropy loss and complement entropy loss, G2Pxy achieves superior effectiveness for unknown class detection and known class classification, which is validated by experiments on bench mark graph datasets. Moreover, G2Pxy does not have specific requirement on the GNN architecture and shows good generalizations.

IROS Conference 2022 Conference Paper

J-RR: Joint Monocular Depth Estimation and Semantic Edge Detection Exploiting Reciprocal Relations

  • Deming Wu
  • Dongchen Zhu
  • Guanghui Zhang
  • Wenjun Shi
  • Xiaolin Zhang
  • Jiamao Li

Depth estimation and semantic edge detection are two key tasks in computer vision, which have made great progress. To date, how to associatively predict the depth and the semantic edge is rarely explored. In this work, we first propose a flexible two-branch framework that can make the two tasks take advantage of each other, achieving a win-win situation. Specifically, for the semantic edge detection branch, an Enhanced Edge Weighting strategy (EEW) is designed, which learns weight information from the by-product of depth branch, depth edge, to enhance edge perception in features. Meanwhile, we make depth estimation benefit from semantic edge detection through introducing Depth Edge Semantic Classification module (DESC). Furthermore, a double reconstruction (D-reconstruction) approach is presented, together with semantic edge-guided disparity smoothing loss to mitigate the ambiguities of the self-supervised manner for depth estimation. Experiments on the Cityscapes dataset demonstrate that our framework outperforms the state-of-the-art method in depth estimation along with a significant improvement in semantic edge detection.

IROS Conference 2022 Conference Paper

Spatiotemporally Enhanced Photometric Loss for Self-Supervised Monocular Depth Estimation

  • Tianyu Zhang
  • Dongchen Zhu
  • Guanghui Zhang
  • Wenjun Shi
  • Yanqing Liu
  • Xiaolin Zhang
  • Jiamao Li

Recovering depth information from a single image is a long-standing challenge, and self-supervised depth estimation methods have gradually attracted attention due to not relying on high-cost ground truth. Constructing an accurate photometric loss based on photometric consistency is crucial for these self-supervised methods to obtain high-quality depth maps. However, the photometric loss in most studies treats all pixels indiscriminately, resulting in poor performance. In this paper, we propose two modules based on the spatial and temporal cues to refine the photometric loss. Delving into the geometric model of photometric consistency, we introduce a depth-aware pixel correspondence module (DPC) inside the monocular depth estimation pipeline. It reduces the uncertainty of photometric errors by applying the homography matrix to the projection of corresponding pixels in far regions instead of the fundamental matrix. Furthermore, we design an omnidirectional auto-masking module (OA) to boost the robustness of our model, which utilizes temporal sequences to generate disturbance poses and hypothetical views to distin-guish dynamic objects with different directions that violate the photometric consistency. Experiments on the KITTI and the Make3d datasets reveal that our framework achieves state-of-the-art performance.

IROS Conference 2021 Conference Paper

Camera Parameters Aware Motion Segmentation Network with Compensated Optical Flow

  • Xianshun Wang
  • Dongchen Zhu
  • Shaojie Xu
  • Wenjun Shi
  • Yanqing Liu
  • Jiamao Li
  • Xiaolin Zhang

Learning to distinguish independent moving objects from the observed optical flow with a moving camera remains challenging. In this work, we first present a novel camera pose compensation (CPC) scheme. With the help of ingenious geometric analysis, it breaks the observed optical flow into patterns that are easier to interpret for the motion segmentation network. Secondly, we further refine such compensation with a camera parameter aware (CPA) module to account for poses’ errors in the CPC processing and enhance the entire network’s tolerance to noises. Additionally, an MMPNet is developed to intensify the identification ability of overall motion patterns. It reaches a larger receptive field with a bottom-up information transmission structure and integrates motion information at different granularities. We demonstrate the benefits of our framework on FlyingThings3D and Monkaa datasets. Without the complement of semantic information, our approach outperforms the top methods for moving objects segmentation.

ICRA Conference 2020 Conference Paper

3DCFS: Fast and Robust Joint 3D Semantic-Instance Segmentation via Coupled Feature Selection

  • Liang Du 0004
  • Jingang Tan
  • Xiangyang Xue 0001
  • Lili Chen
  • Hongkai Wen 0001
  • Jianfeng Feng
  • Jiamao Li
  • Xiaolin Zhang

We propose a novel fast and robust 3D point clouds segmentation framework via coupled feature selection, named 3DCFS, that jointly performs semantic and instance segmentation. Inspired by the human scene perception process, we design a novel coupled feature selection module, named CFSM, that adaptively selects and fuses the reciprocal semantic and instance features from two tasks in a coupled manner. To further boost the performance of the instance segmentation task in our 3DCFS, we investigate a loss function that helps the model learn to balance the magnitudes of the output embedding dimensions during training, which makes calculating the Euclidean distance more reliable and enhances the generalizability of the model. Extensive experiments demonstrate that our 3DCFS outperforms state-of-the-art methods on benchmark datasets in terms of accuracy, speed and computational cost. Codes are available at: https://github.com/Biotan/3DCFS.

IROS Conference 2020 Conference Paper

RegionNet: Region-feature-enhanced 3D Scene Understanding Network with Dual Spatial-aware Discriminative Loss

  • Guanghui Zhang
  • Dongchen Zhu
  • Xiaoqing Ye
  • Wenjun Shi
  • Minghong Chen
  • Jiamao Li
  • Xiaolin Zhang

Neural networks have recently achieved impressive success in semantic and instance segmentation on 2D images. However, their capabilities have not been fully explored to address semantic instance segmentation on unstructured 3D point cloud data. Digging into the regional feature representation to boost point cloud comprehension, we propose a region-feature-enhanced structure consisting of adaptive regional feature complementary (ARFC) module and affinity-based regional relational reasoning (AR 3 ) module. The ARFC module aims to complement low-level features of sparse regions adaptively. The AR 3 module emphasizes on mining the potential reasoning relationships between high-level features based on affinity. Both the ARFC and AR 3 modules are plug-and-play. Besides, a novel dual spatial-aware discriminative loss is proposed to improve the discrimination of instance embedding. Our proposal-free point cloud instance segmentation network (RegionNet) equipped with the region-feature-enhanced structure and dual spatial-aware discriminative loss achieves state-of-the-art performance on S3DIS dataset and ScanNet-v2 dataset.

IROS Conference 2020 Conference Paper

Richer Aggregated Features for Optical Flow Estimation with Edge-aware Refinement

  • Xianshun Wang
  • Dongchen Zhu
  • Jiafei Song
  • Yanqing Liu
  • Jiamao Li
  • Xiaolin Zhang

Recent CNN-based optical flow approaches have a separated structure of feature extraction and flow estimation. The core task of optical flow is finding the corresponding points while rich representation is just the key part of such matching problems. However, the prior work usually pays more attention to the design of flow decoder than the feature extraction. In this paper, we present a novel optical flow estimation network to enrich the feature representation of each pyramid level, with a hierarchical dilated architecture and a bottom-up aggregation scheme. In addition, inspired by edge guided classical methods, we bring the edge-aware idea into our approach and propose an edge-aware refinement (EAR) subnetwork to handle motion boundaries. Using the same decoding structure as PWC-Net, our network outperforms it by a large margin and leads all its derivatives both on KITTI-2012 and KITTI-2015. Further performance analysis proves the effectiveness of proposed ideas.

v2026.09.13