Arrow Research search

Author name cluster

Xiaolong Liu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

8 papers
2 author rows

Possible papers

8

EAAI Journal 2025 Journal Article

Image steganography with high embedding capacity based on multi-target adversarial attack

  • Xiaolong Liu
  • Minghuang Shen
  • Jiayi Liu
  • Qiwen Wu

Deep learning-based steganography techniques utilizing generative adversarial networks have attracted considerable attention due to their ability to produce realistic images that serve as effective carriers for hidden information. However, as the capacity for embedding information increases, the quality of the generated covert images tends to decline significantly. To address these challenges and enhance both the quality of covert images and data-hiding performance, we propose a high-capacity image steganography method known as Multi-Target Adversarial Image Steganography (MTAIS). This method leverages a multi-target adversarial attack technique to effectively conceal high-capacity secret information within images. The proposed scheme adapts the fully connected layer of the recognition model and transforms the undirected adversarial attack into a directed adversarial attack targeting multiple outputs, allowing for fine-tuning of the base model without extensive retraining. We conducted comprehensive experiments to benchmark the proposed scheme against several established deep learning-based steganography schemes. The results indicate that the proposed scheme consistently outperforms its competitors across various evaluation metrics. Notably, our method preserves the quality of the cover image, ensuring visual integrity while achieving a high capacity for embedding secret information. The experimental results underscore the advantages of the proposed scheme in terms of performance and efficiency, establishing it as a robust solution for high-capacity image steganography.

AAAI Conference 2025 Conference Paper

Reverse Region-to-Entity Annotation for Pixel-Level Visual Entity Linking

  • Zhengfei Xu
  • Sijia Zhao
  • Yanchao Hao
  • Xiaolong Liu
  • Lili Li
  • Yuyang Yin
  • Bo Li
  • Xi Chen

Visual Entity Linking (VEL) is a crucial task for achieving fine-grained visual understanding, matching objects within images (visual mentions) to entities in a knowledge base. Previous VEL tasks rely on textual inputs, but writing queries for complex scenes can be challenging. Visual inputs like clicks or bounding boxes offer a more convenient alternative. Therefore, we propose a new task, Pixel-Level Visual Entity Linking (PL-VEL), which uses pixel masks from visual inputs to refer to objects, supplementing reference methods for VEL. To facilitate research on this task, we have constructed the MaskOVEN-Wiki dataset through an entirely automatic reverse region-entity annotation framework. This dataset contains over 5 million annotations aligning pixel-level regions with entity-level labels, which will advance visual understanding towards fine-grained. Moreover, as pixel masks correspond to semantic regions in an image, we enhance previous patch-interacted attention with region-interacted attention by a visual semantic tokenization approach. Manual evaluation results indicate that the reverse annotation framework achieved a 94.8% annotation success rate. Experimental results show that models trained on this dataset improved accuracy by 18 points compared to zero-shot models. Additionally, the semantic tokenization method achieved a 5-point accuracy improvement over the trained baseline.

NeurIPS Conference 2024 Conference Paper

ScaleKD: Strong Vision Transformers Could Be Excellent Teachers

  • Jiawei Fan
  • Chao Li
  • Xiaolong Liu
  • Anbang Yao

In this paper, we question if well pre-trained vision transformer (ViT) models could be used as teachers that exhibit scalable properties to advance cross architecture knowledge distillation research, in the context of adopting mainstream large-scale visual recognition datasets for evaluation. To make this possible, our analysis underlines the importance of seeking effective strategies to align (1) feature computing paradigm differences, (2) model scale differences, and (3) knowledge density differences. By combining three closely coupled components namely *cross attention projector*, *dual-view feature mimicking* and *teacher parameter perception* tailored to address the alignment problems stated above, we present a simple and effective knowledge distillation method, called *ScaleKD*. Our method can train student backbones that span across a variety of convolutional neural network (CNN), multi-layer perceptron (MLP), and ViT architectures on image classification datasets, achieving state-of-the-art knowledge distillation performance. For instance, taking a well pre-trained Swin-L as the teacher model, our method gets 75. 15\%|82. 03\%|84. 16\%|78. 63\%|81. 96\%|83. 93\%|83. 80\%|85. 53\% top-1 accuracies for MobileNet-V1|ResNet-50|ConvNeXt-T|Mixer-S/16|Mixer-B/16|ViT-S/16|Swin-T|ViT-B/16 models trained on ImageNet-1K dataset from scratch, showing 3. 05\%|3. 39\%|2. 02\%|4. 61\%|5. 52\%|4. 03\%|2. 62\%|3. 73\% absolute gains to the individually trained counterparts. Intriguingly, when scaling up the size of teacher models or their pre-training datasets, our method showcases the desired scalable properties, bringing increasingly larger gains to student models. We also empirically show that the student backbones trained by our method transfer well on downstream MS-COCO and ADE20K datasets. More importantly, our method could be used as a more efficient alternative to the time-intensive pre-training paradigm for any target student model on large-scale datasets if a strong pre-trained ViT is available, reducing the amount of viewed training samples up to 195$\times$. The code is available at *https: //github. com/deep-optimization/ScaleKD*.

NeurIPS Conference 2023 Conference Paper

Augmentation-Free Dense Contrastive Knowledge Distillation for Efficient Semantic Segmentation

  • Jiawei Fan
  • Chao Li
  • Xiaolong Liu
  • Meina Song
  • Anbang Yao

In recent years, knowledge distillation methods based on contrastive learning have achieved promising results on image classification and object detection tasks. However, in this line of research, we note that less attention is paid to semantic segmentation. Existing methods heavily rely on data augmentation and memory buffer, which entail high computational resource demands when applying them to handle semantic segmentation that requires to preserve high-resolution feature maps for making dense pixel-wise predictions. In order to address this problem, we present Augmentation-free Dense Contrastive Knowledge Distillation (Af-DCD), a new contrastive distillation learning paradigm to train compact and accurate deep neural networks for semantic segmentation applications. Af-DCD leverages a masked feature mimicking strategy, and formulates a novel contrastive learning loss via taking advantage of tactful feature partitions across both channel and spatial dimensions, allowing to effectively transfer dense and structured local knowledge learnt by the teacher model to a target student model while maintaining training efficiency. Extensive experiments on five mainstream benchmarks with various teacher-student network pairs demonstrate the effectiveness of our approach. For instance, DeepLabV3-Res18|DeepLabV3-MBV2 model trained by Af-DCD reaches 77. 03\%|76. 38\% mIOU on Cityscapes dataset when choosing DeepLabV3-Res101 as the teacher, setting new performance records. Besides that, Af-DCD achieves an absolute mIOU improvement of 3. 26\%|3. 04\%|2. 75\%|2. 30\%|1. 42\% compared with individually trained counterpart on Cityscapes|Pascal VOC|Camvid|ADE20K|COCO-Stuff-164K. Code is available at https: //github. com/OSVAI/Af-DCD.

YNIMG Journal 2023 Journal Article

Differential responses in the mirror neuron system during imitation of individual emotional facial expressions and association with autistic traits

  • Weihua Zhao
  • Qi Liu
  • Xiaolu Zhang
  • Xinwei Song
  • Zhao Zhang
  • Peng Qing
  • Xiaolong Liu
  • Siyu Zhu

The mirror neuron system (MNS), including the inferior frontal gyrus (IFG), inferior parietal lobule (IPL) and superior temporal sulcus (STS) plays an important role in action representation and imitation and may be dysfunctional in autism spectrum disorder (ASD). However, it's not clear how these three regions respond and interact during the imitation of different basic facial expressions and whether the pattern of responses is influenced by autistic traits. Thus, we conducted a natural facial expression (happiness, angry, sadness and fear) imitation task in 100 healthy male subjects where expression intensity was measured using facial emotion recognition software (FaceReader) and MNS responses were recorded using functional near-infrared spectroscopy (fNIRS). Autistic traits were measured using the Autism Spectrum Quotient questionnaire. Results showed that imitation of happy expressions produced the highest expression intensity but a small deactivation in MNS responses, suggesting a lower processing requirement compared to other expressions. A cosine similarity analysis indicated a distinct pattern of MNS responses during imitation of each facial expression with functional intra-hemispheric connectivity between the left IPL and left STS being significantly higher during happy compared to other expressions, while inter-hemispheric connectivity between the left and right IPL differed between imitation of fearful and sad expressions. Furthermore, functional connectivity changes during imitation of each different expression could reliably predict autistic trait scores. Overall, the results provide evidence for distinct patterns of functional connectivity changes between MNS regions during imitation of different emotions which are also associated with autistic traits.

ICLR Conference 2023 Conference Paper

NORM: Knowledge Distillation via N-to-One Representation Matching

  • Xiaolong Liu
  • Lujun Li 0001
  • Chao Li
  • Anbang Yao

Existing feature distillation methods commonly adopt the One-to-one Representation Matching between any pre-selected teacher-student layer pair. In this paper, we present $N$-to-$O$ne $R$epresentation $M$atching (NORM), a new two-stage knowledge distillation method, which relies on a simpleFeature Transform (FT) module consisting of two linear layers. In view of preserving the intact information learnt by the teacher network, during training, our FT module is merely inserted after the last convolutional layer of the student network. The first linear layer projects the student representation to a feature space having $N$ times feature channels than the teacher representation from the last convolutional layer, and the second linear layer contracts the expanded output back to the original feature space. By sequentially splitting the expanded student representation into $N$ non-overlapping feature segments having the same number of feature channels as the teacher's, they can be readily forced to approximate the intact teacher representation simultaneously, formulating a novel many-to-one representation matching mechanism conditioned on a single teacher-student layer pair. After training, such an FT module will be naturally merged into the subsequent fully connected layer thanks to its linear property, introducing no extra parameters or architectural modifications to the student network at inference. Extensive experiments on different visual recognition benchmarks demonstrate the leading performance of our method. For instance, the ResNet18|MobileNet|ResNet50-1/4 model trained by NORM reaches 72.14%|74.26%|68.03% top-1 accuracy on the ImageNet dataset when using a pre-trained ResNet34|ResNet50|ResNet50 model as the teacher, achieving an absolute improvement of 2.01%|4.63%|3.03% against the individually trained counterpart. Code is available at https://github.com/OSVAI/NORM.

IROS Conference 2016 Conference Paper

The design and experiments of a small wheel-legged mobile robot system with two robotic arms

  • Qingkai Chang
  • Xiaolong Liu
  • Wenfu Xu
  • Lei Yan 0011
  • Bingsong Yang

In this paper, we developed a small wheel-legged mobile robot system, which could walk on different road environments using wheels or legs. It is composed of mechanical, sensor and control subsystems. The mechanical subsystem includes a wheel-legged mobile platform, a rigid robotic arm and a flexible arm. The mobile platform provides a variety of movement ways to meet the requirement of different mobility. The rigid arm (denoted by arm-a) is a serial manipulator with 4-DOFs. It can be used to grasp and manipulate payloads. The flexible arm (denoted by arm-b) is a manipulator with continuous curve, and a camera is mounted on arm-b. So it can be used to provide visual inspection and measurement information. The sensor subsystem is composed of ultrasonic sensors mounted on the platform and a WIFI camera mounted on arm-b. It provides measurement information and visual inspection for remote control. The control subsystem includes an embedded controller and a PC computer. The former is developed based on an ARM microprocessor, on which the real-time operation system-RT-Thread system runs. The mission decomposition and trajectory planning algorithms are programed in C language and run in the PC. At last, typical experiments are performed. Experiment results verified the robot's mobility, operation capability and remote-control function.

v2026.09.13