Arrow Research search

Author name cluster

Yue Cao

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

32 papers
2 author rows

Possible papers

32

EAAI Journal 2026 Journal Article

A dual-path lightweight detector with hybrid attention for real-time object detection

  • Boyang Yu
  • Zixuan Li
  • Yue Cao
  • Xu Zhang
  • Wansu Lim
  • William Liu

Unmanned aerial vehicle (UAV)-based object detection presents significant challenges, including pronounced variations in object scale and limited computational resources. To address these issues, this paper proposes the Dual-path and Bimodal-attention-enhanced Network (DBYNet), a real-time detection framework optimized for UAV applications. DBYNet adopts a deployment-oriented design that integrates a dual-path backbone for spatial–semantic feature decoupling, a hybrid attention mechanism for enhanced contextual modeling, and lightweight optimization strategies to improve inference efficiency. Specifically, a shallow lightweight branch preserves fine-grained spatial details, while a deep branch with deformable convolutions captures high-level semantic features, and the proposed hybrid attention combines Overlapping Cross Attention (OCA) and Channel-spatial Bimodal Attention (CAB) to strengthen feature interaction. In addition, Quantization-Aware Training and temperature-aware distillation are employed to reduce model complexity without compromising accuracy. Extensive experiments on the VisDrone2019 dataset demonstrate that DBYNet achieves a favorable accuracy–efficiency trade-off, particularly improving robustness for small, densely distributed, and low-visibility targets in challenging UAV scenarios.

EAAI Journal 2026 Journal Article

An intelligent vision-based method for real-time pig disease identification through postural feature analysis

  • Zhe Yin
  • Yue Cao
  • Hong Feng
  • Qiqi Guo
  • Xuan Wang
  • Zhenyu Liu

The health of pig populations is critical to production efficiency and economic viability. By integrating postural characteristics indicative of disease, a model capable of accommodating complex postural variations can facilitate non-contact, low-cost, real-time disease monitoring in pigs. Building on the baseline You Only Look Once Version 10 (YOLOv10) deep learning model, this study proposes an improved model with three core innovations: a lightweight backbone network, spatial and channel reconstruction convolution, and large-scale keypoint attention. The lightweight backbone enhances feature extraction for subtle postural cues, the attention mechanism strengthens focus on key disease postures, and the optimised feature fusion structure improves feature representation and robustness to complex posture variations. Compared with the baseline model, the proposed method demonstrates significantly improved performance. It achieves a mean average precision of 97. 66 per cent, corresponding to a 3. 9 percentage point increase over the baseline. In particular, the detection precision for African swine fever improves from 94. 7 per cent to 98. 9 per cent, while the harmonic mean score reaches 91. 98 per cent, reflecting a 3. 72 percentage point improvement. Despite these notable gains in accuracy, the proposed method reduces the parameter count by 1. 80 million and the computational complexity by 1. 5 giga floating-point operations, maintaining high computational efficiency without sacrificing detection performance. The results demonstrate the potential practicality of the proposed method for real-time pig disease detection and provide a reliable technical basis for future deployment in intelligent livestock farming systems.

AAAI Conference 2026 Conference Paper

Automatic Translational Correction of Multi-View Coronary Angiography Based on Auto-Annotation Data Generation

  • Yue Cao
  • Zhuo Zhang
  • Shuai Xiao
  • Jialin Li
  • Guipeng Lan
  • Jiabao Wen
  • Jiachen Yang

Multi-view automatic translational correction (ATC) in coronary angiography (CAG) is critical for intraoperative automatic diagnosis, in which deep learning playing a key role. However, heartbeat-induced soft matching errors and costly annotations make it difficult to build high-quality, large-scale datasets for calibration algorithm training. The training of clinical models is difficult to fulfill, as existing datasets differ significantly from real CAG in both style and structure. To address this challenge, we propose a novel high-quality data synthesis method for annotation-free ATC. We fully automated the construction of a labeled, high-fidelity dataset for training matching models. An evolutionary algorithm is introduced for global optimization of translation estimation, mitigating epipolar constraint violations caused by vascular deformation and enabling reliable correction across large viewpoint differences. Furthermore, a theoretical analysis is presented, demonstrating that error propagation between adjacent views is more accurate than direct estimation across distant views. Our experiments on clinical datasets demonstrate that our method not only significantly outperforms weakly supervised learning approaches, but also performs comparably to fully supervised methods. Moreover, it exhibits remarkable multicenter generalizability.

AAAI Conference 2026 Conference Paper

CNM-UNet: Continuous Ordinary Differential Equations for Medical Image Segmentation

  • Tianqi Xu
  • Yashi Zhu
  • Quansong He
  • Yue Cao
  • Kaishen Wang
  • Zhang Yi
  • Tao He

Integrating Ordinary Differential Equations (ODEs) with U-shaped neural networks has emerged as a novel direction in medical image segmentation. Current networks predominantly employ discretization methods incorporating ODEs. However, these methods face inherent trade-offs between model compactness, computational accuracy, and efficiency. Continuous ODE solutions were rarely studied because they face three limitations: high computational costs, long training time, and poor generalization ability. To address these limitations, we propose an innovative Continuous Neural Memory ODE UNet (CNM-UNet), which replaces all hierarchical decoder layers in vanilla UNet with a single Continuous Neural Memory ODEs Block (CNM-Block) decoder, significantly reducing computation costs and improving training efficiency. CNM-UNet leverages ODEs' dynamic properties to establish continuous temporal feature extraction. For alleviating the generalization problem, a DUal SElf-updated (DUSE) strategy based on test-time adaptation principles is introduced to enhance cross-domain generalization. Experimental results demonstrate CNM-UNet's comprehensive advantages in computational capacity, convergence speed, and cross-domain adaptability, offering new insights for practical deployment of continuous ODE methodologies for medical image segmentation.

AAAI Conference 2026 Conference Paper

MAGIC: Mastering Physical Adversarial Generation in Context Through Collaborative LLM Agents

  • Yun Xing
  • Nhat Chung
  • Jie Zhang
  • Yue Cao
  • Ivor Tsang
  • Yang Liu
  • Lei Ma
  • Qing Guo

Physical adversarial attacks in driving scenarios can expose critical vulnerabilities in visual perception models. However, developing such attacks remains non-trivial due to diverse real-world environmental influences. Existing approaches either struggle to generalize to dynamic environments or fail to achieve consistent physical attack performance. To address these challenges, we propose MAGIC (Mastering Physical Adversarial Generation In Context), a novel framework powered by multi-modal LLM agents to automatically understand the scene context during testing time and generate adversarial patches through synergistic interaction of language and vision understanding. Specifically, MAGIC orchestrates three specialized LLM agents: the adv-patch generation agent masters the creation of deceptive patches via strategic prompt manipulation for text-to-image models; the adv-patch deployment agent ensures contextual coherence by determining optimal deployment strategies based on scene understanding; and the self-examination agent completes this trilogy by providing critical oversight and iterative refinement of both processes. We validate our approach with both digital and physical scenarios, i.e., nuImage and real-world scenes, where both statistical and visual results demonstrate that our MAGIC is powerful and effective for attacking widely applied object detection systems, such as YOLO and DETR series.

AAAI Conference 2026 Conference Paper

Robust Noise Modeling for Spike Camera via Time-Interval Quantification and Spike-DSLR Multimodal Dataset in Low-Light Imaging

  • Yue Cao
  • Sizhao Li
  • Liguo Zhang

The inherent differences between spike cameras and traditional frame-based cameras lead to more complex and diverse noise characteristics, particularly under extremely low-light conditions. Existing noise modeling approaches for spike camera predominantly rely on inter-spike intervals (ISI) for noise quantification, which often results in inaccurate noise characterization. Moreover, current datasets for spike camera image reconstruction tasks are either synthetic or lack corresponding high-quality reference images, severely limiting rigorous evaluation of noise modeling methods. To address this limitation, we propose a multimodal noise modeling framework for spike camera that integrates insights from traditional frame-based imaging into spike imaging. Specifically, we introduce a time-interval-based quantification method inspired by the exposure-time concept used in traditional frame-based cameras, enabling accurate noise characterization for spike camera. Furthermore, we present the Spike-DSLR Multimodal Dataset (SDMD), the first real-world dataset capturing aligned multimodal data pairs from spike cameras and Digital Single-Lens Reflex (DSLR) cameras, explicitly designed for evaluating spike camera noise models. Experimental results on SDMD demonstrate that our noise modeling approach significantly enhances spike camera image reconstruction quality under low-light conditions, achieving more than 1.6 dB improvement in PSNR compared to existing state-of-the-art methods. This validates both the necessity and effectiveness of adopting a multimodal perspective in spike camera noise modeling.

NeurIPS Conference 2025 Conference Paper

DepthVanish: Optimizing Adversarial Interval Structures for Stereo-Depth-Invisible Patches

  • Yun Xing
  • Yue Cao
  • Nhat Chung
  • Jie Zhang
  • Ivor Tsang
  • Ming-Ming Cheng
  • Yang Liu
  • Lei Ma

Stereo depth estimation is a critical task in autonomous driving and robotics, where inaccuracies (such as misidentifying nearby objects as distant) can lead to dangerous situations. Adversarial attacks against stereo depth estimation can help revealing vulnerabilities before deployment. Previous works have shown that repeating optimized textures can effectively mislead stereo depth estimation in digital settings. However, our research reveals that these naively repeated textures perform poorly in physical implementations, $\textit{i. e. }$, when deployed as patches, limiting their practical utility for stress-testing stereo depth estimation systems. In this work, for the first time, we discover that introducing regular intervals among the repeated textures, creating a grid structure, significantly enhances the patch attack performance. Through extensive experimentation, we analyze how variations of this novel structure influence the adversarial effectiveness. Based on these insights, we develop a novel stereo depth attack that jointly optimizes both the interval structure and texture elements. Our generated adversarial patches can be inserted into any scenes and successfully attack advanced stereo depth estimation methods of different paradigms, $\textit{i. e. }$, RAFT-Stereo and STTR. Most critically, our patch can also attack commercial RGB-D cameras (Intel RealSense) in real-world conditions, demonstrating their practical relevance for security assessment of stereo systems. The code is officially released at: https: //github. com/WiWiN42/DepthVanish

AAAI Conference 2024 Conference Paper

FedMut: Generalized Federated Learning via Stochastic Mutation

  • Ming Hu
  • Yue Cao
  • Anran Li
  • Zhiming Li
  • Chengwei Liu
  • Tianlin Li
  • Mingsong Chen
  • Yang Liu

Although Federated Learning (FL) enables collaborative model training without sharing the raw data of clients, it encounters low-performance problems caused by various heterogeneous scenarios. Due to the limitation of dispatching the same global model to clients for local training, traditional Federated Average (FedAvg)-based FL models face the problem of easily getting stuck into a sharp solution, which results in training a low-performance global model. To address this problem, this paper presents a novel FL approach named FedMut, which mutates the global model according to the gradient change to generate several intermediate models for the next round of training. Each intermediate model will be dispatched to a client for local training. Eventually, the global model converges into a flat area within the range of mutated models and has a well-generalization compared with the global model trained by FedAvg. Experimental results on well-known datasets demonstrate the effectiveness of our FedMut approach in various data heterogeneity scenarios.

ICLR Conference 2024 Conference Paper

IRAD: Implicit Representation-driven Image Resampling against Adversarial Attacks

  • Yue Cao
  • Tianlin Li
  • Xiaofeng Cao 0002
  • Ivor W. Tsang
  • Yang Liu 0003
  • Qing Guo 0005

We introduce a novel approach to counter adversarial attacks, namely, image resampling. Image resampling transforms a discrete image into a new one, simulating the process of scene recapturing or rerendering as specified by a geometrical transformation. The underlying rationale behind our idea is that image resampling can alleviate the influence of adversarial perturbations while preserving essential semantic information, thereby conferring an inherent advantage in defending against adversarial attacks. To validate this concept, we present a comprehensive study on leveraging image resampling to defend against adversarial attacks. We have developed basic resampling methods that employ interpolation strategies and coordinate shifting magnitudes. Our analysis reveals that these basic methods can partially mitigate adversarial attacks. However, they come with apparent limitations: the accuracy of clean images noticeably decreases, while the improvement in accuracy on adversarial examples is not substantial.We propose implicit representation-driven image resampling (IRAD) to overcome these limitations. First, we construct an implicit continuous representation that enables us to represent any input image within a continuous coordinate space. Second, we introduce SampleNet, which automatically generates pixel-wise shifts for resampling in response to different inputs. Furthermore, we can extend our approach to the state-of-the-art diffusion-based method, accelerating it with fewer time steps while preserving its defense capability. Extensive experiments demonstrate that our method significantly enhances the adversarial robustness of diverse deep models against various attacks while maintaining high accuracy on clean images.

ICRA Conference 2024 Conference Paper

Open X-Embodiment: Robotic Learning Datasets and RT-X Models: Open X-Embodiment Collaboration

  • Abby O'Neill
  • Abdul Rehman
  • Abhiram Maddukuri
  • Abhishek Gupta 0004
  • Abhishek Padalkar
  • Abraham Lee
  • Acorn Pooley
  • Agrim Gupta

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for many applications. Can such a consolidation happen in robotics? Conventionally, robotic learning methods train a separate model for every application, every robot, and even every environment. Can we instead train "generalist" X-robot policy that can be adapted efficiently to new robots, tasks, and environments? In this paper, we provide datasets in standardized data formats and models to make it possible to explore this possibility in the context of robotic manipulation, alongside experimental results that provide an example of effective X-robot policies. We assemble a dataset from 22 different robots collected through a collaboration between 21 institutions, demonstrating 527 skills (160266 tasks). We show that a high-capacity model trained on this data, which we call RT-X, exhibits positive transfer and improves the capabilities of multiple robots by leveraging experience from other platforms. The project website is robotics-transformer-x. github.io.

IROS Conference 2024 Conference Paper

Recurrent Non-Rigid Point Cloud Registration

  • Yue Cao
  • Ziang Cheng
  • Hongdong Li

Non-rigid point cloud registration remains a significant challenge in 3D computer vision due to the complexity of structural deforms, lack of overlaps, and sensitivity to initialization. This paper introduces a framework inspired by the recent success in recurrent architecture, adapted to accommodate the unique characteristics of point clouds. More specifically, we design a recurrent update network block for progressively refining local registration results under a local rigidity assumption, starting from an initial global SE(3) alignment. Through comparison, our method consistently outperforms competing methods in standard metrics, achieving a 33% reduction in EPE on the 4DLoMatch benchmark compared to the second-best method. To the best of our knowledge, the proposed method is the first to successfully demonstrate that the recurrent update strategy can effectively address the non-rigid registration task with large displacement, significant deform, and low overlap. The source code and the model will be released at http://dummy.url/.

ICML Conference 2023 Conference Paper

Continual Learners are Incremental Model Generalizers

  • Jaehong Yoon
  • Sung Ju Hwang
  • Yue Cao

Motivated by the efficiency and rapid convergence of pre-trained models for solving downstream tasks, this paper extensively studies the impact of Continual Learning (CL) models as pre-trainers. We find that, in both supervised and unsupervised CL, the transfer quality of representations does not show a noticeable degradation of fine-tuning performance but rather increases gradually. This is because CL models can learn improved task-general features when easily forgetting task-specific knowledge. Based on this observation, we suggest a new unsupervised CL framework with masked modeling, which aims to capture fluent task-generic representation during training. Furthermore, we propose a new fine-tuning scheme, GLobal Attention Discretization (GLAD), that preserves rich task-generic representation during solving downstream tasks. The model fine-tuned with GLAD achieves competitive performance and can also be used as a good pre-trained model itself. We believe this paper breaks the barriers between pre-training and fine-tuning steps and leads to a sustainable learning framework in which the continual learner incrementally improves model generalization, yielding better transfer to unseen tasks.

IROS Conference 2023 Conference Paper

End-to-End Point Cloud Registration via Rotation Equivariant Descriptors

  • Yue Cao
  • Yujiao Shi 0002
  • Ziang Cheng
  • Hongdong Li

Point cloud registration (PCR) aims to recover the rigid transformation between two noisy, unordered point sets. This task is typically tackled by establishing point-wise correspondences, and solving the rigid transformation between the two sets. Since descriptor-based methods find correspondences by matching the feature space distance, a powerful and rotation-robust point feature extractor is critical to the success of this task. Existing methods assume soft rotation invariance/equivariance through the means of training augmentation, rotational discretization or pre-alignment of patches. In contrast, this paper proposes a new method which generates fully rotation invariant and equivariant descriptors by construction. For each keypoint patch, our network extracts not only a rotation invariant descriptor for establishing corre-spondences, but also a rotation equivariant one. The rotation equivariant descriptor allows relative transformation to be directly recovered from a single correspondence pair, unlike standard methods that require three correspondences. This design significantly reduces iteration number of RANSAC and guarantees high registration recall when the inlier ratio of estimated correspondences is low. Extensive experiments have demonstrated that the proposed method outperforms state-of-art methods in the same category even after much fewer RANSAC iterations.

AAAI Conference 2023 Short Paper

Lightweight Transformer for Multi-Modal Object Detection (Student Abstract)

  • Yue Cao
  • Yanshuo Fan
  • Junchi Bin
  • Zheng Liu

It has become a common practice for many perceptual systems to integrate information from multiple sensors to improve the accuracy of object detection. For example, autonomous vehicles use visible light, and infrared (IR) information to ensure that the car can cope with complex weather conditions. However, the accuracy of the algorithm is usually a trade-off between the computational complexity and memory consumption. In this study, we evaluate the performance and complexity of different fusion operators in multi-modal object detection tasks. On top of that, a Poolformer-based fusion operator (PoolFuser) is proposed to enhance the accuracy of detecting targets without compromising the efficiency of the detection framework.

ICML Conference 2023 Conference Paper

One Transformer Fits All Distributions in Multi-Modal Diffusion at Scale

  • Fan Bao
  • Shen Nie
  • Kaiwen Xue
  • Chongxuan Li
  • Shi Pu 0002
  • Yaole Wang
  • Gang Yue
  • Yue Cao

This paper proposes a unified diffusion framework (dubbed UniDiffuser) to fit all distributions relevant to a set of multi-modal data in one model. Our key insight is – learning diffusion models for marginal, conditional, and joint distributions can be unified as predicting the noise in the perturbed data, where the perturbation levels (i. e. timesteps) can be different for different modalities. Inspired by the unified view, UniDiffuser learns all distributions simultaneously with a minimal modification to the original diffusion model – perturbs data in all modalities instead of a single modality, inputs individual timesteps in different modalities, and predicts the noise of all modalities instead of a single modality. UniDiffuser is parameterized by a transformer for diffusion models to handle input types of different modalities. Implemented on large-scale paired image-text data, UniDiffuser is able to perform image, text, text-to-image, image-to-text, and image-text pair generation by setting proper timesteps without additional overhead. In particular, UniDiffuser is able to produce perceptually realistic samples in all tasks and its quantitative results (e. g. , the FID and CLIP score) are not only superior to existing general-purpose models but also comparable to the bespoken models (e. g. , Stable Diffusion and DALL-E 2) in representative tasks (e. g. , text-to-image generation).

ICML Conference 2023 Conference Paper

Revisiting Discriminative vs. Generative Classifiers: Theory and Implications

  • Chenyu Zheng
  • Guoqiang Wu
  • Fan Bao
  • Yue Cao
  • Chongxuan Li
  • Jun Zhu 0001

A large-scale deep model pre-trained on massive labeled or unlabeled data transfers well to downstream tasks. Linear evaluation freezes parameters in the pre-trained model and trains a linear classifier separately, which is efficient and attractive for transfer. However, little work has investigated the classifier in linear evaluation except for the default logistic regression. Inspired by the statistical efficiency of naive Bayes, the paper revisits the classical topic on discriminative vs. generative classifiers. Theoretically, the paper considers the surrogate loss instead of the zero-one loss in analyses and generalizes the classical results from binary cases to multiclass ones. We show that, under mild assumptions, multiclass naive Bayes requires $O(\log n)$ samples to approach its asymptotic error while the corresponding multiclass logistic regression requires $O(n)$ samples, where $n$ is the feature dimension. To establish it, we present a multiclass $\mathcal{H}$-consistency bound framework and an explicit bound for logistic loss, which are of independent interests. Simulation results on a mixture of Gaussian validate our theoretical findings. Experiments on various pre-trained deep vision models show that naive Bayes consistently converges faster as the number of data increases. Besides, naive Bayes shows promise in few-shot cases and we observe the "two regimes” phenomenon in pre-trained supervised models. Our code is available at https: //github. com/ML-GSAI/Revisiting-Dis-vs-Gen-Classifiers.

NeurIPS Conference 2022 Conference Paper

Could Giant Pre-trained Image Models Extract Universal Representations?

  • Yutong Lin
  • Ze Liu
  • Zheng Zhang
  • Han Hu
  • Nanning Zheng
  • Stephen Lin
  • Yue Cao

Frozen pretrained models have become a viable alternative to the pretraining-then-finetuning paradigm for transfer learning. However, with frozen models there are relatively few parameters available for adapting to downstream tasks, which is problematic in computer vision where tasks vary significantly in input/output format and the type of information that is of value. In this paper, we present a study of frozen pretrained models when applied to diverse and representative computer vision tasks, including object detection, semantic segmentation and video action recognition. From this empirical analysis, our work answers the questions of what pretraining task fits best with this frozen setting, how to make the frozen setting more flexible to various downstream tasks, and the effect of larger model sizes. We additionally examine the upper bound of performance using a giant frozen pretrained model with 3 billion parameters (SwinV2-G) and find that it reaches competitive performance on a varied set of major benchmarks with only one shared frozen base network: 60. 0 box mAP and 52. 2 mask mAP on COCO object detection test-dev, 57. 6 val mIoU on ADE20K semantic segmentation, and 81. 7 top-1 accuracy on Kinetics-400 action recognition. With this work, we hope to bring greater attention to this promising path of freezing pretrained image models.

NeurIPS Conference 2021 Conference Paper

Bootstrap Your Object Detector via Mixed Training

  • Mengde Xu
  • Zheng Zhang
  • Fangyun Wei
  • Yutong Lin
  • Yue Cao
  • Stephen Lin
  • Han Hu
  • Xiang Bai

We introduce MixTraining, a new training paradigm for object detection that can improve the performance of existing detectors for free. MixTraining enhances data augmentation by utilizing augmentations of different strengths while excluding the strong augmentations of certain training samples that may be detrimental to training. In addition, it addresses localization noise and missing labels in human annotations by incorporating pseudo boxes that can compensate for these errors. Both of these MixTraining capabilities are made possible through bootstrapping on the detector, which can be used to predict the difficulty of training on a strong augmentation, as well as to generate reliable pseudo boxes thanks to the robustness of neural networks to labeling error. MixTraining is found to bring consistent improvements across various detectors on the COCO dataset. In particular, the performance of Faster R-CNN~\cite{ren2015faster} with a ResNet-50~\cite{he2016deep} backbone is improved from 41. 7 mAP to 44. 0 mAP, and the accuracy of Cascade-RCNN~\cite{cai2018cascade} with a Swin-Small~\cite{liu2021swin} backbone is raised from 50. 9 mAP to 52. 8 mAP.

IROS Conference 2021 Conference Paper

Run Like a Dog: Learning Based Whole-Body Control Framework for Quadruped Gait Style Transfer

  • Fulong Yin
  • Annan Tang
  • Liangwei Xu
  • Yue Cao
  • Yu Zheng 0001
  • Zhengyou Zhang
  • Xiangyu Chen 0001

In this paper, a learning-based whole-body loco-motion controller is proposed, which enables quadruped robots to perform running in the style of real animals. We use a low-level controller based on multi-rigid body dynamics to calculate desired torques for each joint, while the high-level neural network policy planning the expected gait and foothold. The policy is trained with reinforcement learning, so that the robot can track a variety of trajectories according to the gait patterns recorded from real-world dogs. We transfer the walking and running gait style to quadrupeds in simulation, involving pace, trot, high-speed gallop and natural transitions. The performance is evaluated by the synchronization rate of contact state between the policy result and the recorded sequence. In the experiments, the robot runs steadily at a speed of 2 m/s and showcases a notable synchronization rate of about 80%. Without prior knowledge, the policy demonstrates a realistic foothold distribution that covers the central area of the torso, which is prevalent in running animals.

AAAI Conference 2020 Conference Paper

MultiSumm: Towards a Unified Model for Multi-Lingual Abstractive Summarization

  • Yue Cao
  • Xiaojun Wan
  • Jinge Yao
  • Dian Yu

Automatic text summarization aims at producing a shorter version of the input text that conveys the most important information. However, multi-lingual text summarization, where the goal is to process texts in multiple languages and output summaries in the corresponding languages with a single model, has been rarely studied. In this paper, we present MultiSumm, a novel multi-lingual model for abstractive summarization. The MultiSumm model uses the following training regime: (I) multi-lingual learning that contains language model training, auto-encoder training, translation and backtranslation training, and (II) joint summary generation training. We conduct experiments on summarization datasets for five rich-resource languages: English, Chinese, French, Spanish, and German, as well as two low-resource languages: Bosnian and Croatian. Experimental results show that our proposed model significantly outperforms a multi-lingual baseline model. Specifically, our model achieves comparable or even better performance than models trained separately on each language. As an additional contribution, we construct the first summarization dataset for Bosnian and Croatian, containing 177, 406 and 204, 748 samples, respectively.

NeurIPS Conference 2020 Conference Paper

Parametric Instance Classification for Unsupervised Visual Feature learning

  • Yue Cao
  • Zhenda Xie
  • Bin Liu
  • Yutong Lin
  • Zheng Zhang
  • Han Hu

This paper presents parametric instance classification (PIC) for unsupervised visual feature learning. Unlike the state-of-the-art approaches which do instance discrimination in a dual-branch non-parametric fashion, PIC directly performs a one-branch parametric instance classification, revealing a simple framework similar to supervised classification and without the need to address the information leakage issue. We show that the simple PIC framework can be as effective as the state-of-the-art approaches, i. e. SimCLR and MoCo v2, by adapting several common component settings used in the state-of-the-art approaches. We also propose two novel techniques to further improve effectiveness and practicality of PIC: 1) a sliding-window data scheduler, instead of the previous epoch-based data scheduler, which addresses the extremely infrequent instance visiting issue in PIC and improves the effectiveness; 2) a negative sampling and weight update correction approach to reduce the training time and GPU memory consumption, which also enables application of PIC to almost unlimited training images. We hope that the PIC framework can serve as a simple baseline to facilitate future study. The code and network configurations are available at \url{https: //github. com/bl0/PIC}.

NeurIPS Conference 2020 Conference Paper

RepPoints v2: Verification Meets Regression for Object Detection

  • Yihong Chen
  • Zheng Zhang
  • Yue Cao
  • Liwei Wang
  • Stephen Lin
  • Han Hu

Verification and regression are two general methodologies for prediction in neural networks. Each has its own strengths: verification can be easier to infer accurately, and regression is more efficient and applicable to continuous target variables. Hence, it is often beneficial to carefully combine them to take advantage of their benefits. In this paper, we take this philosophy to improve state-of-the-art object detection, specifically by RepPoints. Though RepPoints provides high performance, we find that its heavy reliance on regression for object localization leaves room for improvement. We introduce verification tasks into the localization prediction of RepPoints, producing RepPoints v2, which proves consistent improvements of about 2. 0 mAP over the original RepPoints on COCO object detection benchmark using different backbones and training methods. RepPoints v2 also achieves 52. 1 mAP on the COCO \texttt{test-dev} by a single model. Moreover, we show that the proposed approach can more generally elevate other object detection frameworks as well as applications such as instance segmentation.

ICLR Conference 2020 Conference Paper

VL-BERT: Pre-training of Generic Visual-Linguistic Representations

  • Weijie Su 0002
  • Xizhou Zhu
  • Yue Cao
  • Bin Li 0025
  • Lewei Lu
  • Furu Wei
  • Jifeng Dai

We introduce a new pre-trainable generic representation for visual-linguistic tasks, called Visual-Linguistic BERT (VL-BERT for short). VL-BERT adopts the simple yet powerful Transformer model as the backbone, and extends it to take both visual and linguistic embedded features as input. In it, each element of the input is either of a word from the input sentence, or a region-of-interest (RoI) from the input image. It is designed to fit for most of the visual-linguistic downstream tasks. To better exploit the generic representation, we pre-train VL-BERT on the massive-scale Conceptual Captions dataset, together with text-only corpus. Extensive empirical analysis demonstrates that the pre-training procedure can better align the visual-linguistic clues and benefit the downstream tasks, such as visual commonsense reasoning, visual question answering and referring expression comprehension. It is worth noting that VL-BERT achieved the first place of single model on the leaderboard of the VCR benchmark.

NeurIPS Conference 2019 Conference Paper

Learning to Optimize in Swarms

  • Yue Cao
  • Tianlong Chen
  • Zhangyang Wang
  • Yang Shen

Learning to optimize has emerged as a powerful framework for various optimization and machine learning tasks. Current such "meta-optimizers" often learn in the space of continuous optimization algorithms that are point-based and uncertainty-unaware. To overcome the limitations, we propose a meta-optimizer that learns in the algorithmic space of both point-based and population-based optimization algorithms. The meta-optimizer targets at a meta-loss function consisting of both cumulative regret and entropy. Specifically, we learn and interpret the update formula through a population of LSTMs embedded with sample- and feature-level attentions. Meanwhile, we estimate the posterior directly over the global optimum and use an uncertainty measure to help guide the learning process. Empirical results over non-convex test functions and the protein-docking application demonstrate that this new meta-optimizer outperforms existing competitors. The codes are publicly available at: https: //github. com/Shen-Lab/LOIS

AAAI Conference 2018 Conference Paper

Unsupervised Domain Adaptation With Distribution Matching Machines

  • Yue Cao
  • Mingsheng Long
  • Jianmin Wang

Domain adaptation generalizes a learning model across source domain and target domain that follow different distributions. Most existing work follows a two-step procedure: first, explores either feature matching or instance reweighting independently, and second, train the transfer classifier separately. In this paper, we show that either feature matching or instance reweighting can only reduce, but not remove, the cross-domain discrepancy, and the knowledge hidden in the relations between the data labels from the source and target domains is important for unsupervised domain adaptation. We propose a new Distribution Matching Machine (DMM) based on the structural risk minimization principle, which learns a transfer support vector machine by extracting invariant feature representations and estimating unbiased instance weights that jointly minimize the cross-domain distribution discrepancy. This leads to a robust transfer learner that performs well against both mismatched features and irrelevant instances. Our theoretical analysis proves that the proposed approach further reduces the generalization error bound of related domain adaptation methods. Comprehensive experiments validate that the DMM approach significantly outperforms competitive methods on standard domain adaptation benchmarks.

AAAI Conference 2017 Conference Paper

Collective Deep Quantization for Efficient Cross-Modal Retrieval

  • Yue Cao
  • Mingsheng Long
  • Jianmin Wang
  • Shichen Liu

Cross-modal similarity retrieval is a problem about designing a retrieval system that supports querying across content modalities, e. g. , using an image to retrieve for texts. This paper presents a compact coding solution for efficient cross-modal retrieval, with a focus on the quantization approach which has already shown the superior performance over the hashing solutions in single-modal similarity retrieval. We propose a collective deep quantization (CDQ) approach, which is the first attempt to introduce quantization in end-to-end deep architecture for cross-modal retrieval. The major contribution lies in jointly learning deep representations and the quantizers for both modalities using carefully-crafted hybrid networks and well-specified loss functions. In addition, our approach simultaneously learns the common quantizer codebook for both modalities through which the crossmodal correlation can be substantially enhanced. CDQ enables efficient and effective cross-modal retrieval using inner product distance computed based on the common codebook with fast distance table lookup. Extensive experiments show that CDQ yields state of the art cross-modal retrieval results on standard benchmarks.

AAAI Conference 2016 Conference Paper

Deep Hashing Network for Efficient Similarity Retrieval

  • Han Zhu
  • Mingsheng Long
  • Jianmin Wang
  • Yue Cao

Due to the storage and retrieval efficiency, hashing has been widely deployed to approximate nearest neighbor search for large-scale multimedia retrieval. Supervised hashing, which improves the quality of hash coding by exploiting the semantic similarity on data pairs, has received increasing attention recently. For most existing supervised hashing methods for image retrieval, an image is first represented as a vector of hand-crafted or machine-learned features, followed by another separate quantization step that generates binary codes. However, suboptimal hash coding may be produced, because the quantization error is not statistically minimized and the feature representation is not optimally compatible with the binary coding. In this paper, we propose a novel Deep Hashing Network (DHN) architecture for supervised hashing, in which we jointly learn good image representation tailored to hash coding and formally control the quantization error. The DHN model constitutes four key components: (1) a subnetwork with multiple convolution-pooling layers to capture image representations; (2) a fully-connected hashing layer to generate compact binary hash codes; (3) a pairwise crossentropy loss layer for similarity-preserving learning; and (4) a pairwise quantization loss for controlling hashing quality. Extensive experiments on standard image retrieval datasets show the proposed DHN model yields substantial boosts over latest state-of-the-art hashing methods.

AAAI Conference 2016 Conference Paper

Deep Quantization Network for Efficient Image Retrieval

  • Yue Cao
  • Mingsheng Long
  • Jianmin Wang
  • Han Zhu
  • Qingfu Wen

Hashing has been widely applied to approximate nearest neighbor search for large-scale multimedia retrieval. Supervised hashing improves the quality of hash coding by exploiting the semantic similarity on data pairs and has received increasing attention recently. For most existing supervised hashing methods for image retrieval, an image is first represented as a vector of hand-crafted or machine-learned features, then quantized by a separate quantization step that generates binary codes. However, suboptimal hash coding may be produced, since the quantization error is not statistically minimized and the feature representation is not optimally compatible with the hash coding. In this paper, we propose a novel Deep Quantization Network (DQN) architecture for supervised hashing, which learns image representation for hash coding and formally control the quantization error. The DQN model constitutes four key components: (1) a sub-network with multiple convolution-pooling layers to capture deep image representations; (2) a fully connected bottleneck layer to generate dimension-reduced representation optimal for hash coding; (3) a pairwise cosine loss layer for similarity-preserving learning; and (4) a product quantization loss for controlling hashing quality and the quantizability of bottleneck representation. Extensive experiments on standard image retrieval datasets show the proposed DQN model yields substantial boosts over latest state-of-the-art hashing methods.

YNIMG Journal 2006 Journal Article

Reduction in V1 activation associated with decreased visibility of a visual target

  • Jie Huang
  • Ming Xiang
  • Yue Cao

The perception of a brief visual target stimulus can be affected by another visual mask stimulus immediately preceding or following the target. The link of this visual masking illusion, with visual cortical activation, offers insights into the neural mechanisms for visual perception. The present study investigated the association of the visibility of a target with cortical activation in humans using psychophysical testing and functional magnetic resonance imaging (fMRI). A visual masking protocol that was suitable for an fMRI study was developed. The event-related fMRI was used to measure activation in primary visual cortex (V1) during visual masking and unmasking stimulation. We found that the visibility of the target stimulus was reduced in the masking condition, due to the presence of mask stimuli, but not in the unmasking condition. We also found that the activation in V1 was modulated by the temporal separation of the mask stimuli from the target and was associated with the visibility of the target that was recorded during psychophysical testing and fMRI. These findings are consistent with what has been observed in the primate visual cortex of monkeys, i. e. , the transient on-response and after-discharge of V1 neurons to the target stimulus were suppressed by forward and backward mask stimuli, respectively.

IJCAI Conference 1999 Conference Paper

SHOP: Simple Hierarchical Ordered Planner

  • Dana Nan
  • Yue Cao
  • Amnon Lotem
  • Hector Munoz-Avila___

SHOP (Simple Hierarchical Ordered Planner) is a domain-independent HTN planning system with the following characteristics. • SHOP plans for tasks in the same order that they will later be executed. This avoids some goalinteraction issues that arise in other HTN planners, so that the planning algorithm is relatively simple. • Since SHOP knows the complete world-state at each step of the planning process, it can use highly expressive domain representations. For example, it can do planning problems that require complex numeric computations. • In our tests, SHOP was several orders of magnitude faster man Blackbox and several times faster than TLpian, even though SHOP is coded in Lisp and the other planners are coded in C.

v2026.09.13