Arrow Research search

Author name cluster

Jian Sun

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

49 papers
2 author rows

Possible papers

49

EAAI Journal 2026 Journal Article

Graph structure consistency and pseudo-label guided source-free domain adaptation for lithology identification

  • Jian Sun
  • Xin Sha
  • Rongjun Zhang
  • Long Ren
  • Zhe Zhang

Existing deep learning-based lithology identification methods encounter several critical bottlenecks: they fail to achieve robust model transfer in source-free scenarios, underutilize the intrinsic topological information of data, and are susceptible to noisy pseudo-labels. To address these challenges, a novel framework termed graph structure consistency and pseudo-label guided source-free domain adaptation (GCPG-SFDA) for lithology identification is proposed. This framework integrates graph-based modeling with the mean teacher framework to build a multi-dimensional system for feature optimization and knowledge transfer. Specifically, tabular lithological data is first transformed into graph structures and processed via graph neural network, enabling the extraction of high-order semantic relationships while preserving original features to enhance target-domain representation. Furthermore, three relationship graphs (teacher graph, student graph, and teacher-student graph) are designed. Graph consistency constraints optimize sample similarities in the feature space, improving the clarity of target domain classification boundaries. Meanwhile, a self-supervised exploration mechanism is implemented to encourage robust feature learning through structural perturbations and teacher-student output synchronization. Comprehensive evaluations indicate that GCPG-SFDA achieves superior performance in lithology identification, offering a robust solution for data-constrained geological tasks.

AAAI Conference 2026 Conference Paper

Transferring Causal Driving Patterns for Generalizable Traffic Simulation with Diffusion-Based Distillation

  • Yuhang Chen
  • Jie Sun
  • Jialin Fan
  • Jian Sun

Traffic simulation is essential for validating the safety and reliability of autonomous driving systems, yet data-driven simulation methods often struggle with distribution shifts, limiting their generalizability across diverse datasets (domains). To address this, we present Causal Driving Pattern Transfer (CDPT), a novel two-stage knowledge distillation framework built upon diffusion model to enhance cross-domain generalizability. In Phase I, we implement hybrid self-distillation within the source domain by integrating feature-, response-, and contrastive-level distillation, which enables the model to decompose complex driving behaviors into their core causal components, including scene-conditioned driven patterns, multi-agent interaction dynamics and casual saliency. In Phase II, we introduce a continual distillation strategy: few-shot samples from the target domain are used to initiate generation of diverse synthetic scenarios, allowing the student model to continually adapt to novel environments without retraining on large-scale data. Extensive experiments demonstrate that CDPT achieves strong generalization in both open-loop and closed-loop simulations, effectively generating realistic, interaction-aware behaviors that are critical for scalable and reliable autonomous driving testing.

AIIM Journal 2025 Journal Article

A prior knowledge-supervised fusion network predicts survival after radiotherapy in patients with advanced gastric cancer

  • Liang Sun
  • Yongxin Lan
  • Jian Sun
  • Pengfei Ji
  • Hongwei Ge
  • Ming Cui
  • Xin Yuan

BACKGROUND AND OBJECTIVE: Predicting overall survival (OS) for advanced gastric cancer patients after radiotherapy is critical for developing an individualized treatment plan. However, existing studies have focused on gastric cancer CT images with a large amount of redundant information, neglecting the role of physicians' prior knowledge in guiding gastric cancer CT image information. We propose a multimodal fusion method based on prior knowledge to predict OS after radiotherapy in advanced gastric cancer patients to assist physicians in clinical diagnosis and treatment. METHODS: A prior knowledge supervised fusion network (PKSFnet) is proposed. Firstly, PKSFnet uses a novel sampling strategy, which enables the input model data to obtain a complete feature space by analyzing the entire patient data space. Afterwards, under the guidance of the multi-domain feature fusion module (MdFF), multimodal information of patients is adaptively fused and mined to improve the prediction performance. RESULTS: The results of the proposed model are superior to those of other unimodal and multimodal state-of-the-art methods. For the segmented survival time classification task, the AUC, specificity, sensitivity, precision of the proposed model are 0.8397, 0.875, 0.7556, and 0.875, respectively. For the survival risk regression task, the C-index and HR of the proposed model are 0.8574 and 4.658 respectively. Ablation experimental results further demonstrate the impact of each module of the proposed model. Finally, we apply the novel sampling strategy to other deep learning models and achieve significant improvement. CONCLUSION: The experimental results have demonstrated that the proposed model can effectively predict OS after radiotherapy in patients with advanced gastric cancer, which demonstrate that the proposed model can facilitate the development and application of robust clinical treatment strategies.

JBHI Journal 2025 Journal Article

A Ubiquitous Platform for Camera-Based Multi-Parameter Vital Signs Monitoring in Hospital ICUs: A Double-Center Clinical Study

  • Huailei Lai
  • Jia Huang
  • Xin Tian
  • Dongfang Yu
  • Yonglong Ye
  • Yongshen Zeng
  • Yukai Huang
  • Guowei Wang

Conventional physiological monitoring relies on multiple contact-based biomedical sensors, such as electrocardiograms, pulse oximeters, and blood pressure (BP) cuffs, which often necessitate the use of cumbersome sensors (e. g. , electrodes, patches, diodes) and extensive wiring. These conventional methods not only impose significant inconvenience on both caregivers and patients but also elevate the risk of patient infections. In this study, we introduce a novel non-contact, multi-parameter vital signs monitoring platform based on the red, green and near Infrared (RG-IR) spectral imaging system, which measures five critical patient parameters in the Intensive Care Unit (ICU), including heart rate (HR), heart rate variability (HRV), respiratory rate (RR), blood oxygen saturation (SpO $_{2}$ ) and BP. Clinical trials were conducted in the ICUs of two hospitals, involving 30 critically ill patients. The clinical outcomes indicate that our system achieves reasonable performance, with mean absolute errors of 1. 89 bpm for HR, 19. 04 ms for SDNN (standard deviation of normal-to-normal intervals), 0. 99 bpm for RR, 2. 75% for SpO $_{2}$ and 5. 67–8. 98 mmHg for BP, providing timely physiological information and marking a significant step toward contactless ICU monitoring. Additionally, the real-time performance of the core algorithms were validated on low-cost embedded chips, demonstrating the system's capability for edge computation and on-board patient monitoring. The proposed system has entered the registration process of the National Medical Products Administration (NMPA) as a Class II medical device (No. 2024007).

NeurIPS Conference 2025 Conference Paper

Coarse-to-Fine 3D Part Assembly via Semantic Super-Parts and Symmetry-Aware Pose Estimation

  • Xinyi Zhang
  • Bingyang Wei
  • Ruixuan Yu
  • Jian Sun

We propose a novel two-stage framework, Coarse-to-Fine Part Assembly (CFPA), for 3D shape assembly from basic parts. Effective part assembly demands precise local geometric reasoning for accurate component assembly, as well as global structural understanding to ensure semantic coherence and plausible configurations. CFPA addresses this challenge by integrating semantic abstraction and symmetry-aware reasoning into a unified pose prediction process. In the first stage, semantic super-parts are constructed via an optimal transport formulation to capture high-level object structure, which is then propagated to individual parts through a dual-range feature propagation mechanism. The second stage refines part poses via cross-stage feature interaction and instance-level geometric encoding, improving spatial precision and coherence. To enable diverse yet valid assemblies, we introduce a symmetry-aware loss that jointly models both self-symmetry and inter-part geometric similarity, allowing for diverse but structurally consistent assemblies. Extensive experiments on the PartNet benchmark demonstrate that CFPA achieves state-of-the-art performance in assembly accuracy, structural consistency, and diversity across multiple categories.

NeurIPS Conference 2025 Conference Paper

DyMoDreamer: World Modeling with Dynamic Modulation

  • Boxuan Zhang
  • Runqing Wang
  • Wei Xiao
  • Weipu Zhang
  • Jian Sun
  • Gao Huang
  • Jie Chen
  • Gang Wang

A critical bottleneck in deep reinforcement learning (DRL) is sample inefficiency, as training high-performance agents often demands extensive environmental interactions. Model-based reinforcement learning (MBRL) mitigates this by building world models that simulate environmental dynamics and generate synthetic experience, improving sample efficiency. However, conventional world models process observations holistically, failing to decouple dynamic objects and temporal features from static backgrounds. This approach is computationally inefficient, especially for visual tasks where dynamic objects significantly influence rewards and decision-making performance. To address this, we introduce DyMoDreamer, a novel MBRL algorithm that incorporates a dynamic modulation mechanism to improve the extraction of dynamic features and enrich the temporal information. DyMoDreamer employs differential observations derived from a novel inter-frame differencing mask, explicitly encoding object-level motion cues and temporal dynamics. Dynamic modulation is modeled as stochastic categorical distributions and integrated into a recurrent state-space model (RSSM), enhancing the model's focus on reward-relevant dynamics. Experiments demonstrate that DyMoDreamer sets a new state-of-the-art on the Atari $100$k benchmark with a $156. 6$\% mean human-normalized score, establishes a new record of $832$ on the DeepMind Visual Control Suite, and gains a $9. 5$\% performance improvement after $1$M steps on the Crafter benchmark.

NeurIPS Conference 2025 Conference Paper

Joint Velocity-Growth Flow Matching for Single-Cell Dynamics Modeling

  • Dongyi Wang
  • Yuanwei Jiang
  • Zhenyi Zhang
  • Xiang Gu
  • Peijie Zhou
  • Jian Sun

Learning the underlying dynamics of single cells from snapshot data has gained increasing attention in scientific and machine learning research. The destructive measurement technique and cell proliferation/death result in unpaired and unbalanced data between snapshots, making the learning of the underlying dynamics challenging. In this paper, we propose joint Velocity-Growth Flow Matching (VGFM), a novel paradigm that jointly learns state transition and mass growth of single-cell populations via flow matching. VGFM builds an ideal single-cell dynamics containing velocity of state and growth of mass, driven by a presented two-period dynamic understanding of the static semi-relaxed optimal transport, a mathematical tool that seeks the coupling between unpaired and unbalanced data. To enable practical usage, we approximate the ideal dynamics using neural networks, forming our joint velocity and growth matching framework. A distribution fitting loss is also employed in VGFM to further improve the fitting performance for snapshot data. Extensive experimental results on both synthetic and real datasets demonstrate that VGFM can capture the underlying biological dynamics accounting for mass and state variations over time, outperforming existing approaches for single-cell dynamics modeling.

JBHI Journal 2025 Journal Article

Re-Visible Dual-Domain Self-Supervised Deep Unfolding Network for MRI Reconstruction

  • Hao Zhang
  • Qi Wang
  • Jian Sun
  • Zhijie Wen
  • Jun Shi
  • Shihui Ying

Magnetic Resonance Imaging (MRI) is widely used in clinical practice, but suffers from prolonged acquisition time. Although deep learning methods have been proposed to accelerate acquisition and demonstrate promising performance, they rely on high-quality fully-sampled datasets for training in a supervised manner. However, such datasets are time-consuming and expensive-to-collect, which constrains their broader applications. On the other hand, self-supervised methods offer an alternative by enabling learning from under-sampled data alone, but most existing methods rely on further partitioned under-sampled k-space data as model's input for training, which causes an input distribution shift between the the training stage and the inference stage. Additionally, their models have not effectively incorporated comprehensive image priors, leading to degraded reconstruction performance. In this paper, we propose a novel re-visible dual-domain self-supervised deep unfolding network to address these issues when only under-sampled datasets are available. Specifically, by incorporating re-visible dual-domain loss, all under-sampled k-space data are utilized during training to mitigate the input distribution shift caused by further partitioning. This design enables the model to implicitly adapt to all under-sampled k-space data as input. Additionally, we design a Deep Unfolding Network based on Chambolle and Pock Proximal Point Algorithm (DUN-CP-PPA) to achieve end-to-end reconstruction. By employing a Spatial-Frequency Feature Extraction (SFFE) block to capture both global and local representations, the model effectively integrates imaging physics with comprehensive image priors to enhance reconstruction performance. Experiments on both single-coil and multi-coil datasets demonstrate that our method outperforms state-of-the-art approaches in terms of reconstruction performance and generalization capability.

NeurIPS Conference 2025 Conference Paper

SpiderSolver: A Geometry-Aware Transformer for Solving PDEs on Complex Geometries

  • KAI QI
  • Fan Wang
  • Zhewen Dong
  • Jian Sun

Transformers have demonstrated effectiveness in solving partial differential equations (PDEs). However, extending them to solve PDEs on complex geometries remains a challenge. In this work, we propose SpiderSolver, a geometry-aware transformer that introduces spiderweb tokenization for handling complex domain geometry and irregularly discretized points. Our method partitions the irregular spatial domain into spiderweb-like patches, guided by the domain boundary geometry. SpiderSolver leverages a coarse-grained attention mechanism to capture global interactions across spiderweb tokens and a fine-grained attention mechanism to refine feature interactions between the domain boundary and its neighboring interior points. We evaluate SpiderSolver on PDEs with diverse domain geometries across seven datasets, including cars, airfoils, blood flow in the human thoracic aorta, as well as canonical cases governed by the Navier-Stokes, Darcy flow, elasticity, and plasticity equations. Experimental results demonstrate that SpiderSolver consistently achieves state-of-the-art performance across different datasets and metrics, with better generalization ability in the OOD setting. The code is available at https: //github. com/Kai-Qi/SpiderSolver.

NeurIPS Conference 2025 Conference Paper

Towards Prospective Medical Image Reconstruction via Knowledge-Informed Dynamic Optimal Transport

  • Taoran Zheng
  • Yan Yang
  • Xing Li
  • Xiang Gu
  • Jian Sun
  • Zongben Xu

Medical image reconstruction from measurement data is a vital but challenging inverse problem. Deep learning approaches have achieved promising results, but often requires paired measurement and high-quality images, which is typically simulated through a forward model, i. e. , retrospective reconstruction. However, training on simulated pairs commonly leads to performance degradation on real prospective data due to the retrospective-to-prospective gap caused by incomplete imaging knowledge in simulation. To address this challenge, this paper introduces imaging Knowledge-Informed Dynamic Optimal Transport (KIDOT), a novel dynamic optimal transport framework with optimality in the sense of preserving consistency with imaging physics in transport, that conceptualizes reconstruction as finding a dynamic transport path. KIDOT learns from unpaired data by modeling reconstruction as a continuous evolution path from measurements to images, guided by an imaging knowledge-informed cost function and transport equation. This dynamic and knowledge-aware approach enhances robustness and better leverages unpaired data while respecting acquisition physics. Theoretically, we demonstrate that KIDOT naturally generalizes dynamic optimal transport, ensuring its mathematical rationale and solution existence. Extensive experiments on MRI and CT reconstruction demonstrate KIDOT's superior performance. Code is available at https: //github. com/TaoranZheng717/KIDOT.

EAAI Journal 2025 Journal Article

Visual-tactile fusion learning for material recognition based on channel switching and dual cross-attention

  • Song Li
  • Wei Sun
  • Qiaokang Liang
  • Jian Sun
  • Hui Yang
  • YuDong Yang

Visual-tactile multimodal object recognition has attracted increasing attention, as information from different modalities can complement each other and enhance recognition performance. However, the inherent heterogeneity between vision and touch presents a key challenge for effective fusion, limiting accuracy and robustness. To address this, we propose a Multimodal Channel-Switching and Dual Cross-Attention Fusion (MCSDCF) method for integrating multisource data. The channel-switching module adaptively weights and exchanges information across modalities, enabling more discriminative and complementary feature representations. To further capture high-level semantic correlations and heterogeneous cues, we introduce a dual crossattention fusion structure that combines intra-modal self-attention with cross-modal mutual attention, reinforcing the quality of fused representations. Extensive experiments on three public benchmark datasets demonstrate the effectiveness of our MCSDCF framework, with multimodal fusion consistently outperforming unimodal baselines. Ablation studies further confirm the individual contributions of the proposed modules to the overall recognition performance.

EAAI Journal 2024 Journal Article

A hybrid data- and model-driven learning framework for remaining useful life prognostics

  • Hongjie Cao
  • Wei Xiao
  • Jian Sun
  • Ming-Gang Gan
  • Gang Wang

The efficient and safe production of machinery equipment relies on the health of its mechanical components, making prognostics and health management (PHM) a critical aspect of production processes. One key PHM measure is the remaining useful life (RUL), which estimates the expected lifespan of a component in a production line before requiring repair or replacement. However, state-of-the-art RUL prediction methods, including data-driven, model-based, and hybrid approaches, face limitations such as incomplete/imprecise physical models, uncertainties in degradation processes, and measurement data noise. To address these limitations, this paper proposes a novel hybrid RUL prediction framework that combines the strengths of data-based and model-driven approaches. The framework includes an exponential model to leverage physical knowledge and a multi-head attention transformer to extract information from data. An extended Kalman filter is used to estimate unknown degradation process parameters and provide physical model information for the prediction process. A regression token is introduced to efficiently fuse the deep learning model and the stochastic filtering method. Numerical tests using real-world rolling bearing degradation datasets demonstrate the superiority of the proposed method over competitive alternatives.

TMLR Journal 2024 Journal Article

A Survey on Out-of-Distribution Detection in NLP

  • Hao Lang
  • Yinhe Zheng
  • Yixuan Li
  • Jian Sun
  • Fei Huang
  • Yongbin Li

Out-of-distribution (OOD) detection is essential for the reliable and safe deployment of machine learning systems in the real world. Great progress has been made over the past years. This paper presents the first review of recent advances in OOD detection with a particular focus on natural language processing approaches. First, we provide a formal definition of OOD detection and discuss several related fields. We then categorize recent algorithms into three classes according to the data they used: (1) OOD data available, (2) OOD data unavailable + in-distribution (ID) label available, and (3) OOD data unavailable + ID label unavailable. Third, we introduce datasets, applications, and metrics. Finally, we summarize existing work and present potential future research topics.

NeurIPS Conference 2024 Conference Paper

Learning 3D Equivariant Implicit Function with Patch-Level Pose-Invariant Representation

  • Xin Hu
  • Xiaole Tang
  • Ruixuan Yu
  • Jian Sun

Implicit neural representation gains popularity in modeling the continuous 3D surface for 3D representation and reconstruction. In this work, we are motivated by the fact that the local 3D patches repeatedly appear on 3D shapes/surfaces if the factor of poses is removed. Based on this observation, we propose the 3D patch-level equivariant implicit function (PEIF) based on the 3D patch-level pose-invariant representation, allowing us to reconstruct 3D surfaces by estimating equivariant displacement vector fields for query points. Specifically, our model is based on the pose-normalized query/patch pairs and enhanced by the proposed intrinsic patch geometry representation, modeling the intrinsic 3D patch geometry feature by learnable multi-head memory banks. Extensive experiments show that our model achieves state-of-the-art performance on multiple surface reconstruction datasets, and also exhibits better generalization to crossdataset shapes and robustness to arbitrary rotations. Our code will be available at https: //github. com/mathXin112/PEIF. git.

TCS Journal 2023 Journal Article

An LP-based approximation algorithm for the generalized traveling salesman path problem

  • Jian Sun
  • Gregory Gutin
  • Ping Li
  • Peihao Shi
  • Xiaoyan Zhang

The traveling salesman problem (TSP) is one of the classic research topics in the field of operations research, graph theory and computer science. In this paper, we propose a generalized model of traveling salesman problem, denoted by generalized traveling salesman path problem. Let G = ( V, E, c ) be a weighted complete graph, in which c is a nonnegative metric cost function on edge set E, i. e. , c: E → R +. The traveling salesman path problem aims to find a Hamiltonian path in G with minimum cost. Compared to the traveling salesman path problem, we are given extra vertex subset V ′ and edge subset E ′ in the problem proposed in this paper; its goal is to construct a path which traverses all the edges in E ′ while only needs to visit each vertex in V ′ exactly once. Based on integer programming, we give a mathematical model of the problem, and design a 1 + 5 2 -approximation algorithm for the problem by combining linear programming rounding strategy and a special graph structure.

NeurIPS Conference 2023 Conference Paper

Constructing Non-isotropic Gaussian Diffusion Model Using Isotropic Gaussian Diffusion Model for Image Editing

  • Xi Yu
  • Xiang Gu
  • Haozhi Liu
  • Jian Sun

Score-based diffusion models (SBDMs) have achieved state-of-the-art results in image generation. In this paper, we propose a Non-isotropic Gaussian Diffusion Model (NGDM) for image editing, which requires editing the source image while preserving the image regions irrelevant to the editing task. We construct NGDM by adding independent Gaussian noises with different variances to different image pixels. Instead of specifically training the NGDM, we rectify the NGDM into an isotropic Gaussian diffusion model with different pixels having different total forward diffusion time. We propose to reverse the diffusion by designing a sampling method that starts at different time for different pixels for denoising to generate images using the pre-trained isotropic Gaussian diffusion model. Experimental results show that NGDM achieves state-of-the-art performance for image editing tasks, considering the trade-off between the fidelity to the source image and alignment with the desired editing target.

AAAI Conference 2023 Conference Paper

Generalized Semantic Segmentation by Self-Supervised Source Domain Projection and Multi-Level Contrastive Learning

  • Liwei Yang
  • Xiang Gu
  • Jian Sun

Deep networks trained on the source domain show degraded performance when tested on unseen target domain data. To enhance the model's generalization ability, most existing domain generalization methods learn domain invariant features by suppressing domain sensitive features. Different from them, we propose a Domain Projection and Contrastive Learning (DPCL) approach for generalized semantic segmentation, which includes two modules: Self-supervised Source Domain Projection (SSDP) and Multi-Level Contrastive Learning (MLCL). SSDP aims to reduce domain gap by projecting data to the source domain, while MLCL is a learning scheme to learn discriminative and generalizable features on the projected data. During test time, we first project the target data by SSDP to mitigate domain shift, then generate the segmentation results by the learned segmentation network based on MLCL. At test time, we can update the projected data by minimizing our proposed pixel-to-pixel contrastive loss to obtain better results. Extensive experiments for semantic segmentation demonstrate the favorable generalization capability of our method on benchmark datasets.

NeurIPS Conference 2023 Conference Paper

Optimal Transport-Guided Conditional Score-Based Diffusion Model

  • Xiang Gu
  • Liwei Yang
  • Jian Sun
  • Zongben Xu

Conditional score-based diffusion model (SBDM) is for conditional generation of target data with paired data as condition, and has achieved great success in image translation. However, it requires the paired data as condition, and there would be insufficient paired data provided in real-world applications. To tackle the applications with partially paired or even unpaired dataset, we propose a novel Optimal Transport-guided Conditional Score-based diffusion model (OTCS) in this paper. We build the coupling relationship for the unpaired or partially paired dataset based on $L_2$-regularized unsupervised or semi-supervised optimal transport, respectively. Based on the coupling relationship, we develop the objective for training the conditional score-based model for unpaired or partially paired settings, which is based on a reformulation and generalization of the conditional SBDM for paired setting. With the estimated coupling relationship, we effectively train the conditional score-based model by designing a ``resampling-by-compatibility'' strategy to choose the sampled data with high compatibility as guidance. Extensive experiments on unpaired super-resolution and semi-paired image-to-image translation demonstrated the effectiveness of the proposed OTCS model. From the viewpoint of optimal transport, OTCS provides an approach to transport data across distributions, which is a challenge for OT on large-scale datasets. We theoretically prove that OTCS realizes the data transport in OT with a theoretical bound.

AAAI Conference 2023 Conference Paper

Prototypical Partial Optimal Transport for Universal Domain Adaptation

  • Yucheng Yang
  • Xiang Gu
  • Jian Sun

Universal domain adaptation (UniDA) aims to transfer knowledge from a labeled source domain to an unlabeled target domain without requiring the same label sets of both domains. The existence of domain and category shift makes the task challenging and requires us to distinguish “known” samples (i.e., samples whose labels exist in both domains) and “unknown” samples (i.e., samples whose labels exist in only one domain) in both domains before reducing the domain gap. In this paper, we consider the problem from the point of view of distribution matching which we only need to align two distributions partially. A novel approach, dubbed mini-batch Prototypical Partial Optimal Transport (m-PPOT), is proposed to conduct partial distribution alignment for UniDA. In training phase, besides minimizing m-PPOT, we also leverage the transport plan of m-PPOT to reweight source prototypes and target samples, and design reweighted entropy loss and reweighted cross-entropy loss to distinguish “known” and “unknown” samples. Experiments on four benchmarks show that our method outperforms the previous state-of-the-art UniDA methods.

NeurIPS Conference 2023 Conference Paper

STORM: Efficient Stochastic Transformer based World Models for Reinforcement Learning

  • Weipu Zhang
  • Gang Wang
  • Jian Sun
  • Yetian Yuan
  • Gao Huang

Recently, model-based reinforcement learning algorithms have demonstrated remarkable efficacy in visual input environments. These approaches begin by constructing a parameterized simulation world model of the real environment through self-supervised learning. By leveraging the imagination of the world model, the agent's policy is enhanced without the constraints of sampling from the real environment. The performance of these algorithms heavily relies on the sequence modeling and generation capabilities of the world model. However, constructing a perfectly accurate model of a complex unknown environment is nearly impossible. Discrepancies between the model and reality may cause the agent to pursue virtual goals, resulting in subpar performance in the real environment. Introducing random noise into model-based reinforcement learning has been proven beneficial. In this work, we introduce Stochastic Transformer-based wORld Model (STORM), an efficient world model architecture that combines the strong sequence modeling and generation capabilities of Transformers with the stochastic nature of variational autoencoders. STORM achieves a mean human performance of $126. 7\%$ on the Atari $100$k benchmark, setting a new record among state-of-the-art methods that do not employ lookahead search techniques. Moreover, training an agent with $1. 85$ hours of real-time interaction experience on a single NVIDIA GeForce RTX 3090 graphics card requires only $4. 3$ hours, showcasing improved efficiency compared to previous methodologies.

ICLR Conference 2023 Conference Paper

VoGE: A Differentiable Volume Renderer using Gaussian Ellipsoids for Analysis-by-Synthesis

  • Angtian Wang
  • Peng Wang 0001
  • Jian Sun
  • Adam Kortylewski
  • Alan L. Yuille

Differentiable rendering allows the application of computer graphics on vision tasks, e.g. object pose and shape fitting, via analysis-by-synthesis, where gradients at occluded regions are important when inverting the rendering process.To obtain those gradients, state-of-the-art (SoTA) differentiable renderers use rasterization to collect a set of nearest components for each pixel and aggregate them based on the viewing distance. In this paper, we propose VoGE, which uses ray tracing to capture nearest components with their volume density distributions on the rays and aggregates via integral of the volume densities based on Gaussian ellipsoids, which brings more efficient and stable gradients. To efficiently render via VoGE, we propose an approximate close-form solution for the volume density aggregation and a coarse-to-fine rendering strategy. Finally, we provide a CUDA implementation of VoGE, which gives a competitive rendering speed in comparison to PyTorch3D. Quantitative and qualitative experiment results show VoGE outperforms SoTA counterparts when applied to various vision tasks, e.g., object pose estimation, shape/texture fitting, and occlusion reasoning. The VoGE code is available at: https://github.com/Angtian/VoGE.

IJCAI Conference 2022 Conference Paper

A Survey on Neural Open Information Extraction: Current Status and Future Directions

  • Shaowen Zhou
  • Bowen Yu
  • Aixin Sun
  • Cheng Long
  • Jingyang Li
  • Jian Sun

Open Information Extraction (OpenIE) facilitates domain-independent discovery of relational facts from large corpora. The technique well suits many open-world natural language understanding scenarios, such as automatic knowledge base construction, open-domain question answering, and explicit reasoning. Thanks to the rapid development in deep learning technologies, numerous neural OpenIE architectures have been proposed and achieve considerable performance improvement. In this survey, we provide an extensive overview of the state-of-the-art neural OpenIE models, their key design decisions, strengths and weakness. Then, we discuss limitations of current solutions and the open issues in OpenIE problem itself. Finally we list recent trends that could help expand its scope and applicability, setting up promising directions for future research in OpenIE. To our best knowledge, this paper is the first review on neural OpenIE.

AAAI Conference 2022 Conference Paper

Anchor DETR: Query Design for Transformer-Based Detector

  • Yingming Wang
  • Xiangyu Zhang
  • Tong Yang
  • Jian Sun

In this paper, we propose a novel query design for the transformer-based object detection. In previous transformerbased detectors, the object queries are a set of learned embeddings. However, each learned embedding does not have an explicit physical meaning and we cannot explain where it will focus on. It is difficult to optimize as the prediction slot of each object query does not have a specific mode. In other words, each object query will not focus on a specific region. To solve these problems, in our query design, object queries are based on anchor points, which are widely used in CNN-based detectors. So each object query focuses on the objects near the anchor point. Moreover, our query design can predict multiple objects at one position to solve the difficulty: “one region, multiple objects”. In addition, we design an attention variant, which can reduce the memory cost while achieving similar or better performance than the standard attention in DETR. Thanks to the query design and the attention variant, the proposed detector that we called Anchor DETR, can achieve better performance and run faster than the DETR with 10× fewer training epochs. For example, it achieves 44. 2 AP with 19 FPS on the MSCOCO dataset when using the ResNet50-DC5 feature for training 50 epochs. Extensive experiments on the MSCOCO benchmark prove the effectiveness of the proposed methods. Code is available at https: //github. com/megvii-research/AnchorDETR.

AAAI Conference 2022 Conference Paper

GALAXY: A Generative Pre-trained Model for Task-Oriented Dialog with Semi-supervised Learning and Explicit Policy Injection

  • Wanwei He
  • Yinpei Dai
  • Yinhe Zheng
  • Yuchuan Wu
  • Zheng Cao
  • Dermot Liu
  • Peng Jiang
  • Min Yang

Pre-trained models have proved to be powerful in enhancing task-oriented dialog systems. However, current pre-training methods mainly focus on enhancing dialog understanding and generation tasks while neglecting the exploitation of dialog policy. In this paper, we propose GALAXY, a novel pre-trained dialog model that explicitly learns dialog policy from limited labeled dialogs and large-scale unlabeled dialog corpora via semi-supervised learning. Specifically, we introduce a dialog act prediction task for policy optimization during pre-training and employ a consistency regularization term to refine the learned representation with the help of unlabeled dialogs. We also implement a gating mechanism to weigh suitable unlabeled dialog samples. Empirical results show that GALAXY substantially improves the performance of task-oriented dialog systems, and achieves new state-of-the-art results on benchmark datasets: In-Car, MultiWOZ2. 0 and Multi- WOZ2. 1, improving their end-to-end combined scores by 2. 5, 5. 3 and 5. 5 points, respectively. We also show that GALAXY has a stronger few-shot ability than existing models under various low-resource settings. For reproducibility, we release the code and data at https: //github. com/siat-nlp/GALAXY.

NeurIPS Conference 2022 Conference Paper

Keypoint-Guided Optimal Transport with Applications in Heterogeneous Domain Adaptation

  • Xiang Gu
  • Yucheng Yang
  • Wei Zeng
  • Jian Sun
  • Zongben Xu

Existing Optimal Transport (OT) methods mainly derive the optimal transport plan/matching under the criterion of transport cost/distance minimization, which may cause incorrect matching in some cases. In many applications, annotating a few matched keypoints across domains is reasonable or even effortless in annotation burden. It is valuable to investigate how to leverage the annotated keypoints to guide the correct matching in OT. In this paper, we propose a novel KeyPoint-Guided model by ReLation preservation (KPG-RL) that searches for the matching guided by the keypoints in OT. To impose the keypoints in OT, first, we propose a mask-based constraint of the transport plan that preserves the matching of keypoint pairs. Second, we propose to preserve the relation of each data point to the keypoints to guide the matching. The proposed KPG-RL model can be solved by the Sinkhorn's algorithm and is applicable even when distributions are supported in different spaces. We further utilize the relation preservation constraint in the Kantorovich Problem and Gromov-Wasserstein model to impose the guidance of keypoints in them. Meanwhile, the proposed KPG-RL model is extended to partial OT setting. As an application, we apply the proposed KPG-RL model to the heterogeneous domain adaptation. Experiments verified the effectiveness of the KPG-RL model.

NeurIPS Conference 2022 Conference Paper

Learning Generalizable Part-based Feature Representation for 3D Point Clouds

  • Xin Wei
  • Xiang Gu
  • Jian Sun

Deep networks on 3D point clouds have achieved remarkable success in 3D classification, while they are vulnerable to geometry variations caused by inconsistent data acquisition procedures. This results in a challenging 3D domain generalization (3DDG) problem, that is to generalize a model trained on source domain to an unseen target domain. Based on the observation that local geometric structures are more generalizable than the whole shape, we propose to reduce the geometry shift by a generalizable part-based feature representation and design a novel part-based domain generalization network (PDG) for 3D point cloud classification. Specifically, we build a part-template feature space shared by source and target domains. Shapes from distinct domains are first organized to part-level features and then represented by part-template features. The transformed part-level features, dubbed aligned part-based representations, are then aggregated by a part-based feature aggregation module. To improve the robustness of the part-based representations, we further propose a contrastive learning framework upon part-based shape representation. Experiments and ablation studies on 3DDA and 3DDG benchmarks justify the efficacy of the proposed approach for domain generalization, compared with the previous state-of-the-art methods. Our code will be available on http: //github. com/weixmath/PDG.

AAAI Conference 2022 Conference Paper

LGD: Label-Guided Self-Distillation for Object Detection

  • Peizhen Zhang
  • Zijian Kang
  • Tong Yang
  • Xiangyu Zhang
  • Nanning Zheng
  • Jian Sun

In this paper, we propose the first self-distillation framework for general object detection, termed LGD (Label-Guided self-Distillation). Previous studies rely on a strong pretrained teacher to provide instructive knowledge that could be unavailable in real-world scenarios. Instead, we generate an instructive knowledge based only on student representations and regular labels. Our framework includes sparse labelappearance encoder, inter-object relation adaptater and intraobject knowledge mapper that jointly form an implicit teacher at training phase, dynamically dependent on labels and evolving student representations. They are trained end-to-end with detector and discarded in inference. Experimentally, LGD obtains decent results on various detectors, datasets, and extensive tasks like instance segmentation. For example in MS- COCO dataset, LGD improves RetinaNet with ResNet-50 under 2× single-scale training from 36. 2% to 39. 0% mAP (+ 2. 8%). It boosts much stronger detectors like FCOS with ResNeXt-101 DCN v2 under 2× multi-scale training from 46. 1% to 47. 9% (+ 1. 8%). Compared with a classical teacherbased method FGFI, LGD not only performs better without requiring pretrained teacher but also reduces 51% training cost beyond inherent student learning. Codes are available at https: //github. com/megvii-research/LGD.

NeurIPS Conference 2022 Conference Paper

Unifying Voxel-based Representation with Transformer for 3D Object Detection

  • Yanwei Li
  • Yilun Chen
  • Xiaojuan Qi
  • Zeming Li
  • Jian Sun
  • Jiaya Jia

In this work, we present a unified framework for multi-modality 3D object detection, named UVTR. The proposed method aims to unify multi-modality representations in the voxel space for accurate and robust single- or cross-modality 3D detection. To this end, the modality-specific space is first designed to represent different inputs in the voxel feature space. Different from previous work, our approach preserves the voxel space without height compression to alleviate semantic ambiguity and enable spatial connections. To make full use of the inputs from different sensors, the cross-modality interaction is then proposed, including knowledge transfer and modality fusion. In this way, geometry-aware expressions in point clouds and context-rich features in images are well utilized for better performance and robustness. The transformer decoder is applied to efficiently sample features from the unified space with learnable positions, which facilitates object-level interactions. In general, UVTR presents an early attempt to represent different modalities in a unified framework. It surpasses previous work in single- or multi-modality entries. The proposed method achieves leading performance in the nuScenes test set for both object detection and the following object tracking task. Code is made publicly available at https: //github. com/dvlab-research/UVTR.

NeurIPS Conference 2021 Conference Paper

Adversarial Reweighting for Partial Domain Adaptation

  • Xiang Gu
  • Xi Yu
  • Yan Yang
  • Jian Sun
  • Zongben Xu

Partial domain adaptation (PDA) has gained much attention due to its practical setting. The current PDA methods usually adapt the feature extractor by aligning the target and reweighted source domain distributions. In this paper, we experimentally find that the feature adaptation by the reweighted distribution alignment in some state-of-the-art PDA methods is not robust to the ``noisy'' weights of source domain data, leading to negative domain transfer on some challenging benchmarks. To tackle the challenge of negative domain transfer, we propose a novel Adversarial Reweighting (AR) approach that adversarially learns the weights of source domain data to align the source and target domain distributions, and the transferable deep recognition network is learned on the reweighted source domain data. Based on this idea, we propose a training algorithm that alternately updates the parameters of the network and optimizes the weights of source domain data. Extensive experiments show that our method achieves state-of-the-art results on the benchmarks of ImageNet-Caltech, Office-Home, VisDA-2017, and DomainNet. Ablation studies also confirm the effectiveness of our approach.

NeurIPS Conference 2021 Conference Paper

Dynamic Grained Encoder for Vision Transformers

  • Lin Song
  • Songyang Zhang
  • Songtao Liu
  • Zeming Li
  • Xuming He
  • Hongbin Sun
  • Jian Sun
  • Nanning Zheng

Transformers, the de-facto standard for language modeling, have been recently applied for vision tasks. This paper introduces sparse queries for vision transformers to exploit the intrinsic spatial redundancy of natural images and save computational costs. Specifically, we propose a Dynamic Grained Encoder for vision transformers, which can adaptively assign a suitable number of queries to each spatial region. Thus it achieves a fine-grained representation in discriminative regions while keeping high efficiency. Besides, the dynamic grained encoder is compatible with most vision transformer frameworks. Without bells and whistles, our encoder allows the state-of-the-art vision transformers to reduce computational complexity by 40%-60% while maintaining comparable performance on image classification. Extensive experiments on object detection and segmentation further demonstrate the generalizability of our approach. Code is available at https: //github. com/StevenGrove/vtpack.

AAAI Conference 2021 Conference Paper

Dynamic Hybrid Relation Exploration Network for Cross-Domain Context-Dependent Semantic Parsing

  • Binyuan Hui
  • Ruiying Geng
  • Qiyu Ren
  • Binhua Li
  • Yongbin Li
  • Jian Sun
  • Fei Huang
  • Luo Si

Semantic parsing has long been a fundamental problem in natural language processing. Recently, cross-domain contextdependent semantic parsing has become a new focus of research. Central to the problem is the challenge of leveraging contextual information of both natural language utterance and database schemas in the interaction history. In this paper, we present a dynamic graph framework that is capable of effectively modelling contextual utterances, tokens, database schemas, and their complicated interaction as the conversation proceeds. The framework employs a dynamic memory decay mechanism that incorporates inductive bias to integrate enriched contextual relation representation, which is further enhanced with a powerful reranking model. At the time of writing, we demonstrate that the proposed framework outperforms all existing models by large margins, achieving new state-of-the-art performance on two large-scale benchmarks, the SParC and CoSQL datasets. Specifically, the model attains a 55. 8% question-match and 30. 8% interaction-match accuracy on SParC, and a 46. 8% question-match and 17. 0% interaction-match accuracy on CoSQL.

NeurIPS Conference 2021 Conference Paper

Instance-Conditional Knowledge Distillation for Object Detection

  • Zijian Kang
  • Peizhen Zhang
  • Xiangyu Zhang
  • Jian Sun
  • Nanning Zheng

Knowledge distillation has shown great success in classification, however, it is still challenging for detection. In a typical image for detection, representations from different locations may have different contributions to detection targets, making the distillation hard to balance. In this paper, we propose a conditional distillation framework to distill the desired knowledge, namely knowledge that is beneficial in terms of both classification and localization for every instance. The framework introduces a learnable conditional decoding module, which retrieves information given each target instance as query. Specifically, we encode the condition information as query and use the teacher's representations as key. The attention between query and key is used to measure the contribution of different features, guided by a localization-recognition-sensitive auxiliary task. Extensive experiments demonstrate the efficacy of our method: we observe impressive improvements under various settings. Notably, we boost RetinaNet with ResNet-50 backbone from $37. 4$ to $40. 7$ mAP ($+3. 3$) under $1\times$ schedule, that even surpasses the teacher ($40. 4$ mAP) with ResNet-101 backbone under $3\times$ schedule. Code has been released on https: //github. com/megvii-research/ICD.

NeurIPS Conference 2021 Conference Paper

Spherical Motion Dynamics: Learning Dynamics of Normalized Neural Network using SGD and Weight Decay

  • Ruosi Wan
  • Zhanxing Zhu
  • Xiangyu Zhang
  • Jian Sun

In this paper, we comprehensively reveal the learning dynamics of normalized neural network using Stochastic Gradient Descent (with momentum) and Weight Decay (WD), named as Spherical Motion Dynamics (SMD). Most related works focus on studying behavior of effective learning rate" in equilibrium" state, i. e. assuming weight norm remains unchanged. However, their discussion on why this equilibrium can be reached is either absent or less convincing. Our work directly explores the cause of equilibrium, as a special state of SMD. Specifically, 1) we introduce the assumptions that can lead to equilibrium state in SMD, and prove equilibrium can be reached in a linear rate regime under given assumptions; 2) we propose ``angular update" as a substitute for effective learning rate to depict the state of SMD, and derive the theoretical value of angular update in equilibrium state; 3) we verify our assumptions and theoretical results on various large-scale computer vision tasks including ImageNet and MSCOCO with standard settings. Experiment results show our theoretical findings agree well with empirical observations. We also show that the behavior of angular update in SMD can produce interesting effect to the optimization of neural network in practice.

AAAI Conference 2021 Conference Paper

Unsupervised Learning of Deterministic Dialogue Structure with Edge-Enhanced Graph Auto-Encoder

  • Yajing Sun
  • Yong Shan
  • Chengguang Tang
  • Yue Hu
  • Yinpei Dai
  • Jing Yu
  • Jian Sun
  • Fei Huang

It is important for task-oriented dialogue systems to discover the dialogue structure (i. e. the general dialogue flow) from dialogue corpora automatically. Previous work models dialogue structure by extracting latent states for each utterance first and then calculating the transition probabilities among states. These two-stage methods ignore the contextual information when calculating the probabilities, which makes the transitions between the states ambiguous. This paper proposes a conversational graph (CG) to represent deterministic dialogue structure where nodes and edges represent the utterance and context information respectively. An unsupervised Edge- Enhanced Graph Auto-Encoder (EGAE) architecture is designed to model local-contextual and global-structural information for conversational graph learning. Furthermore, a selfsupervised objective is introduced with the response selection task to guide the unsupervised learning of the dialogue structure. Experimental results on several public datasets demonstrate that the novel model outperforms several alternatives in aggregating utterances with similar semantics. The effectiveness of the learned dialogue structured is also verified by more than 5% joint accuracy improvement in the downstream task of low resource dialogue state tracking.

NeurIPS Conference 2020 Conference Paper

Decentralized TD Tracking with Linear Function Approximation and its Finite-Time Analysis

  • Gang Wang
  • Songtao Lu
  • Georgios Giannakis
  • Gerald Tesauro
  • Jian Sun

The present contribution deals with decentralized policy evaluation in multi-agent Markov decision processes using temporal-difference (TD) methods with linear function approximation for scalability. The agents cooperate to estimate the value function of such a process by observing continual state transitions of a shared environment over the graph of interconnected nodes (agents), along with locally private rewards. Different from existing consensus-type TD algorithms, the approach here develops a simple decentralized TD tracker by wedding TD learning with gradient tracking techniques. The non-asymptotic properties of the novel TD tracker are established for both independent and identically distributed (i. i. d. ) as well as Markovian transitions through a unifying multistep Lyapunov analysis. In contrast to the prior art, the novel algorithm forgoes the limiting error bounds on the number of agents, which endows it with performance comparable to that of centralized TD methods that are the sharpest known to date.

NeurIPS Conference 2020 Conference Paper

Fine-Grained Dynamic Head for Object Detection

  • Lin Song
  • Yanwei Li
  • Zhengkai Jiang
  • Zeming Li
  • Hongbin Sun
  • Jian Sun
  • Nanning Zheng

The Feature Pyramid Network (FPN) presents a remarkable approach to alleviate the scale variance in object representation by performing instance-level assignments. Nevertheless, this strategy ignores the distinct characteristics of different sub-regions in an instance. To this end, we propose a fine-grained dynamic head to conditionally select a pixel-level combination of FPN features from different scales for each instance, which further releases the ability of multi-scale feature representation. Moreover, we design a spatial gate with the new activation function to reduce computational complexity dramatically through spatially sparse convolutions. Extensive experiments demonstrate the effectiveness and efficiency of the proposed method on several state-of-the-art detection benchmarks. Code is available at https: //github. com/StevenGrove/DynamicHead.

AAAI Conference 2020 Conference Paper

Multi-Point Semantic Representation for Intent Classification

  • Jinghan Zhang
  • Yuxiao Ye
  • Yue Zhang
  • Likun Qiu
  • Bin Fu
  • Yang Li
  • Zhenglu Yang
  • Jian Sun

Detecting user intents from utterances is the basis of natural language understanding (NLU) task. To understand the meaning of utterances, some work focuses on fully representing utterances via semantic parsing in which annotation cost is labor-intentsive. While some researchers simply view this as intent classification or frequently asked questions (FAQs) retrieval, they do not leverage the shared utterances among different intents. We propose a simple and novel multi-point semantic representation framework with relatively low annotation cost to leverage the fine-grained factor information, decomposing queries into four factors, i. e. , topic, predicate, object/condition, query type. Besides, we propose a compositional intent bi-attention model under multi-task learning with three kinds of attention mechanisms among queries, labels and factors, which jointly combines coarse-grained intent and fine-grained factor information. Extensive experiments show that our framework and model significantly outperform several state-of-the-art approaches with an improvement of 1. 35%-2. 47% in terms of accuracy.

NeurIPS Conference 2020 Conference Paper

Rethinking Learnable Tree Filter for Generic Feature Transform

  • Lin Song
  • Yanwei Li
  • Zhengkai Jiang
  • Zeming Li
  • Xiangyu Zhang
  • Hongbin Sun
  • Jian Sun
  • Nanning Zheng

The Learnable Tree Filter presents a remarkable approach to model structure-preserving relations for semantic segmentation. Nevertheless, the intrinsic geometric constraint forces it to focus on the regions with close spatial distance, hindering the effective long-range interactions. To relax the geometric constraint, we give the analysis by reformulating it as a Markov Random Field and introduce a learnable unary term. Besides, we propose a learnable spanning tree algorithm to replace the original non-differentiable one, which further improves the flexibility and robustness. With the above improvements, our method can better capture long range dependencies and preserve structural details with linear complexity, which is extended to several vision tasks for more generic feature transform. Extensive experiments on object detection/instance segmentation demonstrate the consistent improvements over the original version. For semantic segmentation, we achieve leading performance (82. 1% mIoU) on the Cityscapes benchmark without bells-and whistles. Code is available at https: //github. com/StevenGrove/LearnableTreeFilterV2.

IJCAI Conference 2020 Conference Paper

Speeding up Very Fast Decision Tree with Low Computational Cost

  • Jian Sun
  • Hongyu Jia
  • Bo Hu
  • Xiao Huang
  • Hao Zhang
  • Hai Wan
  • Xibin Zhao

Very Fast Decision Tree (VFDT) is one of the most widely used online decision tree induction algorithms, and it provides high classification accuracy with theoretical guarantees. In VFDT, the split-attempt operation is essential for leaf-split. It is computation-intensive since it computes the heuristic measure of all attributes of a leaf. To reduce split-attempts, VFDT tries to split at constant intervals (for example, every 200 examples). However, this mechanism introduces split-delay for split can only happen at fixed intervals, which slows down the growth of VFDT and finally lowers accuracy. To address this problem, we first devise an online incremental algorithm that computes the heuristic measure of an attribute with a much lower computational cost. Then a subset of attributes is carefully selected to find a potential split timing using this algorithm. A split-attempt will be carried out once the timing is verified. By the whole process, computational cost and split-delay are lowered significantly. Comprehensive experiments are conducted using multiple synthetic and real datasets. Compared with state-of-the-art algorithms, our method reduces split-attempts by about 5 to 10 times on average with much lower split-delay, which makes our algorithm run faster and more accurate.

NeurIPS Conference 2019 Conference Paper

DetNAS: Backbone Search for Object Detection

  • Yukang Chen
  • Tong Yang
  • Xiangyu Zhang
  • Gaofeng Meng
  • Xinyu Xiao
  • Jian Sun

Object detectors are usually equipped with backbone networks designed for image classification. It might be sub-optimal because of the gap between the tasks of image classification and object detection. In this work, we present DetNAS to use Neural Architecture Search (NAS) for the design of better backbones for object detection. It is non-trivial because detection training typically needs ImageNetpre-training while NAS systems require accuracies on the target detection task as supervisory signals. Based on the technique of one-shot supernet, which contains all possible networks in the search space, we propose a framework for backbone search on object detection. We train the supernet under the typical detector training schedule: ImageNet pre-training and detection fine-tuning. Then, the architecture search is performed on the trained supernet, using the detection task as the guidance. This framework makes NAS on backbones very efficient. In experiments, we show the effectiveness of DetNAS on various detectors, for instance, one-stage RetinaNetand the two-stage FPN. We empirically find that networks searched on object detection shows consistent superiority compared to those searched on ImageNet classification. The resulting architecture achieves superior performance than hand-crafted networks on COCO with much less FLOPs complexity.

AAAI Conference 2019 Conference Paper

HyperAdam: A Learnable Task-Adaptive Adam for Network Training

  • Shipeng Wang
  • Jian Sun
  • Zongben Xu

Deep neural networks are traditionally trained using humandesigned stochastic optimization algorithms, such as SGD and Adam. Recently, the approach of learning to optimize network parameters has emerged as a promising research topic. However, these learned black-box optimizers sometimes do not fully utilize the experience in human-designed optimizers, therefore have limitation in generalization ability. In this paper, a new optimizer, dubbed as HyperAdam, is proposed that combines the idea of “learning to optimize” and traditional Adam optimizer. Given a network for training, its parameter update in each iteration generated by HyperAdam is an adaptive combination of multiple updates generated by Adam with varying decay rates. The combination weights and decay rates in HyperAdam are adaptively learned depending on the task. HyperAdam is modeled as a recurrent neural network with AdamCell, WeightCell and StateCell. It is justified to be state-of-the-art for various network training, such as multilayer perceptron, CNN and LSTM.

NeurIPS Conference 2019 Conference Paper

Learnable Tree Filter for Structure-preserving Feature Transform

  • Lin Song
  • Yanwei Li
  • Zeming Li
  • Gang Yu
  • Hongbin Sun
  • Jian Sun
  • Nanning Zheng

Learning discriminative global features plays a vital role in semantic segmentation. And most of the existing methods adopt stacks of local convolutions or non-local blocks to capture long-range context. However, due to the absence of spatial structure preservation, these operators ignore the object details when enlarging receptive fields. In this paper, we propose the learnable tree filter to form a generic tree filtering module that leverages the structural property of minimal spanning tree to model long-range dependencies while preserving the details. Furthermore, we propose a highly efficient linear-time algorithm to reduce resource consumption. Thus, the designed modules can be plugged into existing deep neural networks conveniently. To this end, tree filtering modules are embedded to formulate a unified framework for semantic segmentation. We conduct extensive ablation studies to elaborate on the effectiveness and efficiency of the proposed method. Specifically, it attains better performance with much less overhead compared with the classic PSP block and Non-local operation under the same backbone. Our approach is proved to achieve consistent improvements on several benchmarks without bells-and-whistles. Code and models are available at https: //github. com/StevenGrove/TreeFilter-Torch.

NeurIPS Conference 2019 Conference Paper

Neural Diffusion Distance for Image Segmentation

  • Jian Sun
  • Zongben Xu

Diffusion distance is a spectral method for measuring distance among nodes on graph considering global data structure. In this work, we propose a spec-diff-net for computing diffusion distance on graph based on approximate spectral decomposition. The network is a differentiable deep architecture consisting of feature extraction and diffusion distance modules for computing diffusion distance on image by end-to-end training. We design low resolution kernel matching loss and high resolution segment matching loss to enforce the network's output to be consistent with human-labeled image segments. To compute high-resolution diffusion distance or segmentation mask, we design an up-sampling strategy by feature-attentional interpolation which can be learned when training spec-diff-net. With the learned diffusion distance, we propose a hierarchical image segmentation method outperforming previous segmentation methods. Moreover, a weakly supervised semantic segmentation network is designed using diffusion distance and achieved promising results on PASCAL VOC 2012 segmentation dataset.

NeurIPS Conference 2018 Conference Paper

MetaAnchor: Learning to Detect Objects with Customized Anchors

  • Tong Yang
  • Xiangyu Zhang
  • Zeming Li
  • Wenqiang Zhang
  • Jian Sun

We propose a novel and flexible anchor mechanism named MetaAnchor for object detection frameworks. Unlike many previous detectors model anchors via a predefined manner, in MetaAnchor anchor functions could be dynamically generated from the arbitrary customized prior boxes. Taking advantage of weight prediction, MetaAnchor is able to work with most of the anchor-based object detection systems such as RetinaNet. Compared with the predefined anchor scheme, we empirically find that MetaAnchor is more robust to anchor settings and bounding box distributions; in addition, it also shows the potential on the transfer task. Our experiment on COCO detection task shows MetaAnchor consistently outperforms the counterparts in various scenarios.

IJCAI Conference 2017 Conference Paper

Dual Track Multimodal Automatic Learning through Human-Robot Interaction

  • Shuqiang Jiang
  • Weiqing Min
  • Xue Li
  • Huayang Wang
  • Jian Sun
  • Jiaqi Zhou

Human beings are constantly improving their cognitive ability via automatic learning from the interaction with the environment. Two important aspects of automatic learning are the visual perception and knowledge acquisition. The fusion of these two aspects is vital for improving the intelligence and interaction performance of robots. Many automatic knowledge extraction and recognition methods have been widely studied. However, little work focuses on integrating automatic knowledge extraction and recognition into a unified framework to enable jointly visual perception and knowledge acquisition. To solve this problem, we propose a Dual Track Multimodal Automatic Learning (DTMAL) system, which consists of two components: Hybrid Incremental Learning (HIL) from the vision track and Multimodal Knowledge Extraction (MKE) from the knowledge track. HIL can incrementally improve recognition ability of the system by learning new object samples and new object concepts. MKE is capable of constructing and updating the multimodal knowledge items based on the recognized new objects from HIL and other knowledge by exploring the multimodal signals. The fusion of the two tracks is a mutual promotion process and jointly devote to the dual track learning. We have conducted the experiments through human-machine interaction and the experimental results validated the effectiveness of our proposed system.

NeurIPS Conference 2016 Conference Paper

Deep ADMM-Net for Compressive Sensing MRI

  • Yan Yang
  • Jian Sun
  • Huibin Li
  • Zongben Xu

Compressive Sensing (CS) is an effective approach for fast Magnetic Resonance Imaging (MRI). It aims at reconstructing MR image from a small number of under-sampled data in k-space, and accelerating the data acquisition in MRI. To improve the current MRI system in reconstruction accuracy and computational speed, in this paper, we propose a novel deep architecture, dubbed ADMM-Net. ADMM-Net is defined over a data flow graph, which is derived from the iterative procedures in Alternating Direction Method of Multipliers (ADMM) algorithm for optimizing a CS-based MRI model. In the training phase, all parameters of the net, e. g. , image transforms, shrinkage functions, etc. , are discriminatively trained end-to-end using L-BFGS algorithm. In the testing phase, it has computational overhead similar to ADMM but uses optimized parameters learned from the training data for CS-based reconstruction task. Experiments on MRI image reconstruction under different sampling ratios in k-space demonstrate that it significantly improves the baseline ADMM algorithm and achieves high reconstruction accuracies with fast computational speed.

NeurIPS Conference 2016 Conference Paper

R-FCN: Object Detection via Region-based Fully Convolutional Networks

  • Jifeng Dai
  • Yi Li
  • Kaiming He
  • Jian Sun

We present region-based, fully convolutional networks for accurate and efficient object detection. In contrast to previous region-based detectors such as Fast/Faster R-CNN that apply a costly per-region subnetwork hundreds of times, our region-based detector is fully convolutional with almost all computation shared on the entire image. To achieve this goal, we propose position-sensitive score maps to address a dilemma between translation-invariance in image classification and translation-variance in object detection. Our method can thus naturally adopt fully convolutional image classifier backbones, such as the latest Residual Networks (ResNets), for object detection. We show competitive results on the PASCAL VOC datasets (e. g. , 83. 6% mAP on the 2007 set) with the 101-layer ResNet. Meanwhile, our result is achieved at a test-time speed of 170ms per image, 2. 5-20 times faster than the Faster R-CNN counterpart. Code is made publicly available at: https: //github. com/daijifeng001/r-fcn.

NeurIPS Conference 2015 Conference Paper

Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks

  • Shaoqing Ren
  • Kaiming He
  • Ross Girshick
  • Jian Sun

State-of-the-art object detection networks depend on region proposal algorithms to hypothesize object locations. Advances like SPPnet and Fast R-CNN have reduced the running time of these detection networks, exposing region proposal computation as a bottleneck. In this work, we introduce a Region Proposal Network (RPN) that shares full-image convolutional features with the detection network, thus enabling nearly cost-free region proposals. An RPN is a fully-convolutional network that simultaneously predicts object bounds and objectness scores at each position. RPNs are trained end-to-end to generate high-quality region proposals, which are used by Fast R-CNN for detection. With a simple alternating optimization, RPN and Fast R-CNN can be trained to share convolutional features. For the very deep VGG-16 model, our detection system has a frame rate of 5fps (including all steps) on a GPU, while achieving state-of-the-art object detection accuracy on PASCAL VOC 2007 (73. 2% mAP) and 2012 (70. 4% mAP) using 300 proposals per image. Code is available at https: //github. com/ShaoqingRen/faster_rcnn.

JBHI Journal 2014 Journal Article

Limited Correlation Between Conventional Pathologist and Automatic Computer-Assisted Quantification of Hepatic Steatosis due to Difference Between Event-Based and Surface-Based Analysis

  • Meihong Deng
  • Uta Dahmen
  • Jian Sun
  • Hai Huang
  • Christian Sehestedt
  • Andre Homeyer
  • Andrea Schenk
  • Olaf Dirsch

Computer-assisted automatic quantification (CAQ) was developed as an alternative method for the diagnosis of hepatic steatosis in order to compensate for observer-dependent bias. Here, we aim to demonstrate that CAQ can provide an accurate and precise result in analysis of fatty content, but that it is inappropriate to validate CAQ by comparison with conventional pathologist estimation (PE). Male rats were fed with a methionine-choline-deficient plus high-fat diet for three days, one week, or two weeks to induce mild, moderate, or severe steatosis. Samples were collected from all liver lobes. Severity of hepatic steatosis was assessed by an experienced pathologist who estimated the percentage of hepatocytes containing lipid droplets. Fatty content was quantified by PE, CAQ, and biochemical analysis (BA). CAQ, PE, and BA can correctly reflect severe fatty change. However, in the case of mild and moderate steatosis, PE could not reflect the true fatty content ( r between PE and BA was <; 0). The result of CAQ correlated well with that of BA among the various degrees of severity of hepatic steatosis. In conclusion, due to a difference between event-based and surface-based analysis, it is inappropriate to validate the CAQ of hepatic steatosis by comparison with PE.

v2026.09.13