Arrow Research search

Author name cluster

Jin Zhang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

36 papers
2 author rows

Possible papers

36

EAAI Journal 2026 Journal Article

High-precision multimodal vehicle trajectory prediction model based on cross-layer interleaved spatiotemporal attention mechanism

  • Fei Teng
  • Liqiang Jin
  • Junnian Wang
  • Feng Xiao
  • Mengdi Guo
  • Yanbo Zhou
  • Jin Zhang

In increasingly complex traffic environments, spatiotemporal attention mechanisms have made remarkable advancements in scene-level interaction modelling. However, the deep and multi-scale spatiotemporal representations required for safe and efficient decision-making in intelligent vehicles remain underexplored. Aiming to address this limitation, this study proposes a multimodal trajectory prediction model based on a cross-layer interleaved spatiotemporal attention (CLISTA) mechanism. Compared with conventional spatiotemporal attention frameworks, CLISTA more effectively captures multi-scale spatiotemporal interactions in complex traffic scenes through the alternating fusion of spatial and temporal features across network layers via a cross-layer interleaving structure. Firstly, spatial, dynamic and heading conflict risks are derived from the relative motion between the target vehicle and its neighbours and aggregated into a social grid weight matrix, through which the neighbours' collective influence on the target vehicle is quantified. Secondly, spatial and temporal multi-head attention modules are designed within each layer. By integrating an interleaved ‘spatial–temporal’ stacking strategy with cross-layer skip connections, the model facilitates progressive alignment and deep fusion, ranging from local interactions to long-range dependencies. Subsequently, an intention recognition module is developed. A second-order gated bilinear fusion mechanism is introduced to adaptively model higher-order couplings between local neighbour dynamics and global interaction semantics, thereby yielding a multimodal probability distribution over the target vehicle's driving intentions. Lastly, multimodal trajectory predictions are generated by decoding the fused spatiotemporal features together with the inferred intention information. Experimental results on three benchmark datasets—NGSIM (Next Generation Simulation), AD4CHE (Aerial Dataset for China Congested Highway and Expressway), and highD—demonstrate that CLISTA consistently outperforms the baseline methods. Relative to the next-best model, it reduces average/final displacement errors by 16. 67 %/21. 23 %, 12. 99 %/21. 14 % and 10. 53 %/21. 59 % on NGSIM, AD4CHE and HighD, respectively. Overall, CLISTA offers reliable multi-hypothesis trajectory priors for safe and efficient decision-making in complex traffic scenarios.

AIJ Journal 2026 Journal Article

Interactive graph convolutional filtering

  • Jin Zhang
  • Defu Lian
  • Hong Xie
  • Yawen Li
  • Enhong Chen

Interactive Recommender Systems (IRS) have been increasingly used in various domains, including personalized article recommendation, social media, and online advertising. However, IRS faces significant challenges in providing accurate recommendations under limited observations, especially in the context of interactive collaborative filtering. These problems are exacerbated by the cold start problem and data sparsity problem. Existing Multi-Armed Bandit methods, despite their carefully designed exploration strategies, often struggle to provide satisfactory results in the early stages due to the lack of interaction data. Furthermore, these methods are computationally intractable when applied to non-linear models, limiting their applicability. To address these challenges, we propose a novel method, the Interactive Graph Convolutional Filtering model. Our proposed method extends interactive collaborative filtering into the graph model to enhance the performance of collaborative filtering between users and items. We incorporate variational inference techniques to overcome the computational hurdles posed by non-linear models. Furthermore, we employ Bayesian meta-learning methods to effectively address the cold-start problem and derive theoretical regret bounds for our proposed method, ensuring a robust performance guarantee. Extensive experimental results on three real-world datasets validate our method and demonstrate its superiority over existing baselines.

AAAI Conference 2026 Conference Paper

VirtualEnv: A Platform for Embodied AI Research

  • Kabir Swain
  • Sijie Han
  • Ayush Raina
  • Jin Zhang
  • Shuang Li
  • Michael Stopa
  • Antonio Torralba

As large language models (LLMs) continue to improve in reasoning and decision-making, there is a growing need for realistic and interactive environments where their abilities can be rigorously evaluated. We present VirtualEnv, a next-generation simulation platform built on Unreal Engine 5 that enables fine-grained benchmarking of LLMs in embodied and interactive scenarios. VirtualEnv supports rich agent–environment interactions, including object manipulation, navigation, and adaptive multi-agent collaboration, as well as game-inspired mechanics like escape rooms and procedurally generated environments. We provide a user-friendly API built on top of Unreal Engine, allowing researchers to deploy and control LLM-driven agents using natural language instructions. We integrate large-scale LLMs and vision-language models (VLMs), such as GPT-based models, to generate novel environments and structured tasks from multimodal inputs. Our experiments benchmark the performance of several popular LLMs across tasks of increasing complexity, analyzing differences in adaptability, planning, and multi-agent coordination. We also describe our methodology for procedural task generation, task validation, and real-time environment control. VirtualEnv is released as an open-source platform, we aim to advance research at the intersection of AI and gaming, enable standardized evaluation of LLMs in embodied AI settings, and pave the way for future developments in immersive simulations and interactive entertainment.

NeurIPS Conference 2025 Conference Paper

A Single-Loop Gradient Algorithm for Pessimistic Bilevel Optimization via Smooth Approximation

  • Qichao Cao
  • Shangzhi Zeng
  • Jin Zhang

Bilevel optimization has garnered significant attention in the machine learning community recently, particularly regarding the development of efficient numerical methods. While substantial progress has been made in developing efficient algorithms for optimistic bilevel optimization, the study of methods for solving Pessimistic Bilevel Optimization (PBO) remains relatively less explored, especially the design of fully first-order, single-loop gradient-based algorithms. This paper aims to bridge this research gap. We first propose a novel smooth approximation to the PBO problem, using penalization and regularization techniques. Building upon this approximation, we then propose SiPBA (Single-loop Pessimistic Bilevel Algorithm), a new gradient-based method specifically designed for PBO which avoids second-order derivative information or inner-loop iterations for subproblem solving. We provide theoretical validation for the proposed smooth approximation scheme and establish theoretical convergence for the algorithm SiPBA. Numerical experiments on synthetic examples and practical applications demonstrate the effectiveness and efficiency of SiPBA.

EAAI Journal 2025 Journal Article

An on-line global–local defect detection framework for wide cold-rolled strip steel

  • Pan Jiang
  • Zhenying Xu
  • Wei Fan
  • Jin Zhang

Wide cold-rolled strip steel possess distinctive characteristics, including a large surface area, high reflectivity, and multiple defect scales. These attributes have posed a significant challenge for achieving efficient and accurate online detection of small irregular defects in existing research. To address this, this paper proposes a Global–Local Defect Detection Framework (GLDDF) based on machine vision. The GLDDF comprises three parts: an image feature extractor, a global detector and a local detector. A feature extraction backbone network has been introduced to effectively capture multiple, subtle defect features on the steel strip surface. The global detector leverages normalizing flow to estimate the probability density of extracted features and quickly localize anomalous regions based on anomaly scores, effectively reducing computational overhead. And the local detector classifies and localizes defects to improve detection accuracy. The specialized equipment has been developed to overcome high reflection on the strip surface, allowing for on-line collection of high-definition images. A dataset of strip steel images has been produced to validate the algorithm’s effectiveness. The proposed method achieved a mean Average Precision (mAP) of 86. 3% and a processing speed of 75. 7 Frames Per Second (FPS) on the collected dataset, demonstrating a significant improvement in detection speed compared to existing methods. The method has been initially implemented on the cold-rolled strip steel production line of a partner company, yielding promising results.

NeurIPS Conference 2025 Conference Paper

Bilevel Optimization for Adversarial Learning Problems: Sharpness, Generation, and Beyond

  • Risheng Liu
  • Zhu Liu
  • Weihao Mao
  • Wei Yao
  • Jin Zhang

Adversarial learning is a widely used paradigm in machine learning, often formulated as a min-max optimization problem where the inner maximization imposes adversarial constraints to guide the outer learner toward more robust solutions. This framework underlies methods such as Sharpness-Aware Minimization (SAM) and Generative Adversarial Networks (GANs). However, traditional gradient-based approaches to such problems often face challenges in balancing accuracy and efficiency due to second-order complexities. In this paper, we propose a bilevel optimization framework that reformulates these adversarial learning problems by leveraging the tractability of the lower-level problem. The bilevel framework introduces no additional complexity and enables the use of advanced bilevel tools. We further develop a provably convergent single-loop stochastic algorithm that effectively balances learning accuracy and computational cost. Extensive experiments show that our method improves generation quality in terms of FID and JS scores for GANs, and consistently achieves higher accuracy for SAM under label noise and across various backbones, while promoting flatter loss landscapes. Overall, this work provides a practical and theoretically grounded framework for solving adversarial learning tasks through bilevel optimization.

AAAI Conference 2025 Conference Paper

CoMT: A Novel Benchmark for Chain of Multi-modal Thought on Large Vision-Language Models

  • Zihui Cheng
  • Qiguang Chen
  • Jin Zhang
  • Hao Fei
  • Xiaocheng Feng
  • Wanxiang Che
  • Min Li
  • Libo Qin

Large Vision-Language Models (LVLMs) have recently demonstrated amazing success in multi-modal tasks, including advancements in Multi-modal Chain-of-Thought (MCoT) reasoning. Despite these successes, current benchmarks still follow a traditional paradigm with multi-modal input and text-modal output, which leads to significant drawbacks such as missing visual operations and vague expressions. Motivated by this, we introduce a novel Chain of Multi-modal Thought (CoMT) benchmark to address these limitations. Different from the traditional MCoT benchmark, CoMT requires both multi-modal input and multi-modal reasoning output, aiming to mimic human-like reasoning that inherently integrates visual operation. Specifically, CoMT consists of four categories: (1) Visual Creation, (2) Visual Deletion, (3) Visual Update, and (4) Visual Selection to comprehensively explore complex visual operations and concise expression in real scenarios. We evaluate various LVLMs and strategies on CoMT, revealing some key insights into the capabilities and limitations of the current approaches. We hope that CoMT can inspire more research on introducing multi-modal generation into the reasoning process.

ICML Conference 2025 Conference Paper

Efficient Curvature-Aware Hypergradient Approximation for Bilevel Optimization

  • Youran Dong
  • Junfeng Yang
  • Wei Yao
  • Jin Zhang

Bilevel optimization is a powerful tool for many machine learning problems, such as hyperparameter optimization and meta-learning. Estimating hypergradients (also known as implicit gradients) is crucial for developing gradient-based methods for bilevel optimization. In this work, we propose a computationally efficient technique for incorporating curvature information into the approximation of hypergradients and present a novel algorithmic framework based on the resulting enhanced hypergradient computation. We provide convergence rate guarantees for the proposed framework in both deterministic and stochastic scenarios, particularly showing improved computational complexity over popular gradient-based methods in the deterministic setting. This improvement in complexity arises from a careful exploitation of the hypergradient structure and the inexact Newton method. In addition to the theoretical speedup, numerical experiments demonstrate the significant practical performance benefits of incorporating curvature information.

ICLR Conference 2025 Conference Paper

Learning to Plan Before Answering: Self-Teaching LLMs to Learn Abstract Plans for Problem Solving

  • Jin Zhang
  • Flood Sung
  • Zhilin Yang
  • Yang Gao 0029
  • Chongjie Zhang

In the field of large language model (LLM) post-training, the effectiveness of utilizing synthetic data generated by the LLM itself has been well-presented. However, a key question remains unaddressed: what essential information should such self-generated data encapsulate? Existing approaches only produce step-by-step problem solutions, and fail to capture the abstract meta-knowledge necessary for generalization across similar problems. Drawing insights from cognitive science, where humans employ high-level abstraction to simplify complex problems before delving into specifics, we introduce a novel self-training algorithm: LEarning to Plan before Answering (LEPA). LEPA trains the LLM to formulate anticipatory plans, which serve as abstract meta-knowledge for problem-solving, before engaging with the intricacies of problems. This approach not only outlines the solution generation path but also shields the LLM from the distraction of irrelevant details. During data generation, LEPA first crafts an anticipatory plan based on the problem, and then generates a solution that aligns with both the plan and the problem. LEPA refines the plan through self-reflection, aiming to acquire plans that are instrumental in yielding correct solutions. During model optimization, the LLM is trained to predict both the refined plans and the corresponding solutions. By efficiently extracting and utilizing the anticipatory plans, LEPA demonstrates remarkable superiority over conventional algorithms on various challenging natural language reasoning benchmarks.

NeurIPS Conference 2025 Conference Paper

MMCSBench: A Fine-Grained Benchmark for Large Vision-Language Models in Camouflage Scenes

  • Jin Zhang
  • Ruiheng Zhang
  • Zhe Cao
  • Kaizheng Chen

Current camouflaged object detection methods predominantly follow discriminative segmentation paradigms and heavily rely on predefined categories present in the training data, limiting their generalization to unseen or emerging camouflage objects. This limitation is further compounded by the labor-intensive and time-consuming nature of collecting camouflage imagery. Although Large Vision-Language Models (LVLMs) show potential to improve such issues with their powerful generative capabilities, their understanding of camouflage scenes is still insufficient. To bridge this gap, we introduce MMCSBench, the first comprehensive multimodal benchmark designed to evaluate and advance LVLM capabilities in camouflage scenes. MMCSBench comprises 22, 537 images and 76, 843 corresponding image-text pairs across five fine-grained camouflage tasks. Additionally, we propose a new task, Camouflage Efficacy Assessment (CEA), aimed at quantitatively evaluating the camouflage effectiveness of objects in images and enabling automated collection of camouflage images from large-scale databases. Extensive experiments on 26 LVLMs reveal significant shortcomings in models' ability to perceive and interpret camouflage scenes. These findings highlight the fundamental differences between natural and camouflaged visual inputs, offering insights for future research in advancing LVLM capabilities within this challenging domain.

ICLR Conference 2025 Conference Paper

Overcoming Lower-Level Constraints in Bilevel Optimization: A Novel Approach with Regularized Gap Functions

  • Wei Yao
  • Haian Yin
  • Shangzhi Zeng
  • Jin Zhang

Constrained bilevel optimization tackles nested structures present in constrained learning tasks like constrained meta-learning, adversarial learning, and distributed bilevel optimization. However, existing bilevel optimization methods mostly are typically restricted to specific constraint settings, such as linear lower-level constraints. In this work, we overcome this limitation and develop a new single-loop, Hessian-free constrained bilevel algorithm capable of handling more general lower-level constraints. We achieve this by employing a doubly regularized gap function tailored to the constrained lower-level problem, transforming constrained bilevel optimization into an equivalent single-level optimization problem with a single smooth constraint. We rigorously establish the non-asymptotic convergence analysis of the proposed algorithm under the convexity of lower-level problem, avoiding the need for strong convexity assumptions on the lower-level objective or coupling convexity assumptions on lower-level constraints found in existing literature. Additionally, the generality of our method allows for its extension to bilevel optimization with minimax lower-level problem. We evaluate the effectiveness and efficiency of our algorithm on various synthetic problems, typical hyperparameter learning tasks, and generative adversarial network.

ICLR Conference 2025 Conference Paper

qNBO: quasi-Newton Meets Bilevel Optimization

  • Sheng Fang
  • Yongjin Liu
  • Wei Yao
  • Chengming Yu
  • Jin Zhang

Bilevel optimization, which addresses challenges in hierarchical learning tasks, has gained significant interest in machine learning. Implementing gradient descent for bilevel optimization presents computational hurdles, notably the need to compute the exact lower-level solution and the inverse Hessian of the lower-level objective. While these two aspects are inherently connected, existing methods typically handle them separately by solving the lower-level problem and a linear system for the inverse Hessian-vector product. In this paper, we introduce a general framework to tackle these computational challenges in a coordinated manner. Specifically, we leverage quasi-Newton algorithms to accelerate the solution of the lower-level problem while efficiently approximating the inverse Hessian-vector product. Furthermore, by leveraging the superlinear convergence properties of BFGS, we establish a non-asymptotic convergence analysis for the BFGS adaptation within our framework. Numerical experiments demonstrate the comparable or superior performance of our proposed algorithms in real-world learning tasks, including hyperparameter optimization, data hyper-cleaning, and few-shot meta-learning.

EAAI Journal 2025 Journal Article

Semi-supervised contrastive learning for flotation process monitoring with uncertainty-aware prototype optimization

  • Mingxi Ai
  • Jin Zhang
  • Peng Li
  • Jiande Wu
  • Zhaohui Tang
  • Yongfang Xie

Froth flotation is a widely employed mineral beneficiation technique, and effective process monitoring is critical for optimizing mineral separation. However, in the industrial process, labeling froth images to create large labeled datasets is both expensive and time-consuming. Semi-supervised deep learning offers a promising solution, but leveraging unlabeled data remains a significant challenge. To this aim, we propose an uncertainty-aware semi-supervised contrastive learning method. Our approach employs a pseudo-labeling module with dropout to generate pseudo labels and estimate uncertainty. Based on the uncertainty scores, the pseudo-labeled data are split into reliable and unreliable sets. A semi-supervised contrastive learning module is developed to exploit pseudo-label information and learn class-aware representations from the reliable set. Additionally, a dynamically weighted consistency learning module is introduced to explore potential classification information in the unreliable data while preventing the model from being misled by low-confidence predictions. Comparison experiments on industrial zinc flotation data show our method achieves 88. 12% classification accuracy, surpassing the best alternative by a margin of 5. 11%. These results demonstrate that our method generalizes well to unseen test data and outperforms state-of-the-art methods.

IROS Conference 2025 Conference Paper

Stability Enhancement in Variable Morphing Multi-body AUVs for Underwater Structure Maintenance

  • Shuai Kang
  • Yuan He
  • Jin Zhang
  • Yuxi Gao
  • Yunfei Bai
  • Longchuan Li

This paper presents a Variable Morphing Multi-Body AUVs (VMMAUVs) concept, designed for underwater structure maintenance. This robot is capable of dynamically adjusting their structure to adapt to varying operational scenarios. The study explores two key stability mechanisms: buoyancy adjustment and aperture angle control, both aimed at optimizing the metacentric height. Through simulations and experiments with different buoyancy configurations and aperture angles, the results show that the proposed methods significantly enhance the system ’ s stability, enabling faster convergence and better posture retention. The feasibility of the control strategies is validated through various numerical simulations, demonstrating the effectiveness of angle tracking control and buoyancy adjustment in maintaining stability under dynamic oceanic conditions.

EAAI Journal 2025 Journal Article

Super-resolution reconstruction of sequential images based on an active shift via a hybrid attention calibration mechanism

  • Qiang Wu
  • Ziyi Yang
  • Hongfei Zeng
  • Jin Zhang
  • Haojie Xia

Image super-resolution reconstruction converts low-resolution images into high-resolution images, demonstrating extensive potential in processing sequential images. However, most Multi-Image Super-Resolution methods currently face two significant challenges: first, the lack of precision in the shift information between images, as these methods typically rely on algorithms to estimate relative motion. Second, the limited ability to effectively extract subpixel features from low-resolution images directly impacts the richness of details in reconstructed images. This paper proposes a novel active shift-based sequential image super-resolution reconstruction technique to address these issues. This technique integrates hardware control with deep learning algorithms, utilizing a Piezoelectric platform to control camera movement precisely, capturing sequential images with predetermined subpixel shifts, and accurately recording the relative shifts between images. At the algorithmic level, we have designed a hybrid network model that combines a convolutional neural network with a Transformer architecture and integrates channel attention and self-attention mechanisms. This model fully leverages the precise shift information provided by the hardware and significantly enhances the ability to extract image details and overall image quality. Experimental results demonstrate that our method outperforms single-image super-resolution techniques regarding Peak-Signal-to-Noise-Ratio and Structural Similarity Index Measure. To further validate the applicability and effectiveness of this technology, we conducted tests using a resolution test chart, which showed that our technique can increase the resolution of the original imaging system by 25. 6%. Therefore, the strategy combining hardware and software proposed in this paper effectively solves critical issues in Multi-Image Super-Resolution tasks and provides new pathways for image processing technologies.

JBHI Journal 2024 Journal Article

A Comprehensive Privacy-Preserving Federated Learning Scheme With Secure Authentication and Aggregation for Internet of Medical Things

  • Jingwei Liu
  • Jin Zhang
  • Mian Ahmad Jan
  • Rong Sun
  • Lei Liu
  • Sahil Verma
  • Pushpita Chatterjee

Data mining, integration, and utilization are the inevitable trend of the Internet of Medical Things (IoMT) in the context of Big Data. With the increasing demand for data privacy, federated learning has emerged as a new paradigm, which enables distributed joint training of medical data sources without leaving the private domain. However, federated learning is suffering from security threats as the shared local model will reveal original datasets. Privacy leakage is even more fatal in healthcare because medical data contains critically sensitive information. In addition, open wireless channels are susceptible to malicious attacks. To further safeguard the privacy of IoMT, we propose a comprehensive privacy-preserving federated learning scheme with a tactful dropout handling mechanism. The proposed scheme leverages blind masking and certificateless proxy re-encryption (CL-PRE) for secure aggregation, ensuring the confidentiality of the local model and rendering the global model invisible to any parties other than clients. It also provides authentication of uploaded models while protecting identity privacy. Compared with other relevant schemes, our solution has better performance on functional features and efficiency, and is more applicable to IoMT systems with many devices.

ICLR Conference 2024 Conference Paper

Constrained Bi-Level Optimization: Proximal Lagrangian Value Function Approach and Hessian-free Algorithm

  • Wei Yao
  • Chengming Yu
  • Shangzhi Zeng
  • Jin Zhang

This paper presents a new approach and algorithm for solving a class of constrained Bi-Level Optimization (BLO) problems in which the lower-level problem involves constraints coupling both upper-level and lower-level variables. Such problems have recently gained significant attention due to their broad applicability in machine learning. However, conventional gradient-based methods unavoidably rely on computationally intensive calculations related to the Hessian matrix. To address this challenge, we devise a smooth proximal Lagrangian value function to handle the constrained lower-level problem. Utilizing this construct, we introduce a single-level reformulation for constrained BLOs that transforms the original BLO problem into an equivalent optimization problem with smooth constraints. Enabled by this reformulation, we develop a Hessian-free gradient-based algorithm—termed proximal Lagrangian Value function-based Hessian-free Bi-level Algorithm (LV-HBA)—that is straightforward to implement in a single loop manner. Consequently, LV-HBA is especially well-suited for machine learning applications. Furthermore, we offer non-asymptotic convergence analysis for LV-HBA, eliminating the need for traditional strong convexity assumptions for the lower-level problem while also being capable of accommodating non-singleton scenarios. Empirical results substantiate the algorithm's superior practical performance.

EAAI Journal 2024 Journal Article

FCT-Net: A dual-encoding-path network fusing atrous spatial pyramid pooling and transformer for pavement crack detection

  • Bing Xiong
  • Rong Hong
  • Rui Liu
  • Jing Wang
  • Jin Zhang
  • Wei Li
  • Songtao Lv
  • Dongdong Ge

Cracks are a typical form of road damage, and accurate detection of cracks is of great significance for road maintenance work and ensuring traffic safety. Recently, computer vision has gradually been applied in the field of crack segmentation. However, there are still some extremely challenging problems in crack segmentation, such as complex backgrounds, information loss caused by pooling and convolution operations, and insufficient fusion of global and local semantic information. In response to the above problems, this paper proposes a dual-encoding-path network with U-Net architecture called FCT-Net, by fusing channel atrous spatial pyramid pooling (CASPP) and transformer. Specifically, CASPP obtains multi-scale receptive fields by incorporating spatial and channel attention, while refining and extracting local features. Meanwhile, we introduce long-short distance attention to construct a novel transformer with the prominent characteristic of interaction between local and global attention features. In addition, a residual convolution module is designed to enhance the local features of the transformer. Furthermore, we devise a multi-scale attention weight cross fusion module to aggregate the features of the dual encoding branch, for reducing information loss during downsampling and suppress background information. Eventually, we evaluate the performance of FCT-Net by experiments on three public datasets. Extensive experimental results show that FCT-Net achieves higher F1-score and mean intersection over union (mIoU) than state-of-the-art segmentation networks on the DeepCrack537 and CrackLS315 datasets. Meanwhile, it has excellent segmentation performance for cracks in complex scenes, with the highest recall, F1-score, and mIoU respectively as 85. 64%, 81. 67%, and 84. 05% on the CrackTree260 dataset.

NeurIPS Conference 2024 Conference Paper

Generalization Error Bounds for Two-stage Recommender Systems with Tree Structure

  • Jin Zhang
  • Ze Liu
  • Defu Lian
  • Enhong Chen

Two-stage recommender systems play a crucial role in efficiently identifying relevant items and personalizing recommendations from a vast array of options. This paper, based on an error decomposition framework, analyzes the generalization error for two-stage recommender systems with a tree structure, which consist of an efficient tree-based retriever and a more precise yet time-consuming ranker. We use the Rademacher complexity to establish the generalization upper bound for various tree-based retrievers using beam search, as well as for different ranker models under a shifted training distribution. Both theoretical insights and practical experiments on real-world datasets indicate that increasing the branches in tree-based retrievers and harmonizing distributions across stages can enhance the generalization performance of two-stage recommender systems.

JBHI Journal 2024 Journal Article

HST-MRF: Heterogeneous Swin Transformer With Multi-Receptive Field for Medical Image Segmentation

  • Xiaofei Huang
  • Hongfang Gong
  • Jin Zhang

The Transformer has been successfully used in medical image segmentation due to its excellent long-range modeling capabilities. However, patch segmentation is necessary when building a Transformer class model. This process ignores the tissue structure features within patch, resulting in the loss of shallow representation information. In this study, we propose a Heterogeneous Swin Transformer with Multi-Receptive Field (HST-MRF) model that fuses patch information from different receptive fields to solve the problem of loss of feature information caused by patch segmentation. The heterogeneous Swin Transformer (HST) is the core module, which achieves the interaction of multi-receptive field patch information through heterogeneous attention and passes it to the next stage for progressive learning, thus complementing the patch structure information. We also designed a two-stage fusion module, multimodal bilinear pooling (MBP), to assist HST in further fusing multi-receptive field information and combining low-level and high-level semantic information for accurate localization of lesion regions. In addition, we developed adaptive patch embedding (APE) and soft channel attention (SCA) modules to retain more valuable information when acquiring patch embedding and filtering channel features, respectively, thereby improving model segmentation quality. We evaluated HST-MRF on multiple datasets for polyp, skin lesion and breast ultrasound segmentation tasks. Experimental results show that our proposed method outperforms state-of-the-art models and can achieve superior performance. Furthermore, we verified the effectiveness of each module and the benefits of multi-receptive field segmentation in reducing the loss of structural information through ablation experiments and qualitative analysis.

EAAI Journal 2024 Journal Article

Image super-resolution reconstruction using Swin Transformer with efficient channel attention networks

  • Zhenxi Sun
  • Jin Zhang
  • Ziyi Chen
  • Lu Hong
  • Rui Zhang
  • Weishi Li
  • Haojie Xia

Image super-resolution reconstruction (SR) is an important ill-posed problem in low-level vision, which aims to reconstruct high-resolution images from low-resolution images. Although current state-of-the-art methods exhibit impressive performance, their recovery of image detail information and edge information is still unsatisfactory. To address this problem, this paper proposes a shifted window Transformer (Swin Transformer) with an efficient channel attention network (S-ECAN), which combines the attention based on convolutional neural networks and the self-attention of the Swin Transformer to combine the advantages of both and focuses on learning high-frequency features of images. In addition, to solve the problem of Convolutional Neural Network (CNN) based channel attention consumes a large number of parameters to achieve good performance, this paper proposes the Efficient Channel Attention Block (ECAB), which only involves a handful of parameters while bringing clear performance gain. Extensive experimental validation shows that the proposed model can recover more high-frequency details and texture information. The model is validated on Set5, Set14, B100, Urban100, and Manga109 datasets, where it outperforms the state-of-the-art methods by 0. 03–0. 13 dB, 0. 04–0. 09 dB, 0. 01–0. 06 dB, 0. 13–0. 20 dB, and 0. 06–0. 17 dB respectively in terms of objective metrics. Ultimately, the substantial performance gains and enhanced visual results over prior arts validate the effectiveness and competitiveness of our proposed approach, which achieves an improved performance-complexity trade-off.

AAAI Conference 2024 Short Paper

Scene Flow Prior Based Point Cloud Completion with Masked Transformer (Student Abstract)

  • Junzhe Ding
  • Yufei Que
  • Jin Zhang
  • Cheng Wu

It is necessary to explore an effective point cloud completion mechanism that is of great significance for real-world tasks such as autonomous driving, robotics applications, and multi-target tracking. In this paper, we propose a point cloud completion method using a self-supervised transformer model based on the contextual constraints of scene flow. Our method uses the multi-frame point cloud context relationship as a guide to generate a series of token proposals, this priori condition ensures the stability of the point cloud completion. The experimental results show that the method proposed in this paper achieves high accuracy and good stability.

ICML Conference 2024 Conference Paper

SPABA: A Single-Loop and Probabilistic Stochastic Bilevel Algorithm Achieving Optimal Sample Complexity

  • Tianshu Chu
  • Dachuan Xu
  • Wei Yao
  • Jin Zhang

While stochastic bilevel optimization methods have been extensively studied for addressing large-scale nested optimization problems in machine learning, it remains an open question whether the optimal complexity bounds for solving bilevel optimization are the same as those in single-level optimization. Our main result resolves this question: SPABA, an adaptation of the PAGE method for nonconvex optimization in (Li et al. , 2021) to the bilevel setting, can achieve optimal sample complexity in both the finite-sum and expectation settings. We show the optimality of SPABA by proving that there is no gap in complexity analysis between stochastic bilevel and single-level optimization when implementing PAGE. Notably, as indicated by the results of (Dagréou et al. , 2022), there might exist a gap in complexity analysis when implementing other stochastic gradient estimators, like SGD and SAGA. In addition to SPABA, we propose several other single-loop stochastic bilevel algorithms, that either match or improve the state-of-the-art sample complexity results, leveraging our convergence rate and complexity analysis. Numerical experiments demonstrate the superior practical performance of the proposed methods.

AAAI Conference 2023 Conference Paper

Decision-Making Context Interaction Network for Click-Through Rate Prediction

  • Xiang Li
  • Shuwei Chen
  • Jian Dong
  • Jin Zhang
  • Yongkang Wang
  • Xingxing Wang
  • Dong Wang

Click-through rate (CTR) prediction is crucial in recommendation and online advertising systems. Existing methods usually model user behaviors, while ignoring the informative context which influences the user to make a click decision, e.g., click pages and pre-ranking candidates that inform inferences about user interests, leading to suboptimal performance. In this paper, we propose a Decision-Making Context Interaction Network (DCIN), which deploys a carefully designed Context Interaction Unit (CIU) to learn decision-making contexts and thus benefits CTR prediction. In addition, the relationship between different decision-making context sources is explored by the proposed Adaptive Interest Aggregation Unit (AIAU) to improve CTR prediction further. In the experiments on public and industrial datasets, DCIN significantly outperforms the state-of-the-art methods. Notably, the model has obtained the improvement of CTR+2.9%/CPM+2.1%/GMV+1.5% for online A/B testing and served the main traffic of Meituan Waimai advertising system.

JBHI Journal 2023 Journal Article

Privacy-Preserving Multi-Source Domain Adaptation for Medical Data

  • Tianyi Han
  • Xiaoli Gong
  • Fan Feng
  • Jin Zhang
  • Zhe Sun
  • Yu Zhang

Great progress has been made in diagnosing medical diseases based on deep learning. Large-scale medical data are expected to improve deep learning performance further. It is almost impossible for a single institution to collect so much data due to the time-consuming and costly collection and labeling of medical data. Many studies have turned attention to data sharing among multiple medical institutions. However, due to different data acquiring and processing procedures, multiple institutions' medical data is characterized by distribution heterogeneity. Besides, the protection of patient privacy in medical data sharing has also been a common concern. To simultaneously address the problems of heterogeneous data distribution and privacy protection, we propose a novel multi-source source free domain adaptation. When aligning distributed heterogeneous data, our method only require to transfer the pre-trained source models rather than the direct source domain data, thus protecting patients' privacy. In addition, it has the advantages of being efficient and less costly in network resources. The proposed method is evaluated on the multi-site fMRI database Autism Brain Imaging Data Exchange (ABIDE) and yields an average accuracy of 69. 37%. We also analyzed its effectiveness on network resource-saving and conducted additional experiments on Camelyon17 to validate the generalization.

AAAI Conference 2023 Conference Paper

Query-Aware Quantization for Maximum Inner Product Search

  • Jin Zhang
  • Defu Lian
  • Haodi Zhang
  • Baoyun Wang
  • Enhong Chen

Maximum Inner Product Search (MIPS) plays an essential role in many applications ranging from information retrieval, recommender systems to natural language processing. However, exhaustive MIPS is often expensive and impractical when there are a large number of candidate items. The state-of-the-art quantization method of approximated MIPS is product quantization with a score-aware loss, developed by assuming that queries are uniformly distributed in the unit sphere. However, in real-world datasets, the above assumption about queries does not necessarily hold. To this end, we propose a quantization method based on the distribution of queries combined with sampled softmax. Further, we introduce a general framework encompassing the proposed method and multiple quantization methods, and we develop an effective optimization for the proposed general framework. The proposed method is evaluated on three real-world datasets. The experimental results show that it outperforms the state-of-the-art baselines.

AAAI Conference 2022 Conference Paper

Anisotropic Additive Quantization for Fast Inner Product Search

  • Jin Zhang
  • Qi Liu
  • Defu Lian
  • Zheng Liu
  • Le Wu
  • Enhong Chen

Maximum Inner Product Search (MIPS) plays an important role in many applications ranging from information retrieval, recommender systems to natural language processing and machine learning. However, exhaustive MIPS is often expensive and impractical when there are a large number of candidate items. The state-of-the-art approximated MIPS is product quantization with a score-aware loss, which weighs more heavily on items with larger inner product scores. However, it is challenging to extend the score-aware loss for additive quantization due to parallel-orthogonal decomposition of residual error. Learning additive quantization with respect to this loss is important since additive quantization can achieve a lower approximation error than product quantization. To this end, we propose a quantization method called Anisotropic Additive Quantization to combine the scoreaware anisotropic loss and additive quantization. To efficiently update the codebooks in this algorithm, we develop a new alternating optimization algorithm. The proposed algorithm is extensively evaluated on three real-world datasets. The experimental results show that it outperforms the stateof-the-art baselines with respect to approximate search accuracy while guaranteeing a similar retrieval efficiency.

NeurIPS Conference 2022 Conference Paper

CUP: Critic-Guided Policy Reuse

  • Jin Zhang
  • Siyuan Li
  • Chongjie Zhang

The ability to reuse previous policies is an important aspect of human intelligence. To achieve efficient policy reuse, a Deep Reinforcement Learning (DRL) agent needs to decide when to reuse and which source policies to reuse. Previous methods solve this problem by introducing extra components to the underlying algorithm, such as hierarchical high-level policies over source policies, or estimations of source policies' value functions on the target task. However, training these components induces either optimization non-stationarity or heavy sampling cost, significantly impairing the effectiveness of transfer. To tackle this problem, we propose a novel policy reuse algorithm called Critic-gUided Policy reuse (CUP), which avoids training any extra components and efficiently reuses source policies. CUP utilizes the critic, a common component in actor-critic methods, to evaluate and choose source policies. At each state, CUP chooses the source policy that has the largest one-step improvement over the current target policy, and forms a guidance policy. The guidance policy is theoretically guaranteed to be a monotonic improvement over the current target policy. Then the target policy is regularized to imitate the guidance policy to perform efficient policy search. Empirical results demonstrate that CUP achieves efficient transfer and significantly outperforms baseline algorithms.

IJCAI Conference 2022 Conference Paper

Learning Coated Adversarial Camouflages for Object Detectors

  • Yexin Duan
  • Jialin Chen
  • Xingyu Zhou
  • Junhua Zou
  • Zhengyun He
  • Jin Zhang
  • Wu Zhang
  • Zhisong Pan

An adversary can fool deep neural network object detectors by generating adversarial noises. Most of the existing works focus on learning local visible noises in an adversarial "patch" fashion. However, the 2D patch attached to a 3D object tends to suffer from an inevitable reduction in attack performance as the viewpoint changes. To remedy this issue, this work proposes the Coated Adversarial Camouflage (CAC) to attack the detectors in arbitrary viewpoints. Unlike the patch trained in the 2D space, our camouflage generated by a conceptually different training framework consists of 3D rendering and dense proposals attack. Specifically, we make the camouflage perform 3D spatial transformations according to the pose changes of the object. Based on the multi-view rendering results, the top-n proposals of the region proposal network are fixed, and all the classifications in the fixed dense proposals are attacked simultaneously to output errors. In addition, we build a virtual 3D scene to fairly and reproducibly evaluate different attacks. Extensive experiments demonstrate the superiority of CAC over the existing attacks, and it shows impressive performance both in the virtual scene and the real world. This poses a potential threat to the security-critical computer vision systems.

EAAI Journal 2021 Journal Article

A fast X-shaped foreground segmentation network with CompactASPP

  • Jin Zhang
  • Shuaihui Wang
  • Junyang Qiu
  • Xinran Pan
  • Junhua Zou
  • Yexin Duan
  • Zhisong Pan
  • Yang Li

Foreground segmentation models are designed to extract moving objects of varying sizes from the background, which can benefit from representations of various scales. As an effective module for capturing multi-scale contexts, Atrous Spatial Pyramid Pooling (ASPP) convolves a final feature representation via multiple parallel atrous convolutions with different dilation rates. However, as the dilation rate increases, ASPP gradually loses its large-scale modeling ability because the sampling of atrous kernel becomes progressively sparse within the receptive field. To solve this problem, we design a CompactASPP module to convolve feature maps compactly. Without significantly increasing the module size, the CompactASPP can encode multi-scale features from all neurons within the receptive field rather than from neurons in several sparsely distributed positions. Furthermore, we leverage CompactASPP modules to enhance our previous X-Net. The proposed Fast X-Net substantially improves the segmentation speed by over 63. 6% and attains new state-of-the-art performances on CDnet2014, SBI2015 and UCSD benchmarks.

NeurIPS Conference 2021 Conference Paper

Towards Gradient-based Bilevel Optimization with Non-convex Followers and Beyond

  • Risheng Liu
  • Yaohua Liu
  • Shangzhi Zeng
  • Jin Zhang

In recent years, Bi-Level Optimization (BLO) techniques have received extensive attentions from both learning and vision communities. A variety of BLO models in complex and practical tasks are of non-convex follower structure in nature (a. k. a. , without Lower-Level Convexity, LLC for short). However, this challenging class of BLOs is lack of developments on both efficient solution strategies and solid theoretical guarantees. In this work, we propose a new algorithmic framework, named Initialization Auxiliary and Pessimistic Trajectory Truncated Gradient Method (IAPTT-GM), to partially address the above issues. In particular, by introducing an auxiliary as initialization to guide the optimization dynamics and designing a pessimistic trajectory truncation operation, we construct a reliable approximate version of the original BLO in the absence of LLC hypothesis. Our theoretical investigations establish the convergence of solutions returned by IAPTT-GM towards those of the original BLO without LLC. As an additional bonus, we also theoretically justify the quality of our IAPTT-GM embedded with Nesterov's accelerated dynamics under LLC. The experimental results confirm both the convergence of our algorithm without LLC, and the theoretical findings under LLC.

JMLR Journal 2020 Journal Article

Discerning the Linear Convergence of ADMM for Structured Convex Optimization through the Lens of Variational Analysis

  • Xiaoming Yuan
  • Shangzhi Zeng
  • Jin Zhang

Despite the rich literature, the linear convergence of alternating direction method of multipliers (ADMM) has not been fully understood even for the convex case. For example, the linear convergence of ADMM can be empirically observed in a wide range of applications arising in statistics, machine learning, and related areas, while existing theoretical results seem to be too stringent to be satisfied or too ambiguous to be checked and thus why the ADMM performs linear convergence for these applications still seems to be unclear. In this paper, we systematically study the local linear convergence of ADMM in the context of convex optimization through the lens of variational analysis. We show that the local linear convergence of ADMM can be guaranteed without the strong convexity of objective functions together with the full rank assumption of the coefficient matrices, or the full polyhedricity assumption of their subdifferential; and it is possible to discern the local linear convergence for various concrete applications, especially for some representative models arising in statistical learning. We use some variational analysis techniques sophisticatedly; and our analysis is conducted in the most general proximal version of ADMM with Fortin and Glowinski's larger step size so that all major variants of the ADMM known in the literature are covered. [abs] [ pdf ][ bib ] &copy JMLR 2020. ( edit, beta )

YNICL Journal 2019 Journal Article

Disturbed neurovascular coupling in type 2 diabetes mellitus patients: Evidence from a comprehensive fMRI analysis

  • Bo Hu
  • Lin-Feng Yan
  • Qian Sun
  • Ying Yu
  • Jin Zhang
  • Yu-Jie Dai
  • Yang Yang
  • Yu-Chuan Hu

BACKGROUND: Previous studies presumed that the disturbed neurovascular coupling to be a critical risk factor of cognitive impairments in type 2 diabetes mellitus (T2DM), but distinct clinical manifestations were lacked. Consequently, we decided to investigate the neurovascular coupling in T2DM patients by exploring the MRI relationship between neuronal activity and the corresponding cerebral blood perfusion. METHODS: Degree centrality (DC) map and amplitude of low-frequency fluctuation (ALFF) map were used to represent neuronal activity. Cerebral blood flow (CBF) map was used to represent cerebral blood perfusion. Correlation coefficients were calculated to reflect the relationship between neuronal activity and cerebral blood perfusion. RESULTS: At the whole gray matter level, the manifestation of neurovascular coupling was investigated by using 4 neurovascular biomarkers. We compared these biomarkers and found no significant changes. However, at the brain region level, neurovascular biomarkers in T2DM patients were significantly decreased in 10 brain regions. ALFF-CBF in left hippocampus and fractional ALFF-CBF in left amygdala were positively associated with the executive function, while ALFF-CBF in right fusiform gyrus was negatively related to the executive function. The disease severity was negatively related to the memory and executive function. The longer duration of T2DM was related to the milder depression, which suggests T2DM-related depression may not be a physiological condition but be a psychological condition. CONCLUSION: Correlations between neuronal activity and cerebral perfusion maps may be a method for detecting neurovascular coupling abnormalities, which could be used for diagnosis in the future. Trial registry number: This study has been registered in ClinicalTrials.gov (NCT02420470) on April 2, 2015 and published on July 29, 2015.

YNIMG Journal 2019 Journal Article

Neurovascular decoupling in type 2 diabetes mellitus without mild cognitive impairment: Potential biomarker for early cognitive impairment

  • Ying Yu
  • Lin-Feng Yan
  • Qian Sun
  • Bo Hu
  • Jin Zhang
  • Yang Yang
  • Yu-Jie Dai
  • Wu-Xun Cui

Type 2 diabetes mellitus (T2DM) is a significant risk factor for mild cognitive impairment (MCI) and the acceleration of MCI to dementia. The high glucose level induce disturbance of neurovascular (NV) coupling is suggested to be one potential mechanism, however, the neuroimaging evidence is still lacking. To assess the NV decoupling pattern in early diabetic status, 33 T2DM without MCI patients and 33 healthy control subjects were prospectively enrolled. Then, they underwent resting state functional MRI and arterial spin labeling imaging to explore the hub-based networks and to estimate the coupling of voxel-wise cerebral blood flow (CBF)-degree centrality (DC), CBF-mean amplitude of low-frequency fluctuation (mALFF) and CBF- mean regional homogeneity (mReHo). We further evaluated the relationship between NV coupling pattern and cognitive performance (false discovery rate corrected). T2DM without MCI patients displayed significant decrease in the absolute CBF-mALFF, CBF-mReHo coupling of CBFnetwork and in the CBF-DC coupling of DCnetwork. Besides, networks which involved CBF and DC hubs mainly located in the default mode network (DMN). Furthermore, less severe disease and better cognitive performance in T2DM patients were significantly correlated with higher coupling of CBF-DC, CBF-mALFF or CBF-mReHo, especially for the cognitive dimensions of general function and executive function. Thus, coupling of CBF-DC, CBF-mALFF and CBF-mReHo may serve as promising indicators to reflect NV coupling state and to explain the T2DM related early cognitive impairment.

ICML Conference 2015 Conference Paper

Preference Completion: Large-scale Collaborative Ranking from Pairwise Comparisons

  • Dohyung Park
  • Joe Neeman
  • Jin Zhang
  • Sujay Sanghavi
  • Inderjit S. Dhillon

In this paper we consider the collaborative ranking setting: a pool of users each provides a set of pairwise preferences over a small subset of the set of d possible items; from these we need to predict each user’s preferences for items s/he has not yet seen. We do so via fitting a rank r score matrix to the pairwise data, and provide two main contributions: (a) We show that an algorithm based on convex optimization provides good generalization guarantees once each user provides as few as O(r \log^2 d) pairwise comparisons — essentially matching the sample complexity required in the related matrix completion setting (which uses actual numerical as opposed to pairwise information), and also matching a lower bound we establish here. (b) We develop a large-scale non-convex implementation, which we call AltSVM, which trains a factored form of the matrix via alternating minimization (which we show reduces to alternating SVM problems), and scales and parallelizes very well to large problem settings. It also outperforms common baselines on many moderately large popular collaborative filtering datasets in both NDCG and other measures of ranking performance.

YNIMG Journal 2004 Journal Article

Optimizing the fMRI data-processing pipeline using prediction and reproducibility performance metrics: I. A preliminary group analysis

  • Stephen Strother
  • Stephen La Conte
  • Lars Kai Hansen
  • Jon Anderson
  • Jin Zhang
  • Sujit Pulapura
  • David Rottenberg

We argue that published results demonstrate that new insights into human brain function may be obscured by poor and/or limited choices in the data-processing pipeline, and review the work on performance metrics for optimizing pipelines: prediction, reproducibility, and related empirical Receiver Operating Characteristic (ROC) curve metrics. Using the NPAIRS split-half resampling framework for estimating prediction/reproducibility metrics (Strother et al. , 2002), we illustrate its use by testing the relative importance of selected pipeline components (interpolation, in-plane spatial smoothing, temporal detrending, and between-subject alignment) in a group analysis of BOLD-fMRI scans from 16 subjects performing a block-design, parametric-static-force task. Large-scale brain networks were detected using a multivariate linear discriminant analysis (canonical variates analysis, CVA) that was tuned to fit the data. We found that tuning the CVA model and spatial smoothing were the most important processing parameters. Temporal detrending was essential to remove low-frequency, reproducing time trends; the number of cosine basis functions for detrending was optimized by assuming that separate epochs of baseline scans have constant, equal means, and this assumption was assessed with prediction metrics. Higher-order polynomial warps compared to affine alignment had only a minor impact on the performance metrics. We found that both prediction and reproducibility metrics were required for optimizing the pipeline and give somewhat different results. Moreover, the parameter settings of components in the pipeline interact so that the current practice of reporting the optimization of components tested in relative isolation is unlikely to lead to fully optimized processing pipelines.

v2026.09.13