Arrow Research search

Author name cluster

Qi Chen

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

49 papers
2 author rows

Possible papers

49

AAAI Conference 2026 Conference Paper

3D-DRES: Detailed 3D Referring Expression Segmentation

  • Qi Chen
  • Changli Wu
  • Jiayi Ji
  • Yiwei Ma
  • Liujuan Cao

Current 3D visual grounding tasks only process sentence-level detection or segmentation, which critically fails to leverage the rich compositional contextual reasonings within natural language expressions. To address this challenge, we introduce Detailed 3D Referring Expression Segmentation (3D-DRES), a new task that provides a phrase to 3D instance mapping, aiming at enhancing fine-grained 3D vision-language understanding. To support 3D-DRES, we present DetailRefer, a new dataset comprising 55,432 descriptions spanning 11,054 distinct objects. Unlike previous datasets, DetailRefer implements a pioneering phrase-instance annotation paradigm where each referenced noun phrase is explicitly mapped to its corresponding 3D elements. Additionally, we introduce DetailBase, a purposefully streamlined yet effective baseline architecture that supports dual-mode segmentation at both sentence and phrase levels. Our experimental results demonstrate that models trained on DetailRefer not only excel at phrase-level segmentation but also show surprising improvements on traditional 3D-RES benchmarks.

YNIMG Journal 2026 Journal Article

Temporal dynamics of flexible cognitive control

  • Chengyuan Wu
  • Carol A. Seger
  • Yixuan Ku
  • Canhuang Luo
  • Ying Zhou
  • Jiefeng Jiang
  • Qi Chen

In dynamic environments, flexible cognitive control adaptively adjusts processing through proactive mechanisms deployed in advance and reactive mechanisms engaged upon conflict. Previous studies have primarily focused on identifying neural networks supporting specific control components, while less is known about how multiple components interact over time to support adaptive control. To characterize these temporal dynamics, we combined electroencephalography (EEG) recordings with a face-word Stroop paradigm under changing conflict environment. A hierarchical Bayesian model was used to estimate trial-wise learning rate, predicted conflict level, and prediction error, providing computational indices of cognitive control flexibility. Neural correlation analysis indicated that these variables correlated with Theta, Alpha, and Beta oscillations in distinct brain regions. Granger causality analyses revealed connectivity patterns among these regions that varied across different task phase. Furthermore, connections reflecting updates to predicted conflict level prior to stimulus onset indexed individual strength in proactive control, while connections reflecting learning rate updates after stimulus onset indexed reactive control. These findings highlight how oscillatory dynamics coordinate multiple control components and provide new insight into how proactive and reactive control emerge as distinct modes within this interconnected neural architecture of flexible cognitive control.

AAAI Conference 2026 Conference Paper

Tracking the Unstable: Appearance-Guided Motion Modeling for Robust Multi-Object Tracking in UAV-Captured Videos

  • Jianbo Ma
  • Hui Luo
  • Qi Chen
  • Yuankai Qi
  • Yumei Sun
  • Amin Beheshti
  • Jianlin Zhang
  • Ming-Hsuan Yang

Multi-object tracking (MOT) aims to track multiple objects while maintaining consistent identities across frames of a given video. In unmanned aerial vehicle (UAV) recorded videos, frequent viewpoint changes and complex UAV-ground relative motion dynamics pose significant challenges, which often lead to unstable affinity measurement and ambiguous association. Existing methods typically model motion and appearance cues separately, overlooking their spatio-temporal interplay and resulting in suboptimal tracking performance. In this work, we propose AMOT, which jointly exploits appearance and motion cues through two key components: an Appearance-Motion Consistency (AMC) matrix and a Motion-aware Track Continuation (MTC) module. Specifically, the AMC matrix computes bi-directional spatial consistency under the guidance of appearance features, enabling more reliable and context-aware identity association. The MTC module complements AMC by reactivating unmatched tracks through appearance-guided predictions that align with Kalman-based predictions, thereby reducing broken trajectories caused by missed detections. Extensive experiments on three UAV benchmarks, including VisDrone2019, UAVDT, and VT-MOT-UAV, demonstrate that our AMOT outperforms current state-of-the-art methods and generalizes well in a plug-and-play and training-free manner.

EAAI Journal 2026 Journal Article

Underwater engineering image enhancement via turbidity prior and multi-scale adaptive

  • Hu Kong
  • Aimin Yuan
  • Qi Chen
  • Jiusheng Liu
  • Hengya Qiu

Underwater engineering images in highly turbid inland rivers are severely affected by color distortion, low contrast, and blurred details, which limit the effectiveness of structural health monitoring. To address these issues, this paper proposes a Turbidity-aware prior and Multi-scale Adaptive Fusion (TMAF) method. Specifically, the method comprises three innovations: (1) an Intelligent Material-Preserving Adaptive (IMPA) color correction algorithm that eliminates color casts while preserving material authenticity through multi-dimensional feature extraction; (2) a Turbidity-Aware Prior (TAP) module based on statistical analysis of 1008 engineering images for accurate transmittance estimation by fusing local contrast, saturation, and edge density features; and (3) a multi-scale adaptive fusion framework integrating USM, CLAHE, and edge-preserving filtering. Experimental validation on the Inland Underwater Engineering Image (IUCE) dataset and the public Underwater Image Enhancement Benchmark (UIEB) dataset demonstrates that TMAF consistently outperforms nine state-of-the-art methods in both qualitative and quantitative evaluations. The method also significantly improves feature point detection and object detection performance, demonstrating superior robustness and generalization capability, particularly under high-turbidity conditions. This method demonstrates practical utility for image enhancement and subsequent AI-based intelligent detection tasks in complex underwater engineering environments.

AAAI Conference 2026 Conference Paper

UniABG: Unified Adversarial View Bridging and Graph Correspondence for Unsupervised Cross-View Geo-Localization

  • Cuiqun Chen
  • Qi Chen
  • Bin Yang
  • Xingyi Zhang

Cross-view geo-localization (CVGL) matches query images (e.g., drone) to geographically corresponding opposite-view imagery (e.g., satellite). While supervised methods achieve strong performance, their reliance on extensive pairwise annotations limits scalability. Unsupervised alternatives avoid annotation costs but suffer from noisy pseudo-labels due to intrinsic cross-view domain gaps. To address these limitations, we propose UniABG, a novel dual-stage unsupervised cross-view geo-localization framework integrating adversarial view bridging with graph-based correspondence calibration. Our approach first employs View-Aware Adversarial Bridging (VAAB) to model view-invariant features and enhance pseudo-label robustness. Subsequently, Heterogeneous Graph Filtering Calibration (HGFC) refines cross-view associations by constructing dual inter-view structure graphs, achieving reliable view correspondence. Extensive experiments demonstrate state-of-the-art unsupervised performance, showing that UniABG improves Satellite → Drone AP by +10.63% on University-1652 and +16.73% on SUES-200, even surpassing supervised baselines.

EAAI Journal 2025 Journal Article

A camouflage target classification method based on spectral difference enhancement and pixel-pair features in land-based hyperspectral images

  • Jiale Zhao
  • Dan Fang
  • Jiaju Ying
  • Yudan Chen
  • Qi Chen
  • Qianghui Wang
  • Guanglong Wang
  • Bing Zhou

Hyperspectral images are capable of capturing rich spatial and spectral information of targets, rendering them particularly valuable for camouflage target detection and classification applications. However, as camouflage technologies continue to advance, the spectral similarity between camouflaged targets and their backgrounds has become increasingly pronounced, presenting significant challenges for camouflaged target classification in hyperspectral imagery. To overcome this challenge, this paper introduces a novel land-based hyperspectral image classification approach for camouflaged targets, termed Spectral Difference Enhancement and Pixel-Pair Features (SDE-PPF). The proposed methodology initially conducts spectral rearrangement of the hyperspectral image based on target spectral characteristics, which induces oscillatory patterns in the background spectrum. Subsequently, first-order spectral differentiation coupled with nonlinear processing is applied to the rearranged hyperspectral image to effectively amplify subtle spectral differences between targets and backgrounds, thereby improving their discriminability. Following spectral enhancement, the method constructs pixel pairs from the processed hyperspectral image and employs a convolutional neural network to extract pixel-pair features. Network parameters are optimized through comprehensive analysis of pixel-pair sample relationships. During testing, the central pixel is systematically paired with its neighboring pixels, and classification is performed using the trained model. Ultimately, the final classification of each central pixel is determined through a voting mechanism that consolidates all classification results. Comprehensive experiments were performed on four distinct land-based hyperspectral image datasets containing camouflage targets. The experimental results demonstrate that the proposed SDE-PPF method outperforms conventional hyperspectral image classification approaches, achieving remarkable average classification accuracies of 98. 46 %, 99. 05 %, 98. 94 %, and 99. 21 % for detecting camouflage targets against grassland, barren Grassland, withered leaf, and shrubbery backgrounds, respectively. This innovative approach establishes an effective and robust technical solution for camouflage target classification and detection, exhibiting considerable potential for diverse practical applications.

ICRA Conference 2025 Conference Paper

A2DO: Adaptive Anti-Degradation Odometry with Deep Multi-Sensor Fusion for Autonomous Navigation

  • Hui Lai
  • Qi Chen
  • Junping Zhang
  • Jian Pu

Accurate localization is essential for the safe and effective navigation of autonomous vehicles, and Simultaneous Localization and Mapping (SLAM) is a cornerstone technology in this context. However, The performance of the SLAM system can deteriorate under challenging conditions such as low light, adverse weather, or obstructions due to sensor degradation. We present A2DO, a novel end-to-end multi-sensor fusion odometry system that enhances robustness in these scenarios through deep neural networks. A2DO integrates LiDAR and visual data, employing a multilayer, multi-scale feature encoding module augmented by an attention mechanism to mitigate sensor degradation dynamically. The system is pretrained extensively on simulated datasets covering a broad range of degradation scenarios and fine-tuned on a curated set of real-world data, ensuring robust adaptation to complex scenarios. Our experiments demonstrate that A2DO maintains superior localization accuracy and robustness across various degradation conditions, showcasing its potential for practical implementation in autonomous vehicle systems.

NeurIPS Conference 2025 Conference Paper

Are Pixel-Wise Metrics Reliable for Computerized Tomography Reconstruction?

  • Tianyu Lin
  • Xinran Li
  • Chuntung Zhuang
  • Qi Chen
  • Yuanhao Cai
  • Kai Ding
  • Alan Yuille
  • Zongwei Zhou

Widely adopted evaluation metrics for sparse-view CT reconstruction, such as Structural Similarity Index Measure and Peak Signal-to-Noise Ratio, prioritize pixel-wise fidelity but often fail to capture the completeness of critical anatomical structures, particularly small or thin regions that are easily missed. To address this limitation, we propose a suite of novel anatomy-aware evaluation metrics designed to assess structural completeness across anatomical structures, including large organs, small organs, intestines, and vessels. Building on these metrics, we introduce CARE, a Completeness-Aware Reconstruction Enhancement framework that incorporates structural penalties during training to encourage anatomical preservation of significant structures. CARE is model-agnostic and can be seamlessly integrated into analytical, implicit, and generative methods. When applied to these methods, CARE substantially improves structural completeness in CT reconstructions, achieving up to 32% improvement for large organs, 22% for small organs, 40% for intestines, and 36% for vessels.

AAAI Conference 2025 Conference Paper

Attention-Driven GUI Grounding: Leveraging Pretrained Multimodal Large Language Models Without Fine-Tuning

  • Hai-Ming Xu
  • Qi Chen
  • Lei Wang
  • Lingqiao Liu

Recent advancements in Multimodal Large Language Models (MLLMs) have generated significant interest in their ability to autonomously interact with and interpret Graphical User Interfaces (GUIs). A major challenge in these systems is grounding—accurately identifying critical GUI components such as text or icons based on a GUI image and a corresponding text query. Traditionally, this task has relied on fine-tuning MLLMs with specialized training data to predict component locations directly. However, in this paper, we propose a novel Tuning-free Attention-driven Grounding (TAG) method that leverages the inherent attention patterns in pretrained MLLMs to accomplish this task without the need for additional fine-tuning. Our method involves identifying and aggregating attention maps from specific tokens within a carefully constructed query prompt. Applied to MiniCPM-Llama3-V 2.5, a state-of-the-art MLLM, our tuning-free approach achieves performance comparable to tuning-based methods, with notable success in text localization. Additionally, we demonstrate that our attention map-based grounding technique significantly outperforms direct localization predictions from MiniCPM-Llama3-V 2.5, highlighting the potential of using attention maps from pretrained MLLMs and paving the way for future innovations in this domain.

ICRA Conference 2025 Conference Paper

Efficient Scale-Uniform 3D Visual Coverage Algorithm for UAV Based on Elastic Photogrammetric Constraints

  • Jianping Zong
  • Zhongzhi Cao
  • Qi Chen
  • Chuanyu Sun
  • Xiuli Shao
  • Haifeng Li 0008
  • Hongpeng Wang

Unmanned aerial vehicles equipped with modern vision algorithms are crucial for missions such as reconstruction and target acquisition. However, when deployed in the field, undulating terrain can cause significant fluctuations in image scale and degrade the performance of vision algorithms. Instead of developing specialized image processing schemes with limited adaptability, this paper presents a novel 3D visual coverage algorithm that is compatible with existing generic vision algorithms and maintains a uniform image scale for ground targets. In detail, photogrammetric constraints are initially introduced to generate aerial waypoints, and then the negative effects of valley clustering are addressed. Elastic Photogrammetric Constraints (EPC) are further proposed to eliminate valley clustering effects induced by saddle terrain. The experimental results demonstrate that EPC reduces the traversal path length by up to 37. 38 % compared to the previous work, but with a minor trade-off in scale variations.

AAAI Conference 2025 Conference Paper

Enhancing Large Language Model Performance with Gradient-Based Parameter Selection

  • Haoling Li
  • Xin Zhang
  • Xiao Liu
  • Yeyun Gong
  • Yifan Wang
  • Qi Chen
  • Peng Cheng

Large language models (LLMs) have revolutionized numerous fields of research, driving significant advancements in natural language processing, machine translation, and beyond. Although the extensive number of parameters contributes a lot to the great success, existing studies indicate that not all model parameters hold equal importance, which further leads to redundancy during the parameter update process. Recent works for reducing redundant parameter updates for LLMs either lack task-specific data information, may leading to suboptimal model performance, or discard transformer components or insignificant parameters, limiting the model's scalability across different tasks and potentially compromising the LLM structure. To address these issues and further enhance the performance of LLMs, we propose Gradient-Mask Tuning (GMT), a method that selectively updates parameters based on gradient information, which is specific to the target tasks. Specifically, after calculating gradients during back propagation, we measure their absolute values and mask those with small absolute values. Our empirical results in various training paradigms like SFT and DPO for various domains of tasks demonstrate that GMT not only preserves the original network structure but also enhances the potential performance of LLMs. Further analysis indicates that GMT exhibits insensitivity to mask ratio and possesses computational efficiency comparable to vanilla training approach.

ICLR Conference 2025 Conference Paper

Generalization in VAE and Diffusion Models: A Unified Information-Theoretic Analysis

  • Qi Chen
  • Jierui Zhu
  • Florian Shkurti

Despite the empirical success of Diffusion Models (DMs) and Variational Autoencoders (VAEs), their generalization performance remains theoretically underexplored, especially lacking a full consideration of the shared encoder-generator structure. Leveraging recent information-theoretic tools, we propose a unified theoretical framework that provides guarantees for the generalization of both the encoder and generator by treating them as randomized mappings. This framework further enables (1) a refined analysis for VAEs, accounting for the generator's generalization, which was previously overlooked; (2) illustrating an explicit trade-off in generalization terms for DMs that depends on the diffusion time $T$; and (3) providing computable bounds for DMs based solely on the training data, allowing the selection of the optimal $T$ and the integration of such bounds into the optimization process to improve model performance. Empirical results on both synthetic and real datasets illustrate the validity of the proposed theory.

AAAI Conference 2025 Conference Paper

IPDN: Image-enhanced Prompt Decoding Network for 3D Referring Expression Segmentation

  • Qi Chen
  • Changli Wu
  • Jiayi Ji
  • Yiwei Ma
  • Danni Yang
  • Xiaoshuai Sun

3D Referring Expression Segmentation (3D-RES) aims to segment point cloud scenes based on a given expression. However, existing 3D-RES approaches face two major challenges: feature ambiguity and intent ambiguity. Feature ambiguity arises from information loss or distortion during point cloud acquisition due to limitations such as lighting and viewpoint. Intent ambiguity refers to the model's equal treatment of all queries during the decoding process, lacking top-down task-specific guidance. In this paper, we introduce an Image-enhanced Prompt Decoding Network (IPDN), which leverages multi-view images and task-driven information to enhance the model's reasoning capabilities. To address feature ambiguity, we propose the Multi-view Semantic Embedding (MSE) module, which injects multi-view 2D image information into the 3D scene and compensates for potential spatial information loss. To tackle intent ambiguity, we designed a Prompt-Aware Decoder (PAD) that guides the decoding process by deriving task-driven signals from the interaction between the expression and visual features. Comprehensive experiments demonstrate that IPDN outperforms the state-of-the-art by 1.9 and 4.2 points in mIoU metrics on the 3D-RES and 3D-GRES tasks, respectively.

IJCAI Conference 2025 Conference Paper

Localizing Before Answering: A Benchmark for Grounded Medical Visual Question Answering

  • Dung Nguyen
  • Minh Khoi Ho
  • Huy Ta
  • Thanh Tam Nguyen
  • Qi Chen
  • Kumar Rav
  • Quy Duong Dang
  • Satwik Ramchandre

Medical Large Multi-modal Models (LMMs) have demonstrated remarkable capabilities in medical data interpretation. However, these models frequently generate hallucinations contradicting source evidence, particularly due to inadequate localization reasoning. This work reveals a critical limitation in current medical LMMs: instead of analyzing relevant pathological regions, they often rely on linguistic patterns or attend to irrelevant image areas when responding to disease-related queries. To address this, we introduce HEAL-MedVQA (Hallucination Evaluation via Localization MedVQA), a comprehensive benchmark designed to evaluate LMMs' localization abilities and hallucination robustness. HEAL-MedVQA features (i) two innovative evaluation protocols to assess visual and textual shortcut learning, and (ii) a dataset of 67K VQA pairs, with doctor-annotated anatomical segmentation masks for pathological regions. To improve visual reasoning, we propose the Localize-before-Answer (LobA) framework, which trains LMMs to localize target regions of interest and self-prompt to emphasize segmented pathological areas, generating grounded and reliable answers. Experimental results demonstrate that our approach significantly outperforms state-of-the-art biomedical LMMs on the challenging HEAL-MedVQA benchmark, advancing robustness in medical VQA.

TMLR Journal 2025 Journal Article

Meta-Learning Adaptive Loss Functions

  • Christian Raymond
  • Qi Chen
  • Bing Xue
  • Mengjie Zhang

Loss function learning is a new meta-learning paradigm that aims to automate the essential task of designing a loss function for a machine learning model. Existing techniques for loss function learning have shown promising results, often improving a model's training dynamics and final inference performance. However, a significant limitation of these techniques is that the loss functions are meta-learned in an offline fashion, where the meta-objective only considers the very first few steps of training, which is a significantly shorter time horizon than the one typically used for training deep neural networks. This causes significant bias towards loss functions that perform well at the very start of training but perform poorly at the end of training. To address this issue we propose a new loss function learning technique for adaptively updating the loss function online after each update to the base model parameters. The experimental results show that our proposed method consistently outperforms the cross-entropy loss and offline loss function learning techniques on a diverse range of neural network architectures and datasets.

IJCAI Conference 2025 Conference Paper

Multi-Hierarchical Fine-Grained Feature Mapping Driven by Feature Contributions for Molecular Odor Prediction

  • Hongxin Xie
  • Jiande Sun
  • Fanfu Xue
  • Zifei Han
  • Shanshan Feng
  • Qi Chen

Molecular odor prediction involves using a molecule's structure to estimate its odor. While accurate prediction remains challenging, AI models can suggest potential odors. Existing methods, however, often rely on basic descriptors or handcrafted fingerprints, which lack expressive power and hinder effective learning. Furthermore, these methods suffer from severe class imbalance, limiting the training effectiveness of AI models. To address these challenges, we propose a Feature Contribution-driven Hierarchical Multi-Feature Mapping Network (HMFNet). Specifically, we introduce a fine-grained, Local Multi-Hierarchy Feature Extraction module (LMFE) that performs deep feature extraction at the atomic level, capturing detailed features crucial for odor prediction. To enhance the extraction of discriminative atomic features, we integrate a Harmonic Modulated Feature Mapping (HMFM). This module dynamically learns feature importance and frequency modulation, improving the model's capability to capture relevant patterns. Additionally, a Global Multi-Hierarchy Feature Extraction module (GMFE) is designed to learn global features from the molecular graph topology, enabling the model to fully leverage global information and enhance its discriminative power for odor prediction. To further mitigate the issue of class imbalance, we propose a Chemically-Informed Loss (CIL). Experimental results demonstrate that our approach significantly improves performance across various deep learning models, highlighting its potential to advance molecular structure representation and accelerate the development of AI-driven technologies.

NeurIPS Conference 2025 Conference Paper

PanTS: The Pancreatic Tumor Segmentation Dataset

  • Wenxuan Li
  • Xinze Zhou
  • Qi Chen
  • Tianyu Lin
  • Pedro R. A. S. Bassi
  • Xiaoxi Chen
  • Chen Ye
  • Zheren Zhu

PanTS is a large-scale, multi-institutional dataset curated to advance research in pancreatic CT analysis. It contains 36, 390 CT scans from 145 medical centers, with expert-validated, voxel-wise annotations of over 993, 000 anatomical structures, covering pancreatic tumors, pancreas head, body, and tail, and 24 surrounding anatomical structures such as vascular/skeletal structures and abdominal/thoracic organs. Each scan includes metadata such as patient age, sex, diagnosis, contrast phase, in-plane spacing, slice thickness, etc. AI models trained on PanTS achieve significantly better performance in pancreatic tumor detection, localization, and segmentation than those trained on existing public datasets. Our analysis indicates that these gains are directly attributable to the 16× larger-scale tumor annotations and indirectly supported by the 24 additional surrounding anatomical structures. As the largest and most comprehensive resource of its kind, PanTS offers a new benchmark for developing and evaluating AI models in pancreatic CT analysis.

NeurIPS Conference 2025 Conference Paper

RetrievalAttention: Accelerating Long-Context LLM Inference via Vector Retrieval

  • Di Liu
  • Meng Chen
  • Baotong Lu
  • Huiqiang Jiang
  • Zhenhua Han
  • Qianxi Zhang
  • Qi Chen
  • Chengruidong Zhang

Transformer-based Large Language Models (LLMs) have become increasingly important. However, scaling LLMs to longer contexts incurs slow inference speed and high GPU memory consumption for caching key-value (KV) vectors. This paper presents RetrievalAttention, a training-free approach to both accelerate the decoding phase and reduce GPU memory consumption by pre-building KV vector indexes for fixed contexts and maintaining them in CPU memory for efficient retrieval. Unlike conventional KV cache methods, RetrievalAttention integrate approximate nearest neighbor search (ANNS) indexes into attention computation. We observe that off-the-shelf ANNS techniques often fail due to the out-of-distribution (OOD) nature of query and key vectors in attention mechanisms. RetrievalAttention overcomes this with an attention-aware vector index. Our evaluation shows RetrievalAttention achieves near full attention accuracy while accessing only 1-3\% of the data, significantly reducing inference costs. Remarkably, RetrievalAttention enables LLMs with 8B parameters to handle 128K tokens on a single NVIDIA RTX4090 (24GB), achieving a decoding speed of 0. 107 seconds per token.

TMLR Journal 2025 Journal Article

Seeing Beyond Labels: Source-Free Domain Adaptation via Hypothesis Consolidation of Prediction Rationale

  • Yangyang Shu
  • Yuhang Liu
  • Xiaofeng Cao
  • Qi Chen
  • Bowen Zhang
  • Ziqin Zhou
  • Anton van den Hengel
  • Lingqiao Liu

Source-Free Unsupervised Domain Adaptation (SFUDA) is a challenging task where a model needs to be adapted to a new domain without access to target domain labels or source domain data. The primary difficulty in this task is that the model's predictions may be inaccurate, and using these inaccurate predictions for model adaptation can lead to misleading results. To address this issue, this paper proposes a novel approach that considers multiple prediction hypotheses for each sample and investigates the rationale behind each hypothesis. By consolidating these hypothesis rationales, we identify the most likely correct hypotheses, which we then use as a pseudo-labeled set to support a semi-supervised learning procedure for model adaptation. This approach distinguishes itself from conventional semi-supervised learning by relying solely on pseudo-labels rather than ground-truth annotations. To achieve the optimal performance, we propose a three-step adaptation process: model pre-adaptation, hypothesis consolidation, and semi-supervised learning. Extensive experimental results demonstrate that our approach achieves state-of-the-art performance in the SFUDA task and can be easily integrated into existing approaches to improve their performance. The codes are available at \url{https://github.com/GANPerf/HCPR}.

IJCAI Conference 2025 Conference Paper

SketchAgent: Generating Structured Diagrams from Hand-Drawn Sketches

  • Cheng Tan
  • Qi Chen
  • Jingxuan Wei
  • Gaowei Wu
  • Zhangyang Gao
  • Siyuan Li
  • Bihui Yu
  • Ruifeng Guo

Hand-drawn sketches are a natural and efficient medium for capturing and conveying ideas. Despite significant advancements in controllable natural image generation, translating freehand sketches into structured, machine-readable diagrams remains a labor-intensive and predominantly manual task. The primary challenge stems from the inherent ambiguity of sketches, which lack the structural constraints and semantic precision required for automated diagram generation. To address this challenge, we introduce SketchAgent, a multi-agent system designed to automate the transformation of hand-drawn sketches into structured diagrams. SketchAgent integrates sketch recognition, symbolic reasoning, and iterative validation to produce semantically coherent and structurally accurate diagrams, significantly reducing the need for manual effort. To evaluate the effectiveness of our approach, we propose the Sketch2Diagram Benchmark, a comprehensive dataset and evaluation framework encompassing eight diverse diagram categories, such as flowcharts, directed graphs, and model architectures. The dataset comprises over 6, 000 high-quality examples with token-level annotations, standardized preprocessing, and rigorous quality control. By streamlining the diagram generation process, SketchAgent holds great promise for applications in design, education, and engineering, while offering a significant step toward bridging the gap between intuitive sketching and machine-readable diagram generation.

IJCAI Conference 2025 Conference Paper

TSTAI: A Time-varying Brain Effective Connectivity Network Construction Method Combining with Brain Active Information

  • Qi Chen
  • Zhiqiong Wang
  • Jiaxin Li
  • Jinying Tao
  • Junchang Xin

More accurate construction of brain effective conncetivity networks remains a great challenge to achieve accurate auxiliary diagnosis of brain diseases and in-depth exploration of brain function. However, existing methods only consider higher-order or non-stationary assumptions, rather than simultaneously constructing higher-order and non-stationary networks. Among many existing methods, Bayesian network methods demonstrate superior network structure learning ability. In this work, the forward-backward search (FBS) method is optimized by using brain active information, which is improved to a higher-order network structure learning method, called TSTAI. Firstly, in the process of non-stationary network structure learning, two-stage idea is used to search the change points. Then, in the process of learning higher-order network structure, FBS method is combined with two kinds of brain active information to improve the condition set filtering process and scoring function, respectively. Finally, the pruning strategy is used to reduce the search space. Extensive experiments on simulated and real data demonstrate the effectiveness of TSTAI. Through experiments, the TSTAI is compared with state-of-the-art higher-order network construction methods, and the proposed method achieves an improvement of 3. 6% and 17. 4% respectively in the network construction accuracy.

AAAI Conference 2025 Conference Paper

VQTalker: Towards Multilingual Talking Avatars Through Facial Motion Tokenization

  • Tao Liu
  • Ziyang Ma
  • Qi Chen
  • Feilong Chen
  • Shuai Fan
  • Xie Chen
  • Kai Yu

We present VQTalker, a Vector Quantization-based framework for multilingual talking head generation that addresses the challenges of lip synchronization and natural motion across diverse languages. Our approach is grounded in the phonetic principle that human speech comprises a finite set of distinct sound units (phonemes) and corresponding visual articulations (visemes), which often share commonalities across languages. We introduce a facial motion tokenizer based on Group Residual Finite Scalar Quantization (GRFSQ), which creates a discretized representation of facial features. This method enables comprehensive capture of facial movements while improving generalization to multiple languages, even with limited training data. Building on this quantized representation, we implement a coarse-to-fine motion generation process that progressively refines facial animations. Extensive experiments demonstrate that VQTalker achieves state-of-the-art performance in both video-driven and speech-driven scenarios, particularly in multilingual settings. Notably, our method achieves high-quality results at a resolution of 512 × 512 pixels while maintaining a lower bitrate of approximately 11 kbps. Our work opens new possibilities for cross-lingual talking face generation.

AAAI Conference 2024 Conference Paper

3D-STMN: Dependency-Driven Superpoint-Text Matching Network for End-to-End 3D Referring Expression Segmentation

  • Changli Wu
  • Yiwei Ma
  • Qi Chen
  • Haowei Wang
  • Gen Luo
  • Jiayi Ji
  • Xiaoshuai Sun

In 3D Referring Expression Segmentation (3D-RES), the earlier approach adopts a two-stage paradigm, extracting segmentation proposals and then matching them with referring expressions. However, this conventional paradigm encounters significant challenges, most notably in terms of the generation of lackluster initial proposals and a pronounced deceleration in inference speed. Recognizing these limitations, we introduce an innovative end-to-end Superpoint-Text Matching Network (3D-STMN) that is enriched by dependency-driven insights. One of the keystones of our model is the Superpoint-Text Matching (STM) mechanism. Unlike traditional methods that navigate through instance proposals, STM directly correlates linguistic indications with their respective superpoints, clusters of semantically related points. This architectural decision empowers our model to efficiently harness cross-modal semantic relationships, primarily leveraging densely annotated superpoint-text pairs, as opposed to the more sparse instance-text pairs. In pursuit of enhancing the role of text in guiding the segmentation process, we further incorporate the Dependency-Driven Interaction (DDI) module to deepen the network's semantic comprehension of referring expressions. Using the dependency trees as a beacon, this module discerns the intricate relationships between primary terms and their associated descriptors in expressions, thereby elevating both the localization and segmentation capacities. Comprehensive experiments on the ScanRefer benchmark reveal that our model not only sets new performance standards, registering an mIoU gain of 11.7 points but also achieves a staggering enhancement in inference speed, surpassing traditional methods by 95.7 times. The code and models are available at https://github.com/sosppxo/3D-STMN.

AAAI Conference 2024 Conference Paper

CREAD: A Classification-Restoration Framework with Error Adaptive Discretization for Watch Time Prediction in Video Recommender Systems

  • Jie Sun
  • Zhaoying Ding
  • Xiaoshuang Chen
  • Qi Chen
  • Yincheng Wang
  • Kaiqiao Zhan
  • Ben Wang

The watch time is a significant indicator of user satisfaction in video recommender systems. However, the prediction of watch time as a target variable is often hindered by its highly imbalanced distribution with a scarcity of observations for larger target values and over-populated samples for small values. State-of-the-art watch time prediction models discretize the continuous watch time into a set of buckets in order to consider the distribution of watch time. However, it is highly uninvestigated how these discrete buckets should be created from the continuous watch time distribution, and existing discretization approaches suffer from either a large learning error or a large restoration error. To address this challenge, we propose a Classification-Restoration framework with Error-Adaptive-Discretization (CREAD) to accurately predict the watch time. The proposed framework contains a discretization module, a classification module, and a restoration module. It predicts the watch time through multiple classification problems. The discretization process is a key contribution of the CREAD framework. We theoretically analyze the impacts of the discretization on the learning error and the restoration error, and then propose the error-adaptive discretization (EAD) technique to better balance the two errors, which achieves better performance over traditional discretization approaches. We conduct detailed offline evaluations on a public dataset and an industrial dataset, both showing performance gains through the proposed approach. Moreover, We have fully launched our framework to an online video platform, which resulted in a significant increase in users' video watch time by 0.29% through A/B testing. These results highlight the effectiveness of the CREAD framework in watch time prediction in video recommender systems.

NeurIPS Conference 2024 Conference Paper

RG-SAN: Rule-Guided Spatial Awareness Network for End-to-End 3D Referring Expression Segmentation

  • Changli Wu
  • Qi Chen
  • Jiayi Ji
  • Haowei Wang
  • Yiwei Ma
  • You Huang
  • Hao Fei
  • Xiaoshuai Sun

3D Referring Expression Segmentation (3D-RES) aims to segment 3D objects by correlating referring expressions with point clouds. However, traditional approaches frequently encounter issues like over-segmentation or mis-segmentation, due to insufficient emphasis on spatial information of instances. In this paper, we introduce a Rule-Guided Spatial Awareness Network (RG-SAN) by utilizing solely the spatial information of the target instance for supervision. This approach enables the network to accurately depict the spatial relationships among all entities described in the text, thus enhancing the reasoning capabilities. The RG-SAN consists of the Text-driven Localization Module (TLM) and the Rule-guided Weak Supervision (RWS) strategy. The TLM initially locates all mentioned instances and iteratively refines their positional information. The RWS strategy, acknowledging that only target objects have supervised positional information, employs dependency tree rules to precisely guide the core instance’s positioning. Extensive testing on the ScanRefer benchmark has shown that RG-SAN not only establishes new performance benchmarks, with an mIoU increase of 5. 1 points, but also exhibits significant improvements in robustness when processing descriptions with spatial ambiguity. All codes are available at https: //github. com/sosppxo/RG-SAN.

ICRA Conference 2024 Conference Paper

STT: Stateful Tracking with Transformers for Autonomous Driving

  • Longlong Jing
  • Ruichi Yu
  • Xu Chen
  • Zhengli Zhao
  • Shiwei Sheng
  • Colin Graber
  • Qi Chen
  • Qinru Li

Tracking objects in three-dimensional space is critical for autonomous driving. To ensure safety while driving, the tracker must be able to reliably track objects across frames and accurately estimate their states such as velocity and acceleration in the present. Existing works frequently focus on the association task while either neglecting the model’s performance on state estimation or deploying complex heuristics to predict the states. In this paper, we propose STT, a Stateful Tracking model built with Transformers, that can consistently track objects in the scenes while also predicting their states accurately. STT consumes rich appearance, geometry, and motion signals through long term history of detections and is jointly optimized for both data association and state estimation tasks. Since the standard tracking metrics like MOTA and MOTP do not capture the combined performance of the two tasks in the wider spectrum of object states, we extend them with new metrics called S-MOTA and MOTP S that address this limitation. STT achieves competitive real-time performance on the Waymo Open Dataset.

NeurIPS Conference 2024 Conference Paper

Towards Understanding Evolving Patterns in Sequential Data

  • Qiuhao Zeng
  • Long-Kai Huang
  • Qi Chen
  • Charles Ling
  • Boyu Wang

In many machine learning tasks, data is inherently sequential. Most existing algorithms learn from sequential data in an auto-regressive manner, which predicts the next unseen data point based on the observed sequence, implicitly assuming the presence of an \emph{evolving pattern} embedded in the data that can be leveraged. However, identifying and assessing evolving patterns in learning tasks often relies on subjective judgments rooted in the prior knowledge of human experts, lacking a standardized quantitative measure. Furthermore, such measures enable us to determine the suitability of employing sequential models effectively and make informed decisions on the temporal order of time series data, and feature/data selection processes. To address this issue, we introduce the Evolving Rate (EvoRate), which quantitatively approximates the intensity of evolving patterns in the data with Mutual Information. Furthermore, in some temporal data with neural mutual information estimations, we only have snapshots at different timestamps, lacking correspondence, which hinders EvoRate estimation. To tackle this challenge, we propose EvoRate$_\mathcal{W}$, aiming to establish correspondence with optimal transport for estimating the first-order EvoRate. Experiments on synthetic and real-world datasets including images and tabular data validate the efficacy of our EvoRate.

NeurIPS Conference 2024 Conference Paper

Weak-eval-Strong: Evaluating and Eliciting Lateral Thinking of LLMs with Situation Puzzles

  • Qi Chen
  • Bowen Zhang
  • Gang Wang
  • Qi Wu

While advancements in NLP have significantly improved the performance of Large Language Models (LLMs) on tasks requiring vertical thinking, their lateral thinking capabilities remain under-explored and challenging to measure due to the complexity of assessing creative thought processes and the scarcity of relevant data. To address these challenges, we introduce SPLAT, a benchmark leveraging Situation Puzzles to evaluate and elicit LAteral Thinking of LLMs. This benchmark, containing 975 graded situation puzzles across three difficulty levels, employs a new multi-turn player-judge framework instead of the traditional model-based evaluation, which often necessitates a stronger evaluation model. This framework simulates an interactive game where the model (player) asks the evaluation model (judge) questions about an incomplete story to infer the full scenario. The judge answers based on a detailed reference scenario or evaluates if the player's predictions align with the reference one. This approach lessens dependence on more robust evaluation models, enabling the assessment of state-of-the-art LLMs. The experiments demonstrate that a robust evaluation model, such as WizardLM-2, closely matches human judgements in both intermediate question-answering and final scenario accuracy, achieving over 80% agreement--similar to the agreement levels among humans. Furthermore, applying data and reasoning processes from our benchmark to other lateral thinking-related benchmarks, e. g. , RiddleSense and BrainTeaser, leads to performance enhancements. This suggests that our benchmark effectively evaluates and elicits the lateral thinking abilities of LLMs.

AAAI Conference 2024 Conference Paper

WebVLN: Vision-and-Language Navigation on Websites

  • Qi Chen
  • Dileepa Pitawela
  • Chongyang Zhao
  • Gengze Zhou
  • Hsiang-Ting Chen
  • Qi Wu

Vision-and-Language Navigation (VLN) task aims to enable AI agents to accurately understand and follow natural language instructions to navigate through real-world environments, ultimately reaching specific target locations. We recognise a promising opportunity to extend VLN to a comparable navigation task that holds substantial significance in our daily lives, albeit within the virtual realm: navigating websites on the Internet. This paper proposes a new task named Vision-and-Language Navigation on Websites (WebVLN), where we use question-based instructions to train an agent, emulating how users naturally browse websites. Unlike the existing VLN task that only pays attention to vision and instruction (language), the WebVLN agent further considers underlying web-specific content like HTML, which could not be seen on the rendered web pages yet contain rich visual and textual information. Toward this goal, we contribute a dataset, WebVLN-v1, and introduce a novel approach called Website-aware VLN Network (WebVLN-Net), which is built upon the foundation of state-of-the-art VLN techniques. Experimental results show that WebVLN-Net outperforms current VLN and web-related navigation methods. We believe that the introduction of the newWebVLN task and its dataset will establish a new dimension within the VLN domain and contribute to the broader vision-and-language research community. Code is available at: https://github.com/WebVLN/WebVLN.

YNIMG Journal 2023 Journal Article

Alpha oscillations encode Bayesian belief updating underlying attentional allocation in dynamic environments

  • Siying Li
  • Carol A. Seger
  • Jianfeng Zhang
  • Meng Liu
  • Wenshan Dong
  • Wanting Liu
  • Qi Chen

In a dynamic environment, expectations of the future constantly change based on updated evidence and affect the dynamic allocation of attention. To further investigate the neural mechanisms underlying attentional expectancies, we employed a modified Central Cue Posner Paradigm in which the probability of cues being valid (that is, accurately indicated the upcoming target location) was manipulated. Attentional deployment to the cued location (α), which was governed by precision of predictions on previous trials, was estimated using a hierarchical Bayesian model and was included as a regressor in the analyses of electrophysiological (EEG) data. Our results revealed that before the target appeared, alpha oscillations (8∼13 Hz) for high-predictability cues (88 % valid) were significantly predicted by precision-dependent attention (α). This relationship was not observed under low-predictability conditions (69 % and 50 % valid cues). After the target appeared, precision-dependent attention (α) correlated with alpha band oscillations only in the valid cue condition and not in the invalid condition. Further analysis under conditions of significant attentional modulation by precision suggested a separate effect of cue orientation. These results provide new insights on how trial-by-trial Bayesian belief updating relates to alpha band encoding of environmentally-sensitive allocation of visual spatial attention.

NeurIPS Conference 2023 Conference Paper

Model-enhanced Vector Index

  • Hailin Zhang
  • Yujing Wang
  • Qi Chen
  • Ruiheng Chang
  • Ting Zhang
  • Ziming Miao
  • Yingyan Hou
  • Yang Ding

Embedding-based retrieval methods construct vector indices to search for document representations that are most similar to the query representations. They are widely used in document retrieval due to low latency and decent recall performance. Recent research indicates that deep retrieval solutions offer better model quality, but are hindered by unacceptable serving latency and the inability to support document updates. In this paper, we aim to enhance the vector index with end-to-end deep generative models, leveraging the differentiable advantages of deep retrieval models while maintaining desirable serving efficiency. We propose Model-enhanced Vector Index (MEVI), a differentiable model-enhanced index empowered by a twin-tower representation model. MEVI leverages a Residual Quantization (RQ) codebook to bridge the sequence-to-sequence deep retrieval and embedding-based models. To substantially reduce the inference time, instead of decoding the unique document ids in long sequential steps, we first generate some semantic virtual cluster ids of candidate documents in a small number of steps, and then leverage the well-adapted embedding vectors to further perform a fine-grained search for the relevant documents in the candidate virtual clusters. We empirically show that our model achieves better performance on the commonly used academic benchmarks MSMARCO Passage and Natural Questions, with comparable serving latency to dense retrieval solutions.

NeurIPS Conference 2023 Conference Paper

On the Stability-Plasticity Dilemma in Continual Meta-Learning: Theory and Algorithm

  • Qi Chen
  • Changjian Shui
  • Ligong Han
  • Mario Marchand

We focus on Continual Meta-Learning (CML), which targets accumulating and exploiting meta-knowledge on a sequence of non-i. i. d. tasks. The primary challenge is to strike a balance between stability and plasticity, where a model should be stable to avoid catastrophic forgetting in previous tasks and plastic to learn generalizable concepts from new tasks. To address this, we formulate the CML objective as controlling the average excess risk upper bound of the task sequence, which reflects the trade-off between forgetting and generalization. Based on the objective, we introduce a unified theoretical framework for CML in both static and shifting environments, providing guarantees for various task-specific learning algorithms. Moreover, we first present a rigorous analysis of a bi-level trade-off in shifting environments. To approach the optimal trade-off, we propose a novel algorithm that dynamically adjusts the meta-parameter and its learning rate w. r. t environment change. Empirical evaluations on synthetic and real datasets illustrate the effectiveness of the proposed theory and algorithm.

IJCAI Conference 2023 Conference Paper

Self-Supervised Neuron Segmentation with Multi-Agent Reinforcement Learning

  • Yinda Chen
  • Wei Huang
  • Shenglong Zhou
  • Qi Chen
  • Zhiwei Xiong

The performance of existing supervised neuron segmentation methods is highly dependent on the number of accurate annotations, especially when applied to large scale electron microscopy (EM) data. By extracting semantic information from unlabeled data, self-supervised methods can improve the performance of downstream tasks, among which the mask image model (MIM) has been widely used due to its simplicity and effectiveness in recovering original information from masked images. However, due to the high degree of structural locality in EM images, as well as the existence of considerable noise, many voxels contain little discriminative information, making MIM pretraining inefficient on the neuron segmentation task. To overcome this challenge, we propose a decision-based MIM that utilizes reinforcement learning (RL) to automatically search for optimal image masking ratio and masking strategy. Due to the vast exploration space, using single-agent RL for voxel prediction is impractical. Therefore, we treat each input patch as an agent with a shared behavior policy, allowing for multi-agent collaboration. Furthermore, this multi-agent model can capture dependencies between voxels, which is beneficial for the downstream segmentation task. Experiments conducted on representative EM datasets demonstrate that our approach has a significant advantage over alternative self-supervised methods on the task of neuron segmentation. Code is available at https: //github. com/ydchen0806/dbMiM.

NeurIPS Conference 2022 Conference Paper

A Neural Corpus Indexer for Document Retrieval

  • Yujing Wang
  • Yingyan Hou
  • Haonan Wang
  • Ziming Miao
  • Shibin Wu
  • Qi Chen
  • Yuqing Xia
  • Chengmin Chi

Current state-of-the-art document retrieval solutions mainly follow an index-retrieve paradigm, where the index is hard to be directly optimized for the final retrieval target. In this paper, we aim to show that an end-to-end deep neural network unifying training and indexing stages can significantly improve the recall performance of traditional methods. To this end, we propose Neural Corpus Indexer (NCI), a sequence-to-sequence network that generates relevant document identifiers directly for a designated query. To optimize the recall performance of NCI, we invent a prefix-aware weight-adaptive decoder architecture, and leverage tailored techniques including query generation, semantic document identifiers, and consistency-based regularization. Empirical studies demonstrated the superiority of NCI on two commonly used academic benchmarks, achieving +21. 4% and +16. 8% relative enhancement for Recall@1 on NQ320k dataset and R-Precision on TriviaQA dataset, respectively, compared to the best baseline method.

YNICL Journal 2022 Journal Article

Imbalance between the caudate and putamen connectivity in obsessive–compulsive disorder

  • Ziwen Peng
  • Tingxin He
  • Ping Ren
  • Lili Jin
  • Qiong Yang
  • Chuanyong Xu
  • Rongzhen Wen
  • Jierong Chen

BACKGROUND: Compulsive behaviors in obsessive-compulsive disorder (OCD) have been suggested to result from an imbalance in cortico-striatal connectivity. However, the nature of this impairment, the relative involvement of different striatal areas, their imbalance in genetically related but unimpaired individuals, and their relationship with cognitive dysfunction in OCD patients, remain unknown. METHODS: In the current study, striatal (i.e., caudate and putamen) whole-brain connectivity was computed in a sample of OCD patients (OCD, n = 62), unaffected first-degree relatives (UFDR, n = 53) and healthy controls (HC, n = 73) by ROI-based resting-state functional magnetic resonance imaging (rs-fMRI). A behavioral task switch paradigm outside of the scanner was also performed to measure cognitive flexibility in OCD patients. RESULTS: There were significantly increased strengths (Z-transformed Pearson correlation coefficient) in caudate connectivity in OCD patients. A significant correlation between the two types of connectivity strengths in the relevant regions was observed only in the OCD patient group. Furthermore, the caudate connectivity of patients was negatively associated with their task-switch performance. CONCLUSIONS: The imbalance between the caudate and putamen connectivity, arising from the abnormal increase of caudate activity, may serve as a clinical characteristic for obsessive-compulsive disorder.

NeurIPS Conference 2022 Conference Paper

Learning Distinct and Representative Modes for Image Captioning

  • Qi Chen
  • Chaorui Deng
  • Qi Wu

Over the years, state-of-the-art (SoTA) image captioning methods have achieved promising results on some evaluation metrics (e. g. , CIDEr). However, recent findings show that the captions generated by these methods tend to be biased toward the "average" caption that only captures the most general mode (a. k. a, language pattern) in the training corpus, i. e. , the so-called mode collapse problem. Affected by it, the generated captions are limited in diversity and usually less informative than natural image descriptions made by humans. In this paper, we seek to avoid this problem by proposing a Discrete Mode Learning (DML) paradigm for image captioning. Our innovative idea is to explore the rich modes in the training caption corpus to learn a set of "mode embeddings", and further use them to control the mode of the generated captions for existing image captioning models. Specifically, the proposed DML optimizes a dual architecture that consists of an image-conditioned discrete variational autoencoder (CdVAE) branch and a mode-conditioned image captioning (MIC) branch. The CdVAE branch maps each image caption to one of the mode embeddings stored in a learned codebook, and is trained with a pure non-autoregressive generation objective to make the modes distinct and representative. The MIC branch can be simply modified from an existing image captioning model, where the mode embedding is added to the original word embeddings as the control signal. In the experiments, we apply the proposed DML to two widely used image captioning models, Transformer and AoANet. The results show that the learned mode embedding successfully facilitates these models to generate high-quality image captions with different modes, further leading to better performance for both diversity and quality on the MS COCO dataset.

NeurIPS Conference 2022 Conference Paper

On Learning Fairness and Accuracy on Multiple Subgroups

  • Changjian Shui
  • Gezheng Xu
  • Qi Chen
  • Jiaqi Li
  • Charles X. Ling
  • Tal Arbel
  • Boyu Wang
  • Christian Gagné

We propose an analysis in fair learning that preserves the utility of the data while reducing prediction disparities under the criteria of group sufficiency. We focus on the scenario where the data contains multiple or even many subgroups, each with limited number of samples. As a result, we present a principled method for learning a fair predictor for all subgroups via formulating it as a bilevel objective. Specifically, the subgroup specific predictors are learned in the lower-level through a small amount of data and the fair predictor. In the upper-level, the fair predictor is updated to be close to all subgroup specific predictors. We further prove that such a bilevel objective can effectively control the group sufficiency and generalization error. We evaluate the proposed framework on real-world datasets. Empirical evidence suggests the consistently improved fair predictions, as well as the comparable accuracy to the baselines.

ICML Conference 2022 Conference Paper

Optimization-Induced Graph Implicit Nonlinear Diffusion

  • Qi Chen
  • Yifei Wang 0001
  • Yisen Wang 0001
  • Jiansheng Yang
  • Zhouchen Lin

Due to the over-smoothing issue, most existing graph neural networks can only capture limited dependencies with their inherently finite aggregation layers. To overcome this limitation, we propose a new kind of graph convolution, called Graph Implicit Nonlinear Diffusion (GIND), which implicitly has access to infinite hops of neighbors while adaptively aggregating features with nonlinear diffusion to prevent over-smoothing. Notably, we show that the learned representation can be formalized as the minimizer of an explicit convex optimization objective. With this property, we can theoretically characterize the equilibrium of our GIND from an optimization perspective. More interestingly, we can induce new structural variants by modifying the corresponding optimization objective. To be specific, we can embed prior properties to the equilibrium, as well as introducing skip connections to promote training stability. Extensive experiments show that GIND is good at capturing long-range dependencies, and performs well on both homophilic and heterophilic graphs with nonlinear diffusion. Moreover, we show that the optimization-induced variants of our models can boost the performance and improve training stability and efficiency as well. As a result, our GIND obtains significant improvements on both node-level and graph-level tasks.

UAI Conference 2022 Conference Paper

Sublinear time algorithms for greedy selection in high dimensions

  • Qi Chen
  • Kai Liu 0001
  • Ruilong Yao
  • Hu Ding 0003

Greedy selection is a widely used idea for solving many machine learning problems. But greedy selection algorithms often have high complexities and thus may be prohibitive for large-scale data. In this paper, we consider two fundamental optimization problems in machine learning: k-center clustering and convex hull approximation, where they both can be solved via greedy selection. We propose sublinear time algorithms for them through combining the strategies of randomization and greedy selection. Our results are similar in spirit to the linear time stochastic greedy selection algorithms for submodular maximization, but with several important differences. Our runtimes are independent of the number of input data items n. In particular, our runtime for k-center clustering significantly improves upon that of the uniform sampling approach, especially when the dimensionality is high. Our sublinear algorithms can also reduce the computational complexities for various applications, such as data selection and compression, active learning, and topic modeling, etc.

YNICL Journal 2021 Journal Article

Aberrant rich club organization in patients with obsessive-compulsive disorder and their unaffected first-degree relatives

  • Ziwen Peng
  • Xinyi Yang
  • Chuanyong Xu
  • Xiangshu Wu
  • Qiong Yang
  • Zhen Wei
  • Zihan Zhou
  • Tom Verguts

Recent studies suggested that the rich club organization promoting global brain communication and integration of information, may be abnormally increased in obsessive-compulsive disorder (OCD). However, the structural and functional basis of this organization is still not very clear. Given the heritability of OCD, as suggested by previous family-based studies, we hypothesize that aberrant rich club organization may be a trait marker for OCD. In the present study, 32 patients with OCD, 30 unaffected first-degree relatives (FDR) and 32 healthy controls (HC) underwent diffusion tensor imaging (DTI) and functional magnetic resonance imaging (fMRI). We examined the structural rich club organization and its interrelationship with functional coupling. Our results showed that rich club and peripheral connection strength in patients with OCD was lower than in HC, while it was intermediate in FDR. Finally, the coupling between structural and functional connections of the rich club, was decreased in FDR but not in OCD relative to HC, which suggests a buffering mechanism of brain functions in FDR. Overall, our findings suggest that alteration of the rich club organization may reflect a vulnerability biomarker for OCD, possibly buffered by structural and functional coupling of the rich club.

NeurIPS Conference 2021 Conference Paper

Generalization Bounds For Meta-Learning: An Information-Theoretic Analysis

  • Qi Chen
  • Changjian Shui
  • Mario Marchand

We derive a novel information-theoretic analysis of the generalization property of meta-learning algorithms. Concretely, our analysis proposes a generic understanding in both the conventional learning-to-learn framework \citep{amit2018meta} and the modern model-agnostic meta-learning (MAML) algorithms \citep{finn2017model}. Moreover, we provide a data-dependent generalization bound for the stochastic variant of MAML, which is \emph{non-vacuous} for deep few-shot learning. As compared to previous bounds that depend on the square norms of gradients, empirical validations on both simulated data and a well-known few-shot benchmark show that our bound is orders of magnitude tighter in most conditions.

NeurIPS Conference 2021 Conference Paper

PolarStream: Streaming Object Detection and Segmentation with Polar Pillars

  • Qi Chen
  • Sourabh Vora
  • Oscar Beijbom

Recent works recognized lidars as an inherently streaming data source and showed that the end-to-end latency of lidar perception models can be reduced significantly by operating on wedge-shaped point cloud sectors rather then the full point cloud. However, due to use of cartesian coordinate systems these methods represent the sectors as rectangular regions, wasting memory and compute. In this work we propose using a polar coordinate system and make two key improvements on this design. First, we increase the spatial context by using multi-scale padding from neighboring sectors: preceding sector from the current scan and/or the following sector from the past scan. Second, we improve the core polar convolutional architecture by introducing feature undistortion and range stratified convolutions. Experimental results on the nuScenes dataset show significant improvements over other streaming based methods. We also achieve comparable results to existing non-streaming methods but with lower latencies.

NeurIPS Conference 2021 Conference Paper

SPANN: Highly-efficient Billion-scale Approximate Nearest Neighborhood Search

  • Qi Chen
  • Bing Zhao
  • Haidong Wang
  • Mingqin Li
  • ChuanJie Liu
  • Zengzhong Li
  • Mao Yang
  • Jingdong Wang

The in-memory algorithms for approximate nearest neighbor search (ANNS) have achieved great success for fast high-recall search, but are extremely expensive when handling very large scale database. Thus, there is an increasing request for the hybrid ANNS solutions with small memory and inexpensive solid-state drive (SSD). In this paper, we present a simple but efficient memory-disk hybrid indexing and search system, named SPANN, that follows the inverted index methodology. It stores the centroid points of the posting lists in the memory and the large posting lists in the disk. We guarantee both disk-access efficiency (low latency) and high recall by effectively reducing the disk-access number and retrieving high-quality posting lists. In the index-building stage, we adopt a hierarchical balanced clustering algorithm to balance the length of posting lists and augment the posting list by adding the points in the closure of the corresponding clusters. In the search stage, we use a query-aware scheme to dynamically prune the access of unnecessary posting lists. Experiment results demonstrate that SPANN is 2X faster than the state-of-the-art ANNS solution DiskANN to reach the same recall quality 90% with same memory cost in three billion-scale datasets. It can reach 90% recall@1 and recall@10 in just around one millisecond with only about 10% of original memory cost. Code is available at: https: //github. com/microsoft/SPTAG.

AAAI Conference 2021 Conference Paper

StrokeGAN: Reducing Mode Collapse in Chinese Font Generation via Stroke Encoding

  • Jinshan Zeng
  • Qi Chen
  • Yunxin Liu
  • Mingwen Wang
  • Yuan Yao

The generation of stylish Chinese fonts is an important problem involved in many applications. Most of existing generation methods are based on the deep generative models, particularly, the generative adversarial networks (GAN) based models. However, these deep generative models may suffer from the mode collapse issue, which significantly degrades the diversity and quality of generated results. In this paper, we introduce a one-bit stroke encoding to capture the key mode information of Chinese characters and then incorporate it into CycleGAN, a popular deep generative model for Chinese font generation. As a result we propose an efficient method called StrokeGAN, mainly motivated by the observation that the stroke encoding contains amount of mode information of Chinese characters. In order to reconstruct the one-bit stroke encoding of the associated generated characters, we introduce a stroke-encoding reconstruction loss imposed on the discriminator. Equipped with such one-bit stroke encoding and stroke-encoding reconstruction loss, the mode collapse issue of CycleGAN can be significantly alleviated, with an improved preservation of strokes and diversity of generated characters. The effectiveness of StrokeGAN is demonstrated by a series of generation tasks over nine datasets with different fonts. The numerical results demonstrate that StrokeGAN generally outperforms the state-of-the-art methods in terms of content and recognition accuracies, as well as certain stroke error, and also generates more realistic characters.

NeurIPS Conference 2020 Conference Paper

Every View Counts: Cross-View Consistency in 3D Object Detection with Hybrid-Cylindrical-Spherical Voxelization

  • Qi Chen
  • Lin Sun
  • Ernest Cheung
  • Alan L. Yuille

Recent voxel-based 3D object detectors for autonomous vehicles learn point cloud representations either from bird eye view (BEV) or range view (RV, a. k. a. the perspective view). However, each view has its own strengths and weaknesses. In this paper, we present a novel framework to unify and leverage the benefits from both BEV and RV. The widely-used cuboid-shaped voxels in Cartesian coordinate system only benefit learning BEV feature map. Therefore, to enable learning both BEV and RV feature maps, we introduce Hybrid-Cylindrical-Spherical voxelization. Our findings show that simply adding detection on another view as auxiliary supervision will lead to poor performance. We proposed a pair of cross-view transformers to transform the feature maps into the other view and introduce cross-view consistency loss on them. Comprehensive experiments on the challenging NuScenes Dataset validate the effectiveness of our proposed method by virtue of joint optimization and complementary information on both views. Remarkably, our approach achieved mAP of 55. 8%, outperforming all published approaches by at least 3% in overall performance and up to 16. 5% in safety-crucial categories like cyclist.

NeurIPS Conference 2019 Conference Paper

NAT: Neural Architecture Transformer for Accurate and Compact Architectures

  • Yong Guo
  • Yin Zheng
  • Mingkui Tan
  • Qi Chen
  • Jian Chen
  • Peilin Zhao
  • Junzhou Huang

Designing effective architectures is one of the key factors behind the success of deep neural networks. Existing deep architectures are either manually designed or automatically searched by some Neural Architecture Search (NAS) methods. However, even a well-searched architecture may still contain many non-significant or redundant modules or operations (e. g. , convolution or pooling), which may not only incur substantial memory consumption and computation cost but also deteriorate the performance. Thus, it is necessary to optimize the operations inside an architecture to improve the performance without introducing extra computation cost. Unfortunately, such a constrained optimization problem is NP-hard. To make the problem feasible, we cast the optimization problem into a Markov decision process (MDP) and seek to learn a Neural Architecture Transformer (NAT) to replace the redundant operations with the more computationally efficient ones (e. g. , skip connection or directly removing the connection). Based on MDP, we learn NAT by exploiting reinforcement learning to obtain the optimization policies w. r. t. different architectures. To verify the effectiveness of the proposed strategies, we apply NAT on both hand-crafted architectures and NAS based architectures. Extensive experiments on two benchmark datasets, i. e. , CIFAR-10 and ImageNet, demonstrate that the transformed architecture by NAT significantly outperforms both its original form and those architectures optimized by existing methods.

EAAI Journal 2017 Journal Article

A double-region learning algorithm for counting the number of pedestrians in subway surveillance videos

  • Gaoqi He
  • Qi Chen
  • Dongxu Jiang
  • Xingjian Lu
  • Yubo Yuan

Counting pedestrians in surveillance videos has become an urgent safety concern in critical areas. However, surveillance videos of subway spaces suffer from severe crowd occlusion and perspective distortion. In this paper, a novel double-region learning algorithm is presented to overcome these challenges. The main idea of this algorithm is to identify the best two-region boundary and then design a reasonable pedestrian-counting method in each separated region. First, a separate line is obtained via possibility learning, and each frame is divided into a nearby region and a distant region to eliminate the influence of perspective distortion. Second, in the nearby region, we apply the improved aggregate channel feature detection to count the number of pedestrians N 1. In the distant region, we employ the Extreme Learning Machine and Gaussian Process regression methods to estimate the number of pedestrians N 2. Finally, the total number of pedestrians in each frame can be obtained with high accuracy according to N 1 and N 2. We establish a subway pedestrian video dataset about several typical subway stations in Shanghai to validate the algorithm performance. Various experimental results demonstrate that the accuracy of the proposed approach surpasses that of compared methods, which means that our algorithm can meet the management requirements of subway stations.

IROS Conference 2015 Conference Paper

Maintaining constant towing tension between cable ship and burying system under sea waves by hybrid FUZZY P + ID controller

  • Qi Chen
  • Wei Li 0006
  • Xiaohui Wang 0014
  • Yan Li
  • Shuo Li 0001
  • Bin Xian

In this paper, we propose a hybrid FUZZY P + ID controller to stabilize the towing cable tension between a cable ship and a burying system. First, we develop the model of a winch system driven by valve-controlled hydraulic motors and evaluate the step responses yielded by the conventional PID and the proposed FUZZY P+ID controllers using simulations. The comparative studies show that the control performance yielded by the FUZZY P+ID controller is superior. We replace the existing PID controller implemented on the towing winch with the FUZZY P + ID controller for burying cable tasks at speed up to 1 knot under sea waves with significant variations (peak-to-peak) from 1. 5 to 2. 5 meters. The real applications demonstrate that the FUZZY P + ID controller is much more robust than the conventional PID controller.

EAAI Journal 1997 Journal Article

Artificial-neural-network-based fast valving control in a power-generation system

  • Yingduo Han
  • Zonghong Wang
  • Qi Chen
  • Shaohua Tan

This paper presents an artificial-neural-network-based controller to realize fast valving in a power-generation plant. A backpropagation algorithm is used to train the feedforward neural-network controller. The hardware implementation and the test results of the controller on a physical pilot-scale power system set-up are described in detail. Compared with some conventional fast valving methods applied to the same system, test results (both in a computer simulation and on a physical pilot-scale power system set-up) show that the neural-network controller has quite satisfactory generalisation capability, feasibility and reliability, as well as accuracy.

v2026.09.13