Arrow Research search

Author name cluster

Xin Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

172 papers
2 author rows

Possible papers

172

EAAI Journal 2026 Journal Article

A high-precision and efficient method for coal–rock characteristic identification utilizing coal wall temperature field

  • Futao Li
  • Zhongbin Wang
  • Dong Wei
  • Xin Li
  • Lei Si
  • Jinheng Gu
  • Jialiang Dai

Coal-rock characteristic identification is a crucial technology for realizing shearers intelligent control. To enhance the intelligence level of shearers, this paper presents a novel approach for identifying coal–rock characteristics using the temperature field of the coal wall. First, we introduce an enhanced You Only Look Once (YOLO) model, termed Temperature sensitive region-YOLO (TSR-YOLO), specifically designed to extract temperature-sensitive regions within the coal wall temperature field. In terms of structural design, TSR-YOLO innovatively incorporates the Cross Stage Partial FasterNet (C3k2-FasterNet) into the backbone network to accelerate feature extraction and devises the Cross-Stage Partial Kolmogorov–Arnold Network (C3k2-KAN) to enhance detailed feature representation. In the bottleneck network, it integrates the Dynamic Convolution (DynamicConv) module to capture broader and more complex feature, as well as the Variational Overlapping Vision-Generalized Spatial Cross-Stage Partial (VoV-GSCSP) module to enhance computational efficiency and optimize feature extraction performance. Subsequently, we propose a coal–rock characteristic identification method utilizing the ConvNeXt model. To validate its effectiveness, we conduct ablation and comparative experiments using experimental data obtained from infrared images of the coal wall during the shearer cutting process. The results indicate that proposed approach achieves a mean Average Precision (mAP) reaching 98. 1% and an inference speed of 2 ms per image in identifying temperature-sensitive regions of the coal wall temperature field, surpassing other comparative models. Furthermore, the accuracy of coal–rock property identification reaches 97. 6%. This study presents a new approach to coal–rock characteristic identification methods.

EAAI Journal 2026 Journal Article

A multi-view collaborative heterogeneous graph neural network with semantic- and relation-aware for drug-disease association prediction

  • Jinzhou Wu
  • Donglin He
  • Xin Li
  • Rui Wang
  • Yujuan Zhang

Drug repositioning (DR) is crucial for accelerating drug development and reducing costs; computational methods offer efficient alternatives to costly traditional methods. However, existing methods rely on static linear operations for integrating multi-source similarity networks, causing noise accumulation and redundancy. Additionally, meta-path aggregation often fails to dynamically balance intra-path interactions and cross-path heterogeneity, limiting multi-granularity information fusion. To address these, we propose a Multi-view Collaborative Heterogeneous Graph Neural Network with Semantic and Relation-Aware (MSRHGNN) for Drug-Disease Association (DDA) prediction. MSRHGNN jointly encodes similarity and heterogeneous biological network features. It first uses an adaptive dynamic fusion mechanism to integrate multi-source similarity data, leveraging a Graph Transformer to capture richer structural features. Second, within the heterogeneous biological network, the low-order relational view aggregates first-order neighborhood information to capture local topology, while the high-order relational view designs a dual attention collaborative aggregation mechanism: node-level attention driven by central anchor points highlights key interactions within paths, and semantic and relation-aware mechanisms at the cross-path level quantify the consistency within paths and heterogeneity between paths, achieving complementary dynamic fusion. Additionally, MSRHGNN aligns node- and graph-level representations through a multi-view contrastive learning strategy and employs a multi-task balance strategy to alleviate gradient conflicts between the main and auxiliary tasks. Experimental results demonstrate that our method outperforms other baseline models across multiple evaluation metrics on several public datasets, with its stability and robustness validated under cold-start, imbalanced, and noisy conditions. Furthermore, case studies on specific diseases and molecular docking experiments highlight its potential value in practical drug discovery.

TMLR Journal 2026 Journal Article

Accurate Split Learning on Noisy Signals

  • Hang Xu
  • Subhajit Maity
  • Aritra Dutta
  • Xin Li
  • Panos Kalnis

Noise injection is applied in Split Learning to address privacy concerns about data leakage. Previous work protects Split Learning by adding noise to intermediate results during the forward pass. Unfortunately, noisy signals significantly degrade the accuracy of Split Learning training. This paper focuses on improving the training accuracy of Split Learning in the presence of noisy signals while protecting training data from reconstruction attacks. We propose two denoising techniques, namely scaling and random masking. Our theoretical results show that both of our denoising techniques accurately estimate the intermediate variables during the forward pass of Split Learning. Moreover, our experiments with deep neural networks demonstrate that the proposed denoising approaches allow Split Learning to tolerate high noise levels while achieving almost the same accuracy as the noise-free baseline. Interestingly, we show that, after applying our denoising techniques, the resulting network is more resilient to a state-of-the-art attack than the simple noise-injection approach. Our code is publicly available at: github.com/MaitySubhajit/AccurateSL.

EAAI Journal 2026 Journal Article

An end-to-end wavelet-based irregular transformer with gumbel sampling for spatiotemporal welding prediction

  • Changhui Liu
  • Ke Jin
  • Jianzhi Sun
  • Lin Deng
  • Jiewu Leng
  • Xin Li
  • Qian Li
  • Qingcheng Yang

Welding prediction plays a vital role in ensuring assembly precision and minimizing rework in thin-walled structures. Data-driven approaches have attracted increasing attention; however, most existing methods rely on oversimplified input representations, overlook irregular temporal dynamics, and focus solely on single-variable prediction, limiting their applicability in complex welding scenarios. Motivated by these challenges, this study develops an end-to-end Wavelet Irregular Transformer with Gumbel Sampling, designed to achieve accurate spatiotemporal prediction of welding-induced deformation and residual stress. From the artificial intelligence perspective, the model incorporates a large language model-based embedding initializer that compresses and contextualizes step-level simulation parameters, and an adaptive parameterized Gumbel keyframe extractor that dynamically identifies the most informative temporal segments. This design enables efficient learning over ultra-long welding sequences while maintaining high-fidelity temporal representations. From the engineering application perspective, a channel-aware wavelet encoder–decoder is developed to fuse multi-frequency and multi-channel features, improving spatial coherence and capturing coupled stress–strain interactions. Validation on a dedicated thin-plate welding dataset, supplemented by physical experiments, shows that the proposed method achieves superior accuracy, robustness, and computational efficiency compared with optimized encoder–decoder and sequence-modeling baselines. The proposed Wavelet Irregular Transformer with Gumbel Sampling achieves a deformation mean absolute error of 0. 033 mm and a root mean square error of 0. 045 mm on the test set, reducing the deformation mean absolute error by 92. 0% and the root mean square error by 82. 1% compared with the best uniform-sampling baseline, while requiring 76. 7% fewer billion floating-point operations than a full-sequence Transformer.

AAAI Conference 2026 Conference Paper

Beyond Euclidean Assumptions: Geometry-Aware Adaptive Routing for Remote Sensing Segmentation

  • Jie Qiu
  • Dizuo Cao
  • Linwei Dai
  • Xin Li
  • Fan Yang
  • Dong Yu
  • Changying Wang
  • Zongheng Wen

Remote sensing imagery poses a distinct challenge for semantic segmentation due to its inherent fractal complexity and the diversity of geometric structures present in real-world geospatial scenes. Euclidean-based models typically assume spatial uniformity; however, such assumptions often break down when confronted with objects exhibiting markedly different structural characteristics—such as roads versus vegetation—thereby complicating the feature representation process. Hyperbolic space offers a theoretically grounded alternative for modeling such hierarchical and heterogeneous patterns, yet fully replacing Euclidean geometry incurs significant computational overhead. We therefore introduce Geometry-Aware Adaptive Routing (GAAR), a novel module that facilitates geometry-aware routing by dynamically allocating high-level features to either Euclidean or Hyperbolic subspaces through a learnable binary gating mechanism, informed by structural priors learned during training. To further promote routing stability and geometric consistency, we introduce Geometry-Aware Deterministic Regularization (GADR), a regularization strategy that encourages confident, structure-aligned assignments. GAAR is plug-and-play and integrates seamlessly into existing segmentation architectures. Experiments on three challenging Remote Sensing Image Semantic Segmentation (RSISS) benchmarks demonstrate that our approach consistently outperforms state-of-the-art (SOTA) methods, particularly in geometrically complex regions, offering a scalable and effective solution to the limitations of purely Euclidean modeling.

EAAI Journal 2026 Journal Article

Diffusion enhancement-based dual-domain fusion attention motor imagery electroencephalography classification network

  • Xin Li
  • Ji-Hong Wang
  • Ji An
  • Lang-Li Ren
  • Li-Feng Bian
  • Chen Yang

Motor imagery electroencephalography (MI-EEG) signals suffer from low signal-to-noise ratio, strong non-stationarity, and high inter-subject variability—a key obstacle in brain–computer interface (BCI) research. To address small-sample feature extraction challenges and high signal complexity, we propose a Diffusion-augmented Dual-domain Fusion Network (DDF-Net). Specifically, we employ a denoising diffusion probabilistic model separately to pre-train the network and expand MI-EEG datasets; meanwhile, DDF-Net itself introduces frequency-space and temporal dual attention to capture μ/β rhythm representations, and fuses deep spatiotemporal residual and shallow convolutional branches to balance global dependencies and local details. On BCI Competition IV-2a/IV-2b datasets, DDF-Net achieves 90. 66 %/91. 31 % classification accuracy and improved Cohen's κ scores vs. state-of-the-art methods. This confirms its superior performance, stability, and generalizability, providing an effective solution for clinical MI-EEG decoding. Its code and pre-trained weights are available at https: //github. com/HyperSystemAndImageProc/DDF-Net.

JBHI Journal 2026 Journal Article

EAP-LSTM: A Bi-LSTM-Based Deep Learning Framework for Quantitatively Predicting Enhancer Activity in Drosophila and Human Cell Lines

  • Yao Zhang
  • Lichang Dai
  • Yu Dou
  • Xin Li
  • Chang Lu
  • Hao Wu

Enhancer activity plays a critical role in gene regulation, influencing various biological processes such as development and disease progression. Accurate prediction of enhancer activity is essential for understanding the mechanisms underlying gene regulation and enhancer function. This study introduces a novel deep learning framework, EAP-LSTM (Enhancer Activity Prediction based on Bi-LSTM), to quantitatively predict enhancer activity across different species and cell lines. The model integrates multiple feature modules, including Word2Vec-based representations of DNA sequences, reverse complement k-mer, mismatch k-mer features, and epigenomic data. Evaluated on six cell lines, including five human cell lines (A549, HCT116, HepG2, K562, and MCF-7) and one Drosophila cell line (S2), EAP-LSTM consistently outperforms state-of-the-art models, such as DeepSTARR and HEAP, in all datasets. For example, on the K562 dataset, EAP-LSTM achieves a Pearson correlation coefficient (PCC) of 0. 7944, outperforming DeepSTARR and HEAP by 13. 65% and 2. 73%, respectively. In addition, EAP-LSTM demonstrates strong performance in small-sample learning scenarios, showing clear improvements compared with baseline models. Furthermore, the study investigates the role of transcription factor binding sites (TFBSs) within enhancer regions, identifying critical motifs associated with enhancer activity. These findings not only improve enhancer prediction accuracy but also provide valuable insights into the molecular mechanisms underlying enhancer function.

AAAI Conference 2026 Conference Paper

FIA-Edit: Frequency-Interactive Attention for Efficient and High-Fidelity Inversion-Free Text-Guided Image Editing

  • Kaixiang Yang
  • Boyang Shen
  • Xin Li
  • Yuchen Dai
  • Yuxuan Luo
  • Yueran Ma
  • Wei Fang
  • Qiang Li

Text-guided image editing has advanced rapidly with the rise of diffusion models. While flow-based inversion-free methods offer high efficiency by avoiding latent inversion, they often fail to effectively integrate source information, leading to poor background preservation, spatial inconsistencies, and over-editing due to the lack of effective integration of source information. In this paper, we present FIA-Edit, a novel inversion-free framework that achieves high-fidelity and semantically precise edits through a Frequency-Interactive Attention. Specifically, we design two key components: (1) a Frequency Representation Interaction (FRI) module that enhances cross-domain alignment by exchanging frequency components between source and target features within self-attention, and (2) a Feature Injection (FIJ) module that explicitly incorporates source-side queries, keys, values, and text embeddings into the target branch's cross-attention to preserve structure and semantics. Comprehensive and extensive experiments demonstrate that FIA-Edit supports high-fidelity editing at low computational cost (~6s per 512 * 512 image on an RTX 4090) and consistently outperforms existing methods across diverse tasks in visual quality, background fidelity, and controllability. Furthermore, we are the first to extend text-guided image editing to clinical applications. By synthesizing anatomically coherent hemorrhage variations in surgical images, FIA-Edit opens new opportunities for medical data augmentation and delivers significant gains in downstream bleeding classification.

AAAI Conference 2026 Conference Paper

From Tokens to Latent States: Leveraging Pre-trained Language Models for Improving Partially Observable Reinforcement Learning

  • Meiju Li
  • Ruixiang Sun
  • Xin Li
  • Mingzhong Wang

Partially observable Markov decision processes (POMDPs) present significant challenges for reinforcement learning, as agents must learn optimal policies while maintaining belief states over unobserved environment states based on partial observations. We observe a compelling analogy: large language models (LLMs) autoregressively generate token probability distributions based on preceding context, mirroring how belief states are maintained and updated in POMDPs. This insight motivates leveraging the rich prior knowledge embedded in pre-trained LLMs for latent states estimation from observation-action histories. However, two critical challenges emerge: on the one hand, modality misalignment prevents LLMs from directly encoding visual observations and discrete actions; on the other hand, semantic misalignment exists between observation-action sequences and token sequences. To address these challenges, we introduce a novel framework ELSLLM that employs a Johnson-Lindenstrauss projection (JLP) module to transform input dimensions while preserving state similarity with theoretical guarantees, and utilizes modern Hopfield networks (MHN) to store all word embeddings from pre-trained LLMs as a knowledge repository. Through retrieval and querying mechanisms, ELSLLM achieves token-level knowledge alignment without requiring fine-tuning of the pre-trained LLMs. Extensive experiments on partially observable environments demonstrate that ELSLLM achieves state-of-the-art performance, significantly outperforming baseline methods with and without LSTM memory mechanisms. Our work opens new avenues for integrating pre-trained LLMs with reinforcement learning in partially observable settings.

AAAI Conference 2026 Conference Paper

Generative Branching for Mixed-Integer Linear Programming

  • Ruobing Wang
  • Xin Li
  • Yangchuan Wang
  • Zijian Zhang
  • Mingzhong Wang

Branch-and-bound (B&B) is a fundamental algorithmic framework for solving Mixed-Integer Linear Programming (MILP) problems, where branching decisions critically affect solver efficiency. Recent learning-based methods apply imitation learning to select branching variables, but their deterministic predictions limit exploration and generalization. In this paper, we propose a novel framework that formulates branching variable selection as a conditional generative process, exploring deep-level decision features. Our approach leverages diffusion models to enable diverse and exploratory branching score generation, while consistency modeling distills this process into efficient one-step inference conditioned on the B&B state. This mode allows our method to achieve both high-quality and fast branching decisions, significantly improving the overall performance of branch-and-bound solvers. Extensive experiments on challenging cross-scale and cross-category benchmarks demonstrate that our framework consistently outperforms state-of-the-art imitation learning baselines, delivering substantial improvements in solution quality, computational efficiency, and inference speed.

AAAI Conference 2026 Conference Paper

KPDM: Key Phrase Dynamic Masking for Robust Text-to-Image Person Retrieval

  • Shaofeng You
  • Tianle Miao
  • Qihang Chen
  • Xin Li
  • Zhuo Cheng
  • Dapeng Luo

Text-to-image person re-identification (TIReID) aims to retrieve the most relevant pedestrian images from an image gallery based on natural language descriptions. Recent studies have achieved significant performance improvements by leveraging Masked Language Modeling (MLM) to align fine-grained information through local matching. However, in the text feature extraction, randomly masking text tokens may disrupt the semantic relationships between these local tokens, leading to feature misalignment; on the other hand, from an image feature perspective, redundant patches in pedestrian images hinder the information interaction across modalities. Moreover, the presence of noisy image-text pairs further complicates the learning process, as the model may be misled into recognizing incorrect patterns. To address these issues, we propose a robust fine-grained local alignment framework based on Key Phrase Dynamic Mask (KPDM). First, we strengthen the semantic relationships between text tokens by implementing a "adjective + noun'' phrase-level masking strategy, and design a frequency-based masked language loss (FMLM) to supervise fine-grained semantic-level local alignment. Second, we integrate cross-layer importance estimation to highlight key pedestrian image representations while removing redundant image features. Third, we propose a trusted consensus partitioning mechanism, utilizing intra-identity image-text similarity distributions to identify noisy pairs, enhancing the model robustness. Extensive experiments show that our method achieves 67.95% Rank-1 and 51.88% mAP on the RSTPReid dataset, exceeding the previous state-of-the-art by 2.6% and 1%. Furthermore, KPDM achieves Rank-1 accuracies of 75.97% on the CUHK-PEDES dataset and 67.78% on the ICFG-PEDES dataset, outperforming earlier methods.

AAAI Conference 2026 Conference Paper

La La LiDAR: Large-Scale Layout Generation from LiDAR Data

  • Youquan Liu
  • Lingdong Kong
  • Weidong Yang
  • Xin Li
  • Alan Liang
  • Runnan Chen
  • Ben Fei
  • Tongliang Liu

Controllable generation of realistic LiDAR scenes is crucial for applications such as autonomous driving and robotics. While recent diffusion-based models achieve high-fidelity LiDAR generation, they lack explicit control over foreground objects and spatial relationships, limiting their usefulness for scenario simulation and safety validation. To address these limitations, we propose Large-scale Layout-guided LiDAR generation model ("La La LiDAR"), a novel layout-guided generative framework that introduces semantic-enhanced scene graph diffusion with relation-aware contextual conditioning for structured LiDAR layout generation, followed by foreground-aware control injection for complete scene generation. This enables customizable control over object placement while ensuring spatial and semantic consistency. To support our structured LiDAR generation, we introduce Waymo-SG and nuScenes-SG, two large-scale LiDAR scene graph datasets, along with new evaluation metrics for layout synthesis. Extensive experiments demonstrate that La La LiDAR achieves state-of-the-art performance in both LiDAR generation and downstream perception tasks, establishing a new benchmark for controllable 3D scene generation.

AAAI Conference 2026 Conference Paper

LifeAlign: Lifelong Alignment for Large Language Models with Memory-Augmented Focalized Preference Optimization

  • Junsong Li
  • Jie Zhou
  • Bihao Zhan
  • Yutao Yang
  • Qianjun Pan
  • Shilian Chen
  • Tianyu Huai
  • Xin Li

Alignment plays a crucial role in Large Language Models (LLMs) in aligning with human preferences on a specific task/domain. Traditional alignment methods suffer from catastrophic forgetting, where models lose previously learned values when adapting to new preferences or domains. We introduce LifeAlign, a novel framework for lifelong alignment that enables LLMs to maintain consistent human preference alignment across sequential learning tasks without forgetting previously learned values. Our approach consists of two key innovations. First, we propose a focalized preference optimization strategy that aligns LLMs with new preferences while preventing the erosion of alignment acquired from previous tasks. Second, we develop a short-to-long memory consolidation mechanism that merges denoised short-term preference representations into stable long-term memory using intrinsic dimensionality reduction, enabling efficient storage and retrieval of alignment patterns across diverse domains. We evaluate LifeAlign across multiple sequential alignment tasks spanning different domains and preference types. Experimental results demonstrate that our method achieves superior performance in maintaining both preference alignment quality and knowledge retention compared to existing lifelong learning approaches.

AAAI Conference 2026 Conference Paper

Offline Meta-Reinforcement Learning with Flow-Based Task Inference and Adaptive Correction of Feature Overgeneralization

  • Min Wang
  • Xin Li
  • Mingzhong Wang
  • Hasnaa Bennis

Offline meta-reinforcement learning (OMRL) combines the strengths of learning from diverse datasets in offline RL with the adaptability to new tasks of meta-RL, promising safe and efficient knowledge acquisition by RL agents. However, OMRL still suffers extrapolation errors due to out-of-distribution (OOD) actions, compromised by broad task distributions and Markov Decision Process (MDP) ambiguity in meta-RL setups. Existing research indicates that the generalization of the Q network affects the extrapolation error in offline RL. This paper investigates this relationship by decomposing the Q value into feature and weight components, observing that while decomposition enhances adaptability and convergence in the case of high-quality data, it often leads to policy degeneration or collapse in complex tasks. We observe that decomposed Q values introduce a large estimation bias when the feature encounters OOD samples, a phenomenon we term "feature overgeneralization''. To address this issue, we propose FLORA, which identifies OOD samples by modeling feature distributions and estimating their uncertainties. FLORA integrates a return feedback mechanism to adaptively adjust feature components. Furthermore, to learn precise task representations, FLORA explicitly models the complex task distribution using a chain of invertible transformations. We theoretically and empirically demonstrate that FLORA achieves rapid adaptation and meta-policy improvement compared to baselines across various environments.

TMLR Journal 2026 Journal Article

Re:Form --- Reducing Human Priors in Scalable Formal Software Verification with RL in LLMs: A Preliminary Study on Dafny

  • Chuanhao Yan
  • Fengdi Che
  • Xuhan Huang
  • Xu Xu
  • Xin Li
  • Yizhi Li
  • Xingwei Qu
  • Jingzhe Shi

Existing informal language-based (e.g., human language) Large Language Models (LLMs) trained with Reinforcement Learning (RL) face a significant challenge: their verification processes, which provide crucial training signals, are neither reliable nor scalable. In fact, the prevalent large proprietary models could hardly generate verifiable programs. A promising yet largely uncharted alternative is formal language-based reasoning. Grounding LLMs in rigorous formal systems where generative models operate in formal language spaces (e.g., Dafny) enables the automatic and mathematically provable verification of their reasoning processes and outcomes. This capability is pivotal for achieving large-scale, reliable formal software verification. It is a common practice to employ human-annotated chain-of-thought and other human priors to induce the reasoning and coding capabilities of LLMs. Unfortunately, it becomes unacceptably all-consuming to provide such priors for supervising complex programming tasks. In this work, we systematically explore ways to reduce human priors with the formal language, Dafny, as the main environment for our pilot study. Our pipeline mainly relies on introducing an automatic and scalable data curation pipeline, and careful RL designs integrated with feedback from the formal language verifier. We introduce DafnyComp, a benchmark of compositional formal programs with auto-formalized specifications for specification reasoning. Our supervised fine-tuning (SFT) stage enables even small models (e.g., 0.5B) to generate syntactically valid and verifiable Dafny code, surpassing proprietary models. RL with regularization further improves performance, achieving stronger generalization to out-of-domain tasks and outperforming all strong baselines on the challenging DafnyComp benchmark. Anonymized code and models are available at https://github.com/ReFormDafny/ReForm and https://huggingface.co/ReFormDafny.

EAAI Journal 2026 Journal Article

Research of bullheading killing technique based on digital twin

  • Gezhen Mao
  • Jie Zhang
  • Xin Li
  • Sen Yang

Owing to the reliance of traditional bullheading killing technology on manual experience and offline models, it fails to respond to downhole dynamics in real time, leading to increased well-killing risks and test costs. Therefore, based on the process flow of bullheading killing, this study proposes a technical framework and structure for bullheading killing based on a Digital Twin (DT) system. Using Unity Three-Dimensional (Unity 3D) geometric modeling tools and dataset characterization methods, the DT of well control equipment, formations, and well-killing fluids were constructed. Through data preprocessing and a comparison of the learning methods, a data-driven real-time pressure prediction model (based on the Stacking integrated learning model, which is a type of Machine Learning (ML) approach, and also a subset of Artificial Intelligence (AI)) has been identified, which shortens the response time for well-killing parameter adjustment. The experiments showed that the latency of data collection, transmission, processing, and conversion was 68 s, comparing with traditional driller and engineer methods, its Coefficient of Determination (R2) reached 0. 98, and the pressure prediction deviation throughout the entire well-killing process is less than 3 %, verifying the accuracy and real-time performance of the DT system for bullheading killing. This study enhances the safety and efficiency of well-killing operations and provides references and insights for the further application of DT technology, as well as AI technology, in the well control field.

AAAI Conference 2026 Conference Paper

SGAT: Learning Feature Matching with Singularity-enhanced Graph Attention Network

  • Yizhuo Zhang
  • Kun Sun
  • Chang Tang
  • Yuanyuan Liu
  • Xin Li

The task of image feature matching aims to establish correct correspondences between images from two different views. While approaches based on attention mechanisms have demonstrated remarkable advancements in image feature matching, they still encounter substantial limitations. Specifically, current graph attention network approaches face performance bottlenecks in complex scenarios, such as low-texture regions or occlusions. This limitation stems from the self-attention mechanism, which, when lacking effective guidance, can lead to divergent attention weights or incorrect focus on regions with low discriminability, resulting in matching failures in low-texture environments. Inspired by how humans focus on distinctive regions when performing cross-view matching, we enhance attention to singular points in images that are salient, unique and have high cross-view matching potential during information aggregation, thereby improving matching capability. To realize the aforementioned strategies, we develop a novel Singularity-enhanced Graph Attention Network (SGAT). SGAT leverages Co-potentiality and Multi-Scale Singularity as prior guidance, and designs a Singularity-aware Attention mechanism and a Co-potentiality Guided Attention mechanism, specifically enhancing the perception of singularity and matching potential during feature interaction. Experimental results on multiple datasets, including ScanNet1500, demonstrate that our method outperforms current state-of-the-art sparse matching methods. In particular, the improvement is most pronounced in complex scenarios such as low-texture environments, significantly enhancing the accuracy and robustness of image matching and its downstream tasks.

AAAI Conference 2026 Conference Paper

Sonic4D: Spatial Audio Generation for Immersive 4D Scene Exploration

  • Siyi Xie
  • Hanxin Zhu
  • Xinyi Chen
  • Tianyu He
  • Xin Li
  • Zhibo Chen

Recent advancements in 4D generation have demonstrated its remarkable capability in synthesizing photorealistic renderings of dynamic 3D scenes. However, despite achieving impressive visual performance, almost all existing methods overlook the generation of spatial audio aligned with the corresponding 4D scenes, posing a significant limitation to truly immersive audiovisual experiences. To mitigate this issue, we propose Sonic4D, a novel framework that enables spatial audio generation for immersive exploration of 4D scenes. Specifically, our method is composed of three stages: 1) To capture both the dynamic visual content and raw auditory information from a monocular video, we first employ pre-trained expert models to generate the 4D scene and its corresponding monaural audio. 2) Subsequently, to transform the monaural audio into spatial audio, we localize and track the sound sources within the 4D scene, where their 3D spatial coordinates at different timestamps are estimated via a pixel-level visual grounding strategy. 3) Based on the estimated sound source locations, we further synthesize plausible spatial audio that varies across different viewpoints and timestamps using physics-based simulation. Extensive experiments have demonstrated that our proposed method generates realistic spatial audio consistent with the synthesized 4D scene in a training-free manner, significantly enhancing the immersive experience for users.

AAAI Conference 2026 Conference Paper

Test-Time Preference Optimization for Image Restoration

  • Bingchen Li
  • Xin Li
  • Jiaqi Xu
  • Jiaming Guo
  • Wenbo Li
  • Renjing Pei
  • Zhibo Chen

Image restoration (IR) models are typically trained to recover high-quality images using L1 or LPIPS loss. To handle diverse unknown degradations, zero-shot IR methods have also been introduced. However, existing pre-trained and zero-shot IR approaches often fail to align with human preferences, resulting in restored images that may not be favored. This highlights the critical need to enhance restoration quality and adapt flexibly to various image restoration tasks or backbones without requiring model retraining and ideally without labor-intensive preference data collection. In this paper, we propose the first Test-Time Preference Optimization (TTPO) paradigm for image restoration, which enhances perceptual quality, generates preference data on-the-fly, and is compatible with any IR model backbone. Specifically, we design a training-free, three-stage pipeline: (i) generate candidate preference images online using diffusion inversion and denoising based on the initially restored image; (ii) select preferred and dispreferred images using automated preference-aligned metrics or human feedback; and (iii) use the selected preference images as reward signals to guide the diffusion denoising process, optimizing the restored image to better align with human preferences. Extensive experiments across various image restoration tasks and models demonstrate the effectiveness and flexibility of the proposed pipeline.

AAAI Conference 2026 Conference Paper

Towards Affordance-Aware Robotic Dexterous Grasping with Human-like Priors

  • Haoyu Zhao
  • Linghao Zhuang
  • Xingyue Zhao
  • Cheng Zeng
  • Haoran Xu
  • Yuming Jiang
  • Jun CEN
  • Kexiang Wang

A dexterous hand capable of generalizable grasping objects is fundamental for the development of general-purpose embodied AI. However, previous methods focus narrowly on low-level grasp stability metrics, neglecting affordance-aware positioning and human-like poses which are crucial for downstream manipulation. To address these limitations, we propose AffordDex, a novel framework with two-stage training that learns a universal grasping policy with an inherent understanding of both motion priors and object affordances. In the first stage, a trajectory imitator is pre-trained on a large corpus of human hand motions to instill a strong prior for natural movement. In the second stage, a residual module is trained to adapt these general human-like motions to specific object instances. This refinement is critically guided by two components: our Negative Affordance-aware Segmentation (NAA) module, which identifies functionally inappropriate contact regions, and a privileged teacher-student distillation process that ensures the final vision-based policy is highly successful. Extensive experiments demonstrate that AffordDex not only achieves universal dexterous grasping but also remains remarkably human-like in posture and functionally appropriate in contact location. As a result, AffordDex significantly outperforms state-of-the-art baselines across seen objects, unseen instances, and even entirely novel categories.

AAAI Conference 2026 Conference Paper

UniFit: Towards Universal Virtual Try-on with MLLM-Guided Semantic Alignment

  • Wei Zhang
  • Yeying Jin
  • Xin Li
  • Yan Zhang
  • Xiaofeng Cong
  • Cong Wang
  • Fengcai Qiao
  • Zhichao Lian

Image-based virtual try-on (VTON) aims to synthesize photorealistic images of a person wearing specified garments. Despite significant progress, building a universal VTON framework that can flexibly handle diverse and complex tasks remains a major challenge. Recent methods explore multi-task VTON frameworks guided by textual instructions, yet they still face two key limitations: (1) semantic gap between text instructions and reference images, and (2) data scarcity in complex scenarios. To address these challenges, we propose UniFit, a universal VTON framework driven by a Multimodal Large Language Model (MLLM). Specifically, we introduce an MLLM-Guided Semantic Alignment Module (MGSA), which integrates multimodal inputs using an MLLM and a set of learnable queries. By imposing a semantic alignment loss, MGSA captures cross-modal semantic relationships and provides coherent and explicit semantic guidance for the generative process, thereby reducing the semantic gap. Moreover, by devising a two-stage progressive training strategy with a self-synthesis pipeline, UniFit is able to learn complex tasks from limited data. Extensive experiments show that UniFit not only supports a wide range of VTON tasks, including multi-garment and model-to-model try-on, but also achieves state-of-the-art performance.

AAAI Conference 2026 Conference Paper

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation

  • Shi-Xue Zhang
  • Hongfa Wang
  • Duojun Huang
  • Xin Li
  • Xiaobin Zhu
  • Xu-Cheng Yin

Video captions play a crucial role in text-to-video generation tasks, as their quality directly influences the semantic coherence and visual fidelity of the generated videos. Although large vision-language models (VLMs) have demonstrated significant potential in caption generation, existing benchmarks inadequately address fine-grained evaluation, particularly in capturing spatial-temporal details critical for video generation. To address this gap, we introduce the Fine-grained Video Caption Evaluation Benchmark (VCapsBench), the first large-scale fine-grained benchmark comprising 5,677 (5K+) videos and 109,796 (100K+) question-answer pairs. These QA-pairs are systematically annotated across 21 fine-grained dimensions (e.g., camera movement, and shot type) that are empirically proven critical for text-to-video generation. We further introduce three metrics (Accuracy (AR), Inconsistency Rate (IR), Coverage Rate (CR)), and an automated evaluation pipeline leveraging a large language model (LLM) to verify caption quality via contrastive QA-pairs analysis. Our benchmark can advance the development of robust text-to-video models by providing actionable insights for caption optimization.

JBHI Journal 2025 Journal Article

3D Isotropic High-Resolution Fetal Brain MRI Reconstruction From Motion Corrupted Thick Data Based on Physical-Informed Unsupervised Learning

  • Jiangjie Wu
  • Lixuan Chen
  • Zhenghao Li
  • Xin Li
  • Taotao Sun
  • Lihui Wang
  • Rongpin Wang
  • Hongjiang Wei

High-quality 3D fetal brain MRI reconstruction from motion-corrupted 2D slices is crucial for precise clinical diagnosis and advancing our understanding of fetal brain development. This necessitates reliable slice-to-volume registration (SVR) for motion correction and super-resolution reconstruction (SRR) techniques. Traditional approaches have their limitations, but deep learning (DL) offers the potential in enhancing SVR and SRR. However, most of DL methods require large-scale external 3D high-resolution (HR) training datasets, which is challenging in clinical fetal MRI. To address this issue, we propose an unsupervised iterative joint SVR and SRR DL framework for 3D isotropic HR volume reconstruction. Specifically, our method conceptualizes SVR as a function that maps a 2D slice and a 3D target volume to a rigid transformation matrix, aligning the slice to the underlying location within the target volume. This function is parameterized by a convolutional neural network, which is trained by minimizing the difference between the volume slicing at the predicted position and the actual input slice. For SRR, a decoding network embedded within a deep image prior framework, coupled with a comprehensive image degradation model, is used to produce the HR volume. The deep image prior framework offers a local consistency prior to guide the reconstruction of HR volumes. By performing a forward degradation model, the HR volume is optimized by minimizing the loss between the predicted slices and the acquired slices. Experiments on both large-magnitude motion-corrupted simulation data and clinical data have shown that our proposed method outperforms current state-of-the-art fetal brain reconstruction methods.

JBHI Journal 2025 Journal Article

A Dual Domain Collaborative Network for Polyp Segmentation

  • Yao Tong
  • Zuojian Zhou
  • Kongfa Hu
  • Tao Yang
  • Andr´e Kaup
  • Xin Li

Accurate polyp segmentation in colonoscopy images is essential for early colorectal cancer detection but remains a challenging problem due to the limitations in existing methods for optimizing boundary features and aligning cross-level representations. Specifically, the indistinct polyp boundaries and scale variations across different feature levels pose significant challenges for segmentation accuracy. To address these issues, we propose a dual domain collaborative network (DDCNet) that introduces two novel modules: a frequency context enhancement module (FCEM), which operates in the frequency domain to refine high- and low-frequency features, and a cross-level shift recalibrated fusion module (CSFM), which improves multi scale feature alignment in the spatial domain. The FCEM improves boundary precision by adaptively refining high frequency boundary features and enhancing low-frequency contextual information, while the CSFM mitigates cross level feature misalignment by dynamically recalibrating multi-scale features throughout the encoder-decoder architecture. Additionally, we design a hybrid loss function that integrates boundary, cross-entropy, and frequency consistency losses to further boost segmentation performance. Experimental results on three benchmark datasets (Kvasir SEG, CVC-ClinicDB, and CVC-ColonDB) demonstrate that DDCNet achieves state-of-the-art performance, with Dice coefficients of 0. 9343, 0. 9447, and 0. 8155, respectively. These results represent improvements of 1. 0%–1. 5% over the best existing methods. Ablation studies further validate the individual contributions of FCEM, CSFM, and the hybrid loss function. Additionally, we compared the proposed loss function with three commonly used functions.

JAIR Journal 2025 Journal Article

A New Literature Review of 3D Object Detection on Autonomous Driving

  • Peng Zhang
  • Xin Li
  • Xin Lin
  • Liang He

In recent years, the realm of computer vision has experienced a significant surge in the importance of 3D object detection, especially in the context of autonomous driving. The capability to precisely identify the locations, dimensions, and types of key 3D objects surrounding an autonomous vehicle is crucial, rendering 3D object detection a vital component of any advanced perception system. This review delivers an extensive overview of the emerging technologies in 3D object detection tailored for autonomous vehicles. It encompasses a thorough examination, evaluation, and integration of the current research landscape in this domain, staying up-to-date with the latest advancements in 3D object detection and suggesting prospective avenues for future research. Our survey begins by clarifying the principles of 3D object detection and addressing its present challenges in the 3D domain. We then introduce three distinct taxonomies: camera-based, point cloudbased, and multi-modality-based approaches, providing a comprehensive classification of contemporary 3D object detection methodologies from various angles. Diverging from previous reviews, this paper also highlights and scrutinizes common issues and solutions for specific scenarios (such as pedestrian detection, lane lines, roadside cameras, and weather conditions) in object detection. Furthermore, we conduct an in-depth analysis and comparison of different classifications and methods, utilizing various datasets and experimental outcomes. Conclusively, we suggest several potential research directions, offering valuable insights for the ongoing evolution of 3D object detection technology. This review aims to serve as a comprehensive resource for researchers and practitioners in the field, guiding future innovations in 3D object detection for autonomous driving.

EAAI Journal 2025 Journal Article

A novel anomaly detection and classification algorithm for application in tuyere images of blast furnace

  • Yifan Duan
  • Xiaojie Liu
  • Ran Liu
  • Xin Li
  • Hongwei Li
  • Hongyang Li
  • Yanqin Sun
  • Yujie Zhang

Traditional relying on manual experience to assess the tuyere status consumes significant human resources. In the era of intelligent blast furnaces and intensified smelting, this approach struggles to meet the demands for accuracy and real-time assessment, posing challenges to safety and efficiency of blast furnace production. Tuyere images exhibit high feature similarity, and the number of samples is often limited. Therefore, if a simple convolution operation is only used, it will be difficult to discern differences across various images. To address this challenge and cater to the requirements of intelligent tuyere status recognition across different steel enterprises, we designed a novel deep neural network algorithm called ES-SFRNet (Enhanced Sequential: Feature Fusion and Recognition Network), building upon our prior research. The algorithm concurrently modeled tuyere images alongside relevant time series data, comprising three components: Feature pre-extraction, Tuyere status recognition, and Generalization & Robustness. The first two modules focus on feature extraction and fusion of tuyere images, while leveraging edge detection information from the image, we developed a mathematical index A r (Area Ratio) to serve as an auxiliary criterion for tuyere status recognition. Given the model's future scalability and multi-scenario application, the final module focuses on knowledge integration and parameter control. Test results reveal an overall accuracy rate of 99. 3% for the ES-SFRNet algorithm, effectively capturing key parameters to facilitate on-site operations. In comparison to other mainstream object detection algorithms, our algorithm framework excels in tuyere image feature extraction and recognition, which can offer broad applications to Chinese blast furnace ironmaking industry.

EAAI Journal 2025 Journal Article

A Transformer-based self-supervised learning model for fault diagnosis of air-conditioning systems with limited labeled data

  • Mei Hua
  • Ke Yan
  • Xin Li

Despite the great successes of supervised learning-based fault diagnosis techniques for heating, ventilation and air-conditioning (HVAC) systems, their applications are severely limited due to insufficient labeled data accompanied with massive unlabeled data. To address this drawback, a Transformer-based self-supervised representation learning model (TSSRL) is proposed in this study for HVAC fault diagnosis with limited labeled data. Specifically, a customized Transformer model is developed as the feature encoder by embedding a context-attention module on the self-attention module, which enables TSSRL to mine the contextual representations among input data. In addition, a joint data augmentation strategy is designed to improve the diversity of inputs, promoting the pretext tasks to learn more extensive representations from unlabeled data. Meanwhile, two cooperative pretext tasks, namely contrastive similarity matching and data reconstruction, are formulated to extract discriminative representations from unlabeled data. The diagnosis-beneficial representations learned from unlabeled data are used for downstream classification modeling tasks with limited labeled data. Experiments on two benchmark HVAC fault datasets demonstrate the superiority of the proposed TSSRL model over other state-of-the-art HVAC fault diagnosis methods.

AILAW Journal 2025 Journal Article

Adversarial training flat-lattice transformer for named entity recognition of chinese legal texts

  • Jiabao Wang
  • Kaixuan Wang
  • Yang Weng
  • Xin Li

Abstract Judgment documents are the legally binding written conclusion made by the court based on the facts of the case and the law. Due to the use of professional terms and nested combinations of words, potential information of judgment documents has not been deeply excavated. Named Entity Recognition (NER) is a necessary task in Natural Language Processing (NLP), and has been widely introduced into Chinese texts processing for many years. However, the professional terms and nested words lead to the boundaries between entities being blurred, which can not accurately divide entities. In this paper, a new NER model Adversarial Training Flat-Lattice Transformer (AT-Flat) which combines adversarial training and Flat-Lattice Transformer (Flat) is proposed to weaken these problems. In AT-Flat, the Flat combines character and word information to get sequence information, and the CRF is used to output the final entity prediction results. Moreover, the key point is an adversarial training framework introduced to integrate task-shared word boundary information from Chinese Word Segmentation (CWS) task into Chinese NER task. The framework is able to filter out the noise caused by CWS task and further enhance the effect of Chinese NER task. More importantly, experiments on NER task of Chinese traffic accident and financial lending judgment documents show that our method outperforms other state-of-the-art methods. These verify that our method can effectively alleviate the problem of poor NER effect caused by professional terms and nested words. In addition, three public Chinese NER datasets were also used to evaluate our method.

IJCAI Conference 2025 Conference Paper

ARPDL: Adaptive Relational Prior Distribution Loss as an Adapter for Document-Level Relation Extraction

  • Huangming Xu
  • Fu Zhang
  • Jingwei Cheng
  • Xin Li

The goal of document-level relation extraction (DocRE) is to identify relations between entities from multiple sentences. As a multi-label classification task, a common approach is to determine whether there are relations for an entity pair by selecting a multi-label classification threshold, with scores of relations above the threshold predicted as positive and the rest as negative. However, we find that predicting multiple relations for entity pairs causes the decrease of predicted scores in positive classes. This could lead to many positive classes being incorrectly predicted as negative. Additionally, our analysis suggests that fitting the distribution of predicted relations to the prior distribution of relations can help improve prediction performance. However, previous studies have not explored or leveraged the prior distribution of relations. To address these issues and findings, we for the first time propose the idea of incorporating the relational prior distribution into the loss calculation in DocRE tasks. We innovatively propose an Adaptive Relational Prior Distribution Loss (ARPDL), which can adaptively adjust relation prediction scores based on the relational prior distribution. Our designed relational prior distribution component can also be integrated as an adapter into other threshold-based losses to improve prediction performance. Experimental results demonstrate that ARPDL consistently improves the performance of existing DocRE models, achieving new state-of-the-art results. Furthermore, integrating our relational prior distribution adapter into other losses significantly enhances their performance in DocRE tasks, validating the effectiveness and generality of our approach. Code is available at https: //github. com/xhm-code/ARPDL.

AAAI Conference 2025 Conference Paper

Automated Creation of Reusable and Diverse Toolsets for Enhancing LLM Reasoning

  • Zhiyuan Ma
  • Zhenya Huang
  • Jiayu Liu
  • Minmao Wang
  • Hongke Zhao
  • Xin Li

Augmenting large language models (LLMs) with tools significantly enhances their problem-solving potential across multifaceted tasks. However, current tools automatically created by LLMs often serve as a mere summary of specific problems or solutions, which face two main issues: 1) Low reusability: The tools are overly problem-specific and struggle to handle new problems. 2) Limited diversity: The toolsets are too narrow, limiting their application to address a broader range of different problems. In this paper, we propose the Knowledge-grounded Tool Creation with Evolution (KTCE) framework, which aims to craft reusable and comprehensive toolsets for LLMs in a two-stage process. In the first stage (Knowledge-based Tool Creation), we conceptualize tools as a form of executable domain knowledge and propose a problem-knowledge-tool paradigm. Specifically, we leverage LLMs to abstract "knowledge" from "problems" and create a three-layer knowledge tree of topics, concepts, and key points. This hierarchical structure serves as a foundation for inducing atomic "tools" from "knowledge", grounding them in fundamental concepts and enhancing their usability. In the second stage (Tool Evolutionary Search), we evolve the toolsets through several actions including tool selection, mutation, and crossover. This stage mimics the biological evolution process, aiding toolsets in discovering new tools or updating existing ones, thereby increasing the diversity of the toolset. Experiments on challenging mathematical/tabular/scientific reasoning tasks demonstrate that our approach achieves substantial accuracy improvements ranging from 6.23% to 18.49% on average. Moreover, in-depth analyses reveal the superior characteristics of our toolkit, including high reusability, high diversity, and high generalizability on cross-data/LLM performance with low complexity.

NeurIPS Conference 2025 Conference Paper

Diff-ICMH: Harmonizing Machine and Human Vision in Image Compression with Generative Prior

  • Ruoyu Feng
  • Yunpeng Qi
  • Jinming Liu
  • Yixin Gao
  • Xin Li
  • Xin Jin
  • Zhibo Chen

Image compression methods are usually optimized isolatedly for human perception or machine analysis tasks. We reveal fundamental commonalities between these objectives: preserving accurate semantic information is paramount, as it directly dictates the integrity of critical information for intelligent tasks and aids human understanding. Concurrently, enhanced perceptual quality not only improves visual appeal but also, by ensuring realistic image distributions, benefits semantic feature extraction for machine tasks. Based on this insight, we propose Diff-ICMH, a generative image compression framework aiming for harmonizing machine and human vision in image compression. It ensures perceptual realism by leveraging generative priors and simultaneously guarantees semantic fidelity through the incorporation of Semantic Consistency loss (SC loss) during training. Additionally, we introduce the Tag Guidance Module (TGM) that leverages highly semantic image-level tags to stimulate the pre-trained diffusion model's generative capabilities, requiring minimal additional bit rates. Consequently, Diff-ICMH supports multiple intelligent tasks through a single codec and bitstream without any task-specific adaptation, while preserving high-quality visual experience for human perception. Extensive experimental results demonstrate Diff-ICMH's superiority and generalizability across diverse tasks, while maintaining visual appeal for human perception.

IJCAI Conference 2025 Conference Paper

Enhancing Table Recognition with Vision LLMs: A Benchmark and Neighbor-Guided Toolchain Reasoner

  • Yitong Zhou
  • Mingyue Cheng
  • Qingyang Mao
  • Jiahao Wang
  • Feiyang Xu
  • Xin Li

Pre-trained foundation models have recently made significant progress in table-related tasks such as table understanding and reasoning. However, recognizing the structure and content of unstructured tables using Vision Large Language Models (VLLMs) remains under-explored. To bridge this gap, we propose a benchmark based on a hierarchical design philosophy to evaluate the recognition capabilities of VLLMs in training-free scenarios. Through in-depth evaluations, we find that low-quality image input is a significant bottleneck in the recognition process. Drawing inspiration from this, we propose the Neighbor-Guided Toolchain Reasoner (NGTR) framework, which is characterized by integrating diverse lightweight tools for visual operations aimed at mitigating issues with low-quality images. Specifically, we transfer a tool selection experience from a similar neighbor to the input and design a reflection module to supervise the tool invocation process. Extensive experiments on public datasets demonstrate that our approach significantly enhances the recognition capabilities of the vanilla VLLMs. We believe that the benchmark and framework could provide an alternative solution to table recognition.

NeurIPS Conference 2025 Conference Paper

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World?

  • Yuqian Yuan
  • Ronghao Dang
  • Long Li
  • Wentong Li
  • Dian Jiao
  • Xin Li
  • Deli Zhao
  • Fan Wang

The emergence of multimodal large language models (MLLMs) has driven breakthroughs in egocentric vision applications. These applications necessitate persistent, context-aware understanding of objects, as users interact with tools in dynamic and cluttered environments. However, existing embodied benchmarks primarily focus on static scene exploration, emphasizing object's appearance and spatial attributes while neglecting the assessment of dynamic changes arising from users' interactions. capabilities in object-level spatiotemporal reasoning required for real-world interactions. To address this gap, we introduce EOC-Bench, an innovative benchmark designed to systematically evaluate object-centric embodied cognition in dynamic egocentric scenarios. Specially, EOC-Bench features 3, 277 meticulously annotated QA pairs categorized into three temporal categories: Past, Present, and Future, covering 11 fine-grained evaluation dimensions and 3 visual object referencing types. To ensure thorough assessment, we develop a mixed-format human-in-the-loop annotation frameworkBased on EOC-Bench, we conduct comprehensive evaluations of various proprietary, open-source, and object-level MLLMs. EOC-Bench serves as a crucial tool for advancing the embodied object cognitive capabilities of MLLMs, establishing a robust foundation for developing reliable core models for embodied systems.

NeurIPS Conference 2025 Conference Paper

FOCUS: Unified Vision-Language Modeling for Interactive Editing Driven by Referential Segmentation

  • Fan Yang
  • Yousong Zhu
  • Xin Li
  • Yufei Zhan
  • Hongyin Zhao
  • Shurong Zheng
  • Yaowei Wang
  • Ming Tang

Recent Large Vision Language Models (LVLMs) demonstrate promising capabilities in unifying visual understanding and generative modeling, enabling both accurate content understanding and flexible editing. However, current approaches treat \textbf{\textit{"what to see"}} and \textbf{\textit{"how to edit"}} separately: they either perform isolated object segmentation or utilize segmentation masks merely as conditional prompts for local edit generation tasks, often relying on multiple disjointed models. To bridge these gaps, we introduce FOCUS, a unified LVLM that integrates segmentation-aware perception and controllable object-centric generation within an end-to-end framework. FOCUS employs a dual-branch visual encoder to simultaneously capture global semantic context and fine-grained spatial details. In addition, we leverage a MoVQGAN-based visual tokenizer to produce discrete visual tokens that enhance generation quality. To enable accurate and controllable image editing, we propose a progressive multi-stage training pipeline, where segmentation masks are jointly optimized and used as spatial condition prompts to guide the diffusion decoder. This strategy aligns visual encoding, segmentation, and generation modules, effectively bridging segmentation-aware perception with fine-grained visual synthesis. Extensive experiments across three core tasks, including multimodal understanding, referring segmentation accuracy, and controllable image generation, demonstrate that FOCUS achieves strong performance by jointly optimizing visual perception and generative capabilities.

NeurIPS Conference 2025 Conference Paper

From Kolmogorov to Cauchy: Shallow XNet Surpasses KANs

  • Xin Li
  • Xiaotao Zheng
  • Zhihong Xia

We study a shallow variant of XNet, a neural architecture whose activation functions are derived from the Cauchy integral formula. While prior work focused on deep variants, we show that even a single-layer XNet exhibits near-exponential approximation rates—exceeding the polynomial bounds of MLPs and spline-based networks such as Kolmogorov–Arnold Networks (KANs). Empirically, XNet reduces approximation error by over 600× on discontinuous functions, achieves up to 20, 000× lower residuals in physics-informed PDEs, and improves policy accuracy and sample efficiency in PPO-based reinforcement learning—while maintaining comparable or better computational efficiency than KAN baselines. These results demonstrate that expressive approximation can stem from principled activation design rather than depth alone, offering a compact, theoretically grounded alternative for function approximation, scientific computing, and control.

NeurIPS Conference 2025 Conference Paper

FSI-Edit: Frequency and Stochasticity Injection for Flexible Diffusion-Based Image Editing

  • Kaixiang Yang
  • Xin Li
  • Yuxi Li
  • Qiang Li
  • Zhiwei Wang

Latent Diffusion-based Text-to-Image (T2I) is a free image editing tool that typically reverses an image into noise, reconstructs it using its original text prompt, and then generates an edited version under a new target prompt. To preserve unaltered image content, features from the reconstruction are directly injected to replace selected features in the generation. However, this direct replacement often leads to feature incompatibility, compromising editing fidelity and limiting creative flexibility, particularly for non-rigid edits (\emph{e. g. }, structural or pose changes). In this paper, we aim to address these limitations by proposing \textbf{FSI-Edit}, a novel framework using frequency- and stochasticity-based feature injection for flexible image editing. First, FSI-Edit enhances feature consistency by injecting \emph{high-frequency} components of reconstruction features into generation features, mitigating incompatibility while preserving the editing ability for major structures encoded in low-frequency information. Second, it introduces controlled \emph{noise} into the replaced reconstruction features, expanding the generative space to enable diverse non-rigid edits beyond the original image’s constraints. Experiments on non-rigid edits, \emph{e. g. }, addition, deletion, and pose manipulation, demonstrate that FSI-Edit outperforms existing baselines in target alignment, semantic fidelity and visual quality. Our work highlights the critical roles of frequency-aware design and stochasticity in overcoming rigidity in diffusion-based editing.

TIST Journal 2025 Journal Article

How Do Large Language Models Understand Genes and Cells

  • Chen Fang
  • Yidong Wang
  • Yunze Song
  • Qingqing Long
  • Wang Lu
  • Linghui Chen
  • Guihai Feng
  • Yuanchun Zhou

Researching genes and their interactions is crucial for deciphering the fundamental laws of cellular activity, advancing disease treatment, drug discovery, and more. Large language Models (LLMs), with their profound text comprehension and generation capabilities, have made significant strides across various natural science fields. However, their application in cell biology remains limited and a systematic evaluation of their performance is lacking. To address this gap, in this article, we select seven mainstream LLMs and evaluate their performance across nine gene-related problem scenarios. Our findings indicate that LLMs possess a certain level of understanding of genes and cells, but still lag behind domain-specific models in comprehending transcriptional expression profiles. Moreover, we have improved the current method of textual representation of cells, enhancing the LLMs’ ability to tackle cell annotation tasks. We encourage cell biology researchers to leverage LLMs for problem-solving while being mindful of the associated challenges. We release our code and data at https://github.com/epang-ucas/Evaluate_LLMs_to_Genes.

EAAI Journal 2025 Journal Article

Machine learning methods comparison for maritime wireless signal strength prediction

  • Lisha Peng
  • Kun Yang
  • Jianming Wu
  • Chengyuan Wen
  • Tong Peng
  • Xin Li

The 5th generation mobile networks (5G) signal strength prediction techniques based on artificial intelligence (AI) demonstrate numerous advantages, such as easy implementation and high accuracy. However, what kind of machine learning model can provide superior performance in maritime applications is still unknown. In this paper, we developed twelve predictive models, encompassing eight machine learning and four linear regression approaches, for estimating reference signal receiving power (RSRP) and received signal strength indicator (RSSI) in maritime land-to-ship (L2S) scenarios. Our models are rigorously evaluated using nine metrics, trained on 5G band data collected with our self-designed equipment, and refined with carefully selected features to maximize prediction accuracy. The results show that the machine learning models are generally superior to the linear regression models in the fitting and prediction of RSRP and RSSI according to our evaluation and prediction experiments. Especially, rule-based models like decision tree regression (DTR) can accurately learn the impact of both large-scale and small-scale fading on the prediction objects from the data, and also have strong model interpretability. Our work has provided good reference value for future machine learning-based channel modeling and model evaluations.

AAAI Conference 2025 Conference Paper

Multi-Perspective Consolidation Enhanced Cognitive Diagnosis via Conditional Diffusion Model

  • Guanhao Zhao
  • Zhenya Huang
  • Cheng Cheng
  • Yan Zhuang
  • Qingyang Mao
  • Xin Li
  • Shijin Wang
  • Enhong Chen

Cognitive diagnosis, which assesses the learners' competence from learners' interaction logs, plays a vital role in education. It provides a crucial reference for gauging learners' proficiency levels and tailoring future learning activities accordingly. Researchers have proposed numerous cognitive diagnosis models to address this task. Despite their success, these models continue to face the ill-posed problem because of the information loss caused by under-expressive interaction function and incomplete observations. In this paper, we address these challenges by proposing a novel cognitive diagnosis model, DMC-CDM, based on the theoretical premise that cognitive states can be captured with minimal information loss by maximizing the mutual information between observed and potential observations. Specifically, DMC-CDM incorporates a semantic extractor to provide a comprehensive semantic understanding of learners' interaction logs, thereby enhancing current collaborative-based cognitive state representations. It then consolidates multi-perspective observations to capture precise cognitive states by maximizing mutual information between these observations. We conducted extensive experiments on three datasets, and the experimental results demonstrate that our proposed model is both effective and beneficial for downstream applications in education.

EAAI Journal 2025 Journal Article

Multi-pulse superposition for droplet volume control in inkjet printing based on model and data fusion

  • Xiao Yue
  • Xin Li
  • Jiankui Chen
  • Wei Chen
  • Hua Yang
  • Jincheng Gao
  • Zhouping Yin

Inkjet printing technology for fabricating organic light-emitting diode display panels offers advantages such as high material utilization and the capability for large-area manufacturing. When printing display panels with varying resolutions, ejecting droplets of different sizes from the nozzle is often necessary to balance print quality and efficiency. However, due to nozzle size limitations, the volume range of stable droplets produced by a single-pulse driving waveform is relatively narrow, with the maximum volume being less than twice the minimum volume. Therefore, approaches based on superposition of multi-pulse waveforms have attracted attention, but existing studies only implement manual design of waveforms based on experimental laws and rarely involve automatic regulation of multi-pulse waveform parameters, which is not favorable for industrial applications. Based on combining a meniscus vibration model with industrial ejection data, this paper extracts control strategies from historical data using deep reinforcement learning, and recommends initial waveform parameters through a fuzzy system. Then, the multi-pulse waveform parameters are automatically adjusted in real-time based on the observed droplet volume to fuse the droplets at the nozzle, enabling a wider range of droplet volume closed-loop control. Experiments on industrial inkjet printing equipment implemented intelligent closed-loop regulation of multi-pulse driving waveforms, successfully controlling droplets of different sizes such as 2, 4, and 8 picoliter with an error accuracy of less than ± 4 %. This approach applies artificial intelligence algorithms to inkjet printing engineering and intelligently adjusts the multi-pulse waveform parameters to enhance the controllable range of droplet volumes.

IJCAI Conference 2025 Conference Paper

PatternCIR Benchmark and TisCIR: Advancing Zero-Shot Composed Image Retrieval in Remote Sensing

  • Zhechun Liang
  • Tao Huang
  • Fangfang Wu
  • Shiwen Xue
  • Zhenyu Wang
  • Weisheng Dong
  • Xin Li
  • Guangming Shi

Remote sensing composed image retrieval (RSCIR) is a new vision-language task that takes a composed query of an image and text, aiming to search for a target remote sensing image satisfying two conditions from intricate remote sensing imagery. However, the existing attribute-based benchmark Patterncom in RSCIR has significant flaws, including the lack of query text sentences and paired triplets, thus making it unable to evaluate the latest methods. To address this, we propose the Zero-Shot Query Text Generator (ZS-QTG) that can generate full query text sentences based on attributes, and then, by capitalizing on ZS-QTG, we develop the PatternCIR benchmark. PatternCIR rectifies Patterncom’s deficiencies and enables the evaluation of existing methods. Additionally, we explore zero-shot composed image retrieval methods that do not rely on massive pre-collected triplets for training. Existing methods use only the text during retrieval, performing poorly in RSCIR. To improve this, we propose Text-image Sequential Training of Composed Image Retrieval (TisCIR). TisCIR undergoes sequential training of multiple self-masking projection and fine-grained image attention modules, which endows it with the capacity to filter out conflicting information between the image and text, enhancing the retrieval by utilizing both modalities in harmony. TisCIR outperforms existing methods by 12. 40% to 62. 03% on PatternCIR, achieving state-of-the-art performance in RSCIR. The data and code are available here.

IJCAI Conference 2025 Conference Paper

RLMiniStyler: Light-weight RL Style Agent for Arbitrary Sequential Neural Style Generation

  • Jing Hu
  • Chengming Feng
  • Shu Hu
  • Ming-Ching Chang
  • Xin Li
  • Xi Wu
  • Xin Wang

Arbitrary style transfer aims to apply the style of any given artistic image to another content image. Still, existing deep learning-based methods often require significant computational costs to generate diverse stylized results. Motivated by this, we propose a novel reinforcement learning-based framework for arbitrary style transfer RLMiniStyler. This framework leverages a unified reinforcement learning policy to iteratively guide the style transfer process by exploring and exploiting stylization feedback, generating smooth sequences of stylized results while achieving model lightweight. Furthermore, we introduce an uncertainty-aware multi-task learning strategy that automatically adjusts loss weights to adapt to the content and style balance requirements at different training stages, thereby accelerating model convergence. Through a series of experiments across image various resolutions, we have validated the advantages of RLMiniStyler over other state-of-the-art methods in generating high-quality, diverse artistic image sequences at a lower cost. Codes are available at https: //github. com/fengxiaoming520/RLMiniStyler.

AAAI Conference 2025 Conference Paper

Symbolic Neural Ordinary Differential Equations

  • Xin Li
  • Chengli Zhao
  • Xue Zhang
  • Xiaojun Duan

Differential equations are widely used to describe complex dynamical systems with evolving parameters in nature and engineering. Effectively learning a family of maps from the parameter function to the system dynamics is of great significance. In this study, we propose a novel learning framework of symbolic continuous-depth neural networks, termed Symbolic Neural Ordinary Differential Equations (SNODEs), to effectively and accurately learn the underlying dynamics of complex systems. Specifically, our learning framework comprises three stages: initially, pre-training a predefined symbolic neural network via a gradient flow matching strategy; subsequently, fine-tuning this network using Neural ODEs; and finally, constructing a general neural network to capture residuals. In this process, we apply the SNODEs framework to partial differential equation systems through Fourier analysis, achieving resolution-invariant modeling. Moreover, this framework integrates the strengths of symbolism and connectionism, boasting a universal approximation theorem while significantly enhancing interpretability and extrapolation capabilities relative to state-of-the-art baseline methods. We demonstrate this through experiments on several representative complex systems. Therefore, our framework can be further applied to a wide range of scientific problems, such as system bifurcation and control, reconstruction and forecasting, as well as the discovery of new equations.

ICML Conference 2025 Conference Paper

Task-Gated Multi-Expert Collaboration Network for Degraded Multi-Modal Image Fusion

  • Yiming Sun 0003
  • Xin Li
  • Pengfei Zhu 0001
  • Qinghua Hu
  • Dongwei Ren
  • Huiying Xu
  • Xinzhong Zhu

Multi-modal image fusion aims to integrate complementary information from different modalities to enhance perceptual capabilities in applications such as rescue and security. However, real-world imaging often suffers from degradation issues, such as noise, blur, and haze in visible imaging, as well as stripe noise in infrared imaging, which significantly degrades model performance. To address these challenges, we propose a task-gated multi-expert collaboration network (TG-ECNet) for degraded multi-modal image fusion. The core of our model lies in the task-aware gating and multi-expert collaborative framework, where the task-aware gating operates in two stages: degradation-aware gating dynamically allocates expert groups for restoration based on degradation types, and fusion-aware gating guides feature integration across modalities to balance information retention between fusion and restoration tasks. To achieve this, we design a two-stage training strategy that unifies the learning of restoration and fusion tasks. This strategy resolves the inherent conflict in information processing between the two tasks, enabling all-in-one multi-modal image restoration and fusion. Experimental results demonstrate that TG-ECNet significantly enhances fusion performance under diverse complex degradation conditions and improves robustness in downstream applications. The code is available at https: //github. com/LeeX54946/TG-ECNet.

NeurIPS Conference 2025 Conference Paper

The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio

  • Sicong Leng
  • Yun Xing
  • Zesen Cheng
  • Yang Zhou
  • Hang Zhang
  • Xin Li
  • Deli Zhao
  • Shijian Lu

Recent advancements in large multimodal models (LMMs) have significantly enhanced performance across diverse tasks, with ongoing efforts to further integrate additional modalities such as video and audio. However, most existing LMMs remain vulnerable to hallucinations, the discrepancy between the factual multimodal input and the generated textual output, which has limited their applicability in various real-world scenarios. This paper presents the first systematic investigation of hallucinations in LMMs involving the three most common modalities: language, visual, and audio. Our study reveals two key contributors to hallucinations: overreliance on unimodal priors and spurious inter-modality correlations. To address these challenges, we introduce the benchmark The Curse of Multi-Modalities (CMM), which comprehensively evaluates hallucinations in LMMs, providing a detailed analysis of their underlying issues. Our findings highlight key vulnerabilities, including imbalances in modality integration and biases from training data, underscoring the need for balanced cross-modal learning and enhanced hallucination mitigation strategies. Based on our observations and findings, we suggest potential research directions that could enhance the reliability of LMMs.

AAAI Conference 2025 Conference Paper

TIV-Diffusion: Towards Object-Centric Movement for Text-driven Image to Video Generation

  • Xingrui Wang
  • Xin Li
  • Yaosi Hu
  • Hanxin Zhu
  • Chen Hou
  • Cuiling Lan
  • Zhibo Chen

Text-driven Image to Video Generation (TI2V) aims to generate controllable video given the first frame and corresponding textual description. The primary challenges of this task lie in two parts: (i) how to identify the target objects and ensure the consistency between the movement trajectory and the textual description. (ii) how to improve the subjective quality of generated videos. To tackle the above challenges, we propose a new diffusion-based TI2V framework, termed TIV-Diffusion, via object-centric textual-visual alignment, intending to achieve precise control and high-quality video generation based on textual-described motion for different objects. Concretely, we enable our TIV-Diffuion model to perceive the textual-described objects and their motion trajectory by incorporating the fused textual and visual knowledge through scale-offset modulation. Moreover, to mitigate the problems of object disappearance and misaligned objects and motion, we introduce an object-centric textual-visual alignment module, which reduces the risk of misaligned objects/motion by decoupling the objects in the reference image and aligning textual features with each object individually. Based on the above innovations, our TIV-Diffusion achieves state-of-the-art high-quality video generation compared with existing TI2V methods.

AAAI Conference 2025 Conference Paper

Training-Free Image Manipulation Localization Using Diffusion Models

  • Zhenfei Zhang
  • Ming-Ching Chang
  • Xin Li

Image manipulation localization (IML) is a critical technique in media forensics, focusing on identifying tampered regions within manipulated images. Most existing IML methods require extensive training on labeled datasets with both image-level and pixel-level annotations. These methods often struggle with new manipulation types and exhibit low generalizability. In this work, we propose a training-free IML approach using diffusion models. Our method adaptively selects an appropriate number of diffusion timesteps for each input image in the forward process and performs both conditional and unconditional reconstructions in the backward process without relying on external conditions. By comparing these reconstructions, we generate a localization map highlighting regions of manipulation based on inconsistencies. Extensive experiments were conducted using sixteen state-of-the-art (SoTA) methods across six IML datasets. The results demonstrate that our training-free method outperforms SoTA unsupervised and weakly-supervised techniques. Furthermore, our method competes effectively against fully-supervised methods on novel (unseen) manipulation types.

NeurIPS Conference 2024 Conference Paper

Can Large Language Models Analyze Graphs like Professionals? A Benchmark, Datasets and Models

  • Xin Li
  • Weize Chen
  • Qizhi Chu
  • Haopeng Li
  • Zhaojun Sun
  • Ran Li
  • Chen Qian
  • Yiwei Wei

The need to analyze graphs is ubiquitous across various fields, from social networks to biological research and recommendation systems. Therefore, enabling the ability of large language models (LLMs) to process graphs is an important step toward more advanced general intelligence. However, current LLM benchmarks on graph analysis require models to directly reason over the prompts describing graphtopology, and are thus limited to small graphs with only a few dozens of nodes. In contrast, human experts typically write programs based on popular libraries for task solving, and can thus handle graphs with different scales. To this end, a question naturally arises: can LLMs analyze graphs like professionals? In this paper, we introduce ProGraph, a manually crafted benchmark containing 3 categories of graph tasks. The benchmark expects solutions based on programming instead of directly reasoning over raw inputs. Our findings reveal that the performance of current LLMs is unsatisfactory, with the best model achieving only 36% accuracy. To bridge this gap, we propose LLM4Graph datasets, which include crawled documents and auto-generated codes based on 6 widely used graph libraries. By augmenting closed-source LLMs with document retrieval and fine-tuning open-source ones on the codes, we show 11-32% absolute improvements in their accuracies. Our results underscore that the capabilities of LLMs in handling structured data are still under-explored, and show the effectiveness of LLM4Graph in enhancing LLMs’ proficiency of graph analysis. The benchmark, datasets and enhanced open-sourcemodels are available at https: //github. com/BUPT-GAMMA/ProGraph.

NeurIPS Conference 2024 Conference Paper

Continuously Learning, Adapting, and Improving: A Dual-Process Approach to Autonomous Driving

  • Jianbiao Mei
  • Yukai Ma
  • Xuemeng Yang
  • Licheng Wen
  • Xinyu Cai
  • Xin Li
  • Daocheng Fu
  • Bo Zhang

Autonomous driving has advanced significantly due to sensors, machine learning, and artificial intelligence improvements. However, prevailing methods struggle with intricate scenarios and causal relationships, hindering adaptability and interpretability in varied environments. To address the above problems, we introduce LeapAD, a novel paradigm for autonomous driving inspired by the human cognitive process. Specifically, LeapAD emulates human attention by selecting critical objects relevant to driving decisions, simplifying environmental interpretation, and mitigating decision-making complexities. Additionally, LeapAD incorporates an innovative dual-process decision-making module, which consists of an Analytic Process (System-II) for thorough analysis and reasoning, along with a Heuristic Process (System-I) for swift and empirical processing. The Analytic Process leverages its logical reasoning to accumulate linguistic driving experience, which is then transferred to the Heuristic Process by supervised fine-tuning. Through reflection mechanisms and a growing memory bank, LeapAD continuously improves itself from past mistakes in a closed-loop environment. Closed-loop testing in CARLA shows that LeapAD outperforms all methods relying solely on camera input, requiring 1-2 orders of magnitude less labeled data. Experiments also demonstrate that as the memory bank expands, the Heuristic Process with only 1. 8B parameters can inherit the knowledge from a GPT-4 powered Analytic Process and achieve continuous performance improvement. Project page: https: //pjlab-adg. github. io/LeapAD

NeurIPS Conference 2024 Conference Paper

CYCLO: Cyclic Graph Transformer Approach to Multi-Object Relationship Modeling in Aerial Videos

  • Trong-Thuan Nguyen
  • Pha Nguyen
  • Xin Li
  • Jackson Cothren
  • Alper Yilmaz
  • Khoa Luu

Video scene graph generation (VidSGG) has emerged as a transformative approach to capturing and interpreting the intricate relationships among objects and their temporal dynamics in video sequences. In this paper, we introduce the new AeroEye dataset that focuses on multi-object relationship modeling in aerial videos. Our AeroEye dataset features various drone scenes and includes a visually comprehensive and precise collection of predicates that capture the intricate relationships and spatial arrangements among objects. To this end, we propose the novel Cyclic Graph Transformer (CYCLO) approach that allows the model to capture both direct and long-range temporal dependencies by continuously updating the history of interactions in a circular manner. The proposed approach also allows one to handle sequences with inherent cyclical patterns and process object relationships in the correct sequential order. Therefore, it can effectively capture periodic and overlapping relationships while minimizing information loss. The extensive experiments on the AeroEye dataset demonstrate the effectiveness of the proposed CYCLO model, demonstrating its potential to perform scene understanding on drone videos. Finally, the CYCLO method consistently achieves State-of-the-Art (SOTA) results on two in-the-wild scene graph generation benchmarks, i. e. , PVSG and ASPIRe.

NeurIPS Conference 2024 Conference Paper

Decompose, Analyze and Rethink: Solving Intricate Problems with Human-like Reasoning Cycle

  • Shangzi Xue
  • Zhenya Huang
  • Jiayu Liu
  • Xin Lin
  • Yuting Ning
  • Binbin Jin
  • Xin Li
  • Qi Liu

In this paper, we introduce DeAR ( Decompose-Analyze-Rethink ), a framework that iteratively builds a reasoning tree to tackle intricate problems within a single large language model (LLM). Unlike approaches that extend or search for rationales, DeAR is featured by 1) adopting a tree-based question decomposition manner to plan the organization of rationales, which mimics the logical planning inherentin human cognition; 2) globally updating the rationales at each reasoning step through natural language feedback. Specifically, the Decompose stage decomposes the question into simpler sub-questions, storing them as new nodes; the Analyze stage generates and self-checks rationales for sub-questions at each node evel; and the Rethink stage updates parent-node rationales based on feedback from their child nodes. By generating and updating the reasoning process from a more global perspective, DeAR constructs more adaptive and accurate logical structures for complex problems, facilitating timely error correction compared to rationale-extension and search-based approaches such as Tree-of-Thoughts (ToT) and Graph-of-Thoughts (GoT). We conduct extensive experiments on three reasoning benchmarks, including ScienceQA, StrategyQA, and GSM8K, which cover a variety of reasoning tasks, demonstrating that our approach significantly reduces logical errors and enhances performance across various LLMs. Furthermore, we validate that DeAR is an efficient method that achieves a superior trade-off between accuracy and reasoning time compared to ToT and GoT.

EAAI Journal 2024 Journal Article

Dual-domain sampling and feature-domain optimization network for image compressive sensing

  • Xinxin Xiang
  • Fenghua Tong
  • Dawei Zhao
  • Xin Li
  • Shumian Yang

Recently, deep unfolding networks have become the mainstream approach for deep compressed sensing methods due to their good interpretability and high reconstruction performance. However, these existing deep unfolding networks, which usually rely on the pixel domain, face limitations in capturing additional image feature information during the sampling phase and fail to fully exploit feature information during the reconstruction phase. In this paper, we introduce a novel approach called Dual-domain Sampling and Feature-domain Optimization Network (DSFONet) to address these challenges. In the sampling sub-network, we propose a dual-domain sampling technique, by first fusing the pixel domain with the feature domain and then performing physical sampling, to sample richer image feature information. In the reconstruction sub-network, we ensure that the deep network’s increasing depth does not lead to the loss of crucial image feature information. This is achieved by preserving and propagating one-dimensional and multi-dimensional feature information of both inter-stages and sub-stages. Experiment results demonstrate that the proposed DSFONet outperforms the existing state-of-the-art methods.

ICML Conference 2024 Conference Paper

From Fourier to Neural ODEs: Flow Matching for Modeling Complex Systems

  • Xin Li
  • Jingdong Zhang 0001
  • Qunxi Zhu
  • Chengli Zhao
  • Xue Zhang
  • Xiaojun Duan
  • Wei Lin 0003

Modeling complex systems using standard neural ordinary differential equations (NODEs) often faces some essential challenges, including high computational costs and susceptibility to local optima. To address these challenges, we propose a simulation-free framework, called Fourier NODEs (FNODEs), that effectively trains NODEs by directly matching the target vector field based on Fourier analysis. Specifically, we employ the Fourier analysis to estimate temporal and potential high-order spatial gradients from noisy observational data. We then incorporate the estimated spatial gradients as additional inputs to a neural network. Furthermore, we utilize the estimated temporal gradient as the optimization objective for the output of the neural network. Later, the trained neural network generates more data points through an ODE solver without participating in the computational graph, facilitating more accurate estimations of gradients based on Fourier analysis. These two steps form a positive feedback loop, enabling accurate dynamics modeling in our framework. Consequently, our approach outperforms state-of-the-art methods in terms of training time, dynamics prediction, and robustness. Finally, we demonstrate the superior performance of our framework using a number of representative complex systems.

AAAI Conference 2024 Conference Paper

Grab What You Need: Rethinking Complex Table Structure Recognition with Flexible Components Deliberation

  • Hao Liu
  • Xin Li
  • Mingming Gong
  • Bing Liu
  • Yunfei Wu
  • Deqiang Jiang
  • Yinsong Liu
  • Xing Sun

Recently, Table Structure Recognition (TSR) task, aiming at identifying table structure into machine readable formats, has received increasing interest in the community. While impressive success, most single table component-based methods can not perform well on unregularized table cases distracted by not only complicated inner structure but also exterior capture distortion. In this paper, we raise it as Complex TSR problem, where the performance degeneration of existing methods is attributable to their inefficient component usage and redundant post-processing. To mitigate it, we shift our perspective from table component extraction towards the efficient multiple components leverage, which awaits further exploration in the field. Specifically, we propose a seminal method, termed GrabTab, equipped with newly proposed Component Deliberator, to handle various types of tables in a unified framework. Thanks to its progressive deliberation mechanism, our GrabTab can flexibly accommodate to most complex tables with reasonable components selected but without complicated post-processing involved. Quantitative experimental results on public benchmarks demonstrate that our method significantly outperforms the state-of-the-arts, especially under more challenging scenes.

AAAI Conference 2024 Conference Paper

Hyp-OW: Exploiting Hierarchical Structure Learning with Hyperbolic Distance Enhances Open World Object Detection

  • Thang Doan
  • Xin Li
  • Sima Behpour
  • Wenbin He
  • Liang Gou
  • Liu Ren

Open World Object Detection (OWOD) is a challenging and realistic task that extends beyond the scope of standard Object Detection task. It involves detecting both known and unknown objects while integrating learned knowledge for future tasks. However, the level of "unknownness" varies significantly depending on the context. For example, a tree is typically considered part of the background in a self-driving scene, but it may be significant in a household context. We argue that this contextual information should already be embedded within the known classes. In other words, there should be a semantic or latent structure relationship between the known and unknown items to be discovered. Motivated by this observation, we propose Hyp-OW, a method that learns and models hierarchical representation of known items through a SuperClass Regularizer. Leveraging this representation allows us to effectively detect unknown objects using a similarity distance-based relabeling module. Extensive experiments on benchmark datasets demonstrate the effectiveness of Hyp-OW, achieving improvement in both known and unknown detection (up to 6 percent). These findings are particularly pronounced in our newly designed benchmark, where a strong hierarchical structure exists between known and unknown objects.

AAAI Conference 2024 Conference Paper

Improving GNN Calibration with Discriminative Ability: Insights and Strategies

  • Yujie Fang
  • Xin Li
  • Qianyu Chen
  • Mingzhong Wang

The widespread adoption of Graph Neural Networks (GNNs) has led to an increasing focus on their reliability. To address the issue of underconfidence in GNNs, various calibration methods have been developed to gain notable reductions in calibration error. However, we observe that existing approaches generally fail to enhance consistently, and in some cases even deteriorate, GNNs' ability to discriminate between correct and incorrect predictions. In this study, we advocate the significance of discriminative ability and the inclusion of relevant evaluation metrics. Our rationale is twofold: 1) Overlooking discriminative ability can inadvertently compromise the overall quality of the model; 2) Leveraging discriminative ability can significantly inform and improve calibration outcomes. Therefore, we thoroughly explore the reasons why existing calibration methods have ineffectiveness and even degradation regarding the discriminative ability of GNNs. Building upon these insights, we conduct GNN calibration experiments across multiple datasets using a straightforward example model, denoted as DC(GNN). Its excellent performance confirms the potential of integrating discriminative ability as a key consideration in the calibration of GNNs, thereby establishing a pathway toward more effective and reliable network calibration.

AAAI Conference 2024 Conference Paper

Inverse Weight-Balancing for Deep Long-Tailed Learning

  • Wenqi Dang
  • Zhou Yang
  • Weisheng Dong
  • Xin Li
  • Guangming Shi

The performance of deep learning models often degrades rapidly when faced with imbalanced data characterized by a long-tailed distribution. Researchers have found that the fully connected layer trained by cross-entropy loss has large weight-norms for classes with many samples, but not for classes with few samples. How to address the data imbalance problem with both the encoder and the classifier seems an under-researched problem. In this paper, we propose an inverse weight-balancing (IWB) approach to guide model training and alleviate the data imbalance problem in two stages. In the first stage, an encoder and classifier (the fully connected layer) are trained using conventional cross-entropy loss. In the second stage, with a fixed encoder, the classifier is finetuned through an adaptive distribution for IWB in the decision space. Unlike existing inverse image frequency that implements a multiplicative margin adjustment transformation in the classification layer, our approach can be interpreted as an adaptive distribution alignment strategy using not only the class-wise number distribution but also the sample-wise difficulty distribution in both encoder and classifier. Experiments show that our method can greatly improve performance on imbalanced datasets such as CIFAR100-LT with different imbalance factors, ImageNet-LT, and iNaturelists2018.

YNIMG Journal 2024 Journal Article

Low-intensity transcranial ultrasound stimulation improves memory behavior in an ADHD rat model by modulating cortical functional network connectivity

  • Mengran Wang
  • Zhenyu Xie
  • Teng Wang
  • Shuxun Dong
  • Zhenfang Ma
  • Xiangjian Zhang
  • Xin Li
  • Yi Yuan

Working memory in attention deficit hyperactivity disorder (ADHD) is closely related to cortical functional network connectivity (CFNC), such as abnormal connections between the frontal, temporal, occipital cortices and with other brain regions. Low-intensity transcranial ultrasound stimulation (TUS) has the advantages of non-invasiveness, high spatial resolution, and high penetration depth and can improve ADHD memory behavior. However, how it modulates CFNC in ADHD and the CFNC mechanism that improves working memory behavior in ADHD remain unclear. In this study, we observed working memory impairment in ADHD rats, establishing a corresponding relationship between changes in CFNCs and the behavioral state during the working memory task. Specifically, we noted abnormalities in the information transmission and processing capabilities of CFNC in ADHD rats while performing working memory tasks. These abnormalities manifested in the network integration ability of specific areas, as well as the information flow and functional differentiation of CFNC. Furthermore, our findings indicate that TUS effectively enhances the working memory ability of ADHD rats by modulating information transmission, processing, and integration capabilities, along with adjusting the information flow and functional differentiation of CFNC. Additionally, we explain the CFNC mechanism through which TUS improves working memory in ADHD. In summary, these findings suggest that CFNCs are important in working memory behaviors in ADHD.

NeurIPS Conference 2024 Conference Paper

MagR: Weight Magnitude Reduction for Enhancing Post-Training Quantization

  • Aozhong Zhang
  • Naigang Wang
  • Yanxia Deng
  • Xin Li
  • Zi Yang
  • Penghang Yin

In this paper, we present a simple optimization-based preprocessing technique called Weight Magnitude Reduction (MagR) to improve the performance of post-training quantization. For each linear layer, we adjust the pre-trained floating-point weights by solving an $\ell_\infty$-regularized optimization problem. This process greatly diminishes the maximum magnitude of the weights and smooths out outliers, while preserving the layer's output. The preprocessed weights are centered more towards zero, which facilitates the subsequent quantization process. To implement MagR, we address the $\ell_\infty$-regularization by employing an efficient proximal gradient descent algorithm. Unlike existing preprocessing methods that involve linear transformations and subsequent post-processing steps, which can introduce significant overhead at inference time, MagR functions as a non-linear transformation, eliminating the need for any additional post-processing. This ensures that MagR introduces no overhead whatsoever during inference. Our experiments demonstrate that MagR achieves state-of-the-art performance on the Llama family of models. For example, we achieve a Wikitext2 perplexity of 6. 7 on the LLaMA2-70B model for per-channel INT2 weight quantization without incurring any inference overhead.

AAAI Conference 2024 Conference Paper

MetaCARD: Meta-Reinforcement Learning with Task Uncertainty Feedback via Decoupled Context-Aware Reward and Dynamics Components

  • Min Wang
  • Xin Li
  • Leiji Zhang
  • Mingzhong Wang

Meta-Reinforcement Learning (Meta-RL) aims to reveal shared characteristics in dynamics and reward functions across diverse training tasks. This objective is achieved by meta-learning a policy that is conditioned on task representations with encoded trajectory data or context, thus allowing rapid adaptation to new tasks from a known task distribution. However, since the trajectory data generated by the policy may be biased, the task inference module tends to form spurious correlations between trajectory data and specific tasks, thereby leading to poor adaptation to new tasks. To address this issue, we propose the Meta-RL with task unCertAinty feedback through decoupled context-aware Reward and Dynamics components (MetaCARD). MetaCARD distinctly decouples the dynamics and rewards when inferring tasks and integrates task uncertainty feedback from policy evaluation into the task inference module. This design effectively reduces uncertainty in tasks with changes in dynamics or/and reward functions, thereby enabling accurate task identification and adaptation. The experiment results on both Meta-World and classical MuJoCo benchmarks show that MetaCARD significantly outperforms prevailing Meta-RL baselines, demonstrating its remarkable adaptation ability in sophisticated environments that involve changes in both reward functions and dynamics.

AAAI Conference 2024 Conference Paper

MobileInst: Video Instance Segmentation on the Mobile

  • Renhong Zhang
  • Tianheng Cheng
  • Shusheng Yang
  • Haoyi Jiang
  • Shuai Zhang
  • Jiancheng Lyu
  • Xin Li
  • Xiaowen Ying

Video instance segmentation on mobile devices is an important yet very challenging edge AI problem. It mainly suffers from (1) heavy computation and memory costs for frame-by-frame pixel-level instance perception and (2) complicated heuristics for tracking objects. To address these issues, we present MobileInst, a lightweight and mobile-friendly framework for video instance segmentation on mobile devices. Firstly, MobileInst adopts a mobile vision transformer to extract multi-level semantic features and presents an efficient query-based dual-transformer instance decoder for mask kernels and a semantic-enhanced mask decoder to generate instance segmentation per frame. Secondly, MobileInst exploits simple yet effective kernel reuse and kernel association to track objects for video instance segmentation. Further, we propose temporal query passing to enhance the tracking ability for kernels. We conduct experiments on COCO and YouTube-VIS datasets to demonstrate the superiority of MobileInst and evaluate the inference latency on one single CPU core of the Snapdragon 778G Mobile Platform, without other methods of acceleration. On the COCO dataset, MobileInst achieves 31.2 mask AP and 433 ms on the mobile CPU, which reduces the latency by 50% compared to the previous SOTA. For video instance segmentation, MobileInst achieves 35.0 AP and 30.1 AP on YouTube-VIS 2019 & 2021.

NeurIPS Conference 2024 Conference Paper

MOTE-NAS: Multi-Objective Training-based Estimate for Efficient Neural Architecture Search

  • Yu-Ming Zhang
  • Jun-Wei Hsieh
  • Xin Li
  • Ming-Ching Chang
  • Chun-Chieh Lee
  • Kuo-Chin Fan

Neural Architecture Search (NAS) methods seek effective optimization toward performance metrics regarding model accuracy and generalization while facing challenges regarding search costs and GPU resources. Recent Neural Tangent Kernel (NTK) NAS methods achieve remarkable search efficiency based on a training-free model estimate; however, they overlook the non-convex nature of the DNNs in the search process. In this paper, we develop Multi-Objective Training-based Estimate (MOTE) for efficient NAS, retaining search effectiveness and achieving the new state-of-the-art in the accuracy and cost trade-off. To improve NTK and inspired by the Training Speed Estimation (TSE) method, MOTE is designed to model the actual performance of DNNs from macro to micro perspective by draw loss landscape and convergence speed simultaneously. Using two reduction strategies, the MOTE is generated based on a reduced architecture and a reduced dataset. Inspired by evolutionary search, our iterative ranking-based, coarse-to-fine architecture search is highly effective. Experiments on NASBench-201 show MOTE-NAS achieves 94. 32% accuracy on CIFAR-10, 72. 81% on CIFAR-100, and 46. 38% on ImageNet-16-120, outperforming NTK-based NAS approaches. An evaluation-free (EF) version of MOTE-NAS delivers high efficiency in only 5 minutes, delivering a model more accurate than KNAS.

AAAI Conference 2024 Conference Paper

Pushing the Limit of Fine-Tuning for Few-Shot Learning: Where Feature Reusing Meets Cross-Scale Attention

  • Ying-Yu Chen
  • Jun-Wei Hsieh
  • Xin Li
  • Ming-Ching Chang

Due to the scarcity of training samples, Few-Shot Learning (FSL) poses a significant challenge to capture discriminative object features effectively. The combination of transfer learning and meta-learning has recently been explored by pre-training the backbone features using labeled base data and subsequently fine-tuning the model with target data. However, existing meta-learning methods, which use embedding networks, suffer from scaling limitations when dealing with a few labeled samples, resulting in suboptimal results. Inspired by the latest advances in FSL, we further advance the approach of fine-tuning a pre-trained architecture by a strengthened hierarchical feature representation. The technical contributions of this work include: 1) a hybrid design named Intra-Block Fusion (IBF) to strengthen the extracted features within each convolution block; and 2) a novel Cross-Scale Attention (CSA) module to mitigate the scaling inconsistencies arising from the limited training samples, especially for cross-domain tasks. We conducted comprehensive evaluations on standard benchmarks, including three in-domain tasks (miniImageNet, CIFAR-FS, and FC100), as well as two cross-domain tasks (CDFSL and Meta-Dataset). The results have improved significantly over existing state-of-the-art approaches on all benchmark datasets. In particular, the FSL performance on the in-domain FC100 dataset is more than three points better than the latest PMF (Hu et al. 2022).

EAAI Journal 2024 Journal Article

Reinforcement learning control with n-step information for wastewater treatment systems

  • Xin Li
  • Ding Wang
  • Mingming Zhao
  • Junfei Qiao

Wastewater treatment is important for maintaining a balanced urban ecosystem. To ensure the success of wastewater treatment, the tracking error between the crucial variable concentrations and the set point needs to be minimized as much as possible. Since the multiple biochemical reactions are involved, the wastewater treatment system is a nonlinear system with unknown dynamics. For this class of systems, this paper develops an online action dependent heuristic dynamic programming (ADHDP) algorithm combining the temporal difference with λ [TD( λ )], which is called ADHDP( λ ). By introducing the TD( λ ), the future n -step information is considered and the learning efficiency of the ADHDP algorithm is improved. We not only give the implementation process of the ADHDP( λ ) algorithm based on neural networks, but also prove the stability of the algorithm under certain conditions. Finally, the effectiveness of the ADHDP( λ ) algorithm is verified through two nonlinear systems, including a wastewater treatment system and a torsional pendulum system. Simulation results show that the ADHDP( λ ) algorithm has higher learning efficiency compared to the general ADHDP algorithm.

IJCAI Conference 2024 Conference Paper

Revitalizing Real Image Deraining via a Generic Paradigm towards Multiple Rainy Patterns

  • Xin Li
  • Yuxin Feng
  • Fan Zhou
  • Yun Liang
  • Zhuo Su

Synthetic data-driven methods perform well on image rain removal task, but they still face many challenges in real rainfall scenarios due to the complexity and diversity of rainy patterns. In this paper, we propose a new generic paradigm for real image deraining from the perspective of synthesizing data covering more rainy patterns and constructing image rain removal networks with strong generalization performance. Firstly, instead of simply superimposing rain layers, we integrate various rainy patterns and design a phenomenal pipeline that incorporates multiple degradation types. Secondly, we construct a Patterns-aware Rain Removal Network (PRRN), which learns from both synthetic and real data simultaneously. In addition, to eliminate the inevitable distribution differences between synthetic and real data, we design a new Multi-representation Inter-domain Alignment Module (MIAM) in PRRN. By using multiple parallel submodules, MIAM achieves alignment of data domains in multiple feature subspaces. Based on several authoritative objective evaluation metrics, we successfully validate the effectiveness and robustness of the proposed method in real scenarios through extensive experiments carried out on five challenging real datasets.

NeurIPS Conference 2024 Conference Paper

Skill-aware Mutual Information Optimisation for Zero-shot Generalisation in Reinforcement Learning

  • Xuehui Yu
  • Mhairi Dunion
  • Xin Li
  • Stefano V. Albrecht

Meta-Reinforcement Learning (Meta-RL) agents can struggle to operate across tasks with varying environmental features that require different optimal skills (i. e. , different modes of behaviour). Using context encoders based on contrastive learning to enhance the generalisability of Meta-RL agents is now widely studied but faces challenges such as the requirement for a large sample size, also referred to as the $\log$-$K$ curse. To improve RL generalisation to different tasks, we first introduce Skill-aware Mutual Information (SaMI), an optimisation objective that aids in distinguishing context embeddings according to skills, thereby equipping RL agents with the ability to identify and execute different skills across tasks. We then propose Skill-aware Noise Contrastive Estimation (SaNCE), a $K$-sample estimator used to optimise the SaMI objective. We provide a framework for equipping an RL agent with SaNCE in practice and conduct experimental validation on modified MuJoCo and Panda-gym benchmarks. We empirically find that RL agents that learn by maximising SaMI achieve substantially improved zero-shot generalisation to unseen tasks. Additionally, the context encoder trained with SaNCE demonstrates greater robustness to a reduction in the number of available samples, thus possessing the potential to overcome the $\log$-$K$ curse.

AAAI Conference 2024 Conference Paper

SMILEtrack: SiMIlarity LEarning for Occlusion-Aware Multiple Object Tracking

  • Yu-Hsiang Wang
  • Jun-Wei Hsieh
  • Ping-Yang Chen
  • Ming-Ching Chang
  • Hung-Hin So
  • Xin Li

Despite recent progress in Multiple Object Tracking (MOT), several obstacles such as occlusions, similar objects, and complex scenes remain an open challenge. Meanwhile, a systematic study of the cost-performance tradeoff for the popular tracking-by-detection paradigm is still lacking. This paper introduces SMILEtrack, an innovative object tracker that effectively addresses these challenges by integrating an efficient object detector with a Siamese network-based Similarity Learning Module (SLM). The technical contributions of SMILETrack are twofold. First, we propose an SLM that calculates the appearance similarity between two objects, overcoming the limitations of feature descriptors in Separate Detection and Embedding (SDE) models. The SLM incorporates a Patch Self-Attention (PSA) block inspired by the vision Transformer, which generates reliable features for accurate similarity matching. Second, we develop a Similarity Matching Cascade (SMC) module with a novel GATE function for robust object matching across consecutive video frames, further enhancing MOT performance. Together, these innovations help SMILETrack achieve an improved trade-off between the cost (e.g., running speed) and performance (e.g., tracking accuracy) over several existing state-of-the-art benchmarks, including the popular BYTETrack method. SMILETrack outperforms BYTETrack by 0.4-0.8 MOTA and 2.1-2.2 HOTA points on MOT17 and MOT20 datasets. Code is available at http://github.com/pingyang1117/SMILEtrack_official.

NeurIPS Conference 2024 Conference Paper

Stabilize the Latent Space for Image Autoregressive Modeling: A Unified Perspective

  • Yongxin Zhu
  • Bocheng Li
  • Hang Zhang
  • Xin Li
  • Linli Xu
  • Lidong Bing

Latent-based image generative models, such as Latent Diffusion Models (LDMs) and Mask Image Models (MIMs), have achieved notable success in image generation tasks. These models typically leverage reconstructive autoencoders like VQGAN or VAE to encode pixels into a more compact latent space and learn the data distribution in the latent space instead of directly from pixels. However, this practice raises a pertinent question: Is it truly the optimal choice? In response, we begin with an intriguing observation: despite sharing the same latent space, autoregressive models significantly lag behind LDMs and MIMs in image generation. This finding contrasts sharply with the field of NLP, where the autoregressive model GPT has established a commanding presence. To address this discrepancy, we introduce a unified perspective on the relationship between latent space and generative models, emphasizing the stability of latent space in image generative modeling. Furthermore, we propose a simple but effective discrete image tokenizer to stabilize the latent space for image generative modeling by applying K-Means on the latent features of self-supervised learning models. Experimental results show that image autoregressive modeling with our tokenizer (DiGIT) benefits both image understanding and image generation with the next token prediction principle, which is inherently straightforward for GPT models but challenging for other generative models. Remarkably, for the first time, a GPT-style autoregressive model for images outperforms LDMs, which also exhibits substantial improvement akin to GPT when scaling up model size. Our findings underscore the potential of an optimized latent space and the integration of discrete tokenization in advancing the capabilities of image generative models. The code is available at \url{https: //github. com/DAMO-NLP-SG/DiGIT}.

AAAI Conference 2024 Conference Paper

Sunshine to Rainstorm: Cross-Weather Knowledge Distillation for Robust 3D Object Detection

  • Xun Huang
  • Hai Wu
  • Xin Li
  • Xiaoliang Fan
  • Chenglu Wen
  • Cheng Wang

LiDAR-based 3D object detection models inevitably struggle under rainy conditions due to the degraded and noisy scanning signals. Previous research has attempted to address this by simulating the noise from rain to improve the robustness of detection models. However, significant disparities exist between simulated and actual rain-impacted data points. In this work, we propose a novel rain simulation method, termed DRET, that unifies Dynamics and Rainy Environment Theory to provide a cost-effective means of expanding the available realistic rain data for 3D detection training. Furthermore, we present a Sunny-to-Rainy Knowledge Distillation (SRKD) approach to enhance 3D detection under rainy conditions. Extensive experiments on the Waymo-Open-Dataset show that, when combined with the state-of-the-art DSVT model and other classical 3D detectors, our proposed framework demonstrates significant detection accuracy improvements, without losing efficiency. Remarkably, our framework also improves detection capabilities under sunny conditions, therefore offering a robust solution for 3D detection regardless of whether the weather is rainy or sunny.

YNIMG Journal 2024 Journal Article

Voxel-based texture similarity networks reveal individual variability and correlate with biological ontologies

  • Liyuan Lin
  • Zhongyu Chang
  • Yu Zhang
  • Kaizhong Xue
  • Yingying Xie
  • Luli Wei
  • Xin Li
  • Zhen Zhao

The human brain is organized as a complex, hierarchical network. However, the structural covariance patterns among brain regions and the underlying biological substrates of such covariance networks remain to be clarified. The present study proposed a novel individualized structural covariance network termed voxel-based texture similarity networks (vTSNs) based on 76 refined voxel-based textural features derived from structural magnetic resonance images. Validated in three independent longitudinal healthy cohorts (40, 23, and 60 healthy participants, respectively) with two common brain atlases, we found that the vTSN could robustly resolve inter-subject variability with high test-retest reliability. In contrast to the regional-based texture similarity networks (rTSNs) that calculate radiomic features based on region-of-interest information, vTSNs had higher inter- and intra-subject variability ratios and test-retest reliability in connectivity strength and network topological properties. Moreover, the Spearman correlation indicated a stronger association of the gene expression similarity network (GESN) with vTSNs than with rTSNs (vTSN: r = 0.600, rTSN: r = 0.433, z = 39.784, P < 0.001). Hierarchical clustering identified 3 vTSN subnets with differential association patterns with 13 coexpression modules, 16 neurotransmitters, 7 electrophysiology, 4 metabolism, and 2 large-scale structural and 4 functional organization maps. Moreover, these subnets had unique biological hierarchical organization from the subcortex-limbic system to the ventral neocortex and then to the dorsal neocortex. Based on 424 unrelated, qualified healthy subjects from the Human Connectome Project, we found that vTSNs could sensitively represent sex differences, especially for connections in the subcortex-limbic system and between the subcortex-limbic system and the ventral neocortex. Moreover, a multivariate variance component model revealed that vTSNs could explain a significant proportion of inter-subject behavioral variance in cognition (80.0 %) and motor functions (63.4 %). Finally, using 494 healthy adults (aged 19-80 years old) from the Southwest University Adult Lifespan Dataset, the Spearman correlation identified a significant association between aging and vTSN strength, especially within the subcortex-limbic system and between the subcortex-limbic system and the dorsal neocortex. In summary, our proposed vTSN is robust in uncovering individual variability and neurobiological brain processes, which can serve as biologically plausible measures for linking biological processes and human behavior.

AAAI Conference 2024 Conference Paper

Zero-1-to-3: Domain-Level Zero-Shot Cognitive Diagnosis via One Batch of Early-Bird Students towards Three Diagnostic Objectives

  • Weibo Gao
  • Qi Liu
  • Hao Wang
  • Linan Yue
  • Haoyang Bi
  • Yin Gu
  • Fangzhou Yao
  • Zheng Zhang

Cognitive diagnosis seeks to estimate the cognitive states of students by exploring their logged practice quiz data. It plays a pivotal role in personalized learning guidance within intelligent education systems. In this paper, we focus on an important, practical, yet often underexplored task: domain-level zero-shot cognitive diagnosis (DZCD), which arises due to the absence of student practice logs in newly launched domains. Recent cross-domain diagnostic models have been demonstrated to be a promising strategy for DZCD. These methods primarily focus on how to transfer student states across domains. However, they might inadvertently incorporate non-transferable information into student representations, thereby limiting the efficacy of knowledge transfer. To tackle this, we propose Zero-1-to-3, a domain-level zero-shot cognitive diagnosis framework via one batch of early-bird students towards three diagnostic objectives. Our approach initiates with pre-training a diagnosis model with dual regularizers, which decouples student states into domain-shared and domain-specific parts. The shared cognitive signals can be transferred to the target domain, enriching the cognitive priors for the new domain, which ensures the cognitive state propagation objective. Subsequently, we devise a strategy to generate simulated practice logs for cold-start students through analyzing the behavioral patterns from early-bird students, fulfilling the domain-adaption goal. Consequently, we refine the cognitive states of cold-start students as diagnostic outcomes via virtual data, aligning with the diagnosis-oriented goal. Finally, extensive experiments on six real-world datasets highlight the efficacy of our model for DZCD and its practical application in question recommendation. The code is publicly available at https://github.com/bigdata-ustc/Zero-1-to-3.

NeurIPS Conference 2023 Conference Paper

A Bounded Ability Estimation for Computerized Adaptive Testing

  • Yan Zhuang
  • Qi Liu
  • Guanhao Zhao
  • Zhenya Huang
  • Weizhe Huang
  • Zachary Pardos
  • Enhong Chen
  • Jinze Wu

Computerized adaptive testing (CAT), as a tool that can efficiently measure student's ability, has been widely used in various standardized tests (e. g. , GMAT and GRE). The adaptivity of CAT refers to the selection of the most informative questions for each student, reducing test length. Existing CAT methods do not explicitly target ability estimation accuracy since there is no student's true ability as ground truth; therefore, these methods cannot be guaranteed to make the estimate converge to the true with such limited responses. In this paper, we analyze the statistical properties of estimation and find a theoretical approximation of the true ability: the ability estimated by full responses to question bank. Based on this, a Bounded Ability Estimation framework for CAT (BECAT) is proposed in a data-summary manner, which selects a question subset that closely matches the gradient of the full responses. Thus, we develop an expected gradient difference approximation to design a simple greedy selection algorithm, and show the rigorous theoretical and error upper-bound guarantees of its ability estimate. Experiments on both real-world and synthetic datasets, show that it can reach the same estimation accuracy using 15\% less questions on average, significantly reducing test length.

AAAI Conference 2023 Conference Paper

AdaCM: Adaptive ColorMLP for Real-Time Universal Photo-Realistic Style Transfer

  • Tianwei Lin
  • Honglin Lin
  • Fu Li
  • Dongliang He
  • Wenhao Wu
  • Meiling Wang
  • Xin Li
  • Yong Liu

Photo-realistic style transfer aims at migrating the artistic style from an exemplar style image to a content image, producing a result image without spatial distortions or unrealistic artifacts. Impressive results have been achieved by recent deep models. However, deep neural network based methods are too expensive to run in real-time. Meanwhile, bilateral grid based methods are much faster but still contain artifacts like overexposure. In this work, we propose the Adaptive ColorMLP (AdaCM), an effective and efficient framework for universal photo-realistic style transfer. First, we find the complex non-linear color mapping between input and target domain can be efficiently modeled by a small multi-layer perceptron (ColorMLP) model. Then, in AdaCM, we adopt a CNN encoder to adaptively predict all parameters for the ColorMLP conditioned on each input content and style image pair. Experimental results demonstrate that AdaCM can generate vivid and high-quality stylization results. Meanwhile, our AdaCM is ultrafast and can process a 4K resolution image in 6ms on one V100 GPU.

AAAI Conference 2023 Conference Paper

Context-Aware Safe Medication Recommendations with Molecular Graph and DDI Graph Embedding

  • Qianyu Chen
  • Xin Li
  • Kunnan Geng
  • Mingzhong Wang

Molecular structures and Drug-Drug Interactions (DDI) are recognized as important knowledge to guide medication recommendation (MR) tasks, and medical concept embedding has been applied to boost their performance. Though promising performance has been achieved by leveraging Graph Neural Network (GNN) models to encode the molecular structures of medications or/and DDI, we observe that existing models are still defective: 1) to differentiate medications with similar molecules but different functionality; or/and 2) to properly capture the unintended reactions between drugs in the embedding space. To alleviate this limitation, we propose Carmen, a cautiously designed graph embedding-based MR framework. Carmen consists of four components, including patient representation learning, context information extraction, a context-aware GNN, and DDI encoding. Carmen incorporates the visit history into the representation learning of molecular graphs to distinguish molecules with similar topology but dissimilar activity. Its DDI encoding module is specially devised for the non-transitive interaction DDI graphs. The experiments on real-world datasets demonstrate that Carmen achieves remarkable performance improvement over state-of-the-art models and can improve the safety of recommended drugs with a proper DDI graph encoding.

EAAI Journal 2023 Journal Article

Data-driven tracking control design with reinforcement learning involving a wastewater treatment application

  • Ding Wang
  • Xin Li
  • Lingzhi Hu
  • Junfei Qiao

With the increase of urbanization rate, the problem of water shortage and pollution is more and more serious. It is important to improve the efficiency of wastewater treatment to protect the urban ecological environment. The wastewater treatment process involves a variety of biochemical reactions and has strong time-vary dynamics. The concentration design in the wastewater treatment process can be regarded as a tracking control problem for a class of nonlinear systems. In order to solve this problem, this paper develops an intelligent control method with tracking goal representation heuristic dynamic programming (T-GrHDP) by combining the GrHDP with a novel tracking framework. A model network is built by using a dataset consisting of real input and output data of the controlled object, which can overcome the dependence on the system dynamic. In order to improve the learning efficiency of the proposed algorithm, we introduce the goal network to provide more effective information for the critic network. The classical actor-critic scheme in reinforcement learning is used to obtain the approximate optimal control strategy. By introducing some necessary lemmas and assumptions, the convergence of the proposed algorithm is proved. Finally, the T-GrHDP method is successfully applied in two industrial simulations including the wastewater treatment system.

IJCAI Conference 2023 Conference Paper

Dual-view Correlation Hybrid Attention Network for Robust Holistic Mammogram Classification

  • Zhiwei Wang
  • Junlin Xian
  • Kangyi Liu
  • Xin Li
  • Qiang Li
  • Xin Yang

Mammogram image is important for breast cancer screening, and typically obtained in a dual-view form, i. e. , cranio-caudal (CC) and mediolateral oblique (MLO), to provide complementary information for clinical decisions. However, previous methods mostly learn features from the two views independently, which violates the clinical knowledge and ignores the importance of dual-view correlation in the feature learning. In this paper, we propose a dual-view correlation hybrid attention network (DCHA-Net) for robust holistic mammogram classification. Specifically, DCHA-Net is carefully designed to extract and reinvent deep feature maps for the two views, and meanwhile to maximize the underlying correlations between them. A hybrid attention module, consisting of local relation and non-local attention blocks, is proposed to alleviate the spatial misalignment of the paired views in the correlation maximization. A dual-view correlation loss is introduced to maximize the feature similarity between corresponding strip-like regions with equal distance to the chest wall, motivated by the fact that their features represent the same breast tissues, and thus should be highly-correlated with each other. Experimental results on the two public datasets, i. e. , INbreast and CBIS-DDSM, demonstrate that the DCHA-Net can well preserve and maximize feature correlations across views, and thus outperforms previous state-of-the-art methods for classifying a whole mammogram as malignant or not.

NeurIPS Conference 2023 Conference Paper

From Cloze to Comprehension: Retrofitting Pre-trained Masked Language Models to Pre-trained Machine Reader

  • Weiwen Xu
  • Xin Li
  • Wenxuan Zhang
  • Meng Zhou
  • Wai Lam
  • Luo Si
  • Lidong Bing

We present Pre-trained Machine Reader (PMR), a novel method for retrofitting pre-trained masked language models (MLMs) to pre-trained machine reading comprehension (MRC) models without acquiring labeled data. PMR can resolve the discrepancy between model pre-training and downstream fine-tuning of existing MLMs. To build the proposed PMR, we constructed a large volume of general-purpose and high-quality MRC-style training data by using Wikipedia hyperlinks and designed a Wiki Anchor Extraction task to guide the MRC-style pre-training. Apart from its simplicity, PMR effectively solves extraction tasks, such as Extractive Question Answering and Named Entity Recognition. PMR shows tremendous improvements over existing approaches, especially in low-resource scenarios. When applied to the sequence classification task in the MRC formulation, PMR enables the extraction of high-quality rationales to explain the classification process, thereby providing greater prediction explainability. PMR also has the potential to serve as a unified model for tackling various extraction and classification tasks in the MRC formulation.

NeurIPS Conference 2023 Conference Paper

GradOrth: A Simple yet Efficient Out-of-Distribution Detection with Orthogonal Projection of Gradients

  • Sima Behpour
  • Thang Long Doan
  • Xin Li
  • Wenbin He
  • Liang Gou
  • Liu Ren

Detecting out-of-distribution (OOD) data is crucial for ensuring the safe deployment of machine learning models in real-world applications. However, existing OOD detection approaches primarily rely on the feature maps or the full gradient space information to derive OOD scores neglecting the role of \textbf{most important parameters} of the pre-trained network over In-Distribution data. In this study, we propose a novel approach called GradOrth to facilitate OOD detection based on one intriguing observation that the important features to identify OOD data lie in the lower-rank subspace of in-distribution (ID) data. In particular, we identify OOD data by computing the norm of gradient projection on \textit{the subspaces considered \textbf{important} for the in-distribution data}. A large orthogonal projection value (i. e. a small projection value) indicates the sample as OOD as it captures a weak correlation of the in-distribution (ID) data. This simple yet effective method exhibits outstanding performance, showcasing a notable reduction in the average false positive rate at a 95\% true positive rate (FPR95) of up to 8\% when compared to the current state-of-the-art methods.

NeurIPS Conference 2023 Conference Paper

GraphAdapter: Tuning Vision-Language Models With Dual Knowledge Graph

  • Xin Li
  • Dongze Lian
  • Zhihe Lu
  • Jiawang Bai
  • Zhibo Chen
  • Xinchao Wang

Adapter-style efficient transfer learning (ETL) has shown excellent performance in the tuning of vision-language models (VLMs) under the low-data regime, where only a few additional parameters are introduced to excavate the task-specific knowledge based on the general and powerful representation of VLMs. However, most adapter-style works face two limitations: (i) modeling task-specific knowledge with a single modality only; and (ii) overlooking the exploitation of the inter-class relationships in downstream tasks, thereby leading to sub-optimal solutions. To mitigate that, we propose an effective adapter-style tuning strategy, dubbed GraphAdapter, which performs the textual adapter by explicitly modeling the dual-modality structure knowledge (i. e. , the correlation of different semantics/classes in textual and visual modalities) with a dual knowledge graph. In particular, the dual knowledge graph is established with two sub-graphs, i. e. , a textual knowledge sub-graph, and a visual knowledge sub-graph, where the nodes and edges represent the semantics/classes and their correlations in two modalities, respectively. This enables the textual feature of each prompt to leverage the task-specific structure knowledge from both textual and visual modalities, yielding a more effective classifier for downstream tasks. Extensive experimental results on 11 benchmark datasets reveal that our GraphAdapter significantly outperforms the previous adapter-based methods.

YNIMG Journal 2023 Journal Article

Hub architecture of the human structural connectome: Links to aging and processing speed

  • Xin Li
  • Alireza Salami
  • Jonas Persson

The human structural brain network, or connectome, has a rich-club organization with a small number of brain regions showing high network connectivity, called hubs. Hubs are centrally located in the network, energy costly, and critical for human cognition. Aging has been associated with changes in brain structure, function, and cognitive decline, such as processing speed. At a molecular level, the aging process is a progressive accumulation of oxidative damage, which leads to subsequent energy depletion in the neuron and causes cell death. However, it is still unclear how age affects hub connections in the human connectome. The current study aims to address this research gap by constructing structural connectome using fiber bundle capacity (FBC). FBC is derived from Constrained Spherical Deconvolution (CSD) modeling of white-matter fiber bundles, which represents the capacity of a fiber bundle to transfer information. Compared to the raw number of streamlines, FBC is less bias for quantifying connection strength within biological pathways. We found that hubs exhibit longer-distance connections and higher metabolic rates compared to peripheral brain regions, suggesting that hubs are biologically costly. Although the landscape of structural hubs was relatively age-invariant, there were wide-spread age effects on FBC in the connectome. Critically, these age effects were larger in connections within hub compared to peripheral brain connections. These findings were supported by both a cross-sectional sample with wide age-range (N = 137) and a longitudinal sample across 5 years (N = 83). Moreover, our results demonstrated that associations between FBC and processing speed were more concentrated in hub connections than chance level, and FBC in hub connections mediated the age-effects on processing speed. Overall, our findings indicate that structural connections of hubs, which demonstrate greater energy demands, are particular vulnerable to aging. The vulnerability may contribute to age-related impairments in processing speed among older adults.

AAAI Conference 2023 Conference Paper

Learning Compact Features via In-Training Representation Alignment

  • Xin Li
  • Xiangrui Li
  • Deng Pan
  • Yao Qiang
  • Dongxiao Zhu

Deep neural networks (DNNs) for supervised learning can be viewed as a pipeline of the feature extractor (i.e., last hidden layer) and a linear classifier (i.e., output layer) that are trained jointly with stochastic gradient descent (SGD) on the loss function (e.g., cross-entropy). In each epoch, the true gradient of the loss function is estimated using a mini-batch sampled from the training set and model parameters are then updated with the mini-batch gradients. Although the latter provides an unbiased estimation of the former, they are subject to substantial variances derived from the size and number of sampled mini-batches, leading to noisy and jumpy updates. To stabilize such undesirable variance in estimating the true gradients, we propose In-Training Representation Alignment (ITRA) that explicitly aligns feature distributions of two different mini-batches with a matching loss in the SGD training process. We also provide a rigorous analysis of the desirable effects of the matching loss on feature representation learning: (1) extracting compact feature representation; (2) reducing over-adaption on mini-batches via an adaptively weighting mechanism; and (3) accommodating to multi-modalities. Finally, we conduct large-scale experiments on both image and text classifications to demonstrate its superior performance to the strong baselines.

AAAI Conference 2023 Conference Paper

Locate Then Generate: Bridging Vision and Language with Bounding Box for Scene-Text VQA

  • Yongxin Zhu
  • Zhen Liu
  • Yukang Liang
  • Xin Li
  • Hao Liu
  • Changcun Bao
  • Linli Xu

In this paper, we propose a novel multi-modal framework for Scene Text Visual Question Answering (STVQA), which requires models to read scene text in images for question answering. Apart from text or visual objects, which could exist independently, scene text naturally links text and visual modalities together by conveying linguistic semantics while being a visual object in an image simultaneously. Different to conventional STVQA models which take the linguistic semantics and visual semantics in scene text as two separate features, in this paper, we propose a paradigm of "Locate Then Generate" (LTG), which explicitly unifies this two semantics with the spatial bounding box as a bridge connecting them. Specifically, at first, LTG locates the region in an image that may contain the answer words with an answer location module (ALM) consisting of a region proposal network and a language refinement network, both of which can transform to each other with one-to-one mapping via the scene text bounding box. Next, given the answer words selected by ALM, LTG generates a readable answer sequence with an answer generation module (AGM) based on a pre-trained language model. As a benefit of the explicit alignment of the visual and linguistic semantics, even without any scene text based pre-training tasks, LTG can boost the absolute accuracy by +6.06% and +6.92% on the TextVQA dataset and the ST-VQA dataset respectively, compared with a non-pre-training baseline. We further demonstrate that LTG effectively unifies visual and text modalities through the spatial bounding box connection, which is underappreciated in previous methods.

IJCAI Conference 2023 Conference Paper

Negative Flux Aggregation to Estimate Feature Attributions

  • Xin Li
  • Deng Pan
  • Chengyin Li
  • Yao Qiang
  • Dongxiao Zhu

There are increasing demands for understanding deep neural networks' (DNNs) behavior spurred by growing security and/or transparency concerns. Due to multi-layer nonlinearity of the deep neural network architectures, explaining DNN predictions still remains as an open problem, preventing us from gaining a deeper understanding of the mechanisms. To enhance the explainability of DNNs, we estimate the input feature's attributions to the prediction task using divergence and flux. Inspired by the divergence theorem in vector analysis, we develop a novel Negative Flux Aggregation (NeFLAG) formulation and an efficient approximation algorithm to estimate attribution map. Unlike the previous techniques, ours doesn't rely on fitting a surrogate model nor need any path integration of gradients. Both qualitative and quantitative experiments demonstrate a superior performance of NeFLAG in generating more faithful attribution maps than the competing methods. Our code is available at https: //github. com/xinli0928/NeFLAG.

TIST Journal 2023 Journal Article

Relation-aware Graph Convolutional Networks for Multi-relational Network Alignment

  • Yujie Fang
  • Xin Li
  • Rui Ye
  • Xiaoyan Tan
  • Peiyao Zhao
  • Mingzhong Wang

The alignment of multiple multi-relational networks, such as knowledge graphs, is vital for many AI applications. In comparison with existing GCNs which cannot fully utilize relational information of multiple types, we propose a relation-aware graph convolutional network (ERGCN), which is equipped with both entity convolution and relation convolution to learn the entity embeddings and relation embeddings simultaneously. The role discrimination and translation property of knowledge graphs are adopted in the entity convolutional process to incorporate the relation information. To facilitate the relation convolution, we construct quadruples to model the connection between a pair of relations thus to determine their neighborhood, which also enables the relation convolution to be conducted in an efficient way. Thereafter, AERGCN, the alignment framework based on ERGCN, is developed for multi-relational network alignment tasks. Anchors are used to supervise the objective function, which aims at minimizing the distances between anchors and to generate new cross-network triplets to build a bridge between different knowledge graphs at the level of triplet to improve the performance of alignment. Experiments on real-world datasets show that the proposed solutions outperform the competitive baselines in terms of link prediction, entity alignment, and relation alignment.

AAAI Conference 2023 Conference Paper

SelectAugment: Hierarchical Deterministic Sample Selection for Data Augmentation

  • Shiqi Lin
  • Zhizheng Zhang
  • Xin Li
  • Zhibo Chen

Data augmentation (DA) has been extensively studied to facilitate model optimization in many tasks. Prior DA works focus on designing augmentation operations themselves, while leaving selecting suitable samples for augmentation out of consideration. This might incur visual ambiguities and further induce training biases. In this paper, we propose an effective approach, dubbed SelectAugment, to select samples for augmentation in a deterministic and online manner based on the sample contents and the network training status. To facilitate the policy learning, in each batch, we exploit the hierarchy of this task by first determining the augmentation ratio and then deciding whether to augment each training sample under this ratio. We model this process as two-step decision-making and adopt Hierarchical Reinforcement Learning (HRL) to learn the selection policy. In this way, the negative effects of the randomness in selecting samples to augment can be effectively alleviated and the effectiveness of DA is improved. Extensive experiments demonstrate that our proposed SelectAugment significantly improves various off-the-shelf DA methods on image classification and fine-grained image recognition.

AAAI Conference 2023 Conference Paper

The Devil Is in the Frequency: Geminated Gestalt Autoencoder for Self-Supervised Visual Pre-training

  • Hao Liu
  • Xinghua Jiang
  • Xin Li
  • Antai Guo
  • Yiqing Hu
  • Deqiang Jiang
  • Bo Ren

The self-supervised Masked Image Modeling (MIM) schema, following "mask-and-reconstruct" pipeline of recovering contents from masked image, has recently captured the increasing interest in the community, owing to the excellent ability of learning visual representation from unlabeled data. Aiming at learning representations with high semantics abstracted, a group of works attempts to reconstruct non-semantic pixels with large-ratio masking strategy, which may suffer from "over-smoothing" problem, while others directly infuse semantics into targets in off-line way requiring extra data. Different from them, we shift the perspective to the Fourier domain which naturally has global perspective and present a new Masked Image Modeling (MIM), termed Geminated Gestalt Autoencoder (Ge^2-AE) for visual pre-training. Specifically, we equip our model with geminated decoders in charge of reconstructing image contents from both pixel and frequency space, where each other serves as not only the complementation but also the reciprocal constraints. Through this way, more robust representations can be learned in the pre-trained encoders, of which the effectiveness is confirmed by the juxtaposing experimental results on downstream recognition tasks. We also conduct several quantitative and qualitative experiments to investigate the learning behavior of our method. To our best knowledge, this is the first MIM work to solve the visual pre-training through the lens of frequency domain.

YNIMG Journal 2023 Journal Article

The dorsomedial prefrontal cortex represents subjective value across effort-based and risky decision-making

  • Yuan-Wei Yao
  • Kun-Ru Song
  • Nicolas W. Schuck
  • Xin Li
  • Xiao-Yi Fang
  • Jin-Tao Zhang
  • Hauke R. Heekeren
  • Rasmus Bruckner

Decisions that require taking effort costs into account are ubiquitous in real life. The neural common currency theory hypothesizes that a particular neural network integrates different costs (e.g., risk) and rewards into a common scale to facilitate value comparison. Although there has been a surge of interest in the computational and neural basis of effort-related value integration, it is still under debate if effort-based decision-making relies on a domain-general valuation network as implicated in the neural common currency theory. Therefore, we comprehensively compared effort-based and risky decision-making using a combination of computational modeling, univariate and multivariate fMRI analyses, and data from two independent studies. We found that effort-based decision-making can be best described by a power discounting model that accounts for both the discounting rate and effort sensitivity. At the neural level, multivariate decoding analyses indicated that the neural patterns of the dorsomedial prefrontal cortex (dmPFC) represented subjective value across different decision-making tasks including either effort or risk costs, although univariate signals were more diverse. These findings suggest that multivariate dmPFC patterns play a critical role in computing subjective value in a task-independent manner and thus extend the scope of the neural common currency theory.

AAAI Conference 2023 Conference Paper

Transformation-Equivariant 3D Object Detection for Autonomous Driving

  • Hai Wu
  • Chenglu Wen
  • Wei Li
  • Xin Li
  • Ruigang Yang
  • Cheng Wang

3D object detection received increasing attention in autonomous driving recently. Objects in 3D scenes are distributed with diverse orientations. Ordinary detectors do not explicitly model the variations of rotation and reflection transformations. Consequently, large networks and extensive data augmentation are required for robust detection. Recent equivariant networks explicitly model the transformation variations by applying shared networks on multiple transformed point clouds, showing great potential in object geometry modeling. However, it is difficult to apply such networks to 3D object detection in autonomous driving due to its large computation cost and slow reasoning speed. In this work, we present TED, an efficient Transformation-Equivariant 3D Detector to overcome the computation cost and speed issues. TED first applies a sparse convolution backbone to extract multi-channel transformation-equivariant voxel features; and then aligns and aggregates these equivariant features into lightweight and compact representations for high-performance 3D object detection. On the highly competitive KITTI 3D car detection leaderboard, TED ranked 1st among all submissions with competitive efficiency. Code is available at https://github.com/hailanyi/TED.

NeurIPS Conference 2023 Conference Paper

Understanding and Addressing the Pitfalls of Bisimulation-based Representations in Offline Reinforcement Learning

  • Hongyu Zang
  • Xin Li
  • Leiji Zhang
  • Yang Liu
  • Baigui Sun
  • Riashat Islam
  • Remi Tachet des Combes
  • Romain Laroche

While bisimulation-based approaches hold promise for learning robust state representations for Reinforcement Learning (RL) tasks, their efficacy in offline RL tasks has not been up to par. In some instances, their performance has even significantly underperformed alternative methods. We aim to understand why bisimulation methods succeed in online settings, but falter in offline tasks. Our analysis reveals that missing transitions in the dataset are particularly harmful to the bisimulation principle, leading to ineffective estimation. We also shed light on the critical role of reward scaling in bounding the scale of bisimulation measurements and of the value error they induce. Based on these findings, we propose to apply the expectile operator for representation learning to our offline RL setting, which helps to prevent overfitting to incomplete data. Meanwhile, by introducing an appropriate reward scaling strategy, we avoid the risk of feature collapse in representation space. We implement these recommendations on two state-of-the-art bisimulation-based algorithms, MICo and SimSR, and demonstrate performance gains on two benchmark suites: D4RL and Visual D4RL. Codes are provided at \url{https: //github. com/zanghyu/Offline_Bisimulation}.

NeurIPS Conference 2023 Conference Paper

UP-DP: Unsupervised Prompt Learning for Data Pre-Selection with Vision-Language Models

  • Xin Li
  • Sima Behpour
  • Thang Long Doan
  • Wenbin He
  • Liang Gou
  • Liu Ren

In this study, we investigate the task of data pre-selection, which aims to select instances for labeling from an unlabeled dataset through a single pass, thereby optimizing performance for undefined downstream tasks with a limited annotation budget. Previous approaches to data pre-selection relied solely on visual features extracted from foundation models, such as CLIP and BLIP-2, but largely ignored the powerfulness of text features. In this work, we argue that, with proper design, the joint feature space of both vision and text can yield a better representation for data pre-selection. To this end, we introduce UP-DP, a simple yet effective unsupervised prompt learning approach that adapts vision-language models, like BLIP-2, for data pre-selection. Specifically, with the BLIP-2 parameters frozen, we train text prompts to extract the joint features with improved representation, ensuring a diverse cluster structure that covers the entire dataset. We extensively compare our method with the state-of-the-art using seven benchmark datasets in different settings, achieving up to a performance gain of 20\%. Interestingly, the prompts learned from one dataset demonstrate significant generalizability and can be applied directly to enhance the feature extraction of BLIP-2 from other datasets. To the best of our knowledge, UP-DP is the first work to incorporate unsupervised prompt learning in a vision-language model for data pre-selection.

AAAI Conference 2023 Conference Paper

WaveForM: Graph Enhanced Wavelet Learning for Long Sequence Forecasting of Multivariate Time Series

  • Fuhao Yang
  • Xin Li
  • Min Wang
  • Hongyu Zang
  • Wei Pang
  • Mingzhong Wang

Multivariate time series (MTS) analysis and forecasting are crucial in many real-world applications, such as smart traffic management and weather forecasting. However, most existing work either focuses on short sequence forecasting or makes predictions predominantly with time domain features, which is not effective at removing noises with irregular frequencies in MTS. Therefore, we propose WaveForM, an end-to-end graph enhanced Wavelet learning framework for long sequence FORecasting of MTS. WaveForM first utilizes Discrete Wavelet Transform (DWT) to represent MTS in the wavelet domain, which captures both frequency and time domain features with a sound theoretical basis. To enable the effective learning in the wavelet domain, we further propose a graph constructor, which learns a global graph to represent the relationships between MTS variables, and graph-enhanced prediction modules, which utilize dilated convolution and graph convolution to capture the correlations between time series and predict the wavelet coefficients at different levels. Extensive experiments on five real-world forecasting datasets show that our model can achieve considerable performance improvement over different prediction lengths against the most competitive baseline of each dataset.

YNICL Journal 2023 Journal Article

White matter changes underlie hypertension-related cognitive decline in older adults

  • Zilin Li
  • Wenxiao Wang
  • Feng Sang
  • Zhanjun Zhang
  • Xin Li

Hypertension has been well recognized as a risk factor for cognitive impairment and dementia. Although the underlying mechanisms of hypertension-affected cognitive deterioration are not fully understood, white matter changes (WMCs) seem to play an important role. WMCs include low microstructural integrity and subsequent white matter macrostructural lesions, which are common on brain imaging in hypertensive patients and are critical for multiple cognitive domains. This article provides an overview of the impact of hypertension on white matter microstructural and macrostructural changes and its link to cognitive dysfunction. Hypertension may induce microstructural changes in white matter, especially for the long-range fibers such as anterior thalamic radiation (ATR) and inferior fronto-occipital fasciculus (IFOF), and then macrostructural abnormalities affecting different lobes, especially the periventricular area. Different regions' WMCs would further exert different effects to specific cognitive domains and accelerate brain aging. As a modifiable risk factor, hypertension might provide a new perspective for alleviating and delaying cognitive impairment.

NeurIPS Conference 2022 Conference Paper

AttCAT: Explaining Transformers via Attentive Class Activation Tokens

  • Yao Qiang
  • Deng Pan
  • Chengyin Li
  • Xin Li
  • Rhongho Jang
  • Dongxiao Zhu

Transformers have improved the state-of-the-art in various natural language processing and computer vision tasks. However, the success of the Transformer model has not yet been duly explained. Current explanation techniques, which dissect either the self-attention mechanism or gradient-based attribution, do not necessarily provide a faithful explanation of the inner workings of Transformers due to the following reasons: first, attention weights alone without considering the magnitudes of feature values are not adequate to reveal the self-attention mechanism; second, whereas most Transformer explanation techniques utilize self-attention module, the skip-connection module, contributing a significant portion of information flows in Transformers, has not yet been sufficiently exploited in explanation; third, the gradient-based attribution of individual feature does not incorporate interaction among features in explaining the model's output. In order to tackle the above problems, we propose a novel Transformer explanation technique via attentive class activation tokens, aka, AttCAT, leveraging encoded features, their gradients, and their attention weights to generate a faithful and confident explanation for Transformer's output. Extensive experiments are conducted to demonstrate the superior performance of AttCAT, which generalizes well to different Transformer architectures, evaluation metrics, datasets, and tasks, to the baseline methods. Our code is available at: https: //github. com/qiangyao1988/AttCAT.

NeurIPS Conference 2022 Conference Paper

Discrete Compositional Representations as an Abstraction for Goal Conditioned Reinforcement Learning

  • Riashat Islam
  • Hongyu Zang
  • Anirudh Goyal
  • Alex M. Lamb
  • Kenji Kawaguchi
  • Xin Li
  • Romain Laroche
  • Yoshua Bengio

Goal-conditioned reinforcement learning (RL) is a promising direction for training agents that are capable of solving multiple tasks and reach a diverse set of objectives. How to \textit{specify} and \textit{ground} these goals in such a way that we can both reliably reach goals during training as well as generalize to new goals during evaluation remains an open area of research. Defining goals in the space of noisy, high-dimensional sensory inputs is one possibility, yet this poses a challenge for training goal-conditioned agents, or even for generalization to novel goals. We propose to address this by learning compositional representations of goals and processing the resulting representation via a discretization bottleneck, for coarser specification of goals, through an approach we call DGRL. We show that discretizing outputs from goal encoders through a bottleneck can work well in goal-conditioned RL setups, by experimentally evaluating this method on tasks ranging from maze environments to complex robotic navigation and manipulation tasks. Additionally, we show a theoretical result which bounds the expected return for goals not observed during training, while still allowing for specifying goals with expressive combinatorial structure.

YNICL Journal 2022 Journal Article

Exploring brain glucose metabolic patterns in cognitively normal adults at risk of Alzheimer’s disease: A cross-validation study with Chinese and ADNI cohorts

  • Tao-Ran Li
  • Qiu-Yue Dong
  • Xue-Yan Jiang
  • Gui-Xia Kang
  • Xin Li
  • Yun-Yan Xie
  • Jie-Hui Jiang
  • Ying Han

OBJECTIVE: Disease-related metabolic brain patterns have been verified for a variety of neurodegenerative diseases including Alzheimer's disease (AD). This study aimed to explore and validate the pattern derived from cognitively normal controls (NCs) in the Alzheimer's continuum. METHODS: F]florbetapir-PET imaging. Participants were binary-grouped based on β-amyloid (Aβ) status, and the positivity was defined as Aβ+. Voxel-based scaled subprofile model/principal component analysis (SSM/PCA) was used to generate the "at-risk AD-related metabolic pattern (ARADRP)" for NCs. The pattern expression score was obtained and compared between the groups, and receiver operating characteristic curves were drawn. Notably, we conducted cross-validation to verify the robustness and correlation analyses to explore the relationships between the score and AD-related pathological biomarkers. RESULTS: F]florbetapir-PET (p > 0.23). CONCLUSIONS: ARADRP exists for NCs, and the acquired pattern expression score shows a certain ability to discriminate Aβ+ NCs from Aβ- NCs. The SSM/PCA method is expected to be helpful in the ultra-early diagnosis of AD in clinical practice.

IJCAI Conference 2022 Conference Paper

Learning Degradation Uncertainty for Unsupervised Real-world Image Super-resolution

  • Qian Ning
  • Jingzhu Tang
  • Fangfang Wu
  • Weisheng Dong
  • Xin Li
  • Guangming Shi

Acquiring degraded images with paired high-resolution (HR) images is often challenging, impeding the advance of image super-resolution in real-world applications. By generating realistic low-resolution (LR) images with degradation similar to that in real-world scenarios, simulated paired LR-HR data can be constructed for supervised training. However, most of the existing work ignores the degradation uncertainty of the generated realistic LR images, since only one LR image has been generated given an HR image. To address this weakness, we propose learning the degradation uncertainty of generated LR images and sampling multiple LR images from the learned LR image (mean) and degradation uncertainty (variance) and construct LR-HR pairs to train the super-resolution (SR) networks. Specifically, uncertainty can be learned by minimizing the proposed loss based on Kullback-Leibler (KL) divergence. Furthermore, the uncertainty in the feature domain is exploited by a novel perceptual loss; and we propose to calculate the adversarial loss from the gradient information in the SR stage for stable training performance and better visual quality. Experimental results on popular real-world datasets show that our proposed method has performed better than other unsupervised approaches.

AAAI Conference 2022 Conference Paper

Learning Optical Flow with Adaptive Graph Reasoning

  • Ao Luo
  • Fan Yang
  • Kunming Luo
  • Xin Li
  • Haoqiang Fan
  • Shuaicheng Liu

Estimating per-pixel motion between video frames, known as optical flow, is a long-standing problem in video understanding and analysis. Most contemporary optical flow techniques largely focus on addressing the cross-image matching with feature similarity, with few methods considering how to explicitly reason over the given scene for achieving a holistic motion understanding. In this work, taking a fresh perspective, we introduce a novel graph-based approach, called adaptive graph reasoning for optical flow (AGFlow), to emphasize the value of scene/context information in optical flow. Our key idea is to decouple the context reasoning from the matching procedure, and exploit scene information to effectively assist motion estimation by learning to reason over the adaptive graph. The proposed AGFlow can effectively exploit the context information and incorporate it within the matching procedure, producing more robust and accurate results. On both Sintel clean and final passes, our AGFlow achieves the best accuracy with EPE of 1. 43 and 2. 47 pixels, outperforming state-of-the-art approaches by 11. 2% and 13. 6%, respectively. Code is publicly available at https: //github. com/ megvii-research/AGFlow.

AAAI Conference 2022 Conference Paper

Robust Depth Completion with Uncertainty-Driven Loss Functions

  • Yufan Zhu
  • Weisheng Dong
  • Leida Li
  • Jinjian Wu
  • Xin Li
  • Guangming Shi

Recovering a dense depth image from sparse LiDAR scans is a challenging task. Despite the popularity of color-guided methods for sparse-to-dense depth completion, they treated pixels equally during optimization, ignoring the uneven distribution characteristics in the sparse depth map and the accumulated outliers in the synthesized ground truth. In this work, we introduce uncertainty-driven loss functions to improve the robustness of depth completion and handle the uncertainty in depth completion. Specifically, we propose an explicit uncertainty formulation for robust depth completion with Jeffrey’s prior. A parametric uncertain-driven loss is introduced and translated to new loss functions that are robust to noisy or missing data. Meanwhile, we propose a multiscale joint prediction model that can simultaneously predict depth and uncertainty maps. The estimated uncertainty map is also used to perform adaptive prediction on the pixels with high uncertainty, leading to a residual map for refining the completion results. Our method has been tested on KITTI Depth Completion Benchmark and achieved the state-of-the-art robustness performance in terms of MAE, IMAE, and IRMSE metrics.

TMLR Journal 2022 Journal Article

Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

  • Jiahui Yu
  • Yuanzhong Xu
  • Jing Yu Koh
  • Thang Luong
  • Gunjan Baid
  • Zirui Wang
  • Vijay Vasudevan
  • Alexander Ku

We present the Pathways Autoregressive Text-to-Image (Parti) model, which generates high-fidelity photorealistic images and supports content-rich synthesis involving complex compositions and world knowledge. Parti treats text-to-image generation as a sequence-to-sequence modeling problem, akin to machine translation, with sequences of image tokens as the target outputs rather than text tokens in another language. This strategy can naturally tap into the rich body of prior work on large language models, which have seen continued advances in capabilities and performance through scaling data and model sizes. Our approach is simple: First, Parti uses a Transformer-based image tokenizer, ViT-VQGAN, to encode images as sequences of discrete tokens. Second, we achieve consistent quality improvements by scaling the encoder-decoder Transformer model up to 20B parameters, with a new state-of-the-art zero-shot FID score of 7.23 and finetuned FID score of 3.22 on MS-COCO. Our detailed analysis on Localized Narratives as well as PartiPrompts (P2), a new holistic benchmark of over 1600 English prompts, demonstrate the effectiveness of Parti across a wide variety of categories and difficulty aspects. We also explore and highlight limitations of our models in order to define and exemplify key areas of focus for further improvements.

AAAI Conference 2022 Conference Paper

SimSR: Simple Distance-Based State Representations for Deep Reinforcement Learning

  • Hongyu Zang
  • Xin Li
  • Mingzhong Wang

This work explores how to learn robust and generalizable state representation from image-based observations with deep reinforcement learning methods. Addressing the computational complexity, stringent assumptions and representation collapse challenges in existing work of bisimulation metric, we devise Simple State Representation (SimSR) operator. SimSR enables us to design a stochastic approximation method that can practically learn the mapping functions (encoders) from observations to latent representation space. In addition to the theoretical analysis and comparison with the existing work, we experimented and compared our work with recent state-of-the-art solutions in visual MuJoCo tasks. The results shows that our model generally achieves better performance and has better robustness and good generalization.

JBHI Journal 2022 Journal Article

Understanding Dynamics of Pandemic Models to Support Predictions of COVID-19 Transmission: Parameter Sensitivity Analysis of SIR-Type Models

  • Chunfeng Ma
  • Xin Li
  • Zebin Zhao
  • Feng Liu
  • Kun Zhang
  • Adan Wu
  • Xiaowei Nie

Despite efforts made to model and predict COVID-19 transmission, large predictive uncertainty remains. Failure to understand the dynamics of the nonlinear pandemic prediction model is an important reason. To this end, local and multiple global sensitivity analysis approaches are synthetically applied to analyze the sensitivities of parameters and initial state variables and community size (N) in susceptible-infected-recovered (SIR) and its variant susceptible-exposed-infected-recovered (SEIR) models and basic reproduction number ( R0 ), aiming to provide prior information for parameter estimation and suggestions for COVID-19 prevention and control measures. We found that N influences both the maximum number of actively infected cases and the date on which the maximum number of actively infected cases is reached. The high effect of N on maximum actively infected cases and peak date suggests the necessity of isolating the infected cases in a small community. The protection rate and average quarantined time are most sensitive to the infected populations, with a summation of their first-order sensitivity indices greater than 0. 585, and their interactions are also substantial, being 0. 389 and 0. 334, respectively. The high sensitivities and interaction between the protection rate and average quarantined time suggest that protection and isolation measures should always be implemented in conjunction and started as early as possible. These findings provide insights into the predictability of the pandemic models by estimating influential parameters and suggest how to effectively prevent and control epidemic transmission.

ICLR Conference 2022 Conference Paper

Vector-quantized Image Modeling with Improved VQGAN

  • Jiahui Yu
  • Xin Li
  • Jing Yu Koh
  • Han Zhang 0010
  • Ruoming Pang
  • James Qin
  • Alexander Ku
  • Yuanzhong Xu

Pretraining language models with next-token prediction on massive text corpora has delivered phenomenal zero-shot, few-shot, transfer learning and multi-tasking capabilities on both generative and discriminative language tasks. Motivated by this success, we explore a Vector-quantized Image Modeling (VIM) approach that involves pretraining a Transformer to predict rasterized image tokens autoregressively. The discrete image tokens are encoded from a learned Vision-Transformer-based VQGAN (ViT-VQGAN). We first propose multiple improvements over vanilla VQGAN from architecture to codebook learning, yielding better efficiency and reconstruction fidelity. The improved ViT-VQGAN further improves vector-quantized image modeling tasks, including unconditional, class-conditioned image generation and unsupervised representation learning. When trained on ImageNet at 256x256 resolution, we achieve Inception Score (IS) of 175.1 and Fr'echet Inception Distance (FID) of 4.17, a dramatic improvement over the vanilla VQGAN, which obtains 70.6 and 17.04 for IS and FID, respectively. Based on ViT-VQGAN and unsupervised pretraining, we further evaluate the pretrained Transformer by averaging intermediate features, similar to Image GPT (iGPT). This ImageNet-pretrained VIM-L significantly beats iGPT-L on linear-probe accuracy from 60.3% to 73.2% for a similar model size. ViM-L also outperforms iGPT-XL which is trained with extra web image data and larger model size.

NeurIPS Conference 2021 Conference Paper

DeepReduce: A Sparse-tensor Communication Framework for Federated Deep Learning

  • Hang Xu
  • Kelly Kostopoulou
  • Aritra Dutta
  • Xin Li
  • Alexandros Ntoulas
  • Panos Kalnis

Sparse tensors appear frequently in federated deep learning, either as a direct artifact of the deep neural network’s gradients, or as a result of an explicit sparsification process. Existing communication primitives are agnostic to the peculiarities of deep learning; consequently, they impose unnecessary communication overhead. This paper introduces DeepReduce, a versatile framework for the compressed communication of sparse tensors, tailored to federated deep learning. DeepReduce decomposes sparse tensors into two sets, values and indices, and allows both independent and combined compression of these sets. We support a variety of common compressors, such as Deflate for values, or run-length encoding for indices. We also propose two novel compression schemes that achieve superior results: curve fitting-based for values, and bloom filter-based for indices. DeepReduce is orthogonal to existing gradient sparsifiers and can be applied in conjunction with them, transparently to the end-user, to significantly lower the communication overhead. As proof of concept, we implement our approach on TensorFlow and PyTorch. Our experiments with large real models demonstrate that DeepReduce transmits 320% less data than existing sparsifiers, without affecting accuracy. Code is available at https: //github. com/hangxu0304/DeepReduce.

IJCAI Conference 2021 Conference Paper

Explaining Deep Neural Network Models with Adversarial Gradient Integration

  • Deng Pan
  • Xin Li
  • Dongxiao Zhu

Deep neural networks (DNNs) have became one of the most high performing tools in a broad range of machine learning areas. However, the multilayer non-linearity of the network architectures prevent us from gaining a better understanding of the models’ predictions. Gradient based attribution methods (e. g. , Integrated Gradient (IG)) that decipher input features’ contribution to the prediction task have been shown to be highly effective yet requiring a reference input as the anchor for explaining model’s output. The performance of DNN model interpretation can be quite inconsistent with regard to the choice of references. Here we propose an Adversarial Gradient Integration (AGI) method that integrates the gradients from adversarial examples to the target example along the curve of steepest ascent to calculate the resulting contributions from all input features. Our method doesn’t rely on the choice of references, hence can avoid the ambiguity and inconsistency sourced from the reference selection. We demonstrate the performance of our AGI method and compare with competing methods in explaining image classification results. Code is available from https: //github. com/pd90506/AGI.

TIST Journal 2021 Journal Article

Identifying Illicit Drug Dealers on Instagram with Large-scale Multimodal Data Fusion

  • Chuanbo Hu
  • Minglei Yin
  • Bin Liu
  • Xin Li
  • Yanfang Ye

Illicit drug trafficking via social media sites such as Instagram have become a severe problem, thus drawing a great deal of attention from law enforcement and public health agencies. How to identify illicit drug dealers from social media data has remained a technical challenge for the following reasons. On the one hand, the available data are limited because of privacy concerns with crawling social media sites; on the other hand, the diversity of drug dealing patterns makes it difficult to reliably distinguish drug dealers from common drug users. Unlike existing methods that focus on posting-based detection, we propose to tackle the problem of illicit drug dealer identification by constructing a large-scale multimodal dataset named Identifying Drug Dealers on Instagram (IDDIG). Nearly 4,000 user accounts, of which more than 1,400 are drug dealers, have been collected from Instagram with multiple data sources including post comments, post images, homepage bio, and homepage images. We then design a quadruple-based multimodal fusion method to combine the multiple data sources associated with each user account for drug dealer identification. Experimental results on the constructed IDDIG dataset demonstrate the effectiveness of the proposed method in identifying drug dealers (almost 95% accuracy). Moreover, we have developed a hashtag-based community detection technique for discovering evolving patterns, especially those related to geography and drug types.

AAAI Conference 2021 Conference Paper

Improving Adversarial Robustness via Probabilistically Compact Loss with Logit Constraints

  • Xin Li
  • Xiangrui Li
  • Deng Pan
  • Dongxiao Zhu

Convolutional neural networks (CNNs) have achieved stateof-the-art performance on various tasks in computer vision. However, recent studies demonstrate that these models are vulnerable to carefully crafted adversarial samples and suffer from a significant performance drop when predicting them. Many methods have been proposed to improve adversarial robustness (e. g. , adversarial training and new loss functions to learn adversarially robust feature representations). Here we offer a unique insight into the predictive behavior of CNNs that they tend to misclassify adversarial samples into the most probable false classes. This inspires us to propose a new Probabilistically Compact (PC) loss with logit constraints which can be used as a drop-in replacement for crossentropy (CE) loss to improve CNN’s adversarial robustness. Specifically, PC loss enlarges the probability gaps between true class and false classes meanwhile the logit constraints prevent the gaps from being melted by a small perturbation. We extensively compare our method with the state-of-the-art using large scale datasets under both white-box and blackbox attacks to demonstrate its effectiveness. The source codes are available at https: //github. com/xinli0928/PC-LC.

AAAI Conference 2021 Conference Paper

Learning Omni-Frequency Region-adaptive Representations for Real Image Super-Resolution

  • Xin Li
  • Xin Jin
  • Tao Yu
  • Simeng Sun
  • Yingxue Pang
  • Zhizheng Zhang
  • Zhibo Chen

Traditional single image super-resolution (SISR) methods that focus on solving single and uniform degradation (i. e. , bicubic down-sampling), typically suffer from poor performance when applied into real-world low-resolution (LR) images due to the complicated realistic degradations. The key to solving this more challenging real image super-resolution (RealSR) problem lies in learning feature representations that are both informative and content-aware. In this paper, we propose an Omni-frequency Region-adaptive Network (OR- Net) to address both challenges, here we call features of all low, middle and high frequencies omni-frequency features. Specifically, we start from the frequency perspective and design a Frequency Decomposition (FD) module to separate different frequency components to comprehensively compensate the information lost for real LR image. Then, considering the different regions of real LR image have different frequency information lost, we further design a Region-adaptive Frequency Aggregation (RFA) module by leveraging dynamic convolution and spatial attention to adaptively restore frequency components for different regions. The extensive experiments endorse the effective, and scenario-agnostic nature of our OR-Net for RealSR.

YNIMG Journal 2021 Journal Article

Reconstructing seen image from brain activity by visually-guided cognitive representation and adversarial learning

  • Ziqi Ren
  • Jie Li
  • Xuetong Xue
  • Xin Li
  • Fan Yang
  • Zhicheng Jiao
  • Xinbo Gao

Reconstructing perceived stimulus (image) only from human brain activity measured with functional Magnetic Resonance Imaging (fMRI) is a significant task in brain decoding. However, the inconsistent distribution and representation between fMRI signals and visual images cause great 'domain gap'. Moreover, the limited fMRI data instances generally suffer from the issues of low signal noise ratio (SNR), extremely high dimensionality, and limited spatial resolution. Existing methods are often affected by these issues so that a satisfactory reconstruction is still an open problem. In this paper, we show that it is possible to obtain a promising solution by learning visually-guided latent cognitive representations from the fMRI signals, and inversely decoding them to the image stimuli. The resulting framework is called Dual-Variational Autoencoder/ Generative Adversarial Network (D-Vae/Gan), which combines the advantages of adversarial representation learning with knowledge distillation. In addition, we introduce a novel three-stage learning strategy which enables the (cognitive) encoder to gradually distill useful knowledge from the paired (visual) encoder during the learning process. Extensive experimental results on both artificial and natural images have demonstrated that our method could achieve surprisingly good results and outperform the available alternatives.

TIST Journal 2021 Journal Article

Simultaneous Past and Current Social Interaction-aware Trajectory Prediction for Multiple Intelligent Agents in Dynamic Scenes

  • Yanliang Zhu
  • Dongchun Ren
  • Yi Xu
  • Deheng Qian
  • Mingyu Fan
  • Xin Li
  • Huaxia Xia

Trajectory prediction of multiple agents in a crowded scene is an essential component in many applications, including intelligent monitoring, autonomous robotics, and self-driving cars. Accurate agent trajectory prediction remains a significant challenge because of the complex dynamic interactions among the agents and between them and the surrounding scene. To address the challenge, we propose a decoupled attention-based spatial-temporal modeling strategy in the proposed trajectory prediction method. The past and current interactions among agents are dynamically and adaptively summarized by two separate attention-based networks and have proven powerful in improving the prediction accuracy. Moreover, it is optional in the proposed method to make use of the road map and the plan of the ego-agent for scene-compliant and accurate predictions. The road map feature is efficiently extracted by a convolutional neural network, and the features of the ego-agent’s plan is extracted by a gated recurrent network with an attention module based on the temporal characteristic. Experiments on benchmark trajectory prediction datasets demonstrate that the proposed method is effective when the ego-agent plan and the the surrounding scene information are provided and achieves state-of-the-art performance with only the observed trajectories.

ICRA Conference 2021 Conference Paper

Star Topology based Interaction for Robust Trajectory Forecasting in Dynamic Scene

  • Yanliang Zhu
  • Dongchun Ren
  • Deheng Qian
  • Mingyu Fan
  • Xin Li
  • Huaxia Xia

Motion prediction of multiple agents in a dynamic scene is a crucial component in many real applications, including intelligent monitoring and autonomous driving. Due to the complex interactions among the agents and their interactions with the surrounding scene, accurate trajectory prediction is still a great challenge. In this paper, we propose a new method for robust trajectory prediction of multiple intelligent agents in a dynamic scene. The input of the method includes the observed trajectories of all agents, and optionally, the planning of the ego-agent and the surrounding high definition map at every time steps. Given observed trajectories, an efficient approach in a star computational topology is utilized to compute both the spatiotemporal interaction features and the current interaction features between the agents, where the time complexity scales linearly to the number of agents. Moreover, on an autonomous vehicle, the proposed prediction method can make use of the planning of ego-agent to improve the modeling of the interaction between surrounding agents. To increase the robustness to upstream perception noises, at the training stage, we randomly mask out the input data, a. k. a. the points on the observed trajectories of agents and the lane sequence. Experiments on autonomous driving and pedestrian-walking datasets demonstrate that the proposed method is not only effective when the planning of ego-agent and the high definition map are provided, but also achieves state-of-the-art performance with only the observed trajectories.

IJCAI Conference 2021 Conference Paper

Tracklet Proposal Network for Multi-Object Tracking on Point Clouds

  • Hai Wu
  • Qing Li
  • Chenglu Wen
  • Xin Li
  • Xiaoliang Fan
  • Cheng Wang

This paper proposes the first tracklet proposal network, named PC-TCNN, for Multi-Object Tracking (MOT) on point clouds. Our pipeline first generates tracklet proposals, then refines these tracklets and associates them to generate long trajectories. Specifically, object proposal generation and motion regression are first performed on a point cloud sequence to generate tracklet candidates. Then, spatial-temporal features of each tracklet are exploited and their consistency is used to refine the tracklet proposal. Finally, the refined tracklets across multiple frames are associated to perform MOT on the point cloud sequence. The PC-TCNN significantly improves the MOT performance by introducing the tracklet proposal design. On the KITTI tracking benchmark, it attains an MOTA of 91. 75%, outperforming all submitted results on the online leaderboard.

AAAI Conference 2021 Conference Paper

Traffic Shaping in E-Commercial Search Engine: Multi-Objective Online Welfare Maximization

  • Liucheng Sun
  • Chenwei Weng
  • Chengfu Huo
  • Weijun Ren
  • Guochuan Zhang
  • Xin Li

The e-commercial search engine is the primary gateway for customers to find desired products and engage in online shopping. Besides displaying items to optimize for a single objective (i. e. , relevance), ranking items needs to satisfy some other business requirements in practice. Recently, traffic shaping was introduced to incorporate multiple objectives in a constrained optimization framework. However, many practical business requirements can not explicitly represented by linear constraints as in the existing work, and this may limit the scalability of their framework. This paper presents a unified framework from the aspect of multi-objective welfare maximization where we regard all business requirements as objectives to optimize. Our framework can naturally incorporate a wide range of application-driven requirements. In addition to formulating the problem, we design an online traffic splitting algorithm that allows us to flexibly adjust the priorities of different objectives, and it has rigorous theoretical guarantees over the adversarial scenario. We also run experiments on both synthetic and real-world datasets to validate our algorithms.

NeurIPS Conference 2021 Conference Paper

Uncertainty-Driven Loss for Single Image Super-Resolution

  • Qian Ning
  • Weisheng Dong
  • Xin Li
  • Jinjian Wu
  • Guangming Shi

In low-level vision such as single image super-resolution (SISR), traditional MSE or L 1 loss function treats every pixel equally with the assumption that the importance of all pixels is the same. However, it has been long recognized that texture and edge areas carry more important visual information than smooth areas in photographic images. How to achieve such spatial adaptation in a principled manner has been an open problem in both traditional model-based and modern learning-based approaches toward SISR. In this paper, we propose a new adaptive weighted loss for SISR to train deep networks focusing on challenging situations such as textured and edge pixels with high uncertainty. Specifically, we introduce variance estimation characterizing the uncertainty on a pixel-by-pixel basis into SISR solutions so the targeted pixels in a high-resolution image (mean) and their corresponding uncertainty (variance) can be learned simultaneously. Moreover, uncertainty estimation allows us to leverage conventional wisdom such as sparsity prior for regularizing SISR solutions. Ultimately, pixels with large certainty (e. g. , texture and edge pixels) will be prioritized for SISR according to their importance to visual quality. For the first time, we demonstrate that such uncertainty-driven loss can achieve better results than MSE or L 1 loss for a wide range of network architectures. Experimental results on three popular SISR networks show that our proposed uncertainty-driven loss has achieved better PSNR performance than traditional loss functions without any increased computation during testing. The code is available at https: //see. xidian. edu. cn/faculty/wsdong/Projects/UDL-SR. htm

IJCAI Conference 2020 Conference Paper

Beyond Network Pruning: a Joint Search-and-Training Approach

  • Xiaotong Lu
  • Han Huang
  • Weisheng Dong
  • Xin Li
  • Guangming Shi

Network pruning has been proposed as a remedy for alleviating the over-parameterization problem of deep neural networks. However, its value has been recently challenged especially from the perspective of neural architecture search (NAS). We challenge the conventional wisdom of pruning-after-training by proposing a joint search-and-training approach that directly learns a compact network from the scratch. By treating pruning as a search strategy, we present two new insights in this paper: 1) it is possible to expand the search space of networking pruning by associating each filter with a learnable weight; 2) joint search-and-training can be conducted iteratively to maximize the learning efficiency. More specifically, we propose a coarse-to-fine tuning strategy to iteratively sample and update compact sub-network to approximate the target network. The weights associated with network filters will be accordingly updated by joint search-and-training to reflect learned knowledge in NAS space. Moreover, we introduce strategies of random perturbation (inspired by Monte Carlo) and flexible thresholding (inspired by Reinforcement Learning) to adjust the weight and size of each layer. Extensive experiments on ResNet and VGGNet demonstrate the superior performance of our proposed method on popular datasets including CIFAR10, CIFAR100 and ImageNet.

JBHI Journal 2020 Journal Article

Bi-Modality Medical Image Synthesis Using Semi-Supervised Sequential Generative Adversarial Networks

  • Xin Yang
  • Yi Lin
  • Zhiwei Wang
  • Xin Li
  • Kwang-Ting Cheng

In this paper, we propose a bi-modality medical image synthesis approach based on sequential generative adversarial network (GAN) and semi-supervised learning. Our approach consists of two generative modules that synthesize images of the two modalities in a sequential order. A method for measuring the synthesis complexity is proposed to automatically determine the synthesis order in our sequential GAN. Images of the modality with a lower complexity are synthesized first, and the counterparts with a higher complexity are generated later. Our sequential GAN is trained end-to-end in a semi-supervised manner. In supervised training, the joint distribution of bi-modality images are learned from real paired images of the two modalities by explicitly minimizing the reconstruction losses between the real and synthetic images. To avoid overfitting limited training images, in unsupervised training, the marginal distribution of each modality is learned based on unpaired images by minimizing the Wasserstein distance between the distributions of real and fake images. We comprehensively evaluate the proposed model using two synthesis tasks based on three types of evaluate metrics and user studies. Visual and quantitative results demonstrate the superiority of our method to the state-of-the-art methods, and reasonable visual quality and clinical significance. Code is made publicly available at https://github.com/hust-linyi/Multimodal-Medical-Image-Synthesis.

AAAI Conference 2020 Conference Paper

Contextual-Bandit Based Personalized Recommendation with Time-Varying User Interests

  • Xiao Xu
  • Fang Dong
  • Yanghua Li
  • Shaojian He
  • Xin Li

A contextual bandit problem is studied in a highly nonstationary environment, which is ubiquitous in various recommender systems due to the time-varying interests of users. Two models with disjoint and hybrid payoffs are considered to characterize the phenomenon that users’ preferences towards different items vary differently over time. In the disjoint payoff model, the reward of playing an arm is determined by an arm-specific preference vector, which is piecewise-stationary with asynchronous and distinct changes across different arms. An efficient learning algorithm that is adaptive to abrupt reward changes is proposed and theoretical regret analysis is provided to show that a sublinear scaling of regret in the time length T is achieved. The algorithm is further extended to a more general setting with hybrid payoffs where the reward of playing an arm is determined by both an arm-specific preference vector and a joint coefficient vector shared by all arms. Empirical experiments are conducted on real-world datasets to verify the advantages of the proposed learning algorithms against baseline ones in both settings.

IJCAI Conference 2020 Conference Paper

Explainable Recommendation via Interpretable Feature Mapping and Evaluation of Explainability

  • Deng Pan
  • Xiangrui Li
  • Xin Li
  • Dongxiao Zhu

Latent factor collaborative filtering (CF) has been a widely used technique for recommender system by learning the semantic representations of users and items. Recently, explainable recommendation has attracted much attention from research community. However, trade-off exists between explainability and performance of the recommendation where metadata is often needed to alleviate the dilemma. We present a novel feature mapping approach that maps the uninterpretable general features onto the interpretable aspect features, achieving both satisfactory accuracy and explainability in the recommendations by simultaneous minimization of rating prediction loss and interpretation loss. To evaluate the explainability, we propose two new evaluation metrics specifically designed for aspect-level explanation using surrogate ground truth. Experimental results demonstrate a strong performance in both recommendation and explaining explanation, eliminating the need for metadata. Code is available from https: //github. com/pd90506/AMCF.

AAAI Conference 2020 Conference Paper

Hybrid Graph Neural Networks for Crowd Counting

  • Ao Luo
  • Fan Yang
  • Xin Li
  • Dong Nie
  • Zhicheng Jiao
  • Shangchen Zhou
  • Hong Cheng

Crowd counting is an important yet challenging task due to the large scale and density variation. Recent investigations have shown that distilling rich relations among multi-scale features and exploiting useful information from the auxiliary task, i. e. , localization, are vital for this task. Nevertheless, how to comprehensively leverage these relations within a uni- fied network architecture is still a challenging problem. In this paper, we present a novel network structure called Hybrid Graph Neural Network (HyGnn) which targets to relieve the problem by interweaving the multi-scale features for crowd density as well as its auxiliary task (localization) together and performing joint reasoning over a graph. Specifically, HyGnn integrates a hybrid graph to jointly represent the task-specific feature maps of different scales as nodes, and two types of relations as edges: (i) multi-scale relations capturing the feature dependencies across scales and (ii) mutual beneficial relations building bridges for the cooperation between counting and localization. Thus, through message passing, HyGnn can capture and distill richer relations between nodes to obtain more powerful representations, providing robust and accurate results. Our HyGnn performs significantly well on four challenging datasets: ShanghaiTech Part A, ShanghaiTech Part B, UCF CC 50 and UCF QNRF, outperforming the state-ofthe-art algorithms by a large margin.

ICRA Conference 2020 Conference Paper

Modeling and Experiments on the Swallowing and Disgorging Characteristics of an Underwater Continuum Manipulator

  • Haihang Wang
  • He Xu
  • Fengshu Yu
  • Xin Li
  • Chen Yang
  • Siqing Chen
  • Junlong Chen
  • Yonghui Zhang

Soft robots apply compliant materials to perform motions and behaviors not typically achievable by rigid robots. An underwater, compliant, multi-segment continuum manipulator that can bend, swallow, disgorge is developed in this study. The manipulator is driven by McKibben water hydraulic artificial muscle (WHAM). The mechanical properties of the WHAM are tested and analyzed experimentally. The kinematics model, which concerns about the variable diameter structure of the soft grippers, are established to simulate the behaviors of the manipulator among the bending, swallowing and disgorging procedure. A mouth-tongue collaborative soft robot assembled with another single-segment soft robot arm is presented. And its functions are experimentally testified. The distinctive functions were verified according to the experimental results.

AAAI Conference 2020 Conference Paper

Multi-Task Driven Feature Models for Thermal Infrared Tracking

  • Qiao Liu
  • Xin Li
  • Zhenyu He
  • Nana Fan
  • Di Yuan
  • Wei Liu
  • Yongsheng Liang

Existing deep Thermal InfraRed (TIR) trackers usually use the feature models of RGB trackers for representation. However, these feature models learned on RGB images are neither effective in representing TIR objects nor taking fine-grained TIR information into consideration. To this end, we develop a multi-task framework to learn the TIR-specific discriminative features and fine-grained correlation features for TIR tracking. Specifically, we first use an auxiliary classification network to guide the generation of TIR-specific discriminative features for distinguishing the TIR objects belonging to different classes. Second, we design a fine-grained aware module to capture more subtle information for distinguishing the TIR objects belonging to the same class. These two kinds of features complement each other and recognize TIR objects in the levels of inter-class and intra-class respectively. These two feature models are learned using a multi-task matching framework and are jointly optimized on the TIR tracking task. In addition, we develop a large-scale TIR training dataset to train the network for adapting the model to the TIR domain. Extensive experimental results on three benchmarks show that the proposed algorithm achieves a relative gain of 10% over the baseline and performs favorably against the stateof-the-art methods. Codes and the proposed TIR dataset are available at https: //github. com/QiaoLiuHit/MMNet.

AAAI Conference 2020 Conference Paper

On the Learning Property of Logistic and Softmax Losses for Deep Neural Networks

  • Xiangrui Li
  • Xin Li
  • Deng Pan
  • Dongxiao Zhu

Deep convolutional neural networks (CNNs) trained with logistic and softmax losses have made significant advancement in visual recognition tasks in computer vision. When training data exhibit class imbalances, the class-wise reweighted version of logistic and softmax losses are often used to boost performance of the unweighted version. In this paper, motivated to explain the reweighting mechanism, we explicate the learning property of those two loss functions by analyzing the necessary condition (e. g. , gradient equals to zero) after training CNNs to converge to a local minimum. The analysis immediately provides us explanations for understanding (1) quantitative effects of the class-wise reweighting mechanism: deterministic effectiveness for binary classification using logistic loss yet indeterministic for multi-class classification using softmax loss; (2) disadvantage of logistic loss for single-label multi-class classification via one-vs. -all approach, which is due to the averaging effect on predicted probabilities for the negative class (e. g. , non-target classes) in the learning process. With the disadvantage and advantage of logistic loss disentangled, we thereafter propose a novel reweighted logistic loss for multi-class classification. Our simple yet effective formulation improves ordinary logistic loss by focusing on learning hard non-target classes (target vs. non-target class in one-vs. -all) and turned out to be competitive with softmax loss. We evaluate our method on several benchmark datasets to demonstrate its effectiveness.

AAAI Conference 2020 Conference Paper

Point2Node: Correlation Learning of Dynamic-Node for Point Cloud Feature Modeling

  • Wenkai Han
  • Chenglu Wen
  • Cheng Wang
  • Xin Li
  • Qing Li

Fully exploring correlation among points in point clouds is essential for their feature modeling. This paper presents a novel end-to-end graph model, named Point2Node, to represent a given point cloud. Point2Node can dynamically explore correlation among all graph nodes from different levels, and adaptively aggregate the learned features. Specifically, first, to fully explore the spatial correlation among points for enhanced feature description, in a high-dimensional node graph, we dynamically integrate the node’s correlation with self, local, and non-local nodes. Second, to more effectively integrate learned features, we design a data-aware gate mechanism to self-adaptively aggregate features at the channel level. Extensive experiments on various point cloud benchmarks demonstrate that our method outperforms the state-ofthe-art.

AAAI Conference 2020 Conference Paper

Relevance-Promoting Language Model for Short-Text Conversation

  • Xin Li
  • Piji Li
  • Wei Bi
  • Xiaojiang Liu
  • Wai Lam

Despite the effectiveness of sequence-to-sequence framework on the task of Short-Text Conversation (STC), the issue of under-exploitation of training data (i. e. , the supervision signals from query text is ignored) still remains unresolved. Also, the adopted maximization-based decoding strategies, inclined to generating the generic responses or responses with repetition, are unsuited to the STC task. In this paper, we propose to formulate the STC task as a language modeling problem and tailor-make a training strategy to adapt a language model for response generation. To enhance generation performance, we design a relevance-promoting transformer language model, which performs additional supervised source attention after the self-attention to increase the importance of informative query tokens in calculating the token-level representation. The model further refines the query representation with relevance clues inferred from its multiple references during training. In testing, we adopt a randomization-overmaximization strategy to reduce the generation of generic responses. Experimental results on a large Chinese STC dataset demonstrate the superiority of the proposed model on relevance metrics and diversity metrics. 1

ICLR Conference 2020 Conference Paper

Training Recurrent Neural Networks Online by Learning Explicit State Variables

  • Somjit Nath
  • Vincent Liu
  • Alan Chan 0001
  • Xin Li
  • Adam White 0001
  • Martha White

Recurrent neural networks (RNNs) allow an agent to construct a state-representation from a stream of experience, which is essential in partially observable problems. However, there are two primary issues one must overcome when training an RNN: the sensitivity of the learning algorithm's performance to truncation length and and long training times. There are variety of strategies to improve training in RNNs, the mostly notably Backprop Through Time (BPTT) and by Real-Time Recurrent Learning. These strategies, however, are typically computationally expensive and focus computation on computing gradients back in time. In this work, we reformulate the RNN training objective to explicitly learn state vectors; this breaks the dependence across time and so avoids the need to estimate gradients far back in time. We show that for a fixed buffer of data, our algorithm---called Fixed Point Propagation (FPP)---is sound: it converges to a stationary point of the new objective. We investigate the empirical performance of our online FPP algorithm, particularly in terms of computation compared to truncated BPTT with varying truncation levels.

AAAI Conference 2020 Conference Paper

Universal Value Iteration Networks: When Spatially-Invariant Is Not Universal

  • Li Zhang
  • Xin Li
  • Sen Chen
  • Hongyu Zang
  • Jie Huang
  • Mingzhong Wang

In this paper, we first formally define the problem set of spatially invariant Markov Decision Processes (MDPs), and show that Value Iteration Networks (VIN) and its extensions are computationally bounded to it due to the use of the convolution kernel. To generalize VIN to spatially variant MDPs, we propose Universal Value Iteration Networks (UVIN). In comparison with VIN, UVIN automatically learns a flexible but compact network structure to encode the transition dynamics of the problems and support the differentiable planning module. We evaluate UVIN with both spatially invariant and spatially variant tasks, including navigation in regular maze, chessboard maze, and Mars, and Minecraft item syntheses. Results show that UVIN can achieve similar performance as VIN and its extensions on spatially invariant tasks, and significantly outperforms other models on more general problems.

NeurIPS Conference 2020 Conference Paper

Video Object Segmentation with Adaptive Feature Bank and Uncertain-Region Refinement

  • Yongqing Liang
  • Xin Li
  • Navid Jafari
  • Jim Chen

This paper presents a new matching-based framework for semi-supervised video object segmentation (VOS). Recently, state-of-the-art VOS performance has been achieved by matching-based algorithms, in which feature banks are created to store features for region matching and classification. However, how to effectively organize information in the continuously growing feature bank remains under-explored, and this leads to an inefficient design of the bank. We introduced an adaptive feature bank update scheme to dynamically absorb new features and discard obsolete features. We also designed a new confidence loss and a fine-grained segmentation module to enhance the segmentation accuracy in uncertain regions. On public benchmarks, our algorithm outperforms existing state-of-the-arts.

AAAI Conference 2019 Conference Paper

A Unified Model for Opinion Target Extraction and Target Sentiment Prediction

  • Xin Li
  • Lidong Bing
  • Piji Li
  • Wai Lam

Target-based sentiment analysis involves opinion target extraction and target sentiment classification. However, most of the existing works usually studied one of these two sub-tasks alone, which hinders their practical use. This paper aims to solve the complete task of target-based sentiment analysis in an end-to-end fashion, and presents a novel unified model which applies a unified tagging scheme. Our framework involves two stacked recurrent neural networks: The upper one predicts the unified tags to produce the final output results of the primary target-based sentiment analysis; The lower one performs an auxiliary target boundary prediction aiming at guiding the upper network to improve the performance of the primary task. To explore the inter-task dependency, we propose to explicitly model the constrained transitions from target boundaries to target sentiment polarities. We also propose to maintain the sentiment consistency within an opinion target via a gate mechanism which models the relation between the features for the current word and the previous word. We conduct extensive experiments on three benchmark datasets and our framework achieves consistently superior results.

IJCAI Conference 2019 Conference Paper

A Vectorized Relational Graph Convolutional Network for Multi-Relational Network Alignment

  • Rui Ye
  • Xin Li
  • Yujie Fang
  • Hongyu Zang
  • Mingzhong Wang

Alignment of multiple multi-relational networks, such as knowledge graphs, is vital for AI applications. Different from the conventional alignment models, we apply the graph convolutional network (GCN) to achieve more robust network embedding for the alignment task. In comparison with existing GCNs which cannot fully utilize multi-relation information, we propose a vectorized relational graph convolutional network (VR-GCN) to learn the embeddings of both graph entities and relations simultaneously for multi-relational networks. The role discrimination and translation property of knowledge graphs are adopted in the convolutional process. Thereafter, AVR-GCN, the alignment framework based on VR-GCN, is developed for multi-relational network alignment tasks. Anchors are used to supervise the objective function which aims at minimizing the distances between anchors, and to generate new cross-network triplets to build a bridge between different knowledge graphs at the level of triplet to improve the performance of alignment. Experiments on real-world datasets show that the proposed solutions outperform the state-of-the-art methods in terms of network embedding, entity alignment, and relation alignment.

IJCAI Conference 2019 Conference Paper

Decoding EEG by Visual-guided Deep Neural Networks

  • Zhicheng Jiao
  • Haoxuan You
  • Fan Yang
  • Xin Li
  • Han Zhang
  • Dinggang Shen

Decoding visual stimuli from brain activities is an interdisciplinary study of neuroscience and computer vision. With the emerging of Human-AI Collaboration, Human-Computer Interaction, and the development of advanced machine learning models, brain decoding based on deep learning attracts more attention. Electroencephalogram (EEG) is a widely used neurophysiology tool. Inspired by the success of deep learning on image representation and neural decoding, we proposed a visual-guided EEG decoding method that contains a decoding stage and a generation stage. In the classification stage, we designed a visual-guided convolutional neural network (CNN) to obtain more discriminative representations from EEG, which are applied to achieve the classification results. In the generation stage, the visual-guided EEG features are input to our improved deep generative model with a visual consistence module to generate corresponding visual stimuli. With the help of our visual-guided strategies, the proposed method outperforms traditional machine learning methods and deep learning models in the EEG decoding task.

AAAI Conference 2019 Conference Paper

Exploiting Coarse-to-Fine Task Transfer for Aspect-Level Sentiment Classification

  • Zheng Li
  • Ying Wei
  • Yu Zhang
  • Xiang Zhang
  • Xin Li

Aspect-level sentiment classification (ASC) aims at identifying sentiment polarities towards aspects in a sentence, where the aspect can behave as a general Aspect Category (AC) or a specific Aspect Term (AT). However, due to the especially expensive and labor-intensive labeling, existing public corpora in AT-level are all relatively small. Meanwhile, most of the previous methods rely on complicated structures with given scarce data, which largely limits the efficacy of the neural models. In this paper, we exploit a new direction named coarse-to-fine task transfer, which aims to leverage knowledge learned from a rich-resource source domain of the coarse-grained AC task, which is more easily accessible, to improve the learning in a low-resource target domain of the fine-grained AT task. To resolve both the aspect granularity inconsistency and feature mismatch between domains, we propose a Multi-Granularity Alignment Network (MGAN). In MGAN, a novel Coarse2Fine attention guided by an auxiliary task can help the AC task modeling at the same finegrained level with the AT task. To alleviate the feature false alignment, a contrastive feature alignment method is adopted to align aspect-specific feature representations semantically. In addition, a large-scale multi-domain dataset for the AC task is provided. Empirically, extensive experiments demonstrate the effectiveness of the MGAN.

AIIM Journal 2018 Journal Article

Approximate dynamic programming approaches for appointment scheduling with patient preferences

  • Xin Li
  • Jin Wang
  • Richard Y.K. Fung

During the appointment booking process in out-patient departments, the level of patient satisfaction can be affected by whether or not their preferences can be met, including the choice of physicians and preferred time slot. In addition, because the appointments are sequential, considering future possible requests is also necessary for a successful appointment system. This paper proposes a Markov decision process model for optimizing the scheduling of sequential appointments with patient preferences. In contrast to existing models, the evaluation of a booking decision in this model focuses on the extent to which preferences are satisfied. Characteristics of the model are analysed to develop a system for formulating booking policies. Based on these characteristics, two types of approximate dynamic programming algorithms are developed to avoid the curse of dimensionality. Experimental results suggest directions for further fine-tuning of the model, as well as improving the efficiency of the two proposed algorithms.

IJCAI Conference 2018 Conference Paper

Aspect Term Extraction with History Attention and Selective Transformation

  • Xin Li
  • Lidong Bing
  • Piji Li
  • Wai Lam
  • Zhimou Yang

Aspect Term Extraction (ATE), a key sub-task in Aspect-Based Sentiment Analysis, aims to extract explicit aspect expressions from online user reviews. We present a new framework for tackling ATE. It can exploit two useful clues, namely opinion summary and aspect detection history. Opinion summary is distilled from the whole input sentence, conditioned on each current token for aspect prediction, and thus the tailor-made summary can help aspect prediction on this token. On the other hand, the aspect detection history information is distilled from the previous aspect predictions, and it can leverage the coordinate structure and tagging schema constraints to upgrade the aspect prediction. Experimental results over four benchmark datasets clearly demonstrate that our framework can outperform all state-of-the-art methods.

IJCAI Conference 2018 Conference Paper

Automatic Opioid User Detection from Twitter: Transductive Ensemble Built on Different Meta-graph Based Similarities over Heterogeneous Information Network

  • Yujie Fan
  • Yiming Zhang
  • Yanfang Ye
  • Xin Li

Opioid (e. g. , heroin and morphine) addiction has become one of the largest and deadliest epidemics in the United States. To combat such deadly epidemic, in this paper, we propose a novel framework named HinOPU to automatically detect opioid users from Twitter, which will assist in sharpening our understanding toward the behavioral process of opioid addiction and treatment. In HinOPU, to model the users and the posted tweets as well as their rich relationships, we introduce structured heterogeneous information network (HIN) for representation. Afterwards, we use meta-graph based approach to characterize the semantic relatedness over users; we then formulate different similarities over users based on different meta-graphs on HIN. To reduce the cost of acquiring labeled samples for supervised learning, we propose a transductive classification method to build the base classifiers based on different similarities formulated by different meta-graphs. Then, to further improve the detection accuracy, we construct an ensemble to combine different predictions from different base classifiers for opioid user detection. Comprehensive experiments on real sample collections from Twitter are conducted to validate the effectiveness of HinOPU in opioid user detection by comparisons with other alternate methods.

AAAI Conference 2018 Conference Paper

Multi-Scale Bidirectional FCN for Object Skeleton Extraction

  • Fan Yang
  • Xin Li
  • Hong Cheng
  • Yuxiao Guo
  • Leiting Chen
  • Jianping Li

Object skeleton detection is a challenging problem with wide application. Recently, deep Convolutional Neural Networks (CNNs) have substantially improved the performance of the state-of-the-art in this task. However, most of the existing CNN-Based methods are based on a skip-layer structure where low-level and high-level features are combined and learned so as to gather multi-level contextual information. As shallow features are too messy and lack semantic knowledge, they may cause errors and inaccuracy. Therefore, we propose a novel network architecture, Multi-Scale Bidirectional Fully Convolutional Network (MSB-FCN), to better capture and consolidate multi-scale high-level context information for object skeleton detection. Our network uses only deep features to build multi-scale feature representations, and employs a bidirectional structure to collect contextual knowledge. Hence the proposed MSB-FCN has the ability to learn the semantic-level information from different sub-regions. Furthermore, we introduce dense connections into the bidirectional structure of our MSB-FCN to ensure that the learning process at each scale can directly encode information from all other scales. Extensive experiments on various commonly used benchmarks demonstrate that the proposed MSB- FCN has achieved significant improvements over the state-ofthe-art algorithms.

IJCAI Conference 2018 Conference Paper

Non-translational Alignment for Multi-relational Networks

  • Shengnan Li
  • Xin Li
  • Rui Ye
  • Mingzhong Wang
  • Haiping Su
  • Yingzi Ou

Most existing solutions for the alignment of multi-relational networks, such as multi-lingual knowledge bases, are ``translation''-based which facilitate the network embedding via the trans-family, such as TransE. However, they cannot address triangular or other structural properties effectively. Thus, we propose a non-translational approach, which aims to utilize a probabilistic model to offer more robust solutions to the alignment task, by exploring the structural properties as well as leveraging on anchors to project each network onto the same vector space during the process of learning the representation of individual networks. The extensive experiments on four multi-lingual knowledge graphs demonstrate the effectiveness and robustness of the proposed method over a set of state-of-the-art alignment methods.

IJCAI Conference 2017 Conference Paper

A Structural Representation Learning for Multi-relational Networks

  • Lin Liu
  • Xin Li
  • William K. Cheung
  • Chengcheng Xu

Most of the existing multi-relational network embedding methods, e. g. , TransE, are formulated to preserve pair-wise connectivity structures in the networks. With the observations that significant triangular connectivity structures and parallelogram connectivity structures found in many real multi-relational networks are often ignored and that a hard-constraint commonly adopted by most of the network embedding methods is inaccurate by design, we propose a novel representation learning model for multi-relational networks which can alleviate both fundamental limitations. Scalable learning algorithms are derived using the stochastic gradient descent algorithm and negative sampling. Extensive experiments on real multi-relational network datasets of WordNet and Freebase demonstrate the efficacy of the proposed model when compared with the state-of-the-art embedding methods.

IJCAI Conference 2017 Conference Paper

Category-aware Next Point-of-Interest Recommendation via Listwise Bayesian Personalized Ranking

  • Jing He
  • Xin Li
  • Lejian Liao

Next Point-of-interest (POI) recommendation has become an important task for location-based social networks (LBSNs). However, previous efforts suffer from the high computational complexity and the transition pattern between POIs has not been well studied. In this paper, we propose a two-fold approach for next POI recommendation. First, the preferred next category is predicted by using a third-rank tensor optimized by a Listwise Bayesian Personalized Ranking (LBPR) approach. Specifically we introduce two functions, namely Plackett-Luce model and cross entropy, to generate the likelihood of ranking list for posterior computation. Then POI candidates filtered by the predicated category are ranked based on the spatial influence and category ranking influence. Extensive experiments on two real-world datasets demonstrate the significant improvements of our methods over several state-of-the-art methods.

IJCAI Conference 2016 Conference Paper

Aligning Users across Social Networks Using Network Embedding

  • Li Liu
  • William K. Cheung
  • Xin Li
  • Lejian Liao

In this paper, we adopt the representation learning approach to align users across multiple social networks where the social structures of the users are exploited. In particular, we propose to learn a network embedding with the follower-ship/followee-ship of each user explicitly modeled as input/output context vector representations so as to preserve the proximity of users with "similar" followers/followees in the embedded space. For the alignment, we add both known and potential anchor users across the networks to facilitate the transfer of context information across networks. We solve both the network embedding problem and the user alignment problem simultaneously under a unified optimization framework. The stochastic gradient descent and negative sampling algorithms are used to address scalability issues. Extensive experiments on real social network datasets demonstrate the effectiveness and efficiency of the proposed approach compared with several state-of-the-art methods.

IROS Conference 2016 Conference Paper

An egocentric computer vision based co-robot wheelchair

  • Haoxiang Li
  • Mohammed Kutbi
  • Xin Li
  • Changjiang Cai
  • Philippos Mordohai
  • Gang Hua 0001

Motivated by the emerging needs to improve the quality of life for the elderly and disabled individuals who rely on wheelchairs for mobility, and who might have limited or no hand functionality at all, we propose an egocentric computer vision based co-robot wheelchair to enhance their mobility without hand usage. The co-robot wheelchair is built upon a typical commercial power wheelchair. The user can access 360 degrees of motion direction as well as a continuous range of speed without the use of hands via the egocentric computer vision based control we developed. The user wears an egocentric camera and collaborates with the robotic wheelchair by conveying the motion commands with head motions. Compared with previous sip-n-puff, chin-control and tongue-operated solutions to hands-free mobility, this egocentric computer vision based control system provides a more natural human robot interface. Our experiments show that this design is of higher usability and users can quickly learn to control and operate the wheelchair. Besides its convenience in manual navigation, the egocentric camera also supports novel user-robot interaction modes by enabling autonomous navigation towards a detected person or object of interest. User studies demonstrate the usability and efficiency of the proposed egocentric computer vision co-robot wheelchair.

AAAI Conference 2016 Conference Paper

Inferring a Personalized Next Point-of-Interest Recommendation Model with Latent Behavior Patterns

  • Jing He
  • Xin Li
  • Lejian Liao
  • Dandan Song
  • William Cheung

In this paper, we address the problem of personalized next Point-of-interest (POI) recommendation which has become an important and very challenging task in location-based social networks (LBSNs), but not well studied yet. With the conjecture that, under different contextual scenario, human exhibits distinct mobility patterns, we attempt here to jointly model the next POI recommendation under the influence of user’s latent behavior pattern. We propose to adopt a third-rank tensor to model the successive check-in behaviors. By incorporating softmax function to fuse the personalized Markov chain with latent pattern, we furnish a Bayesian Personalized Ranking (BPR) approach and derive the optimization criterion accordingly. Expectation Maximization (EM) is then used to estimate the model parameters. Extensive experiments on two large-scale LB- SNs datasets demonstrate the significant improvements of our model over several state-of-the-art methods.

NeurIPS Conference 2016 Conference Paper

Learning Parametric Sparse Models for Image Super-Resolution

  • Yongbo Li
  • Weisheng Dong
  • Xuemei Xie
  • Guangming Shi
  • Xin Li
  • Donglai Xu

Learning accurate prior knowledge of natural images is of great importance for single image super-resolution (SR). Existing SR methods either learn the prior from the low/high-resolution patch pairs or estimate the prior models from the input low-resolution (LR) image. Specifically, high-frequency details are learned in the former methods. Though effective, they are heuristic and have limitations in dealing with blurred LR images; while the latter suffers from the limitations of frequency aliasing. In this paper, we propose to combine those two lines of ideas for image super-resolution. More specifically, the parametric sparse prior of the desirable high-resolution (HR) image patches are learned from both the input low-resolution (LR) image and a training image dataset. With the learned sparse priors, the sparse codes and thus the HR image patches can be accurately recovered by solving a sparse coding problem. Experimental results show that the proposed SR method outperforms existing state-of-the-art methods in terms of both subjective and objective image qualities.

IJCAI Conference 2016 Conference Paper

Saliency Transfer: An Example-Based Method for Salient Object Detection

  • Xin Li
  • Fan Yang
  • Leiting Chen
  • Hongbin Cai

Over the past decades, numerous theories and studies have demonstrated that salient objects in different scenes often share some properties in common that make them visually stand out from their surroundings, and thus can be processed in finer details. In this paper, we propose a novel method for salient object detection that involves the transfer of the annotations from an existing example onto an input image. Our method, which is based on the low-level saliency features of each pixel, estimates dense pixel-wise correspondences between the input image and an example image, and then integrates high-level concepts to produce an initial saliency map. Finally, a coarse-to-fine optimization framework is proposed to generate uniformly highlighted salient objects. Qualitatively and quantitatively experiments on six popular benchmark datasets validate that our approach greatly outperforms the state-of-the-art algorithms and recently published works.

AAAI Conference 2016 Conference Paper

Social Emotion Classification via Reader Perspective Weighted Model

  • Xin Li
  • Yanghui Rao
  • Yanjia Chen
  • Xuebo Liu
  • Huan Huang

With the development of Web 2. 0, many users express their opinions online. This paper is concerned with the classification of social emotions on varied-scale data sets. Different from traditional models which weight training documents equally, the concept of emotional entropy is proposed to estimate the weight and tackle the issue of noisy documents. The topic assignment is also used to distinguish different emotional senses of the same word. Experimental evaluations using different data sets validate the effectiveness of the proposed social emotion classification model.

IJCAI Conference 2015 Conference Paper

Detecting Promotion Campaigns in Community Question Answering

  • Xin Li
  • Yiqun Liu
  • Min Zhang
  • Shaoping Ma
  • Xuan Zhu
  • Jiashen Sun

With Community Question Answering (CQA) evolving into a quite popular method for information seeking and providing, it also becomes a target for spammers to disseminate promotion campaigns. Although there are a number of quality estimation efforts on the CQA platform, most of these works focus on identifying and reducing lowquality answers, which are mostly generated by impatient or inexperienced answerers. However, a large number of promotion answers appear to provide high-quality information to cheat CQA users in future interactions. Therefore, most existing quality estimation works in CQA may fail to detect these specially designed answers or question-answer pairs. In contrast to these works, we focus on the promotion channels of spammers, which include (shortened) URLs, telephone numbers and social media accounts. Spammers rely on these channels to connect to users to achieve promotion goals so they are irreplaceable for spamming activities. We propose a propagation algorithm to diffuse promotion intents on an “answerer-channel” bipartite graph and detect possible spamming activities. A supervised learning framework is also proposed to identify whether a QA pair is spam based on propagated promotion intents. Experimental results based on more than 6 million entries from a popular Chinese CQA portal show that our approach outperforms a number of existing quality estimation methods for detecting promotion campaigns on both the answer level and QA pair level.

IJCAI Conference 2015 Conference Paper

Multi-Label Classification with Feature-Aware Non-Linear Label Space Transformation

  • Xin Li
  • Yuhong Guo

Multi-label classification with many classes has recently drawn a lot of attention. Existing methods address this problem by performing linear label space transformation to reduce the dimension of label space, and then conducting independent regression for each reduced label dimension. These methods however do not capture nonlinear correlations of the multiple labels and may lead to significant information loss in the process of label space reduction. In this paper, we first propose to exploit kernel canonical correlation analysis (KCCA) to capture nonlinear label correlation information and perform nonlinear label space reduction. Then we develop a novel label space reduction method that explicitly combines linear and nonlinear label space transformations based on CCA and KCCA respectively to address multi-label classification with many classes. The proposed method is a feature-aware label transformation method that promotes the label predictability in the transformed label space from the input features. We conduct experiments on a number of multi-label classification datasets. The proposed approach demonstrates good performance, comparing to a number of stateof-the-art label dimension reduction methods.

IS Journal 2015 Journal Article

Simple is beautiful: Toward light prediction markets

  • Weiyun Chen
  • Xin Li
  • Daniel D. Zeng

This column examines the cognitive load of prediction markets from four steps of user's decision process (that is, timing, pricing, revisiting, and benefit) and discusses how a prediction market mechanism could have a low cognitive load. In the column, the authors propose that fixed-odds betting can be used as a prediction market mechanism with carefully designed event probability estimators. The development of low-cognitive load prediction markets needs the join efforts of computer scientists and economists.

AAAI Conference 2014 Conference Paper

Fraudulent Support Telephone Number Identification Based on Co-Occurrence Information on the Web

  • Xin Li
  • Yiqun Liu
  • Min Zhang
  • Shaoping Ma

“Fraudulent support phones” refers to the misleading telephone numbers placed on Web pages or other media that claim to provide services with which they are not associated. Most fraudulent support phone information is found on search engine result pages (SERPs), and such information substantially degrades the search engine user experience. In this paper, we propose an approach to identify fraudulent support telephone numbers on the Web based on the co-occurrence relations between telephone numbers that appear on SERPs. We start from a small set of seed official support phone numbers and seed fraudulent numbers. Then, we construct a co-occurrence graph according to the co-occurrence relationships of the telephone numbers that appear on Web pages. Additionally, we take the page layout information into consideration on the assumption that telephone numbers that appear in nearby page blocks should be regarded as more closely related. Finally, we develop a propagation algorithm to diffuse the trust scores of seed official support phone numbers and the distrust scores of the seed fraudulent numbers on the co-occurrence graph to detect additional fraudulent numbers. Experimental results based on over 1. 5 million SERPs produced by a popular Chinese commercial search engine indicate that our approach outperforms TrustRank, Anti-TrustRank and Good-Bad Rank algorithms by achieving an AUC value of over 0. 90.

YNIMG Journal 2014 Journal Article

Lack of dystrophin results in abnormal cerebral diffusion and perfusion in vivo

  • Candida L. Goodnough
  • Ying Gao
  • Xin Li
  • Mohammed Q. Qutaish
  • L. Henry Goodnough
  • Joseph Molter
  • David Wilson
  • Chris A. Flask

Dystrophin, the main component of the dystrophin–glycoprotein complex, plays an important role in maintaining the structural integrity of cells. It is also involved in the formation of the blood–brain barrier (BBB). To elucidate the impact of dystrophin disruption in vivo, we characterized changes in cerebral perfusion and diffusion in dystrophin-deficient mice (mdx) by magnetic resonance imaging (MRI). Arterial spin labeling (ASL) and diffusion-weighted MRI (DWI) studies were performed on 2-month-old and 10-month-old mdx mice and their age-matched wild-type controls (WT). The imaging results were correlated with Evan's blue extravasation and vascular density studies. The results show that dystrophin disruption significantly decreased the mean cerebral diffusivity in both 2-month-old (7. 38±0. 30×10-4 mm2/s) and 10-month-old (6. 93±0. 53×10-4 mm2/s) mdx mice as compared to WT (8. 49±0. 24×10-4, 8. 24±0. 25×10-4 mm2/s, respectively). There was also an 18% decrease in cerebral perfusion in 10-month-old mdx mice as compared to WT, which was associated with enhanced arteriogenesis. The reduction in water diffusivity in mdx mice is likely due to an increase in cerebral edema or the existence of large molecules in the extracellular space from a leaky BBB. The observation of decreased perfusion in the setting of enhanced arteriogenesis may be caused by an increase of intracranial pressure from cerebral edema. This study demonstrates the defects in water handling at the BBB and consequently, abnormal perfusion associated with the absence of dystrophin.

YNIMG Journal 2014 Journal Article

Tools for multiple granularity analysis of brain MRI data for individualized image analysis

  • Aigerim Djamanakova
  • Xiaoying Tang
  • Xin Li
  • Andreia V. Faria
  • Can Ceritoglu
  • Kenichi Oishi
  • Argye E. Hillis
  • Marilyn Albert

Voxel-based analysis is widely used for quantitative analysis of brain MRI. While this type of analysis provides the highest granularity level of spatial information (i. e. , each voxel), the sheer number of voxels and noisy information from each voxel often lead to low sensitivity for detection of abnormalities. To ameliorate this issue, granularity reduction is commonly performed by applying isotropic spatial filtering. This study proposes a systematic reduction of the spatial information using ontology-based hierarchical structural relationships. The 254 brain structures were first defined in multiple (n =29) geriatric atlases. The multiple atlases were then applied to T1-weighted MR images of each subject's data for automated brain parcellation and five levels of ontological relationships were established, which further reduced the spatial dimension to as few as 11 structures. At each ontology level, the amount of atrophy was evaluated, providing a unique view of low-granularity analysis. This reduction of spatial information allowed us to investigate the anatomical features of each patient, demonstrated in an Alzheimer's disease group.

IJCAI Conference 2013 Conference Paper

Active Learning with Multi-Label SVM Classification

  • Xin Li
  • Yuhong Guo

Multi-label classification, where each instance is assigned to multiple categories, is a prevalent problem in data analysis. However, annotations of multi-label instances are typically more timeconsuming or expensive to obtain than annotations of single-label instances. Though active learning has been widely studied on reducing labeling effort for single-label problems, current research on multi-label active learning remains in a preliminary state. In this paper, we first propose two novel multi-label active learning strategies, a max-margin prediction uncertainty strategy and a label cardinality inconsistency strategy, and then integrate them into an adaptive framework of multi-label active learning. Our empirical results on multiple multilabel data sets demonstrate the efficacy of the proposed active instance selection strategies and the integrated active learning approach.

ICRA Conference 2012 Conference Paper

Online identification of quality of teleoperator (QoT) for performance improvement of telerobotic operations

  • Yunyi Jia
  • Ning Xi 0001
  • Yunxia Wang
  • Xin Li

In teleoperation studies, most researchers have been researching on the stability and telepresence. Few of them have studied the influence of the operational status of the teleoperator on the telerobotic systems. As a matter of fact, improper and incorrect operations of the teleoperator may decrease the teleoperation efficiency and even result in some serious safety problems even if the stability and telepresence are both guaranteed. Thus, this paper investigates a method to online identify the quality of the teleoperator and then integrate it into the planning and control of the telerobotic system. The method can help improve the performance of the system including efficiency and safety. It is also implemented on a mobile manipulator and the experimental results illustrate the effectiveness of the designed method.

IROS Conference 2011 Conference Paper

Controlling telerobotic operations adaptive to quality of teleoperator and task dexterity

  • Yunyi Jia
  • Ning Xi 0001
  • Fei Wang
  • Yunxia Wang
  • Xin Li

Telerobotic systems have been researched for decades due to their extensive applications in many civilian and military areas. Most research mainly focused on either the stability or telepresence of the telerobotic systems. Few studies have investigated the effects of the confidence of the teleoperator on the performance of the teleoperation. The confidence of the teleoperator is of significant importance to the efficiency and safety of the telerobotic systems. This paper proposes a concept named quality of teleoperator (QoT) to represent the confidence of the decisions and commands generated by the teleoperator. The value of QoT is computed based on a set of mental states of the teleoperator. Based on the QoT, a control adjustment mechanism is designed to enhance the efficiency and safety of the telerobotic systems. Experiments were implemented on a manipulator to demonstrate the effectiveness of the proposed method.

YNIMG Journal 2011 Journal Article

Multi-contrast human neonatal brain atlas: Application to normal neonate development analysis

  • Kenichi Oishi
  • Susumu Mori
  • Pamela K. Donohue
  • Thomas Ernst
  • Lynn Anderson
  • Steven Buchthal
  • Andreia Faria
  • Hangyi Jiang

MRI is a sensitive method for detecting subtle anatomic abnormalities in the neonatal brain. To optimize the usefulness for neonatal and pediatric care, systematic research, based on quantitative image analysis and functional correlation, is required. Normalization-based image analysis is one of the most effective methods for image quantification and statistical comparison. However, the application of this methodology to neonatal brain MRI scans is rare. Some of the difficulties are the rapid changes in T1 and T2 contrasts and the lack of contrast between brain structures, which prohibits accurate cross-subject image registration. Diffusion tensor imaging (DTI), which provides rich and quantitative anatomical contrast in neonate brains, is an ideal technology for normalization-based neonatal brain analysis. In this paper, we report the development of neonatal brain atlases with detailed anatomic information derived from DTI and co-registered anatomical MRI. Combined with a diffeomorphic transformation, we were able to normalize neonatal brain images to the atlas space and three-dimensionally parcellate images into 122 regions. The accuracy of the normalization was comparable to the reliability of human raters. This method was then applied to babies of 37–53 post-conceptional weeks to characterize developmental changes of the white matter, which indicated a posterior-to-anterior and a central-to-peripheral direction of maturation. We expect that future applications of this atlas will include investigations of the effect of prenatal events and the effects of preterm birth or low birth weights, as well as clinical applications, such as determining imaging biomarkers for various neurological disorders.

YNIMG Journal 2011 Journal Article

Quantitative analysis of brain pathology based on MRI and brain atlases—Applications for cerebral palsy

  • Andreia V. Faria
  • Alexander Hoon
  • Elaine Stashinko
  • Xin Li
  • Hangyi Jiang
  • Ameneh Mashayekh
  • Kazi Akhter
  • John Hsu

We have developed a new method to provide a comprehensive quantitative analysis of brain anatomy in cerebral palsy patients, which makes use of two techniques: diffusion tensor imaging and automated 3D whole brain segmentation based on our brain atlas and a nonlinear normalization technique (large-deformation diffeomorphic metric mapping). This method was applied to 13 patients and normal controls. The reliability of the automated segmentation revealed close agreement with the manual segmentation. We illustrate some potential applications for individual characterization and group comparison. This technique also provides a framework for determining the impact of various neuroanatomic features on brain functions.

IS Journal 2010 Journal Article

A Bibliographic Analysis of IEEE Intelligent Systems Publications

  • Zhuo Feng
  • Qingpeng Zhang
  • Xin Li
  • Guanyan Ke
  • Gang Xiong

To better understand the authors and studies in IEEE Intelligent Systems, the authors conducted a bibliographic study on its publications. Specifically, they focus on the most productive and highly cited authors and institutions in IS, the authors and institutions most cited in IS, and the journals and institutions that cited IS articles the most.

YNIMG Journal 2010 Journal Article

Atlas-based analysis of neurodevelopment from infancy to adulthood using diffusion tensor imaging and applications for automated abnormality detection

  • Andreia V. Faria
  • Jiangyang Zhang
  • Kenichi Oishi
  • Xin Li
  • Hangyi Jiang
  • Kazi Akhter
  • Laurent Hermoye
  • Seung-Koo Lee

Quantification of normal brain maturation is a crucial step in understanding developmental abnormalities in brain anatomy and function. The aim of this study was to develop atlas-based tools for time-dependent quantitative image analysis, and to characterize the anatomical changes that occur from 2years of age to adulthood. We used large deformation diffeomorphic metric mapping to register diffusion tensor images of normal participants into the common coordinates and used a pre-segmented atlas to segment the entire brain into 176 structures. Both voxel- and atlas-based analyses reported a structure that showed distinctive changes in terms of its volume and diffusivity measures. In the white matter, fractional anisotropy (FA) linearly increased with age in logarithmic scale, while diffusivity indices, such as apparent diffusion coefficient (ADC), and axial and radial diffusivity, decreased at a different rate in several regions. The average, variability, and the time course of each measured parameter are incorporated into the atlas, which can be used for automated detection of developmental abnormalities. As a demonstration of future application studies, the brainstem anatomy of cerebral palsy patients was evaluated and the altered anatomy was delineated.

YNIMG Journal 2010 Journal Article

Atlas-guided tract reconstruction for automated and comprehensive examination of the white matter anatomy

  • Yajing Zhang
  • Jiangyang Zhang
  • Kenichi Oishi
  • Andreia V. Faria
  • Hangyi Jiang
  • Xin Li
  • Kazi Akhter
  • Pedro Rosa-Neto

Tractography based on diffusion tensor imaging (DTI) is widely used to quantitatively analyze the status of the white matter anatomy in a tract-specific manner in many types of diseases. This approach, however, involves subjective judgment in the tract-editing process to extract only the tracts of interest. This process, usually performed by manual delineation of regions of interest, is also time-consuming, and certain tracts, especially the short cortico-cortical association fibers, are difficult to reconstruct. In this paper, we propose an automated approach for reconstruction of a large number of white matter tracts. In this approach, existing anatomical knowledge about tract trajectories (called the Template ROI Set or TRS) were stored in our DTI-based brain atlas with 130 three-dimensional anatomical segmentations, which were warped non-linearly to individual DTI data. We examined the degree of matching with manual results for selected fibers. We established 30 TRSs to reconstruct 30 prominent and previously well-described fibers. In addition, TRSs were developed to delineate 29 short association fibers that were found in all normal subjects examined in this paper (N=20). Probabilistic maps of the 59 tract trajectories were created from the normal subjects and were incorporated into our image analysis tool for automated tract-specific quantification.

TCS Journal 2010 Journal Article

On exponential time lower bound of Knapsack under backtracking

  • Xin Li
  • Tian Liu

We prove an Ω ( 2 0. 69 n / n ) time lower bound of Knapsack problem under the adaptive priority branching trees (pBT) model. The pBT model is a formal model of algorithms covering backtracking and dynamic programming [M. Alekhnovich, A. Borodin, A. Magen, J. Buresh-Oppenheim, R. Impagliazzo, T. Pitassi, Toward a model for backtracking and dynamic programming, ECCC TR09-038, 2009. Earlier version in Proc 20th IEEE Computational Complexity, 2005, pp. 308–322]. Our result improves the Ω ( 2 0. 5 n / n ) lower bound of M. Alekhovich et al. and the Ω ( 2 0. 66 n / n ) lower bound of Li et al. [X. Li, T. Liu, H. Peng, L. Qian, H. Sun, J. Xu, K. Xu, J. Zhu, Improved exponential time lower bound of Knapsack problem under BT model, in: Proc 4th TAMC 2007, in: LNCS, vol. 4484, 2007, pp. 624–631] through optimized arguments.

YNIMG Journal 2009 Journal Article

Atlas-based whole brain white matter analysis using large deformation diffeomorphic metric mapping: Application to normal elderly and Alzheimer's disease participants

  • Kenichi Oishi
  • Andreia Faria
  • Hangyi Jiang
  • Xin Li
  • Kazi Akhter
  • Jiangyang Zhang
  • John T. Hsu
  • Michael I. Miller

The purpose of this paper is to establish single-participant white matter atlases based on diffusion tensor imaging. As one of the applications of the atlas, automated brain segmentation was performed and the accuracy was measured using Large Deformation Diffeomorphic Metric Mapping (LDDMM). High-quality diffusion tensor imaging (DTI) data from a single-participant were B0-distortion-corrected and transformed to the ICBM-152 atlas or to Talairach coordinates. The deep white matter structures, which have been previously well documented and clearly identified by DTI, were manually segmented. The superficial white matter areas beneath the cortex were defined, based on a population-averaged white matter probability map. The white matter was parcellated into 176 regions based on the anatomical labeling in the ICBM-DTI-81 atlas. The automated parcellation was achieved by warping this parcellation map to normal controls and to Alzheimer's disease patients with severe anatomical atrophy. The parcellation accuracy was measured by a kappa analysis between the automated and manual parcellation at 11 anatomical regions. The kappa values were 0. 70 for both normal controls and patients while the inter-rater reproducibility was 0. 81 (controls) and 0. 82 (patients), suggesting “almost perfect” agreement. A power analysis suggested that the proposed method is suitable for detecting FA and size abnormalities of the white matter in clinical studies.

YNIMG Journal 2009 Journal Article

Landmark-referenced voxel-based analysis of diffusion tensor images of the brainstem white matter tracts

  • Weihong Zhang
  • Xin Li
  • Jiangyang Zhang
  • Andreas Luft
  • Daniel F. Hanley
  • Peter van Zijl
  • Michael I. Miller
  • Laurent Younes

Although DTI can provide detailed information about white matter anatomy, it is not yet straightforward enough to quantify the anatomical information it visualizes. In this study, we developed and tested a new tool to perform brain normalization and voxel-based analysis of DTI data. For the normalization part, manually placed landmarks ensured that the visualized white matter tracts were well-registered among the populations. A standard landmark set in ICBM-152 space and an interface to remap them to subject data were integrated in the procedure. After landmark placement, highly elastic non-linear Large Deformation Diffeomorphic Metric Mapping (LDDMM) was driven by the landmarks to normalize the brainstem anatomy of normal subjects. The approach was then applied to delineate brainstem tract abnormalities in patients with left chronic middle cerebral artery (MCA) stroke. The voxel-based comparison between control and patient groups identified abnormalities in the ipsilesional corticospinal tract and contralesional cerebellar peduncles. We believe that this tool is useful for regional brain normalization of patients with severe anatomical alterations, such as stroke, brain tumor, and lobectomy, for whom standard automated normalization tools may not work properly.

YNIMG Journal 2009 Journal Article

Multi-contrast large deformation diffeomorphic metric mapping for diffusion tensor imaging

  • Can Ceritoglu
  • Kenichi Oishi
  • Xin Li
  • Ming-Chung Chou
  • Laurent Younes
  • Marilyn Albert
  • Constantine Lyketsos
  • Peter C.M. van Zijl

Diffusion tensor imaging (DTI) can reveal detailed white matter anatomy and has the potential to detect abnormalities in specific white matter structures. Such detection and quantification are, however, not straightforward. The voxel-based analysis after image normalization is one of the most widely used methods for quantitative image analyses. To apply this approach to DTI, it is important to examine if structures in the white matter are well registered among subjects, which would be highly dependent on employed algorithms for normalization. In this paper, we evaluate the accuracy of normalization of DTI data using a highly elastic transformation algorithm, called large deformation diffeomorphic metric mapping. After simulation-based validation of the algorithm, DTI data from normal subjects were used to measure the registration accuracy. To examine the impact of morphological abnormalities on the accuracy, the algorithm was also tested using data from Alzheimer's disease (AD) patients with severe brain atrophy. The accuracy level was measured by using manual landmark-based white matter matching and surface-based brain and ventricle matching as gold standard. To improve the accuracy level, cascading and multi-contrast approaches were developed. The accuracy level for the white matter was 1. 88±0. 55 and 2. 19±0. 84 mm for the measured locations in the controls and patients, respectively.

YNIMG Journal 2008 Journal Article

Human brain white matter atlas: Identification and assignment of common anatomical structures in superficial white matter

  • Kenichi Oishi
  • Karl Zilles
  • Katrin Amunts
  • Andreia Faria
  • Hangyi Jiang
  • Xin Li
  • Kazi Akhter
  • Kegang Hua

Structural delineation and assignment are the fundamental steps in understanding the anatomy of the human brain. The white matter has been structurally defined in the past only at its core regions (deep white matter). However, the most peripheral white matter areas, which are interleaved between the cortex and the deep white matter, have lacked clear anatomical definitions and parcellations. We used axonal fiber alignment information from diffusion tensor imaging (DTI) to delineate the peripheral white matter, and investigated its relationship with the cortex and the deep white matter. Using DTI data from 81 healthy subjects, we identified nine common, blade-like anatomical regions, which were further parcellated into 21 subregions based on the cortical anatomy. Four short association fiber tracts connecting adjacent gyri (U-fibers) were also identified reproducibly among the healthy population. We anticipate that this atlas will be useful resource for atlas-based white matter anatomical studies.

YNIMG Journal 2008 Journal Article

Stereotaxic white matter atlas based on diffusion tensor imaging in an ICBM template

  • Susumu Mori
  • Kenichi Oishi
  • Hangyi Jiang
  • Li Jiang
  • Xin Li
  • Kazi Akhter
  • Kegang Hua
  • Andreia V. Faria

Brain registration to a stereotaxic atlas is an effective way to report anatomic locations of interest and to perform anatomic quantification. However, existing stereotaxic atlases lack comprehensive coordinate information about white matter structures. In this paper, white matter-specific atlases in stereotaxic coordinates are introduced. As a reference template, the widely used ICBM-152 was used. The atlas contains fiber orientation maps and hand-segmented white matter parcellation maps based on diffusion tensor imaging (DTI). Registration accuracy by linear and non-linear transformation was measured, and automated template-based white matter parcellation was tested. The results showed a high correlation between the manual ROI-based and the automated approaches for normal adult populations. The atlases are freely available and believed to be a useful resource as a target template and for automated parcellation methods.

YNIMG Journal 2008 Journal Article

Tract probability maps in stereotaxic spaces: Analyses of white matter anatomy and tract-specific quantification

  • Kegang Hua
  • Jiangyang Zhang
  • Setsu Wakana
  • Hangyi Jiang
  • Xin Li
  • Daniel S. Reich
  • Peter A. Calabresi
  • James J. Pekar

Diffusion tensor imaging (DTI) is an exciting new MRI modality that can reveal detailed anatomy of the white matter. DTI also allows us to approximate the 3D trajectories of major white matter bundles. By combining the identified tract coordinates with various types of MR parameter maps, such as T2 and diffusion properties, we can perform tract-specific analysis of these parameters. Unfortunately, 3D tract reconstruction is marred by noise, partial volume effects, and complicated axonal structures. Furthermore, changes in diffusion anisotropy under pathological conditions could alter the results of 3D tract reconstruction. In this study, we created a white matter parcellation atlas based on probabilistic maps of 11 major white matter tracts derived from the DTI data from 28 normal subjects. Using these probabilistic maps, automated tract-specific quantification of fractional anisotropy and mean diffusivity were performed. Excellent correlation was found between the automated and the individual tractography-based results. This tool allows efficient initial screening of the status of multiple white matter tracts.

AAAI Conference 2005 Short Paper

Self-Emergence of Structures in Gene Expression Programming

  • Xin Li

This thesis work aims at improving the problem solving ability of the Gene Expression Programming (GEP) algorithm to fulfill complex data mining tasks by preserving and utilizing the self-emergence of structures during its evolutionary process. The main contributions include the investigation of the constant creation techniques for promoting good functional structures emergent in the evolution, analysis of the limitation with the current implementation scheme of GEP, and introduction of a novel utilization of the emergent structures to achieve a flexible search process for solutions at a higher level.

AAAI Conference 2004 Conference Paper

Identification and Tracing of Ambiguous Names: Discriminative and Generative Approaches

  • Xin Li

A given entity – representing a person, a location or an organization – may be mentioned in text in multiple, ambiguous ways. Understanding natural language requires identifying whether different mentions of a name, within and across documents, represent the same entity. We present two machine learning approaches to this problem, which we call the “Robust Reading” problem. Our first approach is a discriminative approach, trained in a supervised way. Our second approach is a generative model, at the heart of which is a view on how documents are generated and how names (of different entity types) are “sprinkled” into them. In its most general form, our model assumes: (1) a joint distribution over entities (e. g. , a document that mentions “President Kennedy” is more likely to mention “Oswald” or “ White House” than “Roger Clemens”), (2) an “author” model, that assumes that at least one mention of an entity in a document is easily identifiable, and then generates other mentions via (3) an appearance model, governing how mentions are transformed from the “representative” mention. We show that both approaches perform very accurately, in the range of 90% − 95% F1 measure for different entity types, much better than previous approaches to (some aspects of) this problem. Our extensive experiments exhibit the contribution of relational and structural features and, somewhat surprisingly, that the assumptions made within our generative model are strong enough to yield a very powerful approach, that performs better than a supervised approach with limited supervised information.

v2026.09.13