Arrow Research search

Author name cluster

Xi Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

48 papers
2 author rows

Possible papers

48

TMLR Journal 2026 Journal Article

Are vision language models robust to classic uncertainty challenges?

  • Xi Wang
  • Eric Nalisnick

Robustness against uncertain and ambiguous inputs is a critical challenge for deep learning models. While recent advancements in large scale vision language models (VLMs, e.g. GPT-4o) might suggest that increasing model and training dataset size would mitigate this issue, our empirical evaluation shows a more complicated picture. In this work, we sanity check whether modern VLMs pass the two most ``classic'' uncertainty quantification challenges: Anomaly detection and classification under inherently ambiguous conditions, we find that newer and larger VLMs indeed exhibit improved robustness compared to earlier models, but still suffer from a tendency to strictly follow instructions, often causing them to hallucinate confident responses even when faced with unclear or anomalous inputs. Remarkably, for natural images such as ImageNet, this limitation can be overcome without pipeline modifications: simply prompting models to abstain from uncertain predictions enables significant reliability gains, achieving near-perfect robustness in several settings. However, for domain-specific tasks such as galaxy morphology classification, a lack of specialized knowledge prevents reliable uncertainty estimation. Finally, we propose a simple mechanism based on caption diversity to reveal a model’s internal uncertainty, enabling practitioners to predict when models will successfully abstain without relying on labeled data.

AAAI Conference 2026 Conference Paper

Autonomous Vehicle Path Planning by Searching with Differentiable Simulation

  • Asen Nachkov
  • Jan-Nico Zaech
  • Danda Pani Paudel
  • Xi Wang
  • Luc Van Gool

Planning allows an agent to safely refine its actions before executing them in the real world. In autonomous driving, this is crucial to avoid collisions and navigate in complex, dense traffic scenarios. One way to plan is to search for the best action sequence. However, this is challenging when all necessary components – policy, next-state predictor, and critic – have to be learned. Here we propose Differentiable Simulation for Search (DSS), a framework that leverages the differentiable simulator Waymax as both a next state predictor and a critic. It relies on the simulator’s hardcoded dynamics, making state predictions highly accurate, while utilizing the simulator’s differentiability to effectively search across action sequences. Our DSS agent optimizes its actions using gradient descent over imagined future trajectories. We show experimentally that DSS – the combination of planning gradients and stochastic search – significantly improves tracking and path planning accuracy compared to sequence prediction, imitation learning, model-free RL, and other planning methods.

AAAI Conference 2026 Conference Paper

ChipMind: Retrieval-Augmented Reasoning for Long-Context Circuit Design Specifications

  • Changwen Xing
  • SamZaak Wong
  • Xinlai Wan
  • Yanfeng Lu
  • Mengli Zhang
  • Zebin Ma
  • Lei Qi
  • Zhengxiong Li

While Large Language Models (LLMs) demonstrate immense potential for automating integrated circuit (IC) development, their practical deployment is fundamentally limited by restricted context windows. Existing context-extension methods struggle to achieve effective semantic modeling and thorough multi-hop reasoning over extensive, intricate circuit specifications. To address this, we introduce ChipMind, a novel knowledge graph-augmented reasoning framework specifically designed for lengthy IC specifications. ChipMind first transforms circuit specifications into a domain-specific knowledge graph (ChipKG) through the Circuit Semantic-Aware Knowledge Graph Construction methodology. It then leverages the ChipKG-Augmented Reasoning mechanism, combining information-theoretic adaptive retrieval to dynamically trace logical dependencies with intent-aware semantic filtering to prune irrelevant noise, effectively balancing retrieval completeness and precision. Evaluated on an industrial-scale specification reasoning benchmark, ChipMind significantly outperforms state-of-the-art baselines, achieving an average improvement of 34.59% (up to 72.73%). Our framework bridges a critical gap between academic research and practical industrial deployment of LLM-aided Hardware Design (LAD).

AAAI Conference 2026 Conference Paper

Decomposing Prompts, Composing Actions: A Multi-Granularity Prompting Approach for Incremental Action Learning

  • Xinyi Cheng
  • Chenghao Xu
  • Xi Wang
  • Jiexi Yan
  • Yanhua Yang

Continual learning for action recognition is a critical capability for next-generation Extended Reality (XR) systems. Yet it faces a severe real-world challenge: strict user privacy that prohibits data rehearsal. While recent prompt-based continual learning methods show promise, we argue their core 'flat,' single-granularity design fundamentally misaligns with the complexity of human actions. This monolithic architecture fails to model the inherent hierarchical structure and overlooks standard action primitives shared across tasks, resulting in suboptimal performance and hindered knowledge transfer. To overcome this limitation, we propose DPCA, a novel spatio-temporal continual learning framework with multi-granularity adaptive prompting. DPCA learns three synergistic components to resolve this mismatch. First, the task-specific prompter employs a multi-granularity query system to capture the unique, compositional semantics of each action. Second, the task-agnostic prompter learns a globally shared vocabulary of ``action primitives," providing a stable and generalizable knowledge base to mitigate catastrophic forgetting. Finally, we introduce a Dissimilarity Attention Rectification at each granularity level, leveraging a reverse attention mechanism to model class-agnostic background information and effectively alleviating overfitting. The synergy between these components enables robust model adaptation without requiring access to past data. Rigorous experiments on multiple large-scale benchmarks (including NTU RGB+D), under a strict rehearsal-free, few-shot protocol, confirm that DPCA establishes a new state-of-the-art. This advance paves the way for the realization of truly adaptive and privacy-respecting XR systems.

YNICL Journal 2026 Journal Article

Distinct neurologic state in patients with traumatic brain injury and hemorrhagic stroke during the stage of acute disorders of consciousness and the correlation with the neurological prognosis: A multi-modal PET/rs-fMRI study

  • Danjing Yu
  • Kemeng Gao
  • Xiefeng Wang
  • Lin Zhao
  • Yi Sun
  • Zhiyan Shen
  • Yu Wang
  • Ying Wang

PURPOSE: The exact mechanisms underlying the distinct neurological outcomes between Traumatic Brain Injury (TBI) and Hemorrhagic Stroke (HS) remain unclear. Our objective is to assess distinct features of neurologic state between comatose patients with TBI and HS during the stage of acute disorder of consciousness (aDoC) and to identify the correlation of neurologic features with prognosis. METHODS: Data were analyzed from TBI and HS patients examined by positron emission tomography (PET) and resting-state functional magnetic resonance imaging (rs-fMRI) simultaneously. Primary clinical outcomes consisted of the state of consciousness and neurological prognosis. The regional neural activity was assessed by the amplitude of fractional low-frequency fluctuation (fALFF) and regional homogeneity (ReHo) on rs-fMRI scans. The standardized uptake value (SUV) on PET scans quantified neural metabolism. Functional connectivity (FC) and graph theoretic approach (GTA) were employed to compare the FC patterns between TBI and HS. Correlations of PET/rs-fMRI indicators with the prognosis of HS and TBI were identified. RESULTS: Muti-modal PET/rs-fMRI analysis showed more active local neurological state in TBI patients than HS patients, specifically in the right precentral gyrus (PreCG.R), right postcentral gyrus (PoCG.R), right superior temporal gyrus (STG.R) and right middle temporal gyrus (MTG.R). TBI patients demonstrated significantly higher clustering coefficient and nodal efficiency of the sensorimotor network (SMN) along with lower connectivity and network efficiency in the default network (DMN) compared to HS patients. PET/rs-fMRI indicators significantly correlated with the neurological prognosis of TBI and HS. CONCLUSIONS: This study elucidated the underlying mechanisms contributing to the distinct neurologic prognosis between comatose TBI and HS patients, and may contribute to the development of early targeted intervention strategies for specific diseases.

AAAI Conference 2026 Conference Paper

FIXME: Towards End-to-End Benchmarking of LLM-Aided Design Verification

  • Gwok-Waa Wan
  • SamZaak Wong
  • Shengchu Su
  • Chenxu Niu
  • Ning Wang
  • Xinlai Wan
  • Qixiang Chen
  • Mengnv Xing

We introduce FIXME, the first end-to-end and large-scale benchmark for evaluating Large Language Models (LLMs) in hardware design functional verification (FV). Comprising 747 tasks derived from real-world hardware designs, FIXME spans five core FV sub-sets: specification comprehension, reference model generation, testbench generation, assertion design, and RTL debugging. To ensure high data quality, we developed an AI-human collaborative framework for agile data curation and annotation. This process resulted in 25,000 lines of verified RTL, 35,000 lines of enhanced testbenches, and over 1,200 SystemVerilog Assertions. Furthermore, through expert-guided optimization within the multi-agent aided flow, we achieved a remarkable 45.57% improvement in average functional coverage, underscoring the benchmark's robustness. Through evaluation of state-of-the-art LLMs like GPT-4.1, FIXME identifies key limitations and provides actionable insights, advancing the potential of LLM-driven automation in hardware design functional verification.

AAAI Conference 2026 Conference Paper

Hallucination as a Computational Boundary: A Hierarchy of Inevitability and the Oracle Escape

  • Xi Wang
  • Quan Shi
  • Zenghui Ding
  • Jianqing Gao
  • Xianjun Yang

The illusion phenomenon of large language models (LLMs) is the core obstacle to their reliable deployment. This article formalizes the large language model as a probabilistic Turing machine by constructing a "computational necessity hierarchy", and for the first time proves the illusions are inevitable on diagonalization, incomputability, and information theory boundaries supported by the new "learner pump lemma". However, we propose two "escape routes": one is to model Retrieval Enhanced Generations (RAGs) as oracle machines, proving their absolute escape through "computational jumps", providing the first formal theory for the effectiveness of RAGs; The second is to formalize continuous learning as an "internalized oracle" mechanism and implement this path through a novel neural game theory framework.Finally, this article proposes a feasible new principle for artificial intelligence security - Computational Class Alignment (CCA), which requires strict matching between task complexity and the actual computing power of the system, providing theoretical support for the secure application of artificial intelligence.

AAAI Conference 2026 Conference Paper

RABot: Reinforcement-Guided Graph Augmentation for Imbalanced and Noisy Social Bot Detection

  • Longlong Zhang
  • Xi Wang
  • Haotong Du
  • Yangyi Xu
  • Zhuo Liu
  • Yang Liu

Social bot detection is pivotal for safeguarding the integrity of online information ecosystems. Although recent graph neural network (GNN) solutions achieve strong results, they remain hindered by two practical challenges: (i) severe class imbalance arising from the high cost of generating bots, and (ii) topological noise introduced by bots that skillfully mimic human behavior and forge deceptive links. We propose the Reinforcement-guided graph Augmentation social Bot detector (RABot), a multi-granularity graph-augmentation framework that addresses both issues in a unified manner. RABot employs a neighborhood-aware oversampling strategy that linearly interpolates minority-class embeddings within local subgraphs, thereby stabilizing the decision boundary under low-resource regimes. Concurrently, a reinforcement-learning-driven edge-filtering module combines similarity-based edge features with adaptive threshold optimization to excise spurious interactions during message passing, yielding a cleaner topology. Extensive experiments on three real-world benchmarks and four GNN backbones demonstrate that RABot consistently surpasses state-of-the-art baselines. In addition, since its augmentation and filtering modules are orthogonal to the underlying architecture, RABot can be seamlessly integrated into existing GNN pipelines to boost performance with minimal overhead.

AAAI Conference 2026 Conference Paper

TokenPowerBench: Benchmarking the Power Consumption of LLM Inference

  • Chenxu Niu
  • Wei Zhang
  • Jie Li
  • Yongjian Zhao
  • Tongyang Wang
  • Xi Wang
  • Yong Chen

Large language model (LLM) services now answer billions of queries per day, and industry reports show that inference, not training, accounts for more than 90% of total power consumption. However, existing benchmarks focus on either training/fine-tuning or performance of inference and provide little support for power consumption measurement and analysis of inference. We introduce TokenPowerBench, the first lightweight and extensible benchmark designed for LLM-inference power consumption studies. The benchmark combines a declarative configuration interface covering model choice, prompt set, and inference engine, a measurement layer that captures GPU-, node-, and system-level power without specialized power meters, and a phase-aligned metrics pipeline that attributes energy to the prefill and decode stages of every request. These elements make it straightforward to explore the power consumed by an LLM inference run; furthermore, by varying batch size, context length, parallelism strategy and quantization, users can quickly assess how each setting affects joules per token and other energy-efficiency metrics. We evaluate TokenPowerBench on four of the most widely used model series (Llama, Falcon, Qwen, and Mistral). Our experiments cover from 1 billion parameters up to the frontier-scale Llama3-405B model. Furthermore, we release TokenPowerBench as open source to help users to measure power consumption, forecast operating expenses, and meet sustainability targets when deploying LLM services.

AAAI Conference 2026 Conference Paper

Towards Robust Edge Model Adaptation via Elastic Architecture Search

  • Xianhang Chu
  • Xu Yang
  • Kun Wei
  • Xi Wang

Continual test-time adaptation (CTTA) enables online model adjustment under dynamic distribution shifts in real-world environments. However, most existing CTTA frameworks adopt fixed model architectures, lacking the structural flexibility required for deployment across heterogeneous edge devices with varying computational capacities. To address this, we propose an elastic framework for edge CTTA that performs resource-aware dynamic model search based on a pre-trained binary Supernet. This enables architectural flexibility by generating personalized models tailored to the resource constraints of different edge devices. Considering the evolving distribution of unlabeled data on edge devices during deployment, we introduce a pluggable lightweight fine-tuning mechanism. By inserting low-rank adapters into the frozen binary backbone, the model enables continual self-supervised adaptation with minimal computational overhead. In addition, we propose a structure-aware knowledge reflux mechanism that transfers the adaptation experience from fine-tuned edge models back into the Supernet. By distilling knowledge into structurally aligned Supernet paths, future architecture search is improved without requiring retraining. Experiments on multiple benchmarks validate that our method achieves state-of-the-art performance while significantly reducing resource consumption, with re-searched models after knowledge reflux showing further improvements.

EAAI Journal 2025 Journal Article

Convolutional neural network-attention-gate recurrent unit-attention hybrid framework for spindle thermal error modeling with joint feature analysis under complex variable speed conditions

  • Sen Mu
  • Guoqiang Fu
  • Yue Zheng
  • Xi Wang
  • Caijiang Lu
  • Jianzhong Fu

Deep learning-based spindle thermal error modeling and compensation methods can effectively enhance manufacturing precision. The complex variable working conditions and thermal hysteresis effects in machine tool machining bring significant challenges for high-precision thermal error modeling. To address this issue, a hybrid structure network based on the feature extraction capability of Convolutional Neural Network (CNN) and the thermal hysteresis effect resolution capability of deep Gate Recurrent Unit (GRU) is established. A dual-layer attention mechanism is introduced to enhance spatial features and temporal features for improving the model's robustness and accuracy. First, complex variable working conditions result in complexity and nonlinearity of data. CNN is employed to extract spatial features due to its powerful feature extraction capability. A self-attention mechanism is introduced after CNN block to further filter important features. Due to the influence of thermal hysteresis effects, a deep GRU block is established to extract temporal features. A channel attention mechanism is introduced as the final layer of the network to achieve feature selection across different temperature channels. Second, features extracted by two attention mechanism layers are visualized and analyzed using t-distributed stochastic neighbor embedding (t-SNE) algorithm to explain the effectiveness of the dual-layer attention mechanism structure. The probability density distribution of the predicted results is calculated by kernel density estimation. Model performance is analyzed from the perspective of data distribution. Finally, the proposed model is compared with advanced methods under complex working conditions on two machine tools. The effectiveness of the proposed model is further validated through actual cutting compensation.

AAAI Conference 2025 Conference Paper

DCTMamba: Advancing JPEG Image Restoration Through Long-Sequence Modeling and Adaptive Frequency Strategy

  • Xi Wang
  • Xueyang Fu
  • Liang Li
  • Zheng-Jun Zha

Despite the advanced long-sequence modeling of Mamba, which has expanded its applications in image restoration, there remains a lack of exploration combining its strengths with the specific characteristics of JPEG image restoration, where high-frequency components are lost after the Discrete Cosine Transform (DCT). To address this, we introduce DCTMamba, a new framework designed to apply Mamba more effectively to JPEG image restoration. Specifically, our method integrates the Discrete Cosine Transform (DCT) into the Mamba to establish the sequential scanning from lower to higher frequencies, enabling the network to initially reconstruct coarse structures and progressively refine the image with more intricate details. Furthermore, recognizing the variable frequency distributions that arise from DCT transformations across different image sizes, we have developed Scale-Adaptive Normalization to manage these variations adeptly. Comprehensive experiments confirm that DCTMamba significantly outperforms existing solutions, achieving high fidelity in both coarse structures and fine details.CTMamba significantly outperforms existing solutions, achieving high fidelity in both coarse structures and fine details.

NeurIPS Conference 2025 Conference Paper

Dual-Space Semantic Synergy Distillation for Continual Learning of Unlabeled Streams

  • Donghao Sun
  • Xi Wang
  • Xu Yang
  • Kun Wei
  • Cheng Deng

Continual learning from unlabeled data streams while effectively combating catastrophic forgetting poses an intractable challenge. Traditional methods predominantly rely on visual clustering techniques to generate pseudo labels, which are frequently plagued by problems such as noise and suboptimal quality, profoundly affecting the impact on the model evolution. To surmount these obstacles, we introduce an innovative approach that synergistically combines both visual and textual information to generate dual space hybrid pseudo labels for reliable model continual evolution. Specifically, by harnessing the capabilities of large multimodal models, we initially generate generalizable text descriptions for a few representative samples. These descriptions then undergo a `Coarse to Fine' refinement process to capture the subtle nuances between different data points, significantly enhancing the semantic accuracy of the descriptions. Simultaneously, a novel cross-modal hybrid approach seamlessly integrates these fine-grained textual descriptions with visual features, thereby creating a more robust and reliable supervisory signal. Finally, such descriptions are employed to alleviate the catastrophic forgetting issue via a semantic alignment distillation, which capitalizes on the stability inherent in language knowledge to effectively prevent the model from forgetting previously learned information. Comprehensive experiments conducted on a variety of benchmarks demonstrate that our proposed method attains state-of-the-art performance, and ablation studies further substantiate the effectiveness and superiority of the proposed method.

ICML Conference 2025 Conference Paper

How to set AdamW's weight decay as you scale model and dataset size

  • Xi Wang
  • Laurence Aitchison

The scaling of the optimal AdamW weight decay hyperparameter with model and dataset size is critical as we seek to build larger models, but is poorly understood. We show that weights learned by AdamW can be understood as an exponential moving average (EMA) of recent updates. This gives critical insights for how to set the weight decay in AdamW, and how the weight decay should scale with model and dataset size. In particular, the key hyperparameter for an exponential moving average is the EMA timescale. Intuitively, the EMA timescale can be understood as the number of recent iterations the EMA averages over. We find that the optimal timescale, measured in epochs, is roughly constant as we change model and dataset size. Moreover, given a learning rate, there is a one-to-one mapping from the EMA timescale to the weight decay hyperparameter. Thus, if the optimal EMA timescale is constant, that implies that as the dataset size increases, the optimal weight decay should fall and as the model size increases, the optimal weight decay should increase (if we follow the muP recommendation for scaling the learning rate). We validate these scaling rules on ResNet-18 and Vision Transformers trained on CIFAR-10 and ImageNet, and on NanoGPT pre-training on OpenWebText. Finally, we found that as training progresses, muP’s learning rate scaling breaks down for AdamW unless weight decay is scaled appropriately.

ICLR Conference 2025 Conference Paper

KBLaM: Knowledge Base augmented Language Model

  • Xi Wang
  • Taketomo Isazawa
  • Liana Mikaelyan
  • James Hensman

In this paper, we propose Knowledge Base augmented Language Model (KBLAM), a new method for augmenting Large Language Models (LLMs) with external knowledge. KBLAM works with a knowledge base (KB) constructed from a corpus of documents, transforming each piece of knowledge in the KB into continuous key-value vector pairs via pre-trained sentence encoders with linear adapters and integrating them into pre-trained LLMs via a specialized rectangular attention mechanism. Unlike Retrieval-Augmented Generation, KBLAM eliminates external retrieval modules, and unlike in-context learning, its computational overhead scales linearly with KB size rather than quadratically. Our approach enables integrating a large KB of more than 10K triples into an 8B pre-trained LLM of only 8K context window on one single A100 80GB GPU and allows for dynamic updates without model fine-tuning or retraining. Experiments demonstrate KBLAM’s effectiveness in various tasks, including question-answering and open-ended reasoning, while providing interpretable insights into its use of the augmented knowledge. Code and datasets are available at https://github.com/microsoft/KBLaM/

NeurIPS Conference 2025 Conference Paper

LangHOPS: Language Grounded Hierarchical Open-Vocabulary Part Segmentation

  • Yang Miao
  • Jan-Nico Zaech
  • Xi Wang
  • Fabien Despinoy
  • Danda Pani Paudel
  • Luc V Gool

We propose LangHOPS, the first Multimodal Large Language Model (MLLM)-based framework for open-vocabulary object–part instance segmentation. Given an image, LangHOPS can jointly detect and segment hierarchical object and part instances from open-vocabulary candidate categories. Unlike prior approaches that rely on heuristic or learnable visual grouping, our approach grounds object–part hierarchies in language space. It integrates the MLLM into the object-part parsing pipeline to leverage rich knowledge and reasoning capabilities, and link multi-granularity concepts within the hierarchies. We evaluate LangHOPS across multiple challenging scenarios, including in-domain and cross-dataset object-part instance segmentation, and zero-shot semantic segmentation. LangHOPS achieves state-of-the-art results, surpassing previous methods by 5. 5% Average Precision(AP) (in-domain) and 4. 8% (cross-dataset) on the PartImageNet dataset and by 2. 5% mIOU on unseen object parts in ADE20K (zero-shot). Ablation studies further validate the effectiveness of the language-grounded hierarchy and MLLM-driven part query refinement strategy.

JBHI Journal 2025 Journal Article

Multi-Scale Spatio-Temporal Transformer-Based Imbalanced Longitudinal Learning for Glaucoma Forecasting From Irregular Time Series Images

  • Xikai Yang
  • Jian Wu
  • Xi Wang
  • Yuchen Yuan
  • Jinpeng Li
  • Guangyong Chen
  • Ning Li Wang
  • Pheng-Ann Heng

Glaucoma is one of the major eye diseases that leads to progressive optic nerve fiber damage and irreversible blindness, afflicting millions of individuals. Glaucoma forecast is a good solution to early screening and intervention of potential patients, which is helpful to prevent further deterioration of the disease. It leverages a series of historical fundus images of an eye and forecasts the likelihood of glaucoma occurrence in the future. However, the irregular sampling nature and the imbalanced class distribution are two challenges in the development of disease forecasting approaches. To this end, we introduce the Multi-scale Spatio-temporal Transformer Network (MST-former) based on the transformer architecture tailored for sequential image inputs, which can effectively learn representative semantic information from sequential images on both temporal and spatial dimensions. Specifically, we employ a multi-scale structure to extract features at various resolutions, which can largely exploit rich spatial information encoded in each image. Besides, we design a time distance matrix to scale time attention in a non-linear manner, which could effectively deal with the irregularly sampled data. Furthermore, we introduce a temperature-controlled Balanced Softmax Cross-entropy loss to address the class imbalance issue. Extensive experiments on the Sequential fundus Images for Glaucoma Forecast (SIGF) dataset demonstrate the superiority of the proposed MST-former method, achieving an AUC of 96. 6% for glaucoma forecasting. Besides, our method shows excellent generalization capability on the Alzheimer's Disease Neuroimaging Initiative (ADNI) MRI dataset, with an accuracy of 88. 2% for mild cognitive impairment and Alzheimer's disease prediction, outperforming the compared method by a large margin. A series of ablation studies further verify the contribution of our proposed components in addressing the irregular sampled and class imbalanced problems.

IJCAI Conference 2025 Conference Paper

OMS: One More Step Noise Searching to Enhance Membership Inference Attacks for Diffusion Models

  • Xiaomeng Fu
  • Xi Wang
  • Qiao Li
  • Jin Liu
  • Jiao Dai
  • Jizhong Han
  • Xingyu Gao

The data-intensive nature of Diffusion models amplifies the risks of privacy infringements and copyright disputes, particularly when training on extensive unauthorized data scraped from the Internet. Membership Inference Attacks (MIA) aim to determine whether a data sample has been utilized by the target model during training, thereby serving as a pivotal tool for privacy preservation. Current MIA employs the prediction loss to distinguish between training member samples and non-members. These methods assume that, compared to non-members, members, having been encountered by the model during training result in a smaller prediction loss. However, this assumption proves ineffective in diffusion models due to the random noise sampled during the training process. Rather than estimating the loss, our approach examines this random noise and reformulate the MIA as a noise search problem, assuming that members are more feasible to find the noise used in the training process. We formulate this noise search process as an optimization problem and employ the fixed-point iteration to solve it. We analyze current MIA methods through the lens of the noise search framework and reveal that they rely on the first residual as the discriminative metric to differentiate members and non-members. Inspired by this observation, we introduce OMS, which augments existing MIA methods by iterating One More fixed-point Step to include a further residual, i. e. , the second residual. We integrate our method into various MIA methods across different diffusion models. The experimental results validate the efficacy of our proposed approach.

NeurIPS Conference 2025 Conference Paper

Scale-invariant attention

  • Ben Anson
  • Xi Wang
  • Laurence Aitchison

One persistent challenge in LLM research is the development of attention mechanisms that are able to generalise from training on shorter contexts to inference on longer contexts. We propose two conditions that we expect all effective long-context attention mechanisms to have: scale-invariant total attention, and scale-invariant attention sparsity. Under a Gaussian assumption, we show that a simple position-dependent transformation of the attention logits is sufficient for these conditions to hold. Experimentally we find that the resulting scale-invariant attention scheme gives considerable benefits in terms of validation loss when zero-shot generalising from training on short contexts to validation on longer contexts, and is effective at long-context retrieval.

NeurIPS Conference 2025 Conference Paper

StateSpaceDiffuser: Bringing Long Context to Diffusion World Models

  • Nedko Savov
  • Naser Kazemi
  • Deheng Zhang
  • Danda Pani Paudel
  • Xi Wang
  • Luc V Gool

World models have recently gained prominence for action-conditioned visual prediction in complex environments. However, relying on only a few recent observations causes them to lose long-term context. Consequently, within a few steps, the generated scenes drift from what was previously observed, undermining temporal coherence. This limitation, common in state-of-the-art world models, which are diffusion-based, stems from the lack of a lasting environment state. To address this problem, we introduce StateSpaceDiffuser, where a diffusion model is enabled to perform long-context tasks by integrating features from a state-space model, representing the entire interaction history. This design restores long-term memory while preserving the high-fidelity synthesis of diffusion models. To rigorously measure temporal consistency, we develop an evaluation protocol that probes a model’s ability to reinstantiate seen content in extended rollouts. Comprehensive experiments show that StateSpaceDiffuser significantly outperforms a strong diffusion-only baseline, maintaining a coherent visual context for an order of magnitude more steps. It delivers consistent views in both a 2D maze navigation and a complex 3D environment. These results establish that bringing state-space representations into diffusion models is highly effective in demonstrating both visual details and long-term memory. Project page: https: //insait-institute. github. io/StateSpaceDiffuser/

AAAI Conference 2025 Conference Paper

SWEA: Updating Factual Knowledge in Large Language Models via Subject Word Embedding Altering

  • Xiaopeng Li
  • Shasha Li
  • Shezheng Song
  • Huijun Liu
  • Bin Ji
  • Xi Wang
  • Jun Ma
  • Jie Yu

The general capabilities of large language models (LLMs) make them the infrastructure for various AI applications, but updating their inner knowledge requires significant resources. Recent model editing is a promising technique for efficiently updating a small amount of knowledge of LLMs and has attracted much attention. In particular, local editing methods, which directly update model parameters, are proven suitable for updating small amounts of knowledge. Local editing methods update weights by computing least squares closed-form solutions and identify edited knowledge by vector-level matching in inference, which achieve promising results. However, these methods still require a lot of time and resources to complete the computation. Moreover, vector-level matching lacks reliability, and such updates disrupt the original organization of the model's parameters. To address these issues, we propose a detachable and expandable Subject Word Embedding Altering (SWEA) framework, which finds the editing embeddings through token-level matching and adds them to the subject word embeddings in Transformer input. To get these editing embeddings, we propose optimizing then suppressing fusion method, which first optimizes learnable embedding vectors for the editing target and then suppresses the Knowledge Embedding Dimensions (KEDs) to obtain final editing embeddings. We thus propose SWEAOS method for editing factual knowledge in LLMs. We demonstrate the overall state-of-the-art (SOTA) performance of SWEAOS on the CounterFact and zsRE datasets. To further validate the reasoning ability of SWEAOS in editing knowledge, we evaluate it on the more complex RippleEdits benchmark. The results demonstrate that SWEAOS possesses SOTA reasoning ability.

TMLR Journal 2024 Journal Article

Analysis of Classifier-Free Guidance Weight Schedulers

  • Xi Wang
  • Nicolas Dufour
  • Nefeli Andreou
  • Marie-Paule Cani
  • Victoria Fernandez Abrevaya
  • David Picard
  • Vicky Kalogeiton

Classifier-Free Guidance (CFG) enhances the quality and condition adherence of text-to-image diffusion models. It operates by combining the conditional and unconditional predictions using a fixed weight. However, recent works vary the weights throughout the diffusion process, reporting superior results but without providing any rationale or analysis. By conducting comprehensive experiments, this paper provides insights into CFG weight schedulers. Our findings suggest that simple, monotonically increasing weight schedulers consistently lead to improved performances, requiring merely a single line of code. In addition, more complex parametrized schedulers can be optimized for further improvement, but do not generalize across different models and tasks.

ICLR Conference 2024 Conference Paper

Bayesian Low-rank Adaptation for Large Language Models

  • Adam X. Yang
  • Maxime Robeyns
  • Xi Wang
  • Laurence Aitchison

Parameter-efficient fine-tuning (PEFT) has emerged as a new paradigm for cost-efficient fine-tuning of large language models (LLMs), with low-rank adaptation (LoRA) being a widely adopted choice. However, fine-tuned LLMs often become overconfident especially when fine-tuned on small datasets. Bayesian methods, with their inherent ability to estimate uncertainty, serve as potent tools to mitigate overconfidence and enhance calibration. In this work, we introduce Laplace-LoRA, a straightforward yet effective Bayesian method, which applies the Laplace approximation to the LoRA parameters and, considerably boosts the calibration of fine-tuned LLMs.

AAAI Conference 2024 Conference Paper

NeRFail: Neural Radiance Fields-Based Multiview Adversarial Attack

  • Wenxiang Jiang
  • Hanwei Zhang
  • Xi Wang
  • Zhongwen Guo
  • Hao Wang

Adversarial attacks, i.e., generating adversarial perturbations with a small magnitude to deceive deep neural networks, are important for investigating and improving model trustworthiness. Traditionally, the topic was scoped within 2D images without considering 3D multiview information. Benefiting from Neural Radiance Fields (NeRF), one can easily reconstruct a 3D scene with a Multi-Layer Perceptron (MLP) from given 2D views and synthesize photo-realistic renderings of novel vantages. This opens up a door to discussing the possibility of undertaking to attack multiview NeRF network with downstream tasks from different rendering angles, which we denote Neural Radiance Fiels-based multiview adversarial Attack (NeRFail). The goal is, given one scene and a subset of views, to deceive the recognition results of agnostic view angles as well as given views. To do so, we propose a transformation mapping from pixels to 3D points such that our attack generates multiview adversarial perturbations by attacking a subset of images with different views, intending to prevent the downstream classifier from correctly predicting images rendered by NeRF from other views. Experiments show that our multiview adversarial perturbations successfully obfuscate the downstream classifier at both known and unknown views. Notably, when retraining another NeRF on the perturbed training data, we show that the perturbation can be inherited and reproduced. The code can be found at https://github.com/jiang-wenxiang/NeRFail.

YNIMG Journal 2024 Journal Article

Neurophysiological dynamics of metacontrol states: EEG insights into conflict regulation

  • Xi Wang
  • Nasibeh Talebi
  • Xianzhen Zhou
  • Bernhard Hommel
  • Christian Beste

Understanding the neural mechanisms underlying metacontrol and conflict regulation is crucial for insights into cognitive flexibility and persistence. This study employed electroencephalography (EEG), EEG-beamforming and directed connectivity analyses to explore how varying metacontrol states influence conflict regulation at a neurophysiological level. Metacontrol states were manipulated by altering the frequency of congruent and incongruent trials across experimental blocks in a modified flanker task, and both behavioral and electrophysiological measures were analyzed. Behavioral data confirmed the experimental manipulation's efficacy, showing an increase in persistence bias and a reduction in flexibility bias during increased conflict regulation. Electrophysiologically, theta band activity paralleled the behavioral data, suggesting that theta oscillations reflect the mismatch between expected metacontrol bias and actual task demands. Alpha and beta band dynamics differed across experimental blocks, though these changes did not directly mirror behavioral effects. Post-response alpha and beta activity were more pronounced in persistence-biased states, indicating a neural reset mechanism preparing for future cognitive demands. By using a novel artificial neural networks method, directed connectivity analyses revealed enhanced inter-regional communication during persistence states, suggesting stronger top-down control and sensorimotor integration. Overall, theta band activity was closely tied to metacontrol processes, while alpha and beta bands played a role in resetting the neural system for upcoming tasks. These findings provide a deeper understanding of the neural substrates involved in metacontrol and conflict monitoring, emphasizing the distinct roles of different frequency bands in these cognitive processes.

EAAI Journal 2024 Journal Article

Radial basis function neural networks for optimal control with model reduction and transfer learning

  • Anni Zhao
  • Siyuan Xing
  • Xi Wang
  • Jian-Qiao Sun

This paper proposes a method to compute the solutions of linear optimal control expressed in terms of the radial basis function neural networks with Gaussian activation functions for multi-degree-of-freedom dynamic systems. Hamilton–Jacobi–Bellman equation is adopted to formulate the optimal control problem. The radial basis function neural networks are proposed to approximate the value function to solve the Hamilton–Jacobi–Bellman equation with a policy iteration algorithm. A dominant stabilizing control is proposed as an initial control to start the policy iteration that guarantees the convergence of the iteration, particularly for open-loop unstable dynamic systems. The balanced truncation technique is applied to the multi-degree-of-freedom dynamic system to reduce the dimension of the original system, which provides multiple advantages when applying the radial basis function neural networks to unstable dynamic systems in a relatively high dimensional state space. Transfer learning is also adopted to update radial basis function neural networks with experimental data, which results in further control performance improvement. Numerical simulations and experimental studies show that the radial basis function neural networks not only find accurate optimal controls for linear systems, but also offer excellent performance in trajectory tracking and stabilization applications.

JBHI Journal 2024 Journal Article

Single-Cell Heterogeneity-Aware Transformer-Guided Multiple Instance Learning for Cancer Aneuploidy Prediction From Whole Slide Histopathology Images

  • Feiyang Yu
  • Xi Wang
  • Rasoul Sali
  • Ruijiang Li

Aneuploidy is a hallmark of aggressive malignancies associated with therapeutic resistance and poor survival. Measuring aneuploidy requires expensive specialized techniques that are not clinically applicable. Deep learning analysis of routine histopathology slides has revealed associations with genetic mutations. However, existing studies focus on image tiles, and there is no prior work that predicts aneuploidy using single-cell analysis. Here, we present a single-cell heterogeneity-aware and transformer-guided deep learning framework to predict aneuploidy from whole slide histopathology images. First, we perform nuclei segmentation and classification to obtain individual cancer cells, which are clustered into multiple subtypes. The cell subtype distributions are computed to measure cancer cell heterogeneity. Additionally, morphological features of different cell subtypes are extracted. Further, we leverage a multiple instance learning module with Transformer, which encourages the network to focus on the most informative cancer cells. Lastly, a hybrid network is built to unify cell heterogeneity, morphology, and deep features for aneuploidy prediction. We train and validate our method on two public datasets from TCGA: lung adenocarcinoma (LUAD) and head and neck squamous cell carcinoma (HNSC), with 339 and 245 patients. Our model achieves promising performance with AUC of 0. 818 (95% CI: 0. 718–0. 919) and 0. 827 (95% CI: 0. 704–0. 949) on the LUAD and HNSC test sets, respectively. Through extensive ablation and comparison studies, we demonstrate the effectiveness of each component of the model and superior performance over alternative networks. In conclusion, we present a novel deep learning approach to predict aneuploidy from histopathology images, which could inform personalized cancer treatment.

NeurIPS Conference 2024 Conference Paper

SpreadsheetBench: Towards Challenging Real World Spreadsheet Manipulation

  • Zeyao Ma
  • Bohan Zhang
  • Jing Zhang
  • Jifan Yu
  • Xiaokang Zhang
  • Xiaohan Zhang
  • Sijia Luo
  • Xi Wang

We introduce SpreadsheetBench, a challenging spreadsheet manipulation benchmark exclusively derived from real-world scenarios, designed to immerse current large language models (LLMs) in the actual workflow of spreadsheet users. Unlike existing benchmarks that rely on synthesized queries and simplified spreadsheet files, SpreadsheetBench is built from 912 real questions gathered from online Excel forums, which reflect the intricate needs of users. The associated spreadsheets from the forums contain a variety of tabular data such as multiple tables, non-standard relational tables, and abundant non-textual elements. Furthermore, we propose a more reliable evaluation metric akin to online judge platforms, where multiple spreadsheet files are created as test cases for each instruction, ensuring the evaluation of robust solutions capable of handling spreadsheets with varying values. Our comprehensive evaluation of various LLMs under both single-round and multi-round inference settings reveals a substantial gap between the state-of-the-art (SOTA) models and human performance, highlighting the benchmark's difficulty.

JBHI Journal 2023 Journal Article

A Deep Learning Model for Automatic Segmentation of Intraparenchymal and Intraventricular Hemorrhage for Catheter Puncture Path Planning

  • Guoyu Tong
  • Xi Wang
  • Huiyan Jiang
  • Anhua Wu
  • Wen Cheng
  • Xiao Cui
  • Long Bao
  • Ruikai Cai

Intracerebral hemorrhage is the subtype of stroke with the highest mortality rate, especially when it also causes secondary intraventricular hemorrhage. The optimal surgical option for intracerebral hemorrhage remains one of the most controversial areas of neurosurgery. We aim to develop a deep learning model for the automatic segmentation of intraparenchymal and intraventricular hemorrhage for clinical catheter puncture path planning. First, we develop a 3D U-Net embedded with a multi-scale boundary aware module and a consistency loss for segmenting two types of hematoma in computed tomography images. The multi-scale boundary aware module can improve the model's ability to understand the two types of hematoma boundaries. The consistency loss can reduce the probability of classifying a pixel into two categories at the same time. Since different hematoma volumes and locations have different treatments. We also measure hematoma volume, estimate centroid deviation, and compare with clinical methods. Finally, we plan the puncture path and conduct clinical validation. We collected a total of 351 cases, and the test set contained 103 cases. For intraparenchymal hematomas, the accuracy can reach 96 $ \% $ when the proposed method is applied for path planning. For intraventricular hematomas, the proposed model's segmentation efficiency and centroid prediction are superior to other comparable models. Experimental results and clinical practice show that the proposed model has potential for clinical application. In addition, our proposed method has no complicated modules and improves efficiency, with generalization ability.

JMLR Journal 2023 Journal Article

Online Optimization over Riemannian Manifolds

  • Xi Wang
  • Zhipeng Tu
  • Yiguang Hong
  • Yingyi Wu
  • Guodong Shi

Online optimization has witnessed a massive surge of research attention in recent years. In this paper, we propose online gradient descent and online bandit algorithms over Riemannian manifolds in full information and bandit feedback settings respectively, for both geodesically convex and strongly geodesically convex functions. We establish a series of upper bounds on the regrets for the proposed algorithms over Hadamard manifolds. We also find a universal lower bound for achievable regret on Hadamard manifolds. Our analysis shows how time horizon, dimension, and sectional curvature bounds have impact on the regret bounds. When the manifold permits positive sectional curvature, we prove similar regret bound can be established by handling non-constrictive project maps. In addition, numerical studies on problems defined on symmetric positive definite matrix manifold, hyperbolic spaces, and Grassmann manifolds are provided to validate our theoretical findings, using synthetic and real-world data. [abs] [ pdf ][ bib ] [ code ] &copy JMLR 2023. ( edit, beta )

ICLR Conference 2023 Conference Paper

Particle-based Variational Inference with Preconditioned Functional Gradient Flow

  • Hanze Dong
  • Xi Wang
  • Yong Lin
  • Tong Zhang 0001

Particle-based variational inference (VI) minimizes the KL divergence between model samples and the target posterior with gradient flow estimates. With the popularity of Stein variational gradient descent (SVGD), the focus of particle-based VI algorithms has been on the properties of functions in Reproducing Kernel Hilbert Space (RKHS) to approximate the gradient flow. However, the requirement of RKHS restricts the function class and algorithmic flexibility. This paper offers a general solution to this problem by introducing a functional regularization term that encompasses the RKHS norm as a special case. This allows us to propose a new particle-based VI algorithm called preconditioned functional gradient flow (PFG). Compared to SVGD, PFG has several advantages. It has a larger function class, improved scalability in large particle-size scenarios, better adaptation to ill-conditioned distributions, and provable continuous-time convergence in KL divergence. Additionally, non-linear function classes such as neural networks can be incorporated to estimate the gradient flow. Our theory and experiments demonstrate the effectiveness of the proposed framework.

ICLR Conference 2023 Conference Paper

Robustness to corruption in pre-trained Bayesian neural networks

  • Xi Wang
  • Laurence Aitchison

We develop ShiftMatch, a new training-data-dependent likelihood for robustness to corruption in Bayesian neural networks (BNNs). ShiftMatch is inspired by the training-data-dependent “EmpCov” priors from Izmailov et al. (2021a), and efficiently matches test-time spatial correlations to those at training time. Critically, ShiftMatch is designed to leave the neural network’s training time likelihood unchanged, allowing it to use publicly available samples from pre-trained BNNs. Using pre-trained HMC samples, ShiftMatch gives strong performance improvements on CIFAR-10-C, outperforms EmpCov priors (though ShiftMatch uses extra information from a minibatch of corrupted test points), and is perhaps the first Bayesian method capable of convincingly outperforming plain deep ensembles.

AAAI Conference 2023 Conference Paper

SpatialFormer: Semantic and Target Aware Attentions for Few-Shot Learning

  • Jinxiang Lai
  • Siqian Yang
  • Wenlong Wu
  • Tao Wu
  • Guannan Jiang
  • Xi Wang
  • Jun Liu
  • Bin-Bin Gao

Recent Few-Shot Learning (FSL) methods put emphasis on generating a discriminative embedding features to precisely measure the similarity between support and query sets. Current CNN-based cross-attention approaches generate discriminative representations via enhancing the mutually semantic similar regions of support and query pairs. However, it suffers from two problems: CNN structure produces inaccurate attention map based on local features, and mutually similar backgrounds cause distraction. To alleviate these problems, we design a novel SpatialFormer structure to generate more accurate attention regions based on global features. Different from the traditional Transformer modeling intrinsic instance-level similarity which causes accuracy degradation in FSL, our SpatialFormer explores the semantic-level similarity between pair inputs to boost the performance. Then we derive two specific attention modules, named SpatialFormer Semantic Attention (SFSA) and SpatialFormer Target Attention (SFTA), to enhance the target object regions while reduce the background distraction. Particularly, SFSA highlights the regions with same semantic information between pair features, and SFTA finds potential foreground object regions of novel feature that are similar to base categories. Extensive experiments show that our methods are effective and achieve new state-of-the-art results on few-shot classification benchmarks.

AAAI Conference 2023 System Paper

Task2KB: A Public Task-Oriented Knowledge Base

  • Procheta Sen
  • Xi Wang
  • Ruiqing Xu
  • Emine Yilmaz

Search engines and conversational assistants are commonly used to help users complete their every day tasks such as booking travel, cooking, etc. While there are some existing datasets that can be used for this purpose, their coverage is limited to very few domains. In this paper, we propose a novel knowledge base, ‘Task2KB’, which is constructed using data crawled from WikiHow, an online knowledge resource offering instructional articles on a wide range of tasks. Task2KB encapsulates various types of task-related information and attributes, such as requirements, detailed step description, and available methods to complete tasks. Due to its higher coverage compared to existing related knowledge graphs, Task2KB can be highly useful in the development of general purpose task completion assistants.

NeurIPS Conference 2022 Conference Paper

Decoupling Classifier for Boosting Few-shot Object Detection and Instance Segmentation

  • Bin-Bin Gao
  • Xiaochen Chen
  • Zhongyi Huang
  • Congchong Nie
  • Jun Liu
  • Jinxiang Lai
  • Guannan Jiang
  • Xi Wang

This paper focus on few-shot object detection~(FSOD) and instance segmentation~(FSIS), which requires a model to quickly adapt to novel classes with a few labeled instances. The existing methods severely suffer from bias classification because of the missing label issue which naturally exists in an instance-level few-shot scenario and is first formally proposed by us. Our analysis suggests that the standard classification head of most FSOD or FSIS models needs to be decoupled to mitigate the bias classification. Therefore, we propose an embarrassingly simple but effective method that decouples the standard classifier into two heads. Then, these two individual heads are capable of independently addressing clear positive samples and noisy negative samples which are caused by the missing label. In this way, the model can effectively learn novel classes while mitigating the effects of noisy negative samples. Without bells and whistles, our model without any additional computation cost and parameters consistently outperforms its baseline and state-of-the-art by a large margin on PASCAL VOC and MS-COCO benchmarks for FSOD and FSIS tasks. \footnote{\url{https: //csgaobb. github. io/Projects/DCFS}. }

NeurIPS Conference 2022 Conference Paper

Distributed Online Convex Optimization with Compressed Communication

  • Zhipeng Tu
  • Xi Wang
  • Yiguang Hong
  • Lei Wang
  • Deming Yuan
  • Guodong Shi

We consider a distributed online convex optimization problem when streaming data are distributed among computing agents over a connected communication network. Since the data are high-dimensional or the network is large-scale, communication load can be a bottleneck for the efficiency of distributed algorithms. To tackle this bottleneck, we apply the state-of-art data compression scheme to the fundamental GD-based distributed online algorithms. Three algorithms with difference-compressed communication are proposed for full information feedback (DC-DOGD), one-point bandit feedback (DC-DOBD), and two-point bandit feedback (DC-DO2BD), respectively. We obtain regret bounds explicitly in terms of time horizon, compression ratio, decision dimension, agent number, and network parameters. Our algorithms are proved to be no-regret and match the same regret bounds, w. r. t. time horizon, with their uncompressed versions for both convex and strongly convex losses. Numerical experiments are given to validate the theoretical findings and illustrate that the proposed algorithms can effectively reduce the total transmitted bits for distributed online training compared with the uncompressed baseline.

EAAI Journal 2022 Journal Article

Dynamic speed trajectory generation and tracking control for autonomous driving of intelligent high-speed trains combining with deep learning and backstepping control methods

  • Xi Wang
  • Shukai Li
  • Yuan Cao
  • Tianpeng Xin
  • Lixing Yang

The development of autonomous transportation systems has received increasing attention over the last decades. Different from existing automatic train control systems, the decision-making capability in the autonomous driving system enables a train to adapt to the complicated and dynamic circumstances. This paper in particular focuses on the decision-making problem for the autonomous driving of intelligent high-speed trains, and proposes a novel decision-making framework by combining the deep learning and backstepping control methods. By exploiting the deep learning methods, a speed trajectory generator is trained with the actual driving data, and dynamically calculates the reference speed trajectory according to the real-time driving condition. Then, a backstepping controller is designed to track the reference speed trajectory such that the separation, cohesion and alignment requirements for the autonomous driving of high-speed trains are achieved. Simulation experiments are implemented to illustrate the effectiveness of our methods.

NeurIPS Conference 2021 Conference Paper

No-regret Online Learning over Riemannian Manifolds

  • Xi Wang
  • Zhipeng Tu
  • Yiguang Hong
  • Yingyi Wu
  • Guodong Shi

We consider online optimization over Riemannian manifolds, where a learner attempts to minimize a sequence of time-varying loss functions defined on Riemannian manifolds. Though many Euclidean online convex optimization algorithms have been proven useful in a wide range of areas, less attention has been paid to their Riemannian counterparts. In this paper, we study Riemannian online gradient descent (R-OGD) on Hadamard manifolds for both geodesically convex and strongly geodesically convex loss functions, and Riemannian bandit algorithm (R-BAN) on Hadamard homogeneous manifolds for geodesically convex functions. We establish upper bounds on the regrets of the problem with respect to time horizon, manifold curvature, and manifold dimension. We also find a universal lower bound for the achievable regret by constructing an online convex optimization problem on Hadamard manifolds. All the obtained regret bounds match the corresponding results are provided in Euclidean spaces. Finally, some numerical experiments validate our theoretical results.

ICRA Conference 2021 Conference Paper

TT-SLAM: Dense Monocular SLAM for Planar Environments

  • Xi Wang
  • Marc Christie
  • Éric Marchand

This paper proposes a novel visual SLAM method with dense planar reconstruction using a monocular camera: TT-SLAM. The method exploits planar template-based trackers (TT) to compute camera poses and reconstructs a multi-planar scene representation. Multiple homographies are estimated simultaneously by clustering a set of template trackers supported by superpixelized regions. Compared to RANSAC-based multiple homographies method [1], data association and keyframe selection issues are handled by the continuous nature of template trackers. A non-linear optimization process is applied to all the homographies to improve the precision in pose estimation. Experiments show that the proposed method outperforms RANSAC-based multiple homographies method [1] as well as other dense method SLAM techniques such as LSD-SLAM or DPPTAM, and competes with keypoint-based techniques like ORB-SLAM while providing dense planar reconstructions of the environment.

IROS Conference 2020 Conference Paper

Relative Pose Estimation and Planar Reconstruction via Superpixel-Driven Multiple Homographies

  • Xi Wang
  • Marc Christie
  • Éric Marchand

This paper proposes a novel method to simultaneously perform relative camera pose estimation and planar reconstruction of a scene from two RGB images. We start by extracting and matching superpixel information from both images and rely on a novel multi-model RANSAC approach to estimate multiple homographies from superpixels and identify matching planes. Ambiguity issues when performing homography decomposition are handled by proposing a voting system to more reliably estimate relative camera pose and plane parameters. A non-linear optimization process is also proposed to perform bundle adjustment that exploits a joint representation of homographies and works both for image pairs and whole sequences of image (vSLAM). As a result, the approach provides a mean to perform a dense 3D plane reconstruction from two RGB images only without relying on RGB-D inputs or strong priors such as Manhattan assumptions, and can be extented to handle sequences of images. Our results compete with keypointbased techniques such as ORB-SLAM while providing a dense representation and are more precise than direct and semi-direct pose estimation techniques used in LSD-SLAM or DPPTAM.

UAI Conference 2020 Conference Paper

Relaxed Multivariate Bernoulli Distribution and Its Applications to Deep Generative Models

  • Xi Wang
  • Junming Yin

Recent advances in variational auto-encoder (VAE) have demonstrated the possibility of approximating the intractable posterior distribution with a variational distribution parameterized by a neural network. To optimize the variational objective of VAE, the reparameterization trick is commonly applied to obtain a low-variance estimator of the gradient. The main idea of the trick is to express the variational distribution as a differentiable function of parameters and a random variable with a fixed distribution. To extend the reparameterization trick to inference involving discrete latent variables, a common approach is to use a continuous relaxation of the categorical distribution as the approximate posterior. However, when applying continuous relaxation to the multivariate cases, multiple variables are typically assumed to be independent, making it suboptimal in applications where modeling dependency is crucial to the overall performance. In this work, we propose a multivariate generalization of the Relaxed Bernoulli distribution, which can be reparameterized and can capture the correlation between variables via a Gaussian copula. We demonstrate its effectiveness in two tasks: density estimation with Bernoulli VAE and semi-supervised multi-label classification.

JBHI Journal 2020 Journal Article

UD-MIL: Uncertainty-Driven Deep Multiple Instance Learning for OCT Image Classification

  • Xi Wang
  • Fangyao Tang
  • Hao Chen
  • Luyang Luo
  • Ziqi Tang
  • An-Ran Ran
  • Carol Y. Cheung
  • Pheng-Ann Heng

Deep learning has achieved remarkable success in the optical coherence tomography (OCT) image classification task with substantial labelled B-scan images available. However, obtaining such fine-grained expert annotations is usually quite difficult and expensive. How to leverage the volume-level labels to develop a robust classifier is very appealing. In this paper, we propose a weakly supervised deep learning framework with uncertainty estimation to address the macula-related disease classification problem from OCT images with the only volume-level label being available. First, a convolutional neural network (CNN) based instance-level classifier is iteratively refined by using the proposed uncertainty-driven deep multiple instance learning scheme. To our best knowledge, we are the first to incorporate the uncertainty evaluation mechanism into multiple instance learning (MIL) for training a robust instance classifier. The classifier is able to detect suspicious abnormal instances and abstract the corresponding deep embedding with high representation capability simultaneously. Second, a recurrent neural network (RNN) takes instance features from the same bag as input and generates the final bag-level prediction by considering the individually local instance information and globally aggregated bag-level representation. For more comprehensive validation, we built two large diabetic macular edema (DME) OCT datasets from different devices and imaging protocols to evaluate the efficacy of our method, which are composed of 30, 151 B-scans in 1, 396 volumes from 274 patients (Heidelberg-DME dataset) and 38, 976 B-scans in 3, 248 volumes from 490 patients (Triton-DME dataset), respectively. We compare the proposed method with the state-of-the-art approaches, and experimentally demonstrate that our method is superior to alternative methods, achieving volume-level accuracy, F1-score and area under the receiver operating characteristic curve (AUC) of 95. 1%, 0. 939 and 0. 990 on Heidelberg-DME and those of 95. 1%, 0. 935 and 0. 986 on Triton-DME, respectively. Furthermore, the proposed method also yields competitive results on another public age-related macular degeneration OCT dataset, indicating the high potential as an effective screening tool in the clinical practice.

TCS Journal 2019 Journal Article

How many triangles and quadrilaterals are there in an n-dimensional augmented cube?

  • Qiang Dong
  • Xi Wang

The augmented cube is an important variant of hypercube as interconnection topology of parallel computing. In this paper, we examine the numbers of short cycles in augmented cubes, and prove that for n ≥ 3, there are 2 n × ( n − 1 ) triangles and 2 n − 2 × ( 2 n 2 + 5 n − 11 ) quadrilaterals in an n-dimensional augmented cube. This result shows that augmented cubes are promising interconnection networks with superior connectivity and fault-tolerant capability.

IROS Conference 2018 Conference Paper

Optimized Contrast Enhancements to Improve Robustness of Visual Tracking in a SLAM Relocalisation Context

  • Xi Wang
  • Marc Christie
  • Éric Marchand

Robustness of indirect SLAM techniques to light changing conditions remains a central issue in the robotics community. With the change in the illumination of a scene, feature points are either not extracted properly due to low contrasts, or not matched due to large differences in descriptors. In this paper, we propose a multi-layered image representation (MLI) in which each layer holds a contrast enhanced version of the current image in the tracking process in order to improve detection and matching. We show how Mutual Information can be used to compute dynamic contrast enhancements on each layer. We demonstrate how this approach dramatically improves the robustness in dynamic light changing conditions on both synthetic and real environments compared to default ORB-SLAM. This work focalises on the specific case of SLAM relocalisation in which a first pass on a reference video constructs a map, and a second pass with a light changed condition relocalizes the camera in the map.

TCS Journal 2016 Journal Article

An efficient algorithm to construct disjoint path covers of DCell networks

  • Xi Wang
  • Jianxi Fan
  • Xiaohua Jia
  • Cheng-Kuan Lin

Data center networks have been becoming more and more important with the development of cloud computing. For any two integers k ≥ 0 and n ≥ 2, the k-dimensional DCell with n-port switches, D k, n, has been proposed for one of the most important data center networks as a server centric data center network structure. D k, n can support millions of servers with outstanding network capacity and provide good fault tolerance by only using commodity switches. A disjoint path cover has significant applications in data center networks. In this paper, we prove that D k, n is one-to-one r-disjoint path coverable for any integer 1 ≤ r ≤ n + k − 1, except for D 1, 2. Moreover, we propose an O ( t k ) algorithm for finding a one-to-one r-disjoint path cover in D k, n for any integer 1 ≤ r ≤ n + k − 1, where t k is the number of servers in D k, n.

AAAI Conference 2016 Conference Paper

Mitosis Detection in Breast Cancer Histology Images via Deep Cascaded Networks

  • Hao Chen
  • Qi Dou
  • Xi Wang
  • Jing Qin
  • Pheng Heng

The number of mitoses per tissue area gives an important aggressiveness indication of the invasive breast carcinoma. However, automatic mitosis detection in histology images remains a challenging problem. Traditional methods either employ hand-crafted features to discriminate mitoses from other cells or construct a pixel-wise classifier to label every pixel in a sliding window way. While the former suffers from the large shape variation of mitoses and the existence of many mimics with similar appearance, the slow speed of the later prohibits its use in clinical practice. In order to overcome these shortcomings, we propose a fast and accurate method to detect mitosis by designing a novel deep cascaded convolutional neural network, which is composed of two components. First, by leveraging the fully convolutional neural network, we propose a coarse retrieval model to identify and locate the candidates of mitosis while preserving a high sensitivity. Based on these candidates, a fine discrimination model utilizing knowledge transferred from cross-domain is developed to further single out mitoses from hard mimics. Our approach outperformed other methods by a large margin in 2014 ICPR MITOS-ATYPIA challenge in terms of detection accuracy. When compared with the state-of-the-art methods on the 2012 ICPR MITOSIS data (a smaller and less challenging dataset), our method achieved comparable or better results with a roughly 60 times faster speed.

IS Journal 2014 Journal Article

User Recommendations in Reciprocal and Bipartite Social Networks--An Online Dating Case Study

  • Kang Zhao
  • Xi Wang
  • Mo Yu
  • Bo Gao

Many social networks in our daily life are bipartite networks built on reciprocity. How can we make recommendations to others so that the user is interested in and attractive to those other users whom we've recommended? We propose a new collaborative-filtering model to improve user recommendations in bipartite and reciprocal social networks. The model considers a user's taste in picking others and attractiveness in being picked by others. A case study of an online dating network shows that the approach offers good performance in recommending both initial and reciprocal contacts.

ICRA Conference 1989 Conference Paper

Proving the uniform boundedness of some commonly used control schemes for robots

  • Xi Wang
  • Lung-kee Chen

The uniform boundedness of the errors of some commonly used control schemes is considered analytically. Both fixed-point and trajectory control with proportional/derivative feedback are shown. The uniform boundedness of both the feedforward dynamic compensation method and the reduced feedforward compensation method are shown. A simple control scheme is given. An application to walking machines is presented. A technique of choosing Lyapunov functions is proposed to solve these problems. >

v2026.09.13