Arrow Research search

Author name cluster

Xinyu Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

32 papers
2 author rows

Possible papers

32

EAAI Journal 2026 Journal Article

A rapid image-based detection method for coalmine dust concentration and mass dispersion via multi-task deep learning

  • Haoran Fu
  • Xiaoyan Gong
  • Long Chen
  • Xinyu Wang
  • Yuxuan Xue
  • Wansu Yong
  • Hao Feng

Rapid and simultaneous detection of multiple coalmine dust parameters is crucial for accurate dust control. Existing methods often suffer from single detection indicators and delayed responses, making it difficult to effectively implement dust prevention and control measures. This study leverages artificial intelligence to propose an image-based detection method for coalmine dust, built upon multi-task deep learning. The method enables real-time detection of total dust concentration, respiratory dust concentration, and mass dispersion. An image preprocessing strategy is designed, and shared features are selected using the maximal information coefficient and variance inflation factor, with a multidimensional feature library constructed. The proposed architecture integrates a shared layer and a Set Transformer layer, both based on multi-head attention, to optimize inter-task representation consistency and enhance adaptability to dynamic variations in particle count. To jointly optimize the network architecture and training performance, an adaptive loss optimization mechanism and an Optuna-based hyperparameter tuning strategy are introduced. On an independent test set, the method is compared with multi-level baselines and evaluated via ablation studies. A dust particle image detection device is developed, and based on a fully mechanized tunneling face of a coalmine in Shaanxi Province, an experimental platform is constructed for application analysis. The results show that, under coal-dust conditions, the method achieves an average response cycle of 5. 5620 s, and the maximum average relative error across all outputs is 7. 3201%, meeting engineering requirements for real-time performance and detection accuracy. Overall, the method offers robust theoretical and technical support for intelligent dust monitoring.

EAAI Journal 2026 Journal Article

Defect detection of monocrystalline silicon wafers for photovoltaic applications using an improved you only look once version 8 small algorithm

  • Wenbo Bi
  • Xinyu Wang
  • Na Liu
  • Xu Xing
  • Lu Li
  • Hao Liu

Defects on the surface of photovoltaic monocrystalline silicon wafers, such as cracks, corners, and water stains, lead to significant performance degradation and economic losses during manufacturing. To address this, this paper proposes an improved You Only Look Once version 8 small (YOLOv8s) model. The proposed architecture integrates four strategic innovations. First, an Efficient Multi-Scale Convolution (EMSC) module is combined with the Cross-Stage Partial Bottleneck module with two convolutions (C2f) to enhance multi-scale feature extraction capabilities. Second, Spatial Pyramid Pooling-Fast (SPPF) is fused with the Large Separable Kernel Attention (LSKA) module to overcome limitations in processing local details. Third, the Dysample dynamic upsampling operator is introduced to maintain a compact model size while effectively improving detection speed. Finally, the Normalized Wasserstein Distance (NWD) is utilized as the loss function to address the sensitivity of the Intersection over Union (IoU) metric to positional deviations, enhancing precision for small targets. Experimental results demonstrate that the Efficient Lightweight Detection Network (ELDN) achieves superior performance on the validation set with a mean Average Precision (mAP) of 92. 8%. Notably, it exhibits robust generalization on an independent external test set, attaining a mAP of 92. 6%. Validation confirms that YOLOv8s-ELDN consistently outperforms mainstream models. Future research will focus on further optimizing efficiency for deployment on resource-constrained edge devices and addressing defect detection in complex manufacturing environments.

EAAI Journal 2026 Journal Article

Detection of dummy data injection attacks by using particle swarm optimization-attention temporal graph convolutional network model in power system

  • Xinyu Wang
  • Yifan Geng
  • Xiaoyuan Luo
  • Xinping Guan

As a novel deceptive topology attack, the dummy data injection attack (DDIA) poses a critical threat to power system security by exploiting both physical consistency and statistical mimicry of normal operational data, thereby evading traditional distance-based detection methods. Operating under the assumptions of a fully observable Direct Current power flow model with known topology and Gaussian measurement noise, DDIA can induce multi-line overloads while blending into legitimate measurement streams. To address this, this paper proposes a particle swarm optimization-attention temporal graph convolutional network (PSO-ATGCN) framework. The key innovation lies in a synergistic detection paradigm that integrates PSO for adversarial optimization with a topology-aware ATGCN for spatio-temporal feature extraction, specifically designed to counter the stealthy nature of DDIA: a multi-line overload DDIA model that simulates topology-driven attack scenarios; an ATGCN to dynamically capture the complex spatial-temporal dependencies inherent in grid topology and measurements; and a PSO module that concurrently tunes hyperparameters and selects critical features to enhance the model's robustness and attack-discriminative power. Experimental evaluations on IEEE 30-bus, 118-bus, and 300-bus systems demonstrate that the proposed method achieves at least 95. 21% detection accuracy under multi-line overload attacks, with lower false positive rates in noisy environments. Ablation studies confirm that the PSO component contributes a 1. 12% performance gain through feature optimization, while robustness tests validate the framework's superiority against adaptive attacks.

AAAI Conference 2026 Conference Paper

Forecast Then Calibrate: Feature Caching as ODE for Efficient Diffusion Transformers

  • Shikang Zheng
  • Liang Feng
  • Xinyu Wang
  • Qinming Zhou
  • Peiliang Cai
  • Chang Zou
  • Jiacheng Liu
  • Yuqi Lin

Diffusion Transformers (DiTs) have demonstrated exceptional performance in high-fidelity image and video generation. To reduce their substantial computational costs, feature caching techniques have been proposed to accelerate inference by reusing hidden representations from previous timesteps. However, current methods often struggle to maintain generation quality at high acceleration ratios, where prediction errors increase sharply due to the inherent instability of long-step forecasting. In this work, we adopt an ordinary differential equation (ODE) perspective on the hidden-feature sequence, modeling layer representations along the trajectory as a feature-ODE. We attribute the degradation of existing caching strategies to their inability to robustly integrate historical features under large skipping intervals. To address this, we propose FoCa (Forecast-then-Calibrate), which treats feature caching as a feature-ODE solving problem. Extensive experiments on image, video generation, and super-resolution tasks demonstrate the effectiveness of FoCa, especially under aggressive acceleration. Without additional training, FoCa achieves near-lossless speedups of 5.50× on FLUX, 6.45× on HunyuanVideo, 3.17× on Inf-DiT, and maintains high quality with a 4.53× speedup on DiT.

AAAI Conference 2026 Conference Paper

From Sequential to Recursive: Enhancing Decision-Focused Learning with Bidirectional Feedback

  • Xinyu Wang
  • Jinxiao Du
  • Yiyang Peng
  • Wei Ma

Decision-focused learning (DFL) has emerged as a powerful end-to-end alternative to conventional predict-then-optimize (PTO) pipelines by directly optimizing predictive models through downstream decision losses. Existing DFL frameworks are limited by their strictly sequential structure, referred to as sequential DFL (S-DFL). However, S-DFL fails to capture the bidirectional feedback between prediction and optimization in complex interaction scenarios. In view of this, we first time propose recursive decision-focused learning (R-DFL), a novel framework that introduces bidirectional feedback between downstream optimization and upstream prediction. We further extend two distinct differentiation methods: explicit unrolling via automatic differentiation and implicit differentiation based on fixed-point methods, to facilitate efficient gradient propagation in R-DFL. We rigorously prove that both methods achieve comparable gradient accuracy, with the implicit method offering superior computational efficiency. Extensive experiments on both synthetic and real-world datasets, including the newsvendor problem and the bipartite matching problem, demonstrate that R-DFL not only substantially enhances the final decision quality over sequential baselines but also exhibits robust adaptability across diverse scenarios in closed-loop decision-making problems.

YNICL Journal 2026 Journal Article

Natural progression of glioma enhances functional connection with the cerebral cortex through synaptogenesis

  • Jiacheng Lai
  • Yan Bai
  • Hongbo Bao
  • Shuai Wu
  • Xinyu Wang
  • Xia Liang
  • Peng Liang

OBJECTIVES: Understanding the progression mechanisms of glioma holds significant implications for improving clinical management. However, the natural progression patterns of glioma remain poorly understood due to the lack of longitudinal clinical samples from untreated patients. MATERIALS AND METHODS: In this study, we systematically explored the natural progression trajectory of glioma by combining functional magnetic resonance imaging (fMRI) analysis of 24 rare multifocal glioma patients with bioinformatic analysis of single-cell RNA sequencing (scRNA-seq) data obtained from tumor samples of glioma mouse with early, mid, and endpoint lesions. RESULTS: We discovered that larger tumors in multifocal gliomas exhibit stronger functional connectivity with the cerebral cortex and higher degree centrality within brain networks. ScRNA-seq of longitudinal mouse glioma samples revealed progressive activation of synaptic organization and associated regulatory pathways during the natural progression of glioma. CONCLUSION: Our multimodal, cross-scale study demonstrates that the natural progression pattern of glioma macroscopically manifests as functional hyperconnectivity with the cerebral cortex, which is supported by microscale molecular programs driving synaptogenesis. These findings elucidate the characteristics and mechanisms underlying glioma natural progression.

IJCAI Conference 2025 Conference Paper

A Novel Local Search Algorithm for the Vertex Bisection Minimization Problem

  • Rui Sun
  • Xinyu Wang
  • Yiyuan Wang
  • Jiangnan Li
  • Yi Zhou

The vertex bisection minimization problem (VBMP) is a fundamental graph partitioning problem with numerous real-world applications. In this study, we propose a (k, l, S)-cluster guided local search algorithm to address this challenge. First, we propose a novel (k, l, S)-cluster enumeration procedure, which is based on two key concepts: the (k, l, S)-cluster and the local cluster core. The (k, l, S)-cluster limits both the connectivity and distinct boundaries of a given vertex set, and the local cluster core represents the most cohesive substructure within a (k, l, S)-cluster. Building up on the above (k, l, S)-cluster enumeration procedure, we present a novel (k, l, S)-cluster guided perturbation mechanism designed to escape from local optima. Next, we propose a two-manner local search procedure that employs two distinct search models to explore the neighboring search space efficiently. Experimental results demonstrate that the proposed algorithm performs best on nearly all instances.

NeurIPS Conference 2025 Conference Paper

Block-Diagonal LoRA for Eliminating Communication Overhead in Tensor Parallel LoRA Serving

  • Xinyu Wang
  • Jonas Kübler
  • Kailash Budhathoki
  • Yida Wang
  • Matthäus Kleindessner

When serving a single base LLM with several different LoRA adapters simultaneously, the adapters cannot simply be merged with the base model’s weights as the adapter swapping would create overhead and requests using different adapters could not be batched. Rather, the LoRA computations have to be separated from the base LLM computations, and in a multi-device setup the LoRA adapters can be sharded in a way that is well aligned with the base model’s tensor parallel execution, as proposed in S-LoRA. However, the S-LoRA sharding strategy encounters some communication overhead, which may be small in theory, but can be large in practice. In this paper, we propose to constrain certain LoRA factors to be block-diagonal, which allows for an alternative way of sharding LoRA adapters that does not require any additional communication for the LoRA computations. We demonstrate in extensive experiments that our block-diagonal LoRA approach is similarly parameter efficient as standard LoRA (i. e. , for a similar number of parameters it achieves similar downstream performance) and that it leads to significant end-to-end speed-up over S-LoRA. For example, when serving on eight A100 GPUs, we observe up to 1. 79x (1. 23x) end-to-end speed-up with 0. 87x (1. 74x) the number of adapter parameters for Llama-3. 1-70B, and up to 1. 63x (1. 3x) end-to-end speed-up with 0. 86x (1. 73x) the number of adapter parameters for Llama-3. 1-8B.

TIST Journal 2025 Journal Article

Cross-platform Prediction of Depression Treatment Outcome Using Location Sensory Data on Smartphones

  • Soumyashree Sahoo
  • Md. Zakir Hossain
  • Chinmaey Shende
  • Parit Patel
  • Yushuo Niu
  • Reynaldo Morillo
  • Xinyu Wang
  • Shweta Ware

ABSTRACT Currently, depression treatment relies on closely monitoring patients’ response to treatment and adjusting the treatment as needed. Using self-reported or physician-administrated questionnaires to monitor treatment response is, however, subjective, costly and suffers from recall bias. In this paper, we explore using location sensory data collected passively on smartphones to predict treatment outcome. To address heterogeneous data collection on Android and iOS phones, the two predominant smartphone platforms, we explore using domain adaptation techniques to map their data to a common feature space, and then use the data jointly to train machine learning models. We further explore integrating contrastive learning with domain adaptation to augment data and learn feature embeddings. These learned embeddings are then used to train machine learning models to predict depression treatment outcomes. Our evaluation shows that using the embeddings learned by jointly integrating contrastive learning and domain adaptation leads to the best prediction accuracy. In addition, our results show that using location features and baseline self-reported questionnaire score can lead to F1 score up to 0.76. This accuracy is comparable to that obtained using periodic self-reported questionnaires, indicating that using location data is a promising direction for predicting depression treatment outcome. Last, when all location and questionnaire data are used together, the F1 score further increases to 0.79.

NeurIPS Conference 2025 Conference Paper

CymbaDiff: Structured Spatial Diffusion for Sketch-based 3D Semantic Urban Scene Generation

  • Li Liang
  • Bo Miao
  • Xinyu Wang
  • Naveed Akhtar
  • Jordan Vice
  • Ajmal Mian

Outdoor 3D semantic scene generation produces realistic and semantically rich environments for applications such as urban simulation and autonomous driving. However, advances in this direction are constrained by the absence of publicly available, well-annotated datasets. We introduce SketchSem3D, the first large‑scale benchmark for generating 3D outdoor semantic scenes from abstract freehand sketches and pseudo‑labeled annotations of satellite images. SketchSem3D includes two subsets, Sketch-based SemanticKITTI and Sketch-based KITTI-360 (containing LiDAR voxels along with their corresponding sketches and annotated satellite images), to enable standardized, rigorous, and diverse evaluations. We also propose Cylinder Mamba Diffusion (CymbaDiff) that significantly enhances spatial coherence in outdoor 3D scene generation. CymbaDiff imposes structured spatial ordering, explicitly captures cylindrical continuity and vertical hierarchy, and preserves both physical neighborhood relationships and global context within the generated scenes. Extensive experiments on SketchSem3D demonstrate that CymbaDiff achieves superior semantic consistency, spatial realism, and cross-dataset generalization. The code and dataset will be available at here.

IROS Conference 2025 Conference Paper

DGVO: A Dynamically Constrained Gradient Velocity Obstacle Approach for Mobile Robots in Dynamic Environments

  • Bowen Xiao
  • Bo Zhang
  • Danyu Zhang
  • Peiyan Xie
  • Xinyu Wang
  • Ruocheng Li

In this paper, we propose a framework based on velocity obstacles to address dynamic obstacle avoidance problem for constrained mobile robots. The framework establishes a nonlinear mapping from the control domain to the velocity space based on the robot’s kinematic model and input constraints. This mapping defines the Velocity Feasible Region (VFR) as the set of reachable velocities at the next time step. Utilizing the VFR, we propose a gradient field, called the Dynamically Constrained Gradient Velocity Obstacle (DGVO), to represent the feasible motion region for mobile robots. DGVO preserves the original feasible region of the mobile robot. Based on DGVO, we formulate an unconstrained gradient descent optimization problem to compute collision-free velocities in real time. This framework enables real-time online computation of collision-free velocities for any constrained mobile robot, and it exhibits strong robustness to sensor noise. Extensive simulations and real-world experiments have validated the effectiveness of the proposed method. The introduction of the entire work can be found at the following link: https://youtu.be/HrTNTSOhKvE.

IJCAI Conference 2025 Conference Paper

Evaluating and Mitigating Linguistic Discrimination in Large Language Models: Perspectives on Safety Equity and Knowledge Equity

  • Guoliang Dong
  • Haoyu Wang
  • Jun Sun
  • Xinyu Wang

Large language models (LLMs) typically provide multilingual support and demonstrate remarkable capabilities in solving tasks described in different languages. However, LLMs can exhibit linguistic discrimination due to the uneven distribution of training data across languages. That is, LLMs struggle to maintain consistency when handling the same task in different languages, compromising both safety equity and knowledge equity. In this paper, we first systematically evaluate the linguistic discrimination of LLMs from two aspects: safety and quality, using a form of metamorphic testing. The metamorphic relationship we examine is that LLMs are expected to deliver outputs with similar semantics when prompted with inputs that have the same meaning. We conduct this evaluation with two datasets based on four representative LLMs. The results show that LLMs exhibit stronger human alignment capabilities with queries in English, French, Russian, and Spanish compared to queries in Bengali, Georgian, Nepali and Maithili. Moreover, for queries in English, Danish, Czech and Slovenian, LLMs tend to produce responses with a higher quality compared to the other languages. Upon these findings, we propose LDFighter, a similarity-based voting method, to mitigate the linguistic discrimination in LLMs. We comprehensively evaluate LDFighter against a spectrum of queries including benign, harmful, and adversarial prompts. The results show that LDFighter significantly reduces jailbreak success rates and improves response quality. All code, data, and the technical appendix are publicly available at: \url{https: //github. com/dgl-prc/ldfighter}.

AAAI Conference 2025 Conference Paper

Global Attribute-Association Pattern Aggregation for Graph Fraud Detection

  • Mingjiang Duan
  • Da He
  • Tongya Zheng
  • Lingxiang Jia
  • Mingli Song
  • Xinyu Wang
  • Zunlei Feng

Fraud is increasingly prevalent, and its patterns are frequently changing, posing challenges for fraud detection methods such as random forests and Graph Neural Networks (GNNs), which rely on bin-based and mixture features separately. The former may lose crucial graph-associated features, while the latter face incorrect feature fusion. To overcome these limitations, we propose an approach based on attribute-association pattern that leverages the distinct attribute and association patterns differentiating fraudulent from benign behaviors, to enhance fraud detection capabilities. Attribute features are adaptively split into separate bins to eliminate incorrect attribute fusion and combine association patterns through graph neighbor message passing, thereby deriving attribute-association pattern features. Using the learned attribute-association patterns, the fraud patterns between a single pattern and the patterns across the entire graph are globally aggregated. Extensive experiments comparing our approach with 24 methods on 7 datasets demonstrate that the proposed method achieves SOTA performance.

IROS Conference 2025 Conference Paper

Hybrid Transformer-Mamba Model for 3D Semantic Segmentation

  • Xinyu Wang
  • Jinghua Hou
  • Zhe Liu
  • Yingying Zhu

Transformer-based methods have demonstrated remarkable capabilities in 3D semantic segmentation through their powerful attention mechanisms, but the quadratic complexity limits their modeling of long-range dependencies in large-scale point clouds. While recent Mamba-based approaches offer efficient processing with linear complexity, they struggle with feature representation when extracting 3D features. However, effectively combining these complementary strengths remains an open challenge in this field. In this paper, we propose HybridTM, the first hybrid architecture that integrates Transformer and Mamba for 3D semantic segmentation. In addition, we propose the Inner Layer Hybrid Strategy, which combines attention and Mamba at a finer granularity, enabling simultaneous capture of long-range dependencies and fine-grained local features. Extensive experiments demonstrate the effectiveness and generalization of our HybridTM on diverse indoor and outdoor datasets. Furthermore, our HybridTM achieves state-of-the-art performance on ScanNet, ScanNet200, and nuScenes benchmarks. The code will be made available at https://github.com/deepinact/HybridTM.

AAAI Conference 2025 Conference Paper

In2NeCT: Inter-class and Intra-class Neural Collapse Tuning for Semantic Segmentation of Imbalanced Remote Sensing Images

  • Junao Shen
  • Qiyun Hu
  • Tian Feng
  • Xinyu Wang
  • Hui Cui
  • Sensen Wu
  • Wei Zhang

Remote sensing images (RSIs) are frequently characterized by multi-scale inter-class objects and inconsistently distributed objects due to scene limitations, which would cause a significant data imbalance challenging the corresponding semantic segmentation. Recent methods have leveraged various deep learning techniques to capture high-quality representations for RSI semantic segmentation, but are hardly capable of addressing the afore-mentioned challenge given their limited explorations towards the mechanisms behind the representations. The recently discovered Neural Collapse (NC) phenomenon in computer vision models suggests the simplex equiangular tight frame (ETF) as the optimal representation structure, which has motivated us to observe that the optimal structure of last-layer representations is disrupted and inter-class representations for minor classes tend to become closer to each other beacuse of data imbalance. To address these issues, we propose Inter-class and Intra-class Neural Collapse Tuning (In2NeCT) to optimize the representations that satisfy the simplex ETF, which facilitates the discrimination of inter-class representations and the coherence of intra-class representations. Extensive experiments on three datasets demonstrate that our In2NeCT consistently leads to significant improvements in performance and outperforms the state-of-the-art methods.

NeurIPS Conference 2025 Conference Paper

Mamba Modulation: On the Length Generalization of Mamba Models

  • Peng Lu
  • Jerry Huang
  • Qiuhao Zeng
  • Xinyu Wang
  • Boxing Chen
  • Philippe Langlais
  • Yufei Cui

The quadratic complexity of the attention mechanism in Transformer models has motivated the development of alternative architectures with sub-quadratic scaling, such as state-space models. Among these, Mamba has emerged as a leading architecture, achieving state-of-the-art results across a range of language modeling tasks. However, Mamba’s performance significantly deteriorates when applied to contexts longer than those seen during pre-training, revealing a sharp sensitivity to context length extension. Through detailed analysis, we attribute this limitation to the out-of-distribution behavior of its state-space dynamics, particularly within the parameterization of the state transition matrix $A$. Unlike recent works which attribute this sensitivity to the vanished accumulation of discretization time steps, $\exp(-\sum_{t=1}^N{\Delta}_t)$, we establish a connection between state convergence behavior as the input length approaches infinity and the spectrum of the transition matrix $A$, offering a well-founded explanation of its role in length extension. Next, to overcome this challenge, we propose an approach that applies spectrum scaling to pre-trained Mamba models to enable robust long-context generalization by selectively modulating the spectrum of $A$ matrices in each layer. We show that this can significantly improve performance in settings where simply modulating ${\Delta}_t$ fails, validating our insights and providing avenues for better length generalization of state-space models with structured transition matrices.

NeurIPS Conference 2025 Conference Paper

NavBench: Probing Multimodal Large Language Models for Embodied Navigation

  • Yanyuan Qiao
  • Haodong Hong
  • Wenqi Lyu
  • Dong An
  • Siqi Zhang
  • Yutong Xie
  • Xinyu Wang
  • Qi Wu

Multimodal Large Language Models (MLLMs) have demonstrated strong generalization in vision-language tasks, yet their ability to understand and act within embodied environments remains underexplored. We present NavBench, a benchmark to evaluate the embodied navigation capabilities of MLLMs under zero-shot settings. NavBench consists of two components: (1) navigation comprehension, assessed through three cognitively grounded tasks including global instruction alignment, temporal progress estimation, and local observation-action reasoning, covering 3, 200 question-answer pairs; and (2) step-by-step execution in 432 episodes across 72 indoor scenes, stratified by spatial, cognitive, and execution complexity. To support real-world deployment, we introduce a pipeline that converts MLLMs' outputs into robotic actions. We evaluate both proprietary and open-source models, finding that GPT-4o performs well across tasks, while lighter open-source models succeed in simpler cases. Results also show that models with higher comprehension scores tend to achieve better execution performance. Providing map-based context improves decision accuracy, especially in medium-difficulty scenarios. However, most models struggle with temporal understanding, particularly in estimating progress during navigation, which may pose a key challenge.

NeurIPS Conference 2025 Conference Paper

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

  • Ling Fu
  • Zhebin Kuang
  • Jiajun Song
  • Mingxin Huang
  • Biao Yang
  • Yuzhe Li
  • Linghao Zhu
  • Qidi Luo

Scoring the Optical Character Recognition (OCR) capabilities of Large Multimodal Models (LMMs) has witnessed growing interest. Existing benchmarks have highlighted the impressive performance of LMMs in text recognition; however, their abilities in certain challenging tasks, such as text localization, handwritten content extraction, and logical reasoning, remain underexplored. To bridge this gap, we introduce OCRBench v2, a large-scale bilingual text-centric benchmark with currently the most comprehensive set of tasks ($4\times$ more tasks than the previous multi-scene benchmark OCRBench), the widest coverage of scenarios ($31$ diverse scenarios), and thorough evaluation metrics, with $10, 000$ human-verified question-answering pairs and a high proportion of difficult samples. Moreover, we construct a private test set with $1, 500$ manually annotated images. The consistent evaluation trends observed across both public and private test sets validate the OCRBench v2's reliability. After carefully benchmarking state-of-the-art LMMs, we find that most LMMs score below $50$ ($100$ in total) and suffer from five-type limitations, including less frequently encountered text recognition, fine-grained perception, layout perception, complex element parsing, and logical reasoning. The benchmark and evaluation scripts are available at https: //github. com/Yuliang-Liu/MultimodalOCR.

JBHI Journal 2025 Journal Article

SDPR: Prescription Recommendation With Syndrome Differentiation in Traditional Chinese Medicine

  • Wenjing Yue
  • Wendi Ji
  • Xinyu Wang
  • Xin Ma
  • Pengfei Wang
  • Xiaoling Wang

Prescription recommendation is critical for clinical decision support in Traditional Chinese Medicine (TCM), aiming to recommend a herb set based on a patient's symptoms. The core principle of TCM clinical practice, treatment based on syndrome differentiation (SD), follows a four-step progressive process: symptoms to syndromes, therapeutic methods, and herbs. However, existing models oversimplify this process by overlooking therapeutic methods, directly mapping symptoms to herbs or syndromes to herbs, resulting in information loss and reducing the effectiveness of recommended prescriptions. Furthermore, the implicit, sparse, and many-to-many relationships between syndromes and therapeutic methods, coupled with the nonlinear interactions between therapeutic methods and herbs, further hinder the modeling of the complete SD process. To address these challenges, we propose a novel four-partite graph paradigm that explicitly models the four key components of SD and their interactions, preserving critical information at each step and aligning more closely with clinicians' decision-making logic. Building on this, we develop SDPR, an SD-based prescription recommendation model comprising four modules aligned with all SD steps. Then, we integrated them into a multi-task learning framework to fully capture the progressive prescription process. To handle the implicit and complex relationships among syndromes, therapeutic methods, and herbs, we introduce a syndrome-induced pre-training strategy and a therapeutic method-aware contrastive learning framework. Extensive experiments on public and real-world datasets validate SDPR's effectiveness in herb recommendation and prescription retrieval, confirming the strength of the four-partite graph paradigm. Our broader goal is to advance the intelligent development of TCM in healthcare.

AAAI Conference 2025 Conference Paper

VQLTI: Long-Term Tropical Cyclone Intensity Forecasting with Physical Constraints

  • Xinyu Wang
  • Lei Liu
  • Kang Chen
  • Tao Han
  • Bin Li
  • Lei Bai

Tropical cyclone (TC) intensity forecasting is crucial for early disaster warning and emergency decision-making. Numerous researchers have explored deep-learning methods to address computational and post-processing issues in operational forecasting. Regrettably, they exhibit subpar long-term forecasting capabilities. We use two strategies to enhance long-term forecasting. (1) By enhancing the matching between TC intensity and spatial information, we can improve long-term forecasting performance. (2) Incorporating physical knowledge and physical constraints can help mitigate the accumulation of forecasting errors. To achieve the above strategies, we propose the VQLTI framework. VQLTI transfers the TC intensity information to a discrete latent space while retaining the spatial information differences, using large-scale spatial meteorological data as conditions. Furthermore, we leverage the forecast from the weather prediction model FengWu to provide additional physical knowledge for VQLTI. Additionally, we calculate the potential intensity (PI) to impose physical constraints on the latent variables. In the global long-term TC intensity forecasting, VQLTI achieves state-of-the-art results for the 24h to 120h, with the MSW (Maximum Sustained Wind) forecast error reduced by 35.65%-42.51% compared to ECMWF-IFS.

AAAI Conference 2024 Conference Paper

CGMGM: A Cross-Gaussian Mixture Generative Model for Few-Shot Semantic Segmentation

  • Junao Shen
  • Kun Kuang
  • Jiaheng Wang
  • Xinyu Wang
  • Tian Feng
  • Wei Zhang

Few-shot semantic segmentation (FSS) aims to segment unseen objects in a query image using a few pixel-wise annotated support images, thus expanding the capabilities of semantic segmentation. The main challenge lies in extracting sufficient information from the limited support images to guide the segmentation process. Conventional methods typically address this problem by generating single or multiple prototypes from the support images and calculating their cosine similarity to the query image. However, these methods often fail to capture meaningful information for modeling the de facto joint distribution of pixel and category. Consequently, they result in incomplete segmentation of foreground objects and mis-segmentation of the complex background. To overcome this issue, we propose the Cross Gaussian Mixture Generative Model (CGMGM), a novel Gaussian Mixture Models~(GMMs)-based FSS method, which establishes the joint distribution of pixel and category in both the support and query images. Specifically, our method initially matches the feature representations of the query image with those of the support images to generate and refine an initial segmentation mask. It then employs GMMs to accurately model the joint distribution of foreground and background using the support masks and the initial segmentation mask. Subsequently, a parametric decoder utilizes the posterior probability of pixels in the query image, by applying the Bayesian theorem, to the joint distribution, to generate the final segmentation mask. Experimental results on PASCAL-5i and COCO-20i datasets demonstrate our CGMGM's effectiveness and superior performance compared to the state-of-the-art methods.

AAAI Conference 2024 Conference Paper

DGA-GNN: Dynamic Grouping Aggregation GNN for Fraud Detection

  • Mingjiang Duan
  • Tongya Zheng
  • Yang Gao
  • Gang Wang
  • Zunlei Feng
  • Xinyu Wang

Fraud detection has increasingly become a prominent research field due to the dramatically increased incidents of fraud. The complex connections involving thousands, or even millions of nodes, present challenges for fraud detection tasks. Many researchers have developed various graph-based methods to detect fraud from these intricate graphs. However, those methods neglect two distinct characteristics of the fraud graph: the non-additivity of certain attributes and the distinguishability of grouped messages from neighbor nodes. This paper introduces the Dynamic Grouping Aggregation Graph Neural Network (DGA-GNN) for fraud detection, which addresses these two characteristics by dynamically grouping attribute value ranges and neighbor nodes. In DGA-GNN, we initially propose the decision tree binning encoding to transform non-additive node attributes into bin vectors. This approach aligns well with the GNN’s aggregation operation and avoids nonsensical feature generation. Furthermore, we devise a feedback dynamic grouping strategy to classify graph nodes into two distinct groups and then employ a hierarchical aggregation. This method extracts more discriminative features for fraud detection tasks. Extensive experiments on five datasets suggest that our proposed method achieves a 3% ~ 16% improvement over existing SOTA methods. Code is available at https://github.com/AtwoodDuan/DGA-GNN.

NeurIPS Conference 2024 Conference Paper

IaC-Eval: A Code Generation Benchmark for Cloud Infrastructure-as-Code Programs

  • Patrick T. Kon
  • Jiachen Liu
  • Yiming Qiu
  • Weijun Fan
  • Ting He
  • Lei Lin
  • Haoran Zhang
  • Owen M. Park

Infrastructure-as-Code (IaC), an important component of cloud computing, allows the definition of cloud infrastructure in high-level programs. However, developing IaC programs is challenging, complicated by factors that include the burgeoning complexity of the cloud ecosystem (e. g. , diversity of cloud services and workloads), and the relative scarcity of IaC-specific code examples and public repositories. While large language models (LLMs) have shown promise in general code generation and could potentially aid in IaC development, no benchmarks currently exist for evaluating their ability to generate IaC code. We present IaC-Eval, a first step in this research direction. IaC-Eval's dataset includes 458 human-curated scenarios covering a wide range of popular AWS services, at varying difficulty levels. Each scenario mainly comprises a natural language IaC problem description and an infrastructure intent specification. The former is fed as user input to the LLM, while the latter is a general notion used to verify if the generated IaC program conforms to the user's intent; by making explicit the problem's requirements that can encompass various cloud services, resources and internal infrastructure details. Our in-depth evaluation shows that contemporary LLMs perform poorly on IaC-Eval, with the top-performing model, GPT-4, obtaining a pass@1 accuracy of 19. 36%. In contrast, it scores 86. 6% on EvalPlus, a popular Python code generation benchmark, highlighting a need for advancements in this domain. We open-source the IaC-Eval dataset and evaluation framework at https: //github. com/autoiac-project/iac-eval to enable future research on LLM-based IaC code generation.

NeurIPS Conference 2024 Conference Paper

LION: Linear Group RNN for 3D Object Detection in Point Clouds

  • Zhe Liu
  • Jinghua Hou
  • Xinyu Wang
  • Xiaoqing Ye
  • Jingdong Wang
  • Hengshuang Zhao
  • Xiang Bai

The benefit of transformers in large-scale 3D point cloud perception tasks, such as 3D object detection, is limited by their quadratic computation cost when modeling long-range relationships. In contrast, linear RNNs have low computational complexity and are suitable for long-range modeling. Toward this goal, we propose a simple and effective window-based framework built on Linear group RNN (i. e. , perform linear RNN for grouped features) for accurate 3D object detection, called LION. The key property is to allow sufficient feature interaction in a much larger group than transformer-based methods. However, effectively applying linear group RNN to 3D object detection in highly sparse point clouds is not trivial due to its limitation in handling spatial modeling. To tackle this problem, we simply introduce a 3D spatial feature descriptor and integrate it into the linear group RNN operators to enhance their spatial features rather than blindly increasing the number of scanning orders for voxel features. To further address the challenge in highly sparse point clouds, we propose a 3D voxel generation strategy to densify foreground features thanks to linear group RNN as a natural property of auto-regressive models. Extensive experiments verify the effectiveness of the proposed components and the generalization of our LION on different linear group RNN operators including Mamba, RWKV, and RetNet. Furthermore, it is worth mentioning that our LION-Mamba achieves state-of-the-art on Waymo, nuScenes, Argoverse V2, and ONCE datasets. Last but not least, our method supports kinds of advanced linear RNN operators (e. g. , RetNet, RWKV, Mamba, xLSTM and TTT) on small but popular KITTI dataset for a quick experience with our linear RNN-based framework.

AAAI Conference 2024 Conference Paper

Unsupervised Object Interaction Learning with Counterfactual Dynamics Models

  • Jongwook Choi
  • Sungtae Lee
  • Xinyu Wang
  • Sungryull Sohn
  • Honglak Lee

We present COIL (Counterfactual Object Interaction Learning), a novel way of learning skills of object interactions on entity-centric environments. The goal is to learn primitive behaviors that can induce interactions without external reward or any supervision. Existing skill discovery methods are limited to locomotion, simple navigation tasks, or single-object manipulation tasks, mostly not inducing interaction between objects. Unlike a monolithic representation usually used in prior skill learning methods, we propose to use a structured goal representation that can query and scope which objects to interact with, which can serve as a basis for solving more complex downstream tasks. We design a novel counterfactual intrinsic reward through the use of either a forward model or successor features that can learn an interaction skill between a pair of objects given as a goal. Through experiments on continuous control environments such as Magnetic Block and 2.5-D Stacking Box, we demonstrate that an agent can learn object interaction behaviors (e.g., attaching or stacking one block to another) without any external rewards or domain-specific knowledge.

AAAI Conference 2023 Conference Paper

Anomaly Segmentation for High-Resolution Remote Sensing Images Based on Pixel Descriptors

  • Jingtao Li
  • Xinyu Wang
  • Hengwei Zhao
  • Shaoyu Wang
  • Yanfei Zhong

Anomaly segmentation in high spatial resolution (HSR) remote sensing imagery is aimed at segmenting anomaly patterns of the earth deviating from normal patterns, which plays an important role in various Earth vision applications. However, it is a challenging task due to the complex distribution and the irregular shapes of objects, and the lack of abnormal samples. To tackle these problems, an anomaly segmentation model based on pixel descriptors (ASD) is proposed for anomaly segmentation in HSR imagery. Specifically, deep one-class classification is introduced for anomaly segmentation in the feature space with discriminative pixel descriptors. The ASD model incorporates the data argument for generating virtual abnormal samples, which can force the pixel descriptors to be compact for normal data and meanwhile to be diverse to avoid the model collapse problems when only positive samples participated in the training. In addition, the ASD introduced a multi-level and multi-scale feature extraction strategy for learning the low-level and semantic information to make the pixel descriptors feature-rich. The proposed ASD model was validated using four HSR datasets and compared with the recent state-of-the-art models, showing its potential value in Earth vision applications.

NeurIPS Conference 2023 Conference Paper

Polyhedron Attention Module: Learning Adaptive-order Interactions

  • Tan Zhu
  • Fei Dou
  • Xinyu Wang
  • Jin Lu
  • Jinbo Bi

Learning feature interactions can be the key for multivariate predictive modeling. ReLU-activated neural networks create piecewise linear prediction models, and other nonlinear activation functions lead to models with only high-order feature interactions. Recent methods incorporate candidate polynomial terms of fixed orders into deep learning, which is subject to the issue of combinatorial explosion, or learn the orders that are difficult to adapt to different regions of the feature space. We propose a Polyhedron Attention Module (PAM) to create piecewise polynomial models where the input space is split into polyhedrons which define the different pieces and on each piece the hyperplanes that define the polyhedron boundary multiply to form the interactive terms, resulting in interactions of adaptive order to each piece. PAM is interpretable to identify important interactions in predicting a target. Theoretic analysis shows that PAM has stronger expression capability than ReLU-activated networks. Extensive experimental results demonstrate the superior classification performance of PAM on massive datasets of the click-through rate prediction and PAM can learn meaningful interaction effects in a medical problem.

ICRA Conference 2022 Conference Paper

Design of a Biomimetic Tactile Sensor for Material Classification

  • Kevin Dai
  • Xinyu Wang
  • Allison M. Rojas
  • Evan Harber
  • Yu Tian
  • Nicholas Paiva
  • Joseph Gnehm
  • Evan Schindewolf

Tactile sensing typically involves active exploration of unknown surfaces and objects, making it especially effective at processing the characteristics of materials and textures. A key property extracted by human tactile perception in material classification is surface roughness, which relies on measuring vibratory signals using the multi-layered fingertip structure. Existing robotic systems lack tactile sensors that are able to provide high dynamic sensing ranges, perceive material properties, and maintain a low hardware cost. In this work, we introduce the reference design and fabrication procedure of a miniature and low-cost tactile sensor consisting of a biomimetic cutaneous structure, including the artificial fingerprint, dermis, epidermis, and an embedded magnet-sensor structure which serves as a mechanoreceptor for converting mechanical information to digital signals. The presented sensor is capable of detecting high-resolution magnetic field data through the Hall effect and creating high-dimensional time-frequency domain features for material texture classification. Additionally, we investigate the effects of different superficial sensor fingerprint patterns for classifying materials through both simulation and physical experimentation. After extracting time series and frequency domain features, we assess a k-nearest neighbors classifier for distinguishing between different materials. The results from our experiments show that our biomimetic tactile sensors with fingerprint ridges can classify materials with more than 7. 7% higher accuracy and lower variability than ridge-less sensors. These results, along with the low cost and customizability of our sensor, demonstrate high potential for lowering the barrier to entry for a wide array of robotic applications, including modelless tactile sensing for texture classification, material inspection, and object recognition.

YNICL Journal 2022 Journal Article

Evaluating iron deposition in gray matter nuclei of patients with unilateral middle cerebral artery stenosis using quantitative susceptibility mapping

  • Huimin Mao
  • Weiqiang Dou
  • Kunjian Chen
  • Xinyu Wang
  • Xinyi Wang
  • Yu Guo
  • Chao Zhang

Iron mediated oxidative stress is involved in the process of brain injury after long-term ischemia. While increased iron deposition in the affected brain regions was observed in animal models of ischemic stroke, potential changes in the brain iron content in clinical patients with cerebral ischemia remain unclear. Quantitative susceptibility mapping (QSM), a non-invasive magnetic resonance imaging technique, can be used to evaluate iron content in the gray matter (GM) nuclei reliably. In this study, we aimed to quantitatively evaluate iron content changes in GM nuclei of patients with long-term unilateral middle cerebral artery (MCA) stenosis/occlusion-related cerebral ischemia using QSM. Forty-six unilateral MCA stenosis/occlusion patients and 38 age-, sex- and education-matched healthy controls underwent QSM. Clinical variables of history of hypertension, diabetes, hyperlipidemia, hyperhomocysteinemia, smoking, and drinking in all patients were evaluated. The iron-related susceptibility of GM nucleus subregions, including the bilateral caudate nucleus (CN), putamen (PU), globus pallidus (GP), thalamus, substantia nigra (SN), red nucleus, and dentate nucleus, was assessed. Susceptibility was compared between the bilateral GM nuclei in patients and controls. Receiver operating characteristic curve analysis was used to evaluate the efficacy of QSM susceptibility in distinguishing patients with unilateral MCA stenosis/occlusion from healthy controls. Multiple linear regression analysis was used to evaluate the relationship between ipsilateral susceptibility levels and clinical variables. Except for the CN, the susceptibility in most bilateral GM nucleus subregions was comparable in healthy controls, whereas for patients with unilateral MCA stenosis/occlusion, the ipsilateral PU, GP, and SN exhibited significantly higher susceptibility than the contralateral side (all P < 0.05). Compared with controls, susceptibility of the ipsilateral PU, GP, and SN and of contralateral PU in patients were significantly increased (all P < 0.05). The area under the curve (AUC) was greater for the ipsilateral PU than for the GP and SN (AUC = 0.773, 0.662 and 0.681; all P < 0.05). Multiple linear regression analysis showed that the increased susceptibility of the ipsilateral PU was significantly associated with hypertension, of the ipsilateral GP associated with smoking, and of the ipsilateral SN associated with diabetes (all P < 0.05). Our findings provide support for abnormal iron accumulation in the GM nuclei after chronic MCA stenosis/occlusion and its correlation with some cerebrovascular disease risk factors. Therefore, iron deposition in the GM nuclei, as measured by QSM, may be a potential biomarker for long-term cerebral ischemia.

JBHI Journal 2021 Journal Article

Identification of lncRNA Signature Associated With Pan-Cancer Prognosis

  • Guoqing Bao
  • Ran Xu
  • Xiuying Wang
  • Jianxiong Ji
  • Linlin Wang
  • Wenjie Li
  • Qing Zhang
  • Bin Huang

Long noncoding RNAs (lncRNAs) have emerged as potential prognostic markers in various human cancers as they participate in many malignant behaviors. However, the value of lncRNAs as prognostic markers among diverse human cancers is still under investigation, and a systematic signature based on these transcripts that related to pan-cancer prognosis has yet to be reported. In this study, we proposed a framework to incorporate statistical power, biological rationale, and machine learning models for pan-cancer prognosis analysis. The framework identified a 5-lncRNA signature ( ENSG00000206567, PCAT29, ENSG00000257989, LOC388282, and LINC00339 ) from TCGA training studies ( n = 1, 878). The identified lncRNAs are significantly associated (all P $\leq$ 1. 48E-11) with overall survival (OS) of the TCGA cohort ( n = 4, 231). The signature stratified the cohort into low- and high-risk groups with significantly distinct survival outcomes (median OS of 9. 84 years versus 4. 37 years, log-rank P = 1. 48E-38) and achieved a time-dependent ROC/AUC of 0. 66 at 5 years. After routine clinical factors involved, the signature demonstrated better performance for long-term prognostic estimation (AUC of 0. 72). Moreover, the signature was further evaluated on two independent external cohorts (TARGET, n = 1, 122; CPTAC, n = 391; National Cancer Institute) which yielded similar prognostic values (AUC of 0. 60 and 0. 75; log-rank P = 8. 6E-09 and P = 2. 7E-06). An indexing system was developed to map the 5-lncRNA signature to prognoses of pan-cancer patients. In silico functional analysis indicated that the lncRNAs are associated with common biological processes driving human cancers. The five lncRNAs, especially ENSG00000206567, ENSG00000257989 and LOC388282 that never reported before, may serve as viable molecular targets common among diverse cancers.

AAAI Conference 2018 Conference Paper

Event Representations for Automated Story Generation with Deep Neural Nets

  • Lara Martin
  • Prithviraj Ammanabrolu
  • Xinyu Wang
  • William Hancock
  • Shruti Singh
  • Brent Harrison
  • Mark Riedl

Automated story generation is the problem of automatically selecting a sequence of events, actions, or words that can be told as a story. We seek to develop a system that can generate stories by learning everything it needs to know from textual story corpora. To date, recurrent neural networks that learn language models at character, word, or sentence levels have had little success generating coherent stories. We explore the question of event representations that provide a midlevel of abstraction between words and sentences in order to retain the semantic information of the original data while minimizing event sparsity. We present a technique for preprocessing textual story data into event sequences. We then present a technique for automated story generation whereby we decompose the problem into the generation of successive events (event2event) and the generation of natural language sentences from events (event2sentence). We give empirical results comparing different event representations and their effects on event successor generation and the translation of events to natural language.

ICRA Conference 2003 Conference Paper

Two-dimensional signal transmission technology for robotics

  • Hiroyuki Shinoda 0001
  • Naoya Asamura
  • Mitsuhiro Hakozaki
  • Xinyu Wang

The forms of communication available now are categorized into the one or three dimensional. One dimensional communication includes metal wires and optical fibers in which the electro-magnetic field is confined in one dimensional medium. Wireless communication based on RF or optical connection emits electro-magnetic field in 3-D space. Now what if we have "two-dimensional communication" in which signals travels from one point to another point freely in elastic two-dimensional space using electromagnetic field confined in 2-D space? In this paper, we describe such a new technology of 2-D communication brings new paradigm to robotics. The methodologies of machine-design, system-integration, sensing, and computing will be drastically changed. We show architecture of the 2-D signal transmission based on relaying packets between communication chips on a thin sheet, the physical structure of the 2-D signal transmission, the protocols of the signal relay, and the results of the basic experiments.

v2026.09.13