Arrow Research search

Author name cluster

Wei Zhang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

208 papers
2 author rows

Possible papers

208

EAAI Journal 2026 Journal Article

A deep learning architecture for fast simulation of subsurface in-situ pressure profile dynamics based on Conditional Wasserstein Generative Adversarial Network with Gradient Penalty

  • Xiaoyin Peng
  • Bin Yuan
  • Wei Zhang
  • Shuhong Wu
  • Tianyi Fan
  • Baohua Wang
  • Zhihui Chen

Accurate pressure prediction in shale gas reservoirs is crucial for formulating scientific development strategies and optimizing recovery rates. While numerical simulation remains the dominant approach for reservoir simulation in fractured horizontal wells, its computational intensity becomes prohibitive under complex geological conditions. This study presents a novel proxy model based on Conditional Wasserstein Generative Adversarial Network with Gradient Penalty (CWGAN-GP) to address these limitations. By incorporating six critical conditional parameters (time, matrix permeability, fracture permeability, etc.) with fracture morphology characteristics, the model establishes the intrinsic relationship between reservoir conditions and pressure distribution. The integration of Wasserstein distance and gradient penalty theory effectively resolves convergence challenges induced by multi-parameter coupling, reducing training time and number of times compared to conventional Generative Adversarial Network (GAN) architectures. Validated through 60 geological scenarios, the proposed model achieves 98 % prediction accuracy at the trained six reservoir conditions and 92. 3 % accuracy at untrained timesteps under multi-constrained conditions, while demonstrating 2-3 orders of magnitude computational efficiency improvement over numerical simulations. Particularly, its capability to handle time-dependent well control variations and multi-scale geological uncertainties enables reliable applications in history matching and production optimization tasks. This data-driven approach establishes a new paradigm for real-time reservoir management in unconventional resources development, and also provides new ideas for the construction of big data models for dynamic prediction of unconventional reservoirs.

EAAI Journal 2026 Journal Article

A novel two-stage intelligent Kalman filter for maneuvering target tracking

  • Benqi Zhao
  • Gaoliang Peng
  • Wei Zhang
  • Jinghan Wang
  • Shiji Zhang
  • Feng Cheng

As the threat posed by unmanned aerial vehicles (UAVs) continues to grow, the integration of optoelectronic tracking systems with laser emitters has emerged as the most advanced solution. However, in the presence of frequent UAV maneuvers and environmental disturbances, current tracking approaches struggle to ensure both stability and accuracy, which compromises the ability to maintain sustained laser focus on the target. To address this challenge, this paper proposes a novel Two-Stage Intelligent Kalman Filter (TSIKF). By modeling target motion dynamics and environmental disturbance features through dedicated neural networks, we achieve a structured decomposition of uncertainties during tracking. Furthermore, we propose dynamic parameter embedding methods that map feature parameters in real time to the state space model, enabling adaptive, interpretable, and high-precision tracking in complex scenarios. The proposed method is rigorously validated through multiple experiments, including a simulation example, two publicly available UAV datasets, and a real-world tracking system. In all tests, TSIKF consistently outperforms existing state-of-the-art methods in terms of accuracy, stability, and generalizability, while fully meeting the real-time processing requirement. Experimental results demonstrate that TSIKF significantly enhances the system’s ability to maintain focus. This work presents an effective approach to high-dynamic target tracking, and provides theoretical insights and technical support for the development of next-generation intelligent optoelectronic interception systems

EAAI Journal 2026 Journal Article

A whole-life fatigue crack growth rate prediction method based on active learning and physics-informed loss

  • Qixuan Zhang
  • Wei Zhang
  • Rui Huang
  • Xinghui Chen
  • Bingbing Li
  • Fang Wang
  • Yiming Zheng
  • Changyu Zhou

Whole-life fatigue crack growth presents a critical challenge in structural integrity assessment, particularly under complex loading conditions. To address the limitations of standard physics-informed neural networks (PINNs) in capturing the temporal dynamics of fatigue crack growth, this study proposes an active learning-based physics-informed recurrent neural network (AC-PI-RNN). Specifically, a recurrent neural network (RNN) is integrated with a fully connected network, where dynamic features (stress intensity factor range) and static features (stress ratio, load amplitude, and pre-strain) are fused at the RNN input layer to provide comprehensive loading information. To optimize sample selection under data-limited conditions, a query-by-committee active learning strategy is employed. Furthermore, a modified Jones physical model is embedded into the network's loss function to enforce adherence to the underlying physics of fatigue crack growth. Comprehensive evaluations validate the efficacy of the proposed framework, demonstrating enhanced predictive fidelity, robust generalization, and improved physical consistency. A comparative analysis reveals that the AC-PI-RNN significantly outperforms traditional RNN and PINN models, showing a distinct advantage in capturing the complete trajectory of whole-life crack propagation with high precision. The proposed framework provides an effective and interpretable approach for whole-life fatigue crack growth rate prediction under complex loading conditions.

AAAI Conference 2026 Conference Paper

BulletTime4D: Towards High Spatio-Temporal Resolution Dynamic Scene Rendering via Spike-Guided Stereo Vision

  • Yiqian Chang
  • Haoran Xu
  • Qinghong Ye
  • Jianing Li
  • Xuan Wang
  • Wei Zhang
  • Peixi Peng

High spatio‑temporal resolution novel‑view scene rendering is crucial for applications such as sports analysis and scientific experiments. However, existing Dynamic Scene Rendering (DSR) approaches typically rely on conventional RGB cameras with limited frame rates, making it difficult to achieve high spatio‑temporal resolution. In this paper, we present BulletTime4D, a high spatio‑temporal resolution DSR framework, which is the first trial to integrate a spike camera with binocular RGB cameras for dynamic scene reconstruction. Specifically, we first develop a hybrid camera prototype and build a real‑world dynamic scene reconstruction dataset. Then, BulletTime4D presents a multi‑timescale deformation representation by combining low‑frequency spatio‑temporal features with high‑frequency inter‑frame motion features. Finally, a rendering network is designed capable of projecting 4D Gaussians into the spike domain for spike rendering, and a cross‑domain supervision strategy is proposed to achieve high‑frame‑rate texture and color rendering. The results show that BulletTime4D outperforms state‑of‑the‑art methods on both simulated and real‑world datasets. In addition, BulletTime4D can synthesize 300 FPS novel‑view renderings using stereo RGB cameras at 30 FPS and a single spike camera.

JBHI Journal 2026 Journal Article

Decoding Covert Speech from EEG by Functional Areas Spatio-Temporal Transformer

  • Muyun Jiang
  • Wei Zhang
  • Yi Ding
  • Kok Ann Colin Teo
  • LaiGuan Fong
  • Shuailei Zhang
  • Zhiwei Guo
  • Chenyu Liu

Covert speech involves imagining speaking without audible sound or any movements. Decoding covert speech from electroencephalogram (EEG) is challenging due to a limited understanding of neural pronunciation mapping and the low signal-to-noise ratio of the signal. In this study, we developed a large-scale multi-utterance speech EEG dataset from 57 right-handed native English-speaking subjects, each performing covert and overt speech tasks by repeating the same word in five utterances within a ten-second duration. Given the spatio-temporal nature of the neural activation process during speech pronunciation, we developed a Functional Areas Spatio-temporal Transformer (FAST), an effective framework for converting EEG signals into tokens and utilizing transformer architecture for sequence encoding. Our results reveal distinct and interpretable speech neural features by the visualization of FAST-generated activation maps across frontal and temporal brain regions with each word being covertly spoken, providing new insights into the discriminative features of the neural representation of covert speech. This is the first report of such a study, which provides interpretable evidence for speech decoding from EEG. The code for this work has been made public at https://github.com/Jiang-Muyun/FAST

EAAI Journal 2026 Journal Article

Diffusion model-enhanced coral identification: A lightweight multi-scale network for benthic imagery analysis

  • Changen Yang
  • Zhi Zhou
  • Zhuhua Hu
  • Zhaoxuan Lu
  • Yijun Shen
  • Wei Zhang
  • Xi Liang

Artificial intelligence (AI)–based coral monitoring can provide a transformative alternative to expert-dependent and labor-intensive surveys, holding significant ecological value for fragile coral ecosystems. However, coral detection algorithms remain constrained by limited fine-grained taxonomic datasets, edge-device capacity, and the complexity of coral texture feature extraction. To address these challenges, we propose CoralGrad-LiteNet (CG-LiteNet), a lightweight coral detection framework for efficient recognition. Key AI contributions: a Gradient-Aware Hierarchical Feature Fusion Module (GA-HFFM), which employs gradient convolution networks for multi-scale feature extraction, expanding receptive fields to 94. 2% while preserving fine textures; Slim-Backbone and Neck reconstruction, coupled with the proposed Dynamically Anchored Distribution-Aware Head (DADH) and Loss function optimization; and Diffusion Model–based image generation modules that overcome the coral data barrier by constructing the Sanya-Coral dataset and expanding it into Sanya-Coral AI-Enhanced with an 82. 5% scale-increase. Experiments demonstrate that CG-LiteNet surpasses state-of-the-art (SOTA) detectors in this domain. Diffusion-based augmentation yields an average +6. 44% mean Average Precision across intersection over union thresholds from 0. 50 to 0. 95 (mAP50–95) across all coral species, with a peak +14% gain on Favites. CG-LiteNet contains only 2. 1 million (M) parameters and 5. 4 billion floating point operations per second (GFLOPs), reducing size and computation by 18. 6% and 14% versus the baseline, while achieving +3. 6% mAP50 and +2. 8% mAP50–95 on Sanya-Coral, and 87. 1% mAP50 on the AI-Enhanced dataset. Overall, CG-LiteNet provides an efficient and scalable solution for coral detection while pioneering a novel dataset enhancement paradigm to advance fine-grained coral recognition. Code and datasets are available at: https: //github. com/yangchangen-s/CoralGrad-LiteNet.

AAAI Conference 2026 Conference Paper

Divide-and-Conquer Decoupled Network for Cross-Domain Few-Shot Segmentation

  • Runmin Cong
  • Anpeng Wang
  • Bin Wan
  • Cong Zhang
  • Xiaofei Zhou
  • Wei Zhang

Cross-domain few-shot segmentation (CD-FSS) aims to tackle the dual challenge of recognizing novel classes and adapting to unseen domains with limited annotations. However, encoder features often entangle domain-relevant and category-relevant information, limiting both generalization and rapid adaptation to new domains. To address this issue, we propose a Divide-and-Conquer Decoupled Network (DCDNet). In the training stage, to tackle feature entanglement that impedes cross-domain generalization and rapid adaptation, we propose the Adversarial-Contrastive Feature Decomposition (ACFD) module. It decouples backbone features into category-relevant private and domain-relevant shared representations via contrastive learning and adversarial learning. Then, to mitigate the potential degradation caused by the disentanglement, the Matrix-Guided Dynamic Fusion (MGDF) module adaptively integrates base, shared, and private features under spatial guidance, maintaining structural coherence. In addition, in the fine-tuning stage, to enhanced model generalization, the Cross-Adaptive Modulation (CAM) module is placed before the MGDF, where shared features guide private features via modulation ensuring effective integration of domain-relevant information. Extensive experiments on four challenging datasets show that DCDNet outperforms existing CD-FSS methods, setting a new state-of-the-art for cross-domain generalization and few-shot adaptation.

AAAI Conference 2026 Conference Paper

Dual-Path Knowledge-Augmented Contrastive Alignment Network for Spatially Resolved Transcriptomics

  • Wei Zhang
  • Jiajun Chu
  • Xinci Liu
  • Chen Tong
  • Xinyue Li

Spatial Transcriptomics (ST) is a technology that measures gene expression profiles within tissue sections while retaining spatial context. It reveals localized gene expression patterns and tissue heterogeneity, both of which are essential for understanding disease etiology. However, its high cost has driven efforts to predict spatial gene expression from whole slide images. Despite recent advancements, current methods still face significant limitations, such as under-exploitation of high-level biological context, over-reliance on exemplar retrievals, and inadequate alignment of heterogeneous modalities. To address these challenges, we propose DKAN, a novel Dual-path Knowledge-Augmented contrastive alignment Network that predicts spatially resolved gene expression by integrating histopathological images and gene expression profiles through a biologically informed approach. Specifically, we introduce an effective gene semantic representation module that leverages the external gene database to provide additional biological insights, thereby enhancing gene expression prediction. Further, we adopt a unified, one-stage contrastive learning paradigm, seamlessly combining contrastive learning and supervised learning to eliminate reliance on exemplars, complemented with an adaptive weighting mechanism. Additionally, we propose a dual-path contrastive alignment module that employs gene semantic features as dynamic cross-modal coordinators to enable effective heterogeneous feature integration. Through extensive experiments across three public ST datasets, DKAN demonstrates superior performance over state-of-the-art models, establishing a new benchmark for spatial gene expression prediction and offering a powerful tool for advancing biological and clinical research.

EAAI Journal 2026 Journal Article

Engineering application of non-dominated sorting genetic algorithm III: Multi-objective optimization of ultra-high performance concrete for diverse scenarios

  • Wei Zhang
  • Zhenhua Duan
  • Yuqing Wu
  • Chao Liu
  • Yizhou Yao
  • Ahmed Nasr
  • Qingmei Yang
  • Huiyu Xia

This study addresses the technical limitations of conventional mix design methods for ultra-high performance concrete (UHPC) concerning multi-objective synergistic optimization and diverse scenarios adaptability. Leveraging 2824 experimental data points, a comprehensive prediction system was established for mechanical properties, workability and durability. The prediction performance of ten machine learning algorithms was systematically evaluated, and the SHapley Additive exPlanations (SHAP) method was used to elucidate the influence mechanism of crucial features. Furthermore, a comprehensive collaborative optimization framework for UHPC under typical engineering scenarios was developed by integrating the non-dominated sorting genetic algorithm III (NSGA-III) with the technique for order preference by similarity to ideal solution (TOPSIS) decision-making model, and visualization technology was integrated to construct a graphical user interface (GUI) system. The results demonstrate that the NSGA-III algorithm achieved continuous hypervolume (HV) improvement within 500 generations, while the spacing indicator decreased rapidly in the initial iterations, confirming its capability to approximate the actual Pareto front through adaptive crossover-mutation strategies and elitism preservation. The developed ‘data driven - performance prediction - multi objective optimization - decision analysis' technical system, provides a quantifiable and scalable solution for addressing multi-objective optimization challenges in engineering materials.

AAAI Conference 2026 Conference Paper

Evidence-aware Integration and Domain Identification of Spatial Transcriptomics Data

  • Wei Zhang
  • Siyu Yi
  • Lezhi Chen
  • Yifan Wang
  • Ziyue Qiao
  • Yongdao Zhou
  • Wei Ju

Spatial transcriptomics (ST) enables joint profiling of gene expression and spatial positions, thereby revealing spatially resolved biological functions. However, many existing ST analysis methods often fail to explicitly quantify the belief and uncertainty in decisions caused by noisy ST data, making it difficult to handle spots of varying quality in a fine-grained manner. In addition, domain identification is a fundamental and critical task in ST, but commonly used models that separate expression learning and clustering often struggle to learn cluster-friendly latent representations effectively. To address these issues, we propose PREST, a prototype-based evidence-aware integration framework for ST data. PREST performs multi-scale representation learning with fine-grained attention fusion and introduces learnable class prototypes to quantify belief and uncertainty in model decisions. We aim to align overall belief scores with latent semantic information to enhance uncertainty quantification and prototype learning, thereby promoting the learning of clustering-friendly representations. PREST further integrates an uncertainty-aware reconstruction module and spatial regularization to reduce overfitting to unreliable spots and promote denoised, discriminative representations. Extensive experiments on several benchmark datasets validate the effectiveness and superiority of our proposed PREST across various downstream tasks.

AAAI Conference 2026 Conference Paper

Exploring Surround-View Fisheye Camera 3D Object Detection

  • Changcai Li
  • Wenwei Lin
  • Zuoxun Hou
  • Gang Chen
  • Wei Zhang
  • Huihui Zhou
  • Weishi Zheng

In this work, we explore the technical feasibility of implementing end-to-end 3D object detection (3DOD) with surround-view fisheye camera system. Specifically, we first investigate the performance drop incurred when transferring classic pinhole-based 3D object detectors to fisheye imagery. To mitigate this, we then develop two methods that incorporate the unique geometry of fisheye images into mainstream detection frameworks: one based on the bird's-eye-view (BEV) paradigm, named FisheyeBEVDet, and the other on the query-based paradigm, named FisheyePETR. Both methods adopt spherical spatial representations to effectively capture fisheye geometry. In light of the lack of dedicated evaluation benchmarks, we release Fisheye3DOD, a new open dataset synthesized using CARLA and featuring both standard pinhole and fisheye camera arrays. Experiments on Fisheye3DOD demonstrate that our fisheye-compatible modeling improves accuracy by up to 6.2% compared to baseline methods.

TIST Journal 2026 Journal Article

From Hallucination to Certainty: Meta-Knowledge Guided Self-Correcting Large Language Models

  • Wei Zhang
  • Guojun Dai
  • Ding Luo
  • Yan Wang
  • Chen Ye

Recent advancements in Large Language Models (LLMs) have demonstrated remarkable capabilities in natural language understanding and generation. To further enhance their factual grounding and reasoning fidelity, integrating LLMs with Knowledge Graphs (KGs) has emerged as a promising direction. Significant progress has been made in leveraging KGs to augment LLM reasoning through methods like Retrieval-Augmented Generation. However, effectively harnessing the synergy between LLMs and KGs for robust and reliable reasoning still presents critical challenges. Specifically: (1) LLMs struggle to effectively interpret and utilize the structured nature of KGs, due to the discrepancy between their text-based training and KG's symbolic representations; (2) querying and reasoning over structured knowledge in KGs remains inefficient for LLMs, hindering complex inference. To address these limitations, we introduce Meta-Knowledge enhanced Knowledge Graph (MKG), a novel framework that empowers LLMs to effectively leverage structured knowledge from KGs. MKG employs Meta-Knowledge, stored in a multi-store memory with a Self-Correcting Mechanism, to guide LLMs in KG retrieval and reasoning. Our experimental evaluations on complex question answering benchmarks demonstrate that MKG achieves significant performance gains, outperforms the baseline Original LLM, Retrieval-Augmented Generation (RAG), ReAct, GraphRAG and ToG frameworks by 25%, 17%, 11%, 3.3% and 2.6%, respectively.

JBHI Journal 2026 Journal Article

Heterophily-Aware Spectral GCN for Population-Level Brain Disorder Prediction

  • Hao Zhang
  • Liping Wang
  • Yitian Zhao
  • Jianyang Xie
  • Tiyu Fang
  • Ran Song
  • Wei Zhang

Integrating resting-state functional magnetic resonance imaging (rs-fMRI) and phenotypic data is a promising way to build a comprehensive population graph for the prediction of brain disorders using graph neural networks (GNNs). However, existing GNN-based methods face two limitations: the complexity of relationships between subjects poses challenges in constructing a well-defined population graph, and the inherent node heterophily within the population graph is often overlooked. To address them, we propose a population graph with a phenotypic encoder, which leverages rs-fMRI and phenotypic data to model complex relationships between subjects and enables GNN to learn population-level features. We also design a heterophily-aware spectral graph convolution network that incorporates local similarity-based learning to assess node homophily and addresses the heterophily issue. Experiments demonstrate that our method performs well in classifying both Alzheimer's Disease and Autism Spectrum Disorder. In addition, it can distinguish between progressive and stable mild cognitive impairment, facilitating timely interventions for the diseases.

AAAI Conference 2026 Conference Paper

HI-SLAM2: Geometry-Aware Gaussian SLAM for Fast Monocular Scene Reconstruction (Abstract Reprint)

  • Wei Zhang
  • Qing Cheng
  • David Skuddis
  • Niclas Zeller
  • Daniel Cremers
  • Norbert Haala

We present HI-SLAM2, a geometry-aware Gaussian SLAM system that achieves fast and accurate monocular scene reconstruction using only RGB input. Existing Neural SLAM or 3DGS-based SLAM methods often trade off between rendering quality and geometry accuracy, our research demonstrates that both can be achieved simultaneously with RGB input alone. The key idea of our approach is to enhance the ability for geometry estimation by combining easy-to-obtain monocular priors with learning-based dense SLAM, and then using 3D Gaussian splatting as our core map representation to efficiently model the scene. Upon loop closure, our method ensures on-the-fly global consistency through efficient pose graph bundle adjustment and instant map updates by explicitly deforming the 3D Gaussian units based on anchored keyframe updates. Furthermore, we introduce a grid-based scale alignment strategy to maintain improved scale consistency in prior depths for finer depth details. Through extensive experiments on Replica, ScanNet, and ScanNet++, we demonstrate significant improvements over existing Neural SLAM methods and even surpass RGB-D-based methods in both reconstruction and rendering quality.

EAAI Journal 2026 Journal Article

How recycling technologies play stage-specific roles in renewable energy and energy storage systems? Insights from a patent claim analysis

  • Xuefeng Zhao
  • Qianwen Hao
  • Wei Zhang
  • Rui Wang
  • Chengjiang Li

The continued expansion of Renewable Energy and Energy Storage Systems (REESS) has led to increasing challenges related to material consumption and environmental sustainability. Recycling technologies (RT) have emerged as key enablers of resource efficiency and circularity. However, the stage-specific contributions of different RT types within REESS remain poorly understood. This study employs a large language model (LLM) to generate search terms for constructing corpora and classifying patents, subdivides patent subsets, builds a Type & Dependency mechanism for claim analysis, and uses three analytical approaches to elucidate the evolving role of RTs in REESS development. The results reveal three main findings: (1) RTs display stage-specific application patterns, with Chemical RTs dominating the control stage, while Biological RTs remain limited in the generation stage; (2) Each RT type plays a distinct role across stages. Chemical RTs show the strongest impact, Physical RTs offer stable support, and Biological RTs remain emerging with limited but growing potential; (3) RTs exhibit increasing interconnectivity across REESS stages, indicating a shift toward more integrated and circular technological development. These findings contribute to a deeper understanding of RT integration in REESS and offer valuable implications for advancing sustainable energy systems.

EAAI Journal 2026 Journal Article

Image-plane geometric decoding for view-invariant indoor scene reconstruction

  • Mingyang Li
  • Yimeng Fan
  • Changsong Liu
  • Lixue Xu
  • Xin Wang
  • Yanyan Liu
  • Wei Zhang

Volume-based indoor scene reconstruction offers superior generalization and real-time potential. However, existing frameworks rely on weak multi-view geometric constraints, leading to quality degradation as input views decrease. In sparse-view scenarios, these methods often exhibit geometric fragmentation due to the lack of robust priors. To address this, we propose Image-Plane Geometric Decoding Reconstruction (IPDRecon) pipeline, a framework integrating geometric optical principles as inductive bias to systematically exploit single-view spatial information for view-invariant reconstruction. Our approach establishes a structured geometric constraint mechanism through three synergistic modules: the Pixel-level Confidence Encoder (PCE) leverages state–space modeling with diffuse reflection principles to extract distance and position awareness; the Affine Compensation Module (ACM) enforces rigid geometric constraints via affine invariance, enabling accurate recovery of complex structures under sparse views; and the Image-Plane Spatial Decoder (IPSD) employs a multi-source geometric prior fusion strategy to transform traditional back-projection into geometry-aware spatial encoding. Extensive experiments on benchmark datasets (ScanNet V2) demonstrate exceptional stability, achieving 79. 7% Precision and a 0. 722 harmonic mean of precision and recall (F-score). In robustness evaluations averaged on a per-scene basis across the validation set, our method shows remarkable resilience when reducing views from 100 to 60. It maintains a 99. 7% mean performance retention rate, with a per-scene coefficient of variation of 0. 24% and a maximum performance drop of only 0. 42%. These results confirm that our physics-guided approach provides a robust solution for high-fidelity reconstruction in view-limited applications.

JBHI Journal 2026 Journal Article

Img2Gene: Debiased Spatially Resolved Transcriptomics with Biological Context from Pathology Images

  • Wei Zhang
  • Tong Chen
  • Wenxin Xu
  • Collin Sakal
  • Xinyue Li

Spatial transcriptomics integrates morphological information from pathology images with gene expression, providing high-resolution spatial gene expression profiles while preserving tissue architectures in a cost-effective manner. However, the inherent heterogeneity between images and gene expression data, coupled with sparse gene expression distribution, poses significant challenges for accurate and unbiased prediction models. To address these issues, we propose Img2Gene, a debiased framework designed to predict gene expression levels from whole slide images by incorporating biological context. Specifically, we integrate causal analysis into the gene expression prediction task to mitigate data sparsity and achieve unbiased predictions. Furthermore, we employ gene set enrichment analysis to identify highly associated pathway information as biological context and introduce a cross-modal coherence loss to align data from different modalities, fostering enhanced interplay among diverse features and achieving improved accuracy of gene expression prediction. Extensive experiments conducted on four public datasets demonstrate that our method achieves state-of-the-art performance. The pathway data and source code are available at https://github.com/coffeeNtv/Img2Gene.

TIST Journal 2026 Journal Article

Interpretable Structure Learning for Knowledge Components in Education

  • Yuang Wei
  • Yuan-Hao Jiang
  • Changyong Qi
  • Wei Zhang
  • Bo Jiang

Structural relationships among Knowledge Components (KCs) are essential for adaptive learning systems, as they support accurate cognitive diagnosis, personalized path planning, and targeted resource recommendation. However, existing approaches frequently capture correlations instead of reliable directional dependency signals and tend to converge prematurely or become inefficient as graph dimensionality grows. These limitations weaken the reliable modeling of KC-level structure, which in turn reduces interpretability and limits downstream benefits for diagnosis, planning, and recommendation. To this end, we propose a novel structure learning framework that integrates psychometric modeling with structural search. First, we design the I tem R esponse T heory (IRT)-based I nformation C riterion ( IRIC ), an interpretable scoring function that combines information entropy with causal effect estimation grounded in IRT, jointly capturing statistical associations and directionality-sensitive signals under latent ability control. Second, we develop C o- E volutionary O ptimization for S tructural S earch ( CEO-SS ), a multi-population evolutionary algorithm with a game-inspired co-evolution mechanism that balances exploration and exploitation, avoiding premature convergence and showing robust search behavior as graph dimensionality increases within the evaluated benchmarks. Extensive experiments on three types of datasets—including benchmark causal discovery datasets, the public educational dataset, and real-world classroom data—demonstrate that our framework consistently outperforms strong baselines in accuracy and stability, with especially clear gains in adjacency recovery and more modest improvements in edge-direction recovery. In addition, expert evaluation suggests that the learned structures are more diagnostically useful, more actionable for remediation, and more pedagogically plausible than those produced by alternative scoring methods. Overall, the proposed framework provides an interpretable and practically valuable approach to learning KC structures for adaptive learning.

AAAI Conference 2026 Conference Paper

Learning Personalised Human Internal Cognition from External Expressive Behaviours for Real Personality Recognition

  • Xiangyu Kong
  • Hengde Zhu
  • Haoqin Sun
  • Zhihao Guo
  • Jiayan Gu
  • Xinyi Ni
  • Wei Zhang
  • Shizhe Liu

Automatic real personality recognition (RPR) aims to evaluate human real personality traits from their expressive behaviours. However, most existing solutions generally act as external observers to infer observers' personality impressions based on target individuals' expressive behaviours, which significantly deviate from their real personalities and consistently lead to inferior recognition performance. Inspired by the association between real personality and human internal cognition underlying the generation of expressive behaviours, we propose a novel RPR approach that efficiently simulates personalised internal cognition from external short audio-visual behaviours expressed by target individual. The simulated personalised cognition, represented as a set of network weights that enforce the personalised network to reproduce the individual-specific facial reactions, is further encoded as a graph containing two-dimensional node and edge feature matrices, with a novel 2D Graph Neural Network (2D-GNN) proposed for inferring real personality traits from it. To simulate real personality-related cognition, an end-to-end (E2E) strategy is designed to jointly train our cognition simulation, 2D graph construction, and personality recognition modules. Experiments show our approach’s effectiveness in capturing real personality traits with superior computational efficiency.

AAAI Conference 2026 Conference Paper

MAUGen: A Unified Diffusion Approach for Multi-Identity Facial Expression and AU Label Generation

  • Xiangdong Li
  • Ye Lou
  • Ao Gao
  • Wei Zhang
  • Siyang Song

The lack of large-scale, demographically diverse face images with precise Action Unit (AU) occurrence and intensity annotations has long been recognized as a fundamental bottleneck in developing generalizable facial AU recognition systems. In this paper, we propose MAUGen, a diffusion-based multi-modal framework that jointly generates a large collection of photorealistic facial expressions and anatomically consistent AU labels, including both occurrence and intensity, conditioned on a single descriptive text prompt. Our MAUGen involves two key modules: (1) a Multi-modal Representation Learning (MRL) module that captures the relationships among the paired facial textual description, facial identity, facial expression image, and AU activations within a unified latent space; and (2) a Diffusion-based Image-label Generator (DIG) that decodes the obtained joint representation into aligned facial image-label pairs across diverse identities. Under this framework, we introduce the Multi-Identity Facial Action (MIFA), a large-scale multi-modal (i.e., text descriptions, face images with labels) synthetic dataset that features comprehensive AU annotations and identity variations. Extensive experiments demonstrate that MAUGen outperforms existing methods in synthesizing photorealistic, demographically diverse facial images, along with semantically aligned AU labels.

JBHI Journal 2026 Journal Article

Multi-Atlas Brain Network Classification Through Consistency Distillation and Complementary Information Fusion

  • Jiaxing Xu
  • Mengcheng Lan
  • Xia Dong
  • Kai He
  • Wei Zhang
  • Qingtian Bian
  • Yiping Ke

Brain network analysis plays a crucial role in identifying distinctive patterns associated with neurological disorders. Functional magnetic resonance imaging (fMRI) enables the construction of brain networks by analyzing correlations in blood-oxygen-level-dependent (BOLD) signals across different brain regions, known as regions of interest (ROIs). These networks are typically constructed using atlases that parcellate the brain based on various hypotheses of functional and anatomical divisions. However, there is no standard atlas for brain network classification, leading to limitations in detecting abnormalities in disorders. Recent methods leveraging multiple atlases fail to ensure consistency across atlases and lack effective ROI-level information exchange, limiting their efficacy. To address these challenges, we propose the Atlas-Integrated Distillation and Fusion network (AIDFusion), a novel framework designed to enhance brain network classification using fMRI data. AIDFusion introduces a disentangle Transformer to filter out inconsistent atlas-specific information and distill meaningful cross-atlas connections. Additionally, it enforces subject- and population-level consistency constraints to improve cross-atlas coherence. To further enhance feature integration, AIDFusion incorporates an inter-atlas message-passing mechanism that facilitates the fusion of complementary information across brain regions. We evaluate AIDFusion on four resting-state fMRI datasets encompassing different neurological disorders. Experimental results demonstrate its superior classification performance and computational efficiency compared to state-of-the-art methods. Furthermore, a case study highlights AIDFusion’s ability to extract interpretable patterns that align with established neuroscience findings, reinforcing its potential as a robust tool for multi-atlas brain network analysis.

AAAI Conference 2026 Conference Paper

PHPFND: Detecting Fake News via Post-Hoc Processing of LLMs Hallucination

  • Jinke Ma
  • Jiachen Ma
  • Wei Zhang
  • Yong Liu

Large Language Models (LLMs) perform excellently in fake news detection tasks, but their outputs are often accompanied by hallucinations, i.e., generated content that is contradictory to facts. Previous studies have mostly mitigated hallucinations through prompt design. However, this paper reveals that regions in news articles which easily induce hallucinations in LLMs correspond closely to the most challenging regions for fake news detectors. In this paper, we propose a fake news detection framework (PHPFND) based on post-hoc processing of LLMs hallucination. Specifically, our framework includes a hallucination detection module (ISHD) based on information structuring that detects three types of hallucinations in LLMs in a targeted manner, and a hallucination-driven feature enhancement mechanism (HDFE) that incorporates hallucination signals as explicit features into sentence-level encoding and feature fusion to guide the model’s attention toward high-risk regions. Experimental results on two mainstream fake news datasets show that our proposed method significantly outperforms LLM-based baselines.

EAAI Journal 2026 Journal Article

Physics-integrated intelligent method for propeller aerodynamic property predictions of electric aircraft

  • Wei Zhang
  • Ziwei Wang
  • Xiang Li
  • Song Xiang

Aerodynamic properties including thrust and torque are crucial parameters for propellers. Accurate estimations of the aerodynamic properties are of great importance for propeller's design. This paper proposes a physics-integrated intelligent method to estimate the thrust and torque of propellers. First, wind tunnel tests are conducted to collect experimental data at various wind and rotational speeds. Theoretical data is simulated using the fundamental Strip Theory, and the experimental data is augmented with the theoretical data. Next, a deep neural network is developed that integrates both experimental and simulation data. To further optimize the model and reduce reliance on the practical wind tunnel data, the impact of the key parameters on the performance of the deep neural networks is examined. This physics-integrated intelligent model is shown to be effective in predicting aerodynamic parameters for propellers, significantly decreasing the need for large amounts of real wind tunnel data. This method offers a promising solution for designing electric aircraft propellers while minimizing experimental costs compared to other high-fidelity methods.

TMLR Journal 2026 Journal Article

Random Projection-Induced Gaussian Latent Features for Arbitrary Style Transfer

  • Weizhi Lu
  • Zhongzheng Li
  • Dongchen Gao
  • Mingrui Chen
  • Weiyu Li
  • Jinglin Zhang
  • Wei Zhang

The feature transfer technique centered on mean and variance statistics, widely known as AdaIN, lies at the core of current style transfer research. This technique relies on the assumption that latent features for style transfer follow Gaussian distributions. In practice, however, this assumption is often hard to meet, as the features typically exhibit sparse distributions due to the significant spatial correlation inherent in natural images. To tackle this issue, we propose first performing a random projection on the sparse features, and then conducting style transfer on these projections. Statistically, the projections will satisfy or approximate Gaussian distributions, thereby better aligning with AdaIN's requirements and enhancing transfer performance. With the stylized projections, we can further reconstruct them back to the original feature space by leveraging compressed sensing theory, thereby obtaining the stylized features. The entire process constitutes a projection-stylization-reconstruction module, which can be seamlessly integrated into AdaIN without necessitating network retraining. Additionally, our proposed module can also be incorporated into another promising style transfer technique based on cumulative distribution functions, dubbed EFDM. This technique faces limitations when there are substantial differences in sparsity levels between content and style features. By projecting both types of features into dense Gaussian distributions, random projection can reduce their sparsity disparity, thereby improving performance. Experiments demonstrate that the aforementioned performance improvements can be achieved on existing state-of-the-art approaches.

AAAI Conference 2026 Conference Paper

Sample-specific Modality Diagnosis and Cross-modal Enhancement for Incomplete Multimodal Representations

  • Junsong Chen
  • Jiyuan Liu
  • Suyuan Liu
  • Wei Zhang
  • Ao Li
  • En Zhu
  • Xinwang Liu

In multimodal sentiment analysis, modality missingness and quality degradation are common. Existing methods often rely on batch-level modality generation, generation but neglect sample-level missingness, hence their flexibility is limited severely in real-world scenarios. To address this, Sample-specific Modality Diagnosis and Cross-modal Enhancement for Incomplete Multimodal Representations (SMCIR) is proposed. Specifically, The Dynamic Multi-feature Fusion Detector (DMFD) is presented, which detects missingness and severity at the sample-level using indicators such as information entropy, modality similarity, and mutual information. Unlike batch-based methods, the DMFD provides fine-grained detection and adaptive responses, improving sensitivity to modality disturbances. Meanwhile, the Context-aware Modality Completion Generator (CMCG) is developed to restore missing modalities through context-guided reconstruction using multiscale feature fusion and cross-modal attention. In this way, the proposed CMCG method can avoid redundancy and inconsistency, enhancing the consistency and discriminativity of the fused representation. In CMCG, the text modality serves as a stable guide to improve context consistency. Experiments on the CMU-MOSI and CMU-MOSEI datasets show that SMCIR outperforms existing full-modal and non-recovery-based methods, well validating its efficacy and superiority in multimodal learning.

AAAI Conference 2026 Conference Paper

Seeing Is Believing: Rich-Context Hallucination Detection for MLLMs via Backward Visual Grounding

  • Pinxue Guo
  • Chongruo Wu
  • Xinyu Zhou
  • Lingyi Hong
  • Zhaoyu Chen
  • Jinglun Li
  • Kaixun Jiang
  • Sen-Ching Samson Cheung

Multimodal Large Language Models (MLLMs) have unlocked powerful cross-modal capabilities, but still significantly suffer from hallucinations. As such, accurate detection of hallucinations in MLLMs is imperative for ensuring their reliability in practical applications. To this end, guided by the principle of “Seeing is Believing”, we introduce VBackChecker, a novel reference-free hallucination detection framework that verifies the consistency of MLLM-generated responses with visual inputs, by leveraging a pixel-level Grounding LLM equipped with reasoning and referring segmentation capabilities. This referencefree framework not only effectively handles rich-context scenarios, but also offers interpretability. To facilitate this, an innovative pipeline is accordingly designed for generating instruction-tuning data (R-Instruct), featuring richcontext descriptions, grounding masks, and hard negative samples. We further establish R 2 -HalBench, a new hallucination benchmark for MLLMs, which, unlike previous benchmarks, encompasses real-world, rich-context descriptions from 18 MLLMs with high-quality annotations, spanning diverse object-, attribute-, and relationship-level details. VBackChecker outperforms prior complex frameworks and achieves state-of-the-art performance on R^2 -HalBench, even rivaling GPT-4o’s capabilities in hallucination detection. It also surpasses prior methods in the pixel-level grounding task, achieving over a 10% improvement.

AAAI Conference 2026 Conference Paper

SERL: Self-Examining Reinforcement Learning on Open-Domain

  • Weixuan Ou
  • Yanzhao Zheng
  • Shuoshuo Sun
  • Wei Zhang
  • Baohua Dong
  • Hangcheng Zhu
  • Ruohui Huang
  • Gang Yu

Reinforcement Learning (RL) has been shown to improve the capabilities of large language models (LLMs). However, applying RL to open-domain tasks faces two key challenges: (1) the inherent subjectivity of these tasks prevents the verifiable rewards as required by Reinforcement Learning with Verifiable Rewards (RLVR); (2) Reinforcement Learning from Human Feedback (RLHF) relies on external reward mechanisms. To overcome these limitations, we propose Self-Examining Reinforcement Learning (SERL), a novel self-improving framework where the LLM serves as both Actor and Judge. SERL introduces two synergistic reward mechanisms without any external signals. On the one hand, to improve the Actor's capability, we derive rewards from Copeland-style pairwise comparison judgments across a group of generated responses. On the other hand, a self-consistency reward that encourages coherent judgments is proposed to improve the Judge's reliability. This refinement strengthens the Judge, consequently generating a more robust training signal for the Actor. Experiments show that our method outperforms existing self-improvement training methods. SERL improves the LC win rate of Qwen3-8B on AlpacaEval 2.0 from 52.37% to 59.90%. To the best of our knowledge, our method achieves state-of-the-art performance among self-improving approaches. Furthermore, it achieves a performance comparable to significantly larger models like Qwen3-32B, demonstrating superior effectiveness and robustness on open-domain tasks.

AAAI Conference 2026 Conference Paper

TokenPowerBench: Benchmarking the Power Consumption of LLM Inference

  • Chenxu Niu
  • Wei Zhang
  • Jie Li
  • Yongjian Zhao
  • Tongyang Wang
  • Xi Wang
  • Yong Chen

Large language model (LLM) services now answer billions of queries per day, and industry reports show that inference, not training, accounts for more than 90% of total power consumption. However, existing benchmarks focus on either training/fine-tuning or performance of inference and provide little support for power consumption measurement and analysis of inference. We introduce TokenPowerBench, the first lightweight and extensible benchmark designed for LLM-inference power consumption studies. The benchmark combines a declarative configuration interface covering model choice, prompt set, and inference engine, a measurement layer that captures GPU-, node-, and system-level power without specialized power meters, and a phase-aligned metrics pipeline that attributes energy to the prefill and decode stages of every request. These elements make it straightforward to explore the power consumed by an LLM inference run; furthermore, by varying batch size, context length, parallelism strategy and quantization, users can quickly assess how each setting affects joules per token and other energy-efficiency metrics. We evaluate TokenPowerBench on four of the most widely used model series (Llama, Falcon, Qwen, and Mistral). Our experiments cover from 1 billion parameters up to the frontier-scale Llama3-405B model. Furthermore, we release TokenPowerBench as open source to help users to measure power consumption, forecast operating expenses, and meet sustainability targets when deploying LLM services.

AAAI Conference 2026 Conference Paper

UniFit: Towards Universal Virtual Try-on with MLLM-Guided Semantic Alignment

  • Wei Zhang
  • Yeying Jin
  • Xin Li
  • Yan Zhang
  • Xiaofeng Cong
  • Cong Wang
  • Fengcai Qiao
  • Zhichao Lian

Image-based virtual try-on (VTON) aims to synthesize photorealistic images of a person wearing specified garments. Despite significant progress, building a universal VTON framework that can flexibly handle diverse and complex tasks remains a major challenge. Recent methods explore multi-task VTON frameworks guided by textual instructions, yet they still face two key limitations: (1) semantic gap between text instructions and reference images, and (2) data scarcity in complex scenarios. To address these challenges, we propose UniFit, a universal VTON framework driven by a Multimodal Large Language Model (MLLM). Specifically, we introduce an MLLM-Guided Semantic Alignment Module (MGSA), which integrates multimodal inputs using an MLLM and a set of learnable queries. By imposing a semantic alignment loss, MGSA captures cross-modal semantic relationships and provides coherent and explicit semantic guidance for the generative process, thereby reducing the semantic gap. Moreover, by devising a two-stage progressive training strategy with a self-synthesis pipeline, UniFit is able to learn complex tasks from limited data. Extensive experiments show that UniFit not only supports a wide range of VTON tasks, including multi-garment and model-to-model try-on, but also achieves state-of-the-art performance.

EAAI Journal 2025 Journal Article

A cross dual branch guidance network for salient object detection

  • Yiru Wei
  • Zhiliang Zhu
  • Hai Yu
  • Wei Zhang

The effective integration of multi-level contextual information is crucial for deep learning-based salient object detection. However, most existing approaches either adopt the parallel structure or the progressive structure to predict salient objects, which still face challenges in consistently and accurately detecting salient objects of varying scales. In this paper, we propose a novel cross dual branch guidance network to effectively extract the rich semantic features and gradually enhance the saliency map scale-by-scale. Concretely, the parallel branch is guided by the progressive branch to obtain coarse location information of salient objects. In turn, the progressive branch is able to obtain uniform semantics and rich details to enhance saliency map with the guidance of the parallel branch. To obtain the dynamic receptive field, a dynamic sampling module (DSM) is introduced, which can dynamically adjust the sampling positions such that the spatial details of salient objects in complex scenes can be well recognized. In addition, we design a global context module (GCM) to explore the correlation between different parts of salient object or different salient objects, which is favorable for improving the completeness of saliency map. Experiments on five released benchmark datasets demonstrate the effectiveness and superiority of our proposed approach against other state-of-the-art methods.

EAAI Journal 2025 Journal Article

A hybrid architecture of sparse convolutional neural network-transformer for enhanced spatial-geometric feature learning in surface reconstruction

  • Mingyang Li
  • Wei Zhang
  • Yanyan Liu
  • Xiang Feng
  • Changsong Liu
  • Yimeng Fan
  • Lixue Xu

Learning-based methods have garnered significant attention in indoor scene reconstruction tasks. However, researchers have often overlooked the crucial role of the surface prediction stage. Our study specifically focuses on this phase. According to our experiments and analysis, this phase primarily addresses spatial voxel occupancy and geometric structure maintenance. Simple structural designs are insufficient to effectively solve these problems. To address these challenges, we propose a hybrid model, which combines the strengths of Convolution Neural Networks and Transformer architectures for fine reconstruction. Additionally, we introduce several new techniques, including the Sparse Positional Attention mechanism, Sparse Channel Decoding Block, and Mixed Feature Fusion mechanism. These techniques, leveraging the characteristics of sparse computation, enhance feature utilization in both spatial and channel dimensions. With limited training and testing resources, our network achieves optimal results on the ScanNet dataset, improving precision and F-score by 2. 1% and 1. 6%, respectively, and reducing the Chamfer distance to 0. 055 m. To our knowledge, our model is the first use of hybrid structures in the surface prediction phase of an indoor scene reconstruction task. Moreover, we hope that our design and analysis can provide a new paradigm for task network design in this phase.

JBHI Journal 2025 Journal Article

A Knowledge-Guided Multi-modal Neural Network for Breast Cancer Molecular Subtyping

  • Jinlin Ye
  • Yuhan Liu
  • Shangjie Ren
  • Changjun Wang
  • Yidong Zhou
  • Liang Yang
  • Wei Zhang

Precise determination of HER2 subtype is essential for selecting appropriate targeted therapies in breast cancer. However, current HER2 assessment methods remain dependent on invasive tissue biopsies, which are limited by tumor heterogeneity and sampling bias. To address these challenges, this paper proposes a knowledge-guided multi-modal neural network (KMNet) for non-invasive HER2 subtyping by integrating clinical data and ultrasound images. KMNet introduces a Graph-based Clinical Feature encoder (GCF), which constructs a causal graph among clinical indicators based on medical knowledge and extracts high-order feature relationships via the Graph Convolutional Network (GCN). Meanwhile, the Convolutional Neural Network (CNN) and Vision Transformer (ViT)-based hybrid image encoder (CVUIF) captures both local details (calcifications and blood flow) and global dependencies between intra- and peritumoral regions. In addition, the Reduced Dimensional Fusion (RDF) module integrates key information from clinical graph features, ultrasound image features, and structured clinical data to construct a unified multi-modal representation for downstream HER2 subtyping task. Experiments were conducted on the private datasets (HER2USC) and the public datasets (BCW, BCa and SIIM-ISIC). Experimental results demonstrate that KMNet outperformed other reported state-ofthe- art multi-modal algorithms in HER2 subtyping task, offering strong potential for clinical decision support in breast cancer treatment.

EAAI Journal 2025 Journal Article

A multiple knowledge-based evolutionary algorithm for sparse large-scale multi-objective problems

  • Wanting Yang
  • Jianchang Liu
  • Yuanchao Liu
  • Wei Zhang
  • Tianzi Zheng

In real-world applications, sparse large-scale multi-objective optimization problems (LSMOPs) are prevalent. Sparse LSMOPs are the LSMOPs characterized by Pareto-optimal solutions with sparse decision variables. Nonetheless, limited attention has been given to developing general-purpose algorithms for sparse LSMOPs. Existing approaches primarily focus on detecting sparsity, whereas achieving optimal results requires accurate detection of sparse distributions and simultaneous optimization of non-zero variables. Therefore, this paper proposes a multiple knowledge-based sparse large-scale multi-objective evolutionary algorithm. Building on genetic algorithms, the proposed algorithm employs a two-layer encoding scheme, develops a knowledge-driven evolution strategy for optimizing binary vectors, and introduces an association optimization method for optimizing real vectors. The performance of the proposed algorithm is assessed using both benchmark tests and real-world applications. Experimental results demonstrate that the proposed algorithm is competitively effective in solving sparse LSMOPs.

TMLR Journal 2025 Journal Article

AEAP: A Reinforcement Learning Actor Ensemble Algorithm with Adaptive Pruning

  • Wei Zhang
  • Guni Sharon

Actor ensemble reinforcement learning methods have shown promising performance on dense-reward continuous control tasks. However, they exhibit three primary limitations: (1) diversity collapse when using a shared replay buffer, often necessitating carefully tuned regularization terms; (2) computational overhead from maintaining multiple actors; and (3) analytically intractable policy gradients when using stochastic policies in ensembles, requiring approximations that may compromise performance. To address this third limitation, we restrict the ensemble to deterministic policies and propose Actor Ensemble with Adaptive Pruning (AEAP), a multi-actor deterministic policy gradient algorithm that tackles the remaining limitations through a two-stage approach. First, to alleviate diversity collapse, AEAP employs dual-randomized actor selection that decorrelates exploration and learning by randomly choosing different actors for both environment interaction and policy update. This approach also removes reliance on explicit regularization. Second, when convergence to homogeneous policies still occurs over time, computational efficiency is further achieved through adaptive dual-criterion pruning, which progressively removes underperforming or redundant actors based on critic-estimated value and action-space similarity. Although AEAP introduces four additional hyperparameters compared to TD3 (a baseline single-actor deterministic policy gradient algorithm), we provide two domain-agnostic parameter configurations that perform robustly across environments without requiring tuning. AEAP achieves superior or competitive asymptotic performance compared to baselines across six dense-reward MuJoCo tasks. On sparse-reward Fetch benchmarks, AEAP outperforms deterministic policy gradient methods but falls short of SAC (a baseline stochastic policy gradient algorithm) on one of three tasks. When compared to fixed-size multi-actor baselines, AEAP reduces wall-clock time without sacrificing performance, establishing it as an efficient and reliable actor ensemble variant.

TIST Journal 2025 Journal Article

Aspect-Enhanced Explainable Recommendation with Multi-modal Contrastive Learning

  • Hao Liao
  • Shuo Wang
  • Hao Cheng
  • Wei Zhang
  • Jiwei Zhang
  • Mingyang Zhou
  • Kezhong Lu
  • Rui Mao

Explainable recommender systems ( ERS ) aim to enhance users’ trust in the systems by offering personalized recommendations with transparent explanations. This transparency provides users with a clear understanding of the rationale behind the recommendations, fostering a sense of confidence and reliability in the system’s outputs. Generally, the explanations are presented in a familiar and intuitive way, which is in the form of natural language, thus enhancing their accessibility to users. Recently, there has been an increasing focus on leveraging reviews as a valuable source of rich information in both modeling user-item preferences and generating textual interpretations, which can be performed simultaneously in a multi-task framework. Despite the progress made in these review-based recommendation systems, the integration of implicit feedback derived from user-item interactions and user-written text reviews has yet to be fully explored. To fill this gap, we propose a model named SERMON (A s pect-enhanced E xplainable R ecommendation with M ulti-modal C o ntrast Lear n ing). Our model explores the application of multimodal contrastive learning to facilitate reciprocal learning across two modalities, thereby enhancing the modeling of user preferences. Moreover, our model incorporates the aspect information extracted from the review, which provides two significant enhancements to our tasks. Firstly, the quality of the generated explanations is improved by incorporating the aspect characteristics into the explanations generated by a pre-trained model with controlled textual generation ability. Secondly, the commonly used user-item interactions are transformed into user-item-aspect interactions, which we refer to as interaction triple, resulting in a more nuanced representation of user preference. To validate the effectiveness of our model, we conduct extensive experiments on three real-world datasets. The experimental results show that our model outperforms state-of-the-art baselines, with a 2.0% improvement in prediction accuracy and a substantial 24.5% enhancement in explanation quality for the TripAdvisor dataset.

EAAI Journal 2025 Journal Article

Attention based network for real-time road drivable area, lane line detection and scene identification

  • Feng You
  • Yi Xie
  • Siyi Zhang
  • Hao Chen
  • Haiwei Wang
  • Wei Zhang
  • Jianrong Liu

The detection of road drivable areas and lane lines is considered a fundamental component of autonomous driving systems. However, most existing approaches handle these tasks independently, and multi-task networks frequently neglect the inherent correlation between them while failing to differentiate various lane line types. In practice, the delineation of drivable regions is strongly influenced by both lane line characteristics and contextual street scenes. To address these limitations, a novel multi-task network—Real-time Road Drivable Area, Lane Line Detection, and Scene Identification Network (RLSNet)—is proposed. This network is designed to perform simultaneous segmentation of drivable areas, detection of lane lines, and classification of road scenes. Drivable area estimation is optimized through the integration of lane and scene cues, guided by traffic regulations. A Residual Network (ResNet)-based backbone is employed, enhanced with Bidirectional Fusion Attention (BFA) for feature encoding. This is followed by a decoder incorporating a Feature Aggregation Module (FAM) to enable effective semantic–spatial fusion. Lane line detection is further refined using a Bilateral Up-Sampling Decoder (BUSD), while scene understanding is enhanced via a Scene Classification Module (SCM). Extensive experiments conducted on the challenging Berkeley DeepDrive 100K(BDD100K) dataset have demonstrated that RLSNet achieves high accuracy in both drivable area and lane line detection by leveraging the mutual guidance of lane and scene information. Furthermore, the network maintains real-time inference speed at 93 frames per second (FPS), striking a practical balance between semantic fidelity and computational efficiency for real-world deployment. The implementation code has been made publicly available at: https: //github. com/033186ZSY/RLSNet-master.

IROS Conference 2025 Conference Paper

Capsizing-Guided Trajectory Optimization for Autonomous Navigation with Rough Terrain

  • Wei Zhang
  • Yinchuan Wang
  • Wangtao Lu
  • Pengyu Zhang
  • Xiang Zhang
  • Yue Wang
  • Chaoqun Wang

It is a challenging task for ground robots to autonomously navigate in harsh environments due to the presence of non-trivial obstacles and uneven terrain. This requires trajectory planning that balances safety and efficiency. The primary challenge is to generate a feasible trajectory that prevents robot from tip-over while ensuring effective navigation. In this paper, we propose a capsizing-aware trajectory planner (CAP) to achieve trajectory planning on the uneven terrain. The tip-over stability of the robot on rough terrain is analyzed. Based on the tip-over stability, we define the traversable orientation, which indicates the safe range of robot orientations. This orientation is then incorporated into a capsizing-safety constraint for trajectory optimization. We employ a graph-based solver to compute a robust and feasible trajectory while adhering to the capsizing-safety constraint. Extensive simulation and real-world experiments validate the effectiveness and robustness of the proposed method. The results demonstrate that CAP outperforms existing state-of-the-art approaches, providing enhanced navigation performance on uneven terrains.

AIIM Journal 2025 Journal Article

Causal inference model for accurate medical diagnosis in Coronary Artery Bypass Graft operation

  • Qiyi Zhang
  • Wei Zhang
  • Qiang Li
  • Yunpeng Bai
  • Weizhi Nie
  • Keliang Xie

Coronary Artery Bypass Grafting (CABG) is the most commonly performed cardiac surgery. Predicting postoperative complication risks for patients undergoing CABG is crucial for medical professionals. Considering the susceptibility of traditional models to confounding factors and the scarcity of medical data, it is necessary to design a model that can truly capture the cause-and-effect relationship between the disease and its underlying causes and achieve high accuracy even with limited data. In this paper, a novel Causal Inference Operation Risk Predictor (CIORP) is proposed. We construct a Structural Causal Model (SCM) that demonstrates how two confounders influence the model’s predictions. Then we utilize the backdoor adjustment strategy to control potential confounders from pre-operative information and non-causal intraoperative data. In parallel, capitalizing on few-shot learning techniques, we initiate pre-training using categories with ample samples to extract essential features. Subsequently, we fine-tuned our model on sparse sets of labeled data, facilitating accurate predictions in scenarios with limited annotated samples. The experimental outcomes demonstrate that our model surpasses most existing methods in the internal Electronic Health Record (EHR) of CABG patients, effectively predicting low cardiac output, new-onset atrial fibrillation, perioperative myocardial infarction, and cardiac arrest or ventricular fibrillation post-operation. Our work effectively mitigates the impact of confounding factors, allowing the model to make accurate predictions with minimal medical data.

AAAI Conference 2025 Conference Paper

Coherency Improved Explainable Recommendation via Large Language Model

  • Shijie Liu
  • Ruixin Ding
  • Weihai Lu
  • Jun Wang
  • Mo Yu
  • Xiaoming Shi
  • Wei Zhang

Explainable recommender systems are designed to elucidate the explanation behind each recommendation, enabling users to comprehend the underlying logic. Previous works perform rating prediction and explanation generation in a multi-task manner. However, these works suffer from incoherence between predicted ratings and explanations. To address the issue, we propose a novel framework that employs a large language model (LLM) to generate a rating, transforms it into a rating vector, and finally generates an explanation based on the rating vector and user-item information. Moreover, we propose utilizing publicly available LLMs and pre-trained sentiment analysis models to automatically evaluate the coherence without human annotations. Extensive experimental results on three datasets of explainable recommendation show that the proposed framework is effective, outperforming state-of-the-art baselines with improvements of 7.3% in explainability and 4.4% in text quality.

NeurIPS Conference 2025 Conference Paper

Connectome-Based Modelling Reveals Orientation Maps in the Drosophila Optic Lobe

  • Jia Nuo Liew
  • Shenghan Lin
  • Bowen Chen
  • Xiaowei Zhu
  • Wei Zhang
  • Xiaolin Hu

The ability to extract oriented edges from visual input is a core computation across animal vision systems. Orientation maps, long associated with the layered architecture of the mammalian visual cortex, systematically organise neurons by their preferred edge orientation. Despite lacking cortical structures, the Drosophila melanogaster brain contains feature-selective neurons and exhibits complex visual detection capacity, raising the question of whether map-like vision representations can emerge without cortical infrastructure. We integrate a complete fruit fly brain connectome with biologically grounded spiking neuron models to simulate neuroprocessing in the fly visual system. By driving the network with oriented stimuli and analysing downstream responses, we show that coherent orientation maps can emerge from purely connectome-constrained dynamics. These results suggest that species of independent origin could evolve similar visual structures.

YNIMG Journal 2025 Journal Article

Deep diffusion MRI template (DDTemplate): A novel deep learning groupwise diffusion MRI registration method for brain template creation

  • Junyi Wang
  • Xi Zhu
  • Wei Zhang
  • Mubai Du
  • William M. Wells
  • Lauren J O’Donnell
  • Fan Zhang

Diffusion MRI (dMRI) is an advanced imaging technique that enables in-vivo tracking of white matter fiber tracts and estimates the underlying cellular microstructure of brain tissues. Groupwise registration of dMRI data from multiple individuals is an important task for brain template creation and investigation of inter-subject brain variability. However, groupwise registration is a challenging task due to the uniqueness of dMRI data that include multi-dimensional, orientation-dependent signals that describe not only the strength but also the orientation of water diffusion in brain tissues. Deep learning approaches have shown successful performance in standard subject-to-subject dMRI registration. However, no deep learning methods have yet been proposed for groupwise dMRI registration. . In this work, we propose Deep Diffusion MRI Template (DDTemplate), which is a novel deep-learning-based method building upon the popular VoxelMorph framework to take into account dMRI fiber tract information. DDTemplate enables joint usage of whole-brain tissue microstructure and tract-specific fiber orientation information to ensure alignment of white matter fiber tracts and whole brain anatomical structures. We propose a novel deep learning framework that simultaneously trains a groupwise dMRI registration network and generates a population brain template. During inference, the trained model can be applied to register unseen subjects to the learned template. We compare DDTemplate with several state-of-the-art registration methods and demonstrate superior performance on dMRI data from multiple cohorts (adolescents, young adults, and elderly adults) acquired from different scanners. Furthermore, as a testbed task, we perform a between-population analysis to investigate sex differences in the brain, using the popular Tract-Based Spatial Statistics (TBSS) method that relies on groupwise dMRI registration. We find that using DDTemplate can increase the sensitivity in population difference detection, showing the potential of our method's utility in real neuroscientific applications.

EAAI Journal 2025 Journal Article

Denoising diffusion probabilistic model-enabled data augmentation method for intelligent machine fault diagnosis

  • Pengcheng Zhao
  • Wei Zhang
  • Xiaoshan Cao
  • Xiang Li

Bearing fault is one of the main causes of rotating machinery failure. However, collecting sufficient failure data proves challenging in real-world industrial environments due to their complexity. Owing to this constraint, the majority of current methods fail to accurately identify fault types with limited data, thereby impeding timely maintenance efforts. To address this issue, we propose a bearing fault diagnosis method utilizing diffusion model data enhancement in this study. Following data collection, we employ the continuous wavelet transform to convert various one-dimensional vibration data into two-dimensional time series graphs. Subsequently, these feature graphs are partitioned into sets of fault samples. Then, the diffusion model is employed to augment the training samples. These augmented samples are then fed into the convolutional neural network for fault diagnosis, and their diagnostic accuracy is compared with that of the original dataset. The comparative analysis demonstrates that the data augmentation technique, founded on the diffusion model, enhances fault diagnosis accuracy within a small sample training dataset. Ultimately, the efficacy of the proposed method is validated utilizing the Paderborn bearing dataset and juxtaposed against alternative data augmentation techniques. The findings indicate that the diffusion model-based data augmentation method outperforms other techniques in terms of both training accuracy and stability, particularly in scenarios with small sample.

NeurIPS Conference 2025 Conference Paper

EAReranker: Efficient Embedding Adequacy Assessment for Retrieval Augmented Generation

  • Dongyang Zeng
  • Yaping Liu
  • Wei Zhang
  • Shuo Zhang
  • Xinwang Liu
  • Binxing Fang

With the increasing adoption of Retrieval-Augmented Generation (RAG) systems for knowledge-intensive tasks, ensuring the adequacy of retrieved documents has become critically important for generation quality. Traditional reranking approaches face three significant challenges: substantial computational overhead that scales with document length, dependency on plain text that limits application in sensitive scenarios, and insufficient assessment of document value beyond simple relevance metrics. We propose EAReranker, an efficient embedding-based adequacy assessment framework that evaluates document utility for RAG systems without requiring access to original text content. The framework quantifies document adequacy through a comprehensive scoring methodology considering verifiability, coverage, completeness and structural aspects, providing interpretable adequacy classifications for downstream applications. EAReranker employs a Decoder-Only Transformer architecture that introduces embedding dimension expansion method and bin-aware weighted loss, designed specifically to predict adequacy directly from embedding vectors. Our comprehensive evaluation across four public benchmarks demonstrates that EAReranker achieves competitive performance with state-of-the-art plaintext rerankers while maintaining constant memory usage ($\sim$550MB) regardless of input length and processing 2-3x faster than traditional approaches. The semantic bin adequacy prediction accuracy of 92. 85\% LACC@10 and 86. 12\% LACC@25 demonstrates its capability to effectively filter out inadequate documents that could potentially mislead or adversely impact RAG system performance, thereby ensuring only high-utility information serves as generation context. These results establish EAReranker as an efficient and practical solution for enhancing RAG system performance through improved context selection while addressing the computational and privacy challenges of existing methods.

EAAI Journal 2025 Journal Article

Gate-guided spatial-channel reconstruction network: An efficient lightweight framework for steel surface defect detection

  • Wei Zhang

To address the limitations in You Only Look Once version 8 (YOLOv8) for steel surface defect detection, including insufficient generalization of image enhancement, constrained feature representation capability in core modules, and poor adaptability of the loss function to scale variations and sample imbalance, this paper proposes the Gate-guided Spatial-channel Reconstruction Network, an efficient and lightweight improved network. Contrast-Limited Adaptive Histogram Equalization (CLAHE) is introduced to enhance local image details and contrast while reducing noise impact. The Gating Block and the Spatial-channel Reconstruction Block are designed to replace the original C2f (cross-stage partial bottleneck with two convolutions) module in YOLOv8, thereby enhancing feature representation capability and efficiency. The loss function is optimized using Wise-IoU (WIoU) and Slide Loss (SlideLoss) to improve convergence and robustness. The proposed network was evaluated on the Northeastern University Surface Defect Detection (NEU-DET) dataset (200 × 200 pixels) and the Chinese Academy of Sciences Defect Detection (GC10-DET) dataset (2048 × 1000 pixels). It demonstrated high detection accuracy, achieving the mean Average Precision at 50 % (mAP50) of 84. 7 % and 79. 4 %, respectively. Furthermore, the network maintains low complexity with only 3. 6 million parameters and achieves a high detection speed of up to 154 frames per second (FPS). The Gate-guided Spatial-channel Reconstruction Network effectively detects surface defects on hot-rolled steel, achieving state-of-the-art detection accuracy. It successfully meets the requirements for precise and real-time steel surface defect detection under resource-constrained industrial conditions.

NeurIPS Conference 2025 Conference Paper

Guiding Cross-Modal Representations with MLLM Priors via Preference Alignment

  • Pengfei Zhao
  • Rongbo Luan
  • Wei Zhang
  • Peng Wu
  • Sifeng He

Despite Contrastive Language–Image Pre-training (CLIP)'s remarkable capability to retrieve content across modalities, a substantial modality gap persists in its feature space. Intriguingly, we discover that off-the-shelf MLLMs (Multimodal Large Language Models) demonstrate powerful inherent modality alignment properties. While recent MLLM-based retrievers with unified architectures partially mitigate this gap, their reliance on coarse modality alignment mechanisms fundamentally limits their potential. In this work, We introduce MAPLE (Modality-Aligned Preference Learning for Embeddings), a novel framework that leverages the fine-grained alignment priors inherent in MLLM to guide cross-modal representation learning. MAPLE formulates the learning process as reinforcement learning with two key components: (1) Automatic preference data construction using off-the-shelf MLLM, and (2) a new Relative Preference Alignment (RPA) loss, which adapts Direct Preference Optimization (DPO) to the embedding learning setting. Experimental results show that our preference-guided alignment achieves substantial gains in fine-grained cross-modal retrieval, underscoring its effectiveness in handling nuanced semantic distinctions.

NeurIPS Conference 2025 Conference Paper

Improving the Euclidean Diffusion Generation of Manifold Data by Mitigating Score Function Singularity

  • Zichen Liu
  • Wei Zhang
  • Tiejun Li

Euclidean diffusion models have achieved remarkable success in generative modeling across diverse domains, and they have been extended to manifold cases in recent advances. Instead of explicitly utilizing the structure of special manifolds as studied in previous works, in this paper we investigate direct sampling of the Euclidean diffusion models for general manifold-structured data. We reveal the multiscale singularity of the score function in the ambient space, which hinders the accuracy of diffusion-generated samples. We then present an elaborate theoretical analysis of the singularity structure of the score function by decomposing it along the tangential and normal directions of the manifold. To mitigate the singularity and improve the sampling accuracy, we propose two novel methods: (1) Niso-DM, which reduces the scale discrepancies in the score function by utilizing a non-isotropic noise, and (2) Tango-DM, which trains only the tangential component of the score function using a tangential-only loss function. Numerical experiments demonstrate that our methods achieve superior performance on distributions over various manifolds with complex geometries.

AAAI Conference 2025 Conference Paper

In2NeCT: Inter-class and Intra-class Neural Collapse Tuning for Semantic Segmentation of Imbalanced Remote Sensing Images

  • Junao Shen
  • Qiyun Hu
  • Tian Feng
  • Xinyu Wang
  • Hui Cui
  • Sensen Wu
  • Wei Zhang

Remote sensing images (RSIs) are frequently characterized by multi-scale inter-class objects and inconsistently distributed objects due to scene limitations, which would cause a significant data imbalance challenging the corresponding semantic segmentation. Recent methods have leveraged various deep learning techniques to capture high-quality representations for RSI semantic segmentation, but are hardly capable of addressing the afore-mentioned challenge given their limited explorations towards the mechanisms behind the representations. The recently discovered Neural Collapse (NC) phenomenon in computer vision models suggests the simplex equiangular tight frame (ETF) as the optimal representation structure, which has motivated us to observe that the optimal structure of last-layer representations is disrupted and inter-class representations for minor classes tend to become closer to each other beacuse of data imbalance. To address these issues, we propose Inter-class and Intra-class Neural Collapse Tuning (In2NeCT) to optimize the representations that satisfy the simplex ETF, which facilitates the discrimination of inter-class representations and the coherence of intra-class representations. Extensive experiments on three datasets demonstrate that our In2NeCT consistently leads to significant improvements in performance and outperforms the state-of-the-art methods.

YNIMG Journal 2025 Journal Article

Individualized functional connectivity markers for motor and mood symptoms of Parkinson’s disease

  • Louisa Dahmani
  • Yan Bai
  • Wei Zhang
  • Jianxun Ren
  • Shiyi Li
  • Qingyu Hu
  • Xiaoxuan Fu
  • Jianjun Ma

Parkinson's disease (PD) is a complex neurological disorder characterized by many motor and non-motor symptoms. While most studies focus on the motor symptoms of the disease, it is important to identify pathophysiological markers that underlie different facets of the disease. In this case-control study, we sought to discover reliable, individualized functional connectivity markers associated with both motor and mood symptoms of PD. In order to obtain precise functional connectivity measurements, we extensively sampled 166 patients with PD and 51 healthy control participants using functional MRI and characterized functional connectomes at the level of each individual participant. Multiple linear regressions were performed to model the relationship between functional connectivity markers and PD symptoms. We found that a model consisting of 44 functional connections predicted both motor (r = 0.21, p = 0.006) and mood symptoms (depression: r = 0.23, p = 0.006; anxiety: r = 0.21, p = 0.006). Two sets of connections contributed differentially to these predictions. Between-network connections, mainly connecting the sensorimotor and visual large-scale functional networks, substantially contributed to the prediction of motor measures, while within-network connections in the insula and sensorimotor network contributed more so to mood prediction. The middle to posterior insula region played a particularly important role in predicting depression and anxiety scores. We successfully replicated and generalized our findings in three independent PD datasets. Taken together, our findings indicate that sensorimotor and visual network markers are indicative of PD brain pathology, and that distinct subsets of markers are associated with motor and mood symptoms of PD.

NeurIPS Conference 2025 Conference Paper

JavisGPT: A Unified Multi-modal LLM for Sounding-Video Comprehension and Generation

  • Kai Liu
  • Jungang Li
  • Yuchong Sun
  • Shengqiong Wu
  • jianzhang gao
  • Daoan Zhang
  • Wei Zhang
  • Sheng Jin

This paper presents JavisGPT, the first unified multimodal large language model (MLLM) for joint audio-video (JAV) comprehension and generation. JavisGPT has a concise encoder-LLM-decoder architecture, which has a SyncFusion module for spatio-temporal audio-video fusion and synchrony-aware learnable queries to bridge a pretrained JAV-DiT generator. This design enables temporally coherent video-audio understanding and generation from multimodal instructions. We design an effective three-stage training pipeline consisting of multimodal pretraining, audio-video fine-tuning, and large-scale instruction-tuning, to progressively build multimodal comprehension and generation from existing vision-language models. For instruction tuning, we construct JavisInst-Omni, a high-quality instruction dataset with over 200K GPT-4o-curated audio-video-text dialogues that cover diverse and multi-level comprehension and generation scenarios. On JAV comprehension and generation benchmarks, our experiments show that JavisGPT outperforms existing MLLMs, particularly in complex and temporally synchronized settings.

EAAI Journal 2025 Journal Article

Knowledge extraction and alignment for mine ventilation: A knowledge graph construction framework based on large language models

  • Jinyang Dong
  • Junqiao Li
  • Yucheng Li
  • Wei Zhang
  • Zhitao Zhang
  • Chenyang Guo
  • Yu Dang
  • Mei Chen

Mine ventilation corpora contain fragmented entities, attributes, and rule-based knowledge, marked by heterogeneous expressions, implicit structures, and ambiguous terminology. These characteristics hinder systematic modeling and intelligent utilization. To address these challenges, we propose an automated extraction and semantic alignment method based on large language models (LLMs), aiming to construct a high-quality knowledge graph (KG) tailored for mine ventilation. We design an ontology-driven extraction framework for three textual sources — regulations, books, and websites — using prompt engineering and few-shot strategies to extract information, domain knowledge, entity attributes, and rule-based relations in a unified way. For entity alignment, we develop a dual-filtering mechanism that integrates semantic similarity and structural adjacency, and leverage large language models (LLMs) for verification, enabling high-confidence alignment of entities and predicates. We propose a rule-path modeling strategy using the structure “subject entity (with condition) - predicate - object entity (with condition), ” integrating multi-source triples into conditional rule chains. These are mapped into a graph database to support structured knowledge representation. Under few-shot conditions, the extraction accuracy of entity attributes and rule-based relations reached 94% and 99%, respectively. The final alignment of entities and predicates, manually verified, achieved 100% precision. The resulting knowledge graph (KG) comprises 58, 358 entities and 63, 630 edges, demonstrating strong semantic consistency and structural integrity. It provides structured knowledge support for risk warning, question answering (QA), and intelligent decision-making in mine ventilation systems as part of intelligent mining applications.

NeurIPS Conference 2025 Conference Paper

KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation

  • Jiajun Shi
  • Jian Yang
  • Jiaheng Liu
  • Xingyuan Bu
  • Jiangjie Chen
  • Junting Zhou
  • Kaijing Ma
  • Zhoufutu Wen

Recent advancements in large language models (LLMs) underscore the need for more comprehensive evaluation methods to accurately assess their reasoning capabilities. Existing benchmarks are often domain-specific and thus cannot fully capture an LLM’s general reasoning potential. To address this limitation, we introduce the Knowledge Orthogonal Reasoning Gymnasium (KORGym), a dynamic evaluation platform inspired by KOR-Bench and Gymnasium. KORGym offers over fifty games in either textual or visual formats and supports interactive, multi-turn assessments with reinforcement learning scenarios. Using KORGym, we conduct extensive experiments on 19 LLMs and 8 VLMs, revealing consistent reasoning patterns within model families and demonstrating the superior performance of closed-source models. Further analysis examines the effects of modality, reasoning strategies, reinforcement learning techniques, and response length on model performance. We expect KORGym to become a valuable resource for advancing LLM reasoning research and developing evaluation methodologies suited to complex, interactive environments.

NeurIPS Conference 2025 Conference Paper

LIFEBENCH: Evaluating Length Instruction Following in Large Language Models

  • Wei Zhang
  • Zhenhong Zhou
  • Kun Wang
  • Junfeng Fang
  • Rongwu Xu
  • Yuanhe Zhang
  • Rui Wang
  • Ge Zhang

While large language models (LLMs) can solve PhD-level reasoning problems over long context inputs, they still struggle with a seemingly simpler task: following explicit length instructions —e. g. , write a 10, 000-word novel. Additionally, models often generate far too short outputs, terminate prematurely, or even refuse the request. Existing benchmarks focus primarily on evaluating generations quality, but often overlook whether the generations meet length constraints. To this end, we introduce Length Instruction Following Evaluation Benchmark (LIFEBench) to comprehensively evaluate LLMs' ability to follow length instructions across diverse tasks and a wide range of specified lengths. LIFEBench consists of 10, 800 instances across 4 task categories in both English and Chinese, covering length constraints ranging from 16 to 8192 words. We evaluate 26 widely-used LLMs and find that most models reasonably follow short-length instructions but deteriorate sharply beyond a certain threshold. Surprisingly, almost all models fail to reach the vendor-claimed maximum output lengths in practice, as further confirmed by our evaluations extending up to 32K words. Even long-context LLMs, despite their extended input-output windows, counterintuitively fail to improve length-instructions following. Notably, Reasoning LLMs outperform even specialized long-text generation models, achieving state-of-the-art length following. Overall, LIFEBench uncovers fundamental limitations in current LLMs' length instructions following ability, offering critical insights for future progress.

AAAI Conference 2025 Conference Paper

PerReactor: Offline Personalised Multiple Appropriate Facial Reaction Generation

  • Hengde Zhu
  • Xiangyu Kong
  • Weicheng Xie
  • Xin Huang
  • Xilin He
  • Lu Liu
  • Linlin Shen
  • Wei Zhang

In dyadic human-human interactions, individuals may express multiple different facial reactions in response to the same/similar behaviours expressed by their conversational partners depending on their personalised behaviour patterns. As a result, frequently-employed reconstruction loss-based strategies lead the training of previous automatic facial reaction generation (FRG) models to not only suffer from the 'one-to-many mapping' problem, but also fail to comprehensively consider the quality of the generated facial reactions. Besides, none of them considered such personalised behaviour patterns in generating facial reactions. In this paper, we propose the first adversarial FRG model training strategy which jointly learns appropriateness and realism discriminators to provide comprehensive task-specific supervision for training the target facial reaction generators, and reformulates the 'one-to-many (facial reactions) mapping' training problem as a 'one-to-one (distribution) mapping' training task, i.e., the FRG model is trained to output a distribution representing multiple appropriate/plausible facial reaction from each input human behaviour. In addition, our approach also serves as the first offline FRG approach that considers personalised behaviour patterns in generating of target individuals' facial reactions. Experiments show that our PerReactor not only largely outperformed all existing offline solutions for generating more appropriate, diverse and realistic facial reactions, but also is the first approach that can effectively generate personalised appropriate facial reactions.

JMLR Journal 2025 Journal Article

Reinforcement Learning for Infinite-Dimensional Systems

  • Wei Zhang
  • Jr-Shin Li

Interest in reinforcement learning (RL) for large-scale systems, comprising extensive populations of intelligent agents interacting with heterogeneous environments, has surged significantly across diverse scientific domains in recent years. However, the large-scale nature of these systems often leads to high computational costs or reduced performance for most state-of-the-art RL techniques. To address these challenges, we propose a novel RL architecture and derive effective algorithms to learn optimal policies for arbitrarily large systems of agents. In our formulation, we model such systems as parameterized control systems defined on an infinite-dimensional function space. We then develop a moment kernel transform that maps the parameterized system and the value function into a reproducing kernel Hilbert space. This transformation generates a sequence of finite-dimensional moment representations for the RL problem, organized into a filtrated structure. Leveraging this RL filtration, we develop a hierarchical algorithm for learning optimal policies for the infinite-dimensional parameterized system. To enhance the algorithm's efficiency, we incorporate early stopping at each hierarchy, demonstrating the fast convergence property of the algorithm through the construction of a convergent spectral sequence. The performance and efficiency of the proposed algorithm are validated using practical examples in engineering and quantum systems. [abs] [ pdf ][ bib ] &copy JMLR 2025. ( edit, beta )

NeurIPS Conference 2025 Conference Paper

Spike4DGS: Towards High-Speed Dynamic Scene Rendering with 4D Gaussian Splatting via a Spike Camera Array

  • Qinghong Ye
  • Yiqian Chang
  • Jianing Li
  • Haoran Xu
  • Xuan Wang
  • Wei Zhang
  • Yonghong Tian
  • Peixi Peng

Spike camera with high temporal resolution offers a new perspective on high-speed dynamic scene rendering. Most existing rendering methods rely on Neural Radiance Fields (NeRF) or 3D Gaussian Splatting (3DGS) for static scenes using a monocular spike camera. However, these methods struggle with dynamic motion, while a single camera suffers from limited spatial coverage, making it challenging to reconstruct fine details in high-speed scenes. To address these problems, we propose Spike4DGS, the first high-speed dynamic scene rendering framework with 4D Gaussian Splatting using spike camera arrays. Technically, we first build a multi-view spike camera array to validate our solution, then establish both synthetic and real-world multi-view spike-based reconstruction datasets. Then, we design a multi-view spike-based dense initialization module that obtains dense point clouds and camera poses from continuous spike streams. Finally, we propose a spike-pixel synergy constraint supervision to optimize Spike4DGS, incorporating both rendered image quality loss and dynamic spatiotemporal spike loss. The results show that our Spike4DGS outperforms state-of-the-art methods in terms of novel view rendering quality on both synthetic and real-world datasets. More details are available at https: //github. com/Qinghongye/Spike4DGS.

AAAI Conference 2025 Conference Paper

STAIR: Manipulating Collaborative and Multimodal Information for E-Commerce Recommendation

  • Cong Xu
  • Yunhang He
  • Jun Wang
  • Wei Zhang

While the mining of modalities is the focus of most multimodal recommendation methods, we believe that how to fully utilize both collaborative and multimodal information is pivotal in e-commerce scenarios where, as clarified in this work, the user behaviors are rarely determined entirely by multimodal features. In order to combine the two distinct types of information, some additional challenges are encountered: 1) Modality erasure: Vanilla graph convolution, which proves rather useful in collaborative filtering, however erases multimodal information; 2) Modality forgetting: Multimodal information tends to be gradually forgotten as the recommendation loss essentially facilitates the learning of collaborative information. To this end, we propose a novel approach named STAIR, which employs a novel stepwise graph convolution to enable a co-existence of collaborative and multimodal information in e-commerce recommendation. Besides, it starts with the raw multimodal features as an initialization, and the forgetting problem can be significantly alleviated through constrained embedding updates. As a result, STAIR achieves state-of-the-art recommendation performance on three public e-commerce datasets with minimal computational and memory costs.

IROS Conference 2025 Conference Paper

Tracking Highly Dynamic Humanoid Motion with Dynamic IMU Measurement Fusion

  • Jeronimo Cox
  • Wei Zhang
  • Tomonari Furukawa

Inertial sensing estimation methods allows human motion tracking in the absence of optical tracking and joint encoders, but the methods are rather developed for quasistatic motion due to the limited motion capability of humanoids. This paper presents a new method that tracks highly dynamic motion using Inertial Measurement Unit (IMU) measurements. Unlike conventional methods dependent on quasistatic motion for inclination correction with the measured gravity vector, the proposed method uses accelerometers to correct the rotational rate. This is achieved by placing sensors on the ends of links, and converting the acceleration measured at the ends to angular rate based on centrifugal forces. Measuring human motions of low and high intensities is used to identify any strengths and weaknesses of the proposed method with different applications. The proposed technique maintains an acceptable error for both quasistatic and highly dynamic motions and can be used to accurately visualize measured motions.

EAAI Journal 2025 Journal Article

Underwater moving target detection and tracking based on enhanced you only look once and deep simple online and realtime tracking strategy

  • Bing Sun
  • Wei Zhang
  • Cheng Xing
  • Yingyao Li

Addressing the challenges posed by factors like light attenuation, turbid water quality, large target scale changes, and dense distribution in underwater moving target images, we propose an enhanced strategy combining YOLOv5 (You Only Look Once version 5) and Deep SORT (Simple Online and Realtime Tracking) for target detection and tracking. Initially, various image enhancement techniques are employed to improve detection performance, while a Deep Convolutional Generative Adversarial Networks (DCGAN) approach augments the small underwater image dataset. Subsequently, we enhance the YOLOv5 method by integrating a convolutional attention module into the detection network to enhance target salience in large-scale scenes. Additionally, a tiny target detection head is introduced to enhance the detector’s capability in adapting to scale changes, particularly for small objects. Finally, the Deep SORT algorithm is integrated for tracking detected targets. Experimental comparisons with state-of-the-art methods validate the efficacy of the proposed approach. Our method achieves a detection accuracy of 94. 3% on a self-made underwater fish target dataset, with a detection speed of 71FPS and robust tracking performance. Moreover, its applicability extends to various underwater targets, demonstrating superior performance.

JBHI Journal 2025 Journal Article

Video Object Segmentation with Optimal Frame Auto-selection Based on Prior Knowledge for Midbrain Assessment in Transcranial Ultrasound

  • Xinyi Wang
  • Sai Kit LAM
  • Hongyu KANG
  • Yu Sun
  • Chao HOU
  • Shuai Li
  • Xin Sun
  • Fangxian LI

Transcranial sonography (TCS) provides a non-invasive means of assessing movement disorders such as Parkinson's disease (PD). However, current TCS-based evaluations rely heavily on manual operation by experienced physicians, making the process time-consuming and physician-dependent. For the first time, we aimed to develop a hybrid pipeline for real-time video object segmentation (VOS) and automatic optimal frame selection. Eighty-three standardized TCS real-time data comprising 1, 992 midbrain frames from Beijing Tiantan Hospital were collected. We adopted three state-of-the-art VOS models (STCN, RDE-VOS, and XMEM) and incorporated anatomical priors to guide optimal frame selection. Specifically, we leveraged the anatomical trend of midbrain morphology to estimate the midbrain radius at the optimal frame and selected the frame where the VOS-segmented midbrain best matched this estimate. The XMEM-based pipeline achieved high segmentation performance (Jaccard: 0. 85, Boundary Accuracy: 0. 95, Dice: 0. 92) and optimal frame selection (Distance: 4. 87; Jaccard: 0. 92), with efficiency (51. 05 FPS, 0. 56 s/patient, 661. 55 MB). Subgroup analyses confirmed robustness across image quality and PD conditions. Assessment of a junior physician's selection suggests potential to reduce the expertise gap in optimal frame selection. The proposed hybrid pipeline offers an automated tool for midbrain assessment using TCS, which may help reduce physicians' workload and minimize subjectivity, particularly supporting junior physicians in mitigating the expertise-demanding nature of TCS. This approach may serve as a foundation for more promising TCS-based assessments in the future, contributing to broader adoption of non-invasive ultrasound techniques in PD evaluation.

EAAI Journal 2024 Journal Article

An effective method for small object detection in low-resolution images

  • Rudong Jing
  • Wei Zhang
  • Yanyan Liu
  • Wenlin Li
  • Yuming Li
  • Changsong Liu

Having a tiny scale and few identifiable features, small objects are particularly difficult to detect, especially when the image resolution is not high. Further, large-scale variation across object instances places a burden on small object detection. In view of the insufficient detection ability of YOLOF (You Only Look One-level Feature) detector for small objects at low resolution, we propose an effective method, named YOLOFs, to boost small object detection precision in the case of object instances scale variation. The proposed three key modules (feature fusion module, visual perception module and feature encoding module) expand the receptive field of the network and greatly improve small object detection precision. Extensive experiments on the COCO, VOC, and Fire datasets prove the effectiveness of our method while keeping the detector simple and accurate. Without bells and whistles, our method outperforms YOLOF 5. 1 AP (average precision), 6. 2 AP, and 4. 9 AP on small object detection at 704 × 704, 608 × 608, and 512 × 512 resolutions, respectively, and it also achieves better performance on the VOC and Fire datasets. Our code will be made publicly available.

YNICL Journal 2024 Journal Article

Association between clinical features and decreased degree centrality and variability in dynamic functional connectivity in the obsessive–compulsive disorder

  • Changjun Teng
  • Wei Zhang
  • Da Zhang
  • XiaoMeng Shi
  • Xin Wu
  • Huifen Qiao
  • Chengbin Guan
  • Xiao Hu

Neuroimaging studies have indicated widespread brain structural and functional disruptions in patients with obsessive-compulsive disorder (OCD). However, the underlying mechanism of these changes remains unclear. A total of 45 patients with OCD and 42 healthy controls (HC) were enrolled. The study investigated local degree centrality (DC) abnormalities and employed abnormal regions of DC as seeds to investigate variability in dynamic functional connectivity (dFC) in the whole brain using a sliding window approach to analyze resting-state functional magnetic resonance imaging. The relationship between abnormal DC and dFC as well as the clinical features of OCD were examined using correlation analysis. Our findings suggested decreased DC in the bilateral thalamus, bilateral precuneus, and bilateral cuneus in OCD patients and a nominally negative correlation between the DC value in the thalamus and illness severity measured using the Yale-Brown Obsessive Compulsive Scale (Y-BOCS). In addition, seed-based dFC analysis showed that compared to measurements in the HC, the patients had decreased dFC variability between the left thalamus and the left cuneus and right lingual gyrus, and between the bilateral cuneus and bilateral postcentral gyrus, and a nominally positive correlation between the duration of illness and dFC variability between the left cuneus and left postcentral gyrus. These results indicated that OCD patients had decreased hub importance in the bilateral thalamus and cuneus throughout the entire brain. This reduction was associated with impaired coupling with dynamic function in the visual cortex and sensorimotor network and provided novel insights into the neurophysiological mechanisms underlying OCD.

AAAI Conference 2024 Conference Paper

CGMGM: A Cross-Gaussian Mixture Generative Model for Few-Shot Semantic Segmentation

  • Junao Shen
  • Kun Kuang
  • Jiaheng Wang
  • Xinyu Wang
  • Tian Feng
  • Wei Zhang

Few-shot semantic segmentation (FSS) aims to segment unseen objects in a query image using a few pixel-wise annotated support images, thus expanding the capabilities of semantic segmentation. The main challenge lies in extracting sufficient information from the limited support images to guide the segmentation process. Conventional methods typically address this problem by generating single or multiple prototypes from the support images and calculating their cosine similarity to the query image. However, these methods often fail to capture meaningful information for modeling the de facto joint distribution of pixel and category. Consequently, they result in incomplete segmentation of foreground objects and mis-segmentation of the complex background. To overcome this issue, we propose the Cross Gaussian Mixture Generative Model (CGMGM), a novel Gaussian Mixture Models~(GMMs)-based FSS method, which establishes the joint distribution of pixel and category in both the support and query images. Specifically, our method initially matches the feature representations of the query image with those of the support images to generate and refine an initial segmentation mask. It then employs GMMs to accurately model the joint distribution of foreground and background using the support masks and the initial segmentation mask. Subsequently, a parametric decoder utilizes the posterior probability of pixels in the query image, by applying the Bayesian theorem, to the joint distribution, to generate the final segmentation mask. Experimental results on PASCAL-5i and COCO-20i datasets demonstrate our CGMGM's effectiveness and superior performance compared to the state-of-the-art methods.

JBHI Journal 2024 Journal Article

DAI-Net: Dual Adaptive Interaction Network for Coordinated Medication Recommendation

  • Xin Zou
  • Xiao He
  • Xiao Zheng
  • Wei Zhang
  • Jiajia Chen
  • Chang Tang

Medication recommendation is a productive task for AI-driven healthcare systems, which can assist clinicians in prescribing judicious and effective treatments. However, existing medication recommendation methods omit two key pieces of information: Coarse-grained interaction information between distinct types of symptoms in a patient's medical history and corresponding medication representations can serve as attention for predicting the current medication combinations of the patient. Fine-grained interaction information between medication substructure representations and different types of symptoms can facilitate the construction of molecular-level disentangled medication representations. To address this dilemma, we propose a novel D ual A daptive I nteraction Net work (DAI-Net), which encodes comprehensive interaction knowledge between patients' multifaceted health records and medication molecules to improve the performance of medication recommendation and heighten interpretability of the model. Specifically, we design a symptom-aware medication matching module to extract coordinated associations between patient symptoms and medication molecules, coarse-grained interaction learning. The medication embeddings are utilized to transform patient-medication matching properties into a symptom-substructure matching matrix for fine-grained interaction. The patient's Longitudinal representation is employed as a query to decode both symptom-medication and symptom-substructure matching information for coordinated medication representation. DAI-Net is an end-to-end recommendation model. Extensive experiments on the real-world EHR datasets, i. e. , the public benchmark MIMIC-III, MIMIC-IV, and eICU, demonstrate that the proposed DAI-Net achieves competitive performance compared to other state-of-the-art ones, with an average improvement of 1. 8%, 2. 1% in Jaccard on MIMIC-III and -IV dataset.

AIIM Journal 2024 Journal Article

Detecting mental and physical disorders using multi-task learning equipped with knowledge graph attention network

  • Wei Zhang
  • Ling Kong
  • Soobin Lee
  • Yan Chen
  • Guangxu Zhang
  • Hao Wang
  • Min Song

Mental and physical disorders (MPD) are inextricably linked in many medical cases; psychosomatic diseases can be induced by mental concerns and psychological discomfort can ensue from physiological diseases. However, existing medical informatics studies focus on identifying mental or physical disorders from a unilateral perspective. Consequently, no existing domain knowledge base, corpus, or detection modeling approach considers mental as well as physical aspects concurrently. This paper proposes a joint modeling approach to detect MPD. First, we crawl through online medical consultation records of patients from websites and build an MPD knowledge ontology by extracting the core conceptual features of the text. Based on the ontology, an MPD knowledge graph containing 12, 673 nodes and 82, 195 relations is obtained using term matching with a domain thesaurus of each concept. Subsequently, an MPD corpus with fine-grained severities (None, Mild, Moderate, Severe, Dangerous) and 8909 records is constructed by formulating MPD classification criteria and a data annotation process under the guidance of domain experts. Taking the knowledge graph and corpus as the dataset, we design a multi-task learning model to detect the MPD severity, in which a knowledge graph attention network (KGAT) is embedded to better extract knowledge features. Experiments are performed to demonstrate the effectiveness of our model. Furthermore, we employ ontology-based and centrality-based methods to discover additional potential inferred knowledge, which can be captured by KGAT so as to improve the prediction performance and interpretability of our model. Our dataset has been made publicly available, so it can be further used as a medical informatics reference in the fields of psychosomatic medicine, psychiatrics, physical co-morbidity, and so on.

AAMAS Conference 2024 Conference Paper

Distance-Aware Attentive Framework for Multi-Agent Collaborative Perception in Presence of Pose Error

  • Binyu Zhao
  • Wei Zhang
  • Zhaonian Zou

Multi-agent collaborative perception exchanges information to promote holistic perception, especially for remote and invisible areas that are limited by detection range and occlusion. Due to imperfect localization in practice, it usually suffers from pose estimation error, which can cause spatial message misalignment and performance degradation. Unlike most existing methods using additional module or procedure to correct pose error, we propose a novel framework, DistAtt, to suppress pose error and mine useful information simultaneously. It mainly consists of distance-aware feature sampling and cross-agent feature aggregation. The former utilizes diverse pooling kernels to downsample the intermediate features to different multiple granularities, and the latter utilizes specially designed attention mechanism to learn the most critical information. Furthermore, it adopts compensation strategy for more stable optimization. Experimental results show that DistAtt significantly suppresses the effect of localization noise and achieves outperformed performance when pose error exists.

NeurIPS Conference 2024 Conference Paper

Enhancing LLM’s Cognition via Structurization

  • Kai Liu
  • Zhihang Fu
  • Chao Chen
  • Wei Zhang
  • Rongxin Jiang
  • Fan Zhou
  • Yaowu Chen
  • Yue Wu

When reading long-form text, human cognition is complex and structurized. While large language models (LLMs) process input contexts through a causal and sequential perspective, this approach can potentially limit their ability to handle intricate and complex inputs effectively. To enhance LLM’s cognition capability, this paper presents a novel concept of context structurization. Specifically, we transform the plain, unordered contextual sentences into well-ordered and hierarchically structurized elements. By doing so, LLMs can better grasp intricate and extended contexts through precise attention and information-seeking along the organized structures. Extensive evaluations are conducted across various model architectures and sizes (including a series of auto-regressive LLMs as well as BERT-like masking models) on a diverse set of NLP tasks (e. g. , context-based question-answering, exhaustive hallucination evaluation, and passage-level dense retrieval). Empirical results show consistent and significant performance gains afforded by a single-round structurization. In particular, we boost the open-sourced LLaMA2-70B model to achieve comparable performance against GPT-3. 5-Turbo as the halluci- nation evaluator. Besides, we show the feasibility of distilling advanced LLMs’ language processing abilities to a smaller yet effective StruXGPT-7B to execute structurization, addressing the practicality of our approach. Code is available at https: //github. com/alibaba/struxgpt.

AAAI Conference 2024 Conference Paper

Gaussian Process Neural Additive Models

  • Wei Zhang
  • Brian Barr
  • John Paisley

Deep neural networks have revolutionized many fields, but their black-box nature also occasionally prevents their wider adoption in fields such as healthcare and finance where interpretable and explainable models are required. The recent development of Neural Additive Models (NAMs) poses a major step in the direction of interpretable deep learning for tabular datasets. In this paper, we propose a new subclass of NAMs that utilize a single-layer neural network construction of the Gaussian process via random Fourier features, which we call Gaussian Process Neural Additive Models (GP-NAM). GP-NAMs have the advantage of a convex objective function and number of trainable parameters that grows linearly with feature dimensions. It suffers no loss in performance compared with deeper NAM approaches because GPs are well-suited to learning complex non-parametric univariate functions. We demonstrate the performance of GP-NAM on several tabular datasets, showing that it achieves comparable performance in both classification and regression tasks with a massive reduction in the number of parameters.

NeurIPS Conference 2024 Conference Paper

Graph-enhanced Optimizers for Structure-aware Recommendation Embedding Evolution

  • Cong Xu
  • Jun Wang
  • Jianyong Wang
  • Wei Zhang

Embedding plays a key role in modern recommender systems because they are virtual representations of real-world entities and the foundation for subsequent decision-making models. In this paper, we propose a novel embedding update mechanism, Structure-aware Embedding Evolution (SEvo for short), to encourage related nodes to evolve similarly at each step. Unlike GNN (Graph Neural Network) that typically serves as an intermediate module, SEvo is able to directly inject graph structural information into embedding with minimal computational overhead during training. The convergence properties of SEvo along with its potential variants are theoretically analyzed to justify the validity of the designs. Moreover, SEvo can be seamlessly integrated into existing optimizers for state-of-the-art performance. Particularly SEvo-enhanced AdamW with moment estimate correction demonstrates consistent improvements across a spectrum of models and datasets, suggesting a novel technical route to effectively utilize graph structural information beyond explicit GNN modules.

TIST Journal 2024 Journal Article

Improving Faithfulness and Factuality with Contrastive Learning in Explainable Recommendation

  • Haojie Zhuang
  • Wei Zhang
  • Weitong Chen
  • Jian Yang
  • Quan Z. Sheng

Recommender systems have become increasingly important in navigating the vast amount of information and options available in various domains. By tailoring and personalizing recommendations to user preferences and interests, these systems improve the user experience, efficiency, and satisfaction. With a growing demand for transparency and understanding of recommendation outputs, explainable recommender systems have gained growing attention in recent years. Additionally, as user reviews could be considered the rationales behind why the user likes (or dislikes) the products, generating informative and reliable reviews alongside recommendations has thus emerged as a research focus in explainable recommendation. However, the model-generated reviews might contain factually inconsistent contents (i.e., the hallucination issue), which would thus compromise the recommendation rationales. To address this issue, we propose a contrastive learning framework to improve the faithfulness and factuality in explainable recommendation in this article. We further develop different strategies of generating positive and negative examples for contrastive learning, such as back-translation or synonym substitution for positive examples, and editing positive examples or utilizing model-generated texts for negative examples. Our proposed method optimizes the model to distinguish faithful explanations (i.e., positive examples) and unfaithful ones with factual errors (i.e., negative examples), which thus drives the model to generate faithful reviews as explanations while avoiding inconsistent contents. Extensive experiments and analysis on three benchmark datasets show that our proposed model outperforms other review generation baselines in faithfulness and factuality. In addition, the proposed contrastive learning component could be easily incorporated into other explainable recommender systems in a plug-and-play manner.

ICML Conference 2024 Conference Paper

Interpreting and Improving Large Language Models in Arithmetic Calculation

  • Wei Zhang
  • Chaoqun Wan
  • Yonggang Zhang 0003
  • Yiu-ming Cheung
  • Xinmei Tian 0001
  • Xu Shen 0001
  • Jieping Ye

Large language models (LLMs) have demonstrated remarkable potential across numerous applications and have shown an emergent ability to tackle complex reasoning tasks, such as mathematical computations. However, even for the simplest arithmetic calculations, the intrinsic mechanisms behind LLMs remains mysterious, making it challenging to ensure reliability. In this work, we delve into uncovering a specific mechanism by which LLMs execute calculations. Through comprehensive experiments, we find that LLMs frequently involve a small fraction ($<$5%) of attention heads, which play a pivotal role in focusing on operands and operators during calculation processes. Subsequently, the information from these operands is processed through multi-layer perceptrons (MLPs), progressively leading to the final solution. These pivotal heads/MLPs, though identified on a specific dataset, exhibit transferability across different datasets and even distinct tasks. This insight prompted us to investigate the potential benefits of selectively fine-tuning these essential heads/MLPs to boost the LLMs’ computational performance. We empirically find that such precise tuning can yield notable enhancements on mathematical prowess, without compromising the performance on non-mathematical tasks. Our work serves as a preliminary exploration into the arithmetic calculation abilities inherent in LLMs, laying a solid foundation to reveal more intricate mathematical tasks.

AAAI Conference 2024 Conference Paper

LaneGraph2Seq: Lane Topology Extraction with Language Model via Vertex-Edge Encoding and Connectivity Enhancement

  • Renyuan Peng
  • Xinyue Cai
  • Hang Xu
  • Jiachen Lu
  • Feng Wen
  • Wei Zhang
  • Li Zhang

Understanding road structures is crucial for autonomous driving. Intricate road structures are often depicted using lane graphs, which include centerline curves and connections forming a Directed Acyclic Graph (DAG). Accurate extraction of lane graphs relies on precisely estimating vertex and edge information within the DAG. Recent research highlights Transformer-based language models' impressive sequence prediction abilities, making them effective for learning graph representations when graph data are encoded as sequences. However, existing studies focus mainly on modeling vertices explicitly, leaving edge information simply embedded in the network. Consequently, these approaches fall short in the task of lane graph extraction. To address this, we introduce LaneGraph2Seq, a novel approach for lane graph extraction. It leverages a language model with vertex-edge encoding and connectivity enhancement. Our serialization strategy includes a vertex-centric depth-first traversal and a concise edge-based partition sequence. Additionally, we use classifier-free guidance combined with nucleus sampling to improve lane connectivity. We validate our method on prominent datasets, nuScenes and Argoverse 2, showcasing consistent and compelling results. Our LaneGraph2Seq approach demonstrates superior performance compared to state-of-the-art techniques in lane graph extraction.

AAAI Conference 2024 Conference Paper

Latent Space Editing in Transformer-Based Flow Matching

  • Vincent Tao Hu
  • Wei Zhang
  • Meng Tang
  • Pascal Mettes
  • Deli Zhao
  • Cees Snoek

This paper strives for image editing via generative models. Flow Matching is an emerging generative modeling technique that offers the advantage of simple and efficient training. Simultaneously, a new transformer-based U-ViT has recently been proposed to replace the commonly used UNet for better scalability and performance in generative modeling. Hence, Flow Matching with a transformer backbone offers the potential for scalable and high-quality generative modeling, but their latent structure and editing ability are as of yet unknown. Hence, we adopt this setting and explore how to edit images through latent space manipulation. We introduce an editing space, which we call u-space, that can be manipulated in a controllable, accumulative, and composable manner. Additionally, we propose a tailored sampling solution to enable sampling with the more efficient adaptive step-size ODE solvers. Lastly, we put forth a straightforward yet powerful method for achieving fine-grained and nuanced editing using text prompts. Our framework is simple and efficient, all while being highly effective at editing images while preserving the essence of the original content. Our code will be publicly available at https://taohu.me/lfm/

NeurIPS Conference 2024 Conference Paper

Locally Private and Robust Multi-Armed Bandits

  • Xingyu Zhou
  • Wei Zhang

We study the interplay between local differential privacy (LDP) and robustness to Huber corruption and possibly heavy-tailed rewards in the context of multi-armed bandits (MABs). We consider two different practical settings: LDP-then-Corruption (LTC) where each user's locally private response might be further corrupted during the data collection process, and Corruption-then-LDP (CTL) where each user's raw data may be corrupted such that the LDP mechanism will only be applied to the corrupted data. To start with, we present the first tight characterization of the mean estimation error in high probability under both LTC and CTL settings. Leveraging this new result, we then present an almost tight characterization (up to log factor) of the minimax regret in online MABs and sub-optimality in offline MABs under both LTC and CTL settings, respectively. Our theoretical results in both settings are also corroborated by a set of systematic simulations. One key message in this paper is that LTC is a more difficult setting that leads to a worse performance guarantee compared to the CTL setting (in the minimax sense). Our sharp understanding of LTC and CTL also naturally allows us to give the first tight performance bounds for the most practical setting where corruption could happen both before and after the LDP mechanism. As an important by-product, we also give the first correct and tight regret bound for locally private and heavy-tailed online MABs, i. e. , without Huber corruption, by identifying a fundamental flaw in the state-of-the-art.

TIST Journal 2024 Journal Article

Mitigating the Impact of Inaccurate Feedback in Dynamic Learning-to-Rank: A Study of Overlooked Interesting Items

  • Chenhao Zhang
  • Weitong Chen
  • Wei Zhang
  • Miao Xu

Dynamic Learning-to-Rank (DLTR) is a method of updating a ranking policy in real time based on user feedback, which may not always be accurate. Although previous DLTR work has achieved fair and unbiased DLTR under inaccurate feedback, they face the tradeoff between fairness and user utility and also have limitations in the setting of feeding items. Existing DLTR works improve ranking utility by eliminating bias from inaccurate feedback on observed items, but the impact of another pervasive form of inaccurate feedback, overlooked or ignored interesting items, remains unclear. For example, users may browse the rankings too quickly to catch interesting items or miss interesting items because the snippets are not optimized enough. This phenomenon raises two questions: (i) Will overlooked interesting items affect the ranking results? and (ii) Is it possible to improve utility without sacrificing fairness if these effects are eliminated? These questions are particularly relevant for small and medium-sized retailers who are just starting out and may have limited data, leading to the use of inaccurate feedback to update their models. In this article, we find that inaccurate feedback in the form of overlooked interesting items has a negative impact on DLTR performance in terms of utility. To address this, we treat the overlooked interesting items as noise and propose a novel DLTR method, the Co-teaching Rank (CoTeR), that has good utility and fairness performance when inaccurate feedback is present in the form of overlooked interesting items. Our solution incorporates a co-teaching-based component with a customized loss function and data sampling strategy, as well as a mean pooling strategy to further accommodate newly added products without historical data. Through experiments, we demonstrate that CoTeR not only enhances utilities but also preserves ranking fairness and can smoothly handle newly introduced items.

JBHI Journal 2024 Journal Article

Predictive Modeling for Hospital Readmissions for Patients With Heart Disease: An Updated Review From 2012–2023

  • Wei Zhang
  • Weihan Cheng
  • Koichi Fujiwara
  • Richard Evans
  • Chengyan Zhu

Hospital readmissions are a major concern for healthcare leaders, policy makers, and patients, resulting in adverse health outcomes and imposing an increased burden on hospital resources. This review aims to synthesize existing literature on predictive models focused on patients diagnosed with heart disease, which is known for its high readmission rates. Seven databases (i. e. , Web of Science, Scopus, PubMed, ProQuest, Ovid, Cochrane Library and EBSCO) were consulted resulting in the inclusion of 56 eligible studies. Among these, 44 focused on model development, 7 on model validation, 4 on model improvement, and 1 on model implementation. Data were extracted on readmission types, data sources, modeling methods, and predictors, while assessments were conducted to analyze the quality of the studies. Findings showed that readmission types were significantly influenced by policy decisions, data predominantly originated from hospitals, and the prevalent modeling methods used were regression and single-layer machine learning techniques. The most important clinical predictors were related to comorbidities and complications, while the key demographic predictors were age and race. The study found that, despite advancements during the last decade, several limitations exist in current research, particularly in addressing attrition bias and handling missing data. Future research should, therefore, focus on optimizing readmission types, enhancing model generalization, using interpretable models, and emphasizing model implementation.

NeurIPS Conference 2024 Conference Paper

Real-time Stereo-based 3D Object Detection for Streaming Perception

  • Changcai Li
  • Zonghua Gu
  • Gang Chen
  • Libo Huang
  • Wei Zhang
  • Huihui Zhou

The ability to promptly respond to environmental changes is crucial for the perception system of autonomous driving. Recently, a new task called streaming perception was proposed. It jointly evaluate the latency and accuracy into a single metric for video online perception. In this work, we introduce StreamDSGN, the first real-time stereo-based 3D object detection framework designed for streaming perception. StreamDSGN is an end-to-end framework that directly predicts the 3D properties of objects in the next moment by leveraging historical information, thereby alleviating the accuracy degradation of streaming perception. Further, StreamDSGN applies three strategies to enhance the perception accuracy: (1) A feature-flow-based fusion method, which generates a pseudo-next feature at the current moment to address the misalignment issue between feature and ground truth. (2) An extra regression loss for explicit supervision of object motion consistency in consecutive frames. (3) A large kernel backbone with a large receptive field for effectively capturing long-range spatial contextual features caused by changes in object positions. Experiments on the KITTI Tracking dataset show that, compared with the strong baseline, StreamDSGN significantly improves the streaming average precision by up to 4. 33%. Our code is available at https: //github. com/weiyangdaren/streamDSGN-pytorch.

AAAI Conference 2024 Conference Paper

Referred by Multi-Modality: A Unified Temporal Transformer for Video Object Segmentation

  • Shilin Yan
  • Renrui Zhang
  • Ziyu Guo
  • Wenchao Chen
  • Wei Zhang
  • Hongyang Li
  • Yu Qiao
  • Hao Dong

Recently, video object segmentation (VOS) referred by multi-modal signals, e.g., language and audio, has evoked increasing attention in both industry and academia. It is challenging for exploring the semantic alignment within modalities and the visual correspondence across frames. However, existing methods adopt separate network architectures for different modalities, and neglect the inter-frame temporal interaction with references. In this paper, we propose MUTR, a Multi-modal Unified Temporal transformer for Referring video object segmentation. With a unified framework for the first time, MUTR adopts a DETR-style transformer and is capable of segmenting video objects designated by either text or audio reference. Specifically, we introduce two strategies to fully explore the temporal relations between videos and multi-modal signals. Firstly, for low-level temporal aggregation before the transformer, we enable the multi-modal references to capture multi-scale visual cues from consecutive video frames. This effectively endows the text or audio signals with temporal knowledge and boosts the semantic alignment between modalities. Secondly, for high-level temporal interaction after the transformer, we conduct inter-frame feature communication for different object embeddings, contributing to better object-wise correspondence for tracking along the video. On Ref-YouTube-VOS and AVSBench datasets with respective text and audio references, MUTR achieves +4.2% and +8.7% J&F improvements to state-of-the-art methods, demonstrating our significance for unified multi-modal VOS. Code is released at https://github.com/OpenGVLab/MUTR.

YNIMG Journal 2024 Journal Article

Revealing the spatiotemporal brain dynamics of covert speech compared with overt speech: A simultaneous EEG-fMRI study

  • Wei Zhang
  • Muyun Jiang
  • Kok Ann Colin Teo
  • Raghavan Bhuvanakantham
  • LaiGuan Fong
  • Wei Khang Jeremy Sim
  • Zhiwei Guo
  • Chuan Huat Vince Foo

Covert speech (CS) refers to speaking internally to oneself without producing any sound or movement. CS is involved in multiple cognitive functions and disorders. Reconstructing CS content by brain-computer interface (BCI) is also an emerging technique. However, it is still controversial whether CS is a truncated neural process of overt speech (OS) or involves independent patterns. Here, we performed a word-speaking experiment with simultaneous EEG-fMRI. It involved 32 participants, who generated words both overtly and covertly. By integrating spatial constraints from fMRI into EEG source localization, we precisely estimated the spatiotemporal dynamics of neural activity. During CS, EEG source activity was localized in three regions: the left precentral gyrus, the left supplementary motor area, and the left putamen. Although OS involved more brain regions with stronger activations, CS was characterized by an earlier event-locked activation in the left putamen (peak at 262 ms versus 1170 ms). The left putamen was also identified as the only hub node within the functional connectivity (FC) networks of both OS and CS, while showing weaker FC strength towards speech-related regions in the dominant hemisphere during CS. Path analysis revealed significant multivariate associations, indicating an indirect association between the earlier activation in the left putamen and CS, which was mediated by reduced FC towards speech-related regions. These findings revealed the specific spatiotemporal dynamics of CS, offering insights into CS mechanisms that are potentially relevant for future treatment of self-regulation deficits, speech disorders, and development of BCI speech applications.

JBHI Journal 2024 Journal Article

S2VQ-VAE: Semi-Supervised Vector Quantised-Variational AutoEncoder for Automatic Evaluation of Trail Making Test

  • Zeshen Tang
  • Shiyu Tang
  • Haoran Wang
  • Renren Li
  • Xiaochen Zhang
  • Wei Zhang
  • Xiao Yuan
  • Yaning Zang

Background: Computer-aided detection of cognitive impairment garnered increasing attention, offering older adults in the community access to more objective, ecologically valid, and convenient cognitive assessments using multimodal sensing technology on digital devices. Methodology: In this study, we aimed to develop an automated method for screening cognitive impairment, building on paper- and electronic TMTs. We proposed a novel deep representation learning approach named Semi-Supervised Vector Quantised-Variational AutoEncoder (S2VQ-VAE). Within S2VQ-VAE, we incorporated intra- and inter-class correlation losses to disentangle class-related factors. These factors were then combined with various real-time obtainable features (including demographic, time-related, pressure-related, and jerk-related features) to create a robust feature engineering block. Finally, we identified the light gradient boosting machine as the optimal classifier. The experiments were conducted on a dataset collected from older adults in the community. Results: The experimental results showed that the proposed multi-type feature fusion method outperformed the conventional method used in paper-based TMTs and the existing VAE-based feature extraction in terms of screening performance. Conclusions: In conclusion, the proposed deep representation learning method significantly enhances the cognitive diagnosis capabilities of behavior-based TMTs and streamlines large-scale community-based cognitive impairment screening while reducing the workload of professional healthcare staff.

AAAI Conference 2024 Conference Paper

Symbolic Cognitive Diagnosis via Hybrid Optimization for Intelligent Education Systems

  • Junhao Shen
  • Hong Qian
  • Wei Zhang
  • Aimin Zhou

Cognitive diagnosis assessment is a fundamental and crucial task for student learning. It models the student-exercise interaction, and discovers the students' proficiency levels on each knowledge attribute. In real-world intelligent education systems, generalization and interpretability of cognitive diagnosis methods are of equal importance. However, most existing methods can hardly make the best of both worlds due to the complicated student-exercise interaction. To this end, this paper proposes a symbolic cognitive diagnosis~(SCD) framework to simultaneously enhance generalization and interpretability. The SCD framework incorporates the symbolic tree to explicably represent the complicated student-exercise interaction function, and utilizes gradient-based optimization methods to effectively learn the student and exercise parameters. Meanwhile, the accompanying challenge is that we need to tunnel the discrete symbolic representation and continuous parameter optimization. To address this challenge, we propose to hybridly optimize the representation and parameters in an alternating manner. To fulfill SCD, it alternately learns the symbolic tree by derivative-free genetic programming and learns the student and exercise parameters via gradient-based Adam. The extensive experimental results on various real-world datasets show the superiority of SCD on both generalization and interpretability. The ablation study verifies the efficacy of each ingredient in SCD, and the case study explicitly showcases how the interpretable ability of SCD works.

AAAI Conference 2023 Conference Paper

3D-TOGO: Towards Text-Guided Cross-Category 3D Object Generation

  • Zutao Jiang
  • Guansong Lu
  • Xiaodan Liang
  • Jihua Zhu
  • Wei Zhang
  • Xiaojun Chang
  • Hang Xu

This article has been updated and an error has been fixed in published paper. An Erratum to this article was published on 6 September 2023. Text-guided 3D object generation aims to generate 3D objects described by user-defined captions, which paves a flexible way to visualize what we imagined. Although some works have been devoted to solving this challenging task, these works either utilize some explicit 3D representations (e.g., mesh), which lack texture and require post-processing for rendering photo-realistic views; or require individual time-consuming optimization for every single case. Here, we make the first attempt to achieve generic text-guided cross-category 3D object generation via a new 3D-TOGO model, which integrates a text-to-views generation module and a views-to-3D generation module. The text-to-views generation module is designed to generate different views of the target 3D object given an input caption. prior-guidance, caption-guidance and view contrastive learning are proposed for achieving better view-consistency and caption similarity. Meanwhile, a pixelNeRF model is adopted for the views-to-3D generation module to obtain the implicit 3D neural representation from the previously-generated views. Our 3D-TOGO model generates 3D objects in the form of the neural radiance field with good texture and requires no time-cost optimization for every single caption. Besides, 3D-TOGO can control the category, color and shape of generated 3D objects with the input caption. Extensive experiments on the largest 3D object dataset (i.e., ABO) are conducted to verify that 3D-TOGO can better generate high-quality 3D objects according to the input captions across 98 different categories, in terms of PSNR, SSIM, LPIPS and CLIP-score, compared with text-NeRF and Dreamfields.

AAAI Conference 2023 Conference Paper

Adaptive Low-Precision Training for Embeddings in Click-Through Rate Prediction

  • Shiwei Li
  • Huifeng Guo
  • Lu Hou
  • Wei Zhang
  • Xing Tang
  • Ruiming Tang
  • Rui Zhang
  • Ruixuan Li

Embedding tables are usually huge in click-through rate (CTR) prediction models. To train and deploy the CTR models efficiently and economically, it is necessary to compress their embedding tables. To this end, we formulate a novel quantization training paradigm to compress the embeddings from the training stage, termed low-precision training (LPT). Also, we provide theoretical analysis on its convergence. The results show that stochastic weight quantization has a faster convergence rate and a smaller convergence error than deterministic weight quantization in LPT. Further, to reduce accuracy degradation, we propose adaptive low-precision training (ALPT) which learns the step size (i.e., the quantization resolution). Experiments on two real-world datasets confirm our analysis and show that ALPT can significantly improve the prediction accuracy, especially at extremely low bit width. For the first time in CTR models, we successfully train 8-bit embeddings without sacrificing prediction accuracy.

AAAI Conference 2023 Conference Paper

FlowFace: Semantic Flow-Guided Shape-Aware Face Swapping

  • Hao Zeng
  • Wei Zhang
  • Changjie Fan
  • Tangjie Lv
  • Suzhen Wang
  • Zhimeng Zhang
  • Bowen Ma
  • Lincheng Li

In this work, we propose a semantic flow-guided two-stage framework for shape-aware face swapping, namely FlowFace. Unlike most previous methods that focus on transferring the source inner facial features but neglect facial contours, our FlowFace can transfer both of them to a target face, thus leading to more realistic face swapping. Concretely, our FlowFace consists of a face reshaping network and a face swapping network. The face reshaping network addresses the shape outline differences between the source and target faces. It first estimates a semantic flow (i.e. face shape differences) between the source and the target face, and then explicitly warps the target face shape with the estimated semantic flow. After reshaping, the face swapping network generates inner facial features that exhibit the identity of the source face. We employ a pre-trained face masked autoencoder (MAE) to extract facial features from both the source face and the target face. In contrast to previous methods that use identity embedding to preserve identity information, the features extracted by our encoder can better capture facial appearances and identity information. Then, we develop a cross-attention fusion module to adaptively fuse inner facial features from the source face with the target facial attributes, thus leading to better identity preservation. Extensive quantitative and qualitative experiments on in-the-wild faces demonstrate that our FlowFace outperforms the state-of-the-art significantly.

YNIMG Journal 2023 Journal Article

Genome-wide association study of cerebellar white matter microstructure and genetic overlap with common brain disorders

  • Bang-Sheng Wu
  • Yi-Jun Ge
  • Wei Zhang
  • Shi-Dong Chen
  • Shi-Tong Xiang
  • Ya-Ru Zhang
  • Ya-Nan Ou
  • Yu-Chao Jiang

BACKGROUND: The cerebellum is recognized as being involved in neurocognitive and motor functions with communication with extra-cerebellar regions relying on the white matter integrity of the cerebellar peduncles. However, the genetic determinants of cerebellar white matter integrity remain largely unknown. METHODS: We conducted a genome-wide association analysis of cerebellar white matter microstructure using diffusion tensor imaging data from 25,415 individuals from UK Biobank. The integrity of cerebellar white matter microstructure was measured as fractional anisotropy (FA) and mean diffusivity (MD). Identification of independent genomic loci, functional annotation, and tissue and cell-type analysis were conducted with FUMA. The linkage disequilibrium score regression (LDSC) was used to calculate genetic correlations between cerebellar white matter microstructure and regional brain volumes and brain-related traits. Furthermore, the conditional/conjunctional false discovery rate (condFDR/conjFDR) framework was employed to identify the shared genetic basis between cerebellar white matter microstructure and common brain disorders. RESULTS: ) and 86 genes associated with cerebellar white matter microstructure. Further functional enrichment analysis implicated the involvement of GABAergic neurons and cholinergic pathways. Significant polygenetic overlap between cerebellar white matter tracts and their anatomically connected or adjacent brain regions was detected. In addition, we report the overall genetic correlation and specific loci shared between cerebellar white matter microstructural integrity and brain-related traits, including movement, cognitive, psychiatric, and cerebrovascular categories. CONCLUSIONS: Collectively, this study represents a step forward in understanding the genetics of cerebellar white matter microstructure and its shared genetic etiology with common brain disorders.

IJCAI Conference 2023 Conference Paper

Learning to Binarize Continuous Features for Neuro-Rule Networks

  • Wei Zhang
  • Yongxiang Liu
  • Zhuo Wang
  • Jianyong Wang

Neuro-Rule Networks (NRNs) emerge as a promising neuro-symbolic method, enjoyed by the ability to equate fully-connected neural networks with logic rules. To support learning logic rules consisting of boolean variables, converting input features into binary representations is required. Different from discrete features that could be directly transformed by one-hot encodings, continuous features need to be binarized based on some numerical intervals. Existing studies usually select the bound values of intervals based on empirical strategies (e. g. , equal-width interval). However, it is not optimal since the bounds are fixed and cannot be optimized to accommodate the ultimate training target. In this paper, we propose AutoInt, an approach that automatically binarizes continuous features and enables the intervals to be optimized with NRNs in an end-to-end fashion. Specifically, AutoInt automatically selects an interval for a given continuous feature in a soft manner to enable a differentiable learning procedure of interval-related parameters. Moreover, it introduces an additional soft K-means clustering loss to make the interval centres approach the original feature value distribution, thus reducing the risk of overfitting intervals. We conduct comprehensive experiments on public datasets and demonstrate the effectiveness of AutoInt in boosting the performance of NRNs.

AAAI Conference 2023 Conference Paper

Multi-Level Confidence Learning for Trustworthy Multimodal Classification

  • Xiao Zheng
  • Chang Tang
  • Zhiguo Wan
  • Chengyu Hu
  • Wei Zhang

With the rapid development of various data acquisition technologies, more and more multimodal data come into being. It is important to integrate different modalities which are with high-dimensional features for boosting final multimodal data classification task. However, existing multimodal classification methods mainly focus on exploiting the complementary information of different modalities, while ignoring the learning confidence during information fusion. In this paper, we propose a trustworthy multimodal classification network via multi-level confidence learning, referred to as MLCLNet. Considering that a large number of feature dimensions could not contribute to final classification performance but disturb the discriminability of different samples, we propose a feature confidence learning mechanism to suppress some redundant features, as well as enhancing the expression of discriminative feature dimensions in each modality. In order to capture the inherent sample structure information implied in each modality, we design a graph convolutional network branch to learn the corresponding structure preserved feature representation and generate modal-specific initial classification labels. Since samples from different modalities should share consistent labels, a cross-modal label fusion module is deployed to capture the label correlations of different modalities. In addition, motivated the ideally orthogonality of final fused label matrix, we design a label confidence loss to supervise the network for learning more separable data representations. To the best of our knowledge, MLCLNet is the first work which integrates both feature and label-level confidence learning for multimodal classification. Extensive experiments on four multimodal medical datasets are conducted to validate superior performance of MLCLNet when compared to other state-of-the-art methods.

JBHI Journal 2023 Journal Article

Multi-Level Constrained Intra and Inter Subject Feature Representation for Facial Video Based BVP Signal Measurement

  • Bin Li
  • Wei Zhang
  • Hong Fu
  • Hao Liu
  • Feng Xu

Facial video-based blood volume pulse (BVP) signal measurement holds great potential for remote health monitoring, while existing methods have issues with convolutional kernel perceptual field constraints. This article proposes an end-to-end multi-level constrained spatiotemporal representation structure for facial video-based BVP signal measurement. First, an intra- and inter-subject feature representation is proposed to strengthen the BVP-related features generation at high, semantic, and shallow levels, respectively. Second, the global-local association is presented to enhance BVP signal period pattern learning, and the global temporal features are introduced into the local spatial convolution of each frame by adaptive kernel weights. Finally, the multi-dimensional fused features are mapped to one-dimensional BVP signals by the task-oriented signal estimator. The experimental results on the publicly available MMSE-HR dataset demonstrate that the proposed structure overperforms state-of-the-art methods (e. g. , AutoHR) in BVP signal measurement, with a 20% and 40% reduction in mean absolute error and root mean squared error, respectively. The proposed structure would be a powerful tool for telemedical and non-contact heart health monitoring.

ICRA Conference 2023 Conference Paper

Obstacle-Aware Topological Planning over Polyhedral Representation for Quadrotors

  • Junjie Gao 0001
  • Fenghua He 0001
  • Wei Zhang
  • Yu Yao 0004

In this paper, we propose a novel mapping-planning framework for autonomous quadrotor navigation. First, a polyhedron-based mapping algorithm is presented to fully exploit the information of the onboard sensor data. Polyhedra are generated to approximate the segmented clusters of occupied voxels. Then, customized data structures are designed to extract information for motion planning in real time. With complete knowledge of the shape, position, and number of the observed obstacles, we can conveniently generate smooth trajectories with sufficient obstacle clearance along the most desired direction. Before searching for the initial path, a local topological graph is constructed to keep the path expanding in the most favorable topology class. The following path search is segmented based on the graph vertices, which allows fast convergence. The refined trajectory is obtained after smoothing, and large deviations are penalized in the formulated optimization problem to preserve the original clearance. Finally, we analyze and validate the proposed framework through extensive simulations and real-world quadrotor flights.

AAAI Conference 2023 Conference Paper

OMPQ: Orthogonal Mixed Precision Quantization

  • Yuexiao Ma
  • Taisong Jin
  • Xiawu Zheng
  • Yan Wang
  • Huixia Li
  • Yongjian Wu
  • Guannan Jiang
  • Wei Zhang

To bridge the ever-increasing gap between deep neural networks' complexity and hardware capability, network quantization has attracted more and more research attention. The latest trend of mixed precision quantization takes advantage of hardware's multiple bit-width arithmetic operations to unleash the full potential of network quantization. However, existing approaches rely heavily on an extremely time-consuming search process and various relaxations when seeking the optimal bit configuration. To address this issue, we propose to optimize a proxy metric of network orthogonality that can be efficiently solved with linear programming, which proves to be highly correlated with quantized model accuracy and bit-width. Our approach significantly reduces the search time and the required data amount by orders of magnitude, but without a compromise on quantization accuracy. Specifically, we achieve 72.08% Top-1 accuracy on ResNet-18 with 6.7Mb parameters, which does not require any searching iterations. Given the high efficiency and low data dependency of our algorithm, we use it for the post-training quantization, which achieves 71.27% Top-1 accuracy on MobileNetV2 with only 1.5Mb parameters.

NeurIPS Conference 2023 Conference Paper

OpenLane-V2: A Topology Reasoning Benchmark for Unified 3D HD Mapping

  • Huijie Wang
  • Tianyu Li
  • Yang Li
  • Li Chen
  • Chonghao Sima
  • Zhenbo Liu
  • Bangjun Wang
  • Peijin Jia

Accurately depicting the complex traffic scene is a vital component for autonomous vehicles to execute correct judgments. However, existing benchmarks tend to oversimplify the scene by solely focusing on lane perception tasks. Observing that human drivers rely on both lanes and traffic signals to operate their vehicles safely, we present OpenLane-V2, the first dataset on topology reasoning for traffic scene structure. The objective of the presented dataset is to advance research in understanding the structure of road scenes by examining the relationship between perceived entities, such as traffic elements and lanes. Leveraging existing datasets, OpenLane-V2 consists of 2, 000 annotated road scenes that describe traffic elements and their correlation to the lanes. It comprises three primary sub-tasks, including the 3D lane detection inherited from OpenLane, accompanied by corresponding metrics to evaluate the model’s performance. We evaluate various state-of-the-art methods, and present their quantitative and qualitative results on OpenLane-V2 to indicate future avenues for investigating topology reasoning in traffic scenes.

NeurIPS Conference 2023 Conference Paper

Reading Relevant Feature from Global Representation Memory for Visual Object Tracking

  • Xinyu Zhou
  • Pinxue Guo
  • Lingyi Hong
  • Jinglun Li
  • Wei Zhang
  • Weifeng Ge
  • Wenqiang Zhang

Reference features from a template or historical frames are crucial for visual object tracking. Prior works utilize all features from a fixed template or memory for visual object tracking. However, due to the dynamic nature of videos, the required reference historical information for different search regions at different time steps is also inconsistent. Therefore, using all features in the template and memory can lead to redundancy and impair tracking performance. To alleviate this issue, we propose a novel tracking paradigm, consisting of a relevance attention mechanism and a global representation memory, which can adaptively assist the search region in selecting the most relevant historical information from reference features. Specifically, the proposed relevance attention mechanism in this work differs from previous approaches in that it can dynamically choose and build the optimal global representation memory for the current frame by accessing cross-frame information globally. Moreover, it can flexibly read the relevant historical information from the constructed memory to reduce redundancy and counteract the negative effects of harmful information. Extensive experiments validate the effectiveness of the proposed method, achieving competitive performance on five challenging datasets with 71 FPS.

AAAI Conference 2023 Conference Paper

RSPT: Reconstruct Surroundings and Predict Trajectory for Generalizable Active Object Tracking

  • Fangwei Zhong
  • Xiao Bi
  • Yudi Zhang
  • Wei Zhang
  • Yizhou Wang

Active Object Tracking (AOT) aims to maintain a specific relation between the tracker and object(s) by autonomously controlling the motion system of a tracker given observations. It is widely used in various applications such as mobile robots and autonomous driving. However, Building a generalizable active tracker that works robustly across various scenarios remains a challenge, particularly in unstructured environments with cluttered obstacles and diverse layouts. To realize this, we argue that the key is to construct a state representation that can model the geometry structure of the surroundings and the dynamics of the target. To this end, we propose a framework called RSPT to form a structure-aware motion representation by Reconstructing Surroundings and Predicting the target Trajectory. Moreover, we further enhance the generalization of the policy network by training in the asymmetric dueling mechanism. Empirical results show that RSPT outperforms existing methods in unseen environments, especially those with cluttered obstacles and diverse layouts. We also demonstrate good sim-to-real transfer when deploying RSPT in real-world scenarios.

AAAI Conference 2023 Conference Paper

Self-Decoupling and Ensemble Distillation for Efficient Segmentation

  • Yuang Liu
  • Wei Zhang
  • Jun Wang

Knowledge distillation (KD) is a promising teacher-student learning paradigm that transfers information from a cumbersome teacher to a student network. To avoid the training cost of a large teacher network, the recent studies propose to distill knowledge from the student itself, called Self-KD. However, due to the limitations of the performance and capacity of the student, the soft-labels or features distilled by the student barely provide reliable guidance. Moreover, most of the Self-KD algorithms are specific to classification tasks based on soft-labels, and not suitable for semantic segmentation. To alleviate these contradictions, we revisit the label and feature distillation problem in segmentation, and propose Self-Decoupling and Ensemble Distillation for Efficient Segmentation (SDES). Specifically, we design a decoupled prediction ensemble distillation (DPED) algorithm that generates reliable soft-labels with multiple expert decoders, and a decoupled feature ensemble distillation (DFED) mechanism to utilize more important channel-wise feature maps for encoder learning. The extensive experiments on three public segmentation datasets demonstrate the superiority of our approach and the efficacy of each component in the framework through the ablation study.

AAAI Conference 2023 Conference Paper

SpatialFormer: Semantic and Target Aware Attentions for Few-Shot Learning

  • Jinxiang Lai
  • Siqian Yang
  • Wenlong Wu
  • Tao Wu
  • Guannan Jiang
  • Xi Wang
  • Jun Liu
  • Bin-Bin Gao

Recent Few-Shot Learning (FSL) methods put emphasis on generating a discriminative embedding features to precisely measure the similarity between support and query sets. Current CNN-based cross-attention approaches generate discriminative representations via enhancing the mutually semantic similar regions of support and query pairs. However, it suffers from two problems: CNN structure produces inaccurate attention map based on local features, and mutually similar backgrounds cause distraction. To alleviate these problems, we design a novel SpatialFormer structure to generate more accurate attention regions based on global features. Different from the traditional Transformer modeling intrinsic instance-level similarity which causes accuracy degradation in FSL, our SpatialFormer explores the semantic-level similarity between pair inputs to boost the performance. Then we derive two specific attention modules, named SpatialFormer Semantic Attention (SFSA) and SpatialFormer Target Attention (SFTA), to enhance the target object regions while reduce the background distraction. Particularly, SFSA highlights the regions with same semantic information between pair features, and SFTA finds potential foreground object regions of novel feature that are similar to base categories. Extensive experiments show that our methods are effective and achieve new state-of-the-art results on few-shot classification benchmarks.

EAAI Journal 2022 Journal Article

An efficient fire and smoke detection algorithm based on an end-to-end structured network

  • Yuming Li
  • Wei Zhang
  • Yanyan Liu
  • Rudong Jing
  • Changsong Liu

Detection transformer (DETR) combines convolutional neural network (CNN) with transformer, providing a more advanced idea. In this paper, an object detection model based on DETR is proposed for fire and smoke detection. Compared with other methods that based on deep learning, the proposed one simplifies the pipeline of detection and builds an end-to-end detector. At the same time, the original DETR model usually requires long training time and large amount of computation, resulting in relatively poor performance in detection speed and accuracy, and it shows to be not friendly for small or early fire detection. Therefore, when designing the proposed network model, a normalization-based attention module is added in the feature extraction stage to highlight the effective features, being beneficial to the process of encoding. A multiscale deformable attention is also used in the encoder–decoder structure, accelerating therefore, the process of convergence during the training of the model, which includes the enhancement of the detection of small objects. Also, considering the cost of computation, the number of layers in the encoder–decoder structure is redefined to reduce the complexity of the model, and that also reduces the requirements of the application equipment. Detailed experiments are conducted on three self-built datasets and two public video sets. The results show that, the proposed method has an excellent performance on all of the datasets considered here.

AAAI Conference 2022 Conference Paper

Comprehensive Regularization in a Bi-directional Predictive Network for Video Anomaly Detection

  • Chengwei Chen
  • Yuan Xie
  • Shaohui Lin
  • Angela Yao
  • Guannan Jiang
  • Wei Zhang
  • Yanyun Qu
  • Ruizhi Qiao

Video anomaly detection aims to automatically identify unusual objects or behaviours by learning from normal videos. Previous methods tend to use simplistic reconstruction or prediction constraints, which leads to the insufficiency of learned representations for normal data. As such, we propose a novel bi-directional architecture with three consistency constraints to comprehensively regularize the prediction task from pixelwise, cross-modal, and temporal-sequence levels. First, predictive consistency is proposed to consider the symmetry property of motion and appearance in forwards and backwards time, which ensures the highly realistic appearance and motion predictions at the pixel-wise level. Second, association consistency considers the relevance between different modalities and uses one modality to regularize the prediction of another one. Finally, temporal consistency utilizes the relationship of the video sequence and ensures that the predictive network generates temporally consistent frames. During inference, the pattern of abnormal frames is unpredictable and will therefore cause higher prediction errors. Experiments show that our method outperforms advanced anomaly detectors and achieves state-of-the-art results on UCSD Ped2, CUHK Avenue, and ShanghaiTech datasets.

IJCAI Conference 2022 Conference Paper

Data-Efficient Backdoor Attacks

  • Pengfei Xia
  • Ziqiang Li
  • Wei Zhang
  • Bin Li

Recent studies have proven that deep neural networks are vulnerable to backdoor attacks. Specifically, by mixing a small number of poisoned samples into the training set, the behavior of the trained model can be maliciously controlled. Existing attack methods construct such adversaries by randomly selecting some clean data from the benign set and then embedding a trigger into them. However, this selection strategy ignores the fact that each poisoned sample contributes inequally to the backdoor injection, which reduces the efficiency of poisoning. In this paper, we formulate improving the poisoned data efficiency by the selection as an optimization problem and propose a Filtering-and-Updating Strategy (FUS) to solve it. The experimental results on CIFAR-10 and ImageNet-10 indicate that the proposed method is effective: the same attack success rate can be achieved with only 47% to 75% of the poisoned sample volume compared to the random selection strategy. More importantly, the adversaries selected according to one setting can generalize well to other settings, exhibiting strong transferability. The prototype code of our method is now available at https: //github. com/xpf/Data-Efficient-Backdoor-Attacks.

NeurIPS Conference 2022 Conference Paper

DetCLIP: Dictionary-Enriched Visual-Concept Paralleled Pre-training for Open-world Detection

  • Lewei Yao
  • Jianhua Han
  • Youpeng Wen
  • Xiaodan Liang
  • Dan Xu
  • Wei Zhang
  • Zhenguo Li
  • Chunjing Xu

Open-world object detection, as a more general and challenging goal, aims to recognize and localize objects described by arbitrary category names. The recent work GLIP formulates this problem as a grounding problem by concatenating all category names of detection datasets into sentences, which leads to inefficient interaction between category names. This paper presents DetCLIP, a paralleled visual-concept pre-training method for open-world detection by resorting to knowledge enrichment from a designed concept dictionary. To achieve better learning efficiency, we propose a novel paralleled concept formulation that extracts concepts separately to better utilize heterogeneous datasets (i. e. , detection, grounding, and image-text pairs) for training. We further design a concept dictionary (with descriptions) from various online sources and detection datasets to provide prior knowledge for each concept. By enriching the concepts with their descriptions, we explicitly build the relationships among various concepts to facilitate the open-domain learning. The proposed concept dictionary is further used to provide sufficient negative concepts for the construction of the word-region alignment loss, and to complete labels for objects with missing descriptions in captions of image-text pair data. The proposed framework demonstrates strong zero-shot detection performances, e. g. , on the LVIS dataset, our DetCLIP-T outperforms GLIP-T by 9. 9% mAP and obtains a 13. 5% improvement on rare categories compared to the fully-supervised model with the same backbone as ours.

AAAI Conference 2022 Conference Paper

DIRL: Domain-Invariant Representation Learning for Generalizable Semantic Segmentation

  • Qi Xu
  • Liang Yao
  • Zhengkai Jiang
  • Guannan Jiang
  • Wenqing Chu
  • Wenhui Han
  • Wei Zhang
  • Chengjie Wang

Model generalization to the unseen scenes is crucial to realworld applications, such as autonomous driving, which requires robust vision systems. To enhance the model generalization, domain generalization through learning the domaininvariant representation has been widely studied. However, most existing works learn the shared feature space within multi-source domains but ignore the characteristic of the feature itself (e. g. , the feature sensitivity to the domain-specific style). Therefore, we propose the Domain-invariant Representation Learning (DIRL) for domain generalization which utilizes the feature sensitivity as the feature prior to guide the enhancement of the model generalization capability. The guidance reflects in two folds: 1) Feature re-calibration that introduces the Prior Guided Attention Module (PGAM) to emphasize the insensitive features and suppress the sensitive features. 2): Feature whiting that proposes the Guided Feature Whiting (GFW) to remove the feature correlations which are sensitive to the domain-specific style. We construct the domain-invariant representation which suppresses the effect of the domain-specific style on the quality and correlation of the features. As a result, our method is simple yet effective, and can enhance the robustness of various backbone networks with little computational cost. Extensive experiments over multiple domains generalizable segmentation tasks show the superiority of our approach to other methods.

AAAI Conference 2022 Conference Paper

Explore Inter-contrast between Videos via Composition for Weakly Supervised Temporal Sentence Grounding

  • Jiaming Chen
  • Weixin Luo
  • Wei Zhang
  • Lin Ma

Weakly supervised temporal sentence grounding aims to temporally localize the target segment corresponding to a given natural language query, where it provides video-query pairs without temporal annotations during training. Most existing methods use the fused visual-linguistic feature to reconstruct the query, where the least reconstruction error determines the target segment. This work introduces a novel approach that explores the inter-contrast between videos in a composed video by selecting components from two different videos and fusing them into a single video. Such a straightforward yet effective composition strategy provides the temporal annotations at multiple composed positions, resulting in numerous videos with temporal ground-truths for training the temporal sentence grounding task. A transformer framework is introduced with multi-tasks training to learn a compact but efficient visual-linguistic space. The experimental results on the public Charades-STA and ActivityNet- Caption dataset demonstrate the effectiveness of the proposed method, where our approach achieves comparable performance over the state-of-the-art weakly-supervised baselines. The code is available at https: //github. com/PPjmchen/ Composition WSTG.

YNIMG Journal 2022 Journal Article

Hemodynamic and metabolic correspondence of resting-state voxel-based physiological metrics in healthy adults

  • Shengwen Deng
  • Crystal G. Franklin
  • Michael O'Boyle
  • Wei Zhang
  • Betty L. Heyl
  • Paul A. Jerabek
  • Hanzhang Lu
  • Peter T. Fox

Voxel-based physiological (VBP) variables derived from blood oxygen level dependent (BOLD) fMRI time-course variations include: amplitude of low frequency fluctuations (ALFF), fractional amplitude of low frequency fluctuations (fALFF) and regional homogeneity (ReHo). Although these BOLD-derived variables can detect between-group (e. g. disease vs control) spatial pattern differences, physiological interpretations are not well established. The primary objective of this study was to quantify spatial correspondences between BOLD VBP variables and PET measurements of cerebral metabolic rate and hemodynamics, being well-validated physiological standards. To this end, quantitative, whole-brain PET images of metabolic rate of glucose (MRGlu; 18FDG) and oxygen (MRO2; 15OO), blood flow (BF; H2 15O) and blood volume (BV; C15O) were obtained in 16 healthy controls. In the same subjects, BOLD time-courses were obtained for computation of ALFF, fALFF and ReHo images. PET variables were compared pair-wise with BOLD variables. In group-averaged, across-region analyses, ALFF corresponded significantly only with BV (R = 0. 64; p < 0. 0001). fALFF corresponded most strongly with MRGlu (R = 0. 79; p < 0. 0001), but also significantly (p < 0. 0001) with MRO2 (R = 0. 68), BF (R = 0. 68) and BV (R=0. 68). ReHo performed similarly to fALFF, with significant strong correspondence (p < 0. 0001) with MRGlu (R = 0. 78), MRO2 (R = 0. 54), and, but less strongly with BF (R = 0. 50) and BV (R=0. 50). Mutual information analyses further clarified these physiological interpretations. When conditioned by BV, ALFF retained no significant MRGlu, MRO2 or BF information. When conditioned by MRGlu, fALFF and ReHo retained no significant MRO2, BF or BV information. Of concern, however, the strength of PET-BOLD correspondences varied markedly by brain region, which calls for future investigation on physiological interpretations at a regional and per-subject basis.

AAAI Conference 2022 Conference Paper

LCTR: On Awakening the Local Continuity of Transformer for Weakly Supervised Object Localization

  • Zhiwei Chen
  • Changan Wang
  • Yabiao Wang
  • Guannan Jiang
  • Yunhang Shen
  • Ying Tai
  • Chengjie Wang
  • Wei Zhang

Weakly supervised object localization (WSOL) aims to learn object localizer solely by using image-level labels. The convolution neural network (CNN) based techniques often result in highlighting the most discriminative part of objects while ignoring the entire object extent. Recently, the transformer architecture has been deployed to WSOL to capture the longrange feature dependencies with self-attention mechanism and multilayer perceptron structure. Nevertheless, transformers lack the locality inductive bias inherent to CNNs and therefore may deteriorate local feature details in WSOL. In this paper, we propose a novel framework built upon the transformer, termed LCTR (Local Continuity TRansformer), which targets at enhancing the local perception capability of global features among long-range feature dependencies. To this end, we propose a relational patch-attention module (RPAM), which considers cross-patch information on a global basis. We further design a cue digging module (CDM), which utilizes local features to guide the learning trend of the model for highlighting the weak local responses. Finally, comprehensive experiments are carried out on two widely used datasets, i. e. , CUB-200-2011 and ILSVRC, to verify the effectiveness of our method.

NeurIPS Conference 2022 Conference Paper

Leveraging the Hints: Adaptive Bidding in Repeated First-Price Auctions

  • Wei Zhang
  • Yanjun Han
  • Zhengyuan Zhou
  • Aaron Flores
  • Tsachy Weissman

With the advent and increasing consolidation of e-commerce, digital advertising has very recently replaced traditional advertising as the main marketing force in the economy. In the past four years, a particularly important development in the digital advertising industry is the shift from second-price auctions to first-price auctions for online display ads. This shift immediately motivated the intellectually challenging question of how to bid in first-price auctions, because unlike in second-price auctions, bidding one's private value truthfully is no longer optimal. Following a series of recent works in this area, we consider a differentiated setup: we do not make any assumption about other bidders' maximum bid (i. e. it can be adversarial over time), and instead assume that we have access to a hint that serves as a prediction of other bidders' maximum bid, where the prediction is learned through some blackbox machine learning model. We consider two types of hints: one where a single point-prediction is available, and the other where a hint interval (representing a type of confidence region into which others' maximum bid falls) is available. We establish minimax optimal regret bounds for both cases and highlight the quantitatively different behavior between the two settings. We also provide improved regret bounds when the others' maximum bid exhibits the further structure of sparsity. Finally, we complement the theoretical results with demonstrations using real bidding data.

AAAI Conference 2022 Conference Paper

Multi-Knowledge Aggregation and Transfer for Semantic Segmentation

  • Yuang Liu
  • Wei Zhang
  • Jun Wang

As a popular deep neural networks (DNN) compression technique, knowledge distillation (KD) has attracted increasing attentions recently. Existing KD methods usually utilize one kind of knowledge in an intermediate layer of DNN for classification tasks to transfer useful information from cumbersome teacher networks to compact student networks. However, this paradigm is not very suitable for semantic segmentation, a comprehensive vision task based on both pixel-level and contextual information, since it cannot provide rich information for distillation. In this paper, we propose a novel multi-knowledge aggregation and transfer (MKAT) framework to comprehensively distill knowledge within an intermediate layer for semantic segmentation. Specifically, the proposed framework consists of three parts: Independent Transformers and Encoders module (ITE), Auxiliary Prediction Branch (APB), and Mutual Label Calibration (MLC) mechanism, which can take advantage of abundant knowledge from intermediate features. To demonstrate the effectiveness of our proposed approach, we conduct extensive experiments on three segmentation datasets: Pascal VOC, Cityscapes, and CamVid, showing that MKAT outperforms the other KD methods.

JBHI Journal 2022 Journal Article

NMFLRR: Clustering scRNA-Seq Data by Integrating Nonnegative Matrix Factorization With Low Rank Representation

  • Wei Zhang
  • Xiaoli Xue
  • Xiaoying Zheng
  • Zizhu Fan

Fast-developing single-cell technologies create unprecedented opportunities to reveal cell heterogeneity and diversity. Accurate classification of single cells is a critical prerequisite for recovering the mechanisms of heterogeneity. However, the scRNA-seq profiles we obtained at present have high dimensionality, sparsity, and noise, which pose challenges for existing clustering methods in grouping cells that belong to the same subpopulation based on transcriptomic profiles. Although many computational methods have been proposed developing novel and effective computational methods to accurately identify cell types remains a considerable challenge. We present a new computational framework to identify cell types by integrating low-rank representation (LRR) and nonnegative matrix factorization (NMF); this framework is named NMFLRR. The LRR captures the global properties of original data by using nuclear norms, and a locality constrained graph regularization term is introduced to characterize the data’s local geometric information. The similarity matrix and low-dimensional features of data can be simultaneously obtained by applying the alternating direction method of multipliers (ADMM) algorithm to handle each variable alternatively in an iterative way. We finally obtained the predicted cell types by using a spectral algorithm based on the optimized similarity matrix. Nine real scRNA-seq datasets were used to test the performance of NMFLRR and fifteen other competitive methods, and the accuracy and robustness of the simulation results suggest the NMFLRR is a promising algorithm for the classification of single cells. The simulation code is freely available at: https://github.com/wzhangwhu/NMFLRR_code.

JBHI Journal 2022 Journal Article

Reinforcement Learning Based Diagnosis and Prediction for COVID-19 by Optimizing a Mixed Cost Function From CT Images

  • Siying Chen
  • Minghui Liu
  • Pan Deng
  • Jiali Deng
  • Yi Yuan
  • Xuan Cheng
  • Tianshu Xie
  • Libo Xie

A novel coronavirus disease (COVID-19) is a pandemic disease has caused 4 million deaths and more than 200 million infections worldwide (as of August 4, 2021). Rapid and accurate diagnosis of COVID-19 infection is critical to controlling the spread of the epidemic. In order to quickly and efficiently detect COVID-19 and reduce the threat of COVID-19 to human survival, we have firstly proposed a detection framework based on reinforcement learning for COVID-19 diagnosis, which constructs a mixed loss function that can integrate the advantages of multiple loss functions. This paper uses the accuracy of the validation set as the reward value, and obtains the initial model for the next epoch by searching the model corresponding to the maximum reward value in each epoch. We also have proposed a prediction framework that integrates multiple detection frameworks using parameter sharing to predict the progression of patients' disease without additional training. This paper also constructed a higher-quality version of the CT image dataset containing 247 cases screened by professional physicians, and obtained more excellent results on this dataset. Meanwhile, we used the other two COVID-19 datasets as external verifications, and still achieved a high accuracy rate without additional training. Finally, the experimental results show that our classification accuracy can reach 98. 31%, and the precision, sensitivity, specificity, and AUC (Area Under Curve) are 98. 82%, 97. 99%, 98. 67%, and 0. 989, respectively. The accuracy of external verification can reach 93. 34% and 91. 05%. What's more, the accuracy of our prediction framework is 91. 54%. A large number of experiments demonstrate that our proposed method is effective and robust for COVID-19 detection and prediction.

NeurIPS Conference 2022 Conference Paper

Robustness to Unbounded Smoothness of Generalized SignSGD

  • Michael Crawshaw
  • Mingrui Liu
  • Francesco Orabona
  • Wei Zhang
  • Zhenxun Zhuang

Traditional analyses in non-convex optimization typically rely on the smoothness assumption, namely requiring the gradients to be Lipschitz. However, recent evidence shows that this smoothness condition does not capture the properties of some deep learning objective functions, including the ones involving Recurrent Neural Networks and LSTMs. Instead, they satisfy a much more relaxed condition, with potentially unbounded smoothness. Under this relaxed assumption, it has been theoretically and empirically shown that the gradient-clipped SGD has an advantage over the vanilla one. In this paper, we show that clipping is not indispensable for Adam-type algorithms in tackling such scenarios: we theoretically prove that a generalized SignSGD algorithm can obtain similar convergence rates as SGD with clipping but does not need explicit clipping at all. This family of algorithms on one end recovers SignSGD and on the other end closely resembles the popular Adam algorithm. Our analysis underlines the critical role that momentum plays in analyzing SignSGD-type and Adam-type algorithms: it not only reduces the effects of noise, thus removing the need for large mini-batch in previous analyses of SignSGD-type algorithms, but it also substantially reduces the effects of unbounded smoothness and gradient norms. To the best of our knowledge, this work is the first one showing the benefit of Adam-type algorithms compared with non-adaptive gradient algorithms such as gradient descent in the unbounded smoothness setting. We also compare these algorithms with popular optimizers on a set of deep learning tasks, observing that we can match the performance of Adam while beating others.

AAAI Conference 2022 Conference Paper

Visual Consensus Modeling for Video-Text Retrieval

  • Shuqiang Cao
  • Bairui Wang
  • Wei Zhang
  • Lin Ma

In this paper, we propose a novel method to mine the commonsense knowledge shared between the video and text modalities for video-text retrieval, namely visual consensus modeling. Different from the existing works, which learn the video and text representations and their complicated relationships solely based on the pairwise video-text data, we make the first attempt to model the visual consensus by mining the visual concepts from videos and exploiting their co-occurrence patterns within the video and text modalities with no reliance on any additional concept annotations. Specifically, we build a shareable and learnable graph as the visual consensus, where the nodes denoting the mined visual concepts and the edges connecting the nodes representing the co-occurrence relationships between the visual concepts. Extensive experimental results on the public benchmark datasets demonstrate that our proposed method, with the ability to effectively model the visual consensus, achieves state-of-the-art performance on the bidirectional video-text retrieval task. Our code is available at https: //github. com/sqiangcao99/VCM.

AAAI Conference 2022 Conference Paper

Weakly-Supervised Salient Object Detection Using Point Supervision

  • Shuyong Gao
  • Wei Zhang
  • Yan Wang
  • Qianyu Guo
  • Chenglong Zhang
  • Yangji He
  • Wenqiang Zhang

Current state-of-the-art saliency detection models rely heavily on large datasets of accurate pixel-wise annotations, but manually labeling pixels is time-consuming and laborintensive. There are some weakly supervised methods developed for alleviating the problem, such as image label, bounding box label, and scribble label, while point label still has not been explored in this field. In this paper, we propose a novel weakly-supervised salient object detection method using point supervision. To infer the saliency map, we first design an adaptive masked flood filling algorithm to generate pseudo labels. Then we develop a transformer-based pointsupervised saliency detection model to produce the first round of saliency maps. However, due to the sparseness of the label, the weakly supervised model tends to degenerate into a general foreground detection model. To address this issue, we propose a Non-Salient Suppression (NSS) method to optimize the erroneous saliency maps generated in the first round and leverage them for the second round of training. Moreover, we build a new point-supervised dataset (P- DUTS) by relabeling the DUTS dataset. In P-DUTS, there is only one labeled point for each salient object. Comprehensive experiments on five largest benchmark datasets demonstrate our method outperforms the previous state-of-the-art methods trained with the stronger supervision and even surpass several fully supervised state-of-the-art models. The code is available at: https: //github. com/shuyonggao/PSOD.

NeurIPS Conference 2022 Conference Paper

Wukong: A 100 Million Large-scale Chinese Cross-modal Pre-training Benchmark

  • Jiaxi Gu
  • Xiaojun Meng
  • Guansong Lu
  • Lu Hou
  • Niu Minzhe
  • Xiaodan Liang
  • Lewei Yao
  • Runhui Huang

Vision-Language Pre-training (VLP) models have shown remarkable performance on various downstream tasks. Their success heavily relies on the scale of pre-trained cross-modal datasets. However, the lack of large-scale datasets and benchmarks in Chinese hinders the development of Chinese VLP models and broader multilingual applications. In this work, we release a large-scale Chinese cross-modal dataset named Wukong, which contains 100 million Chinese image-text pairs collected from the web. Wukong aims to benchmark different multi-modal pre-training methods to facilitate the VLP research and community development. Furthermore, we release a group of models pre-trained with various image encoders (ViT-B/ViT-L/SwinT) and also apply advanced pre-training techniques into VLP such as locked-image text tuning, token-wise similarity in contrastive learning, and reduced-token interaction. Extensive experiments and a benchmarking of different downstream tasks including a new largest human-verified image-text test dataset are also provided. Experiments show that Wukong can serve as a promising Chinese pre-training dataset and benchmark for different cross-modal learning methods. For the zero-shot image classification task on 10 datasets, $Wukong_\text{ViT-L}$ achieves an average accuracy of 73. 03%. For the image-text retrieval task, it achieves a mean recall of 71. 6% on AIC-ICC which is 12. 9% higher than WenLan 2. 0. Also, our Wukong models are benchmarked on downstream tasks with other variants on multiple datasets, e. g. , Flickr8K-CN, Flickr-30K-CN, COCO-CN, et al. More information can be referred to https: //wukong-dataset. github. io/wukong-dataset/.

JBHI Journal 2021 Journal Article

Attention-Guided Deep Neural Network With Multi-Scale Feature Fusion for Liver Vessel Segmentation

  • Qingsen Yan
  • Bo Wang
  • Wei Zhang
  • Chuan Luo
  • Wei Xu
  • Zhengqing Xu
  • Yanning Zhang
  • Qinfeng Shi

Liver vessel segmentation is fast becoming a key instrument in the diagnosis and surgical planning of liver diseases. In clinical practice, liver vessels are normally manual annotated by clinicians on each slice of CT images, which is extremely laborious. Several deep learning methods exist for liver vessel segmentation, however, promoting the performance of segmentation remains a major challenge due to the large variations and complex structure of liver vessels. Previous methods mainly using existing UNet architecture, but not all features of the encoder are useful for segmentation and some even cause interferences. To overcome this problem, we propose a novel deep neural network for liver vessel segmentation, called LVSNet, which employs special designs to obtain the accurate structure of the liver vessel. Specifically, we design Attention-Guided Concatenation (AGC) module to adaptively select the useful context features from low-level features guided by high-level features. The proposed AGC module focuses on capturing rich complemented information to obtain more details. In addition, we introduce an innovative multi-scale fusion block by constructing hierarchical residual-like connections within one single residual block, which is of great importance for effectively linking the local blood vessel fragments together. Furthermore, we construct a new dataset containing 40 thin thickness cases (0. 625 mm) which consist of CT volumes and annotated vessels. To evaluate the effectiveness of the method with minor vessels, we also propose an automatic stratification method to split major and minor liver vessels. Extensive experimental results demonstrate that the proposed LVSNet outperforms previous methods on liver vessel segmentation datasets. Additionally, we conduct a series of ablation studies that comprehensively support the superiority of the underlying concepts.

AAAI Conference 2021 Conference Paper

Circles are like Ellipses, or Ellipses are like Circles? Measuring the Degree of Asymmetry of Static and Contextual Word Embeddings and the Implications to Representation Learning

  • Wei Zhang
  • Murray Campbell
  • Yang Yu
  • Sadhana Kumaravel

Human judgments of word similarity have been a popular method of evaluating the quality of word embedding. But it fails to measure the geometry properties such as asymmetry. For example, it is more natural to say “Ellipses are like Circles” than “Circles are like Ellipses”. Such asymmetry has been observed from the word evocation experiment, where one word is used to recall another. This association data have been understudied for measuring embedding quality. In this paper, we use three well-known evocation datasets for the purpose and study both static embedding as well as contextual embedding, such as BERT. To fight for the dynamic nature of BERT embedding, we probe BERT’s conditional probabilities as a language model, using a large number of Wikipedia contexts to derive a theoretically justifiable Bayesian asymmetry score. The result shows that the asymmetry judgment and similarity judgments disagree, and asymmetry judgment aligns with its strong performance on “extrinsic evaluations”. This is the first time we can show contextual embeddings’s strength on intrinsic evaluation, and the asymmetry judgment provides a new perspective to evaluate contextual embedding and new insights for representation learning.

NeurIPS Conference 2021 Conference Paper

CodeNet: A Large-Scale AI for Code Dataset for Learning a Diversity of Coding Tasks

  • Ruchir Puri
  • David Kung
  • Geert Janssen
  • Wei Zhang
  • Giacomo Domeniconi
  • Vladimir Zolotov
  • Julian T Dolby
  • Jie Chen

Over the last several decades, software has been woven into the fabric of every aspect of our society. As software development surges and code infrastructure of enterprise applications ages, it is now more critical than ever to increase software development productivity and modernize legacy applications. Advances in deep learning and machine learning algorithms have enabled breakthroughs in computer vision, speech recognition, natural language processing and beyond, motivating researchers to leverage AI techniques to improve software development efficiency. Thus, the fast-emerging research area of “AI for Code” has garnered new interest and gathered momentum. In this paper, we present a large-scale dataset \textit{CodeNet}, consisting of over 14 million code samples and about 500 million lines of code in 55 different programming languages, which is aimed at teaching AI to code. In addition to its large scale, CodeNet has a rich set of high-quality annotations to benchmark and help accelerate research in AI techniques for a variety of critical coding tasks, including code similarity and classification, code translation between a large variety of programming languages, and code performance (runtime and memory) improvement techniques. Additionally, CodeNet provides sample input and output test sets for 98. 5\% of the code samples, which can be used as an oracle for determining code correctness and potentially guide reinforcement learning for code quality improvements. As a usability feature, we provide several pre-processing tools in CodeNet to transform source code into representations that can be readily used as inputs into machine learning models. Results of code classification and code similarity experiments using the CodeNet dataset are provided as a reference. We hope that the scale, diversity and rich, high-quality annotations of CodeNet will offer unprecedented research opportunities at the intersection of AI and Software Engineering.

JBHI Journal 2021 Journal Article

DeepUWF: An Automated Ultra-Wide-Field Fundus Screening System via Deep Learning

  • Wei Zhang
  • Xiujuan Zhao
  • Yuanyuan Chen
  • Jie Zhong
  • Zhang Yi

The emerging ultra-wide field of view (UWF) fundus color imaging is a powerful tool for fundus screening. However, manual screening is labor-intensive and subjective. Based on 2644 UWF images, a set of early fundus abnormal screening system named DeepUWF is developed. DeepUWF includes an abnormal fundus screening subsystem and a disease diagnosis subsystem for three kinds of fundus diseases (retinal tear & retinal detachment, diabetic retinopathy and pathological myopia). The components in the system are composed of a set of excellent convolutional neural networks and two custom classifiers. However, the contrast of UWF images used in the research is low, which seriously limits the extraction of fine features of UWF images by depth model. Therefore, the high specificity and low sensitivity of prediction results have always been difficult problems in research. In order to solve this problem, six kinds of image preprocessing techniques are adopted, and their effects on the prediction performance of fundus abnormal and three kinds of fundus diseases models are studied. A variety of experimental indicators are used to evaluate the algorithms for validity and reliability. The experimental results show that these preprocessing methods are helpful to improve the learning ability of the networks and achieve good sensitivity and specificity. Without ophthalmologists, DeepUWF has potential application value, which is helpful for fundus health screening and workflow improvement.

IJCAI Conference 2021 Conference Paper

Drop Redundant, Shrink Irrelevant: Selective Knowledge Injection for Language Pretraining

  • Ningyu Zhang
  • Shumin Deng
  • Xu Cheng
  • Xi Chen
  • Yichi Zhang
  • Wei Zhang
  • Huajun Chen

Previous research has demonstrated the power of leveraging prior knowledge to improve the performance of deep models in natural language processing. However, traditional methods neglect the fact that redundant and irrelevant knowledge exists in external knowledge bases. In this study, we launched an in-depth empirical investigation into downstream tasks and found that knowledge-enhanced approaches do not always exhibit satisfactory improvements. To this end, we investigate the fundamental reasons for ineffective knowledge infusion and present selective injection for language pretraining, which constitutes a model-agnostic method and is readily pluggable into previous approaches. Experimental results on benchmark datasets demonstrate that our approach can enhance state-of-the-art knowledge injection methods.

AAAI Conference 2021 Conference Paper

Exploiting Relationship for Complex-scene Image Generation

  • Tianyu Hua
  • Hongdong Zheng
  • Yalong Bai
  • Wei Zhang
  • Xiao-Ping Zhang
  • Tao Mei

The significant progress on Generative Adversarial Networks (GANs) has facilitated realistic single-object image generation based on language input. However, complex-scene generation (with various interactions among multiple objects) still suffers from messy layouts and object distortions, due to diverse configurations in layouts and appearances. Prior methods are mostly object-driven and ignore their inter-relations that play a significant role in complex-scene images. This work explores relationship-aware complex-scene image generation, where multiple objects are inter-related as a scene graph. With the help of relationships, we propose three major updates in the generation framework. First, reasonable spatial layouts are inferred by jointly considering the semantics and relationships among objects. Compared to standard location regression, we show relative scales and distances serve a more reliable target. Second, since the relations between objects significantly influence an object’s appearance, we design a relation-guided generator to generate objects reflecting their relationships. Third, a novel scene graph discriminator is proposed to guarantee the consistency between the generated image and the input scene graph. Our method tends to synthesize plausible layouts and objects, respecting the interplay of multiple objects in an image. Experimental results on Visual Genome and HICO-DET datasets show that our proposed method significantly outperforms prior arts in terms of IS and FID metrics. Based on our user study and visual inspection, our method is more effective in generating logical layout and appearance for complex-scenes.

AAAI Conference 2021 Conference Paper

Graph-Based Tri-Attention Network for Answer Ranking in CQA

  • Wei Zhang
  • Zeyuan Chen
  • Chao Dong
  • Wen Wang
  • Hongyuan Zha
  • Jianyong Wang

In community-based question answering (CQA) platforms, automatic answer ranking for a given question is critical for finding potentially popular answers in early times. The mainstream approaches learn to generate answer ranking scores based on the matching degree between question and answer representations as well as the influence of respondents. However, they encounter two main limitations: (1) Correlations between answers in the same question are often overlooked. (2) Question and respondent representations are built independently of specific answers before affecting answer representations. To address the limitations, we devise a novel graph-based tri-attention network, namely GTAN, which has two innovations. First, GTAN proposes to construct a graph for each question and learn answer correlations from each graph through graph neural networks (GNNs). Second, based on the representations learned from GNNs, an alternating tri-attention method is developed to alternatively build target-aware respondent representations, answerspecific question representations, and context-aware answer representations by attention computation. GTAN finally integrates the above representations to generate answer ranking scores. Experiments on three real-world CQA datasets demonstrate GTAN significantly outperforms state-of-the-art answer ranking methods, validating the rationality of the network architecture.

NeurIPS Conference 2021 Conference Paper

IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning

  • Pan Lu
  • Liang Qiu
  • Jiaqi Chen
  • Tanglin Xia
  • Yizhou Zhao
  • Wei Zhang
  • Zhou Yu
  • Xiaodan Liang

Current visual question answering (VQA) tasks mainly consider answering human-annotated questions for natural images. However, aside from natural images, abstract diagrams with semantic richness are still understudied in visual understanding and reasoning research. In this work, we introduce a new challenge of Icon Question Answering (IconQA) with the goal of answering a question in an icon image context. We release IconQA, a large-scale dataset that consists of 107, 439 questions and three sub-tasks: multi-image-choice, multi-text-choice, and filling-in-the-blank. The IconQA dataset is inspired by real-world diagram word problems that highlight the importance of abstract diagram understanding and comprehensive cognitive reasoning. Thus, IconQA requires not only perception skills like object recognition and text understanding, but also diverse cognitive reasoning skills, such as geometric reasoning, commonsense reasoning, and arithmetic reasoning. To facilitate potential IconQA models to learn semantic representations for icon images, we further release an icon dataset Icon645 which contains 645, 687 colored icons on 377 classes. We conduct extensive user studies and blind experiments and reproduce a wide range of advanced VQA methods to benchmark the IconQA task. Also, we develop a strong IconQA baseline Patch-TRM that applies a pyramid cross-modal Transformer with input diagram embeddings pre-trained on the icon dataset. IconQA and Icon645 are available at https: //iconqa. github. io.

IJCAI Conference 2021 Conference Paper

Mental Models of AI Agents in a Cooperative Game Setting (Extended Abstract)

  • Katy Ilonka Gero
  • Zahra Ashktorab
  • Casey Dugan
  • Qian Pan
  • James Johnson
  • Werner Geyer
  • Maria Ruiz
  • Sarah Miller

As more and more forms of AI become prevalent, it becomes increasingly important to understand how people develop mental models of these systems. In this work we study people's mental models of an AI agent in a cooperative word guessing game. We run a study in which people play the game with an AI agent while ``thinking out loud''; through thematic analysis we identify features of the mental models developed by participants. In a large-scale study we have participants play the game with the AI agent online and use a post-game survey to probe their mental model. We find that those who win more often have better estimates of the AI agent's abilities. We present three components---global knowledge, local knowledge, and knowledge distribution---for modeling AI systems and propose that understanding the underlying technology is insufficient for developing appropriate conceptual models---analysis of behavior is also necessary.

EAAI Journal 2021 Journal Article

Microblog sentiment analysis via embedding social contexts into an attentive LSTM

  • Jing Yang
  • Xiaomei Zou
  • Wei Zhang
  • Hongyu Han

With the rise of microblogging services like Twitter and Sina Weibo, users are able to post various contents on breaking news, public events, or products conveniently and swiftly. These massive contents carry users’ mass sentiment and opinions on various topics, which are a kind of useful and timely source. Traditional microblog sentiment analysis methods often assume that microblogs are independent and identically distributed, they ignore the fact that the microblogs are networked data. Although some methods take the relations between microblogs into consideration, they only use shallow network features which are not sufficient, such as neighbors. Besides, these methods are content-based methods because they cannot use social context information in the prediction stage. To solve this problem, in this paper we use a deep learning method to fully capture the features of microblog relations including both the implicit and explicit ones and use these features to promote microblog sentiment analysis results. Specifically, we first construct a graph which models the relations between microblogs inspired by sentiment consistency and emotional contagion theories. Then we embed the microblog graph and get a continuous vector representation for social contexts of each microblog. After that, we propose a novel neural network to integrate social context knowledge with text information. To handle the problem that different words have different contributions to the classification result, we introduce the attention mechanism into our model. We conduct experiments on three publicly released datasets. The experimental results show that our proposed model can outperform state-of-the-art methods consistently and significantly.

AAAI Conference 2021 Conference Paper

Multi-modal Multi-label Emotion Recognition with Heterogeneous Hierarchical Message Passing

  • Dong Zhang
  • Xincheng Ju
  • Wei Zhang
  • Junhui Li
  • Shoushan Li
  • Qiaoming Zhu
  • Guodong Zhou

As an important research issue in affective computing community, multi-modal emotion recognition has become a hot topic in the last few years. However, almost all existing studies perform multiple binary classification for each emotion with focus on complete time series data. In this paper, we focus on multi-modal emotion recognition in a multilabel scenario. In this scenario, we consider not only the label-to-label dependency, but also the feature-to-label and modality-to-label dependencies. Particularly, we propose a heterogeneous hierarchical message passing network to effectively model above dependencies. Furthermore, we propose a new multi-modal multi-label emotion dataset based on partial time-series content to show predominant generalization of our model. Detailed evaluation demonstrates the effectiveness of our approach.

AAAI Conference 2021 Conference Paper

Non-asymptotic Convergence of Adam-type Reinforcement Learning Algorithms under Markovian Sampling

  • Huaqing Xiong
  • Tengyu Xu
  • Yingbin Liang
  • Wei Zhang

Despite the wide applications of Adam in reinforcement learning (RL), the theoretical convergence of Adam-type RL algorithms has not been established. This paper provides the first such convergence analysis for two fundamental RL algorithms of policy gradient (PG) and temporal difference (TD) learning that incorporate AMSGrad updates (a standard alternative of Adam in theoretical analysis), referred to as PG- AMSGrad and TD-AMSGrad, respectively. Moreover, our analysis focuses on Markovian sampling for both algorithms. We show that under general nonlinear function approximation, PG-AMSGrad with a constant stepsize converges to a neighborhood of a stationary point at the rate of O(1/T) (where T denotes the number of iterations), and with a diminishing stepsize converges exactly to a stationary point at the rate of O(log2 T/ √ T). Furthermore, under linear function approximation, TD-AMSGrad with a constant stepsize converges to a neighborhood of the global optimum at the rate of O(1/T), and with a diminishing stepsize converges exactly to the global optimum at the rate of O(log T/ √ T). Our study develops new techniques for analyzing the Adamtype RL algorithms under Markovian sampling.

NeurIPS Conference 2021 Conference Paper

One Million Scenes for Autonomous Driving: ONCE Dataset

  • Jiageng Mao
  • Niu Minzhe
  • ChenHan Jiang
  • hanxue liang
  • Jingheng Chen
  • Xiaodan Liang
  • Yamin Li
  • Chaoqiang Ye

Current perception models in autonomous driving have become notorious for greatly relying on a mass of annotated data to cover unseen cases and address the long-tail problem. On the other hand, learning from unlabeled large-scale collected data and incrementally self-training powerful recognition models have received increasing attention and may become the solutions of next-generation industry-level powerful and robust perception models in autonomous driving. However, the research community generally suffered from data inadequacy of those essential real-world scene data, which hampers the future exploration of fully/semi/self-supervised methods for 3D perception. In this paper, we introduce the ONCE (One millioN sCenEs) dataset for 3D object detection in the autonomous driving scenario. The ONCE dataset consists of 1 million LiDAR scenes and 7 million corresponding camera images. The data is selected from 144 driving hours, which is 20x longer than the largest 3D autonomous driving dataset available (\eg nuScenes and Waymo), and it is collected across a range of different areas, periods and weather conditions. To facilitate future research on exploiting unlabeled data for 3D detection, we additionally provide a benchmark in which we reproduce and evaluate a variety of self-supervised and semi-supervised methods on the ONCE dataset. We conduct extensive analyses on those methods and provide valuable observations on their performance related to the scale of used data. Data, code, and more information are available at \href{https: //once-for-auto-driving. github. io/index. html}{http: //www. once-for-auto-driving. com}.

NeurIPS Conference 2021 Conference Paper

Post-Training Quantization for Vision Transformer

  • Zhenhua Liu
  • Yunhe Wang
  • Kai Han
  • Wei Zhang
  • Siwei Ma
  • Wen Gao

Recently, transformer has achieved remarkable performance on a variety of computer vision applications. Compared with mainstream convolutional neural networks, vision transformers are often of sophisticated architectures for extracting powerful feature representations, which are more difficult to be developed on mobile devices. In this paper, we present an effective post-training quantization algorithm for reducing the memory storage and computational costs of vision transformers. Basically, the quantization task can be regarded as finding the optimal low-bit quantization intervals for weights and inputs, respectively. To preserve the functionality of the attention mechanism, we introduce a ranking loss into the conventional quantization objective that aims to keep the relative order of the self-attention results after quantization. Moreover, we thoroughly analyze the relationship between quantization loss of different layers and the feature diversity, and explore a mixed-precision quantization scheme by exploiting the nuclear norm of each attention map and output feature. The effectiveness of the proposed method is verified on several benchmark models and datasets, which outperforms the state-of-the-art post-training quantization algorithms. For instance, we can obtain an 81. 29% top-1 accuracy using DeiT-B model on ImageNet dataset with about 8-bit quantization. Code will be available at https: //gitee. com/mindspore/models/tree/master/research/cv/VT-PTQ.

NeurIPS Conference 2021 Conference Paper

Scalable Rule-Based Representation Learning for Interpretable Classification

  • Zhuo Wang
  • Wei Zhang
  • Ning Liu
  • Jianyong Wang

Rule-based models, e. g. , decision trees, are widely used in scenarios demanding high model interpretability for their transparent inner structures and good model expressivity. However, rule-based models are hard to optimize, especially on large data sets, due to their discrete parameters and structures. Ensemble methods and fuzzy/soft rules are commonly used to improve performance, but they sacrifice the model interpretability. To obtain both good scalability and interpretability, we propose a new classifier, named Rule-based Representation Learner (RRL), that automatically learns interpretable non-fuzzy rules for data representation and classification. To train the non-differentiable RRL effectively, we project it to a continuous space and propose a novel training method, called Gradient Grafting, that can directly optimize the discrete model using gradient descent. An improved design of logical activation functions is also devised to increase the scalability of RRL and enable it to discretize the continuous features end-to-end. Exhaustive experiments on nine small and four large data sets show that RRL outperforms the competitive interpretable approaches and can be easily adjusted to obtain a trade-off between classification accuracy and model complexity for different scenarios. Our code is available at: https: //github. com/12wang3/rrl.

ICML Conference 2021 Conference Paper

SCC: an efficient deep reinforcement learning agent mastering the game of StarCraft II

  • Xiangjun Wang
  • Junxiao Song
  • Penghui Qi
  • Peng Peng
  • Zhenkun Tang
  • Wei Zhang
  • Weimin Li
  • Xiongjun Pi

AlphaStar, the AI that reaches GrandMaster level in StarCraft II, is a remarkable milestone demonstrating what deep reinforcement learning can achieve in complex Real-Time Strategy (RTS) games. However, the complexities of the game, algorithms and systems, and especially the tremendous amount of computation needed are big obstacles for the community to conduct further research in this direction. We propose a deep reinforcement learning agent, StarCraft Commander (SCC). With order of magnitude less computation, it demonstrates top human performance defeating GrandMaster players in test matches and top professional players in a live event. Moreover, it shows strong robustness to various human strategies and discovers novel strategies unseen from human plays. In this paper, we’ll share the key insights and optimizations on efficient imitation learning and reinforcement learning for StarCraft II full game.

JBHI Journal 2021 Journal Article

SCCLRR: A Robust Computational Method for Accurate Clustering Single Cell RNA-Seq Data

  • Wei Zhang
  • Yuanyuan Li
  • Xiufen Zou

Single-cell RNA transcriptome data present a tremendous opportunity for studying the cellular heterogeneity. Identifying subpopulations based on scRNA-seq data is a hot topic in recent years, although many researchers have been focused on designing elegant computational methods for identifying new cell types; however, the performance of these methods is still unsatisfactory due to the high dimensionality, sparsity and noise of scRNA-seq data. In this study, we propose a new cell type detection method by learning a robust and accurate similarity matrix, named SCCLRR. The method simultaneously captures both global and local intrinsic properties of data based on a low rank representation (LRR) framework mathematical model. The integrated normalized Euclidean distance and cosine similarity are used to balance the intrinsic linear and nonlinear manifold of data in the local regularization term. To solve the non-convex optimization model, we present an iterative optimization procedure using the alternating direction method of multipliers (ADMM) algorithm. We evaluate the performance of the SCCLRR method on nine real scRNA-seq datasets and compare it with seven state-of-the-art methods. The simulation results show that the SCCLRR outperforms other methods and is robust and effective for clustering scRNA-seq data. (The code of SCCLRR is free available for academic https://github.com/wzhangwhu/SCCLRR ).

NeurIPS Conference 2021 Conference Paper

SODA10M: A Large-Scale 2D Self/Semi-Supervised Object Detection Dataset for Autonomous Driving

  • Jianhua Han
  • Xiwen Liang
  • Hang Xu
  • Kai Chen
  • Lanqing Hong
  • Jiageng Mao
  • Chaoqiang Ye
  • Wei Zhang

Aiming at facilitating a real-world, ever-evolving and scalable autonomous driving system, we present a large-scale dataset for standardizing the evaluation of different self-supervised and semi-supervised approaches by learning from raw data, which is the first and largest dataset to date. Existing autonomous driving systems heavily rely on `perfect' visual perception models (i. e. , detection) trained using extensive annotated data to ensure safety. However, it is unrealistic to elaborately label instances of all scenarios and circumstances (i. e. , night, extreme weather, cities) when deploying a robust autonomous driving system. Motivated by recent advances of self-supervised and semi-supervised learning, a promising direction is to learn a robust detection model by collaboratively exploiting large-scale unlabeled data and few labeled data. Existing datasets (i. e. , BDD100K, Waymo) either provide only a small amount of data or covers limited domains with full annotation, hindering the exploration of large-scale pre-trained models. Here, we release a Large-Scale 2D Self/semi-supervised Object Detection dataset for Autonomous driving, named as SODA10M, containing 10 million unlabeled images and 20K images labeled with 6 representative object categories. To improve diversity, the images are collected within 27833 driving hours under different weather conditions, periods and location scenes of 32 different cities. We provide extensive experiments and deep analyses of existing popular self-supervised and semi-supervised approaches, and some interesting findings in autonomous driving scope. Experiments show that SODA10M can serve as a promising pre-training dataset for different self-supervised learning methods, which gives superior performance when finetuning with different downstream tasks (i. e. , detection, semantic/instance segmentation) in autonomous driving domain. This dataset has been used to hold the ICCV2021 SSLAD challenge. More information can refer to https: //soda-2d. github. io.

AAAI Conference 2021 Conference Paper

Unsupervised Active Learning via Subspace Learning

  • Changsheng Li
  • Kaihang Mao
  • Lingyan Liang
  • Dongchun Ren
  • Wei Zhang
  • Ye Yuan
  • Guoren Wang

Unsupervised active learning has been an active research topic in machine learning community, with the purpose of choosing representative samples to be labelled in an unsupervised manner. Previous works usually take the minimization of data reconstruction loss as the criterion to select representative samples, by which the original inputs can be better approximated. However, data are often drawn from low-dimensional subspaces embedded in an arbitrary highdimensional space in many scenarios, thus it might severely bring in noise if attempting to precisely reconstruct all entries of one observation, leading to a suboptimal solution. In view of this, this paper proposes a novel unsupervised Active Learning model via Subspace Learning, called ALSL. In contrast to previous approaches, ALSL aims to discover low-rank structures of data, and then perform sample selection based on the learnt low-rank representations. To this end, we devise two different strategies and propose two corresponding formulations to select samples with and under low-rank sample representations, respectively. Since the proposed formulations involve several non-smooth regularization terms, we develop a simple but effective optimization procedure to solve them. Extensive experiments are performed on five publicly available datasets, and experimental results demonstrate the proposed first formulation achieves comparable performance with the state-of-the-arts, while the second formulation significantly outperforms them, achieving a 13% improvement over the second best baseline at most.

IJCAI Conference 2021 Conference Paper

Weakly-Supervised Spatio-Temporal Anomaly Detection in Surveillance Video

  • Jie Wu
  • Wei Zhang
  • Guanbin Li
  • Wenhao Wu
  • Xiao Tan
  • Yingying Li
  • Errui Ding
  • Liang Lin

In this paper, we introduce a novel task, referred to as Weakly-Supervised Spatio-Temporal Anomaly Detection (WSSTAD) in surveillance video. Specifically, given an untrimmed video, WSSTAD aims to localize a spatio-temporal tube (i. e. , a sequence of bounding boxes at consecutive times) that encloses the abnormal event, with only coarse video-level annotations as supervision during training. To address this challenging task, we propose a dual-branch network which takes as input the proposals with multi-granularities in both spatial-temporal domains. Each branch employs a relationship reasoning module to capture the correlation between tubes/videolets, which can provide rich contextual information and complex entity relationships for the concept learning of abnormal behaviors. Mutually-guided Progressive Refinement framework is set up to employ dual-path mutual guidance in a recurrent manner, iteratively sharing auxiliary supervision information across branches. It impels the learned concepts of each branch to serve as a guide for its counterpart, which progressively refines the corresponding branch and the whole framework. Furthermore, we contribute two datasets, i. e. , ST-UCF-Crime and STRA, consisting of videos containing spatio-temporal abnormal annotations to serve as the benchmarks for WSSTAD. We conduct extensive qualitative and quantitative evaluations to demonstrate the effectiveness of the proposed approach and analyze the key factors that contribute more to handle this task.

IS Journal 2020 Journal Article

A Data-Analytics Approach for Risk Evaluation in Peer-to-Peer Lending Platforms

  • Feng He
  • Yuelei Li
  • Tiecheng Xu
  • Libo Yin
  • Wei Zhang
  • Xiaotao Zhang

The goal of this article is to investigate the roles of individual behavior characteristics and Internet finance industry risk in the light of bank run theory for P2P. We know that risk evaluation is clearly important for peer-to-peer (P2P) lending platforms in China, as during the last two years, the industry has experienced thousands of platform crashes. Traditional approaches to evaluate enterprise risk are increasingly ineffective in this industry, due to the difficulty of assessing the real information. In addition, the Internet business model makes it possible to record new kinds of information. By applying a data-driven analytics method, we build an intelligent risk evaluation model for P2P platforms that have comparable targeting platforms. The case study shows that our risk evaluation method can generate early warning signals regarding platform or industry risk, which is able to provide effective supporting for P2P business in practice.

NeurIPS Conference 2020 Conference Paper

A Decentralized Parallel Algorithm for Training Generative Adversarial Nets

  • Mingrui Liu
  • Wei Zhang
  • Youssef Mroueh
  • Xiaodong Cui
  • Jarret Ross
  • Tianbao Yang
  • Payel Das

Generative Adversarial Networks (GANs) are a powerful class of generative models in the deep learning community. Current practice on large-scale GAN training utilizes large models and distributed large-batch training strategies, and is implemented on deep learning frameworks (e. g. , TensorFlow, PyTorch, etc. ) designed in a centralized manner. In the centralized network topology, every worker needs to either directly communicate with the central node or indirectly communicate with all other workers in every iteration. However, when the network bandwidth is low or network latency is high, the performance would be significantly degraded. Despite recent progress on decentralized algorithms for training deep neural networks, it remains unclear whether it is possible to train GANs in a decentralized manner. The main difficulty lies at handling the nonconvex-nonconcave min-max optimization and the decentralized communication simultaneously. In this paper, we address this difficulty by designing the \textbf{first gradient-based decentralized parallel algorithm} which allows workers to have multiple rounds of communications in one iteration and to update the discriminator and generator simultaneously, and this design makes it amenable for the convergence analysis of the proposed decentralized algorithm. Theoretically, our proposed decentralized algorithm is able to solve a class of non-convex non-concave min-max problems with provable non-asymptotic convergence to first-order stationary point. Experimental results on GANs demonstrate the effectiveness of the proposed algorithm.

IJCAI Conference 2020 Conference Paper

Analysis of Q-learning with Adaptation and Momentum Restart for Gradient Descent

  • Bowen Weng
  • Huaqing Xiong
  • Yingbin Liang
  • Wei Zhang

Existing convergence analyses of Q-learning mostly focus on the vanilla stochastic gradient descent (SGD) type of updates. Despite the Adaptive Moment Estimation (Adam) has been commonly used for practical Q-learning algorithms, there has not been any convergence guarantee provided for Q-learning with such type of updates. In this paper, we first characterize the convergence rate for Q-AMSGrad, which is the Q-learning algorithm with AMSGrad update (a commonly adopted alternative of Adam for theoretical analysis). To further improve the performance, we propose to incorporate the momentum restart scheme to Q-AMSGrad, resulting in the so-called Q-AMSGradR algorithm. The convergence rate of Q-AMSGradR is also established. Our experiments on a linear quadratic regulator problem demonstrate that the two proposed Q-learning algorithms outperform the vanilla Q-learning with SGD updates. The two algorithms also exhibit significantly better performance than the DQN learning method over a batch of Atari 2600 games.

AAAI Conference 2020 System Paper

Automatic Car Damage Assessment System: Reading and Understanding Videos as Professional Insurance Inspectors

  • Wei Zhang
  • Yuan Cheng
  • Xin Guo
  • Qingpei Guo
  • Jian Wang
  • Qing Wang
  • Chen Jiang
  • Meng Wang

We demonstrate a car damage assessment system in car insurance field based on artificial intelligence techniques, which can exempt insurance inspectors from checking cars on site and help people without professional knowledge to evaluate car damages when accidents happen. Unlike existing approaches, we utilize videos instead of photos to interact with users to make the whole procedure as simple as possible. We adopt object and video detection and segmentation techniques in computer vision, and take advantage of multiple frames extracted from videos to achieve high damage recognition accuracy. The system uploads video streams captured by mobile devices, recognizes car damage on the cloud asynchronously and then returns damaged components and repair costs to users. The system evaluates car damages and returns results automatically and effectively in seconds, which reduces laboratory costs and decreases insurance claim time significantly.

NeurIPS Conference 2020 Conference Paper

Finite-Time Analysis for Double Q-learning

  • Huaqing Xiong
  • Lin Zhao
  • Yingbin Liang
  • Wei Zhang

Although Q-learning is one of the most successful algorithms for finding the best action-value function (and thus the optimal policy) in reinforcement learning, its implementation often suffers from large overestimation of Q-function values incurred by random sampling. The double Q-learning algorithm proposed in~\citet{hasselt2010double} overcomes such an overestimation issue by randomly switching the update between two Q-estimators, and has thus gained significant popularity in practice. However, the theoretical understanding of double Q-learning is rather limited. So far only the asymptotic convergence has been established, which does not characterize how fast the algorithm converges. In this paper, we provide the first non-asymptotic (i. e. , finite-time) analysis for double Q-learning. We show that both synchronous and asynchronous double Q-learning are guaranteed to converge to an $\epsilon$-accurate neighborhood of the global optimum by taking $\tilde{\Omega}\left(\left( \frac{1}{(1-\gamma)^6\epsilon^2}\right)^{\frac{1}{\omega}} +\left(\frac{1}{1-\gamma}\right)^{\frac{1}{1-\omega}}\right)$ iterations, where $\omega\in(0, 1)$ is the decay parameter of the learning rate, and $\gamma$ is the discount factor. Our analysis develops novel techniques to derive finite-time bounds on the difference between two inter-connected stochastic processes, which is new to the literature of stochastic approximation.

IJCAI Conference 2020 Conference Paper

Global Structure and Local Semantics-Preserved Embeddings for Entity Alignment

  • Hao Nie
  • Xianpei Han
  • Le Sun
  • Chi Man Wong
  • Qiang Chen
  • Suhui Wu
  • Wei Zhang

Entity alignment (EA) aims to identify entities located in different knowledge graphs (KGs) that refer to the same real-world object. To learn the entity representations, most EA approaches rely on either translation-based methods which capture the local relation semantics of entities or graph convolutional networks (GCNs), which exploit the global KG structure. Afterward, the aligned entities are identified based on their distances. In this paper, we propose to jointly leverage the global KG structure and entity-specific relational triples for better entity alignment. Specifically, a global structure and local semantics preserving network is proposed to learn entity representations in a coarse-to-fine manner. Experiments on several real-world datasets show that our method significantly outperforms other entity alignment approaches and achieves the new state-of-the-art performance.

AAAI Conference 2020 Conference Paper

Improving Domain-Adapted Sentiment Classification by Deep Adversarial Mutual Learning

  • Qianming Xue
  • Wei Zhang
  • Hongyuan Zha

Domain-adapted sentiment classification refers to training on a labeled source domain to well infer document-level sentiment on an unlabeled target domain. Most existing relevant models involve a feature extractor and a sentiment classifier, where the feature extractor works towards learning domaininvariant features from both domains, and the sentiment classifier is trained only on the source domain to guide the feature extractor. As such, they lack a mechanism to use sentiment polarity lying in the target domain. To improve domainadapted sentiment classification by learning sentiment from the target domain as well, we devise a novel deep adversarial mutual learning approach involving two groups of feature extractors, domain discriminators, sentiment classifiers, and label probers. The domain discriminators enable the feature extractors to obtain domain-invariant features. Meanwhile, the label prober in each group explores document sentiment polarity of the target domain through the sentiment prediction generated by the classifier in the peer group, and guides the learning of the feature extractor in its own group. The proposed approach achieves the mutual learning of the two groups in an end-to-end manner. Experiments on multiple public datasets indicate our method obtains the state-of-theart performance, validating the effectiveness of mutual learning through label probers.

AAAI Conference 2020 Conference Paper

Improving Neural Relation Extraction with Positive and Unlabeled Learning

  • Zhengqiu He
  • Wenliang Chen
  • Yuyi Wang
  • Wei Zhang
  • Guanchun Wang
  • Min Zhang

We present a novel approach to improve the performance of distant supervision relation extraction with Positive and Unlabeled (PU) Learning. This approach first applies reinforcement learning to decide whether a sentence is positive to a given relation, and then positive and unlabeled bags are constructed. In contrast to most previous studies, which mainly use selected positive instances only, we make full use of unlabeled instances and propose two new representations for positive and unlabeled bags. These two representations are then combined in an appropriate way to make bag-level prediction. Experimental results on a widely used real-world dataset demonstrate that this new approach indeed achieves significant and consistent improvements as compared to several competitive baselines.

NeurIPS Conference 2020 Conference Paper

Kernel Based Progressive Distillation for Adder Neural Networks

  • Yixing Xu
  • Chang Xu
  • Xinghao Chen
  • Wei Zhang
  • Chunjing Xu
  • Yunhe Wang

Adder Neural Networks (ANNs) which only contain additions bring us a new way of developing deep neural networks with low energy consumption. Unfortunately, there is an accuracy drop when replacing all convolution filters by adder filters. The main reason here is the optimization difficulty of ANNs using $\ell_1$-norm, in which the estimation of gradient in back propagation is inaccurate. In this paper, we present a novel method for further improving the performance of ANNs without increasing the trainable parameters via a progressive kernel based knowledge distillation (PKKD) method. A convolutional neural network (CNN) with the same architecture is simultaneously initialized and trained as a teacher network, features and weights of ANN and CNN will be transformed to a new space to eliminate the accuracy drop. The similarity is conducted in a higher-dimensional space to disentangle the difference of their distributions using a kernel based method. Finally, the desired ANN is learned based on the information from both the ground-truth and teacher, progressively. The effectiveness of the proposed method for learning ANN with higher performance is then well-verified on several benchmarks. For instance, the ANN-50 trained using the proposed PKKD method obtains a 76. 8\% top-1 accuracy on ImageNet dataset, which is 0. 6\% higher than that of the ResNet-50.

AAAI Conference 2020 Conference Paper

Knowledge Graph Alignment Network with Gated Multi-Hop Neighborhood Aggregation

  • Zequn Sun
  • Chengming Wang
  • Wei Hu
  • Muhao Chen
  • Jian Dai
  • Wei Zhang
  • Yuzhong Qu

Graph neural networks (GNNs) have emerged as a powerful paradigm for embedding-based entity alignment due to their capability of identifying isomorphic subgraphs. However, in real knowledge graphs (KGs), the counterpart entities usually have non-isomorphic neighborhood structures, which easily causes GNNs to yield different representations for them. To tackle this problem, we propose a new KG alignment network, namely AliNet, aiming at mitigating the non-isomorphism of neighborhood structures in an end-to-end manner. As the direct neighbors of counterpart entities are usually dissimilar due to the schema heterogeneity, AliNet introduces distant neighbors to expand the overlap between their neighborhood structures. It employs an attention mechanism to highlight helpful distant neighbors and reduce noises. Then, it controls the aggregation of both direct and distant neighborhood information using a gating mechanism. We further propose a relation loss to refine entity representations. We perform thorough experiments with detailed ablation studies and analyses on five entity alignment datasets, demonstrating the effectiveness of AliNet.

AAAI Conference 2020 Conference Paper

Learning Long- and Short-Term User Literal-Preference with Multimodal Hierarchical Transformer Network for Personalized Image Caption

  • Wei Zhang
  • Yue Ying
  • Pan Lu
  • Hongyuan Zha

Personalized image caption, a natural extension of the standard image caption task, requires to generate brief image descriptions tailored for users’ writing style and traits, and is more practical to meet users’ real demands. Only a few recent studies shed light on this crucial task and learn static user representations to capture their long-term literal-preference. However, it is insufficient to achieve satisfactory performance due to the intrinsic existence of not only long-term user literal-preference, but also short-term literal-preference which is associated with users’ recent states. To bridge this gap, we develop a novel multimodal hierarchical transformer network (MHTN) for personalized image caption in this paper. It learns short-term user literal-preference based on users’ recent captions through a short-term user encoder at the low level. And at the high level, the multimodal encoder integrates target image representations with short-term literalpreference, as well as long-term literal-preference learned from user IDs. These two encoders enjoy the advantages of the powerful transformer networks. Extensive experiments on two real datasets show the effectiveness of considering two types of user literal-preference simultaneously and better performance over the state-of-the-art models.

YNICL Journal 2020 Journal Article

Mapping typical and hypokinetic dysarthric speech production network using a connected speech paradigm in functional MRI

  • Shalini Narayana
  • Megan B. Parsons
  • Wei Zhang
  • Crystal Franklin
  • Katherine Schiller
  • Asim F. Choudhri
  • Peter T. Fox
  • Mark S. LeDoux

O-water PET in the same group of typical speakers and another HKD cohort (n = 10, 2 females). The fMRI overt connected speech paradigm did not result in excessive motion artifacts and successfully identified the same brain areas demonstrated in the PET studies in the two cohorts. The SPN derived in fMRI demonstrated significant spatial overlap with the corresponding PET derived maps (typical speakers: r = 0.52; speakers with HKD: r = 0.43) and identified the components of the neural circuit of speech production belonging to the feedforward and feedback subsystems. The fMRI study in speakers with HKD identified significantly decreased activity in critical feedforward (bilateral dorsal premotor and motor cortices) and feedback (auditory and somatosensory areas) subsystems replicating previous PET study findings in this cohort. These results demonstrate that the overt connected speech paradigm is feasible during fMRI and can accurately localize the neural substrates of typical and disordered speech production. Our fMRI paradigm should prove useful for study of motor speech and voice disorders, including stuttering, apraxia of speech, dysarthria, and spasmodic dysphonia.

NeurIPS Conference 2020 Conference Paper

Model Rubik’s Cube: Twisting Resolution, Depth and Width for TinyNets

  • Kai Han
  • Yunhe Wang
  • Qiulin Zhang
  • Wei Zhang
  • Chunjing Xu
  • Tong Zhang

To obtain excellent deep neural architectures, a series of techniques are carefully designed in EfficientNets. The giant formula for simultaneously enlarging the resolution, depth and width provides us a Rubik’s cube for neural networks. So that we can find networks with high efficiency and excellent performance by twisting the three dimensions. This paper aims to explore the twisting rules for obtaining deep neural networks with minimum model sizes and computational costs. Different from the network enlarging, we observe that resolution and depth are more important than width for tiny networks. Therefore, the original method, \ie the compound scaling in EfficientNet is no longer suitable. To this end, we summarize a tiny formula for downsizing neural architectures through a series of smaller models derived from the EfficientNet-B0 with the FLOPs constraint. Experimental results on the ImageNet benchmark illustrate that our TinyNet performs much better than the smaller version of EfficientNets using the inversed giant formula. For instance, our TinyNet-E achieves a 59. 9\% Top-1 accuracy with only 24M FLOPs, which is about 1. 9\% higher than that of the previous best MobileNetV3 with similar computational cost. Code will be available at \url{https: //github. com/huawei-noah/CV-Backbones/tree/master/tinynet}, and \url{https: //gitee. com/mindspore/mindspore/tree/master/model_zoo/research/cv/tinynet}.

TIST Journal 2020 Journal Article

Multi-Task Learning for Entity Recommendation and Document Ranking in Web Search

  • Jizhou Huang
  • Haifeng Wang
  • Wei Zhang
  • Ting Liu

Entity recommendation, providing users with an improved search experience by proactively recommending related entities to a given query, has become an indispensable feature of today’s Web search engine. Existing studies typically only consider the query issued at the current timestep while ignoring the in-session user search behavior (short-term search history) or historical user search behavior across all sessions (long-term search history) when generating entity recommendations. As a consequence, they may fail to recommend entities of interest relevant to a user’s actual information need. In this work, we believe that both short-term and long-term search history convey valuable evidence that could help understand the user’s search intent behind a query, and take both of them into consideration for entity recommendation. Furthermore, there has been little work on exploring whether the use of other companion tasks in Web search such as document ranking as auxiliary tasks could improve the performance of entity recommendation. To this end, we propose a multi-task learning framework with deep neural networks (DNNs) to jointly learn and optimize two companion tasks in Web search engines: entity recommendation and document ranking, which can be easily trained in an end-to-end manner. Specifically, we regard document ranking as an auxiliary task to improve the main task of entity recommendation, where the representations of queries, sessions, and users are shared across all tasks and optimized by the multi-task objective during training. We evaluate our approach using large-scale, real-world search logs of a widely-used commercial Web search engine. We also performed extensive ablation experiments over a number of facets of the proposed multi-task DNN model to figure out their relative importance. The experimental results show that both short-term and long-term search history can bring significant improvements in recommendation effectiveness, and the combination of both outperforms using either of them individually. In addition, the experiments show that the performance of both entity recommendation and document ranking can be significantly improved, which demonstrates the effectiveness of using multi-task learning to jointly optimize the two companion tasks in Web search.

NeurIPS Conference 2020 Conference Paper

Online Decision Based Visual Tracking via Reinforcement Learning

  • Ke Song
  • Wei Zhang
  • Ran Song
  • Yibin Li

A deep visual tracker is typically based on either object detection or template matching while each of them is only suitable for a particular group of scenes. It is straightforward to consider fusing them together to pursue more reliable tracking. However, this is not wise as they follow different tracking principles. Unlike previous fusion-based methods, we propose a novel ensemble framework, named DTNet, with an online decision mechanism for visual tracking based on hierarchical reinforcement learning. The decision mechanism substantiates an intelligent switching strategy where the detection and the template trackers have to compete with each other to conduct tracking within different scenes that they are adept in. Besides, we present a novel detection tracker which avoids the common issue of incorrect proposal. Extensive results show that our DTNet achieves state-of-the-art tracking performance as well as good balance between accuracy and efficiency. The project website is available at https: //vsislab. github. io/DTNet/.

NeurIPS Conference 2020 Conference Paper

Residual Distillation: Towards Portable Deep Neural Networks without Shortcuts

  • Guilin Li
  • Junlei Zhang
  • Yunhe Wang
  • Chuanjian Liu
  • Matthias Tan
  • Yunfeng Lin
  • Wei Zhang
  • Jiashi Feng

By transferring both features and gradients between different layers, shortcut connections explored by ResNets allow us to effectively train very deep neural networks up to hundreds of layers. However, the additional computation costs induced by those shortcuts are often overlooked. For example, during online inference, the shortcuts in ResNet-50 account for about 40 percent of the entire memory usage on feature maps, because the features in the preceding layers cannot be released until the subsequent calculation is completed. In this work, for the first time, we consider training the CNN models with shortcuts and deploying them without. In particular, we propose a novel joint-training framework to train plain CNN by leveraging the gradients of the ResNet counterpart. During forward step, the feature maps of the early stages of plain CNN are passed through later stages of both itself and the ResNet counterpart to calculate the loss. During backpropagation, gradients calculated from a mixture of these two parts are used to update the plainCNN network to solve the gradient vanishing problem. Extensive experiments on ImageNet/CIFAR10/CIFAR100 demonstrate that the plainCNN network without shortcuts generated by our approach can achieve the same level of accuracy as that of the ResNet baseline while achieving about $1. 4\times $ speed-up and $1. 25\times$ memory reduction. We also verified the feature transferability of our ImageNet pretrained plain-CNN network by fine-tuning it on MIT 67 and Caltech 101. Our results show that the performance of the plain-CNN is slightly higher than that of its baseline ResNet-50 on these two datasets. The codes are in: \href{https: //github. com/leoozy/JointRD_Neurips2020}{https: //github. com/leoozy/JointRD\_Neurips2020}

NeurIPS Conference 2020 Conference Paper

ScaleCom: Scalable Sparsified Gradient Compression for Communication-Efficient Distributed Training

  • Chia-Yu Chen
  • Jiamin Ni
  • Songtao Lu
  • Xiaodong Cui
  • Pin-Yu Chen
  • Xiao Sun
  • Naigang Wang
  • Swagath Venkataramani

Large-scale distributed training of Deep Neural Networks (DNNs) on state-of-the-art platforms are expected to be severely communication constrained. To overcome this limitation, numerous gradient compression techniques have been proposed and have demonstrated high compression ratios. However, most existing compression methods do not scale well to large scale distributed systems (due to gradient build-up) and / or lack evaluations in large datasets. To mitigate these issues, we propose a new compression technique, Scalable Sparsified Gradient Compression (ScaleComp), that (i) leverages similarity in the gradient distribution amongst learners to provide a commutative compressor and keep communication cost constant to worker number and (ii) includes low-pass filter in local gradient accumulations to mitigate the impacts of large batch size training and significantly improve scalability. Using theoretical analysis, we show that ScaleComp provides favorable convergence guarantees and is compatible with gradient all-reduce techniques. Furthermore, we experimentally demonstrate that ScaleComp has small overheads, directly reduces gradient traffic and provides high compression rates (70-150X) and excellent scalability (up to 64-80 learners and 10X larger batch sizes over normal training) across a wide range of applications (image, language, and speech) without significant accuracy loss.

AAAI Conference 2020 Conference Paper

SM-NAS: Structural-to-Modular Neural Architecture Search for Object Detection

  • Lewei Yao
  • Hang Xu
  • Wei Zhang
  • Xiaodan Liang
  • Zhenguo Li

The state-of-the-art object detection method is complicated with various modules such as backbone, RPN, feature fusion neck and RCNN head, where each module may have different designs and structures. How to leverage the computational cost and accuracy trade-off for the structural combination as well as the modular selection of multiple modules? Neural architecture search (NAS) has shown great potential in finding an optimal solution. Existing NAS works for object detection only focus on searching better design of a single module such as backbone or feature fusion neck, while neglecting the balance of the whole system. In this paper, we present a twostage coarse-to-fine searching strategy named Structural-to- Modular NAS (SM-NAS) for searching a GPU-friendly design of both an efficient combination of modules and better modular-level architecture for object detection. Specifically, Structural-level searching stage first aims to find an efficient combination of different modules; Modular-level searching stage then evolves each specific module and pushes the Pareto front forward to a faster task-specific network. We consider a multi-objective search where the search space covers many popular designs of detection methods. We directly search a detection backbone without pre-trained models or any proxy task by exploring a fast training from scratch strategy. The resulting architectures dominate state-of-the-art object detection systems in both inference time and accuracy and demonstrate the effectiveness on multiple detection datasets, e. g. halving the inference time with additional 1% mAP improvement compared to FPN and reaching 46% mAP with the similar inference time of MaskRCNN.

YNIMG Journal 2020 Journal Article

Targeting brain functions from the scalp: Transcranial brain atlas based on large-scale fMRI data synthesis

  • Yihan Jiang
  • Zheng Li
  • Yang Zhao
  • Xiang Xiao
  • Wei Zhang
  • Peipei Sun
  • Yihong Yang
  • Chaozhe Zhu

Transcranial brain mapping techniques, such as functional near-infrared spectroscopy (fNIRS) and transcranial magnetic stimulation (TMS), have been playing an increasingly important role in studies of human brain functions. Given a brain function of interest, fNIRS probes and TMS coils should be properly placed on the scalp to ensure that the function is effectively measured or modulated. However, since brain activity is inside the skull and invisible to the researcher during placement, this blind targeting may cause the device to partially or completely miss the functional target, resulting in inconsistent experimental results and divergent clinical outcomes, especially when participants’ structural MRI data are not available. To address this issue, we propose here a framework for targeting a designated function directly from the scalp. First, a functional brain atlas for the targeted brain function is constructed via a meta-analysis of large-scale functional magnetic resonance imaging datasets. Second, the functional brain atlas is presented on the scalp surface by using a transcranial mapping previously established from an structural MRI dataset (n ​= ​114), resulting in a novel functional transcranial brain atlas (fTBA). Finally, a low-cost, portable scalp-navigation system is used to localize the transcranial device on the individual’s scalp with the guidance of the fTBA. To demonstrate the feasibility of the targeting framework, both fNIRS and TMS mapping experiments were conducted. The results show that fTBA-guided fNIRS positioning can detect functional activity with high sensitivity and specificity for working memory and motor systems; Moreover, compared with traditional TMS targeting approaches (e. g. the International 10–20 System and the conventional 5-cm rule), the fTBA suggested motor stimulation site is closesr to both the motor hotspot and the center of gravity of motor evoked potentials (MEP-COG). In summary, the proposed method unblinds the transcranial function targeting process using prior information, providing an effective and straightforward approach to transcranial brain mapping studies, especially those without participants’ structural MRI data.

AAAI Conference 2020 Conference Paper

Transparent Classification with Multilayer Logical Perceptrons and Random Binarization

  • Zhuo Wang
  • Wei Zhang
  • Ning Liu
  • Jianyong Wang

Models with transparent inner structure and high classification performance are required to reduce potential risk and provide trust for users in domains like health care, finance, security, etc. However, existing models are hard to simultaneously satisfy the above two properties. In this paper, we propose a new hierarchical rule-based model for classi- fication tasks, named Concept Rule Sets (CRS), which has both a strong expressive ability and a transparent inner structure. To address the challenge of efficiently learning the nondifferentiable CRS model, we propose a novel neural network architecture, Multilayer Logical Perceptron (MLLP), which is a continuous version of CRS. Using MLLP and the Random Binarization (RB) method we proposed, we can search the discrete solution of CRS in continuous space using gradient descent and ensure the discrete CRS acts almost the same as the corresponding continuous MLLP. Experiments on 12 public data sets show that CRS outperforms the state-of-theart approaches and the complexity of the learned CRS is close to the simple decision tree.

AAAI Conference 2020 Conference Paper

ZoomNet: Part-Aware Adaptive Zooming Neural Network for 3D Object Detection

  • Zhenbo Xu
  • Wei Zhang
  • Xiaoqing Ye
  • Xiao Tan
  • Wei Yang
  • Shilei Wen
  • Errui Ding
  • Ajin Meng

3D object detection is an essential task in autonomous driving and robotics. Though great progress has been made, challenges remain in estimating 3D pose for distant and occluded objects. In this paper, we present a novel framework named ZoomNet for stereo imagery-based 3D detection. The pipeline of ZoomNet begins with an ordinary 2D object detection model which is used to obtain pairs of leftright bounding boxes. To further exploit the abundant texture cues in rgb images for more accurate disparity estimation, we introduce a conceptually straight-forward module – adaptive zooming, which simultaneously resizes 2D instance bounding boxes to a unified resolution and adjusts the camera intrinsic parameters accordingly. In this way, we are able to estimate higher-quality disparity maps from the resized box images then construct dense point clouds for both nearby and distant objects. Moreover, we introduce to learn part locations as complementary features to improve the resistance against occlusion and put forward the 3D fitting score to better estimate the 3D detection quality. Extensive experiments on the popular KITTI 3D detection dataset indicate ZoomNet surpasses all previous state-of-the-art methods by large margins (improved by 9. 4% on APbv (IoU=0. 7) over pseudo-LiDAR). Ablation study also demonstrates that our adaptive zooming strategy brings an improvement of over 10% on AP3d (IoU=0. 7). In addition, since the official KITTI benchmark lacks fine-grained annotations like pixel-wise part locations, we also present our KFG dataset by augmenting KITTI with detailed instance-wise annotations including pixel-wise part location, pixel-wise disparity, etc. . Both the KFG dataset and our codes will be publicly available at https: //github. com/detectRecog/ZoomNet.

YNIMG Journal 2019 Journal Article

Acute stress alters the ‘default’ brain processing

  • Wei Zhang
  • Mahur M. Hashemi
  • Reinoud Kaldewaij
  • Saskia B.J. Koch
  • Christian Beckmann
  • Floris Klumpers
  • Karin Roelofs

Active adaptation to acute stress is essential for coping with daily life challenges. The stress hormone cortisol, as well as large scale re-allocations of brain resources have been implicated in this adaptation. Stress-induced shifts between large-scale brain networks, including salience (SN), central executive (CEN) and default mode networks (DMN), have however been demonstrated mainly under task-conditions. It remains unclear whether such network shifts also occur in the absence of ongoing task-demands, and most critically, whether these network shifts are predictive of individual variation in the magnitude of cortisol stress-responses. In a sample of 335 healthy participants, we investigated stress-induced functional connectivity changes (delta-FC) of the SN, CEN and DMN, using resting-state fMRI data acquired before and after a socially evaluated cold-pressor test and a mental arithmetic task. To investigate which network changes are associated with acute stress, we evaluated the association between cortisol increase and delta-FC of each network. Stress-induced cortisol increase was associated with increased connectivity within the SN, but with decreased coupling of DMN at both local (within network) and global (synchronization with brain regions also outside the network) levels. These findings indicate that acute stress prompts immediate connectivity changes in large-scale resting-state networks, including the SN and DMN in the absence of explicit ongoing task-demands. Most interestingly, this brain reorganization is coupled with individuals’ cortisol stress-responsiveness. These results suggest that the observed stress-induced network reorganization might function as a neural mechanism determining individual stress reactivity and, therefore, it could serve as a promising marker for future studies on stress resilience and vulnerability.

IJCAI Conference 2019 Conference Paper

End-to-End Multi-Perspective Matching for Entity Resolution

  • Cheng Fu
  • Xianpei Han
  • Le Sun
  • Bo Chen
  • Wei Zhang
  • Suhui Wu
  • Hao Kong

Entity resolution (ER) aims to identify data records referring to the same real-world entity. Due to the heterogeneity of entity attributes and the diversity of similarity measures, one main challenge of ER is how to select appropriate similarity measures for different attributes. Previous ER methods usually employ heuristic similarity selection algorithms, which are highly specialized to specific ER problems and are hard to be generalized to other situations. Furthermore, previous studies usually perform similarity learning and similarity selection independently, which often result in error propagation and are hard to be optimized globally. To resolve the above problems, this paper proposes an end-to-end multi-perspective entity matching model, which can adaptively select optimal similarity measures for heterogenous attributes by jointly learning and selecting similarity measures in an end-to-end way. Experiments on two real-world datasets show that our method significantly outperforms previous ER methods.

IJCAI Conference 2019 Conference Paper

Exploring the Task Cooperation in Multi-goal Visual Navigation

  • Yuechen Wu
  • Zhenhuan Rao
  • Wei Zhang
  • Shijian Lu
  • Weizhi Lu
  • Zheng-Jun Zha

Learning to adapt to a series of different goals in visual navigation is challenging. In this work, we present a model-embedded actor-critic architecture for the multi-goal visual navigation task. To enhance the task cooperation in multi-goal learning, we introduce two new designs to the reinforcement learning scheme: inverse dynamics model (InvDM) and multi-goal co-learning (MgCl). Specifically, InvDM is proposed to capture the navigation-relevant association between state and goal, and provide additional training signals to relieve the sparse reward issue. MgCl aims at improving the sample efficiency and supports the agent to learn from unintentional positive experiences. Extensive results on the interactive platform AI2-THOR demonstrate that the proposed method converges faster than state-of-the-art methods while producing more direct routes to navigate to the goal. The video demonstration is available at: https: //youtube. com/channel/UCtpTMOsctt3yPzXqe_JMD3w/videos.

AAAI Conference 2019 Conference Paper

Hierarchical Photo-Scene Encoder for Album Storytelling

  • Bairui Wang
  • Lin Ma
  • Wei Zhang
  • Wenhao Jiang
  • Feng Zhang

In this paper, we propose a novel model with a hierarchical photo-scene encoder and a reconstructor for the task of album storytelling. The photo-scene encoder contains two subencoders, namely the photo and scene encoders, which are stacked together and behave hierarchically to fully exploit the structure information of the photos within an album. Specifically, the photo encoder generates semantic representation for each photo while exploiting temporal relationships among them. The scene encoder, relying on the obtained photo representations, is responsible for detecting the scene changes and generating scene representations. Subsequently, the decoder dynamically and attentively summarizes the encoded photo and scene representations to generate a sequence of album representations, based on which a story consisting of multiple coherent sentences is generated. In order to fully extract the useful semantic information from an album, a reconstructor is employed to reproduce the summarized album representations based on the hidden states of the decoder. The proposed model can be trained in an end-to-end manner, which results in an improved performance over the state-of-the-arts on the public visual storytelling (VIST) dataset. Ablation studies further demonstrate the effectiveness of the proposed hierarchical photo-scene encoder and reconstructor.

NeurIPS Conference 2019 Conference Paper

Hybrid 8-bit Floating Point (HFP8) Training and Inference for Deep Neural Networks

  • Xiao Sun
  • Jungwook Choi
  • Chia-Yu Chen
  • Naigang Wang
  • Swagath Venkataramani
  • Vijayalakshmi (Viji) Srinivasan
  • Xiaodong Cui
  • Wei Zhang

Reducing the numerical precision of data and computation is extremely effective in accelerating deep learning training workloads. Towards this end, 8-bit floating point representations (FP8) were recently proposed for DNN training. However, its applicability was demonstrated on a few selected models only and significant degradation is observed when popular networks such as MobileNet and Transformer are trained using FP8. This degradation is due to the inherent precision requirement difference in the forward and backward passes of DNN training. Using theoretical insights, we propose a hybrid FP8 (HFP8) format and DNN end-to-end distributed training procedure. We demonstrate, using HFP8, the successful training of deep learning models across a whole spectrum of applications including Image Classification, Object Detection, Language and Speech without accuracy degradation. Finally, we demonstrate that, by using the new 8 bit format, we can directly quantize a pre-trained model down to 8-bits without losing accuracy by simply fine-tuning batch normalization statistics. These novel techniques enable a new generations of 8-bit hardware that are robust for building and deploying neural network models.

IJCAI Conference 2019 Conference Paper

MSR: Multi-Scale Shape Regression for Scene Text Detection

  • Chuhui Xue
  • Shijian Lu
  • Wei Zhang

State-of-the-art scene text detection techniques predict quadrilateral boxes that are prone to localization errors while dealing with straight or curved text lines of different orientations and lengths in scenes. This paper presents a novel multi-scale shape regression network (MSR) that is capable of locating text lines of different lengths, shapes and curvatures in scenes. The proposed MSR detects scene texts by predicting dense text boundary points that inherently capture the location and shape of text lines accurately and are also more tolerant to the variation of text line length as compared with the state of the arts using proposals or segmentation. Additionally, the multi-scale network extracts and fuses features at different scales which demonstrates superb tolerance to the text scale variation. Extensive experiments over several public datasets show that the proposed MSR obtains superior detection performance for both curved and straight text lines of different lengths and orientations.

YNIMG Journal 2019 Journal Article

Network analysis reveals disrupted functional brain circuitry in drug-naive social anxiety disorder

  • Xun Yang
  • Jin Liu
  • Yajing Meng
  • Mingrui Xia
  • Zaixu Cui
  • Xi Wu
  • Xinyu Hu
  • Wei Zhang

Social anxiety disorder (SAD) is a common and disabling condition characterized by excessive fear and avoidance of public scrutiny. Psychoradiology studies have suggested that the emotional and behavior deficits in SAD are associated with abnormalities in regional brain function and functional connectivity. However, little is known about whether intrinsic functional brain networks in patients with SAD are topologically disrupted. Here, we collected resting-state fMRI data from 33 drug-naive patients with SAD and 32 healthy controls (HC), constructed functional networks with 34 predefined regions based on previous meta-analytic research with task-based fMRI in SAD, and performed network-based statistic and graph-theory analyses. The network-based statistic analysis revealed a single connected abnormal circuitry including the frontolimbic circuit (termed the “fear circuit”, including the dorsolateral prefrontal cortex, ventral medial prefrontal cortex and insula) and posterior cingulate/occipital areas supporting perceptual processing. In this single altered network, patients with SAD had higher functional connectivity than HC. At the global level, graph-theory analysis revealed that the patients exhibited a lower normalized characteristic path length than HC, which suggests a disorder-related shift of network topology toward randomized configurations. SAD-related deficits in nodal degree, efficiency and participation coefficient were detected in the parahippocampal gyrus, posterior cingulate cortex, dorsolateral prefrontal cortex, insula and the calcarine sulcus. Aspects of abnormal connectivity were associated with anxiety symptoms. These findings highlight the aberrant topological organization of functional brain network organization in SAD, which provides insights into the neural mechanisms underlying excessive fear and avoidance of social interactions in patients with debilitating social anxiety.

AIJ Journal 2019 Journal Article

Syntax-aware entity representations for neural relation extraction

  • Zhengqiu He
  • Wenliang Chen
  • Zhenghua Li
  • Wei Zhang
  • Hao Shao
  • Min Zhang

Distantly supervised relation extraction has been widely used to find novel relational facts between entities from text, and can be easily scaled to very large corpora. Previous studies on neural relation extraction treat this task as a multi-instance learning problem, and encode the sentences in low-dimensional spaces via neural networks. Although great progress has been made, they seldom consider the information represented by entities, which are of great significance to relation extraction. In this article, we propose several methods based on different tree-based models to learn syntax-aware entity representations for neural relation extraction. First, we encode the context of entities on dependency trees as sentence-level entity embedding based on tree-structured neural network models. Then, we utilize inter-sentence attention mechanism to obtain sentence bag level entity embedding over all sentences containing the specified entity pair. Finally, we combine both sentence embedding and entity embedding for relation classification. Experimental results on a widely used real-world dataset indicate that our system performs better than the state-of-the-art systems of relation extraction.

AAAI Conference 2018 Conference Paper

AdaComp: Adaptive Residual Gradient Compression for Data-Parallel Distributed Training

  • Chia-Yu Chen
  • Jungwook Choi
  • Daniel Brand
  • Ankur Agrawal
  • Wei Zhang
  • Kailash Gopalakrishnan

Highly distributed training of Deep Neural Networks (DNNs) on future compute platforms (offering 100 of TeraOps/s of computational capacity) is expected to be severely communication constrained. To overcome this limitation, new gradient compression techniques are needed that are computationally friendly, applicable to a wide variety of layers seen in Deep Neural Networks and adaptable to variations in network architectures as well as their hyper-parameters. In this paper we introduce a novel technique - the Adaptive Residual Gradient Compression (AdaComp) scheme. AdaComp is based on localized selection of gradient residues and automatically tunes the compression rate depending on local activity. We show excellent results on a wide spectrum of state of the art Deep Learning models in multiple domains (vision, speech, language), datasets (MNIST, CIFAR10, ImageNet, BN50, Shakespeare), optimizers (SGD with momentum, Adam) and network parameters (number of learners, minibatch-size etc.). Exploiting both sparsity and quantization, we demonstrate end-to-end compression rates of ∼200× for fully-connected and recurrent layers, and ∼40× for convolutional layers, without any noticeable degradation in model accuracies.

AAAI Conference 2018 Conference Paper

Adversarial Learning for Chinese NER From Crowd Annotations

  • YaoSheng Yang
  • Meishan Zhang
  • Wenliang Chen
  • Wei Zhang
  • Haofen Wang
  • Min Zhang

To quickly obtain new labeled data, we can choose crowdsourcing as an alternative way at lower cost in a short time. But as an exchange, crowd annotations from non-experts may be of lower quality than those from experts. In this paper, we propose an approach to performing crowd annotation learning for Chinese Named Entity Recognition (NER) to make full use of the noisy sequence labels from multiple annotators. Inspired by adversarial learning, our approach uses a common Bi-LSTM and a private Bi-LSTM for representing annotatorgeneric and -specific information. The annotator-generic information is the common knowledge for entities easily mastered by the crowd. Finally, we build our Chinese NE tagger based on the LSTM-CRF model. In our experiments, we create two data sets for Chinese NER tasks from two domains. The experimental results show that our system achieves better scores than strong baseline systems.

AAAI Conference 2018 Conference Paper

Co-Attending Free-Form Regions and Detections With Multi-Modal Multiplicative Feature Embedding for Visual Question Answering

  • Pan Lu
  • Hongsheng Li
  • Wei Zhang
  • Jianyong Wang
  • Xiaogang Wang

Recently, the Visual Question Answering (VQA) task has gained increasing attention in artificial intelligence. Existing VQA methods mainly adopt the visual attention mechanism to associate the input question with corresponding image regions for effective question answering. The free-form region based and the detection-based visual attention mechanisms are mostly investigated, with the former ones attending free-form image regions and the latter ones attending prespecified detection-box regions. We argue that the two attention mechanisms are able to provide complementary information and should be effectively integrated to better solve the VQA problem. In this paper, we propose a novel deep neural network for VQA that integrates both attention mechanisms. Our proposed framework effectively fuses features from free-form image regions, detection boxes, and question representations via a multi-modal multiplicative feature embedding scheme to jointly attend question-related free-form image regions and detection boxes for more accurate question answering. The proposed method is extensively evaluated on two publicly available datasets, COCO-QA and VQA, and outperforms state-of-the-art approaches. Source code is available at https: //github. com/lupantech/dual-mfa-vqa.

AAAI Conference 2018 Conference Paper

Consistent and Specific Multi-View Subspace Clustering

  • Shirui Luo
  • Changqing Zhang
  • Wei Zhang
  • Xiaochun Cao

Multi-view clustering has attracted intensive attention due to the effectiveness of exploiting multiple views of data. However, most existing multi-view clustering methods only aim to explore the consistency or enhance the diversity of different views. In this paper, we propose a novel multi-view subspace clustering method (CSMSC), where consistency and specificity are jointly exploited for subspace representation learning. We formulate the multi-view self-representation property using a shared consistent representation and a set of specific representations, which better fits the real-world datasets. Specifically, consistency models the common properties among all views, while specificity captures the inherent difference in each view. In addition, to optimize the nonconvex problem, we introduce a convex relaxation and develop an alternating optimization algorithm to recover the corresponding data representations. Experimental evaluations on four benchmark datasets demonstrate that the proposed approach achieves better performance over several state-of-thearts.

IJCAI Conference 2018 Conference Paper

Convolutional Memory Blocks for Depth Data Representation Learning

  • Keze Wang
  • Liang Lin
  • Chuangjie Ren
  • Wei Zhang
  • Wenxiu Sun

Compared to natural RGB images, data captured by 3D / depth sensors (e. g. , Microsoft Kinect) have different properties, e. g. , less discriminable in appearance due to lacking color / texture information. Applying convolutional neural networks (CNNs) on these depth data would lead to unsatisfying learning efficiency, i. e. , requiring large amounts of annotated training data for convergence. To address this issue, this paper proposes a novel memory network module, called Convolutional Memory Block (CMB), which empowers CNNs with the memory mechanism on handling depth data. Different from the existing memory networks that store long / short term dependency from sequential data, our proposed CMB focuses on modeling the representative dependency (correlation) among non-sequential samples. Specifically, our CMB consists of one internal memory (i. e. , a set of feature maps) and three specific controllers, which enable a powerful yet efficient memory manipulation mechanism. In this way, the internal memory, being implicitly aggregated from all previous inputted samples, can learn to store and utilize representative features among the samples. Furthermore, we employ our CMB to develop a concise framework for predicting articulated pose from still depth images. Comprehensive evaluations on three public benchmarks demonstrate significant superiority (about 6%) of our framework over all the compared methods. More importantly, thanks to the enhanced learning efficiency, our framework can still achieve satisfying results using 50% less training data.

NeurIPS Conference 2018 Conference Paper

Evolutionary Stochastic Gradient Descent for Optimization of Deep Neural Networks

  • Xiaodong Cui
  • Wei Zhang
  • Zoltán Tüske
  • Michael Picheny

We propose a population-based Evolutionary Stochastic Gradient Descent (ESGD) framework for optimizing deep neural networks. ESGD combines SGD and gradient-free evolutionary algorithms as complementary algorithms in one framework in which the optimization alternates between the SGD step and evolution step to improve the average fitness of the population. With a back-off strategy in the SGD step and an elitist strategy in the evolution step, it guarantees that the best fitness in the population will never degrade. In addition, individuals in the population optimized with various SGD-based optimizers using distinct hyper-parameters in the SGD step are considered as competing species in a coevolution setting such that the complementarity of the optimizers is also taken into account. The effectiveness of ESGD is demonstrated across multiple applications including speech recognition, image recognition and language modeling, using networks with a variety of deep architectures.

IJCAI Conference 2018 Conference Paper

Image-level to Pixel-wise Labeling: From Theory to Practice

  • Tiezhu Sun
  • Wei Zhang
  • Zhijie Wang
  • Lin Ma
  • Zequn Jie

Conventional convolutional neural networks (CNNs) have achieved great success in image semantic segmentation. Existing methods mainly focus on learning pixel-wise labels from an image directly. In this paper, we advocate tackling the pixel-wise segmentation problem by considering the image-level classification labels. Theoretically, we analyze and discuss the effects of image-level labels on pixel-wise segmentation from the perspective of information theory. In practice, an end-to-end segmentation model is built by fusing the image-level and pixel-wise labeling networks. A generative network is included to reconstruct the input image and further boost the segmentation model training with an auxiliary loss. Extensive experimental results on benchmark dataset demonstrate the effectiveness of the proposed method, where good image-level labels can significantly improve the pixel-wise segmentation accuracy.

IJCAI Conference 2018 Conference Paper

Improving Entity Recommendation with Search Log and Multi-Task Learning

  • Jizhou Huang
  • Wei Zhang
  • Yaming Sun
  • Haifeng Wang
  • Ting Liu

Entity recommendation, providing search users with an improved experience by assisting them in finding related entities for a given query, has become an indispensable feature of today's Web search engine. Existing studies typically only consider the query issued at the current time step while ignoring the in-session preceding queries. Thus, they typically fail to handle the ambiguous queries such as "apple" because the model could not understand which apple (company or fruit) is talked about. In this work, we believe that the in-session contexts convey valuable evidences that could facilitate the semantic modeling of queries, and take that into consideration for entity recommendation. Furthermore, in order to better model the semantics of queries, we learn the model in a multi-task learning setting where the query representation is shared across entity recommendation and context-aware ranking. We evaluate our approach using large-scale, real-world search logs of a widely used commercial Web search engine. The experimental results show that incorporating context information significantly improves entity recommendation, and learning the model in a multi-task learning setting could bring further improvements.

IJCAI Conference 2018 Conference Paper

Learning Sequential Correlation for User Generated Textual Content Popularity Prediction

  • Wen Wang
  • Wei Zhang
  • Jun Wang
  • Junchi Yan
  • Hongyuan Zha

Popularity prediction of user generated textual content is critical for prioritizing information in the web, which alleviates heavy information overload for ordinary readers. Most previous studies model each content instance separately for prediction and thus overlook the sequential correlations between instances of a specific user. In this paper, we go deeper into this problem based on the two observations for each user, i. e. , sequential content correlation and sequential popularity correlation. We propose a novel deep sequential model called User Memory-augmented recurrent Attention Network (UMAN). This model encodes the two correlations by updating external user memories which is further leveraged for target text representation learning and popularity prediction. The experimental results on several real-world datasets validate the benefits of considering these correlations and demonstrate UMAN achieves best performance among several strong competitors.

IJCAI Conference 2018 Conference Paper

Master-Slave Curriculum Design for Reinforcement Learning

  • Yuechen Wu
  • Wei Zhang
  • Ke Song

Curriculum learning is often introduced as a leverage to improve the agent training for complex tasks, where the goal is to generate a sequence of easier subasks for an agent to train on, such that final performance or learning speed is improved. However, conventional curriculum is mainly designed for one agent with fixed action space and sequential simple-to-hard training manner. Instead, we present a novel curriculum learning strategy by introducing the concept of master-slave agents and enabling flexible action setting for agent training. Multiple agents, referred as master agent for the target task and slave agents for the subtasks, are trained concurrently within different action spaces by sharing a perception network with an asynchronous strategy. Extensive evaluation on the VizDoom platform demonstrates the joint learning of master agent and slave agents mutually benefit each other. Significant improvement is obtained over A3C in terms of learning speed and performance.

AAAI Conference 2018 Conference Paper

R 3: Reinforced Ranker-Reader for Open-Domain Question Answering

  • Shuohang Wang
  • Mo Yu
  • Xiaoxiao Guo
  • Zhiguo Wang
  • Tim Klinger
  • Wei Zhang
  • Shiyu Chang
  • Gerry Tesauro

In recent years researchers have achieved considerable success applying neural network methods to question answering (QA). These approaches have achieved state of the art results in simplified closed-domain settings1 such as the SQuAD (Rajpurkar et al. 2016) dataset, which provides a preselected passage, from which the answer to a given question may be extracted. More recently, researchers have begun to tackle open-domain QA, in which the model is given a question and access to a large corpus (e. g. , wikipedia) instead of a pre-selected passage (Chen et al. 2017a). This setting is more complex as it requires large-scale search for relevant passages by an information retrieval component, combined with a reading comprehension model that “reads” the passages to generate an answer to the question. Performance in this setting lags well behind closed-domain performance. In this paper, we present a novel open-domain QA system called Reinforced Ranker-Reader (R3 ), based on two algorithmic innovations. First, we propose a new pipeline for open-domain QA with a Ranker component, which learns to rank retrieved passages in terms of likelihood of extracting the ground-truth answer to a given question. Second, we propose a novel method that jointly trains the Ranker along with an answer-extraction Reader model, based on reinforcement learning. We report extensive experimental results showing that our method significantly improves on the state of the art for multiple open-domain QA datasets. 2

AAAI Conference 2018 Conference Paper

SEE: Syntax-Aware Entity Embedding for Neural Relation Extraction

  • Zhengqiu He
  • Wenliang Chen
  • Zhenghua Li
  • Meishan Zhang
  • Wei Zhang
  • Min Zhang

Distant supervised relation extraction is an efficient approach to scale relation extraction to very large corpora, and has been widely used to find novel relational facts from plain text. Recent studies on neural relation extraction have shown great progress on this task via modeling the sentences in lowdimensional spaces, but seldom considered syntax information to model the entities. In this paper, we propose to learn syntax-aware entity embedding for neural relation extraction. First, we encode the context of entities on a dependency tree as sentence-level entity embedding based on tree-GRU. Then, we utilize both intra-sentence and inter-sentence attentions to obtain sentence set-level entity embedding over all sentences containing the focus entity pair. Finally, we combine both sentence embedding and entity embedding for relation classi- fication. We conduct experiments on a widely used real-world dataset and the experimental results show that our model can make full use of all informative instances and achieve stateof-the-art performance of relation extraction.

YNIMG Journal 2018 Journal Article

Spatio-temporal modeling of connectome-scale brain network interactions via time-evolving graphs

  • Jing Yuan
  • Xiang Li
  • Jinhe Zhang
  • Liao Luo
  • Qinglin Dong
  • Jinglei Lv
  • Yu Zhao
  • Xi Jiang

Many recent literature studies have revealed interesting dynamics patterns of functional brain networks derived from fMRI data. However, it has been rarely explored how functional networks spatially overlap (or interact) and how such connectome-scale network interactions temporally evolve. To explore these unanswered questions, this paper presents a novel framework for spatio-temporal modeling of connectome-scale functional brain network interactions via two main effective computational methodologies. First, to integrate, pool and compare brain networks across individuals and their cognitive states under task performances, we designed a novel group-wise dictionary learning scheme to derive connectome-scale consistent brain network templates that can be used to define the common reference space of brain network interactions. Second, the temporal dynamics of spatial network interactions is modeled by a weighted time-evolving graph, and then a data-driven unsupervised learning algorithm based on the dynamic behavioral mixed-membership model (DBMM) is adopted to identify behavioral patterns of brain networks during the temporal evolution process of spatial overlaps/interactions. Experimental results on the Human Connectome Project (HCP) task fMRI data showed that our methods can reveal meaningful, diverse behavior patterns of connectome-scale network interactions. In particular, those networks’ behavior patterns are distinct across HCP tasks such as motor, working memory, language and social tasks, and their dynamics well correspond to the temporal changes of specific task designs. In general, our framework offers a new approach to characterizing human brain function by quantitative description for the temporal evolution of spatial overlaps/interactions of connectome-scale brain networks in a standard reference space.

NeurIPS Conference 2017 Conference Paper

Can Decentralized Algorithms Outperform Centralized Algorithms? A Case Study for Decentralized Parallel Stochastic Gradient Descent

  • Xiangru Lian
  • Ce Zhang
  • Huan ZHang
  • Cho-Jui Hsieh
  • Wei Zhang
  • Ji Liu

Most distributed machine learning systems nowadays, including TensorFlow and CNTK, are built in a centralized fashion. One bottleneck of centralized algorithms lies on high communication cost on the central node. Motivated by this, we ask, can decentralized algorithms be faster than its centralized counterpart? Although decentralized PSGD (D-PSGD) algorithms have been studied by the control community, existing analysis and theory do not show any advantage over centralized PSGD (C-PSGD) algorithms, simply assuming the application scenario where only the decentralized network is available. In this paper, we study a D-PSGD algorithm and provide the first theoretical analysis that indicates a regime in which decentralized algorithms might outperform centralized algorithms for distributed stochastic gradient descent. This is because D-PSGD has comparable total computational complexities to C-PSGD but requires much less communication cost on the busiest node. We further conduct an empirical study to validate our theoretical analysis across multiple frameworks (CNTK and Torch), different network configurations, and computation platforms up to 112 GPUs. On network configurations with low bandwidth or high latency, D-PSGD can be up to one order of magnitude faster than its well-optimized centralized counterparts.

JBHI Journal 2017 Journal Article

Chair Rise Peak Power in Daily Life Measured With a Pendant Sensor Associates With Mobility, Limitation in Activities, and Frailty in Old People

  • Wei Zhang
  • G. Ruben H. Regterschot
  • Hilde Geraedts
  • Heribert Baldus
  • Wiebren Zijlstra

The aim of this study was to analyze the clinical relevance of sensor-based daily life chair rise performance measured in old people. A pendant-sensor was worn during standardized tests and in daily life to detect chair rise transfers and analyze transfer peak power. Linear correlations between mean, median, 25th, and 75th percentile transfer peak powers in daily life and mean peak power in standardized tests were evaluated with Pearson correlation ( r). Associations between transfer peak powers in different experiments and outcomes of a clinical mobility test [timed-up-and-go (TUG)], a test of limitation in activities [Groningen activity restriction scale (GARS)], and a frailty test [Groningen frailty indicator (GFI)] were evaluated with Spearman correlation (ρ). Twenty-five old people (70-85 years) participated in the study. The results showed that chair rise peak powers assessed based upon one-week of daily life activities significantly correlated with peak power measured in standardized tests (r: [0. 66, 0. 74], p <; 0. 01). Chair rise peak power in daily life significantly associated with TUG scores (ρ: [-0. 71, -0. 58], p <; 0. 01), GARS (ρ: [-0. 62, -0. 48], p <; 0. 05), and GFI (ρ: [-0. 52, -0. 43], p <; 0. 05). Chair rise peak powers in daily life had stronger associations with clinical measurements than standardized tests. In addition, chair rise peak powers measured in old people using assistive devices was significantly lower compared to those not using assistive devices. These results indicate usefulness of the pendant-sensor-based chair rise performance analysis in continuous monitoring and assessment of mobility, limitations in activities and frailty associated variables in old people's daily life.

TIST Journal 2017 Journal Article

Energy-Efficient Mobile Video Streaming

  • Wei Zhang
  • Rui Fan
  • Yonggang Wen
  • Fang Liu

Video streaming is one of the most widely used mobile applications today, and it also accounts for a large fraction of mobile battery usage. Much of the energy consumption is for wireless data transmission and is highly correlated to network bandwidth conditions. In periods of poor connectivity, up to 90% of mobile energy can be used for wireless data transfer. In this article, we study the problem of energy-efficient mobile video streaming. We make use of the observed correlation between bandwidth and user location, and also observe that a user’s location is predictable in many situations, such as when commuting to a known destination. Based on the user’s predicted locations and bandwidth conditions, we optimize wireless transmission times to achieve high quality video playback while minimizing energy use. We propose an optimal offline algorithm for this problem, which runs in O ( Tk ) time, where T is the duration of the video and k is the size of the video buffer. We also propose LAWS, a Location AWare Streaming algorithm. LAWS learns from historical location-aware bandwidth conditions and predicts future bandwidths along a planned route to make online wireless download decisions. We evaluate LAWS using real bandwidth traces, and show that LAWS closely approximates the performance of the optimal offline algorithm, achieving 90.6% of the optimal performance on average, and 97% in certain cases. LAWS also outperforms three popular strategies used in practice by, on average, 69%, 63%, and 38%, respectively. Lastly, we show that LAWS is able to deal with noisy data and can attain the stated performance after sampling bandwidth conditions only five times.

IJCAI Conference 2017 Conference Paper

Learning to Explain Entity Relationships by Pairwise Ranking with Convolutional Neural Networks

  • Jizhou Huang
  • Wei Zhang
  • Shiqi Zhao
  • Shiqiang Ding
  • Haifeng Wang

Providing a plausible explanation for the relationship between two related entities is an important task in some applications of knowledge graphs, such as in search engines. However, most existing methods require a large number of manually labeled training data, which cannot be applied in large-scale knowledge graphs due to the expensive data annotation. In addition, these methods typically rely on costly handcrafted features. In this paper, we propose an effective pairwise ranking model by leveraging clickthrough data of a Web search engine to address these two problems. We first construct large-scale training data by leveraging the query-title pairs derived from clickthrough data of a Web search engine. Then, we build a pairwise ranking model which employs a convolutional neural network to automatically learn relevant features. The proposed model can be easily trained with backpropagation to perform the ranking task. The experiments show that our method significantly outperforms several strong baselines.

YNICL Journal 2017 Journal Article

Metrics of brain network architecture capture the impact of disease in children with epilepsy

  • Michael J. Paldino
  • Wei Zhang
  • Zili D. Chu
  • Farahnaz Golriz

BACKGROUND AND OBJECTIVE: Epilepsy is associated with alterations in the structural framework of the cerebral network. The aim of this study was to measure the potential of global metrics of network architecture derived from resting state functional MRI to capture the impact of epilepsy on the developing brain. METHODS: Pediatric patients were retrospectively identified with: 1. Focal epilepsy; 2. Brain MRI at 3 Tesla, including resting state functional MRI; 3. Full scale IQ measured by a pediatric neuropsychologist. The cerebral cortex was parcellated into approximately 700 gray matter network nodes. The strength of a connection between two nodes was defined as the correlation between their resting BOLD signal time series. The following global network metrics were then calculated: clustering coefficient, transitivity, modularity, path length, and global efficiency. Epilepsy duration was used as an index for the cumulative impact of epilepsy on the brain. RESULTS: : 0.0001). Specifically, modularity and to a lesser extent path length and global efficiency were independently associated with epilepsy duration. CONCLUSIONS: We observed that a machine learning algorithm accurately predicted epilepsy duration based on global metrics of network architecture derived from resting state fMRI. These findings suggest that network metrics have the potential to form the basis for statistical models that translate quantitative imaging data into patient-level markers of cognitive deterioration.

IJCAI Conference 2017 Conference Paper

Model Accuracy and Runtime Tradeoff in Distributed Deep Learning: A Systematic Study

  • Suyog Gupta
  • Wei Zhang
  • Fei Wang

Deep learning with a large number of parame-ters requires distributed training, where model accuracy and runtime are two important factors to be considered. However, there has been no systematic study of the tradeoff between these two factors during the model training process. This paper presents Rudra, a parameter server based distributed computing framework tuned for training large-scale deep neural networks. Using variants of the asynchronous stochastic gradient descent algorithm we study the impact of synchronization protocol, stale gradient updates, minibatch size, learning rates, and number of learners on runtime performance and model accuracy. We introduce a new learningrate modulation strategy to counter the effect of stale gradients and propose a new synchronization protocol that can effectively bound the staleness in gradients, improve runtime performance and achieve good model accuracy. Our empirical investigation reveals a principled approach for distributed training of neural networks: the mini-batch size per learner should be reduced as more learners are added to the system to preserve the model accuracy. We validate this approach using commonly-used image classification benchmarks: CIFAR10 and ImageNet.

EAAI Journal 2017 Journal Article

The effect of genetic algorithm learning with a classifier system in limit order markets

  • Lijian Wei
  • Xiong Xiong
  • Wei Zhang
  • Xue-Zhong He
  • Yongjie Zhang

By introducing a genetic algorithm with a classifier system as a learning mechanism for uninformed traders into a dynamic limit order market with asymmetric information, this paper examines the effect of the learning on traders’ trading behavior, market liquidity and efficiency. We show that the learning is effective and valuable with respect to information acquisition, forecasting, buy–sell order choice accuracies, and profit opportunity for uninformed traders. It improves information dissemination efficiency and reduces the information advantage of informed traders and hence the value of the private information. In particular, the learning and information become more valuable with higher volatility, less informed traders, and longer information lag. Furthermore, the learning makes not only uninformed but also informed traders submit more limit orders and hence increases market liquidity supply.

IJCAI Conference 2016 Conference Paper

Collaborative Multi-Level Embedding Learning from Reviews for Rating Prediction

  • Wei Zhang
  • Quan Yuan
  • Jiawei Han
  • Jianyong Wang

We investigate the problem of personalized review-based rating prediction which aims at predicting users' ratings for items that they have not evaluated by using their historical reviews and ratings. Most of existing methods solve this problem by integrating topic model and latent factor model to learn interpretable user and items factors. However, these methods cannot utilize word local context information of reviews. Moreover, it simply restricts user and item representations equivalent to their review representations, which may bring some irrelevant information in review text and harm the accuracy of rating prediction. In this paper, we propose a novel Collaborative Multi-Level Embedding (CMLE) model to address these limitations. The main technical contribution of CMLE is to integrate word embedding model with standard matrix factorization model through a projection level. This allows CMLE to inherit the ability of capturing word local context information from word embedding model and relax the strict equivalence requirement by projecting review embedding to user and item embeddings. A joint optimization problem is formulated and solved through an efficient stochastic gradient ascent algorithm. Empirical evaluations on real datasets show CMLE outperforms several competitive methods and can solve the two limitations well.

IJCAI Conference 2016 Conference Paper

Model-Based Deep Hand Pose Estimation

  • Xingyi Zhou
  • Qingfu Wan
  • Wei Zhang
  • Xiangyang Xue
  • Yichen Wei

Previous learning based hand pose estimation methods does not fully exploit the prior information in hand model geometry. Instead, they usually rely a separate model fitting step to generate valid hand poses. Such a post processing is inconvenient and sub-optimal. In this work, we propose a model based deep learning approach that adopts a forward kinematics based layer to ensure the geometric validity of estimated poses. For the first time, we show that embedding such a non-linear generative process in deep learning is feasible for hand pose estimation. Our approach is verified on challenging public datasets and achieves state-of-the-art performance.

IJCAI Conference 2016 Conference Paper

Staleness-Aware Async-SGD for Distributed Deep Learning

  • Wei Zhang
  • Suyog Gupta
  • Xiangru Lian
  • Ji Liu

Deep neural networks have been shown to achieve state-of-the-art performance in several machine learning tasks. Stochastic Gradient Descent (SGD) is the preferred optimization algorithm for training these networks and asynchronous SGD (ASGD) has been widely adopted for accelerating the training of large-scale deep networks in a distributed computing environment. However, in practice it is quite challenging to tune the training hyperparameters (such as learning rate) when using ASGD so as achieve convergence and linear speedup, since the stability of the optimization algorithm is strongly influenced by the asynchronous nature of parameter updates. In this paper, we propose a variant of the ASGD algorithm in which the learning rate is modulated according to the gradient staleness and provide theoretical guarantees for convergence of this algorithm. Experimental verification is performed on commonly-used image classification benchmarks: CIFAR10 and Imagenet to demonstrate the superior effectiveness of the proposed approach, compared to SSGD (Synchronous SGD) and the conventional ASGD algorithm.

IJCAI Conference 2015 Conference Paper

Prior-Based Dual Additive Latent Dirichlet Allocation for User-Item Connected Documents

  • Wei Zhang
  • Jianyong Wang

User-item connected documents, such as customer reviews for specific items in online shopping website and user tips in location-based social networks, have become more and more prevalent recently. Inferring the topic distributions of user-item connected documents is beneficial for many applications, including document classification and summarization of users and items. While many different topic models have been proposed for modeling multiple text, most of them cannot account for the dual role of user-item connected documents (each document is related to one user and one item simultaneously) in topic distribution generation process. In this paper, we propose a novel probabilistic topic model called Prior-based Dual Additive Latent Dirichlet Allocation (PDA-LDA). It addresses the dual role of each document by associating its Dirichlet prior for topic distribution with user and item topic factors, which leads to a document-level asymmetric Dirichlet prior. In the experiments, we evaluate PDA-LDA on several real datasets and the results demonstrate that our model is effective in comparison to several other models, including held-out perplexity on modeling text and document classification application.

JBHI Journal 2014 Journal Article

Computer-Aided Bleeding Detection in WCE Video

  • Yanan Fu
  • Wei Zhang
  • Mrinal Mandal
  • Max Q.-H. Meng

Wireless capsule endoscopy (WCE) can directly take digital images in the gastrointestinal tract of a patient. It has opened a new chapter in small intestine examination. However, a major problem associated with this technology is that too many images need to be manually examined by clinicians. Currently, there is no standard for capsule endoscopy image interpretation and classification. Most state-of-the-art CAD methods often suffer from poor performance, high computational cost, or multiple empirical thresholds. In this paper, a new method for rapid bleeding detection in the WCE video is proposed. We group pixels through superpixel segmentation to reduce the computational complexity while maintaining high diagnostic accuracy. Feature of each superpixel is extracted using the red ratio in RGB space and fed into support vector machine for classification. Also, the influence of edge pixels has been removed in this paper. Comparative experiments show that our algorithm is superior to the existing methods in terms of sensitivity, specificity, and accuracy.

YNIMG Journal 2014 Journal Article

Concurrent TMS to the primary motor cortex augments slow motor learning

  • Shalini Narayana
  • Wei Zhang
  • William Rogers
  • Casey Strickland
  • Crystal Franklin
  • Jack L. Lancaster
  • Peter T. Fox

Transcranial magnetic stimulation (TMS) has shown promise as a treatment tool, with one FDA approved use. While TMS alone is able to up- (or down-) regulate a targeted neural system, we argue that TMS applied as an adjuvant is more effective for repetitive physical, behavioral and cognitive therapies, that is, therapies which are designed to alter the network properties of neural systems through Hebbian learning. We tested this hypothesis in the context of a slow motor learning paradigm. Healthy right-handed individuals were assigned to receive 5Hz TMS (TMS group) or sham TMS (sham group) to the right primary motor cortex (M1) as they performed daily motor practice of a digit sequence task with their non-dominant hand for 4weeks. Resting cerebral blood flow (CBF) was measured by H2 15O PET at baseline and after 4weeks of practice. Sequence performance was measured daily as the number of correct sequences performed, and modeled using a hyperbolic function. Sequence performance increased significantly at 4weeks relative to baseline in both groups. The TMS group had a significant additional improvement in performance, specifically, in the rate of skill acquisition. In both groups, an improvement in sequence timing and transfer of skills to non-trained motor domains was also found. Compared to the sham group, the TMS group demonstrated increases in resting CBF specifically in regions known to mediate skill learning namely, the M1, cingulate cortex, putamen, hippocampus, and cerebellum. These results indicate that TMS applied concomitantly augments behavioral effects of motor practice, with corresponding neural plasticity in motor sequence learning network. These findings are the first demonstration of the behavioral and neural enhancing effects of TMS on slow motor practice and have direct application in neurorehabilitation where TMS could be applied in conjunction with physical therapy.

JAAMAS Journal 2014 Journal Article

Flocking of partially-informed multi-agent systems avoiding obstacles with arbitrary shape

  • Jiaojie Li
  • Wei Zhang
  • Yupu Yang

Abstract In this paper, we study the flocking problem of multi-agent systems with obstacle avoidance, in the situation when only a fraction of the agents have information on the obstacles. Obstacles of arbitrary shape are allowed, no matter if their boundary is smooth or non-smooth, and no matter it they are convex or non-convex. A novel geometry representation rule is proposed to transfer obstacles to a dense obstacle-agents lattice structure. Non-convex regions of the obstacles are detected and supplemented using a geometric rule. The uninformed agents can detect a section of the obstacles boundary using only a range position sensor. We prove that with the proposed protocol, uninformed agents which maintain a joint path with any informed agent can avoid obstacles that move uniformly and assemble around a point along with the informed agents. Eventually all the assembled agents reach consensus on their velocity. In the entire flocking process, no distinct pair of agents collide with each other, nor collide with obstacles. The assembled agents are guaranteed not to be lost in any non-convex region of the obstacles within a distance constraint. Numerical simulations demonstrate the flocking algorithm with obstacle avoidance both in 2D and 3D space. The situation when every agent is informed is considered as a special case.

YNICL Journal 2014 Journal Article

Independent contribution of individual white matter pathways to language function in pediatric epilepsy patients

  • Michael J. Paldino
  • Kara Hedges
  • Wei Zhang

BACKGROUND AND PURPOSE: Patients with epilepsy and malformations of cortical development (MCDs) are at high risk for language and other cognitive impairment. Specific impairments, however, are not well correlated with the extent and locale of dysplastic cortex; such findings highlight the relevance of aberrant cortico-cortical interactions, or connectivity, to the clinical phenotype. The goal of this study was to determine the independent contribution of well-described white matter pathways to language function in a cohort of pediatric patients with epilepsy. MATERIALS AND METHODS: Patients were retrospectively identified from an existing database of pediatric epilepsy patients with the following inclusion criteria: 1. diagnosis of MCDs, 2. DTI performed at 3 T, and 3. language characterized by a pediatric neurologist. Diffusion Toolkit and Trackvis (http://www.trackvis.org) were used for segmentation and analysis of the following tracts: corpus callosum, corticospinal tracts, inferior longitudinal fasciculi (ILFs), inferior fronto-occipital fasciculi (IFOFs), uncinate fasciculi (UFs), and arcuate fasciculi (AFs). Mean diffusivity (MD) and fractional anisotropy (FA) were calculated for each tract. Wilcoxon rank sum test (corrected for multiple comparisons) was used to assess potential differences in tract parameters between language-impaired and language-intact patients. In a separate analysis, a machine learning algorithm (random forest approach) was applied to measure the independent contribution of the measured diffusion parameters for each tract to the clinical phenotype (language impairment). In other words, the importance of each tract parameter was measured after adjusting for the contribution of all other tracts. RESULTS: Thirty-three MCD patients were included (age range: 3-18 years). Twenty-one patients had intact language, twelve had language impairment. All tracts were identified bilaterally in all patients except for the AF, which was not identified on the right in 10 subjects and not identified on the left in 11 subjects. MD and/or FA within the left AF, UF, ILF, and IFOF differed between language-intact and language-impaired groups. However, only parameters related to the left uncinate, inferior fronto-occipital, and arcuate fasciculi were independently associated with the clinical phenotype. CONCLUSIONS: Scalar metrics derived from the left uncinate, inferior fronto-occipital, and arcuate fasciculi were independently associated with language function. These results support the importance of these pathways in human language function in patients with MCDs.

AAAI Conference 2014 Conference Paper

Semantic Segmentation Using Multiple Graphs with Block-Diagonal Constraints

  • Ke Zhang
  • Wei Zhang
  • Sheng Zeng
  • Xiangyang Xue

In this paper we propose a novel method for image semantic segmentation using multiple graphs. The multiview affinity graph is constructed by leveraging the consistency between semantic space and multiple visual spaces. With block-diagonal constraints, we enforce the affinity matrix to be sparse such that the pairwise potential for dissimilar superpixels is close to zero. By a divide-and-conquer strategy, the optimization for learning affinity matrix is decomposed into several subproblems that can be solved in parallel. Using the neighborhood relationship between superpixels and the consistency between affinity matrix and labelconfidence matrix, we infer the semantic label for each superpixel of unlabeled images by minimizing an objective whose closed form solution can be easily obtained. Experimental results on two real-world image datasets demonstrate the effectiveness of our method.

IJCAI Conference 2013 Conference Paper

Dimensionality Reduction with Generalized Linear Models

  • Mo Chen
  • Wei Li
  • Wei Zhang
  • Xiaogang Wang

In this paper, we propose a general dimensionality reduction method for data generated from a very broad family of distributions and nonlinear functions based on the generalized linear model, called Generalized Linear Principal Component Analysis (GLPCA). Data of different domains often have very different structures. These data can be modeled by different distributions and reconstruction functions. For example, real valued data can be modeled by the Gaussian distribution with a linear reconstruction function, whereas binary valued data may be more appropriately modeled by the Bernoulli distribution with a logit or probit function. Based on general linear models, we propose a unified framework for extracting features from data of different domains. A general optimization algorithm based on natural gradient ascent on distribution manifold is proposed for obtaining the maximum likelihood solutions. We also present some specific algorithms derived from this framework to deal with specific data modeling problems such as document modeling. Experimental results of these algorithms on several data sets are shown for the validation of GLPCA.

IJCAI Conference 2013 Conference Paper

Integrating Semantic Relatedness and Words' Intrinsic Features for Keyword Extraction

  • Wei Zhang
  • Wei Feng
  • Jianyong Wang

Keyword extraction attracts much attention for its significant role in various natural language processing tasks. While some existing methods for keyword extraction have considered using single type of semantic relatedness between words or inherent attributes of words, almost all of them ignore two important issues: 1) how to fuse multiple types of semantic relations between words into a uniform semantic measurement and automatically learn the weights of the edges between the words in the word graph of each document, and 2) how to integrate the relations between words and words’ intrinsic features into a unified model. In this work, we tackle the two issues based on the supervised random walk model. We propose a supervised ranking based method for keyword extraction, which is called SEAFARER1. It can not only automatically learn the weights of the edges in the unified graph of each document which includes multiple semantic relations but also combine the merits of semantic relations of edges and intrinsic attributes of nodes together. We conducted extensive experimental study on an established benchmark and the experimental results demonstrate that SEAFARER outperforms the state-of-the-art supervised and unsupervised methods.

IJCAI Conference 2013 Conference Paper

Multi-View Embedding Learning for Incompletely Labeled Data

  • Wei Zhang
  • Ke Zhang
  • Pan Gu
  • Xiangyang Xue

In many applications, the data may be high dimensional, represented by multiple features, and associated with more than one labels. Embedding learning is an effective strategy for dimensionality reduction and for nearest neighbor search in massive datasets. We propose a novel method to seek compact embedding that allows efficient retrieval with incompletely-labeled multi-view data. Based on multi-graph Laplacian, we achieve the optimal combination of heterogeneous features to effectively describe data, which exploits the feature correlations between different views. We learn the embedding that preserves the neighborhood context in the original spaces, and obtain the complete labels simultaneously. Inter-label correlations are sufficiently leveraged in the proposed framework. Our goal is to find the maps from multiple input spaces to the compact embedding space and to the semantic concept space at the same time. There is semantic gap between the input multi-view feature spaces and the semantic concept space; and the compact embedding space can be looked on as the bridge between the above spaces. Experimental evaluation on three real-world datasets demonstrates the effectiveness of the proposed method.

IJCAI Conference 2013 Conference Paper

Sparse Reconstruction for Weakly Supervised Semantic Segmentation

  • Ke Zhang
  • Wei Zhang
  • Yingbin Zheng
  • Xiangyang Xue

We propose a novel approach to semantic segmentation using weakly supervised labels. In traditional fully supervised methods, superpixel labels are available for training; however, it is not easy to obtain enough labeled superpixels to learn a satisfying model for semantic segmentation. By contrast, only image-level labels are necessary in weakly supervised methods, which makes them more practical in real applications. In this paper we develop a new way of evaluating classification models for semantic segmentation given weekly supervised labels. For a certain category, provided the classi- fication model parameter, we firstly learn the basis superpixels by sparse reconstruction, and then evaluate the parameters by measuring the reconstruction errors among negative and positive superpixels. Based on Gaussian Mixture Models, we use Iterative Merging Update (IMU) algorithm to obtain the best parameters for the classification models. Experimental results on two real-world datasets show that the proposed approach outperforms the existing weakly supervised methods, and it also competes with state-of-the-art fully supervised methods.

IJCAI Conference 2011 Conference Paper

Entity Linking with Effective Acronym Expansion, Instance Selection, and Topic Modeling

  • Wei Zhang
  • Yan Chuan Sim
  • Jian Su
  • Chew Lim Tan

Entity linking maps name mentions in the documents to entries in a knowledge base through resolving the name variations and ambiguities. In this paper, we propose three advancements for entity linking. Firstly, expanding acronyms can effectively reduce the ambiguity of the acronym mentions. However, only rule-based approaches relying heavily on the presence of text markers have been used for entity linking. In this paper, we propose a supervised learning algorithm to expand more complicated acronyms encountered, which leads to 15. 1% accuracy improvement over state-of-the-art acronym expansion methods. Secondly, as entity linking annotation is expensive and labor intensive, to automate the annotation process without compromise of accuracy, we propose an instance selection strategy to effectively utilize the automatically generated annotation. In our selection strategy, an informative and diverse set of instances are selected for effective disambiguation. Lastly, topic modeling is used to model the semantic topics of the articles. These advancements give statistical significant improvement to entity linking individually. Collectively they lead the highest performance on KBP-2010 task.

YNIMG Journal 2011 Journal Article

Functional neuroimaging of the baboon during concurrent image-guided transcranial magnetic stimulation

  • Felipe S. Salinas
  • C. Ákos Szabó
  • Wei Zhang
  • Lisa Jones
  • M. Michelle Leland
  • Hsiao-Ying Wey
  • Timothy Q. Duong
  • Peter T. Fox

Transcranial magnetic stimulation (TMS) has well-established applications in basic neuroscience and promising applications in neurological and psychiatric disorders. However the underlying mechanisms of TMS-induced alterations in brain function are not well understood. As a result, treatment design parameters are determined ad hoc and not informed by any coherent theory or model. Once the mechanisms underlying TMS's modulatory effects on brain systems are better understood and modeled, TMS's potential as a therapeutic and/or investigative tool will be more readily explored and exploited. An animal model is better suited to study different TMS variables, therefore we developed a baboon model to facilitate testing of some of the current theoretical models of TMS interactions with brain regions. We have demonstrated the feasibility of this approach by successfully imaging cerebral blood flow (CBF) changes with H2 15O positron emission tomography imaging during high-frequency, suprathreshold repetitive TMS in the primary motor cortex of five healthy, adult baboons.

IJCAI Conference 2011 Conference Paper

Multi-Kernel Multi-Label Learning with Max-Margin Concept Network

  • Wei Zhang
  • Xiangyang Xue
  • Jianping Fan
  • Xiaojing Huang
  • Bin Wu
  • Mingjie Liu

In this paper, a novel method is developed for enabling Multi-Kernel Multi-Label Learning. Inter-label dependency and similarity diversity are simultaneously leveraged in the proposed method. A concept network is constructed to capture the inter-label correlations for classifier training. Maximal margin approach is used to effectively formulate the feature-label associations and the label-label correlations. Specific kernels are learned not only for each label but also for each pair of the inter-related labels. By learning the eigenfunctions of the kernels, the similarity between a new data point and the training samples can be computed in the online mode. Our experimental results on real datasets (web pages, images, music, and bioinformatics) have demonstrated the effectiveness of our method.

YNIMG Journal 2010 Journal Article

Selective aberrant functional connectivity of resting state networks in social anxiety disorder

  • Wei Liao
  • Huafu Chen
  • Yuan Feng
  • Dante Mantini
  • Claudio Gentili
  • Zhengyong Pan
  • Jurong Ding
  • Xujun Duan

Several functional MRI (fMRI) activation studies have highlighted specific differences in brain response in social anxiety disorder (SAD) patients. Little is known, so far, about the changes in the functional architecture of resting state networks (RSNs) in SAD during resting state. We investigated statistical differences in RSNs on 20 SAD and 20 controls using independent component analysis. A diffuse impact on widely distributed RSNs and selective changes of RSN intrinsic functional connectivity were observed in SAD. Functional connectivity was decreased in the somato-motor (primary and motor cortices) and visual (primary visual cortex) networks, increased in a network including medial prefrontal cortex which is thought to be involved in self-referential processes, and increased or decreased in the default mode network (posterior cingulate cortex/precuneus, bilateral inferior parietal gyrus, angular gyrus, middle temporal gyrus, and superior and medial frontal gyrus) which has been suggested to be involved in episodic memory, and self-projection, the dorsal attention network (middle and superior occipital gyrus, inferior and superior parietal gyrus, and middle and superior frontal gyrus) which is thought to mediate goal-directed top-down processing, the core network (insula-cingulate cortices) which is associated with task control function, and the central-executive network (fronto-parietal cortices). A relationship between functional connectivity and disease severity was found in specific regions of RSNs, including medial and lateral prefrontal cortex, as well as parietal and occipital regions. Our results might supply a novel way to look into neuro-pathophysiological mechanisms in SAD patients.

IS Journal 2007 Journal Article

Can Irrational Investors Survive? A Social-Computing Perspective

  • Yongjie Zhang
  • Wei Zhang

Standard financial theory includes a rigorous theoretical system for equilibrium asset pricing. Among the assumptions this system is founded on is the necessity of investors' homogeneity and rationality. The proposed agent-based model accounts for interactions between irrational and rational investors, expanding on existing work and offering hope for irrational-investor survival in artificial stock markets.

IROS Conference 2006 Conference Paper

A Tripodic Biomimetic Underwater Microrobots Utilizing ICPF Actuators

  • Wei Zhang
  • Shuxiang Guo
  • Kinji Asaka

In this paper, to deal with the locomotion in underwater environment, we designed a novel type of leg, inspired by stick insect legs. The biomimetic leg is composed of two segments of ICPF (ionic conducting polymer film) actuators with a 2-DOF motion. It has simple structure, simple control method, and good localization characteristic. We designed and developed a prototype microrobot with 3 units of the legs. The dimension of the microrobot is 48*20*17 mm 3 and its dried weight is 0. 65 g. We have measured the speeds with different control signals. The experimental results show that it can arrive at a speed of 4. 7 mm/s with a control signal of 5 Hz and 10 V. And results indicate that the novel type leg has a good speed performance and localization in aquatic environment. The microrobot using ICPF actuators has powerful applications in the biomedical and naval fields

NeurIPS Conference 2005 Conference Paper

A Computational Model of Eye Movements during Object Class Detection

  • Wei Zhang
  • Hyejin Yang
  • Dimitris Samaras
  • Gregory Zelinsky

We present a computational model of human eye movements in an ob- ject class detection task. The model combines state-of-the-art computer vision object class detection methods (SIFT features trained using Ad- aBoost) with a biologically plausible model of human eye movement to produce a sequence of simulated fixations, culminating with the acqui- sition of a target. We validated the model by comparing its behavior to the behavior of human observers performing the identical object class detection task (looking for a teddy bear among visually complex non- target objects). We found considerable agreement between the model and human data in multiple eye movement measures, including number of fixations, cumulative probability of fixating the target, and scanpath distance.

NeurIPS Conference 2005 Conference Paper

The Role of Top-down and Bottom-up Processes in Guiding Eye Movements during Visual Search

  • Gregory Zelinsky
  • Wei Zhang
  • Bing Yu
  • Xin Chen
  • Dimitris Samaras

To investigate how top-down (TD) and bottom-up (BU) information is weighted in the guidance of human search behavior, we manipulated the proportions of BU and TD components in a saliency-based model. The model is biologically plausible and implements an artificial retina and a neuronal population code. The BU component is based on feature- contrast. The TD component is defined by a feature-template match to a stored target representation. We compared the model’s behavior at differ- ent mixtures of TD and BU components to the eye movement behavior of human observers performing the identical search task. We found that a purely TD model provides a much closer match to human behavior than any mixture model using BU information. Only when biological con- straints are removed (e. g. , eliminating the retina) did a BU/TD mixture model begin to approximate human behavior.

YNIMG Journal 2003 Journal Article

Sexual dimorphism and asymmetries in the gray–white composition of the human cerebrum

  • John S Allen
  • Hanna Damasio
  • Thomas J Grabowski
  • Joel Bruss
  • Wei Zhang

Using high resolution MRI scans and automated tissue segmentation, gray and white matter (GM, WM) volumes of the frontal, temporal, parietal, and occipital lobes, cingulate gyrus, and insula were calculated. Subjects included 23 male and 23 female healthy, right-handed subjects. For all structures, male volumes were greater than female, but the gray/white (G/W) ratio was consistently higher across structures in women than men. Sexual dimorphism was greater for WM than GM: most of the G/W ratio sex differences can be attributed to variation in WM volume. The corpus callosum, although larger in men, is less sexually dimorphic than the WM as a whole. Several regions demonstrate pair-wise asymmetries in G/W ratio and WM volume. Both the cingulate gyrus and insula exhibit strong asymmetries. The left cingulate gyrus is significantly larger than the right, and the G/W ratio of the left insula is significantly greater than that of the right. Although statistically significant sex differences and asymmetries are present at this level of analysis, we argue that researchers should be wary of ascribing cognitive functional significance to these patterns at this time. This is not to say, however, that these patterns are not important for understanding the natural history of the human brain, and its evolution and development.

EAAI Journal 1994 Journal Article

Verification of real-time programs by a knowledge-based strategy

  • Wei Zhang
  • Jiren Liu
  • Huatian Li

A verification system for real-time programs must provide a means of showing that the programs are logically correct and of proving that all timing constraints are met. To prove a program correct, according to traditional methods, all possible execution sequences (i. e. traces) must be shown to satisfy the specifications. That, however, will lead to the exponential explosion of the traces. This paper adopts an Artificial Intelligence technique to verify real-time programs, and proposes a method by which the real-time programs are proved just in a so-called “Reasonable Trace Space” (RTS). The RTS is a subset of the set of all traces (in short, the ATS) and is determined by a knowledge base which is provided by the program designer or a software verification expert. If a real-time program is proved correct in the RTS (which is much smaller than the ATS), it can be concluded that it is error free in the sense of the knowledge base. A knowledge-based trace-generation algorithm which enumerates all reasonable traces and a trace-verification algorithm which decides whether a trace is correct or not, are presented. Their performances are also analyzed.

v2026.09.13