Arrow Research search

Author name cluster

Yan Zhang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

128 papers
2 author rows

Possible papers

128

EAAI Journal 2026 Journal Article

A Center-Focused Transformer for hyperspectral image classification

  • Chaoxu Yang
  • Jia Duan
  • Xi Liu
  • Lianchong Zhang
  • Jiangbing Sun
  • Yan Zhang
  • Wei Ren

In recent years, transformer-based methods have achieved remarkable progress in hyperspectral image classification (HSIC). However, they often rely heavily on extensive training samples to achieve optimal performance. Moreover, these methods frequently fail to adequately capture diverse local spectral–spatial correlations and multi-granular features inherent in hyperspectral images (HSIs). Crucially, existing approaches often overlook the pivotal role of the target center pixel. Their attention mechanisms tend to focus on irrelevant background regions, thereby reducing feature discriminability and degrading classification accuracy. To address these challenges, we propose a novel Center-Focused Transformer (CFT) framework that seamlessly integrates multi-scale spectral–spatial fusion for HSIC. Our framework comprises three key components. First, the Spectral–Spatial Fusion (SSF) mechanism integrates local and global dependencies by employing PCA alongside a Superpixel Graph Feature Extraction (SGFE) block. Second, the Multi-Granular Feature Enhancement (MGFE) approach strengthens spectral–spatial interactions through patch augmentation, a HybridConv block, and a Multi-Scale CBAM (MS-CBAM) block. Finally, the Focus Center Transformer (FCT) strategy explicitly emphasizes the importance of the central pixel for precise classification by incorporating Gaussian Positional Embedding (GPE) and cross-layer aggregation. Extensive experiments on four public datasets demonstrate that the proposed CFT consistently outperforms state-of-the-art methods, highlighting its potential for practical engineering applications. • A novel Center-Focused Transformer (CFT) framework is proposed for hyperspectral image classification. • The CFT model integrates a Spectral–Spatial Fusion (SSF) mechanism to effectively capture local and global dependencies. • A Multi-Granular Feature Enhancement (MGFE) approach is introduced to model multi-scale features in both spectral and spatial dimensions. • A Focus Center Transformer (FCT) strategy with Gaussian positional embedding is proposed to improve classification accuracy. • The CFT consistently outperforms state-of-the-art methods on four public datasets, showcasing its potential for engineering applications.

AAAI Conference 2026 Conference Paper

Beyond Counting: Evaluating Abstract and Emotional Reasoning in Vision-Language Models

  • Yuan Zhou
  • Yan Zhang
  • Jianlong Chang
  • Xin Gu
  • Ying Wang
  • Kun Ding
  • Guangwen Yang
  • Shiming Xiang

Despite the rapid progress of Vision Language Models (VLMs), existing benchmarks still concentrate on coarse-grained object recognition or simple relational reasoning, leaving the fine-grained and higher-order reasoning abilities of these systems largely unexamined. To bridge this critical evaluation gap, we introduce EmojiGrid, a novel diagnostic benchmark specifically designed to probe these fine-grained and higher-order skills. Leveraging the universal and semantically rich nature of emojis, we synthesize a grid‑based visual dataset paired with 29,000+ QA pairs. Each pair is explicitly anchored in a three-level cognitive taxonomy comprising (i) Perception and Information Extraction, (ii) Relational and Structural Reasoning, and (iii) Abstraction and Advanced Cognition. These dimensions further decompose into nine categories covering a broad range of cognitive skills, including counting, spatial relations, compositional logic, semantic sentiment, and related higher-order reasoning tasks. Our extensive evaluation of 25 state-of-the-art open-source and proprietary VLMs reveals a significant performance gap between foundational perceptual tasks and higher-level cognitive abilities, particularly in abstraction and advanced emotional reasoning. Notably, all models struggle with compositional logic, spatial consistency, and especially emotional and semantic understanding. EmojiGrid provides a quantifiable, fine-grained benchmark to diagnose VLM limitations and guides future progress toward models that can truly perceive, reason about, and interpret complex, symbol-rich visual scenes.

JBHI Journal 2026 Journal Article

BLADE: Breast Lesion Analysis with Domain Expertise for DCE-MRI Diagnosis

  • Zhitao Wei
  • Yi Dai
  • Yanting Liang
  • Chinting Wong
  • Yanfen Cui
  • Xiaobo Chen
  • Zhihe Zhao
  • Xiaodong Zheng

Dynamic Contrast-Enhanced Magnetic Reso nance Imaging (DCE-MRI) is pivotal in breast cancer diag nosis, yet radiologists face challenges in interpreting its complex data due to the lack of robust automated tools. Current lesion diagnosis systems struggle with limited datasets and insufficient integration of domain knowledge. To overcome these limitations, we propose Breast Lesion Analysis with DomainExpertise(BLADE), anoveldiagnosis framework that synergizes deep learning with clinical ex pertise. BLADE leverages a pre-trained vertical foundation model (optimized via Momentum Contrast on 2. 1 million MRI slices) as its encoder, ensuring robust feature extraction. Crucially, the system incorporates prior multi-phasic hemodynamic knowledge to emulate radiologists' diagnos tic reasoning and introduces a Breast Imaging Reporting and Data System (BI-RADS)-based constraint during training to align predictions with clinical standards. Extensive experiments demonstrate that BLADE outperforms state of-the-art methods, achieving an Area Under the Curve (AUC) of 0. 9228 and 0. 9553 on two external test datasets, respectively. Notably, BLADE significantly enhances clin ical workflow; when used as an assistive tool, BLADE improves diagnostic accuracy by 14. 31%, surpassing stan daloneperformanceofclinicians. This workbridgesthegap between AI-driven analysis and clinical practice in breast MRI interpretation. The source code is available at https://github.com/GDPHMediaLab/BLADE.

AAMAS Conference 2026 Conference Paper

Cross-Domain Alignment with Fine Geometric Perception for Detail-Preserving Point Cloud Completion

  • Chen Huang
  • Haobo Ma
  • Yan Zhang
  • Chao Yang
  • Jianhua Song

Point cloud completion involves inferring and reconstructing the full structure of an object or scene from incomplete 3D point cloud data. Deep learning-based methods typically use encoder-decoder architecturestolearngeometricpriorsfrompartialinputsforreconstruction. However, these methods often prioritize global features over local geometric details, leading to coarse completions lacking high-frequency information. Sequential application of such models can also cause error accumulation and increased computational costs. To address these issues, we propose CAM-FGP, a Cross-domain Alignment Method with Fine Geometric Perception, designed to enhance structural integrity and restore details, especially in regions with missing geometry. CAM-FGP first employs a Fine Geometry Detail Extraction Network (FGDE) to gather highresolution local details from visible point clouds while integrating low-resolution global information to reinforce the missing areas’ structure. Then, aHierarchicalOptimalTransportNetwork(HOTN) aligns multi-source point cloud distributions, improving the transferability of local geometric features. Lastly, CAM-FGP utilizes a multi-stage hidden state completion and fusion strategy to merge local and global features. This approach preserves continuity, reduces memory-induced information loss, and lowers computational costs. CAM-FGP achieves state-of-the-art performance on several benchmark datasets, demonstrating its superiority in point cloud completion.

YNIMG Journal 2026 Journal Article

Distinct frontal lobe subregions mediate the emergence and reporting of visual consciousness

  • Yan Zhang
  • Xiarong Li
  • Zhenlan Jin
  • Junjun Zhang
  • Ling Li

Persistent debate surrounds whether the frontal lobe supports the emergence or reporting of consciousness, raising the hypothesis that distinct frontal subregions may support these processes. We addressed this by combining electroencephalography (EEG) with eye-tracking in Report and No-Report paradigms. Eye-movement features distinguished conscious and unconscious trials in the no-report task. Event-related potential analyses showed that the Visual Awareness Negativity (VAN) was independent of reporting, whereas P3b occurred only with explicit reports. Importantly, the frontal Dorsal Attention Network (DAN) supported the emergence of consciousness, independent of post-perceptual reporting, as shown by multivoxel pattern analysis showing that a classifier's ability to decode visual consciousness generalized bidirectionally between report and no-report tasks. In contrast, frontal components of the Default Mode Network (DMN) and Frontoparietal Control Network (FPN) encoded visual consciousness only when explicit reports were required, indicating roles in reporting. These findings demonstrate a functional dissociation within the frontal lobe and refine the anatomical framework for the neural basis of visual consciousness.

AAAI Conference 2026 Conference Paper

Graph-Driven Domain Co-Adaptation for Cross-Domain Image Quality Assessment

  • Shun Zhu
  • Xichen Yang
  • Yan Zhang
  • Tianshu Wang
  • Zhongyuan Mao
  • Tianyin Li
  • Zhuoyan Sun
  • Xiaobo Shen

As a typical information medium, images are widely utilized across various scenarios. Measuring image quality accurately is meaningful for the subsequent usability of images. However, significant variations exist in image types and distortion types in different scenarios. And, acquiring labeled images for each specific scenario is time-consuming and labor-intensive. Consequently, designing cross-domain image quality assessment (IQA) that generalizes across different scenarios remains a substantial challenge. Existing cross-domain IQA methods primarily focus on content relevance while neglecting distortion differences, leading to limited applicability while distortion fluctuates. To address these limitations, a graph-driven domain co-adaptation framework for cross-domain IQA (GDCIQA) is proposed. Firstly, a graph knowledge sharing (GKS) module that constructs graphs via inter-domain distortion relevance has been proposed. GKS employs graph neural networks to update quality-aware features in the source domain by leveraging target-domain representations. Secondly, the proposed co-adaptation learning (CAL) mechanism can enable joint optimization of different modules, which ensures comprehensive sharing of quality-aware and distortion-related information. Finally, a domain adaptation framework has been designed to train models effectively on labeled source images, yielding target-domain-optimized IQA models. Experimental results demonstrate that GDCIQA achieves higher accuracy and stability in cross-domain scenarios. The proposed GKS and CAL can advance cross-domain IQA research.

JBHI Journal 2026 Journal Article

IIMCNet: Intra- and Inter-Modality Correlation Network for Hybrid EEG-fNIRS Brain-Computer Interface

  • Xiaoyang Yuan
  • Yan Zhang
  • Peter Rolfe

Hybrid Brain-Computer Interface (BCI) enhances accuracy and reliability by leveraging the complementary information provided by multi-modality signal fusion. EEG-fNIRS, a fusion of electroencephalogram (EEG) and functional near-infrared spectroscopy (fNIRS), have emerged as the suitable techniques for real-world BCI applications due to their portability and economic viability. Existing methods typically focus on the high-level feature representation with late-fusion or early-fusion strategies during the recognition tasks. However, they usually overlook the joint feature extraction of both intra-modality and inter-modality, which is crucial for optimizing BCI performance. In this study, we introduce an Intra- and Inter-modality Correlation Network (IIMCNet) to integrate both the inherent features derived from individual modalities: EEG, deoxygenated hemoglobin (HbR), and oxygenated hemoglobin (HbO), as well as the cross-modality features between EEG-HbR, EEG-HbO, and HbR-HbO data. The intra-modality correlation features are generated using a late fusion method (Intra-net), which combines the uni-modality features extracted by E-Net and f-Net. Concurrently, the inter-modality correlation features are extracted using an early fusion method (Inter-net). Inter-net is consist of three dilated convolution-based C-Nets that focus on neurovascular coupling across modalities. Finally, three intra-modality features, three inter-modality features, and the concatenate hybrid feature are fed into deep supervision module to enhance robustness and accuracy. Experiment results demonstrate the IIMCNet exhibits superior performance compared to methods that rely solely on either intra-modality or inter-modality correlation networks. Furthermore, IIMCNet outperforms other state-of-the-art methods in motor imagery and mental arithmetic tasks, respectively.

AAAI Conference 2026 Conference Paper

LookFlow: Training-Free and Efficient High-Resolution Image Synthesis via Dynamic Lookahead Guidance Flow

  • Yuan Zhou
  • Yan Zhang
  • Jianlong Chang
  • Xin Gu
  • Ying Wang
  • Kun Ding
  • Guangwen Yang
  • Shiming Xiang

Rectification flow Transformers (RFTs) have shown promising performance in diffusion-based image synthesis but are typically confined to lower-resolution scenarios, limiting their ability to generate high-resolution images. Existing resolution extrapolation approaches often suffer from excessive computational overhead, resulting in prolonged inference times. We propose LookFlow, a training-free high-resolution synthesis framework that accelerates inference while preserving visual quality. Building on pretrained text-to-image RFTs, LookFlow employs a dynamic lookahead guidance flow mechanism to refine high-resolution velocity predictions by leveraging multi-timestep lookahead information extracted from a low-resolution flow. Additionally, reusing temporally similar features across consecutive timesteps drastically reduces computation and significantly decreases inference time overhead. Extensive experiments on COCO demonstrate that LookFlow robustly scales resolutions from 4× to 25×, achieving up to a maximum speedup of 2.01× while maintaining competitive visual fidelity.

AAAI Conference 2026 Conference Paper

Mitigating Error Accumulation in Knowledge Editing for Multi-Hop Question Answering

  • Jiaxin Guo
  • Hao Sun
  • Wenhao Zhang
  • Xuanbo Fan
  • Yan Zhang

Knowledge editing (KE) has emerged as an effective approach for updating factual information in large language models (LLMs) without the need for full retraining. Most of the existing methods for addressing the "ripple effect" in KE adopt a chain-structured reasoning process, making them vulnerable to error accumulation from early incorrect steps. Moreover, their conflict detection mechanisms are often susceptible to the LLM's inherent confirmation bias, further undermining the reliability of the editing process. To overcome these challenges, we propose Tree of Editing (ToE), a tree-structured, retrieval-enhanced knowledge editing framework designed to support robust reasoning under factual updates. ToE expands reasoning paths using a breadth-first strategy combined with score-guided beam search, enabling diverse and error-tolerant inference. Besides, we introduce an observer to objectively update knowledge, avoiding the bias caused by LLMs' over-confidence. Experimental results on two benchmarks, namely MQuAKE-CF (targeting ripple-aware editing) and DUNE (free-form editing), demonstrate that ToE framework significantly outperforms existing methods.

AAMAS Conference 2026 Conference Paper

Multimodal Emotion Recognition in Conversation via Large Language Models and Global-Local Cross-Domain Graphs

  • Haobo Ma
  • Chen Huang
  • Yan Zhang
  • Chao Yang
  • Jianhua Song

Multimodal Emotion Recognition in Conversation (MERC) aims to identify emotions in target utterances using multimodal data and has garnered significant interest due to its applications in conversational AI. Recognition accuracy hinges on effectively integrating multimodal cues and contextual information. However, local noise and global outliers often impair performance, while traditional approaches based on simple feature concatenation struggle to capture complex cross-modal interactions. To address these challenges, we propose LLM-EmoGraph, a novel framework that combines large language models (LLMs) with a global-local cross-domain graph architecture. Specifically, LLM-EmoGraph leverages multimodal masking strategies, a large-scale cross-domain multi-graph pretraining to improve transferability across modalities and graph structures. Then, LLM-EmoGraph further introduces an adaptive dual-scale feature fusion strategy to align semantic features across text, speech, and visual inputs. In addition, a weakly supervised hierarchical emotion classification scheme enhanced by LLMs boosts robustness and accuracy. Experiments on two benchmark datasets show that LLM-EmoGraph significantly outperforms existing methods.

JBHI Journal 2026 Journal Article

SAVLT: Structure-Aware Vision-Language Tuning for Multi-Center Cervical OCT Diagnosis

  • Mi Yin
  • Yuchen Pei
  • Yixiong Zou
  • Yan Zhang
  • Yutao Ma

Cervical optical coherence tomography (OCT) enables micrometer-scale visualization of tissue, yet trustworthy diagnosis under limited supervision remains challenging. While vision-language models (VLMs) offer a solution by leveraging consistent clinical semantics, their adaptation is hindered by confounding artifacts that masquerade as biological structures. This phenomenon hijacks global attention, obscuring the pathological layer degradation that is crucial for diagnosis. To overcome this limitation, we propose SAVLT, a structure-aware tuning framework that adapts VLMs via parameter-efficient fine-tuning. To shift from global matching to anatomical grounding, SAVLT introduces a region-aware spatial attention (RaSA) module that enforces spatial constraints. RaSA purifies visual representations from non-biological noise, restoring the model's focus on intra-tissue structural integrity. Furthermore, a dual-constraint objective couples image-text alignment with learnable visual prototypes to stabilize optimization against multi-center domain shifts and prompt variations. Validated across a multi-center cohort and two external test datasets, SAVLT delivers robust few-shot generalization and clinical interpretability, establishing a reliable paradigm for deploying foundation models in heterogeneous OCT imaging. Source code is publicly available at https://github.com/rabbit-my/SAVLT.

AAAI Conference 2026 Conference Paper

SegMem-RAG: Adaptive Memory for Retrieval-Augmented Generation in Open-Ended Knowledge Environments

  • Xuanbo Fan
  • Tianqi Zhao
  • Yi Cheng
  • Chi Xiu
  • Jiaxin Guo
  • Boci Peng
  • Bingjing Xu
  • Jessica Zhang

Retrieval-Augmented Generation (RAG) improves the factual accuracy of large language models by grounding responses in external content. However, most RAG systems assume access to static and well-organized corpora with fixed retrieval logic. In practice, real-world sources are heterogeneous and unlabeled, including user-uploaded documents, manuals, and datasets. Effective access in such settings requires adaptive and self-directed retrieval behavior. We present SegMem‑RAG, a memory-augmented RAG framework that learns to route queries across multiple unlabeled corpora based on experience. It incrementally updates a structured memory and uses self-reflection to guide retrieval over time without supervision. Experimental results demonstrate that SegMem‑RAG significantly outperforms recent baselines in generation quality on multi-corpus QA tasks.

AAAI Conference 2026 Conference Paper

SIAM: Towards Generalizable Articulated Object Modeling via Single Robot-Object Interaction

  • Yuyan Liu
  • Li Zhang
  • Di Wu
  • Yan Zhang
  • Anran Huang
  • Zhi Wang
  • Liu Liu
  • Dan Guo

Articulated object modeling, which represents interconnected rigid bodies with their geometry, part segmentation, articulation tree, and physical properties, is crucial for robotic perception and manipulation. Recently existing methods like SAGCI leverage Interactive Perception (IP) to refine models through robot interaction. However, SAGCI suffers from prior-dependency (requiring initialization), neglects kinematic/dynamic constraints, and generates non-watertight meshes. To overcome these limitations, we propose SIAM, a novel framework for efficient and generalizable Single-Interaction Articulated Modeling. Given an initial point cloud, SIAM first enables minimal robot interaction to trigger object motion. It then precisely segments parts by analyzing point cloud differences pre- and post-interaction. For joint parameter estimation, we introduce an optimization incorporating novel kinematic energy constraints, enhancing physical consistency. Finally, we reconstruct a high-quality, topologically watertight mesh by learning 3D Gaussian Primitives from multi-view RGB-D observations under deformation. Extensive experiments on the PartNet-Mobility benchmark demonstrate state-of-the-art articulation modeling performance. Successful real-world deployment with an xArm robot further validates the framework's practicality and transferability. SIAM achieves accurate, prior-free modeling with significantly reduced interaction cost.

AAAI Conference 2026 Conference Paper

UniFit: Towards Universal Virtual Try-on with MLLM-Guided Semantic Alignment

  • Wei Zhang
  • Yeying Jin
  • Xin Li
  • Yan Zhang
  • Xiaofeng Cong
  • Cong Wang
  • Fengcai Qiao
  • Zhichao Lian

Image-based virtual try-on (VTON) aims to synthesize photorealistic images of a person wearing specified garments. Despite significant progress, building a universal VTON framework that can flexibly handle diverse and complex tasks remains a major challenge. Recent methods explore multi-task VTON frameworks guided by textual instructions, yet they still face two key limitations: (1) semantic gap between text instructions and reference images, and (2) data scarcity in complex scenarios. To address these challenges, we propose UniFit, a universal VTON framework driven by a Multimodal Large Language Model (MLLM). Specifically, we introduce an MLLM-Guided Semantic Alignment Module (MGSA), which integrates multimodal inputs using an MLLM and a set of learnable queries. By imposing a semantic alignment loss, MGSA captures cross-modal semantic relationships and provides coherent and explicit semantic guidance for the generative process, thereby reducing the semantic gap. Moreover, by devising a two-stage progressive training strategy with a self-synthesis pipeline, UniFit is able to learn complex tasks from limited data. Extensive experiments show that UniFit not only supports a wide range of VTON tasks, including multi-garment and model-to-model try-on, but also achieves state-of-the-art performance.

ICML Conference 2025 Conference Paper

A Theory for Conditional Generative Modeling on Multiple Data Sources

  • Rongzhen Wang
  • Yan Zhang
  • Chenyu Zheng
  • Chongxuan Li
  • Guoqiang Wu

The success of large generative models has driven a paradigm shift, leveraging massive multi-source data to enhance model capabilities. However, the interaction among these sources remains theoretically underexplored. This paper provides a first attempt to fill this gap by rigorously analyzing multi-source training in conditional generative modeling, where each condition represents a distinct data source. Specifically, we establish a general distribution estimation error bound in average total variation distance for conditional maximum likelihood estimation based on the bracketing number. Our result shows that when source distributions share certain similarities and the model is expressive enough, multi-source training guarantees a sharper bound than single-source training. We further instantiate the general theory on conditional Gaussian estimation and deep generative models including autoregressive and flexible energy-based models, by characterizing their bracketing numbers. The results highlight that the number of sources and similarity among source distributions improve the advantage of multi-source training. Simulations and real-world experiments validate our theory.

TIST Journal 2025 Journal Article

Advancing Session-Based Recommendations with Atten-Mixer+: Dynamic and Adaptive Multi-Level Intent Mining

  • Peiyan Zhang
  • Jiayan Guo
  • Chaozhuo Li
  • Liying Kang
  • Jaeboum Kim
  • Jie Xu
  • Xi Zhang
  • Yan Zhang

Session-Based Recommendation (SBR) systems, traditionally reliant on complex Graph Neural Networks (GNNs), often face challenges with marginal performance improvements despite increased model complexity. In this article, we dissect the classical GNN-based SBR models and empirically find that the sophisticated GNN propagations might be redundant, given the readout module plays a significant role in GNN-based models. Based on this observation, we introduce Atten-Mixer+, an advanced iteration of our previously developed Multi-Level Attention Mixture Network (Atten-Mixer). Atten-Mixer+ forgoes GNN propagation in favor of a dynamic and adaptive readout process, tailored to the unique characteristics of each session. Different from the vanilla version, Atten-Mixer+ features the Adaptive Intent Scaler (AIS) layer, which dynamically determines the depth of multi-level user intent analysis and a soft allocation approach for generating user intent queries across entire user interaction sequences. This innovative design allows Atten-Mixer+ to capture a nuanced and comprehensive understanding of user behaviors, overcoming the limitations of fixed-length analysis. Empirical evaluations on benchmark datasets highlight Atten-Mixer+’s superior efficiency and effectiveness, marking a significant step forward in the predictive accuracy of SBR systems.

YNIMG Journal 2025 Journal Article

Altered brain network dynamics during rumination in remitted depression

  • Su Shu
  • Wenwen Ou
  • Mohan Ma
  • Hairuo He
  • Qianqian Zhang
  • Mei Huang
  • Wentao Chen
  • Aoqian Deng

Rumination is a known risk factor for depression relapse. Understanding its neurobiological mechanisms during depression remission can inform strategies to prevent relapse, yet the temporal dynamics of brain networks during rumination in remitted depression remain unclear. Here, we collected rumination induction fMRI data from 42 patients with remitted depression and 41 healthy controls (HCs). Using an energy landscape approach, we investigated the temporal dynamics of brain networks during rumination. The appearance frequency (AF) and transition frequency (TF) metrics were defined to quantify the dynamic properties of brain states. Patients during remission showed higher levels of rumination than HCs. Both groups exhibited four brain states during rumination, which consisted of complementary network group activation (states 1 and 2, states 3 and 4). In patients, the AFs of and reciprocal TFs between states 1 and 2 during rumination were significantly increased, while AFs of states 3 and 4 and reciprocal TFs involving states 1-3, 1-4, 2-3, and 2-4 were decreased, both when compared to HCs and relative to patients themselves during distraction. Moreover, we found that for patients, the AF of state 1 was negatively correlated with rumination levels and marginally positively associated with attention, while the AF of state 2 was negatively associated with performance on attention tasks. Our study revealed altered dynamic characteristics of brain states composed of network groups during rumination in remitted depression. Additionally, the findings suggest that heightened self-focus linked to rumination may impair the brain's ability to efficiently allocate attentional resources.

TMLR Journal 2025 Journal Article

Any-Property-Conditional Molecule Generation with Self-Criticism using Spanning Trees

  • Alexia Jolicoeur-Martineau
  • Aristide Baratin
  • Kisoo Kwon
  • Boris Knyazev
  • Yan Zhang

Generating novel molecules is challenging, with most representations of molecules leading to generative models producing many invalid molecules. Spanning Tree-based Graph Generation (STGG) is a promising approach to ensure the generation of valid molecules, outperforming state-of-the-art generative models for unconditional generation. In practice, it is desirable to generate molecules conditional on one or multiple target properties rather than unconditionally. Thus, we extend STGG to multi-property conditional generation. Our approach, STGG+, incorporates a modern Transformer architecture, random masking of properties during training (enabling conditioning on any subset of properties and classifier-free guidance), an auxiliary property-prediction loss (allowing the model to self-criticize molecules and select the best ones), and other improvements. We show that STGG+ achieves state-of-the-art performance on in-distribution and out-of-distribution conditional generation, as well as reward maximization.

AAAI Conference 2025 Conference Paper

APKGC: Noise-enhanced Multi-Modal Knowledge Graph Completion with Attention Penalty

  • Yue Jian
  • Xiangyu Luo
  • Zhifei Li
  • Miao Zhang
  • Yan Zhang
  • Kui Xiao
  • Xiaoju Hou

Multimodal knowledge graphs (MMKG) store structured world knowledge enriched with multimodal descriptive information. However, MMKG often faces the challenge of incompleteness. The primary objective of multimodal knowledge graph completion (MMKGC) is to predict missing entities within MMKG. Current MMKGC methods struggle with addressing the issue of over-trust attention and how to enhance the robustness of the model. To overcome these problems, we introduce APKGC, a noise-enhanced multimodal method for knowledge graph completion with attention penalty. APKGC effectively adjusts the attention scores in the language model and alleviates over-trust attention through a specifically designed attention penalty module. Additionally, an adaptive noise sampling module is proposed to supplement the entity's multimodal information, thereby enhancing the model's robustness. Experimental evaluation demonstrates that APKGC excels in overcoming these challenges. Compared to the existing state-of-the-art MMKGC model, APKGC improves Hit@1 by 3.3% on the DB15K dataset and by 3.4% on the MKG-W dataset.

JBHI Journal 2025 Journal Article

Automatic Brain Segmentation for PET/MR Dual-Modal Images Through a Cross-Fusion Mechanism

  • Hongyan Tang
  • Zhenxing Huang
  • Wenbo Li
  • Yaping Wu
  • Jianmin Yuan
  • Yang Yang
  • Yan Zhang
  • Jing Qin

The precise segmentation of different brain regions and tissues is usually a prerequisite for the detection and diagnosis of various neurological disorders in neuroscience. Considering the abundance of functional and structural dual-modality information for positron emission tomography/magnetic resonance (PET/MR) images, we propose a novel 3D whole-brain segmentation network with a cross-fusion mechanism introduced to obtain 45 brain regions. Specifically, the network processes PET and MR images simultaneously, employing UX-Net and a cross-fusion block for feature extraction and fusion in the encoder. We test our method by comparing it with other deep learning-based methods, including 3DUXNET, SwinUNETR, UNETR, nnFormer, UNet3D, NestedUNet, ResUNet, and VNet. The experimental results demonstrate that the proposed method achieves better segmentation performance in terms of both visual and quantitative evaluation metrics and achieves more precise segmentation in three views while preserving fine details. In particular, the proposed method achieves superior quantitative results, with a Dice coefficient of 85. 73% $\pm$ 0. 01%, a Jaccard index of 76. 68% $\pm$ 0. 02%, a sensitivity of 85. 00% $\pm$ 0. 01%, a precision of 83. 26% $\pm$ 0. 03% and a Hausdorff distance (HD) of 4. 4885 $\pm$ 14. 85%. Moreover, the distribution and correlation of the SUV in the volume of interest (VOI) are also evaluated (PCC > 0. 9), indicating consistency with the ground truth and the superiority of the proposed method. In future work, we will utilize our whole-brain segmentation method in clinical practice to assist doctors in accurately diagnosing and treating brain diseases.

JBHI Journal 2025 Journal Article

Blind Source Separation-Embedded Electroencephalogram Microstate Trajectory Modeling for Generalized Anxiety Disorder Identification

  • Hongzuo Chu
  • Guanyi Lv
  • Mohan Ma
  • Lingsi Zeng
  • Fanyu Meng
  • Hui Liang
  • Yan Zhang
  • Shuang Liu

Generalized anxiety disorder (GAD) diagnosis remains challenging due to lacking reliable biomarkers. Electroencephalography(EEG) microstate analysis shows promise in detecting GAD-related neural dynamics, but its clinical application is limited by insufficient spatial resolution and sensitivity. To address this challenge, we propose a novel framework integrating fast independent component analysis (FastICA) with microstate analysis to enhance spatial specificity in EEG signal decomposition, consequently, GAD can be more accurately identified. By isolating dominant independent components and projecting them onto the channels with the highest weights, our method effectively reduces volume conduction effects and signal mixing across channels, thereby sharpening the spatial topography of EEG microstates and improving spatial resolution. In a cohort of 28 GAD patients and 28 healthy controls, the FastICA-enhanced microstate features exhibited stronger intergroup differences in key parameters-including significantly increased occurrence, coverage, and duration of microstate A*-and revealed altered transition probabilities (e. g. , C*→B*, p = 0. 045), indicating improved discriminative power for anxiety-specific patterns. Furthermore, classification using a Support Vector Machine (SVM) with enhanced features achieved improved sensitivity (3. 6% increase) and precision (5. 5% increase) compared to the standard microstate approach. These findings underscore the potential of blind source separation techniques to refine EEG-based biomarkers for anxiety disorders. Our work not only advances the technical resolution of microstate analysis but also provides a clinically translatable pathway for objective GAD diagnosis. Future studies could extend this framework to broader psychiatric conditions and explore multimodal machine learning models for enhanced robustness.

YNICL Journal 2025 Journal Article

Brain network dynamics during rumination relate to relapse of depression

  • Su Shu
  • Yumeng Ju
  • Mi Wang
  • Wenwen Ou
  • Mohan Ma
  • Qianqian Zhang
  • Mei Huang
  • Hairuo He

BACKGROUND: Rumination is a maladaptive cognitive style and a risk factor for relapse of depression. However, the clinically relevant pattern of dynamic network reconfiguration during rumination in remitted depression and its implication in relapse remained unclear. METHODS: We employed a rumination induction neuroimaging paradigm in which subjects would be guided into an active rumination state and a distraction state. Forty-two patients with remitted depression were involved. Participants underwent assessments of rumination behavior and imaging tasks, and were then monitored for two year to assess the potential relapse of depression. A time-resolved community detection approach was applied to investigate the temporal dynamics of brain networks, and the dynamic network properties including flexibility and integration were analyzed. RESULTS: = 0.036). Moreover, elastic net regression indicated that dynamic network features could predict two-year relapse outcomes with moderate accuracy (AUC = 0.70). CONCLUSIONS: Our findings reveal a potential mechanistic link between the brain network dynamics during rumination and relapse of depression, shedding light on the intricate relationship between cognitive-affective processes, neural dynamics, and the potential vulnerability to depression recurrence.

AAAI Conference 2025 Conference Paper

BUFF: Bayesian Uncertainty Guided Diffusion Probabilistic Model for Single Image Super-Resolution

  • Zihao He
  • Shengchuan Zhang
  • Runze Hu
  • Yunhang Shen
  • Yan Zhang

Super-resolution (SR) techniques are critical for enhancing image quality, particularly in scenarios where high-resolution imagery is essential yet limited by hardware constraints. Existing diffusion models for SR have relied predominantly on Gaussian models for noise generation, which often fall short when dealing with the complex and variable texture inherent in natural scenes. To address these deficiencies, we introduce the Bayesian Uncertainty Guided Diffusion Probabilistic Model (BUFF). BUFF distinguishes itself by incorporating a Bayesian network to generate high-resolution uncertainty masks. These masks guide the diffusion process, allowing for the adjustment of noise intensity in a manner that is both context-aware and adaptive. This novel approach not only enhances the fidelity of super-resolved images to their original high-resolution counterparts but also significantly mitigates artifacts and blurring in areas characterized by complex textures and fine details. The model demonstrates exceptional robustness against complex noise patterns and showcases superior adaptability in handling textures and edges within images. Empirical evidence, supported by visual results, illustrates the model's robustness, especially in challenging scenarios, and its effectiveness in addressing common SR issues such as blurring. Experimental evaluations conducted on the DIV2K dataset reveal that BUFF achieves a notable improvement, with a +0.61 increase compared to baseline in SSIM on BSD100, surpassing traditional diffusion approaches by an average additional +0.20dB PSNR gain. These findings underscore the potential of Bayesian methods in enhancing diffusion processes for SR, paving the way for future advancements in the field.

IJCAI Conference 2025 Conference Paper

DGCPL: Dual Graph Distillation for Concept Prerequisite Relation Learning

  • Miao Zhang
  • Jiawei Wang
  • Jinying Han
  • Kui Xiao
  • Zhifei Li
  • Yan Zhang
  • Hao Chen
  • Shihui Wang

Concept prerequisite relations determine the learning order of knowledge concepts in one domain, which has an important impact on teachers' course design and students' personalized learning. Current research usually predicts concept prerequisite relations from the perspective of knowledge, and rarely pays attention to the role of learners' learning behavior. We propose a Dual Graph Distillation Method for Concept Prerequisite Relation Learning (DGCPL). Specifically, DGCPL constructs a dual graph structure from both the knowledge and learning behavior perspectives, and captures the high-order knowledge features and learning behavior features through the concept-resource hypergraph and the learning behavior graph respectively. In addition, we introduce a gated knowledge distillation to fuse the structural information of concept nodes in the two graphs, so as to obtain a more comprehensive concept embedding representation and achieve accurate prediction of prerequisite relations. On three public benchmark datasets, we compare DGCPL with eight graph-based baseline methods and five traditional classification baseline methods. The experimental results show that DGCPL achieves state-of-the-art performance in learning concept prerequisite relations. Our code is available at https: //github. com/wisejw/DGCPL.

EAAI Journal 2025 Journal Article

Dual-path information enhanced pyramid Unet for COVID-19 lung infection segmentation

  • Yan Zhang
  • Qi Mao
  • Yi Tian
  • Wenfeng Wang
  • Lijia Ren
  • Haibo Li

The coronavirus disease 2019 (COVID-19) pandemic has brought computer-aided diagnosis into the spotlight. COVID-19 computed tomography (CT) images often have redundant background, and the proportion of infected area and healthy areas is unbalanced, which may confuse the model and make it difficult to identify the infected areas, resulting in the inability of the model to make accurate determinations. The low contrast between infected and normal areas makes the boundaries difficult to distinguish. The infected areas are small, irregular and unevenly distributed. Thus, models often suffers from incomplete and inadequate segmentation. To solve these challenges, a dual-path information enhanced pyramid Unet network (DIEP-Unet) was proposed for COVID-19 infection segmentation. First, an approach of coarse multi-scale feature map (CMFM) was proposed, which included a Fourier image boundary enhancement (FIBE) module to enhance the boundary features of infected areas, and a multi-scale information extraction and fusion (MSIE) module to locate the position of the infected areas. Second, the hybrid attention global context awareness (HAGCA) module collected hybrid attention information and multi-scale features from different branches, which can better aggregate information between the encoder and the decoder. Experimental results revealed that the proposed method achieves an accuracy of 0. 9982 and a specificity of 0. 9989. Numerical examples and ablation studies demonstrate that the proposed DIEP-Unet achieves high accuracy and outperforms existing segmentation methods. Therefore, the proposed method has the potential applications in the detection, localization, and labeling of other diseased areas.

AAAI Conference 2025 Conference Paper

Feature Denoising Diffusion Model for Blind Image Quality Assessment

  • Xudong Li
  • Yan Zhang
  • Yunhang Shen
  • Ke Li
  • Runze Hu
  • Xiawu Zheng
  • Sicheng Zhao

Blind Image Quality Assessment (BIQA) aims to evaluate image quality in line with human perception, without reference benchmarks. Currently, deep learning BIQA methods typically depend on using features from high-level tasks for transfer learning. However, the inherent differences between BIQA and these high-level tasks inevitably introduce noise into the quality-aware features. In this paper, we take an initial step toward exploring the diffusion model for feature denoising in BIQA, namely Perceptual Feature Diffusion for IQA (PFD-IQA), which aims to remove noise from quality-aware features. Specifically, 1) we propose a Perceptual Prior Discovery and Aggregation module to establish two auxiliary tasks to discover potential low-level features in images that are used to aggregate perceptual textual prompt conditions for the diffusion model. 2) we propose a Perceptual Conditional Feature Refinement strategy, which matches noisy features to predefined denoising trajectories and then performs exact feature denoising based on textual prompt conditions. By incorporating a lightweight denoiser and requiring only a few feature denoising steps (e.g., just five iterations), our PFD-IQA framework achieves superior performance across eight standard BIQA datasets, validating its effectiveness.

NeurIPS Conference 2025 Conference Paper

Joint Modeling of fMRI and EEG Imaging Using Ordinary Differential Equation-Based Hypergraph Neural Networks

  • Yan Zhang
  • Yang Gao
  • Min Li

Fusing multimodal brain imaging has been a hot topic since different modalities of brain imaging can provide complementary information. However, due to the size of simultaneous recorded fMRI-EEG dataset being limited and the substantial discrepancy between hemodynamic responses of fMRI and neural oscillations of EEG, the joint modeling of fMRI and EEG images is a rarely explored area and has not yielded satisfactory results. Existing studies have also indicated that the relationships between region of interest (ROI) are not one-to-one when synchronizing fMRI and EEG. Current graph-based multimodal modeling methods overlook those information. Based on this, we propose a hypergraph based fMRI-EEG modeling framework for asynchronous fMRI-EEG data named FE-NET. To the best of our knowledge, this is the first attempt to jointly model asynchronous EEG and fMRI data as Neural ODEs based hypergraph. Extensive experiments have demonstrated that the proposed FE-NET outperforms many state-of-the-art brain imaging modeling methods. Meanwhile, compared to simultaneously recorded fMRI-EEG data, asynchronously acquired fMRI-EEG data is less costly, which demonstrates the practical applicability of our method.

AAAI Conference 2025 Conference Paper

Learning Concept Prerequisite Relation via Global Knowledge Relation Optimization

  • Miao Zhang
  • Jiawei Wang
  • Kui Xiao
  • Shihui Wang
  • Yan Zhang
  • Hao Chen
  • Zhifei Li

Learning concept prerequisite relations helps better master and build a logically coherent knowledge structure. Many studies use graph neural networks to create heterogeneous knowledge networks that enhance concept representations. However, different types of relations in these networks can influence each other. Existing research often focuses solely on concept relations, neglecting other types of knowledge connections. To address this issue, this paper proposes a novel concept prerequisite relation learning model, named the Global Knowledge Relation Optimization Model(GKROM). Specifically, we capture the impact of different knowledge relation types on document and concept semantic representations separately, integrating the document and concept semantic representations. Then, we introduce multi-objective learning to optimize the knowledge relation network from a global perspective. Through the above optimization, GKROM learns richer semantic representations for concepts and documents, improving the accuracy of concept prerequisite relation learning. Extensive experiments on public datasets demonstrate the effectiveness of our GKROM, achieving state-of-the-art performance in concept prerequisite relation learning.

NeurIPS Conference 2025 Conference Paper

LTD-Bench: Evaluating Large Language Models by Letting Them Draw

  • Liuhao Lin
  • Ke Li
  • Zihan Xu
  • Yuchen Shi
  • Yulei Qin
  • Yan Zhang
  • Xing Sun
  • Rongrong Ji

Current evaluation paradigms for large language models (LLMs) represent a critical blind spot in AI research—relying on opaque numerical metrics that conceal fundamental limitations in spatial reasoning while providing no intuitive understanding of model capabilities. This deficiency creates a dangerous disconnect between reported performance and practical abilities, particularly for applications requiring physical world understanding. We introduce LTD-Bench, a breakthrough benchmark that transforms LLM evaluation from abstract scores to directly observable visual outputs by requiring models to generate drawings through dot matrices or executable code. This approach makes spatial reasoning limitations immediately apparent even to non-experts, bridging the fundamental gap between statistical performance and intuitive assessment. LTD-Bench implements a comprehensive methodology with complementary generation tasks (testing spatial imagination) and recognition tasks (assessing spatial perception) across three progressively challenging difficulty levels, methodically evaluating both directions of the critical language-spatial mapping. Our extensive experiments with state-of-the-art models expose an alarming capability gap: even LLMs achieving impressive results on traditional benchmarks demonstrate profound deficiencies in establishing bidirectional mappings between language and spatial concepts—a fundamental limitation that undermines their potential as genuine world models. Furthermore, LTD-Bench's visual outputs enable powerful diagnostic analysis, offering a potential approach to investigate model similarity. Our dataset and codes are available at https: //github. com/walktaster/LTD-Bench.

IJCAI Conference 2025 Conference Paper

MAGE: Multimodal Alignment and Generation Enhancement via Bridging Visual and Semantic Spaces

  • Shaojun E
  • Yuchen Yang
  • Jiaheng Wu
  • Yan Zhang
  • Tiejun Zhao
  • Ziyan Chen

In the latest advancements in multimodal learning, effectively addressing the spatial and semantic losses of visual data after encoding remains a critical challenge. This is because the performance of large multimodal models is positively correlated with the coupling between visual encoders and large language models. Existing approaches often face issues such as vector gaps or semantic disparities, resulting in information loss during the propagation process. To address these issues, we propose MAGE (Multimodal Alignment and Generation Enhancement), a novel framework that bridges the semantic spaces of vision and text through an innovative alignment mechanism. By introducing the Intelligent Alignment Network (IAN), MAGE achieves dimensional and semantic alignment. To reduce the gap between synonymous heterogeneous data, we employ a training strategy that combines cross-entropy and mean squared error, significantly enhancing the alignment effect. Moreover, to enhance MAGE’s “Any-to-Any” capability, we developed a fine-tuning dataset for multimodal tool-calling instructions to expand the model’s output capability boundaries. Finally, our proposed multimodal large model architecture, MAGE, achieved significantly better performance compared to similar works across various evaluation benchmarks, including MME, MMBench, and SEED. Complete code and appendix are available at: https: //github. com/GTCOM-NLP/MAGE

NeurIPS Conference 2025 Conference Paper

MLEP: Multi-granularity Local Entropy Patterns for Generalized AI-generated Image Detection

  • Lin Yuan
  • Xiaowan Li
  • Yan Zhang
  • Jiawei Zhang
  • Hongbo Li
  • Xinbo Gao

Advances in image generation technologies have raised growing concerns about their potential misuse, particularly in producing misinformation and deepfakes. This creates an urgent demand for effective methods to detect AI-generated images (AIGIs). While progress has been made, achieving reliable performance across diverse generative models and scenarios remains challenging due to the absence of source-invariant features and the limited generalization of existing approaches. In this study, we investigate the potential of using image entropy as a discriminative cue for AIGI detection and propose Multi-granularity Local Entropy Patterns (MLEP), a set of feature maps computed based on Shannon entropy from shuffled small patches at multiple image scales. MLEP effectively captures pixel dependencies across scales and dimensions while disrupting semantic content, thereby reducing potential content bias. Based on MLEP, we can easily build a robust CNN-based classifier capable of detecting AIGIs with enhanced reliability. Extensive experiments in an open-world setting, involving images synthesized by 32 distinct generative models, demonstrate that our approach achieves substantial improvements over state-of-the-art methods in both accuracy and generalization. Our code and models are available at https: //www. github. com/fkeufss/MLEP/.

IJCAI Conference 2025 Conference Paper

Modality-Fair Preference Optimization for Trustworthy MLLM Alignment

  • Songtao Jiang
  • Yan Zhang
  • Ruizhe Chen
  • Tianxiang Hu
  • Yeying Jin
  • Qinglin He
  • Yang Feng
  • Jian Wu

Multimodal large language models (MLLMs) have achieved remarkable success across various tasks. However, separate training of visual and textual encoders often results in a misalignment of the modality. Such misalignment may lead models to generate content that is absent from the input image, a phenomenon referred to as hallucination. These inaccuracies severely undermine the trustworthiness of MLLMs in real-world applications. Despite attempts to optimize text preferences to mitigate this issue, our initial investigation indicates that the trustworthiness of MLLMs remains inadequate. Specifically, these models tend to provide preferred answers even when the input image is heavily distorted. Analysis of visual token attention also indicates that the model focuses primarily on the surrounding context rather than the key object referenced in the question. These findings highlight a misalignment between the modalities, where answers inadequately leverage input images. Motivated by our findings, we propose Modality-Fair Preference Optimization (MFPO), which comprises three components: the construction of a multimodal preference dataset in which dispreferred images differ from originals solely in key regions; an image reward loss function encouraging the model to generate answers better aligned with the input images; and an easy-to-hard iterative alignment strategy to stabilize joint modality training. Extensive experiments on three trustworthiness benchmarks demonstrate that MFPO significantly enhances the trustworthiness of MLLMs. In particular, it enables the 7B models to attain trustworthiness levels on par with, or even surpass, those of the 13B, 34B, and larger models.

EAAI Journal 2025 Journal Article

MSPFNet: Multi-scale perceptual focusing network for scrap steel segmentation and classification

  • Yunfeng Xu
  • Changda Liu
  • Jiakui Zhong
  • Wei Mei
  • Yan Zhang

Scrap steel recycling is one of the important measures to reduce carbon emissions and promote sustainable development. However, due to scrap steel data has the characteristics of a large number and variety of targets in the same image, disorderly arrangement, varying sizes, and mutual backgrounds, traditional manual rating methods are inefficient and highly influenced by subjective factors. Moreover, many efficient vision transformers methods currently face challenges in scrap steel segmentation and classification. To address these problems, we propose multi-scale perceptual focusing network for scrap steel segmentation and classification (MSPFNet). MSPFNet utilizes multi-scale feature extraction and perception mechanisms to focus on the key regions of scrap steel images, thereby improving the accuracy of scrap steel segmentation and classification. We apply MSPFNet to both public datasets (ImageNet-1K, ADE20K and Cityscapes) and scrap steel dataset in real-world application scenarios, experimental results show that MSPFNet achieves better performance than other existing methods with the similar parameter magnitude. MSPFNet-B (The base model of MSPFNet, MSPFNet is divided into three levels according to model size: tiny, small, and base) achieves mIoU (mean Intersection over Union) of 69. 8% on the scrap steel dataset, surpassing other state-of-the-art segmentation models with the similar parameter magnitude, such as Segformer-B3 (The B3 model of Segformer, Segformer is divided into six levels according to model size: B0, B1, B2, B3, B4 and B5) and SegNeXt-B (The base model of SegNeXt) by 2. 1% and 4. 8%, respectively. Furthermore, extensive experiments were conducted on public datasets (ImageNet-1K, ADE20K and Cityscapes), showing that MSPFNet outperforms other existing methods with the similar parameter magnitude. We tested MSPFNet in a real-life scenario at a steel plant, and the model accurately segmented different levels of scrap steel. This research provides a new solution for the scrap steel recycling and processing industry, with the potential to improve the efficiency and quality of scrap steel recycling and achieve automation in the process.

ICLR Conference 2025 Conference Paper

ProtPainter: Draw or Drag Protein via Topology-guided Diffusion

  • Zhengxi Lu
  • Shizhuo Cheng
  • Tintin Jiang
  • Yan Zhang
  • Min Zhang 0069

Recent advances in protein backbone generation have achieved promising results under structural, functional, or physical constraints. However, existing methods lack the flexibility for precise topology control, limiting navigation of the backbone space. We present $\textbf{ProtPainter}$, a diffusion-based approach for generating protein backbones conditioned on 3D curves. ProtPainter follows a two-stage process: curve-based sketching and sketch-guided backbone generation. For the first stage, we propose $\textbf{CurveEncoder}$, which predicts secondary structure annotations from a curve to parametrize sketch generation. For the second stage, the sketch guides the generative process in Denoising Diffusion Probabilistic Modeling (DDPM) to generate backbones. During the process, we further introduce a fusion scheduling scheme, Helix-Gating, to control the scaling factors. To evaluate, we propose the first benchmark for topology-conditioned protein generation, introducing Protein Restoration Task and a new metric, self-consistency Topology Fitness (scTF). Experiments demonstrate ProtPainter's ability to generate topology-fit (scTF $>$ 0.8) and designable (scTM $>$ 0.5) backbones, with drawing and dragging tasks showcasing its flexibility and versatility.

IJCAI Conference 2025 Conference Paper

SCOUT: Semi-supervised Camouflaged Object Detection by Utilizing Text and Adaptive Data Selection

  • Weiqi Yan
  • Lvhai Chen
  • Shengchuan Zhang
  • Yan Zhang
  • Liujuan Cao

The difficulty of pixel-level annotation has significantly hindered the development of the Camouflaged Object Detection (COD) field. To save on annotation costs, previous works leverage the semi-supervised COD framework that relies on a small number of labeled data and a large volume of unlabeled data. We argue that there is still significant room for improvement in the effective utilization of unlabeled data. To this end, we introduce a Semi-supervised Camouflaged Object Detection by Utilizing Text and Adaptive Data Selection (SCOUT). It includes an Adaptive Data Augment and Selection (ADAS) module and a Text Fusion Module (TFM). The ADSA module selects valuable data for annotation through an adversarial augment and sampling strategy. The TFM module further leverages the selected valuable data by combining camouflage-related knowledge and text-visual interaction. To adapt to this work, we build a new dataset, namely RefTextCOD. Extensive experiments show that the proposed method surpasses previous semi-supervised methods in the COD field and achieves state-of-the-art performance. Our code will be released at https: //github. com/Heartfirey/UCOD-DPL.

JBHI Journal 2025 Journal Article

Synergistic Drug Combination Prediction via Dual-Level Feature Aggregation and Knowledge Graph-Based Deep Neural Network

  • Ying Zuo
  • Yan Zhang
  • Li Wang
  • Jianping Yu
  • Jiawei Luo
  • Qiu Xiao

Identifying synergistic drug combinations is a critical but difficult challenge in cancer treatment, owing to the sheer complexity and enormous number of possible drug combinations. However, most existing computational methods rely on a single data perspective and often overlooking the complexity of interactions between different biological entities. Furthermore, they fail to fully integrate the intrinsic properties of drugs and cell lines with the broader biological relationships that play a crucial role in drug synergy. To address these challenges, we propose a novel framework called LGSyn that integrates two types of information: local features, including molecular fingerprints, descriptors, and gene expression profiles, as well as global features that encompass broader biological interactions, including drug-protein, protein-cell line, protein-protein, and cell line-tissue interactions. By combining these two types of features, LGSyn leverages the full spectrum of biological knowledge to predict drug synergy. In LGSyn, we developed three fusion strategies to effectively integrate local and global information and identify the most suitable strategy. The resulting fused feature vectors are then fed into a deep neural network for training and synergy prediction. Experimental results demonstrate that the proposed method outperforms current state-of-the-art models, achieving superior accuracy and stability in drug synergy prediction.

AAAI Conference 2025 Conference Paper

Towards Macro-AUC Oriented Imbalanced Multi-Label Continual Learning

  • Yan Zhang
  • Guoqiang Wu
  • Bingzheng Wang
  • Teng Pang
  • Haoliang Sun
  • Yilong Yin

In Continual Learning (CL), while existing work primarily focuses on the multi-class classification task, there has been limited research on Multi-Label Learning (MLL). In practice, MLL datasets are often class-imbalanced, making it inherently challenging, a problem that is even more acute in CL. Due to its sensitivity to imbalance, Macro-AUC is an appropriate and widely used measure in MLL. However, there is no research to optimize Macro-AUC in MLCL specifically. To fill this gap, in this paper, we propose a new memory replay-based method to tackle the imbalance issue for Macro-AUC-oriented MLCL. Specifically, inspired by recent theory work, we propose a new Reweighted Label-Distribution-Aware Margin (RLDAM) loss. Furthermore, to be compatible with the RLDAM loss, a new memory-updating strategy named Weight Retain Updating (WRU) is proposed to maintain the numbers of positive and negative instances of the original dataset in memory. Theoretically, we provide superior generalization analyses of the RLDAM-based algorithm in terms of Macro-AUC, separately in batch MLL and MLCL settings. This is the first work to offer theoretical generalization analyses in MLCL to our knowledge. Finally, a series of experimental results illustrate the effectiveness of our method over several baselines.

IJCAI Conference 2025 Conference Paper

Towards Micro-Action Recognition with Limited Annotations: An Asynchronous Pseudo Labeling and Training Approach

  • Yan Zhang
  • Lechao Cheng
  • Yaxiong Wang
  • Zhun Zhong
  • Meng Wang

Micro-Action Recognition (MAR) aims to classify subtle human actions in video. However, annotating MAR datasets is particularly challenging due to the subtlety of actions. To this end, we introduce the setting of Semi-Supervised MAR (SSMAR), where only a part of samples are labeled. We first evaluate traditional Semi-Supervised Learning (SSL) methods to SSMAR and find that these methods tend to overfit on inaccurate pseudo-labels, leading to error accumulation and degraded performance. This issue primarily arises from the common practice of directly using the predictions of classifier as pseudo-labels to train the model. To solve this issue, we propose a novel framework, called Asynchronous Pseudo Labeling and Training (APLT), which explicitly separates the pseudo-labeling process from model training. Specifically, we introduce a semi-supervised clustering method during the offline pseudo-labeling phase to generate more accurate pseudo-labels. Moreover, a self-adaptive thresholding strategy is proposed to dynamically filter noisy labels of different classes. We then build a memory-based prototype classifier based on the filtered pseudo-labels, which is fixed and used to guide the subsequent model training phase. By alternating the two pseudo-labeling and model training phases in an asynchronous manner, the model can not only be learned with more accurate pseudo-labels but also avoid the overfitting issue. Experiments on three MAR datasets show that our APLT largely outperforms state-of-the-art SSL methods. For instance, APLT improves accuracy by 14. 5% over FixMatch on the MA-12 dataset when using only 50% labeled data. Code is available at https: //github. com/zy-hfut/APLT

AAAI Conference 2025 Conference Paper

Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal Clues

  • Yan Zhang
  • Gangyan Zeng
  • Huawen Shen
  • Daiqing Wu
  • Yu Zhou
  • Can Ma

Video text-based visual question answering (TextVQA) is a practical task that aims to answer questions by jointly reasoning textual and visual information in a given video. Inspired by the development of TextVQA in image domain, existing Video TextVQA approaches leverage a language model (e.g. T5) to process text-rich multiple frames and generate answers auto-regressively. Nevertheless, the spatio-temporal relationships among visual entities (including scene text and objects) will be disrupted and models are susceptible to interference from unrelated information, resulting in irrational reasoning and inaccurate answering. To tackle these challenges, we propose the TEA (stands for "Track the Answer'') method that better extends the generative TextVQA framework from image to video. TEA recovers the spatio-temporal relationships in a complementary way and incorporates OCR-aware clues to enhance the quality of reasoning questions. Extensive experiments on several public Video TextVQA datasets validate the effectiveness and generalization of our framework. TEA outperforms existing TextVQA methods, video-language pretraining methods and video large language models by great margins. The code will be publicly released.

EAAI Journal 2025 Journal Article

Transfer learning framework integrating attention mechanism and domain adaptation for Low Earth Orbit satellite network traffic prediction

  • Yan Zhang
  • Yong Wang
  • Qingsong Zhao
  • Yadi Zhai
  • Zhi Lin
  • Luda Zhao
  • Yihua Hu

Traffic prediction is a crucial prerequisite for planning and even network security in Low Earth Orbit (LEO) satellite networks (LSNs). This paper designed a transfer learning framework for LSN traffic prediction that leverages attention mechanism and domain adaptation. Firstly, by integrating five-dimensional data including global population distribution, local time coefficient, the Internet penetration rate of the country, daily data volume of a single Internet user, and global aeronautical traffic demand, a traffic model that could characterize the traffic situation of the area covered by LEO satellite within a specific time range was constructed. Considering the problem of insufficient online traffic data, knowledge was transferred from terrestrial network traffic (source domain) to satellite network traffic (target domain) by incorporating the Domain-Adversarial Neural Network (DANN) method to tackle the data distribution discrepancies between the source and target domains. Finally, by combining DANN with the attention mechanism, the domain-invariant features of the source domain and the target domain were extracted to predict satellite network traffic accurately. Experimental results show that compared to baseline models, the error of the proposed framework in terms of root mean square error measurement is reduced by 9. 57% to 33. 47% and 18. 85% to 38. 99% in the two simulated LSN traffic scenarios. Moreover, this framework has low computational complexity among other transfer learning models, which can lay a foundation for subsequent satellite traffic planning and network security.

NeurIPS Conference 2025 Conference Paper

When Semantics Mislead Vision: Mitigating Large Multimodal Models Hallucinations in Scene Text Spotting and Understanding

  • Yan Shu
  • Hangui Lin
  • Yexin Liu
  • Yan Zhang
  • Gangyan Zeng
  • Yan Li
  • Yu Zhou
  • Ser Nam Lim

Large Multimodal Models (LMMs) have achieved impressive progress in visual perception and reasoning. However, when confronted with visually ambiguous or non-semantic scene text, they often struggle to accurately spot and understand the content, frequently generating semantically plausible yet visually incorrect answers, which we refer to as semantic hallucination. In this work, we investigate the underlying causes of semantic hallucination and identify a key finding: Transformer layers in LLM with stronger attention focus on scene text regions are less prone to producing semantic hallucinations. Thus, we propose a training-free semantic hallucination mitigation framework comprising two key components: (1) ZoomText, a coarse-to-fine strategy that identifies potential text regions without external detectors; and (2) Grounded Layer Correction, which adaptively leverages the internal representations from layers less prone to hallucination to guide decoding, correcting hallucinated outputs for non-semantic samples while preserving the semantics of meaningful ones. To enable rigorous evaluation, we introduce TextHalu-Bench, a benchmark of 1, 740 samples spanning both semantic and non-semantic cases, with manually curated question–answer pairs designed to probe model hallucinations. Extensive experiments demonstrate that our method not only effectively mitigates semantic hallucination but also achieves strong performance on public benchmarks for scene text spotting and understanding.

NeurIPS Conference 2025 Conference Paper

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs

  • Xudong Li
  • Mengdan Zhang
  • Peixian Chen
  • Xiawu Zheng
  • Yan Zhang
  • Jingyuan Zheng
  • Yunhang Shen
  • Ke Li

Multi-modal Large Language Models (MLLMs) excel at single-image tasks but struggle with multi-image understanding due to cross-modal misalignment, leading to hallucinations (context omission, conflation, and misinterpretation). Existing methods using Direct Preference Optimization (DPO) constrain optimization to a solitary image reference within the input sequence, neglecting holistic context modeling. To address this, we propose Context-to-Cue Direct Preference Optimization (CcDPO), a multi-level preference optimization framework that enhances per-image perception in multi-image settings by zooming into visual clues—from sequential context to local details. Our approach features two sequentially dependent components: (i) Context-Level Optimization: By introducing low-cost sequence preference pairs, we optimize the model to distinguish between complete and disrupted multi-image contexts, thereby correcting cognitive biases in MLLMs’ multi-image understanding. (ii) Needle-Level Optimization: By integrating region-specific visual prompts with multimodal preference supervision, we direct the model’s attention to critical visual details, effectively suppressing perceptual biases toward fine-grained visual information. To support scalable optimization, we also construct MultiScope-42k, an automatically generated multi-image dataset with hierarchical preference pairs. Experiments show that CcDPO significantly reduces hallucinations and yields consistent performance gains across general single- and multi-image tasks. Codes are available at https: //github. com/LXDxmu/CcDPO.

EAAI Journal 2024 Journal Article

A three-stage pavement image crack detection framework with positive sample augmentation

  • Qingsong Song
  • Liming Liu
  • Na Lu
  • Yan Zhang
  • Ravie Chandren Muniyandi
  • Yisheng An

Most pavement crack detection methods based on deep learning rely too much on pixel-wise labels, and are facing with sample imbalance problem. This paper proposes a three-stage pavement crack detection framework with positive sample augmentation by an autoencoder-deep convolutional generative adversarial network (E-DCGAN). Positive samples are the images labeled as crack in the training datasets. Firstly, crack regions are located using graph based on visual saliency (GBVS). Secondly, a convolutional neural network (CNN) model is designed to recognize crack pixels in the located crack regions. The E-DCGAN is explored to generate crack images, which are then added as augmented positive samples in the training dataset of the CNN model. The trained CNN model takes the located crack regions as inputs, outputs the recognized crack pixels. Finally, fine crack detection is realized through region growing. Crack contours are completely depicted pixel-wise based on the recognized crack pixels and the Breadth-first search. Experiments are carried out on two benchmark datasets and in practical applications. The results showed that the framework could detect cracks at the pixel level. Its performance is improved compared with standard fully convolutional network (FCN)-based and U-shaped CNN (U-Net)-based crack detection methods. In addition, the standard FCN-based and U-Net-based methods require accurate pixel-wise annotations, while the proposed framework only requires image-wise annotations. It can be concluded that the proposed framework not only effectively avoids the dependence on pixel-level annotations, improves the sample imbalance issues, but also achieves competitive detection performance over the FCN and U-Net based methods.

AAAI Conference 2024 Conference Paper

DiffAIL: Diffusion Adversarial Imitation Learning

  • Bingzheng Wang
  • Guoqiang Wu
  • Teng Pang
  • Yan Zhang
  • Yilong Yin

Imitation learning aims to solve the problem of defining reward functions in real-world decision-making tasks. The current popular approach is the Adversarial Imitation Learning (AIL) framework, which matches expert state-action occupancy measures to obtain a surrogate reward for forward reinforcement learning. However, the traditional discriminator is a simple binary classifier and doesn't learn an accurate distribution, which may result in failing to identify expert-level state-action pairs induced by the policy interacting with the environment. To address this issue, we propose a method named diffusion adversarial imitation learning (DiffAIL), which introduces the diffusion model into the AIL framework. Specifically, DiffAIL models the state-action pairs as unconditional diffusion models and uses diffusion loss as part of the discriminator's learning objective, which enables the discriminator to capture better expert demonstrations and improve generalization. Experimentally, the results show that our method achieves state-of-the-art performance and significantly surpasses expert demonstration on two benchmark tasks, including the standard state-action setting and state-only settings.

JBHI Journal 2024 Journal Article

Dual-Teacher Feature Distillation: A Transfer Learning Method for Insomniac PSG Staging

  • Lijuan Duan
  • Yan Zhang
  • Zhaoyang Huang
  • Bian Ma
  • Wenjian Wang
  • Yuanhua Qiao

Insomnia is the most common sleep disorder linked with adverse long-term medical and psychiatric outcomes. Automatic sleep staging plays a crucial role in aiding doctors to diagnose insomnia disorder. Only a few studies have been conducted to develop automatic sleep staging methods for insomniacs, and most of them have utilized transfer learning methods, which involve pre-training models on healthy individuals and then fine-tuning them on insomniacs. Unfortunately, significant differences in feature distribution between the two subject groups impede the transfer performance, highlighting the need to effectively integrate the features of healthy subjects and insomniacs. In this paper, we propose a dual-teacher cross-domain knowledge transfer method based on the feature-based knowledge distillation to improve the performance of sleep staging for insomniacs. Specifically, the insomnia teacher directly learns from insomniacs and feeds the corresponding domain-specific features into the student network, while the health domain teacher guide the student network to learn domain-generic features. During the training process, we adopt the OFD (Overhaul of Feature Distillation) method to build the health domain teacher. We conducted the experiments to validate the proposed method, using the Sleep-EDF database as the source domain and the CAP-Database as the target domain. The results demonstrate that our method surpasses advanced techniques, achieving an average sleep staging accuracy of 80. 56% on the CAP-Database. Furthermore, our method exhibits promising performance on the private dataset.

ICLR Conference 2024 Conference Paper

Graph Neural Networks for Learning Equivariant Representations of Neural Networks

  • Miltiadis Kofinas
  • Boris Knyazev 0001
  • Yan Zhang
  • Yunlu Chen
  • Gertjan J. Burghouts
  • Efstratios Gavves
  • Cees G. M. Snoek
  • David W. Zhang

Neural networks that process the parameters of other neural networks find applications in domains as diverse as classifying implicit neural representations, generating neural network weights, and predicting generalization errors. However, existing approaches either overlook the inherent permutation symmetry in the neural network or rely on intricate weight-sharing patterns to achieve equivariance, while ignoring the impact of the network architecture itself. In this work, we propose to represent neural networks as computational graphs of parameters, which allows us to harness powerful graph neural networks and transformers that preserve permutation symmetry. Consequently, our approach enables a single model to encode neural computational graphs with diverse architectures. We showcase the effectiveness of our method on a wide range of tasks, including classification and editing of implicit neural representations, predicting generalization performance, and learning to optimize, while consistently outperforming state-of-the-art methods. The source code is open-sourced at https://github.com/mkofinas/neural-graphs.

ICML Conference 2024 Conference Paper

Improved Generalization of Weight Space Networks via Augmentations

  • Aviv Shamsian
  • Aviv Navon
  • David W. Zhang
  • Yan Zhang
  • Ethan Fetaya
  • Gal Chechik
  • Haggai Maron

Learning in deep weight spaces (DWS), where neural networks process the weights of other neural networks, is an emerging research direction, with applications to 2D and 3D neural fields (INRs, NeRFs), as well as making inferences about other types of neural networks. Unfortunately, weight space models tend to suffer from substantial overfitting. We empirically analyze the reasons for this overfitting and find that a key reason is the lack of diversity in DWS datasets. While a given object can be represented by many different weight configurations, typical INR training sets fail to capture variability across INRs that represent the same object. To address this, we explore strategies for data augmentation in weight spaces and propose a MixUp method adapted for weight spaces. We demonstrate the effectiveness of these methods in two setups. In classification, they improve performance similarly to having up to 10 times more data. In self-supervised contrastive learning, they yield substantial 5-10% gains in downstream classification.

TMLR Journal 2024 Journal Article

PopulAtion Parameter Averaging (PAPA)

  • Alexia Jolicoeur-Martineau
  • Emy Gervais
  • Kilian Fatras
  • Yan Zhang
  • Simon Lacoste-Julien

Ensemble methods combine the predictions of multiple models to improve performance, but they require significantly higher computation costs at inference time. To avoid these costs, multiple neural networks can be combined into one by averaging their weights. However, this usually performs significantly worse than ensembling. Weight averaging is only beneficial when different enough to benefit from combining them, but similar enough to average well. Based on this idea, we propose PopulAtion Parameter Averaging (PAPA): a method that combines the generality of ensembling with the efficiency of weight averaging. PAPA leverages a population of diverse models (trained on different data orders, augmentations, and regularizations) while slowly pushing the weights of the networks toward the population average of the weights. We also propose PAPA variants (PAPA-all, and PAPA-2) that average weights rarely rather than continuously; all methods increase generalization, but PAPA tends to perform best. PAPA reduces the performance gap between averaging and ensembling, increasing the average accuracy of a population of models by up to 0.8% on CIFAR-10, 1.9% on CIFAR-100, and 1.6% on ImageNet when compared to training independent (non-averaged) models.

TIST Journal 2024 Journal Article

Proposal Semantic Relationship Graph Network for Temporal Action Detection

  • Shaowen Su
  • Yan Zhang
  • Minggang Gan

Temporal action detection, a critical task in video activity understanding, is typically divided into two stages: proposal generation and classification. However, most existing methods overlook the importance of information transfer among proposals during classification, often treating each proposal in isolation, which hampers accurate label prediction. In this article, we propose a novel method for inferring semantic relationships both within and between action proposals, guiding the fusion of action proposal features accordingly. Building on this approach, we introduce the Proposal Semantic Relationship Graph Network (PSRGN), an end-to-end model that leverages intra-proposal semantic relationship graphs to extract cross-scale temporal context and an inter-proposal semantic relationship graph to incorporate complementary neighboring information, significantly improving proposal feature quality and overall detection performance. This is the first method to apply graph structure learning in temporal action detection, adaptively constructing the inter-proposal semantic graph. Extensive experiments on two datasets demonstrate the effectiveness of our approach, achieving state-of-the-art (SOTA). Code and results are available at http://github.com/Riiick2011/PSRGN.

ICRA Conference 2024 Conference Paper

Representing Robot Geometry as Distance Fields: Applications to Whole-body Manipulation

  • Yiming Li
  • Yan Zhang
  • Amirreza Razmjoo
  • Sylvain Calinon

In this work, we propose a novel approach to represent robot geometry as distance fields (RDF) that extends the principle of signed distance fields (SDFs) to articulated kinematic chains. Our method employs a combination of Bernstein polynomials to encode the signed distance for each robot link with high accuracy and efficiency while ensuring the mathematical continuity and differentiability of SDFs. We further leverage the kinematics chain of the robot to produce the SDF representation in joint space, allowing robust distance queries in arbitrary joint configurations. The proposed RDF representation is differentiable and smooth in both task and joint spaces, enabling its direct integration to optimization problems. Additionally, the 0-level set of the robot corresponds to the robot surface, which can be seamlessly integrated into whole-body manipulation tasks. We conduct various experiments in both simulations and with 7-axis Franka Emika robots, comparing against baseline methods, and demonstrating its effectiveness in collision avoidance and whole-body manipulation tasks. Project page: https://sites.google.com/view/lrdf/home

YNICL Journal 2024 Journal Article

Right superior frontal gyrus: A potential neuroimaging biomarker for predicting short-term efficacy in schizophrenia

  • Yongfeng Yang
  • Xueyan Jin
  • Yongjiang Xue
  • Xue Li
  • Yi Chen
  • Ning Kang
  • Wei Yan
  • Peng Li

Antipsychotic drug treatment for schizophrenia (SZ) can alter brain structure and function, but it is unclear if specific regional changes are associated with treatment outcome. Therefore, we examined the effects of antipsychotic drug treatment on regional grey matter (GM) density, white matter (WM) density, and functional connectivity (FC) as well as associations between regional changes and treatment efficacy. SZ patients (n = 163) and health controls (HCs) (n = 131) were examined by structural magnetic resonance imaging (sMRI) at baseline, and a subset of SZ patients (n = 77) were re-examined after 8 weeks of second-generation antipsychotic treatment to assess changes in regional GM and WM density. In addition, 88 SZ patients and 81 HCs were examined by resting-state functional MRI (rs-fMRI) at baseline and the patients were re-examined post-treatment to examine FC changes. The Positive and Negative Syndrome Scale (PANSS) and MATRICS Consensus Cognitive Battery (MCCB) were applied to measure psychiatric symptoms and cognitive impairments in SZ. SZ patients were then stratified into response and non-response groups according to PANSS score change (≥50 % decrease or <50 % decrease, respectively). The GM density of the right cingulate gyrus, WM density of the right superior frontal gyrus (SFG) plus 5 other WM tracts were reduced in the response group compared to the non-response group. The FC values between the right anterior cingulate and paracingulate gyrus and left thalamus were reduced in the entire SZ group (n = 88) after treatment, while FC between the right inferior temporal gyrus (ITG) and right medial superior frontal gyrus (SFGmed) was increased in the response group. There were no significant changes in regional FC among the non-response group after treatment and no correlations with symptom or cognition test scores. These findings suggest that the right SFG is a critical target of antipsychotic drugs and that WM density and FC alterations within this region could be used as potential indicators in predicting the treatment outcome of antipsychotics of SZ.

NeurIPS Conference 2024 Conference Paper

RLE: A Unified Perspective of Data Augmentation for Cross-Spectral Re-Identification

  • Tan Lei
  • Yukang Zhang
  • Keke Han
  • Pingyang Dai
  • Yan Zhang
  • Yongjian Wu
  • Rongrong Ji

This paper makes a step towards modeling the modality discrepancy in the cross-spectral re-identification task. Based on the Lambertain model, we observe that the non-linear modality discrepancy mainly comes from diverse linear transformations acting on the surface of different materials. From this view, we unify all data augmentation strategies for cross-spectral re-identification as mimicking such local linear transformations and categorize them into moderate transformation and radical transformation. By extending the observation, we propose a Random Linear Enhancement (RLE) strategy which includes Moderate Random Linear Enhancement (MRLE) and Radical Random Linear Enhancement (RRLE) to push the boundaries of both types of transformation. Moderate Random Linear Enhancement is designed to provide diverse image transformations that satisfy the original linear correlations under constrained conditions, whereas Radical Random Linear Enhancement seeks to generate local linear transformations directly without relying on external information. The experimental results not only demonstrate the superiority and effectiveness of RLE but also confirm its great potential as a general-purpose data augmentation for cross-spectral re-identification.

AAAI Conference 2024 Conference Paper

Semi-Supervised Blind Image Quality Assessment through Knowledge Distillation and Incremental Learning

  • Wensheng Pan
  • Timin Gao
  • Yan Zhang
  • Xiawu Zheng
  • Yunhang Shen
  • Ke Li
  • Runze Hu
  • Yutao Liu

Blind Image Quality Assessment (BIQA) aims to simulate human assessment of image quality. It has a great demand for labeled data, which is often insufficient in practice. Some researchers employ unsupervised methods to address this issue, which is challenging to emulate the human subjective system. To this end, we introduce a unified framework that combines semi-supervised and incremental learning to address the mentioned issue. Specifically, when training data is limited, semi-supervised learning is necessary to infer extensive unlabeled data. To facilitate semi-supervised learning, we use knowledge distillation to assign pseudo-labels to unlabeled data, preserving analytical capability. To gradually improve the quality of pseudo labels, we introduce incremental learning. However, incremental learning can lead to catastrophic forgetting. We employ Experience Replay by selecting representative samples during multiple rounds of semi-supervised learning, to alleviate forgetting and ensure model stability. Experimental results show that the proposed approach achieves state-of-the-art performance across various benchmark datasets. After being trained on the LIVE dataset, our method can be directly transferred to the CSIQ dataset. Compared with other methods, it significantly outperforms unsupervised methods on the CSIQ dataset with a marginal performance drop (-0.002) on the LIVE dataset. In conclusion, our proposed method demonstrates its potential to tackle the challenges in real-world production processes.

JBHI Journal 2024 Journal Article

Semi-Supervised Disease Classification Based on Limited Medical Image Data

  • Yan Zhang
  • Chun Li
  • Zhaoxia Liu
  • Ming Li

Inrecent years, significant progress has been made in the field of learning from positive and unlabeled examples (PU learning), particularly in the context of advancing image and text classification tasks. However, applying PU learning to semi-supervised disease classification remains a formidable challenge, primarily due to the limited availability of labeled medical images. In the realm of medical image-aided diagnosis algorithms, numerous theoretical and practical obstacles persist. The research on PU learning for medical image-assisted diagnosis holds substantial importance, as it aims to reduce the time spent by professional experts in classifying images. Unlike natural images, medical images are typically accompanied by a scarcity of annotated data, while an abundance of unlabeled cases exists. Addressing these challenges, this paper introduces a novel generative model inspired by Hölder divergence, specifically designed for semi-supervised disease classification using positive and unlabeled medical image data. In this paper, we present a comprehensive formulation of the problem and establish its theoretical feasibility through rigorous mathematical analysis. To evaluate the effectiveness of our proposed approach, we conduct extensive experiments on five benchmark datasets commonly used in PU medical learning: BreastMNIST, PneumoniaMNIST, BloodMNIST, OCTMNIST, and AMD. The experimental results clearly demonstrate the superiority of our method over existing approaches based on KL divergence. Notably, our approach achieves state-of-the-art performance on all five disease classification benchmarks. By addressing the limitations imposed by limited labeled data and harnessing the untapped potential of unlabeled medical images, our novel generative model presents a promising direction for enhancing semi-supervised disease classification in the field of medical image analysis.

ICML Conference 2024 Conference Paper

Unsupervised Concept Discovery Mitigates Spurious Correlations

  • Md Rifat Arefin
  • Yan Zhang
  • Aristide Baratin
  • Francesco Locatello
  • Irina Rish
  • Dianbo Liu
  • Kenji Kawaguchi

Models prone to spurious correlations in training data often produce brittle predictions and introduce unintended biases. Addressing this challenge typically involves methods relying on prior knowledge and group annotation to remove spurious correlations, which may not be readily available in many applications. In this paper, we establish a novel connection between unsupervised object-centric learning and mitigation of spurious correlations. Instead of directly inferring subgroups with varying correlations with labels, our approach focuses on discovering concepts: discrete ideas that are shared across input samples. Leveraging existing object-centric representation learning, we introduce CoBalT: a concept balancing technique that effectively mitigates spurious correlations without requiring human labeling of subgroups. Evaluation across the benchmark datasets for sub-population shifts demonstrate superior or competitive performance compared state-of-the-art baselines, without the need for group annotation. Code is available at https: //github. com/rarefin/CoBalT

NeurIPS Conference 2023 Conference Paper

Cascading Bandits: Optimizing Recommendation Frequency in Delayed Feedback Environments

  • Dairui Wang
  • Junyu Cao
  • Yan Zhang
  • Wei Qi

Delayed feedback is a critical problem in dynamic recommender systems. In practice, the feedback result often depends on the frequency of recommendation. Most existing online learning literature fails to consider optimization of the recommendation frequency, and regards the reward from each successfully recommended message to be equal. In this paper, we consider a novel cascading bandits setting, where individual messages from a selected list are sent to a user periodically. Whenever a user does not like a message, she may abandon the system with a probability positively correlated with the recommendation frequency. A learning agent needs to learn both the underlying message attraction probabilities and users' abandonment probabilities through the randomly delayed feedback. We first show a dynamic programming solution to finding the optimal message sequence in deterministic scenarios, in which the reward is allowed to vary with different messages. Then we propose a polynomial time UCB-based offline learning algorithm, and discuss its performance by characterizing its regret bound. For the online setting, we propose a learning algorithm which allows adaptive content for a given user. Numerical experiment on AmEx dataset confirms the effectiveness of our algorithms.

EAAI Journal 2023 Journal Article

Classification and rating of steel scrap using deep learning

  • Wenguang Xu
  • Pengcheng Xiao
  • Liguang Zhu
  • Yan Zhang
  • Jinbao Chang
  • Rong Zhu
  • Yunfeng Xu

To address the issues of high human interference and low efficiency in traditional manual methods for classifying and rating steel scrap, we propose the development of CSBFNet, a deep learning-based model for multi-category steel scrap classification and rating. Firstly, we built a 1: 3 physical model of steel scrap quality inspection to simulate the unloading of a truck. We used a high-resolution vision sensor to capture the morphological characteristics of various steel scraps. Next, we trained the CSBFNet model using this data to obtain characteristic information for classifying and judging various types of scrap steel. Finally, we tested and improved the CSBFNet model at a Chinese steel mill. The results demonstrate that the model can effectively determine the automatic rating for different grades of scrap. The average accuracy rate of all types of steel scrap reaches 92. 4% for the full category, with an mAP of 90. 7%. Compared to traditional artificial quality detection methods, it has clear advantages in accuracy and fairness. This model solves the problem of evaluating the quality of steel scrap in the recycling process.

ICML Conference 2023 Conference Paper

CrossSplit: Mitigating Label Noise Memorization through Data Splitting

  • Jihye Kim
  • Aristide Baratin
  • Yan Zhang
  • Simon Lacoste-Julien

We approach the problem of improving robustness of deep learning algorithms in the presence of label noise. Building upon existing label correction and co-teaching methods, we propose a novel training procedure to mitigate the memorization of noisy labels, called CrossSplit, which uses a pair of neural networks trained on two disjoint parts of the labeled dataset. CrossSplit combines two main ingredients: (i) Cross-split label correction. The idea is that, since the model trained on one part of the data cannot memorize example-label pairs from the other part, the training labels presented to each network can be smoothly adjusted by using the predictions of its peer network; (ii) Cross-split semi-supervised training. A network trained on one part of the data also uses the unlabeled inputs of the other part. Extensive experiments on CIFAR-10, CIFAR-100, Tiny-ImageNet and mini-WebVision datasets demonstrate that our method can outperform the current state-of-the-art in a wide range of noise ratios. The project page is at https: //rlawlgul. github. io/.

AAAI Conference 2023 Conference Paper

Data-Efficient Image Quality Assessment with Attention-Panel Decoder

  • Guanyi Qin
  • Runze Hu
  • Yutao Liu
  • Xiawu Zheng
  • Haotian Liu
  • Xiu Li
  • Yan Zhang

Blind Image Quality Assessment (BIQA) is a fundamental task in computer vision, which however remains unresolved due to the complex distortion conditions and diversified image contents. To confront this challenge, we in this paper propose a novel BIQA pipeline based on the Transformer architecture, which achieves an efficient quality-aware feature representation with much fewer data. More specifically, we consider the traditional fine-tuning in BIQA as an interpretation of the pre-trained model. In this way, we further introduce a Transformer decoder to refine the perceptual information of the CLS token from different perspectives. This enables our model to establish the quality-aware feature manifold efficiently while attaining a strong generalization capability. Meanwhile, inspired by the subjective evaluation behaviors of human, we introduce a novel attention panel mechanism, which improves the model performance and reduces the prediction uncertainty simultaneously. The proposed BIQA method maintains a light-weight design with only one layer of the decoder, yet extensive experiments on eight standard BIQA datasets (both synthetic and authentic) demonstrate its superior performance to the state-of-the-art BIQA methods, i.e., achieving the SRCC values of 0.875 (vs. 0.859 in LIVEC) and 0.980 (vs. 0.969 in LIVE). Checkpoints, logs and code will be available at https://github.com/narthchin/DEIQT.

YNIMG Journal 2023 Journal Article

In vivo labeling and quantitative imaging of neuronal populations using MRI

  • Shana Li
  • Xiang Xu
  • Canjun Li
  • Ziyan Xu
  • Ke Wu
  • Qiong Ye
  • Yan Zhang
  • Xiaohua Jiang

The study of neural circuits, which underlies perception, cognition, emotion, and behavior, is essential for understanding the mammalian brain, a complex organ consisting of billions of neurons. To study the structure and function of the brain, in vivo neuronal labeling and imaging techniques are crucial as they provide true physiological information that ex vivo methods cannot offer. In this paper, we present a new strategy for in vivo neuronal labeling and quantification using MRI. We demonstrate the efficacy of this method by delivering the oatp1a1 gene to the target neurons using rAAV2-retro virus. OATP1A1 protein expression on the neuronal membrane increased the uptake of a specific MRI contrast agent (Gd-EOB-DTPA), leading to hyperintense signals on T1W images of labeled neuronal populations. We also used dynamic contrast enhancement-based methods to obtain quantitative information on labeled neuronal populations in vivo.

ICML Conference 2023 Conference Paper

Unlocking Slot Attention by Changing Optimal Transport Costs

  • Yan Zhang
  • David W. Zhang
  • Simon Lacoste-Julien
  • Gertjan J. Burghouts
  • Cees G. M. Snoek

Slot attention is a powerful method for object-centric modeling in images and videos. However, its set-equivariance limits its ability to handle videos with a dynamic number of objects because it cannot break ties. To overcome this limitation, we first establish a connection between slot attention and optimal transport. Based on this new perspective we propose MESH (Minimize Entropy of Sinkhorn): a cross-attention module that combines the tiebreaking properties of unregularized optimal transport with the speed of regularized optimal transport. We evaluate slot attention using MESH on multiple object-centric learning benchmarks and find significant improvements over slot attention in every setting.

JBHI Journal 2022 Journal Article

A Fully Deep Learning Paradigm for Pneumoconiosis Staging on Chest Radiographs

  • Wenjian Sun
  • Dongsheng Wu
  • Yang Luo
  • Lu Liu
  • Hongjing Zhang
  • Shuang Wu
  • Yan Zhang
  • Chenglong Wang

Pneumoconiosis staging has been a very challenging task, both for certified radiologists and computer-aided detection algorithms. Although deep learning has shown proven advantages in the detection of pneumoconiosis, it remains challenging in pneumoconiosis staging due to the stage ambiguity of pneumoconiosis and noisy samples caused by misdiagnosis when they are used in training deep learning models. In this article, we propose a fully deep learning pneumoconiosis staging paradigm that comprises a segmentation procedure and a staging procedure. The segmentation procedure extracts lung fields in chest radiographs through an Asymmetric Encoder-Decoder Network (AED-Net) that can mitigate the domain shift between multiple datasets. The staging procedure classifies the lung fields into four stages through our proposed deep log-normal label distribution learning and focal staging loss. The two cascaded procedures can effectively solve the problem of model overfitting caused by stage ambiguity and noisy labels of pneumoconiosis. Besides, we collect a clinical chest radiograph dataset of pneumoconiosis from the certified radiologist's diagnostic reports. The experimental results on this novel pneumoconiosis dataset confirm that the proposed deep pneumoconiosis staging paradigm achieves an Accuracy of 90. 4%, a Precision of 84. 8%, a Sensitivity of 78. 4%, a Specificity of 95. 6%, an F1-score of 80. 9% and an Area Under the Curve (AUC) of 96%. In particular, we achieve 68. 4% Precision, 76. 5% Sensitivity, 95% Specificity, 72. 2% F1-score and 89% AUC on the early pneumoconiosis ‘stage-1’.

YNICL Journal 2022 Journal Article

Abnormal patterns of regional homogeneity and functional connectivity across the adolescent first-episode, adult first-episode and adult chronic schizophrenia

  • Yongfeng Yang
  • Yuqing Sun
  • Yuliang Zhang
  • Xueyan Jin
  • Zheng Li
  • Minli Ding
  • Han Shi
  • Qing Liu

Functional deficits in schizophrenia (SZ) are observed prior to the onset of psychosis and differ at different stages of SZ. However, there is a paucity of studies focused on adolescent first-episode SZ (AOS), adult first-episode SZ (AFES), and adult chronic SZ (CHSZ). In this study, we investigated regional activity and corresponding functional connectivity alterations that have aimed to compare the three disease stages simultaneously. The subjects comprised 49 patients with AOS, 57 patients with AFES, 51 patients with CHSZ, 41 adolescent healthy controls, and 138 adult healthy controls. We compared regional homogeneity (ReHo) between patients at each disease stage with matched healthy controls. We focused on the shared brain regions that showed significant differences between SZ patients at the three different disease stages and healthy controls. Further analysis was conducted to explore whether the patterns of the whole brain functional connectivity alterations were similar. The putamen and medial frontal gyrus (MFG) showed consistently abnormal patterns in AOS, AFES, and CHSZ. Commonly decreased ReHo values in the MFG and increased ReHo values in the bilateral putamen were found in AOS, AFES, and CHSZ. Functional connectivity of MFG remained common abnormality in different SZ stage. In conclusion, ReHo abnormalities in the MFG and the putamen may be common abnormal patterns of brain function in the three different stages of SZ. The vmPFC-dlPFC FC abnormality common occurs in adolescence and adulthood.. This study may provide a more comprehensive understanding of the neurodevelopmental abnormality across the AOS, AFES, and CHSZ.

TCS Journal 2022 Journal Article

Encoding safety in CLL

  • Yan Zhang
  • Zhaohui Zhu
  • Jinjin Zhang

Graphical representation (representing logical specifications by means of one or several labeled transition systems) is a typical connection between process algebra and temporal logic, the two main paradigms that specify and reason about reactive concurrent systems. In this paper, we encode (represent “graphically”) a fragment of Action-based CTL, proposed by Lüttgen and Vogler, in the process calculus CLL R. In this way, safety properties can be described easily, and usual process operators (parallel, choice, etc.), logical operators (conjunction and disjunction) and temporal operators (always, until, etc.) can be mixed freely in CLL R.

JBHI Journal 2022 Journal Article

Metabolic and Transcriptional Analysis of Recombinant Saccharomyces Cerevisiae for Xylose Fermentation: A Feasible and Efficient Approach

  • Xin-Chi Shi
  • Yan Zhang
  • Ting Wang
  • Xiang-Chen Wang
  • Hai-Bin Lv
  • Pedro Laborda
  • Ting-Ting Duan

Lignocellulose is an abundant xylose-containing biomass found in agricultural wastes, and has arisen as a suitable alternative to fossil fuels for the production of bioethanol. Although Saccharomyces cerevisiae has been thoroughly used for the production of bioethanol, its potential to utilize lignocellulose remains poorly understood. In this work, xylose-metabolic genes of Pichia stipitis and Candida tropicalis, under the control of different promoters, were introduced into S. cerevisiae. RNA-seq analysis was use to examine the response of S. cerevisiae metabolism to the introduction of xylose-metabolic genes. The use of the PGK1 promoter to drive xylitol dehydrogenase (XDH) expression, instead of the TEF1 promoter, improved xylose utilization in “XR-pXDH” strain by overexpressing xylose reductase (XR) and XDH form C. tropicalis, enhancing the production of xylitol (13. 66 $\pm$ 0. 54 g/L after 6 days fermentation). Overexpression of xylulokinase and XR/XDH from P. stipitis remarkably decreased xylitol accumulation (1. 13 $\pm$ 0. 06 g/L and 0. 89 $\pm$ 0. 04 g/L xylitol, respectively) and increased ethanol production (196. 14 $\%$ and 148. 50 $\%$ increases during the xylose utilization stage, respectively), in comparison with the results of XR-pXDH. This result may be produced due to the enhanced xylose transport, Embden-Meyerhof and pentose phosphate pathways, as well as alleviated oxidative stress. The low xylose consumption rate in these recombinant as well as alleviated strains comparing with P. stipitis and C. tropicalis may be explained by the insufficient supplementation of NADPH and NAD $^+$. The results obtained in this work provide new insights on the potential utilization of xylose using bioengineered S. cerevisiae strains.

ICLR Conference 2022 Conference Paper

Multiset-Equivariant Set Prediction with Approximate Implicit Differentiation

  • Yan Zhang
  • David W. Zhang
  • Simon Lacoste-Julien
  • Gertjan J. Burghouts
  • Cees G. M. Snoek

Most set prediction models in deep learning use set-equivariant operations, but they actually operate on multisets. We show that set-equivariant functions cannot represent certain functions on multisets, so we introduce the more appropriate notion of multiset-equivariance. We identify that the existing Deep Set Prediction Network (DSPN) can be multiset-equivariant without being hindered by set-equivariance and improve it with approximate implicit differentiation, allowing for better optimization while being faster and saving memory. In a range of toy experiments, we show that the perspective of multiset-equivariance is beneficial and that our changes to DSPN achieve better results in most cases. On CLEVR object property prediction, we substantially improve over the state-of-the-art Slot Attention from 8% to 77% in one of the strictest evaluation metrics because of the benefits made possible by implicit differentiation.

IJCAI Conference 2021 Conference Paper

Keep the Structure: A Latent Shift-Reduce Parser for Semantic Parsing

  • Yuntao Li
  • Bei Chen
  • Qian Liu
  • Yan Gao
  • Jian-Guang Lou
  • Yan Zhang
  • Dongmei Zhang

Traditional end-to-end semantic parsing models treat a natural language utterance as a holonomic structure. However, hierarchical structures exist in natural languages, which also align with the hierarchical structures of logical forms. In this paper, we propose a latent shift-reduce parser, called LASP, which decomposes both natural language queries and logical form expressions according to their hierarchical structures and finds local alignment between them to enhance semantic parsing. LASP consists of a base parser and a shift-reduce splitter. The splitter dynamically separates an NL query into several spans. The base parser converts the relevant simple spans into logical forms, which are further combined to obtain the final logical form. We conducted empirical studies on two datasets across different domains and different types of logical forms. The results demonstrate that the proposed method significantly improves the performance of semantic parsing, especially on unseen scenarios.

AAAI Conference 2021 Conference Paper

Learning from History: Modeling Temporal Knowledge Graphs with Sequential Copy-Generation Networks

  • Cunchao Zhu
  • Muhao Chen
  • Changjun Fan
  • Guangquan Cheng
  • Yan Zhang

Large knowledge graphs often grow to store temporal facts that model the dynamic relations or interactions of entities along the timeline. Since such temporal knowledge graphs often suffer from incompleteness, it is important to develop time-aware representation learning models that help to infer the missing temporal facts. While the temporal facts are typically evolving, it is observed that many facts often show a repeated pattern along the timeline, such as economic crises and diplomatic activities. This observation indicates that a model could potentially learn much from the known facts appeared in history. To this end, we propose a new representation learning model for temporal knowledge graphs, namely CyGNet, based on a novel timeaware copy-generation mechanism. CyGNet is not only able to predict future facts from the whole entity vocabulary, but also capable of identifying facts with repetition and accordingly predicting such future facts with reference to the known facts in the past. We evaluate the proposed method on the knowledge graph completion task using five benchmark datasets. Extensive experiments demonstrate the effectiveness of CyGNet for predicting future facts with repetition as well as de novo fact prediction.

AAAI Conference 2021 Conference Paper

Partial-Label and Structure-constrained Deep Coupled Factorization Network

  • Yan Zhang
  • Zhao Zhang
  • Yang Wang
  • Zheng Zhang
  • Li Zhang
  • Shuicheng Yan
  • Meng Wang

In this paper, we technically propose an enriched prior guided framework, called Dual-constrained Deep Semi-Supervised Coupled Factorization Network (DS2 CF-Net), for discovering hierarchical coupled data representation. To extract hidden deep features, DS2 CF-Net is formulated as a partial-label and geometrical structure-constrained framework. Specifically, DS2 CF-Net designs a deep factorization architecture using multilayers of linear transformations, which can coupled update both the basis vectors and new representations in each layer. To enable learned deep representations and coefficients to be discriminative, we also consider enriching the supervised prior by joint deep coefficients-based label prediction and then incorporate the enriched prior information as additional label and structure constraints. The label constraint can enable the intra-class samples to have same coordinate in feature space, and the structure constraint forces the coefficients in each layer to be block-diagonal so that the enriched prior using the self-expressive label propagation are more accurate. Our network also integrates the adaptive dualgraph learning to retain the local structures of both data and feature manifolds in each layer. Extensive experiments on image datasets demonstrate the effectiveness of DS2 CF-Net for representation learning and clustering.

NeurIPS Conference 2020 Conference Paper

Better Set Representations For Relational Reasoning

  • Qian Huang
  • Horace He
  • Abhay Singh
  • Yan Zhang
  • Ser Nam Lim
  • Austin R. Benson

Incorporating relational reasoning into neural networks has greatly expanded their capabilities and scope. One defining trait of relational reasoning is that it operates on a set of entities, as opposed to standard vector representations. Existing end-to-end approaches for relational reasoning typically extract entities from inputs by directly interpreting the latent feature representations as a set. We show that these approaches do not respect set permutational invariance and thus have fundamental representational limitations. To resolve this limitation, we propose a simple and general network module called Set Refiner Network (SRN). We first use synthetic image experiments to demonstrate how our approach effectively decomposes objects without explicit supervision. Then, we insert our module into existing relational reasoning models and show that respecting set invariance leads to substantial gains in prediction performance and robustness on several relational reasoning tasks. Code can be found at github. com/CUAI/BetterSetRepresentations.

IS Journal 2020 Journal Article

Collaborative Generative Hashing for Marketing and Fast Cold-Start Recommendation

  • Yan Zhang
  • Ivor W. Tsang
  • Lixin Duan

Cold-start has being a critical issue in recommender systems with the explosion of data in e-commerce. Most existing studies proposed to alleviate the cold-start problem are also known as hybrid recommender systems that learn representations of users and items by combining user-item interactive and user/item content information. However, previous hybrid methods regularly suffered poor efficiency bottlenecking in online recommendations with large-scale items, because they were designed to project users and items into continuous latent space where the online recommendation is expensive. To this end, we propose a collaborative generated hashing (CGH) framework to improve the efficiency by denoting users and items as binary codes, then fast hashing search techniques can be used to speed up the online recommendation. In addition, the proposed CGH can generate potential users or items for marketing application where the generative network is designed with the principle of minimum description length, which is used to learn compact and informative binary codes. Extensive experiments on two public datasets show the advantages for recommendations in various settings over competing baselines and analyze its feasibility in marketing application.

AAAI Conference 2020 Conference Paper

Learning from Positive and Unlabeled Data without Explicit Estimation of Class Prior

  • Chenguang Zhang
  • Yuexian Hou
  • Yan Zhang

Learning a classifier from positive and unlabeled data may occur in various applications. It differs from the standard classification problems by the absence of labeled negative examples in the training set. So far, two main strategies have typically been used for this issue: the likely negative examplesbased strategy and the class prior-based strategy, in which the likely negative examples or the class prior is required to be obtained in a preprocessing step. In this paper, a new strategy based on the Bhattacharyya coefficient is put forward, which formalizes this learning problem as an optimization problem and does not need a preprocessing step. We first show that with the given positive class conditional probability density function (PDF) and the mixture PDF of both the positive class and the negative class, the class prior can be estimated by minimizing the Bhattacharyya coefficient of the positive class with respect to the negative class. We then show how to use this result in an implicit mixture model of restricted Boltzmann machines to estimate the positive class conditional PDF and the negative class conditional PDF directly to obtain a classifier without the explicit estimation of the class prior. Many experiments on real and synthetic datasets illustrated the superiority of the proposed approach.

IJCAI Conference 2020 Conference Paper

Model-theoretic Characterizations of Existential Rule Languages

  • Heng Zhang
  • Yan Zhang
  • Guifei Jiang

Existential rules, a. k. a. dependencies in databases, and Datalog+/- in knowledge representation and reasoning recently, are a family of important logical languages widely used in computer science and artificial intelligence. Towards a deep understanding of these languages in model theory, we establish model-theoretic characterizations for a number of existential rule languages such as (disjunctive) embedded dependencies, tuple-generating dependencies (TGDs), (frontier-)guarded TGDs and linear TGDs. All these characterizations hold for the class of arbitrary structures, and most of them also work on the class of finite structures. As a natural application of these results, complexity bounds for the rewritability of above languages are also identified.

YNIMG Journal 2020 Journal Article

Neural mechanisms of AVPR1A RS3-RS1 haplotypes that impact verbal learning and memory

  • Yan Zhang
  • Dan Zhu
  • Peng Zhang
  • Wei Li
  • Wen Qin
  • Feng Liu
  • Jiayuan Xu
  • Qiang Xu

Converging evidence from both human and animal studies has highlighted the pervasive role of the neuropeptide arginine vasopressin (AVP), which is mediated by arginine vasopressin receptor 1A (AVPR1A), in both social and nonsocial learning and memory. However, the effect of genetic variants in AVPR1A on verbal learning and memory is unknown. The hippocampus is a heterogeneous structure that consists of several anatomically and functionally distinct subfields, and it is the principal target structure for the memory-enhancing effect of AVP. We tested the hypothesis that genetic variants in the RS3 and RS1 repeat polymorphisms may influence verbal learning and memory performance evaluated by the California Verbal Learning Test-II (CVLT-II) by modulating the gray matter volume (GMV) and resting-state functional connectivity (rsFC) of whole hippocampus and its subfields in a large cohort of young healthy subjects (n = 1001). Using a short/long classification scheme for the repeat length of RS3 and RS1, we found that the individuals carrying more short alleles of RS3-RS1 haplotypes had poorer learning and memory performance compared to that of those carrying more long alleles. We also revealed that individuals carrying more short alleles exhibited a significantly smaller GMV in the left cornu ammonis (CA)2/3 and weaker rsFC of the left CA2/3-bilateral thalamic (primarily in medial prefrontal subfields) compared to those carrying more long alleles. Furthermore, multiple mediation analysis confirmed that these two hippocampal imaging measures jointly and fully mediated the relationship between the genetic variants in AVPR1A RS3-RS1 haplotypes and the individual differences in verbal learning and memory performance. Our results suggest that genetic variants in AVPR1A RS3-RS1 haplotypes may affect verbal learning and memory performance in part by modulating the left hippocampal CA2/3 structure and its rsFC with the thalamus.

AIIM Journal 2020 Journal Article

Quantitative knowledge presentation models of traditional Chinese medicine (TCM): A review

  • Xiaoli Chu
  • Bingzhen Sun
  • Qingchun Huang
  • Shouping Peng
  • Yingyan Zhou
  • Yan Zhang

Modern computer technology sheds light on new ways of innovating Traditional Chinese Medicine (TCM). One method that gets increasing attention is the quantitative research method, which makes use of data mining and artificial intelligence technology as well as the mathematical principles in the research on rationales, academic viewpoints of famous doctors of TCM, dialectical treatment by TCM, clinical technology of TCM, the patterns of TCM prescriptions, clinical curative effects of TCM and other aspects. This paper reviews the methods, means, progress and achievements of quantitative research on TCM. In the core database of the Web of Science, "Traditional Chinese Medicine", "Computational Science" and "Mathematical Computational Biology" are selected as the main retrieval fields, and the retrieval time interval from 1999 to 2019 is used to collect relevant literature. It is found that researchers from China Academy of Chinese Medical Sciences, Zhejiang University, Chinese Academy of Sciences and other institutes have opened up new methods of research on TCM since 2009, with quantitative methods and knowledge presentation models. The adopted tools mainly consist of text mining, knowledge discovery, technologies of the TCM database, data mining and drug discovery through TCM calculation, etc. In the future, research on quantitative models of TCM will focus on solving the heterogeneity and incompleteness of big data of TCM, establishing standardized treatment systems, and promoting the development of modernization and internationalization of TCM.

AAAI Conference 2020 Conference Paper

SK-Net: Deep Learning on Point Cloud via End-to-End Discovery of Spatial Keypoints

  • Weikun Wu
  • Yan Zhang
  • David Wang
  • Yunqi Lei

Since the PointNet was proposed, deep learning on point cloud has been the concentration of intense 3D research. However, existing point-based methods usually are not adequate to extract the local features and the spatial pattern of a point cloud for further shape understanding. This paper presents an end-to-end framework, SK-Net, to jointly optimize the inference of spatial keypoint with the learning of feature representation of a point cloud for a specific point cloud task. One key process of SK-Net is the generation of spatial keypoints (Skeypoints). It is jointly conducted by two proposed regulating losses and a task objective function without knowledge of Skeypoint location annotations and proposals. Specifically, our Skeypoints are not sensitive to the location consistency but are acutely aware of shape. Another key process of SK-Net is the extraction of the local structure of Skeypoints (detail feature) and the local spatial pattern of normalized Skeypoints (pattern feature). This process generates a comprehensive representation, pattern-detail (PD) feature, which comprises the local detail information of a point cloud and reveals its spatial pattern through the part district reconstruction on normalized Skeypoints. Consequently, our network is prompted to effectively understand the correlation between different regions of a point cloud and integrate contextual information of the point cloud. In point cloud tasks, such as classification and segmentation, our proposed method performs better than or comparable with the state-of-the-art approaches. We also present an ablation study to demonstrate the advantages of SK-Net.

AAAI Conference 2020 Conference Paper

Towards Universal Languages for Tractable Ontology Mediated Query Answering

  • Heng Zhang
  • Yan Zhang
  • Jia-Huai You
  • Zhiyong Feng
  • Guifei Jiang

An ontology language for ontology mediated query answering (OMQA-language) is universal for a family of OMQAlanguages if it is the most expressive one among this family. In this paper, we focus on three families of tractable OMQAlanguages, including first-order rewritable languages and languages whose data complexity of the query answering is in AC0 or PTIME. On the negative side, we prove that there is, in general, no universal language for each of these families of languages. On the positive side, we propose a novel property, the locality, to approximate the first-order rewritability, and show that there exists a language of disjunctive embedded dependencies that is universal for the family of OMQAlanguages with locality. All of these results apply to OMQA with query languages such as conjunctive queries, unions of conjunctive queries and acyclic conjunctive queries.

NeurIPS Conference 2019 Conference Paper

Deep Set Prediction Networks

  • Yan Zhang
  • Jonathon Hare
  • Adam Prugel-Bennett

Current approaches for predicting sets from feature vectors ignore the unordered nature of sets and suffer from discontinuity issues as a result. We propose a general model for predicting sets that properly respects the structure of sets and avoids this problem. With a single feature vector as input, we show that our model is able to auto-encode point sets, predict the set of bounding boxes of objects in an image, and predict the set of attributes of these objects.

AAAI Conference 2019 Conference Paper

EA Reader: Enhance Attentive Reader for Cloze-Style Question Answering via Multi-Space Context Fusion

  • Chengzhen Fu
  • Yan Zhang

Query-document semantic interactions are essential for the success of many cloze-style question answering models. Recently, researchers have proposed several attention-based methods to predict the answer by focusing on appropriate subparts of the context document. In this paper, we design a novel module to produce the query-aware context vector, named Multi-Space based Context Fusion (MSCF), with the following considerations: (1) interactions are applied across multiple latent semantic spaces; (2) attention is measured at bit level, not at token level. Moreover, we extend MSCF to the multi-hop architecture. This unified model is called Enhanced Attentive Reader (EA Reader). During the iterative inference process, the reader is equipped with a novel memory update rule and maintains the understanding of documents through read, update and write operations. We conduct extensive experiments on four real-world datasets. Our results demonstrate that EA Reader outperforms state-of-the-art models.

JAIR Journal 2019 Journal Article

Polynomial and Exponential Bounded Logic Programs with Function Symbols: Some New Decidable Classes

  • Vernon Asuncion
  • Yan Zhang
  • Heng Zhang
  • Ruixuan Li

A logic program with function symbols is called finitely ground if there is a finite propositional logic program whose stable models are exactly the same as the stable models of this program. Finite groundability is an important property for logic programs with function symbols because it makes feasible to compute such programs' stable models using traditional ASP solvers. In this paper, we introduce new decidable classes of finitely ground programs called poly-bounded and k-EXP-bounded programs, which, to the best of our knowledge, strictly contain all other decidable classes of finitely ground programs discovered so far in the literature. We also study the relevant complexity properties for these classes of programs. We prove that the membership complexities for poly-bounded and k-EXP-bounded programs are EXPTIME-complete and (k+1)-EXPTIME-complete, respectively.

AAMAS Conference 2018 Conference Paper

CityScope Andorra: A Multi-level Interactive and Tangible Agent-based Visualization

  • Arnaud Grignard
  • N�ria Maci�
  • Luis Alonso Pastor
  • Ariel Noyman
  • Yan Zhang
  • Kent Larson

This study proposes a novel information visualization approach developed and deployed in the state of Andorra. We present a framework to analyze and represent the flow of people through a multi-level interactive and tangible agent-based visualization. The presented framework, developed to understand Andorra visitor behavior, is embedded in the MIT CityScope framework used for civic engagement, urban development, and decision making.

AAAI Conference 2018 Conference Paper

COSINE: Community-Preserving Social Network Embedding From Information Diffusion Cascades

  • Yuan Zhang
  • Tianshu Lyu
  • Yan Zhang

This paper studies the problem of social network embedding without relying on network structures that are usually not observed in many cases. We address that the information diffusion process across networks naturally reflects rich proximity relationships between users. Meanwhile, social networks contain multiple communities regularizing communication pathways for information propagation. Based on the above observations, we propose a probabilistic generative model, called COSINE, to learn community-preserving social network embeddings from the recurrent and time-stamped social contagion logs, namely information diffusion cascades. The learned embeddings therefore capture the high-order user proximities in social networks. Leveraging COSINE, we are able to discover underlying social communities and predict temporal dynamics of social contagion. Experimental results on both synthetic and real-world datasets show that our proposed model significantly outperforms the existing approaches.

AAMAS Conference 2018 Conference Paper

Real-time Machine Learning Prediction of an Agent-Based Model for Urban Decision-making

  • Yan Zhang
  • Arnaud Grignard
  • Kevin Lyons
  • Alexander Aubuchon
  • Kent Larson

CityMatrix is an urban decision support system that has been developed to facilitate more collaborative and evidence-based urban decision-making for experts and non-experts. Machine learning techniques have been applied to achieve real-time prediction of an agent-based model (ABM) of city traffic. The prediction with a shallow convolutional neural network (CNN) is significantly faster than performing the original ABM, and has enough accuracy for decision-making. The result is a versatile, quick, accurate, and computationally efficient approach to provide real-time feedback and optimization for urban decision-making.

EAAI Journal 2018 Journal Article

Sparsity-based inverse halftoning via semi-coupled multi-dictionary learning and structural clustering

  • Yan Zhang
  • Erhu Zhang
  • Wanjun Chen
  • Yajun Chen
  • Jinghong Duan

Inverse halftoning is the restoration of a continuous-tone image from its halftone version, which is a critical process for halftone transform, digital archive management and high precision identification of halftone. In this paper, a novel inverse halftoning method based on semi-coupled multi-dictionary learning is proposed to address the cross-style image restoration from halftone images to continuous-tone images. By using semi-coupled multi-dictionary learning, multiple dictionary pairs and their corresponding mapping functions between continuous-tone image and its halftone version could be simultaneously learned. The learned multiple dictionary pairs can well represent the structure characteristics of halftone images and continuous-tone images, respectively. In addition, the mapping functions learned by semi-coupled manner can bridge the gap between the two different style images of halftone image and continuous-tone image. Unlike the existed methods, the proposed method could effectively relax the assumption of the same sparse coding coefficients in coupled dictionary learning. To obtain more accurate mapping functions, a structural clustering method for cross-style image patches is proposed by using SUSAN (smallest univalue segment assimilating nucleus) filtering and HOG (histogram of oriented gradient) features, which can capture the similar structure features from halftone images and continuous-tone images, and thus improve the classification accurate rate of halftone image patches. The experimental results demonstrate that the proposed method can restore higher quality continuous-tone images than that produced by the state-of-the-art methods, which not only reduce the screen noise in smooth regions, but also provide well fine details and clear edges.

AIJ Journal 2017 Journal Article

A progression semantics for first-order logic programs

  • Yi Zhou
  • Yan Zhang

In this paper, we propose a progression semantics for first-order normal logic programs, and show that it is equivalent to the well-known stable model (answer set) semantics. The progressional definition sheds new insights into Answer Set Programming (ASP), for instance, its relationships to Datalog, First-Order Logic (FOL) and Satisfiability Modulo Theories (SMT). As an example, we extend the notion of boundedness in Datalog for ASP, and show that it coincides with the notions of recursion-freeness and loop-freeness under program equivalence. In addition, we prove that boundedness precisely captures first-order definability for normal logic programs on arbitrary structures. Finally, we show that the progressional definition suggests an alternative translation from ASP to SMT, which yields a new way of implementing first-order ASP.

AAAI Conference 2017 Conference Paper

Discrete Personalized Ranking for Fast Collaborative Filtering from Implicit Feedback

  • Yan Zhang
  • Defu Lian
  • Guowu Yang

Personalized ranking is usually considered as an ultimate goal of recommendation systems, but it suffers from efficiency issues when making recommendations. To this end, we propose a learning-based hashing framework called Discrete Personalized Ranking (DPR), to map users and items to a Hamming space, where user-item affinity can be efficiently calculated via Hamming distance. Due to the existence of discrete constraints, it is possible to exploit a two-stage learning procedure for learning binary codes according to most existing methods. This two-stage procedure consists of relaxed optimization by discarding discrete constraints and subsequent binary quantization. However, such a procedure has been shown resulting in a large quantization loss, so that longer binary codes would be required. To this end, DPR directly tackles the discrete optimization problem of personalized ranking. And the balance and un-correlation constraints of binary codes are imposed to derive compact but informatics binary codes. Based on the evaluation on several datasets, the proposed framework shows consistent superiority to the competing baselines even though only using shorter binary code.

AAAI Conference 2017 Conference Paper

Polynomially Bounded Logic Programs with Function Symbols: A New Decidable

  • Vernon Asuncion
  • Yan Zhang
  • Heng Zhang

A logic program with function symbols is called finitely ground if there is a finite propositional logic program whose stable models are exactly the same as the stable models of this program. Finite groundability is an important property for logic programs with function symbols because it makes feasible to compute such program’s stable models using traditional ASP solvers. In this paper, we introduce a new decidable class of finitely ground programs called POLY-bounded programs, which, to the best of our knowledge, strictly contains all decidable classes of finitely ground programs discovered so far in the literature. We also study the related complexity property for this class of programs. We prove that deciding whether a program is POLY-bounded is EXPTIMEcomplete.

EAAI Journal 2016 Journal Article

Deep neural network for halftone image classification based on sparse auto-encoder

  • Yan Zhang
  • Erhu Zhang
  • Wanjun Chen

To restore high quality continuous tone images from each class of halftone images, halftone image fine classification is the key problem. In this paper, a novel feature learning method is proposed for classifying 14 kinds of halftone images produced by the most well-known halftoning algorithms. This study employs the stacked sparse auto-encoders (SAE) trained with unsupervised learning for extracting features of halftone images, and then uses softmax regression with supervised learning for fine-tuning the deep neural network and classifying halftone images. In order to reduce the run-time of deep neural network and improve the image correct classification rate, we propose an effective patch extraction method for testing halftone images by measuring the mean and variance of local entropy in a patch. Halftone image classification is determined by the classification results of all effective patches inside an image via majority voting (MV). The experimental results demonstrate that our proposed method achieves an average correct classification rate (ACCR) of over 99. 44% for 14 kinds of halftone images on two public image sets. Compared with state-of-the-art LMS–Bayes and M 10 – ML methods, the proposed SAE-MV method can distinguish the most categories of halftone images and achieve competitive ACCR, meanwhile, demonstrate better generalization performance.

IJCAI Conference 2016 Conference Paper

Expressive Completeness of Existential Rule Languages for Ontology-Based Query Answering

  • Heng Zhang
  • Yan Zhang
  • Jia-Huai You

Existential rules, also known as data dependencies in Databases, have been recently rediscovered as a promising family of languages for Ontology-based Query Answering. In this paper, we prove that disjunctive embedded dependencies exactly capture the class of recursively enumerable ontologies in Ontology-based Conjunctive Query Answering (OCQA). Our expressive completeness result does not rely on any built-in linear order on the database. To establish the expressive completeness, we introduce a novel semantic definition for OCQA ontologies. We also show that neither the class of disjunctive tuple-generating dependencies nor the class of embedded dependencies is expressively complete for recursively enumerable OCQA ontologies.

AAAI Conference 2016 Conference Paper

Query Answering with Inconsistent Existential Rules under Stable Model Semantics

  • Hai Wan
  • Heng Zhang
  • Peng Xiao
  • Haoran Huang
  • Yan Zhang

Classical inconsistency-tolerant query answering relies on selecting maximal components of an ABox/database which are consistent with the ontology. However, some rules in ontologies might be unreliable if they are extracted from ontology learning or written by unskillful knowledge engineers. In this paper we present a framework of handling inconsistent existential rules under stable model semantics, which is defined by a notion called rule repairs to select maximal components of the existential rules. Surprisingly, for R-acyclic existential rules with R-stratified or guarded existential rules with strati- fied negations, both the data complexity and combined complexity of query answering under the rule repair semantics remain the same as that under the conventional query answering semantics. This leads us to propose several approaches to handle the rule repair semantics by calling answer set programming solvers. An experimental evaluation shows that these approaches have good scalability of query answering under rule repairs on realistic cases.

AAAI Conference 2015 Conference Paper

Existential Rule Languages with Finite Chase: Complexity and Expressiveness

  • Heng Zhang
  • Yan Zhang
  • Jia-Huai You

Finite chase, or alternatively chase termination, is an important condition to ensure the decidability of existential rule languages. In the past few years, a number of rule languages with finite chase have been studied. In this work, we propose a novel approach for classifying the rule languages with finite chase. Using this approach, a family of decidable rule languages, which extend the existing languages with the finite chase property, are naturally defined. We then study the complexity of these languages. Although all of them are tractable for data complexity, we show that their combined complexity can be arbitrarily high. Furthermore, we prove that all the rule languages with finite chase that extend the weakly acyclic language are of the same expressiveness as the weakly acyclic one, while rule languages with higher combined complexity are in general more succinct than those with lower combined complexity.

AIJ Journal 2015 Journal Article

Ordered completion for logic programs with aggregates

  • Vernon Asuncion
  • Yin Chen
  • Yan Zhang
  • Yi Zhou

We consider the problem of translating first-order answer set programs with aggregates into first-order sentences with the same type of aggregates. In particular, we show that, on finite structures, normal logic programs with convex aggregates, which cover both monotone and antimonotone aggregates as well as the aggregates appearing in most benchmark programs, can always be captured in first-order logic with the same type of aggregates by introducing auxiliary predicates. More precisely, we prove that every finite stable model of a normal program with convex aggregates is corresponding to a classical model of its enhanced ordered completion. This translation then suggests an alternative way for computing the stable models of such kind of programs. We report some experimental results, which demonstrate that our solver GROCv2 is comparable to the state-of-the-art answer set solvers. We further show that convex aggregates form a maximal class for this purpose. That is, we can always construct a normal logic program under any given non-convex aggregate context and prove that it can never be translated into first-order sentences with the same type of aggregates unless NP = coNP.

AAAI Conference 2014 Conference Paper

Computing General First-Order Parallel and Prioritized Circumscription

  • Hai Wan
  • Zhanhao Xiao
  • Zhenfeng Yuan
  • Heng Zhang
  • Yan Zhang

This paper focuses on computing general first-order parallel and prioritized circumscription with varying constants. We propose linear translations from general first-order circumscription to first-order theories under stable model semantics over arbitrary structures, including Trv for parallel circumscription and Trs v for conjunction of parallel circumscriptions (further for prioritized circumscription). To improve the efficiency, we give an optimization Γ∃ to reduce logic programs in size when eliminating existential quantifiers during the translations. Based on these results, a general first-order circumscription solver, named cfo2lp, is developed by calling answer set programming (ASP) solvers. Using circuit diagnosis problem and extended stable marriage problem as benchmarks, we compare cfo2lp with a propositional circumscription solver circ2dlp and an ASP solver with complex optimization metasp on efficiency. Experimental results demonstrate that for problems represented by first-order circumscription naturally and intuitively, cfo2lp can compute all solutions over finite structures. We also apply our approach to description logics with circumscription and repairs in inconsistent databases, which can be handled effectively.

IS Journal 2014 Journal Article

How Effective Are the Prevailing Attack-Defense Models for Cybersecurity Anyway?

  • Daojing He
  • Sammy Chan
  • Yan Zhang
  • Chunming Wu
  • Bing Wang

Attack-defense models play an important role in the design of cybersecurity systems. Here, the authors review some traditional and prevailing attack-defense models along with their weaknesses. Then, they survey some recently proposed paradigm shifts to these models based on which more effective security strategies can be designed. Further, they provide some suggestions on how to adopt the new models, and present challenges that need to be addressed in this field.

JBHI Journal 2014 Journal Article

Lightweight and Confidential Data Discovery and Dissemination for Wireless Body Area Networks

  • Daojing He
  • Sammy Chan
  • Yan Zhang
  • Haomiao Yang

As a special sensor network, a wireless body area network (WBAN) provides an economical solution to real-time monitoring and reporting of patients' physiological data. After a WBAN is deployed, it is sometimes necessary to disseminate data into the network through wireless links to adjust configuration parameters of body sensors or distribute management commands and queries to sensors. A number of such protocols have been proposed recently, but they all focus on how to ensure reliability and overlook security vulnerabilities. Taking into account the unique features and application requirements of a WBAN, this paper presents the design, implementation, and evaluation of a secure, lightweight, confidential, and denial-of-service-resistant data discovery and dissemination protocol for WBANs to ensure the data items disseminated are not altered or tampered. Based on multiple one-way key hash chains, our protocol provides instantaneous authentication and can tolerate node compromise. Besides the theoretical analysis that demonstrates the security and performance of the proposed protocol, this paper also reports the experimental evaluation of our protocol in a network of resource-limited sensor nodes, which shows its efficiency in practice. In particular, extensive security analysis shows that our protocol is provably secure.

KR Conference 2014 Conference Paper

Logic Programs with Ordered Disjunction: First-order Semantics and Expressiveness

  • Vernon Asuncion
  • Yan Zhang
  • Heng Zhang

Logic programs with ordered disjunction (LPODs) (Brewka 2002) generalize normal logic programs by combining alternative and ranked options in the heads of rules. It has been showed that LPODs are useful in a number of areas including game theory, policy languages, planning and argumentations. In this paper, we extend propositional LPODs to the first-order case, where a classical second-order formula is defined to capture the stable model semantics of the underlying first-order LPODs. We then develop a progression semantics that is equivalent to the stable model semantics but naturally represents the reasoning procedure of LPODs. We show that on finite structures, every LPOD can be translated to a firstorder sentence, which provides a basis for computing stable models of LPODs. We further study the complexity and expressiveness of LPODs and prove that almost positive LPODs precisely capture first-order normal logic programs, which indicates that ordered disjunction itself and constraints are sufficient to represent negation as failure. A ← not C B ← not D A ← not C C ← not D, not B B ← not C, not A B ← not D B ← not C, not A C ← not D, not B. Then the class of stable models of Π consists of all stable models of these four split programs, which is {{A, B}, {B}, {C}}. Then by integrating proper preference relation among these stable models, the preferred stable models can be obtained for an LPOD. There have been several extensions of LPODs in recent years: Karger et al (2008) extended LPODs by allowing both ordered and unordered disjunction in the heads of rules; Confalonieri et al (2010) recently defined a possibilistic semantics for LPODs in order to handle uncertainty; and Cabalar (2011) also proposed a direct translation from LPODs to normal logic programs via the logic of Here-and-There. It has been argued that LPODs provide a natural way to deal with preference in reasoning that are useful in various applications such as game theory, policy languages, planning and argumentations (Brewka 2002; Cabalar 2011; Confalonieri et al. 2010). On the other hand, in recent years, Answer Set Programming (ASP) has been generalized to arbitrary first-order sentences (Ferraris, Lee, and Lifschitz 2011). One challenging research along this direction is to establish proper logical and computational foundations for promoting useful functionalities in existing ASP paradigm to the first-order level. A number of topics in this aspect have been investigated and relevant properties revealed, e. g., (Asuncion et al. 2012; Asuncion, Zhang, and Zhou 2013; Lee and Meng 2011; Babb and Lee 2012). One major advantage of first-order ASP is that it provides a succinct declarative language, in which the underlying problem constraints (rules) may be completely separated from concrete problem instances, and hence more flexible for problem representation and modeling (Lin and Zhou 2011). In this paper, we study the semantics and expressiveness of LPODs on the first-order level. We make the following main contributions towards this topic: 1. Following the style of general stable model semantics

IJCAI Conference 2013 Conference Paper

Definability of Horn Revision from Horn Contraction

  • Zhiqiang Zhuang
  • Maurice Pagnucco
  • Yan Zhang

In the AGM framework [Alchourrón and Makinson, 1985], a revision function can be defined directly through constructions like systems of spheres, epistemic entrenchment, etc. , or indirectly through a contraction operation via the Levi identity. A recent trend is to construct AGM style contraction and revision functions that operate under Horn logic. A direct construction of Horn revision is given in [Delgrande and Peppas, 2011]. However, it is unknown whether Horn revision can be defined indirectly from Horn contraction. In this paper, we address this problem by obtaining a model-based Horn revision through the model-based Horn contraction studied in [Zhuang and Pagnucco, 2012]. Our result shows that, under proper restrictions, Horn revision is definable through Horn contraction via the Levi identity.

IJCAI Conference 2013 Conference Paper

First-Order Expressibility and Boundedness of Disjunctive Logic Programs

  • Heng Zhang
  • Yan Zhang

In this paper, the fixed point semantics developed in [Lobo et al. , 1992] is generalized to disjunctive logic programs with default negation and over arbitrary structures, and proved to coincide with the stable model semantics. By using the tool of ultraproducts, a preservation theorem, which asserts that a disjunctive logic program without default negation is bounded with respect to the proposed semantics if and only if it has a first-order equivalent, is then obtained. For the disjunctive logic programs with default negation, a sufficient condition assuring the first-order expressibility is also proposed.

EAAI Journal 2013 Journal Article

Robust digital watermarking in PDTDFB domain based on least squares support vector machine

  • Hong-Ying Yang
  • Xiang-Yang Wang
  • Yan Zhang
  • Miao E-nuo

Geometric distortion is known as one of the most difficult attacks to resist, for it can desynchronize the location of the watermark and hence causes incorrect watermark detection. It is a challenging work to design a robust image watermarking scheme against geometric distortions. Based on the least squares support vector machine (LS-SVM) geometric distortions correction, we propose a new image watermarking scheme in shiftable complex directional pyramid (PDTDFB) domain with good visual quality and reasonable resistance toward geometric distortions in this paper. Firstly, the PDTDFB decomposition is performed on the original host image. Then, the corresponding lowpass subband is divided into small blocks. Finally, the digital watermark is embedded into host image by modulating the selected lowpass PDTDFB coefficients in small blocks. The main steps of digital watermark detecting procedure include: (1) the PDTDFB decomposition is performed on the test images, and some low-order Gaussian–Hermite moment energy of highpass subbands are computed, which are regarded as the effective feature vectors; (2) the appropriate kernel function is selected for training, and a LS-SVM training model can be obtained; (3) the watermarked image is corrected with the well trained LS-SVM model; and (4) the digital watermark is extracted from the corrected watermarked image. Experimental results show that the proposed image watermarking is not only invisible and robust against common image processing operations such as filtering, noise adding, and JPEG compression etc, but also robust against the geometrical distortions.

TCS Journal 2012 Journal Article

Almost optimal distributed M2M multicasting in wireless mesh networks

  • Qin Xin
  • Fredrik Manne
  • Yan Zhang
  • Xin Wang

Wireless Mesh Networking (WMN) is an emerging communication paradigm to enable resilient, cost-efficient and reliable services for the future-generation wireless networks. In this paper, we study the problem of multipoint-to-multipoint (M2M) multicasting in a WMN which aims to use the minimum number of time slots to exchange messages among a group of k mesh nodes in a multi-hop WMN with n mesh nodes. We study the M2M multicasting problem in a distributed environment where each participant only knows that there are k participants and it does not know who are other k − 1 participants among n mesh nodes. It is known that the computation of an optimal M2M multicasting schedule isNP-hard. We present a fully distributed deterministic algorithm for such an M2M multicasting problem and analyze its time complexity. We show that if the maximum hop distance between any two out of the k participants is d, then the studied M2M multicasting problem can be solved in time O ( d log 2 n + k log 3 n log k ) with a polynomial-time computation, which is an almost optimal scheme due to the lower bound Ω ( d + k log n log k ) given by Chlebus et al. (2009) [5]. Our algorithm also improves the currently best known result with running time O ( d log 2 n + k log 4 n ) by Gąsieniec et al. (2006) [13]. In this paper, we also propose a distributed deterministic algorithm which accomplishes the M2M multicasting in time O ( d + k ) with a polynomial-time computation in unit disk graphs. This is an asymptotically optimal algorithm in the sense that there exists a WMN topology, e. g. , a line, a ring, a star or a complete graph, in which the M2M multicasting cannot be completed in less than Ω ( d + k ) units of time.

KR Conference 2012 Short Paper

Forgetting in Logic Programs under Strong Equivalence

  • Yisong Wang
  • Yan Zhang
  • Yi Zhou
  • Mingyi Zhang

Wang (2008) then proposed a semantic forgetting for consistent disjunctive logic programs, which preserves equivalence but not strong equivalence. They specifically indicated the importance of preserving strong equivalence in logic programming forgetting and raised this issue as a future work. Wong (2009) proposed two forgetting operators for disjunctive logic programs. Although Wong’s forgetting indeed preserves strong equivalence, it may lose the intuition of weakening under various circumstances (see Related Work for details). In addition to preserving strong equivalence, expressibility is another desired criterion for logic programming forgetting. Ideally we would expect that the result of forgetting some atoms from a logic program is still expressible by a logic program. Finally, we believe that as a way of weakening, forgetting in logic programs should obey some common intuitions shared by forgetting in classical logics. For instance, forgetting something from a logic program should lead to a weaker program in certain sense. On the other hand, such weakening should only be associated to the relevant information that has been forgotten. For this purpose, Zhang and Zhou (2009) proposed four forgetting postulates to formalize these common intuitions and showed that forgetting in classical propositional logic and modal logic S5 can be precisely captured by these postulates. Interestingly, none of previous forgetting notions in logic programs actually satisfies Zhang-Zhou’s postulates. In summary, we consider the following criteria that a forgetting notion in logic program should meet: • Expressibility. The result of forgetting in an arbitrary logic program should also be expressible via an arbitrary logic program; • Preserving strong equivalence. Two strongly equivalent programs should remain strongly equivalence after forgetting the same set of atoms; • Satisfying common intuitions of forgetting. Preferably, forgetting in logic programs should be semantically characterized by Zhang-Zhou’s four forgetting postulates. In this paper we present a comprehensive study on forgetting in the context of arbitrary logic programs (propositional theories) under answer set semantics. In our approach, a program Π is viewed as a theory of the logic of here-andthere (or simply called HT logic), then forgetting a set V of In this paper, we propose a semantic forgetting for arbitrary logic programs (or propositional theories) under answer set semantics, called HT-forgetting. The HTforgetting preserves strong equivalence in the sense that strongly equivalent logic programs will remain strongly equivalent after forgetting the same set of atoms. The result of an HT-forgetting is always expressible by a logic program, and in particular, the result of an HT-forgetting in a Horn program is expressible in a Horn program; and a representation theorem shows that HT-forgetting can be precisely characterized by Zhang-Zhou’s four forgetting postulates under the logic of here-and-there. We also reveal underlying connections between HTforgetting and classical forgetting, and provide complexity results for decision problems.

AIJ Journal 2012 Journal Article

Ordered completion for first-order logic programs on finite structures

  • Vernon Asuncion
  • Fangzhen Lin
  • Yan Zhang
  • Yi Zhou

In this paper, we propose a translation from normal first-order logic programs under the stable model semantics to first-order sentences on finite structures. The translation is done through, what we call, ordered completion which is a modification of Clarkʼs completion with some auxiliary predicates added to keep track of the derivation order. We show that, on finite structures, classical models of the ordered completion of a normal logic program correspond exactly to the stable models of the program. We also extend this result to normal programs with constraints and choice rules. From a theoretical viewpoint, this work clarifies the relationships between normal logic programming under the stable model semantics and classical first-order logic. It follows that, on finite structures, every normal program can be defined by a first-order sentence if new predicates are allowed. This is a tight result as not every normal logic program can be defined by a first-order sentence if no extra predicates are allowed or when infinite structures are considered. Furthermore, we show that the result cannot be extended to disjunctive logic programs, assuming that NP ≠ coNP. From a practical viewpoint, this work leads to a new type of ASP solver by grounding on a programʼs ordered completion instead of the program itself. We report on a first implementation of such a solver based on several optimization techniques. Our experimental results show that our solver compares favorably to other major ASP solvers on the Hamiltonian Circuit program, especially on large domains.

AAAI Conference 2012 Conference Paper

Ordered Completion for Logic Programs with Aggregates

  • Vernon Asuncion
  • Yan Zhang
  • Yi Zhou

In this paper, we show that first-order logic programs with monotone aggregates under the stable model semantics can be captured in classical first-order logic. More precisely, we extend the notion of ordered completion for logic programs with a large variety of aggregates so that every stable model of a program with aggregates corresponds to a classical model of its enhanced ordered completion, and vice versa.

AAAI Conference 2011 Conference Paper

Bounded Forgetting

  • Yi Zhou
  • Yan Zhang

The result of forgetting some predicates in a first-order sentence may not exist in the sense that it might not be captured by any first-order sentences. This, indeed, severely restricts the usage of forgetting in applications. To address this issue, we propose a notion called k-forgetting, also called bounded forgetting in general, for any fixed number k. We present several equivalent characterizations of bounded forgetting and show that the result of bounded forgetting, on one hand, can always be captured by a single first-order sentence, and on the other hand, preserves the information that we are concerned with.

AIJ Journal 2011 Journal Article

Loop-separable programs and their first-order definability

  • Yin Chen
  • Fangzhen Lin
  • Yan Zhang
  • Yi Zhou

An answer set program with variables is first-order definable on finite structures if the set of its finite answer sets can be captured by a first-order sentence. Characterizing classes of programs that are first-order definable on finite structures is theoretically challenging and of practical relevance to answer set programming. In this paper, we identify a non-trivial class of answer set programs called loop-separable programs and show that they are first-order definable on finite structures.

AAAI Conference 2011 Conference Paper

Progression Semantics for Disjunctive Logic Programs

  • Yi Zhou
  • Yan Zhang

In this paper, we extend the progression semantics for firstorder disjunctive logic programs and show that it coincides with the stable model semantics. Based on it, we further show how disjunctive answer set programming is related to Satisfiability Modulo Theories.

IJCAI Conference 2011 Conference Paper

Translating First-Order Theories into Logic Programs

  • Heng Zhang
  • Yan Zhang
  • Mingsheng Ying
  • Yi Zhou

This paper focuses on computing first-order theories under either stable model semantics or circumscription. A reduction from first-order theories to logic programs under stable model semantics over finite structures is proposed, and an embedding of circumscription into stable model semantics is also given. Having such reduction and embedding, reasoning problems represented by first-order theories under these two semantics can then be handled by using existing answer set solvers. The effectiveness of this approach in computing hard problems beyond NP is demonstrated by some experiments.

IS Journal 2010 Journal Article

Context-Aware Middleware and Intelligent Agents for Smart Environments

  • Hamid R. Arabnia
  • Wai-Chi Fang
  • Changhoon Lee
  • Yan Zhang

The present article issues the latest research and development in context-aware middleware and intelligent agents for smart environments. Smart environments-smart homes, smart offices, smart schools, and so on-represent advanced communication and computing environments featuring continually evolving everyday objects for nonexpert users. Smart environments have rapidly emerged as an exciting new paradigm that tends to include different research fields such as ubiquitous, pervasive, and grid computing. Such environments aim to provide computing and communication services in a far more convenient, seamless, and enjoyable way. Users will be able to easily, conveniently, and remotely access and control all information and appliances in their environment, using various services resulting from the integrated cooperation of possibly heterogeneous communication-enabled objects. However, realizing the services' advantages will require appropriate middleware support to facilitate context-dependent intelligent agents, thus leveraging cost-effective design and implementation of smart-environment applications.

AAAI Conference 2010 Conference Paper

First-Order Indefinability of Answer Set Programs on Finite Structures

  • Yin Chen
  • Yan Zhang
  • Yi Zhou

An answer set program with variables is first-order definable on finite structures if the set of its finite answer sets can be captured by a first-order sentence, otherwise this program is first-order indefinable on finite structures. In this paper, we study the problem of first-order indefinability of answer set programs. We provide an Ehrenfeucht-Fraı̈ssé gametheoretic characterization for the first-order indefinability of answer set programs on finite structures. As an application of this approach, we show that the well-known finding Hamiltonian cycles program is not first-order definable on finite structures. We then define two notions named the 0-1 property and unbounded cycles or paths under the answer set semantics, from which we develop two sufficient conditions that may be effectively used in proving a program’s first-order indefinability on finite structures under certain circumstances.

KR Conference 2010 Conference Paper

Forgetting Revisited

  • Yi Zhou
  • Yan Zhang

In this paper, we propose an alternative notion, called weak forgetting, of forgetting a set of predicates in a first-order theory. One important feature of this new notion is that the result of weak forgetting is always first-order expressible. In contrast, this is not the case for the traditional notion of forgetting, called strong forgetting, introduced by Lin and Reiter. As a consequence, these two notions are not exactly the same. Interestingly, we prove that they coincide when the result of strong forgetting is first-order expressible. We also present a representation theorem to characterize weak forgetting from different aspects.

KR Conference 2010 Conference Paper

On the Progression Semantics and Boundedness of Answer Set Programs

  • Yan Zhang
  • Yi Zhou

In this paper, we propose a progression semantics for firstorder answer set programs. Based on this new semantics, we are able to define the notion of boundedness for answer set programming. We prove that boundedness coincides with the notions of recursion-free and loop-free under program equivalence, and is also equivalent to first-order definability of answer set programs on arbitrary structures.

AAAI Conference 2010 Conference Paper

Ordered Completion for First-Order Logic Programs on Finite Structures

  • Vernon Asuncion
  • Fangzhen Lin
  • Yan Zhang
  • Yi Zhou

In this paper, we propose a translation from normal first-order logic programs under the answer set semantics to first-order theories on finite structures. Specifically, we introduce ordered completions which are modifications of Clark’s completions with some extra predicates added to keep track of the derivation order, and show that on finite structures, classical models of the ordered-completion of a normal logic program correspond exactly to the answer sets (stable models) of the logic program.

AIJ Journal 2009 Journal Article

Knowledge forgetting: Properties and applications

  • Yan Zhang
  • Yi Zhou

In this paper we study a formal notion of knowledge forgetting in S5 modal logic. We propose four postulates and prove that these postulates precisely characterize both semantic and logical properties of knowledge forgetting. We then investigate possible applications of knowledge forgetting in various epistemic reasoning scenarios. In particular, we show that different forms of knowledge updates may be represented via knowledge forgetting. We also demonstrate how knowledge forgetting can be used in formalizing and reasoning about knowledge games with bounded memory.

AAMAS Conference 2008 Conference Paper

Partial Goal Satisfaction and Goal Change

  • Yi Zhou
  • Leon van der Torre
  • Yan Zhang

Partial implication semantics in the context of a background theory has been introduced to formalize partial goal satisfaction in the context of beliefs. In this paper, we introduce strong partial implication prohibiting redundancies and weak partial implication allowing side effects, we study their semantic as well as complexity properties, and we apply the three notions of partial implication to goal change in the context of beliefs.

IJCAI Conference 2007 Conference Paper

  • Yan Zhang

Although epistemic logic programming has an enhanced capacity to handle complex incomplete information reasoning and represent agents' epistemic behaviours, it embeds a significantly higher computational complexity than non-disjunctive and disjunctive answer set programming. In this paper, we investigate some important properties of epistemic logic programs. In particular, we show that Lee and Lifschitz's result on loop formulas for disjunctive logic programs can be extended to a special class of epistemic logic programs. We also study the polysize model property for epistemic logic programs. Based on these discoveries, we identify two non-trivial classes of epistemic logic programs whose consistency checking complexity is reduced from PSPACE-complete to NP-complete and \Sigma_{2}^{P}-complete respectively. We observe that many important applications on epistemic representation fall into these two classes of epistemic logic programs.

ECAI Conference 2006 Conference Paper

CTL Model Update: Semantics, Computations and Implementation

  • Yulin Ding
  • Yan Zhang

Minimal change is a fundamental principle for modeling system dynamics. In this paper, we study the issue of minimal change for Computational Tree Logic (CTL) model update. We first propose five primitive operations which capture the basic update of the CTL model, and then define the minimal change criteria for CTL model update based on these primitive operations. We provide essential semantic and computational characterizations for our CTL model update approach. We develop a formal algorithm to implement this update that employs the underlying minimal change principle. We also present a CTL model update example using the well known microwave oven scenario.

AIJ Journal 2006 Journal Article

Solving logic program conflict through strong and weak forgettings

  • Yan Zhang
  • Norman Y. Foo

We consider how to forget a set of atoms in a logic program. Intuitively, when a set of atoms is forgotten from a logic program, all atoms in the set should be eliminated from this program in some way, and other atoms related to them in the program might also be affected. We define notions of strong and weak forgettings in logic programs to capture such intuition, reveal their close connections to the notion of forgetting in classical propositional theories, and provide a precise semantic characterization for them. Based on these notions, we then develop a general framework for conflict solving in logic programs. We investigate various semantic properties and features in relation to strong and weak forgettings and conflict solving in the proposed framework. We argue that many important conflict solving problems can be represented within this framework. In particular, we show that all major logic program update approaches can be transformed into our framework, under which each approach becomes a specific conflict solving case with certain constraints. We also study essential computational properties of strong and weak forgettings and conflict solving in the framework.

AAAI Conference 2005 Conference Paper

A Unified Framework for Representing Logic Program Updates

  • Yan Zhang

As a promising formulation to represent and reason about agents’ dynamic behavious, logic program updates have been considerably studied recently. While similarities and differences between various approaches were discussed and evaluated by researchers, there is a lack of method to represent different logic program update approaches under a common framework. In this paper, we continue our study on a general framework for logic program conflict solving based on notions of strong and weak forgettings (Zhang, Foo, & Wang 2005). We show that all major logic program update approaches can be transformed into our framework, under which each update approach becomes a specific conflict solving case with certain constraints. We also investigate related computational properties for these transformations.

AIJ Journal 2005 Journal Article

Knowledge updates: Semantics and complexity issues

  • Chitta Baral
  • Yan Zhang

We consider the problem of updating of an agent's knowledge. We propose a formal method of knowledge update on the basis of the semantics of modal logic S5. In our method, an update is specified according to the minimal change on both the agent's actual world and knowledge. We discuss general minimal change properties of knowledge update and show that our knowledge update operator satisfies all the update postulates of Katsuno and Mendelzon. We characterize several specific forms of knowledge update which have important applications in reasoning about change of agents' knowledge. We also examine the persistence property of knowledge and ignorance associated with knowledge update. We then investigate the computational complexity of model checking for knowledge update. We first show that in general the model checking for knowledge update is Σ 2 P -complete. We then identify a subclass of knowledge update problems that has polynomial time complexity for model checking. We point out that some important knowledge update problems belong to this subclass. We further address another interesting subclass of knowledge update problems for which the complexity of model checking is NP-complete.

IJCAI Conference 2005 Conference Paper

Solving Logic Program Conflict through Strong and Weak Forgettings

  • Yan Zhang
  • Norman Foo
  • Kewen

We consider how to forget a set of atoms in a logic program. Intuitively, when a set of atoms is forgotten from a logic program, all atoms in the set should be eliminated from this program in some way, and other atoms related to them in the program might also be affected. We define notions of strong and weak forgettings in logic programs to capture such intuition and reveal their close connections to the notion of forgetting in classical propositional theories. Based on these notions, we then propose a framework for conflict solving in logic programs, which is general enough to represent many important conflict solving problems. We also study some essential semantic and computational properties in relation to strong and weak forgettings and conflict solving in our framework.

KR Conference 2004 Conference Paper

Reasoning about Knowledge by Variable Forgetting

  • Guanfeng Lv
  • Kaile Su
  • Yan Zhang

In this paper, we investigate knowledge reasoning within a simple framework called knowledge structure. We use variable forgetting as a basic operation for one agent to reason about its own or other agents’ knowledge. In our framework, two notions namely agents’ observable variables and the weakest sufficient condition play important roles in knowledge reasoning. Given a background knowledge base and a set of observable variables for an agent, we show that the notion of agent knowing a formula can be defined as a weakest sufficient condition of the formula on the set of observable variables for the agent under background knowledge base. Moreover, we show how to capture the notion of common knowledge by using a generalized notion of weakest sufficient condition. We also discuss possible applications of our framework in some interesting domains such as the automated analysis of the well-known muddy children puzzle and the verification of the revised Needham-Schroeder protocol.

IJCAI Conference 2003 Conference Paper

Minimal Change and Maximal Coherence for Epistemic Logic Program Updates

  • Yan Zhang

We consider the problem of updating nonmonotonic knowledge bases represented by epistemic logic programs where disjunctive information and notions of knowledge and beliefs can be explicitly expressed. We propose a formulation for epistemic logic program updates based on a principle called minimal change and maximal coherence. The central feature of our approach is that during an update procedure, contradictory information is removed on a basis of minimal change under the semantics of epistemic logic programs and then coherent information is maximally retained in the update result. By using our approach, we can characterize an update result in both semantic and syntactic forms. We show that our approach handles update sequences and satisfies the consistency requirement. We also investigate important semantic properties of our update approach such as reduction, persistence and preservation.

TCS Journal 1999 Journal Article

Specifying causality in action theories: a default logic approach

  • Yan Zhang

Recent research on reasoning about action has shown that the traditional logic form of domain constraints is problematic to represent ramifications of actions that are related to causality of domains. To handle this problem properly, as proposed by some researchers, it is necessary to describe causal relations of domains explicitly in action theories. In this paper, we address this problem from a new point of view. Specifically, unlike other researchers viewing causal relations as some kind of inference rules, we distinguish causal relations between defeasible and non-defeasible cases. It turns out that a causal theory in our formalism can be specified by using Reiter's default logic. Based on this idea, we propose a causality-based minimal change approach for representing effects of actions, and argue that our approach provides more plausible solutions for the ramification and qualification problems compared with other related work. We also describe a logic programming approximation to compute causal theories of actions which provides an implementational basis for our approach.

IJCAI Conference 1997 Conference Paper

Towards Generalized Rule-based Updates

  • Yan Zhang
  • Norman Y Foo

Recent work on rule-based updates provided new frameworks for updates in more general knowledge domains [Marek and Truszczriski, 1994; Baral, 1994; Przymusinski and Turner, 1995]. In this paper, we consider a simple generalization of rule-based updates where incomplete knowledge bases are allowed and update rules may contain two types of negations. It turns out that previous methods cannot deal with this generalized rule-based update properly. To overcome the difficulty, we argue that necessary preferences between update rules and inertia rules must be taken into account in update specifications. From this motivation, we propose prioritized logic programs (PLPs) by adding preferences into extended logic programs [Gelfond and Lifschitz, 1991]. Formal semantics of PLPs is provided in terms of the answer set semantics of extended logic programs. We then show that the procedure of generalized rule-based update can be formalized in the framework of PLPs. The minimal change property of the update is also investigated.

AAAI Conference 1996 Conference Paper

Updating Knowledge Bases with Disjunctive Information

  • Yan Zhang

It is well known that the minimal change principle was widely used in knowledge base updates. However, recent research has shown that conventional minimal change methods, eg. the PMA (Winslett 1988), are generally problematic for updating knowledge bases with disjunctive information. In this paper, we propose two different approaches to deal with this problem - one is called the nae’ pzirnalcharage with excepqons (MCE), the other is called the miraimal charage with maximal disjunctive inclusions (MCD). The first method is syntax-based, while the second is modeltheoretic. We show that these two approaches are equivalent for propositional knowledge base updates, and the second method is also appropriate for first order knowledge base updates. We then prove that our new update approaches still satisfy the standard Katsuno and Mendelzon’ s update postulates.

IJCAI Conference 1993 Conference Paper

Reasoning About Persistence: A Theory of Actions

  • Yan Zhang
  • Norman Y. Foo

Winslett proposed a method for reasoning about action called the possible models approach (PMA). The PMA successfully removed the major difficulty manifested by Ginsberg and Smith's possible worlds approach (PWA). In this paper, we show that Winslett's PMA fails to solve the frame and ramification problems for some actions, as does the PWA. From this observation, we classify actions as definite and indefinite, and find that, in general, the PMA is not appropriate for both definite and indefinite actions. We propose a new approach to formalize actions based on persistence. We compare our approach with the PMA in detail, and show that our new formalization can avoid the problems in the PMA and PWA in most cases, and give more intuitive results for reasoning about action, regardless of whether the action is definite or indefinite.

v2026.09.13