Arrow Research search

Author name cluster

Zheng Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

99 papers
2 author rows

Possible papers

99

EAAI Journal 2026 Journal Article

A lightweight and attention-enhanced framework for robust pavement defect detection

  • Xiaoyan Li
  • Ning Zhang
  • Yue Pan
  • Yaowen Lv
  • Xiping Xu
  • Zheng Wang

Accurately detecting pavement anomalies, a critical task within structural health monitoring (SHM), is essential for infrastructure safety and automated monitoring systems. However, existing deep learning based object detectors, including state-of-the-art You Only Look Once (YOLO) variants, often struggle with defects such as elongated and low-contrast potholes due to irregular geometry and limited spatial context awareness. In this study, we propose an Efficient and CA enhanced You Only Look Once framework (EC-YOLO), an improved deep learning based object detection network designed to address these challenges. The proposed model builds upon the YOLOv11 architecture and introduces two major enhancements: (1) replacing the shallow backbone with EfficientNet-B0 for superior fine-grained feature extraction, and (2) integrating a Coordinate Attention (CA) module into the large-object detection head to capture long-range spatial dependencies. Extensive experiments on the Urban Digital Twins dataset demonstrate that EC-YOLO achieves state-of-the-art performance, attaining 96. 5% mean Average Precision (mAP)@0. 5 and 71. 3% mAP@0. 5: 0. 95. After deployment engine optimization, the model maintains real-time inference at 225. 4 frames per second (FPS) on an NVIDIA Jetson Orin Nano with only 1. 7 giga floating point operations (GFLOPs). Ablation studies further verify the contribution of each component. Moreover, EC-YOLO exhibits strong generalization by outperforming existing models on the Urban Digital Twins for Intelligent Road Inspection (UDTIRI) external benchmark. Overall, deployment verification on the Jetson platform confirms that EC-YOLO is a robust, lightweight, and effective solution for practical road defect inspection in resource-constrained environments.

EAAI Journal 2026 Journal Article

A metaheuristic-driven categorical boosting framework with interpretability for high-precision prediction of mechanical properties in corroded reinforced concrete beams

  • Yuzhuo Zhang
  • Zheng Wang
  • Jinlong Liu
  • Yalin Li
  • Zhenqin Huang
  • Xiaohu Yu

The degradation of mechanical properties in corroded reinforced concrete (RC) beams presents a major challenge for assessing structural durability. To address this issue, this study proposes an integrated machine learning (ML) framework to predict the mechanical properties of such beams. First, a database of 464 samples was established, including 12 input parameters and 2 output parameters, followed by correlation analysis of the inputs. On this basis, the applicability of existing design codes and empirical models was evaluated. Subsequently, eight ML models were trained, with their hyperparameters optimized via Bayesian optimization (BO) to enhance prediction accuracy. The Categorical Boosting (CatBoost) model was identified as the most accurate, and its hyperparameters were further optimized using Particle Swarm Optimization (PSO) and Genetic Algorithm (GA) for improved performance. Results show the PSO-optimized CatBoost model achieves the highest prediction accuracy to date: for the flexural strength test set, the coefficient of determination (R 2 ) is 0. 984 and root mean square error (RMSE) is 3. 1602; for the deflection test set, R 2 is 0. 975 and RMSE is 0. 6259. Compared with design codes, flexural strength test set R 2 increases by 27. 3 % and RMSE decreases by 72. 8 %; versus traditional models like Support Vector Regression (SVR), R 2 rises by 5. 4 % and RMSE drops by 43. 5 %. Additionally, SHapley Additive exPlanations (SHAP) analysis reveals geometric parameters (beam height, beam width) dominate flexural strength, while elastic stiffness and beam length drive deflection. Finally, a user-friendly graphical user interface (GUI) was developed for rapid mechanical property assessment of corroded RC beams.

JBHI Journal 2026 Journal Article

CAMM: Confidence-Aligned Multiview Multimodal Fusion for Brain Disorders Prediction With Imaging Transcriptomics

  • Haoran Luo
  • Zhoujie Fan
  • Wei Li
  • Hong Liang
  • Chen Jason Zhang
  • Xiaoyong Wei
  • Zheng Wang
  • Shan Cong

Brain disorder prediction can be enhanced by models that capture not only imaging phenotypes but also their underlying molecular context. Neuroimaging provides detailed structural and functional information, yet it offers limited insight into the gene-regulated processes driving these alterations. Transcriptomic atlases offer such molecular insights but are rarely available at the subject level due to invasive sampling. To address this gap, we propose CAMM, a confidence-aware multi-modal framework that integrates transcriptomic priors with imaging features to embed molecular context before fusion. CAMM further introduces a unified confidence calibration–regularization strategy that adapts modality contributions at the sample level, ensuring that information from high-confidence samples is leveraged to improve predictions for low-confidence samples, thereby enhancing robustness. Applied to large neuroimaging cohorts, CAMM consistently surpasses state-of-the-art baselines and identifies biologically meaningful biomarkers, demonstrating how transcriptomic priors can bridge molecular mechanisms and imaging for interpretable precision modeling of brain disorders.

AAAI Conference 2026 Conference Paper

D-GARA: A Dynamic Benchmarking Framework for GUI Agent Robustness in Real-World Anomalies

  • Sen Chen
  • Tong Zhao
  • Yi Bin
  • Fei Ma
  • Wenqi Shao
  • Zheng Wang

Developing intelligent agents capable of operating a wide range of Graphical User Interfaces (GUIs) with human-level proficiency is a key milestone on the path toward Artificial General Intelligence. While most existing datasets and benchmarks for training and evaluating GUI agents are static and idealized, failing to reflect the complexity and unpredictability of real-world environments, particularly the presence of anomalies. To bridge this research gap, we propose D-GARA, a dynamic benchmarking framework, to evaluate Android GUI agent robustness in real-world anomalies. D-GARA introduces a diverse set of real-world anomalies that GUI agents commonly face in practice, including interruptions such as permission dialogs, battery warnings, and update prompts. Based on D-GARA framework, we construct and annotate a benchmark featuring commonly used Android applications with embedded anomalies to support broader community research. Comprehensive experiments and results demonstrate substantial performance degradation in state-of-the-art GUI agents when exposed to anomaly-rich environments, highlighting the need for robustness-aware learning. D-GARA is modular and extensible, supporting the seamless integration of new tasks, anomaly types, and interaction scenarios to meet specific evaluation goals.

JBHI Journal 2026 Journal Article

Deep Learning-Based Vitiligo Activity Evaluation Using Wood's Lamp Imaging: A Clinical Decision Support

  • Zheng Wang
  • Zixuan Nie
  • Hui Hu
  • Kaibin Lin
  • Chong Wang
  • Hongyang Fu
  • Jianglin Zhang

Vitiligo is an autoimmune disorder characterized by heterogeneous and unpredictable depigmentation, which poses substantial challenges for objective disease monitoring and treatment evaluation in routine clinical practice. To address the lack of automated and quantitative follow-up tools, we designed and validated an end-to-end deep learning–based system using Wood's lamp imaging to support lesion localization, longitudinal tracking, and activity assessment. The framework employs Mask R-CNN for automated lesion detection and segmentation, followed by quantitative pigmentation-state analysis and disease activity evaluation using t distributed stochastic neighbor embedding (t SNE) and the Vitiligo Disease Activity (VIDA) score. Disease progression is further characterized through risk, correlation, and survival analyses. Quantitatively, the Mask R CNN model achieved robust performance, with a Dice coefficient of 90. 5%, a mean intersection over-union of 83. 1%, and an area under the curve of 0. 9567 on external validation, supporting reliable lesion delineation under Wood's lamp imaging. When benchmarked against representative segmentation backbones (U-Net, U-Net++, U2-Net, and DeepLabV3), the Mask R CNN–based pipeline yielded the highest mean IoU (83. 1 ± 3. 9%) while maintaining a competitive Dice score (90. 5 ± 1. 38%). t-SNE analysis effectively separated pigmentation states, revealing distinct patterns of hyperpigmentation, hypopigmentation, and apigmentation that were consistent with clinical activity. Longitudinal survival analysis demonstrated that stable lesions derived significant benefit from sustained treatment, whereas active lesions exhibited elevated early risk. Age-stratified analysis identified the highest relapse risk in patients aged 11–20 years, while patients aged 71–80 years showed reduced disease stability. Regression analysis further indicated that, during the stable phase, increased associated with hypopigmentation subsequent was associated with subsequent pigment regeneration, underscoring stage-dependent treatment effects. Integrating deep learning based lesion analysis with clinically grounded activity and longitudinal risk modeling enables precise, objective monitoring of vitiligo and supports personalized, stage-aware treatment decision-making.

AAAI Conference 2026 Conference Paper

Hyper-Opinion Vagueness Quantification for Robust Multimodal Learning

  • Disen Hu
  • Xun Jiang
  • Xiaofeng Cao
  • Zheng Wang
  • Jingkuan Song
  • Heng Tao Shen
  • Xing Xu

Robust Multimodal Learning (RML) aims to address the issues of unreliable predictions of multimodal models. Nevertheless, previous RML works often struggle to distinguish between different categories that rely on identical intra-modal cues, making ambiguous predictions. We defined this degree of ``uncertain'' in extracting discriminative features of a multimodal model as vagueness. Neglecting such vagueness, as previous RML works commonly do, will undermine the ability to extract unique semantics of each category in multimodal models, further resulting in worse robustness under disturbances that affect semantic representations. Additionally, this vagueness will lead the parameter updating processes towards unreliable fusion, thus diverting the learning processes of the multimodal model from learning unique features of each category. Based on the above insight, we propose a novel robust multimodal learning approach, termed Hyper-Opinion Quantifying Vagueness (HOQV). Specifically, we first introduce hyper-opinion to capture and quantify the vagueness of multimodal learning in discriminating representations of different categories. Moreover, to mitigate the interference in parameter updating of unreliable representations with high vagueness, we also design the Hyper-Opinion Gradient Modulation to guide the optimization processes. We evaluate our HOQV on six datasets with different disturbances, including noise and adversarial attack, and demonstrate that our proposed method achieves state-of-the-art performance consistently.

AAAI Conference 2026 Conference Paper

PMPGuard: Catching Pseudo-Matched Pairs in Remote Sensing Image–Text Retrieval

  • Pengxiang Ouyang
  • Qing Ma
  • Zheng Wang
  • Cong Bai

Remote sensing (RS) image–text retrieval faces significant challenges in real-world datasets due to the presence of Pseudo-Matched Pairs (PMPs), semantically mismatched or weakly aligned image–text pairs, which hinder the learning of reliable cross-modal alignments. To address this issue, we propose a novel retrieval framework that leverages Cross-Modal Gated Attention and a Positive–Negative Awareness Attention mechanism to mitigate the impact of such noisy associations. The gated module dynamically regulates cross-modal information flow, while the awareness mechanism explicitly distinguishes informative (positive) cues from misleading (negative) ones during alignment learning. Extensive experiments on three benchmark RS datasets, i.e., RSICD, RSITMD, and RS5M, demonstrate that our method consistently achieves state-of-the-art performance, highlighting its robustness and effectiveness in handling real-world mismatches and PMPs in RS image–text retrieval tasks.

AAAI Conference 2026 Conference Paper

Reliable-View 2D-3D Key-Part Aligned Transformer with Reinforced Masking for 3D Point Cloud Understanding

  • Xianglong Jin
  • Zheng Wang
  • Rong Wang
  • Feiping Nie

Self-supervised 3D point cloud understanding is crucial for scene understanding, where Masked Autoencoders (MAE) have achieved excellent performance in point cloud representation learning. However, existing MAE-style methods fail to consider spatial-semantic variations in masking strategies, and joint learning with multi-view images often overlooks view redundancy. To address these challenges, we propose an MAE framework enhanced with reliable multi-view 2D-3D Key-part alignment and Reinforced masking, named as KR-MAE. Our approach comprises three key innovations: Reinforced Masking (RM) strategically samples visible tokens based on semantic saliency to enhance reconstruction fidelity; Reliable Multi-View Selector (RVS) dynamically refines the most informative image subset by filtering occluded or low-texture views, mitigating detrimental redundancy; Reliable-view 2D-3D Key-part Aligned Transformer (KAT) establishes semantic-aligned correspondence between salient 3D point cloud parts and reliable multi-view 2D image patches, leveraging rich texture cues from 2D images to compensate for sparse geometry in point cloud. Extensive experiments on 3D classification and segmentation benchmarks demonstrate that KR-MAE achieves state-of-the-art performance, surpassing prior multi-modal methods.

AAAI Conference 2026 Conference Paper

S2-Boost: Synergistic Semantic Boosting for Coarse-to-Fine Ensemble Learning

  • Guanxiong He
  • Zheng Wang
  • Jie Wang
  • Liaoyuan Tang
  • Rong Wang
  • Feiping Nie

Neuroscientific evidence reveals that human visual recognition is not an instantaneous event but a hierarchical process, where the brain constructs a holistic perception by progressively integrating simple features like edges or texture into complex scenes. Ensemble learning successfully utilizes this principle, yet existing methods typically integrate models at the decision level, neglecting the rich, complementary information within the feature space itself and thus fundamentally limiting their potential. To address this, we introduce Synergistic Semantic Boosting (S2-Boosting), a framework that employs a self-supervised hierarchical semantic learning module to decompose an image into complementary, semantically meaningful parts autonomously. These parts guide a boosting procedure where a sequence of specialized learners, each focusing on a specific semantic partition, collaboratively corrects the ensemble's errors. We further present encouraging results on real-world image datasets, highlighting the intrinsic interpretability, paving the way for more robust and transparent models.

AAAI Conference 2026 Conference Paper

SimROD: A Simple Baseline for Raw Object Detection with Global and Local Enhancements

  • Haiyang Xie
  • Xi Shen
  • Shihua Huang
  • Qirui Wang
  • Zheng Wang

Most visual models are designed for sRGB images, yet RAW data offers significant advantages for object detection by preserving sensor information before ISP processing. This enables improved detection accuracy and more efficient hardware designs by bypassing the ISP. However, RAW object detection is challenging due to limited training data, unbalanced pixel distributions, and sensor noise. To address this, we propose SimROD, a lightweight and effective approach for RAW object detection. We introduce a Global Gamma Enhancement (GGE) module, which applies a learnable global gamma transformation with only four parameters, improving feature representation while keeping the model efficient. Additionally, we leverage the green channel's richer signal to enhance local details, aligning with the human eye’s sensitivity and Bayer filter design. Extensive experiments on multiple RAW object detection datasets and detectors demonstrate that SimROD outperforms state-of-the-art methods like RAW-Adapter and DIAP while maintaining efficiency. Our work highlights the potential of RAW data for real-world object detection.

AAAI Conference 2026 Conference Paper

Towards Federated Clustering: A Client-wise Private Graph Aggregation Framework

  • Guanxiong He
  • Zheng Wang
  • Jie Wang
  • Liaoyuan Tang
  • Rong Wang
  • Feiping Nie

Federated clustering addresses the critical challenge of extracting patterns from decentralized, unlabeled data. However, it is hampered by the flaw that current approaches are forced to accept a compromise between performance and privacy: transmitting embedding representations risks sensitive data leakage, while sharing only abstract cluster prototypes leads to diminished model accuracy. To resolve this dilemma, we propose Structural Privacy-Preserving Federated Graph Clustering (SPP-FGC), a novel algorithm that innovatively leverages local structural graphs as the primary medium for privacy-preserving knowledge sharing, thus moving beyond the limitations of conventional techniques. Our framework operates on a clear client-server logic; on the client-side, each participant constructs a private structural graph that captures intrinsic data relationships, which the server then securely aggregates and aligns to form a comprehensive global graph from which a unified clustering structure is derived. The framework offers two distinct modes to suit different needs. SPP-FGC is designed as an efficient one-shot method that completes its task in a single communication round, ideal for rapid analysis. For more complex, unstructured data like images, SPP-FGC+ employs an iterative process where clients and the server collaboratively refine feature representations to achieve superior downstream performance. Extensive experiments demonstrate that our framework achieves state-of-the-art performance, improving clustering accuracy by up to 10% (NMI) over federated baselines while maintaining provable privacy guarantees.

IROS Conference 2025 Conference Paper

A New Unsupervised Infrared and Visible Image Fusion Method Based on Salient Object Segmentation under Poor Illumination

  • Zheng Wang
  • Haifeng Ji
  • Baoliang Wang
  • Zhiyao Huang

This work aims to propose a new unsupervised infrared and visible image fusion method based on salient object segmentation, which can obtain a fused image with more information on salient object and realize the salient object segmentation under poor illumination. The new method can be divided into four steps: (1) A new superpixel segmentation method based on simple linear iterative clustering (SLIC) with K-means subdivision is used to initially process the infrared and visible image, which has better superpixel segmentation quality. (2) A new improved Density Peaks Clustering (DPC) based on superpixel is used to realize the salient object segmentation of the infrared image, which is improved to be automatically selecting the cluster centers with less computation cost. (3) A new GrabCut strategy using the eroded and dilated salient object regions of the infrared image to predetermine the foreground and background respectively is used to achieve the salient object segmentation of the visible image, which can be totally automatic with better salient object segmentation quality. (4) An image fusion strategy is used to realize the final image fusion, which treats the salient object region and background respectively. Experiments were carried out under different poor illumination scenes in the real world. The experimental results show that the new infrared and visible image fusion method is successful with Q AB/F greater than 0. 69. In addition, the provided superpixel segmentation method, salient object segmentation method and new GrabCut strategy are also effective. The research results provide an effective infrared and visible image fusion thought and three useful methods, which can provide a good reference for researchers. And, the research work reveals the application potential of DPC on image fusion and salient object segmentation, and broadens the application fields of DPC.

AAAI Conference 2025 Conference Paper

Balancing Privacy and Performance: A Many-in-One Approach for Image Anonymization

  • Xuemei Jia
  • Jiawei Du
  • Hui Wei
  • Ruinian Xue
  • Zheng Wang
  • Hongyuan Zhu
  • Jun Chen

The effective utilization of data through Deep Neural Networks (DNNs) has profoundly influenced various aspects of society. The growing demand for high-quality, particularly personalized, data has spurred research efforts to prevent data leakage and protect privacy in recent years. Early privacy-preserving methods primarily relied on instance-wise modifications, such as erasing or obfuscating essential features for de-identification. However, this approach highlights an inherent trade-off: minimal modification offers insufficient privacy protection, while excessive modification significantly degrades task performance. In this paper, we propose a novel Recombining for Obfuscation (FRO) approach to address this trade-off. Unlike existing methods that generate one anonymized instance by perturbing the original data on a one-to-one basis, our FRO approach generates an anonymized instance by reassembling mixed ID-related features from multiple original data sources on a many-in-one basis. Instead of introducing additional noise for de-identification, our approach leverages the existing non-polluted features from other instances to anonymize data. Extensive experiments on identity identification tasks demonstrate that FRO outperforms previous state-of-the-art methods, not only in utility performance but also in visual anonymization.

AIIM Journal 2025 Journal Article

CATI: A medical context-enhanced framework for diagnosis code assignment in the UK Biobank study

  • Yue Shen
  • Jie Wang
  • Zhe Wang
  • Zhihao Shi
  • Hanzhu Chen
  • Zheng Wang
  • Yukang Jiang
  • Xiaopu Wang

Diagnosis codes are standard code format of diseases or medical conditions. This study is aimed at assigning diagnosis codes to patients in large-scale biobanks, particularly addressing the issue of missing codes for some patients. This is crucial for downstream disease-related tasks. While recent methods primarily rely on structured biobank data for code assignment, they often overlook the valuable medical context provided by textual information in the biobanks and hierarchical structure of the disease coding system. To address this gap, we have developed CATI, a medical context-enhanced framework for diagnosis Code Assignment by integrating Textual details derived from key features and disease hIerarchy. The study is based on the UK Biobank data and considers Phecodes and ICD-10 codes as standard disease formats. We start by representing ten informative codified features using their formal names and then integrate them into CATI as text embeddings, achieved through prompt tuning on the pre-trained language model BioBERT. Recognizing the hierarchical structure of diagnosis codes, we have developed a novel convolution layer in our method that effectively propagates logits between adjacent diagnosis codes. Evaluation results demonstrate that CATI outperforms existing state-of-the-art methods in terms of both Phecodes and ICD-10 codes, boasting at least a 5. 16% improvement in average AUROC for unseen disease codes and an 8. 68% rise in average AUPRC for disease codes with training instances ranging in (1000, 10000]. This framework contributes to the formation of well-defined cohorts for downstream studies and offers a unique perspective for addressing complex healthcare tasks by incorporating vital medical context.

NeurIPS Conference 2025 Conference Paper

Controlling Thinking Speed in Reasoning Models

  • Zhengkai Lin
  • Zhihang Fu
  • Ze Chen
  • Chao Chen
  • Liang Xie
  • Wenxiao Wang
  • Deng Cai
  • Zheng Wang

Human cognition is theorized to operate in two modes: fast, intuitive System 1 thinking and slow, deliberate System 2 thinking. While current Large Reasoning Models (LRMs) excel at System 2 thinking, their inability to perform fast thinking leads to high computational overhead and latency. In this work, we enable LRMs to approximate human intelligence through dynamic thinking speed adjustment, optimizing accuracy-efficiency trade-offs. Our approach addresses two key questions: (1) how to control thinking speed in LRMs, and (2) when to adjust it for optimal performance. For the first question, we identify the steering vector that governs slow-fast thinking transitions in LRMs' representation space. Using this vector, we achieve the first representation editing-based test-time scaling effect, outperforming existing prompt-based scaling methods. For the second question, we apply real-time difficulty estimation to signal reasoning segments of varying complexity. Combining these techniques, we propose the first reasoning strategy that enables fast processing of easy steps and deeper analysis for complex reasoning. Without any training or additional cost, our plug-and-play method yields an average +1. 3\% accuracy with -8. 6\% token usage across leading LRMs and advanced reasoning benchmarks. All of our algorithms are implemented based on vLLM and are expected to support broader applications and inspire future research.

AAAI Conference 2025 Conference Paper

Effects of Momentum in Implicit Bias of Gradient Flow for Diagonal Linear Networks

  • Bochen Lyu
  • He Wang
  • Zheng Wang
  • Zhanxing Zhu

This paper targets on the regularization effect of momentum-based methods in regression settings and analyzes the popular diagonal linear networks to precisely characterize the implicit bias of continuous versions of heavy-ball (HB) and Nesterov's method of accelerated gradients (NAG). We show that, HB and NAG exhibit different implicit bias compared to GD for diagonal linear networks, which is different from the one for classic linear regression problem where momentum-based methods share the same implicit bias with GD. Specifically, the role of momentum in the implicit bias of GD is twofold: (a) HB and NAG induce extra initialization mitigation effects similar to SGD that are beneficial for generalization of sparse regression; (b) the implicit regularization effects of HB and NAG also depend on the initialization of gradients explicitly, which may not be benign for generalization. As a result, whether HB and NAG have better generalization properties than GD jointly depends on the aforementioned twofold effects determined by various parameters such as learning rate, momentum factor, and integral of gradients. Our findings highlight the potential beneficial role of momentum and can help understand its advantages in practice such as when it will lead to better generalization performance.

IROS Conference 2025 Conference Paper

Fast Policy: Accelerating Visuomotor Policies without Re-training

  • Tongshu Wu
  • Zheng Wang

Diffusion models are increasingly employed in visuomotor policies to achieve promising performance of behavior cloning. However, the slow inference caused by iterative denoising is a notorious disadvantage, which greatly limits its application in resource-limited and real-time interactive robot systems. The prevailing strategy to this problem is distillation, but it still requires considerable resources to retrain a student model. To this end, we take another training-free view to develop a novel Fast Policy (termed FP), which can be regarded as a powerful and accelerated alternative to Diffusion Policy for learning visuomotor robot control. Specifically, our comprehensive study of UNet encoder shows that its features change little during inference, prompting us to reuse encoder features in non-critical denoising steps. In addition, we design strategies based on Fourier energy to screen critical and non-critical steps dynamically according to different tasks. Importantly, to mitigate performance degradation caused by the repeated use of non-critical steps, we further introduce a noise correction strategy. Our FP is evaluated on multiple simulation benchmarks and the comparison results with existing speed-up methods demonstrate our effectiveness and superiority with state-of-the-art success rates in visuomotor inference speed. The code is available at https://github.com/xwccchong/Fast-Policy

NeurIPS Conference 2025 Conference Paper

Heavy-Ball Momentum Method in Continuous Time and Discretization Error Analysis

  • Bochen Lyu
  • Xiaojing Zhang
  • Fangyi Zheng
  • He Wang
  • Zheng Wang
  • Zhanxing Zhu

This paper establishes a continuous time approximation, a piece-wise continuous differential equation, for the discrete Heavy-Ball (HB) momentum method with explicit discretization error. Investigating continuous differential equations has been a promising approach for studying the discrete optimization methods. Despite the crucial role of momentum in gradient-based optimization methods, the gap between the original dynamics and the continuous time approximations due to the discretization error has not been comprehensively bridged yet. In this work, we study the HB momentum method in continuous time while putting more focus on the discretization error to provide additional theoretical tools to this area. In particular, we design a first-order piece-wise continuous differential equation, where we add a number of counter terms to account for the discretization error explicitly. As a result, we provide a continuous time model for the HB momentum method that allows the control of discretization error to arbitrary order of the learning rate. As an application, we leverage it to find a new implicit regularization of the directional smoothness and investigate the implicit bias of HB for diagonal linear networks, indicating how our results can be used in deep learning. Our theoretical findings are further supported by numerical experiments.

AAAI Conference 2025 Conference Paper

Language Pre-training Guided Masking Representation Learning for Time Series Classification

  • Liaoyuan Tang
  • Zheng Wang
  • Jie Wang
  • Guanxiong He
  • Zhezheng Hao
  • Rong Wang
  • Feiping Nie

The representation learning of time series has a wide range of downstream tasks and applications in many practical scenarios. However, due to the complexity, spatiotemporality, and continuity of sequential stream data, compared with the representation learning of structural data such as images/videos, the time series self-supervised representation learning is even more challenging. Besides, the direct application of existing contrastive learning and masked autoencoder based approaches to time series representation learning encounters inherent theoretical limitations, such as ineffective augmentation and masking strategies. To this end, we propose a Language Pre-training guided Masking Representation Learning (LPMRL) for times series classification. Specifically, we first propose a novel language pre-training guided masking encoder for adaptively sampling semantic spatiotemporal patches via natural language descriptions and improving the discriminability of latent representations. Furthermore, we present the dual-information contrastive learning mechanism to explore both local and global information by meticulously designing high-quality hard negative samples of time series data samples. As a result, we also design various experiments, such as visualization of masking position and distribution and reconstruction error to verify the reasonability of proposed language guided masking technique. Last, we evaluate the performance of proposed representation learning via classification task conducted on 106 time series datasets, which demonstrates the effectiveness of proposed method.

NeurIPS Conference 2025 Conference Paper

Less is More: Improving LLM Alignment via Preference Data Selection

  • Xun Deng
  • Han Zhong
  • Rui Ai
  • Fuli Feng
  • Zheng Wang
  • Xiangnan He

Direct Preference Optimization (DPO) has emerged as a promising approach for aligning large language models with human preferences. While prior work mainly extends DPO from the aspect of the objective function, we instead improve DPO from the largely overlooked but critical aspect of data selection. Specifically, we address the issue of parameter shrinkage caused by noisy data by proposing a novel margin-maximization principle for dataset curation in DPO training. To further mitigate the noise in different reward models, we propose a Bayesian Aggregation approach that unifies multiple margin sources (external and implicit) into a single preference probability. Extensive experiments in diverse settings demonstrate the consistently high data efficiency of our approach. Remarkably, by using just 10\% of the Ultrafeedback dataset, our approach achieves 3\% to 8\% improvements across various Llama, Mistral, and Qwen models on the AlpacaEval2 benchmark. Furthermore, our approach seamlessly extends to iterative DPO, yielding a roughly 3\% improvement with 25\% online data, revealing the high redundancy in this presumed high-quality data construction manner. These results highlight the potential of data selection strategies for advancing preference optimization.

IJCAI Conference 2025 Conference Paper

Optimizing Personalized Federated Learning Through Adaptive Layer-Wise Learning

  • Weihang Chen
  • Cheng Yang
  • Jie Ren
  • Zhiqiang Li
  • Zheng Wang

Real-life deployment of federated Learning (FL) often faces non-IID data, which leads to poor accuracy and slow convergence. Personalized FL (pFL) tackles these issues by tailoring local models to individual data sources and using weighted aggregation methods for client-specific learning. However, existing pFL methods often fail to provide each local model with global knowledge on demand while maintaining low computational overhead. Additionally, local models tend to over-personalize their data during the training process, potentially dropping previously acquired global information. We propose FLAYER, a novel layer-wise learning method for pFL that optimizes local model personalization performance. FLAYER considers the different roles and learning abilities of neural network layers of individual local models. It incorporates global information for each local model as needed to initialize the local model cost-effectively. It then dynamically adjusts learning rates for each layer during local training, optimizing the personalized learning process for each local model while preserving global knowledge. Additionally, to enhance global representation in pFL, FLAYER selectively uploads parameters for global aggregation in a layer-wise manner. We evaluate FLAYER on four representative datasets in computer vision and natural language processing domains. Compared to eight state-of-the-art pFL methods, FLAYER improves the inference accuracy, on average, by 5. 20% (up to 14. 29%). Code is available at https: //github. com/lancasterJie/FLAYER/.

AAAI Conference 2025 Conference Paper

Pioneering Explainable Video Fact-Checking with a New Dataset and Multi-role Multimodal Model Approach

  • Kaipeng Niu
  • Danni Xu
  • Bingjian Yang
  • Wenxuan Liu
  • Zheng Wang

Existing video fact-checking datasets often lack detailed evidence and explanations, compromising the reliability and interpretability of fact-checking methods. To address these gaps, we developed a novel dataset featuring comprehensive annotations for each news item, including veracity labels, the rationales behind these labels, and supporting evidence. This dataset significantly enhances models' ability to accurately identify and explain video content. We also present an explainable automatic framework 3MFact, utilizing Multi-role Multimodal Models for video Fact-checking. Our framework iteratively gathers and synthesizes online evidence to progressively determine the veracity label, generating three key outputs: veracity label, rationale, and supported evidence. We aim for this work to be a pioneering effort, providing robust support for the field of video fact-checking.

IROS Conference 2025 Conference Paper

PneuChip: A Compact Pneumatic Controller for Large-scale Soft Artificial Muscles

  • Zheng Wang
  • Zhe Liu
  • Yimo Wang
  • Hongying Zhang

Pneumatic soft actuators are known for their versatility and reliability; however, their control presents a major challenge as systems scale beyond tens of actuators. Traditional rigid pneumatic valves add bulk, weight, and complexity, while most soft valves fail to generate programmable independent output states. We propose PneuChip, a compact pneumatic controller designed for large-scale soft actuators. The PneuChip functions as a two-dimensional array with rows and columns, controlled by m + n input signals to generate ${2^{m + n}} - {2^m} - {2^n} + 2$ distinct output states, enabling programmable control over m × n soft actuators. To validate its effectiveness, we implemented PneuChip in a muscular-skeletal robotic arm comprising 24 Miura-Ori inspired, negative pressure-actuated artificial muscles and a rigid two-link skeleton connected by a ball joint. A 4×6 PneuChip was fabricated and integrated to control the arm’s 3 degree of freedoms (DOFs) motion. Within a 120° rotation range, the robot arm achieved 946 distinct positions with smooth state transitions, paving the way for future applications in trajectory tracking and dexterous manipulation. The compact design and high controllability of PneuChip promise to notably simplify complex pneumatic systems, significantly enhancing the practicality of large-scale soft robots for various applications.

EAAI Journal 2025 Journal Article

Progressive Enhancement Dehazing for object detection in extreme weather

  • Zhiying Li
  • Junhao Wu
  • Shuyuan Lin
  • Zheng Wang
  • Xiaobo Jin
  • Guanggang Geng
  • Feiran Huang
  • Jian Weng

Existing object detection technologies have made significant advances in perception systems for autonomous driving. However, accurately detecting objects in extreme weather conditions, especially fog, remains a significant challenge. On the one hand, current methods struggle to balance image enhancement and object detection, leading to the neglect of essential information that could improve detection accuracy. On the other hand, the lack of mature datasets in this area limits many studies to synthetic fog data generated from the prior-based atmospheric scattering model, which inevitably restricts further performance improvements. To address these challenges, we propose the Progressive Enhancement Dehazing You Only Look Once (PED-YOLO) method, which processes images in real time under foggy conditions to improve object detection. We design a novel progressive image processing module that follows a unique paradigm of progressive supervised learning, which gradually processes the image from small size to large size, to effectively overcome the challenge of processing large-sized images at once. Moreover, we develop a small convolutional network module that focuses on each channel of the image and enables more accurate adaptive prediction of the parameters of the filters. In addition, we develop a novel style transfer model to generate simulated fog images that more closely resemble real fog images and train them together to bring the synthetic fog domain closer to the real fog domain and improve the generalizability. We evaluate our method extensively on several popular datasets, and the experimental results show the superior performance of PED-YOLO, highlighting its potential to advance autonomous driving.

AAAI Conference 2025 Conference Paper

ProtCLIP: Function-Informed Protein Multi-Modal Learning

  • Hanjing Zhou
  • Mingze Yin
  • Wei Wu
  • Mingyang Li
  • Kun Fu
  • Jintai Chen
  • Jian Wu
  • Zheng Wang

Multi-modality pre-training paradigm that aligns protein sequences and biological descriptions has learned general protein representations and achieved promising performance in various downstream applications. However, these works were still unable to replicate the extraordinary success of language-supervised visual foundation models due to the ineffective usage of aligned protein-text paired data and the lack of an effective function-informed pre-training paradigm. To address these issues, this paper curates a large-scale protein-text paired dataset called ProtAnno with a property-driven sampling strategy, and introduces a novel function-informed protein pre-training paradigm. Specifically, the sampling strategy determines selecting probability based on the sample confidence and property coverage, balancing the data quality and data quantity in face of large-scale noisy data. Furthermore, motivated by significance of the protein specific functional mechanism, the proposed paradigm explicitly model protein static and dynamic functional segments by two segment-wise pre-training objectives, injecting fine-grained information in a function-informed manner. Leveraging all these innovations, we develop ProtCLIP, a multi-modality foundation model that comprehensively represents function-aware protein embeddings. On 22 different protein benchmarks within 5 types, including protein functionality classification, mutation effect prediction, cross-modal transformation, semantic similarity inference and protein-protein interaction prediction, our ProtCLIP consistently achieves SOTA performance, with remarkable improvements of 75% on average in five cross-modal transformation benchmarks, 59.9% in GO-CC and 39.7% in GO-BP protein function prediction. The experimental results verify the extraordinary potential of ProtCLIP serving as the protein multi-modality foundation model.

AAAI Conference 2025 Conference Paper

Rethinking Cancer Gene Identification Through Graph Anomaly Analysis

  • Yilong Zang
  • Lingfei Ren
  • Yue Li
  • Zhikang Wang
  • David Antony Selby
  • Zheng Wang
  • Sebastian Josef Vollmer
  • Hongzhi Yin

Graph neural networks (GNNs) have shown promise in integrating protein-protein interaction (PPI) networks for identifying cancer genes in recent studies. However, due to the insufficient modeling of the biological information in PPI networks, more faithfully depiction of complex protein interaction patterns for cancer genes within the graph structure remains largely unexplored. This study takes a pioneering step toward bridging biological anomalies in protein interactions caused by cancer genes to statistical graph anomaly. We find a unique graph anomaly exhibited by cancer genes, namely weight heterogeneity, which manifests as significantly higher variance in edge weights of cancer gene nodes within the graph. Additionally, from the spectral perspective, we demonstrate that the weight heterogeneity could lead to the "flattening out" of spectral energy, with a concentration towards the extremes of the spectrum. Building on these insights, we propose the HIerarchical-Perspective Graph Neural Network (HIPGNN) that not only determines spectral energy distribution variations on the spectral perspective, but also perceives detailed protein interaction context on the spatial perspective. Extensive experiments are conducted on two reprocessed datasets STRINGdb and CPDB, and the experimental results demonstrate the superiority of HIPGNN.

AAAI Conference 2025 Conference Paper

Synergy of GFlowNet and Protein Language Model Makes a Diverse Antibody Designer

  • Mingze Yin
  • Hanjing Zhou
  • Yiheng Zhu
  • Jialu Wu
  • Wei Wu
  • Mingyang Li
  • Kun Fu
  • Zheng Wang

Antibodies defend our health by binding to antigens with high specificity and potentiality, primarily relying on the Complementarity-Determining Region (CDR). Yet, current experimental methods of discovering new antibody CDRs are heavily time-consuming. Computational design could alleviate this burden; especially, protein language models have proven quite beneficial in many recent studies. However, most existing models solely focus on antibody potentiality and struggle to encapsulate the diverse range of plausible CDR candidates, limiting their effectiveness in real-world scenarios as binding is only one factor in the multitude of drug-forming criteria. In this paper, we introduce PG-AbD, a framework uniting Generative Flow Networks (GFlowNets) and pretrained Protein Language Models (PLMs) to successfully generate highly potent, diverse and novel antibody candidates. We innovatively construct a Products of Experts (PoE) composed by the global-distribution-modeling PLM and the local-distribution-modeling Potts Model to serve as the reward function of GFlowNet. The joint training paradigm is introduced, where PoE is trained by contrastive divergence with the negative samples generated by GFlowNet, and then guides GFlowNet to sample diverse antibody candidates. We evaluate PG-AbD on extensive antibody design benchmarks. It significantly outperforms existing methods in diversity (13.5% on RabDab, 31.1% on SabDab) while maintaining optimal potential and novelty. Generated antibodies are also found to form stable, regular 3D structures with their corresponding antigens, demonstrating the great potential of PG-AbD to accelerate real-world antibody discovery.

AAAI Conference 2025 Conference Paper

The Parables of the Mustard Seed and the Yeast: Extremely Low-Budget, High-Performance Nighttime Semantic Segmentation

  • Shiqin Wang
  • Xin Xu
  • Haoyang Chen
  • Kui Jiang
  • Zheng Wang

Nighttime Semantic Segmentation (NSS) is essential to many cutting-edge vision applications. However, existing technologies overly rely on massive labeled data, whose annotation is time-consuming and laborious. In this paper, we pioneer a new task focusing on exploring the potential of training strategy and framework design with limited annotation to achieve high-performance NSS. Insufficient information at very low labeling budgets can easily lead to under-optimization or overfitting of the model. Our solution comprises two main components: i) a novel region-based active sampling strategy called Contextual-Aware Region Query (CARQ), which identifies highly informative target nighttime regions for labeling; and ii) an innovative Fragmentation Synergy Active Domain Adaptation framework (FS-ADA), which progressively broadcasts the limited annotation to the unlabeled regions, achieving high performance with a minimal annotation budget. Extensive experiments demonstrate that our method outperforms state-of-the-art UDA-NSS & ADA-SS methods across four day-to-nighttime benchmarks, and generalizes well to foggy, rainy, & snowy scenes. In particular only with 1% target nighttime data annotation, our method is on par with the mainstream fully-supervised methods on the BDD100K-Night val dataset.

AAAI Conference 2025 Conference Paper

TokenMatcher: Diverse Tokens Matching for Unsupervised Visible-Infrared Person Re-Identification

  • Xiao Wang
  • Lekai Liu
  • Bin Yang
  • Mang Ye
  • Zheng Wang
  • Xin Xu

Unsupervised visible-infrared person re-identification (US-VI-ReID) seeks to match infrared and visible images of the same individual without the use of annotations. Current methods typically derive cross-modal correspondences through a single global feature matching process for generating pseudo labels and learning modality-invariant features. However, this matching approach is hindered by both intra-modality and inter-modality discrepancies, which result in imprecise measurements. As a consequence, the clustering of individuals with single global feature is often incomplete and unreliable, leading to suboptimal performance in cross-modal clustering tasks. To address these challenges and to extract cross-modality discriminative identity information, we propose a TokenMatcher, which encompasses three key components: Diverse Tokens Matching (DTM), Diverse Tokens Neighbor Learning (DTNL), and the Homogeneous Fusion (HF) Module. DTM utilizes multiple class tokens within the visual transformer framework to capture diverse embedding representations, thereby facilitating the integration of fine-grained information essential for reliable cross-modality correspondences. DTNL enhances the intra-modality and inter-modality consistency among diverse tokens by refining neighborhood sets with insights from neighboring tokens and camera information, promoting robust neighborhood learning and fostering discriminative identity information. Additionally, the HF module consolidates clusters of the same identity while effectively separating those of different identities. Extensive experiments conducted on the publicly available SYSU-MM01 and RegDB datasets demonstrate the efficacy of the proposed method.

AAAI Conference 2025 Conference Paper

VEGAS: Towards Visually Explainable and Grounded Artificial Social Intelligence

  • Hao Li
  • Hao Fei
  • Zechao Hu
  • Zhengwei Yang
  • Zheng Wang

Social Intelligence Queries (Social-IQ) serve as the primary multimodal benchmark for evaluating a model’s social intelligence level. While impressive multiple-choice question (MCQ) accuracy is achieved by current solutions, increasing evidence shows that they are largely, and in some cases entirely, dependent on language modality, overlooking visual context. Additionally, the closed-set nature further prevents the exploration of whether and to what extent the reasoning path behind selection is correct. To address these limitations, we propose the Visually Explainable and Grounded Artificial Social Intelligence (VEGAS) model. As a generative multimodal model, VEGAS leverages open-ended answering to provide explainable responses, which enhances the clarity and evaluation of reasoning paths. To enable visually grounded answering, we propose a novel sampling strategy to provide the model with more relevant visual frames. We then enhance the model’s interpretation of these frames through Generalist Instruction Fine-Tuning (GIFT), which aims to: i) learn multimodal language transformations for fundamental emotional social traits, and ii) establish multimodal joint reasoning capabilities. Extensive experiments, comprising modality ablation, open-ended assessments, and supervised MCQ evaluations, consistently show that VEGAS effectively utilizes visual information in reasoning to produce correct and also credible answers. We expect this work to offer a new perspective on Social-IQ and advance the development of human-like social AI.

NeurIPS Conference 2025 Conference Paper

Yggdrasil: Bridging Dynamic Speculation and Static Runtime for Latency-Optimal Tree-Based LLM Decoding

  • Yue Guan
  • Changming Yu
  • Shihan Fang
  • Weiming Hu
  • Zaifeng Pan
  • Zheng Wang
  • Zihan Liu
  • Yangjie Zhou

Speculative decoding improves LLM inference by generating and verifying multiple tokens in parallel, but existing systems suffer from suboptimal performance due to a mismatch between dynamic speculation and static runtime assumptions. We present Yggdrasil, a co-designed system that enables latency-optimal speculative decoding through context-aware tree drafting and compiler-friendly execution. Yggdrasil introduces an equal-growth tree structure for static graph compatibility, a latency-aware optimization objective for draft selection, and stage-based scheduling to reduce overhead. Yggdrasil supports unmodified LLMs and achieves up to $3. 98\times$ speedup over state-of-the-art baselines across multiple hardware setups.

NeurIPS Conference 2024 Conference Paper

Bridge-IF: Learning Inverse Protein Folding with Markov Bridges

  • Yiheng Zhu
  • Jialu Wu
  • Qiuyi Li
  • Jiahuan Yan
  • Mingze Yin
  • Wei Wu
  • Mingyang Li
  • Jieping Ye

Inverse protein folding is a fundamental task in computational protein design, which aims to design protein sequences that fold into the desired backbone structures. While the development of machine learning algorithms for this task has seen significant success, the prevailing approaches, which predominantly employ a discriminative formulation, frequently encounter the error accumulation issue and often fail to capture the extensive variety of plausible sequences. To fill these gaps, we propose Bridge-IF, a generative diffusion bridge model for inverse folding, which is designed to learn the probabilistic dependency between the distributions of backbone structures and protein sequences. Specifically, we harness an expressive structure encoder to propose a discrete, informative prior derived from structures, and establish a Markov bridge to connect this prior with native sequences. During the inference stage, Bridge-IF progressively refines the prior sequence, culminating in a more plausible design. Moreover, we introduce a reparameterization perspective on Markov bridge models, from which we derive a simplified loss function that facilitates more effective training. We also modulate protein language models (PLMs) with structural conditions to precisely approximate the Markov bridge process, thereby significantly enhancing generation performance while maintaining parameter-efficient training. Extensive experiments on well-established benchmarks demonstrate that Bridge-IF predominantly surpasses existing baselines in sequence recovery and excels in the design of plausible proteins with high foldability. The code is available at https: //github. com/violet-sto/Bridge-IF.

AILAW Journal 2024 Journal Article

Causality-inspired legal provision selection with large language model-based explanation

  • Zheng Wang
  • Yuanzhi Ding
  • Caiyuan Wu
  • Yuzhen Guo
  • Wei Zhou

Abstract Accurate identification of legal provisions is crucial for adjudicating criminal cases, but the complexity and volume of legal texts pose significant challenges for legal professionals. This paper addresses these challenges by introducing a novel legal provision selection framework that transforms the task from a simple classification problem into a sophisticated system combining semantic matching with causal relationship learning. Leveraging large language models, our approach enhances the understanding and interpretation of legal language, by extracting nuanced features from legal texts for deeper contextual comprehension. Additionally, integrating causal learning aligns with the inherent causality in legal reasoning, improving model interpretability and mitigating data bias. Our method demonstrates superior accuracy and robustness through extensive experiments on the CAIL2018 dataset and its subsets. This research significantly advances legal AI applications, promoting efficiency and fairness in the criminal justice system by providing precise and reliable legal provision selection.

AAAI Conference 2024 Conference Paper

Contributing Dimension Structure of Deep Feature for Coreset Selection

  • Zhijing Wan
  • Zhixiang Wang
  • Yuran Wang
  • Zheng Wang
  • Hongyuan Zhu
  • Shin'ichi Satoh

Coreset selection seeks to choose a subset of crucial training samples for efficient learning. It has gained traction in deep learning, particularly with the surge in training dataset sizes. Sample selection hinges on two main aspects: a sample's representation in enhancing performance and the role of sample diversity in averting overfitting. Existing methods typically measure both the representation and diversity of data based on similarity metrics, such as L2-norm. They have capably tackled representation via distribution matching guided by the similarities of features, gradients, or other information between data. However, the results of effectively diverse sample selection are mired in sub-optimality. This is because the similarity metrics usually simply aggregate dimension similarities without acknowledging disparities among the dimensions that significantly contribute to the final similarity. As a result, they fall short of adequately capturing diversity. To address this, we propose a feature-based diversity constraint, compelling the chosen subset to exhibit maximum diversity. Our key lies in the introduction of a novel Contributing Dimension Structure (CDS) metric. Different from similarity metrics that measure the overall similarity of high-dimensional features, our CDS metric considers not only the reduction of redundancy in feature dimensions, but also the difference between dimensions that contribute significantly to the final similarity. We reveal that existing methods tend to favor samples with similar CDS, leading to a reduced variety of CDS types within the coreset and subsequently hindering model performance. In response, we enhance the performance of five classical selection methods by integrating the CDS constraint. Our experiments on three datasets demonstrate the general effectiveness of the proposed method in boosting existing methods.

EAAI Journal 2024 Journal Article

Dynamic scheduling for multi-level air defense with contingency situations based on Human-Intelligence collaboration

  • Rugang Tang
  • Xin Ning
  • Zheng Wang
  • Jiaqi Fan
  • Shichao Ma

Resource scheduling is an important part of military operation, especially in key-point air defense under saturation attack. Many achievements have been made in the field of radar resource scheduling and multi-aircraft scheduling. However, there is little research on the integrated scheduling of detection, tracking and attack, which will greatly increase the resource utilization to deal with the resource shortage problem caused by saturation attacks on key places. In this paper, we propose to autonomously accomplish real-time resource dispatching through end-to-end deep reinforcement learning (DRL), while allowing the commander’s intervention to accomplish a variety of complex tactics. First, an integrated scheduling model of detection, tracking and interception is proposed and transformed into a sequential decision problem by introducing a disjunctive graph and a graph neural network (GNN) to extract node features. Subsequently, the Proximal Policy Optimization (PPO) algorithm is applied to learn the air defense environment (ADE), which is modeled as an Markov decision process (MDP). Benefitting from the powerful generalization capability of the policy network, our algorithm can adapt to scheduling missions of different sizes. Moreover, we propose a novel Human-Intelligence collaborative dynamic scheduling framework for emergency response. Simulation results indicate that our algorithm generates high-quality scheduling policies for defense resources, exhibiting superior performance than existing methods. In addition, the dynamic scheduling performance of the human and intelligence collaboration approach in response to multiple contingencies is proven.

IJCAI Conference 2024 Conference Paper

Expressiveness is Effectiveness: Self-supervised Fashion-aware CLIP for Video-to-Shop Retrieval

  • Likai Tian
  • Zhengwei Yang
  • Zechao Hu
  • Hao Li
  • Yifang Yin
  • Zheng Wang

The rise of online shopping and social media has spurred the Video-to-Shop Retrieval (VSR) task, which involves identifying fashion items (e. g. , clothing) in videos and matching them with identical products provided by stores. In real-world scenarios, human movement in dynamic video scenes can cause substantial morphological alterations of fashion items with aspects of occlusion, shifting viewpoints (parallax), and partial visibility (truncation). This results in those high-quality frames being overwhelmed by a vast of redundant ones, which makes the retrieval less effectiveness. To this end, this paper introduces a framework, named Self-supervised Fashion-aware CLIP (SF-CLIP), for effective VSR. The SF-CLIP enables the discovery of salient frames with high fashion expressiveness via generating pseudo-labels from three key aspects of fashion expressiveness to assess occlusion, parallax, and truncation. With such pseudo-labels, the ability of CLIP is expanded to facilitate the discovery of salient frames. Furthermore, to encompass the comprehensive representations among salient frames, a dual-branch graph-based fusion module is proposed to extract and integrate inter-frame features. Extensive experiments demonstrate the superiority of SF-CLIP over the state-of-the-arts.

IJCAI Conference 2024 Conference Paper

FedPFT: Federated Proxy Fine-Tuning of Foundation Models

  • Zhaopeng Peng
  • Xiaoliang Fan
  • Yufan Chen
  • Zheng Wang
  • Shirui Pan
  • Chenglu Wen
  • Ruisheng Zhang
  • Cheng Wang

Adapting Foundation Models (FMs) for down- stream tasks through Federated Learning (FL) emerges a promising strategy for protecting data privacy and valuable FMs. Existing methods fine- tune FM by allocating sub-FM to clients in FL, however, leading to suboptimal performance due to insufficient tuning and inevitable error accumula- tions of gradients. In this paper, we propose Feder- ated Proxy Fine-Tuning (FedPFT), a novel method enhancing FMs adaptation in downstream tasks through FL by two key modules. First, the sub-FM construction module employs a layer-wise com- pression approach, facilitating comprehensive FM fine-tuning across all layers by emphasizing those crucial neurons. Second, the sub-FM alignment module conducts a two-step distillations—layer- level and neuron-level—before and during FL fine- tuning respectively, to reduce error of gradient by accurately aligning sub-FM with FM under theo- retical guarantees. Experimental results on seven commonly used datasets (i. e. , four text and three vi- sion) demonstrate the superiority of FedPFT. Our code is available at https: //github. com/pzp-dzd/FedPFT.

ICLR Conference 2024 Conference Paper

Graph-constrained diffusion for End-to-End Path Planning

  • Dingyuan Shi
  • Yongxin Tong
  • Zimu Zhou
  • Ke Xu 0001
  • Zheng Wang
  • Jieping Ye

Path planning underpins various applications such as transportation, logistics, and robotics. Conventionally, path planning is formulated with explicit optimization objectives such as distance or time. However, real-world data reveals that user intentions are hard-to-model, suggesting a need for data-driven path planning that implicitly incorporates the complex user intentions. In this paper, we propose GDP, a diffusion-based model for end-to-end data-driven path planning. It effectively learns path patterns via a novel diffusion process that incorporates constraints from road networks, and plans paths as conditional path generation given the origin and destination as prior evidence. GDP is the first solution that bypasses the traditional search-based frameworks, a long-standing performance bottleneck in path planning. We validate the efficacy of GDP on two real-world datasets. Our GDP beats strong baselines by 14.2% ~ 43.5% and achieves state-of-the-art performances.

JBHI Journal 2024 Journal Article

GREMI: An Explainable Multi-Omics Integration Framework for Enhanced Disease Prediction and Module Identification

  • Hong Liang
  • Haoran Luo
  • Zhiling Sang
  • Miao Jia
  • Xiaohan Jiang
  • Zheng Wang
  • Shan Cong
  • Xiaohui Yao

Multi-omics integration has demonstrated promising performance in complex disease prediction. However, existing research typically focuses on maximizing prediction accuracy, while often neglecting the essential task of discovering meaningful biomarkers. This issue is particularly important in biomedicine, as molecules often interact rather than function individually to influence disease outcomes. To this end, we propose a two-phase framework named GREMI to assist multi-omics classification and explanation. In the prediction phase, we propose to improve prediction performance by employing a graph attention architecture on sample-wise co-functional networks to incorporate biomolecular interaction information for enhanced feature representation, followed by the integration of a joint-late mixed strategy and the true-class-probability block to adaptively evaluate classification confidence at both feature and omics levels. In the interpretation phase, we propose a multi-view approach to explain disease outcomes from the interaction module perspective, providing a more intuitive understanding and biomedical rationale. We incorporate Monte Carlo tree search (MCTS) to explore local-view subgraphs and pinpoint modules that highly contribute to disease characterization from the global-view. Extensive experiments demonstrate that the proposed framework outperforms state-of-the-art methods in seven different classification tasks, and our model effectively addresses data mutual interference when the number of omics types increases. We further illustrate the functional- and disease-relevance of the identified modules, as well as validate the classification performance of discovered modules using an independent cohort.

AAAI Conference 2024 Conference Paper

Open-Vocabulary Video Relation Extraction

  • Wentao Tian
  • Zheng Wang
  • Yuqian Fu
  • Jingjing Chen
  • Lechao Cheng

A comprehensive understanding of videos is inseparable from describing the action with its contextual action-object interactions. However, many current video understanding tasks prioritize general action classification and overlook the actors and relationships that shape the nature of the action, resulting in a superficial understanding of the action. Motivated by this, we introduce Open-vocabulary Video Relation Extraction (OVRE), a novel task that views action understanding through the lens of action-centric relation triplets. OVRE focuses on pairwise relations that take part in the action and describes these relation triplets with natural languages. Moreover, we curate the Moments-OVRE dataset, which comprises 180K videos with action-centric relation triplets, sourced from a multi-label action classification dataset. With Moments-OVRE, we further propose a cross-modal mapping model to generate relation triplets as a sequence. Finally, we benchmark existing cross-modal generation models on the new task of OVRE. Our code and dataset are available at https://github.com/Iriya99/OVRE.

IJCAI Conference 2024 Conference Paper

Perturbation Guiding Contrastive Representation Learning for Time Series Anomaly Detection

  • Liaoyuan Tang
  • Zheng Wang
  • Guanxiong He
  • Rong Wang
  • Feiping Nie

Time series anomaly detection is a critical task with applications in various domains. Due to annotation challenges, self-supervised methods have become the mainstream approach for time series anomaly detection in recent years. However, current contrastive methods categorize data perturbations into binary classes, normal or anomaly, which lack clarity on the specific impact of different perturbation methods. Inspired by the hypothesis that "the higher the probability of misclassifying perturbation types, the higher the probability of anomalies", we propose PCRTA, our approach firstly devises a perturbation classifier to learn the pseudo-labels of data perturbations. Furthermore, for addressing "class collapse issue" in contrastive learning, we propose a perturbation guiding positive and negative samples selection strategy by introducing learnable perturbation classification networks. Extensive experiments on six realworld datasets demonstrate the significant superiority of our model over thirteen state-of-the-art competitors, and obtains average 5. 14%, 8. 24% improvement in F1 score and AUC-PR, respectively.

EAAI Journal 2024 Journal Article

Plant leaf disease identification by parameter-efficient transformer with adapter

  • Xingshi Xu
  • Guangyuan Yang
  • Yunfei Wang
  • Yuying Shang
  • Zhixin Hua
  • Zheng Wang
  • Huaibo Song

Accurate identification of plant disease is of crucial importance for agricultural production. Existing research often tailor dedicated leaf disease recognition models for each plant species to achieve superior recognition accuracy and robustness. However, these approaches are labour-intensive and time-consuming. In addition, when integrating multiple models on a single device, numerous model parameters need to be stored. To address the aforementioned problems, a model named PDNet was designed for the identification of plant leaf disease, advancing the concept of using the same network with almost the same weights for different plant species. First, an Adapter block was utilized to implement a parameter-efficient training. The proposed method could add only 1. 39% additional parameters when developing a new model for a specific leaf disease identification task, significantly outperforming the existing 100% fully fine-tuned models. Moreover, to enhance the identification accuracy, Overlapping patch embedding (OPE) was introduced into PDNet to avoid the destruction of essential information on plant leaves. Subsequently, Angular Softmax Loss (A-Softmax Loss) was employed to achieve fine-grained recognition of similar diseases. The proposed method on six plant disease datasets achieved the identification accuracies of 95. 16%, 96. 74%, 98. 26%, 91. 16%, 98. 67% and 99. 58%, respectively. It demonstrated excellent performance under Few-Shot conditions. The recognition accuracy remained above 75% with only three training samples for each disease. This study provides a novel method and paradigm for plant leaf disease identification, offering valuable insights for the field.

AAAI Conference 2024 Conference Paper

PrefAce: Face-Centric Pretraining with Self-Structure Aware Distillation

  • Siyuan Hu
  • Zheng Wang
  • Peng Hu
  • Xi Peng
  • Jie Wu
  • Hongyuan Zhu
  • Yew Soon Ong

Video-based facial analysis is important for autonomous agents to understand human expressions and sentiments. However, limited labeled data is available to learn effective facial representations. This paper proposes a novel self-supervised face-centric pretraining framework, called PrefAce, which learns transferable video facial representation without labels. The self-supervised learning is performed with an effective landmark-guided global-local tube distillation. Meanwhile, a novel instance-wise update FaceFeat Cache is built to enforce more discriminative and diverse representations for downstream tasks. Extensive experiments demonstrate that the proposed framework learns universal instance-aware facial representations with fine-grained landmark details from videos. The point is that it can transfer across various facial analysis tasks, e.g., Facial Attribute Recognition (FAR), Facial Expression Recognition (FER), DeepFake Detection (DFD), and Lip Synchronization (LS). Our framework also outperforms the state-of-the-art on various downstream tasks, even in low data regimes. Code is available at https://github.com/siyuan-h/PrefAce.

ICRA Conference 2024 Conference Paper

RBI-RRT*: Efficient Sampling-based Path Planning for High-dimensional State Space

  • Fang Chen
  • Yu Zheng
  • Zheng Wang
  • Wanchao Chi
  • Sicong Liu

Sampling-based planning algorithms such as RRT have been proved to be efficient in solving path planning problems for robotic systems. Various improvements to the RRT algorithm have been presented to improve the performance of the extension and convergence of the random trees, such as Informed RRT*. However, with the growth of spatial dimensions, the time consumption of randomly sampling the entire state space and incrementally rewiring the random trees raises drastically before a feasible solution is found. In this paper, to enhance the convergence performance of optimal solutions, we present Reconstructed Bi-directional Informed RRT* (RBI-RRT*) path planning algorithm. The algorithm acts as RRT-Connect to rapidly find a feasible solution, which helps compress the sampling space as Informed RRT* does. After the random trees are transformed into RRT* structure by the reconstruction process in RBI-RRT*, the algorithm continues to find the near-optimal path. A series of simulations and real-world robot experiments were conducted to evaluate the algorithm against existing planning algorithms. Compared to Informed RRT* Connect, RBI-RRT* reduced the computation time of achieving a specific cost by 22. 1% on average in simulations and 11. 2% in the real-world robotic arm experiments. The results show that RBI-RRT* is more efficient in high-dimensional planning problems.

NeurIPS Conference 2024 Conference Paper

ReFT: Representation Finetuning for Language Models

  • Zhengxuan Wu
  • Aryaman Arora
  • Zheng Wang
  • Atticus Geiger
  • Dan Jurafsky
  • Christopher D. Manning
  • Christopher Potts

Parameter-efficient finetuning (PEFT) methods seek to adapt large neural models via updates to a small number of weights. However, much prior interpretability work has shown that representations encode rich semantic information, suggesting that editing representations might be a more powerful alternative. We pursue this hypothesis by developing a family of Representation Finetuning (ReFT) methods. ReFT methods operate on a frozen base model and learn task-specific interventions on hidden representations. We define a strong instance of the ReFT family, Low-rank Linear Subspace ReFT (LoReFT), and we identify an ablation of this method that trades some performance for increased efficiency. Both are drop-in replacements for existing PEFTs and learn interventions that are 15x--65x more parameter-efficient than LoRA. We showcase LoReFT on eight commonsense reasoning tasks, four arithmetic reasoning tasks, instruction-tuning, and GLUE. In all these evaluations, our ReFTs deliver the best balance of efficiency and performance, and almost always outperform state-of-the-art PEFTs. Upon publication, we will publicly release our generic ReFT training library.

NeurIPS Conference 2024 Conference Paper

Revisiting Adversarial Patches for Designing Camera-Agnostic Attacks against Person Detection

  • Hui Wei
  • Zhixiang Wang
  • Kewei Zhang
  • Jiaqi Hou
  • Yuanwei Liu
  • Hao Tang
  • Zheng Wang

Physical adversarial attacks can deceive deep neural networks (DNNs), leading to erroneous predictions in real-world scenarios. To uncover potential security risks, attacking the safety-critical task of person detection has garnered significant attention. However, we observe that existing attack methods overlook the pivotal role of the camera, involving capturing real-world scenes and converting them into digital images, in the physical adversarial attack workflow. This oversight leads to instability and challenges in reproducing these attacks. In this work, we revisit patch-based attacks against person detectors and introduce a camera-agnostic physical adversarial attack to mitigate this limitation. Specifically, we construct a differentiable camera Image Signal Processing (ISP) proxy network to compensate for the physical-to-digital transition gap. Furthermore, the camera ISP proxy network serves as a defense module, forming an adversarial optimization framework with the attack module. The attack module optimizes adversarial patches to maximize effectiveness, while the defense module optimizes the conditional parameters of the camera ISP proxy network to minimize attack effectiveness. These modules engage in an adversarial game, enhancing cross-camera stability. Experimental results demonstrate that our proposed Camera-Agnostic Patch (CAP) attack effectively conceals persons from detectors across various imaging hardware, including two distinct cameras and four smartphones.

IJCAI Conference 2024 Conference Paper

Sparse Multi-Relational Graph Convolutional Network for Multi-type Object Trajectory Prediction

  • Jianhui Zhang
  • Jun Yao
  • Liqi Yan
  • Yanhong Xu
  • Zheng Wang

Object trajectory prediction is a hot research issue with wide applications in video surveillance and autonomous driving. The previous studies consider the interaction sparsity mainly among the pedestrians instead of multi-type of objects, which brings new types of interactions and consequently superfluous ones. This paper proposes a Multi-type Object Trajectory Prediction (MOTP) method with a Sparse Multi-relational Graph Convolutional Network (SMGCN) and a novel multi-round Global Temporal Aggregation (GTA). MOTP introduces a novel adaptive sparsification and multi-scale division method to model interactions among multitype of objects. It further incorporates a Sparse Multi-relational Temporal Graph to capture the temporal division of multi-type trajectories, along with a multi-round Global Temporal Aggregation (GTA) mechanism to mitigate error accumulation, and enhances the trajectory prediction accuracy. The extensive evaluation on the ETH, UCY and SDD datasets shows that our method outperforms the typical state-of-the-art works by significant margins. Codes will be available in https: //github. com/ sounio/SMGCN.

ECAI Conference 2024 Conference Paper

TDCL: Dense Semantic Contrastive Learning for Vision-Language Tracking

  • Zheng Wang
  • Xiankang He
  • Kaiyang Lan
  • Ying Cui
  • Dongyan Guo

Traditional single-object tracking tasks are undergoing a new wave of transformation, especially with the emergence of the lack of semantics, which has led to the rise of the vision-language tracking task. However, previous approaches that combine the visual tracker with natural language descriptions tend to rely on a global representation of the text description, considering less about the fine-grained connections between the text description and the visual appearance. This paper proposes to utilize a bi-directional cross-attention module to capture the connections between language and visual features, which are further projected as dense semantic representations for alignment. In order to keep the semantic consistency between the search region and the coupled natural language and align the fused feature, this paper proposes a novel dense semantic contrastive learning loss to bridge the semantic gap between text and visual modalities and align them in a dense form. The proposed framework achieves promising results in tracking datasets that contain natural language descriptions, such as TNL2K, and OTB99-LANG. Our approach provides a novel solution for representing and aligning cross-modal information for the single object tracking task and may inspire further research in this field.

NeurIPS Conference 2024 Conference Paper

The Implicit Bias of Gradient Descent toward Collaboration between Layers: A Dynamic Analysis of Multilayer Perceptions

  • Zheng Wang
  • Geyong Min
  • Wenjie Ruan

The implicit bias of gradient descent has long been considered the primary mechanism explaining the superior generalization of over-parameterized neural networks without overfitting, even when the training error is zero. However, the implicit bias toward adversarial robustness has rarely been considered in the research community, although it is crucial for the trustworthiness of machine learning models. To fill this gap, in this paper, we explore whether consecutive layers collaborate to strengthen adversarial robustness during gradient descent. By quantifying this collaboration between layers using our proposed concept, co-correlation, we demonstrate a monotonically increasing trend in co-correlation, which implies a decreasing trend in adversarial robustness during gradient descent. Additionally, we observe different behaviours between narrow and wide neural networks during gradient descent. We conducted extensive experiments that verified our proposed theorems.

TMLR Journal 2024 Journal Article

Unleashing the Potential of Acquisition Functions in High-Dimensional Bayesian Optimization

  • Jiayu Zhao
  • Renyu Yang
  • SHENGHAO QIU
  • Zheng Wang

Bayesian optimization (BO) is widely used to optimize expensive-to-evaluate black-box functions. It first builds a surrogate for the objective and quantifies its uncertainty. It then decides where to sample by maximizing an acquisition function (AF) defined by the surrogate model. However, when dealing with high-dimensional problems, finding the global maximum of the AF becomes increasingly challenging. In such cases, the manner in which the AF maximizer is initialized plays a pivotal role. An inappropriate initialization can severely limit the potential of AF. This paper investigates a largely understudied problem concerning the impact of AF maximizer initialization on exploiting AFs' capability. Our large-scale empirical study shows that the widely used random initialization strategy may fail to harness the potential of an AF. Based on this observation, we propose a better initialization approach by employing multiple heuristic optimizers to leverage the historical data of black-box optimization to generate initial points for an AF maximizer. We evaluate our approach with a variety of heavily studied synthetic test functions and real-world applications. Experimental results show that our techniques, while simple, can significantly enhance the standard BO and outperform state-of-the-art methods by a large margin in most test cases.

ICML Conference 2024 Conference Paper

Unveiling and Harnessing Hidden Attention Sinks: Enhancing Large Language Models without Training through Attention Calibration

  • Zhongzhi Yu
  • Zheng Wang
  • Yonggan Fu
  • Huihong Shi
  • Khalid Shaikh 0002
  • Yingyan Celine Lin

Attention is a fundamental component behind the remarkable achievements of large language models (LLMs). However, our current understanding of the attention mechanism, especially regarding how attention distributions are established, remains limited. Inspired by recent studies that explore the presence of attention sink in the initial token, which receives disproportionately large attention scores despite their lack of semantic importance, this work delves deeper into this phenomenon. We aim to provide a more profound understanding of the existence of attention sinks within LLMs and to uncover ways to enhance the achievable accuracy of LLMs by directly optimizing the attention distributions, without the need for weight finetuning. Specifically, this work begins with comprehensive visualizations of the attention distributions in LLMs during inference across various inputs and tasks. Based on these visualizations, to the best of our knowledge, we are the first to discover that (1) attention sinks occur not only at the start of sequences but also within later tokens of the input, and (2) not all attention sinks have a positive impact on the achievable accuracy of LLMs. Building upon our findings, we propose a training-free Attention Calibration Technique (ACT) that automatically optimizes the attention distributions on the fly during inference in an input-adaptive manner. Extensive experiments validate that ACT consistently enhances the accuracy of various LLMs across different applications. Specifically, ACT achieves an average improvement of up to $7. 30%$ in accuracy across different datasets when applied to Llama-30B.

ICML Conference 2024 Conference Paper

When Linear Attention Meets Autoregressive Decoding: Towards More Effective and Efficient Linearized Large Language Models

  • Haoran You
  • Yichao Fu
  • Zheng Wang
  • Amir Yazdanbakhsh
  • Yingyan Celine Lin

Autoregressive Large Language Models (LLMs) have achieved impressive performance in language tasks but face two significant bottlenecks: (1) quadratic complexity in the attention module as the number of tokens increases, and (2) limited efficiency due to the sequential processing nature of autoregressive LLMs during generation. While linear attention and speculative decoding offer potential solutions, their applicability and synergistic potential for enhancing autoregressive LLMs remain uncertain. We conduct the first comprehensive study on the efficacy of existing linear attention methods for autoregressive LLMs, integrating them with speculative decoding. We introduce an augmentation technique for linear attention that ensures compatibility with speculative decoding, enabling more efficient training and serving of LLMs. Extensive experiments and ablation studies involving seven existing linear attention models and five encoder/decoder-based LLMs consistently validate the effectiveness of our augmented linearized LLMs. Notably, our approach achieves up to a 6. 67 reduction in perplexity on the LLaMA model and up to a 2$\times$ speedup during generation compared to prior linear attention methods. Codes and models are available at https: //github. com/GATECH-EIC/Linearized-LLM.

IJCAI Conference 2023 Conference Paper

Don't Ignore Alienation and Marginalization: Correlating Fraud Detection

  • Yilong Zang
  • Ruimin Hu
  • Zheng Wang
  • Danni Xu
  • Jia Wu
  • Dengshi Li
  • Junhang Wu
  • Lingfei Ren

The anonymity of online networks makes tackling fraud increasingly costly. Thanks to the superiority of graph representation learning, graph-based fraud detection has made significant progress in recent years. However, upgrading fraudulent strategies produces more advanced and difficult scams. One common strategy is synergistic camouflage —— combining multiple means to deceive others. Existing methods mostly investigate the differences between relations on individual frauds, that neglect the correlation among multi-relation fraudulent behaviors. In this paper, we design several statistics to validate the existence of synergistic camouflage of fraudsters by exploring the correlation among multi-relation interactions. From the perspective of multi-relation, we find two distinctive features of fraudulent behaviors, i. e. , alienation and marginalization. Based on the finding, we propose COFRAUD, a correlation-aware fraud detection model, which innovatively incorporates synergistic camouflage into fraud detection. It captures the correlation among multi-relation fraudulent behaviors. Experimental results on two public datasets demonstrate that COFRAUD achieves significant improvements over state-of-the-art methods.

NeurIPS Conference 2023 Conference Paper

Dynamic Tensor Decomposition via Neural Diffusion-Reaction Processes

  • Zheng Wang
  • Shikai Fang
  • Shibo Li
  • Shandian Zhe

Tensor decomposition is an important tool for multiway data analysis. In practice, the data is often sparse yet associated with rich temporal information. Existing methods, however, often under-use the time information and ignore the structural knowledge within the sparsely observed tensor entries. To overcome these limitations and to better capture the underlying temporal structure, we propose Dynamic EMbedIngs fOr dynamic Tensor dEcomposition (DEMOTE). We develop a neural diffusion-reaction process to estimate dynamic embeddings for the entities in each tensor mode. Specifically, based on the observed tensor entries, we build a multi-partite graph to encode the correlation between the entities. We construct a graph diffusion process to co-evolve the embedding trajectories of the correlated entities and use a neural network to construct a reaction process for each individual entity. In this way, our model can capture both the commonalities and personalities during the evolution of the embeddings for different entities. We then use a neural network to model the entry value as a nonlinear function of the embedding trajectories. For model estimation, we combine ODE solvers to develop a stochastic mini-batch learning algorithm. We propose a stratified sampling method to balance the cost of processing each mini-batch so as to improve the overall efficiency. We show the advantage of our approach in both simulation studies and real-world applications. The code is available at https: //github. com/wzhut/Dynamic-Tensor-Decomposition-via-Neural-Diffusion-Reaction-Processes.

AAAI Conference 2023 Conference Paper

FedGS: Federated Graph-Based Sampling with Arbitrary Client Availability

  • Zheng Wang
  • Xiaoliang Fan
  • Jianzhong Qi
  • Haibing Jin
  • Peizhen Yang
  • Siqi Shen
  • Cheng Wang

While federated learning has shown strong results in opti- mizing a machine learning model without direct access to the original data, its performance may be hindered by in- termittent client availability which slows down the conver- gence and biases the final learned model. There are significant challenges to achieve both stable and bias-free training un- der arbitrary client availability. To address these challenges, we propose a framework named Federated Graph-based Sam- pling (FEDGS), to stabilize the global model update and mitigate the long-term bias given arbitrary client availabil- ity simultaneously. First, we model the data correlations of clients with a Data-Distribution-Dependency Graph (3DG) that helps keep the sampled clients data apart from each other, which is theoretically shown to improve the approximation to the optimal model update. Second, constrained by the far- distance in data distribution of the sampled clients, we fur- ther minimize the variance of the numbers of times that the clients are sampled, to mitigate long-term bias. To validate the effectiveness of FEDGS, we conduct experiments on three datasets under a comprehensive set of seven client availability modes. Our experimental results confirm FEDGS’s advantage in both enabling a fair client-sampling scheme and improving the model performance under arbitrary client availability. Our code is available at https://github.com/WwZzz/FedGS.

IJCAI Conference 2023 Conference Paper

From Generation to Suppression: Towards Effective Irregular Glow Removal for Nighttime Visibility Enhancement

  • Wanyu Wu
  • Wei Wang
  • Zheng Wang
  • Kui Jiang
  • Xin Xu

Most existing Low-Light Image Enhancement (LLIE) methods are primarily designed to improve brightness in dark regions, which suffer from severe degradation in nighttime images. However, these methods have limited exploration in another major visibility damage, the glow effects in real night scenes. Glow effects are inevitable in the presence of artificial light sources and cause further diffused blurring when directly enhanced. To settle this issue, we innovatively consider the glow suppression task as learning physical glow generation via multiple scattering estimation according to the Atmospheric Point Spread Function (APSF). In response to the challenges posed by uneven glow intensity and varying source shapes, an APSF-based Nighttime Imaging Model with Near-field Light Sources (NIM-NLS) is specifically derived to design a scalable Light-aware Blind Deconvolution Network (LBDN). The glow-suppressed result is then brightened via a Retinex-based Enhancement Module (REM). Remarkably, the proposed glow suppression method is based on zero-shot learning and does not rely on any paired or unpaired training data. Empirical evaluations demonstrate the effectiveness of the proposed method in both glow suppression and low-light enhancement tasks.

AAAI Conference 2023 Conference Paper

HOTCOLD Block: Fooling Thermal Infrared Detectors with a Novel Wearable Design

  • Hui Wei
  • Zhixiang Wang
  • Xuemei Jia
  • Yinqiang Zheng
  • Hao Tang
  • Shin'ichi Satoh
  • Zheng Wang

Adversarial attacks on thermal infrared imaging expose the risk of related applications. Estimating the security of these systems is essential for safely deploying them in the real world. In many cases, realizing the attacks in the physical space requires elaborate special perturbations. These solutions are often impractical and attention-grabbing. To address the need for a physically practical and stealthy adversarial attack, we introduce HotCold Block, a novel physical attack for infrared detectors that hide persons utilizing the wearable Warming Paste and Cooling Paste. By attaching these readily available temperature-controlled materials to the body, HotCold Block evades human eyes efficiently. Moreover, unlike existing methods that build adversarial patches with complex texture and structure features, HotCold Block utilizes an SSP-oriented adversarial optimization algorithm that enables attacks with pure color blocks and explores the influence of size, shape, and position on attack performance. Extensive experimental results in both digital and physical environments demonstrate the performance of our proposed HotCold Block. Code is available: https://github.com/weihui1308/HOTCOLDBlock.

YNIMG Journal 2023 Journal Article

Metabolic and functional substrates of impulsive decision-making in individuals with heroin addiction after prolonged methadone maintenance treatment

  • Qian Lv
  • Miao Zhang
  • Haifeng Jiang
  • Yilin Liu
  • Shaoling Zhao
  • Xiaomin Xu
  • Wenlei Zhang
  • Tianzhen Chen

F-FDG PET). Subjects receiving MMT exhibited significantly elevated self-reported impulsivity, and computational modeling revealed a marked impulsive decision bias manifested as switching more frequently without available evidence. Moreover, this impulsive decision bias was associated with the dose and duration of methadone use, irrelevant to the duration of heroin use. During the task, the switch-related hypoactivation in the left rostral middle frontal gyrus was correlated with the impulsive decision bias while the function of reward sensitivity was intact in subjects receiving MMT. Using prior brain-wide receptor density data, we found that the highest variance of regional metabolic abnormalities was explained by the spatial distribution of μ-opioid receptors among 10 types of neurotransmitter receptors. Heightened impulsivity in individuals receiving prolonged MMT is manifested as atypical choice bias and noise in decision-making processes, which is further driven by deficits in top-down cognitive control, other than reward sensitivity. Our findings uncover multifaceted mechanisms underlying elevated impulsivity in subjects receiving MMT, which might provide insights for developing complementary therapies to improve retention during MMT.

ICRA Conference 2023 Conference Paper

Origami Folding Enhances Modularity and Mechanical Efficiency of Soft Actuators

  • Zheng Wang
  • Yazhou Song
  • Zhongkui Wang
  • Hongying Zhang

Soft robots have long been attractive to robotic engineers due to their remarkable dexterity; however, reports that standardize soft actuators into modularized off-shelf devices akin to rigid robots are still rare, and the mechanical efficiency of existing designs is still limited. This work identifies origami folding to enable the design of LEGO-like modularized soft actuators with high mechanical efficiency in terms of payload capability and workspace. Herein, three modularized origami actuators that can generate translational, bending, and twisting motion are designed, prototyped, and tested. The translational actuator can contract to 40% of its original length, and the twisting and bending actuators can exert 31° and 52° angular motions, respectively. The translational actuator can exert a blocked force of about 821 times self-weight. The motion of origami soft actuators is accurately modeled using rigid body kinematics, and complex systems built by them are captured by homogeneous transformation. Finally, the modularized design and efficient kinematic model are verified on a manipulator and a reconfigurable letter. Benefiting from the unprecedented modularity and mechanical efficiency, these LEGO-like origami actuators are promising for practical applications like food handling and healthcare.

AAAI Conference 2023 Conference Paper

Store and Fetch Immediately: Everything Is All You Need for Space-Time Video Super-resolution

  • Mengshun Hu
  • Kui Jiang
  • Zhixiang Nie
  • Jiahuan Zhou
  • Zheng Wang

Existing space-time video super-resolution (ST-VSR) methods fail to achieve high-quality reconstruction since they fail to fully explore the spatial-temporal correlations, long-range components in particular. Although the recurrent structure for ST-VSR adopts bidirectional propagation to aggregate information from the entire video, collecting the temporal information between the past and future via one-stage representations inevitably loses the long-range relations. To alleviate the limitation, this paper proposes an immediate storeand-fetch network to promote long-range correlation learning, where the stored information from the past and future can be refetched to help the representation of the current frame. Specifically, the proposed network consists of two modules: a backward recurrent module (BRM) and a forward recurrent module (FRM). The former first performs backward inference from future to past, while storing future super-resolution (SR) information for each frame. Following that, the latter performs forward inference from past to future to super-resolve all frames, while storing past SR information for each frame. Since FRM inherits SR information from BRM, therefore, spatial and temporal information from the entire video sequence is immediately stored and fetched, which allows drastic improvement for ST-VSR. Extensive experiments both on ST-VSR and space video super-resolution (S-VSR) as well as time video super-resolution (T-VSR) have demonstrated the effectiveness of our proposed method over other state-of-the-art methods on public datasets. Code is available https://github.com/hhhhhumengshun/SFI-STVR

NeurIPS Conference 2023 Conference Paper

Streaming Factor Trajectory Learning for Temporal Tensor Decomposition

  • Shikai Fang
  • Xin Yu
  • Shibo Li
  • Zheng Wang
  • Mike Kirby
  • Shandian Zhe

Practical tensor data is often along with time information. Most existing temporal decomposition approaches estimate a set of fixed factors for the objects in each tensor mode, and hence cannot capture the temporal evolution of the objects' representation. More important, we lack an effective approach to capture such evolution from streaming data, which is common in real-world applications. To address these issues, we propose Streaming Factor Trajectory Learning (SFTL) for temporal tensor decomposition. We use Gaussian processes (GPs) to model the trajectory of factors so as to flexibly estimate their temporal evolution. To address the computational challenges in handling streaming data, we convert the GPs into a state-space prior by constructing an equivalent stochastic differential equation (SDE). We develop an efficient online filtering algorithm to estimate a decoupled running posterior of the involved factor states upon receiving new data. The decoupled estimation enables us to conduct standard Rauch-Tung-Striebel smoothing to compute the full posterior of all the trajectories in parallel, without the need for revisiting any previous data. We have shown the advantage of SFTL in both synthetic tasks and real-world applications.

IJCAI Conference 2022 Conference Paper

DANet: Image Deraining via Dynamic Association Learning

  • Kui Jiang
  • Zhongyuan Wang
  • Zheng Wang
  • Peng Yi
  • Junjun Jiang
  • Jinsheng Xiao
  • Chia-Wen Lin

Rain streaks and background components in a rainy input are highly correlated, making the deraining task a composition of the rain streak removal and background restoration. However, the correlation of these two components is barely considered, leading to unsatisfied deraining results. To this end, we propose a dynamic associated network (DANet) to achieve the association learning between rain streak removal and background recovery. There are two key aspects to fulfill the association learning: 1) DANet unveils the latent association knowledge between rain streak prediction and background texture recovery, and leverages it as an extra prior via an associated learning module (ALM) to promote the texture recovery. 2) DANet introduces the parametric association constraint for enhancing the compatibility of deraining model with background reconstruction, enabling it to be automatically learned from the training data. Moreover, we observe that the sampled rainy image enjoys the similar distribution to the original one. We thus propose to learn the rain distribution at the sampling space, and exploit super-resolution to reconstruct high-frequency background details for computation and memory reduction. Our proposed DANet achieves the approximate deraining performance to the state-of-the-art MPRNet but only requires 52. 6\% and 23\% inference time and computational cost, respectively.

AAAI Conference 2022 Conference Paper

Degrade Is Upgrade: Learning Degradation for Low-Light Image Enhancement

  • Kui Jiang
  • Zhongyuan Wang
  • Zheng Wang
  • Chen Chen
  • Peng Yi
  • Tao Lu
  • Chia-Wen Lin

Low-light image enhancement aims to improve an image’s visibility while keeping its visual naturalness. Different from existing methods tending to accomplish the relighting task directly by ignoring the fidelity and naturalness recovery, we investigate the intrinsic degradation and relight the lowlight image while refining the details and color in two steps. Inspired by the color image formulation (diffuse illumination color plus environment illumination color), we first estimate the degradation from low-light inputs to simulate the distortion of environment illumination color, and then refine the content to recover the loss of diffuse illumination color. To this end, we propose a novel Degradation-to-Refinement Generation Network (DRGN). Its distinctive features can be summarized as 1) A novel two-step generation network for degradation learning and content refinement. It is not only superior to one-step methods, but also capable of synthesizing sufficient paired samples to benefit the model training; 2) A multi-resolution fusion network to represent the target information (degradation or contents) in a multi-scale cooperative manner, which is more effective to address the complex unmixing problems. Extensive experiments on both the enhancement task and joint detection task have verified the effectiveness and efficiency of our proposed method, surpassing the SOTA by 0. 70dB on average and 3. 18% in mAP, respectively. The code will be available soon.

AAAI Conference 2022 Conference Paper

ELMA: Energy-Based Learning for Multi-Agent Activity Forecasting

  • Yuke Li
  • Pin Wang
  • Lixiong Chen
  • Zheng Wang
  • Ching-Yao Chan

This paper describes an energy-based learning method that predicts the activities of multiple agents simultaneously. It aims to forecast both upcoming actions and paths of all agents in a scene based on their past activities, which can be jointly formulated by a probabilistic model over time. Learning this model is challenging because: 1) it has a large number of time-dependent variables that must scale with the forecast horizon and the number of agents; 2) distribution functions have to contain multiple modes in order to capture the spatiotemporal complexities of each agent’s activities. To address these challenges, we put forth a novel Energy-based Learning approach for Multi-Agent activity forecasting (ELMA) to estimate this complex model via maximum log-likelihood estimation. Specifically, by sampling from a sequence of factorized marginalized multi-modal distributions, ELMA generates the possible future actions efficiently. Moreover, by graph-based representations, ELMA also explicitly resolves the spatio-temporal dependencies of all agents’ activities in a single pass. Our experiments on two large-scale datasets prove that ELMA outperforms recent leading studies by an obvious margin.

NeurIPS Conference 2022 Conference Paper

Infinite-Fidelity Coregionalization for Physical Simulation

  • Shibo Li
  • Zheng Wang
  • Robert Kirby
  • Shandian Zhe

Multi-fidelity modeling and learning is important in physical simulation related applications. It can leverage both low-fidelity and high-fidelity examples for training so as to reduce the cost of data generation yet still achieving good performance. While existing approaches only model finite, discrete fidelities, in practice, the feasible fidelity choice is often infinite, which can correspond to a continuous mesh spacing or finite element length. In this paper, we propose Infinite Fidelity Coregionalization (IFC). Given the data, our method can extract and exploit rich information within infinite, continuous fidelities to bolster the prediction accuracy. Our model can interpolate and/or extrapolate the predictions to novel fidelities that are not covered by the training data. Specifically, we introduce a low-dimensional latent output as a continuous function of the fidelity and input, and multiple it with a basis matrix to predict high-dimensional solution outputs. We model the latent output as a neural Ordinary Differential Equation (ODE) to capture the complex relationships within and integrate information throughout the continuous fidelities. We then use Gaussian processes or another ODE to estimate the fidelity-varying bases. For efficient inference, we reorganize the bases as a tensor, and use a tensor-Gaussian variational posterior approximation to develop a scalable inference algorithm for massive outputs. We show the advantage of our method in several benchmark tasks in computational physics.

IJCAI Conference 2022 Conference Paper

Rainy WCity: A Real Rainfall Dataset with Diverse Conditions for Semantic Driving Scene Understanding

  • Xian Zhong
  • Shidong Tu
  • Xianzheng Ma
  • Kui Jiang
  • Wenxin Huang
  • Zheng Wang

Scene understanding in adverse weather conditions (e. g. rainy and foggy days) has drawn increasing attention, arising some specific benchmarks and algorithms. However, scene segmentation under rainy weather is still challenging and under-explored due to the following limitations on the datasets and methods: 1) Manually synthetic rainy samples with empirically settings and human subjective assumptions; 2) Limited rainy conditions, including the rain patterns, intensity, and degradation factors; 3) Separated training manners for image deraining and semantic segmentation. To break these limitations, we pioneer a real, comprehensive, and well-annotated scene understanding dataset under rainy weather, named Rainy WCity. It covers various rain patterns and their bring-in negative visual effects, covering wiper, droplet, reflection, refraction, shadow, windshield-blurring, etc. In addition, to alleviate dependence on paired training samples, we design an unsupervised contrastive learning network for real image deraining and the final rainy scene semantic segmentation via multi-task joint optimization. A comprehensive comparison analysis is also provided, which shows that scene understanding in rainy weather is a largely open problem. Finally, we summarize our general observations, identify open research challenges, and point out future directions.

IJCAI Conference 2021 Conference Paper

Federated Learning with Fair Averaging

  • Zheng Wang
  • Xiaoliang Fan
  • Jianzhong Qi
  • Chenglu Wen
  • Cheng Wang
  • Rongshan Yu

Fairness has emerged as a critical problem in federated learning (FL). In this work, we identify a cause of unfairness in FL -- conflicting gradients with large differences in the magnitudes. To address this issue, we propose the federated fair averaging (FedFV) algorithm to mitigate potential conflicts among clients before averaging their gradients. We first use the cosine similarity to detect gradient conflicts, and then iteratively eliminate such conflicts by modifying both the direction and the magnitude of the gradients. We further show the theoretical foundation of FedFV to mitigate the issue conflicting gradients and converge to Pareto stationary solutions. Extensive experiments on a suite of federated datasets confirm that FedFV compares favorably against state-of-the-art methods in terms of fairness, accuracy and efficiency. The source code is available at https: //github. com/WwZzz/easyFL.

AAAI Conference 2021 Conference Paper

Learning to Attack Real-World Models for Person Re-identification via Virtual-Guided Meta-Learning

  • Fengxiang Yang
  • Zhun Zhong
  • Hong Liu
  • Zheng Wang
  • Zhiming Luo
  • Shaozi Li
  • Nicu Sebe
  • Shin'ichi Satoh

Recent advances in person re-identification (re-ID) have led to impressive retrieval accuracy. However, existing re-ID models are challenged by the adversarial examples crafted by adding quasi-imperceptible perturbations. Moreover, re- ID systems face the domain shift issue that training and testing domains are not consistent. In this study, we argue that learning powerful attackers with high universality that works well on unseen domains is an important step in promoting the robustness of re-ID systems. Therefore, we introduce a novel universal attack algorithm called “MetaAttack” for person re-ID. MetaAttack can mislead re-ID models on unseen domains by a universal adversarial perturbation. Specifically, to capture common patterns across different domains, we propose a meta-learning scheme to seek the universal perturbation via the gradient interaction between meta-train and meta-test formed by two datasets. We also take advantage of a virtual dataset (PersonX), instead of real ones, to conduct meta-test. This scheme not only enables us to learn with more comprehensive variation factors but also mitigates the negative effects caused by biased factors of real datasets. Experiments on three large-scale re-ID datasets demonstrate the effectiveness of our method in attacking re-ID models on unseen domains. Our final visualization results reveal some new properties of existing re-ID systems, which can guide us in designing a more robust re- ID model. Code and supplemental material are available at https: //github. com/FlyingRoastDuck/MetaAttack AAAI21.

IJCAI Conference 2021 Conference Paper

Location Predicts You: Location Prediction via Bi-direction Speculation and Dual-level Association

  • Xixi Li
  • Ruimin Hu
  • Zheng Wang
  • Toshihiko Yamasaki

Location prediction is of great importance in location-based applications for the construction of the smart city. To our knowledge, existing models for location prediction focus on the users' preference on POIs from the perspective of the human side. However, modeling users' interests from the historical trajectory is still limited by the data sparsity. Additionally, most of existing methods predict the next location according to the individual data independently. But the data sparsity makes it difficult to mine explicit mobility patterns or capture the casual behavior for each user. To address the issues above, we propose a novel Bi-direction Speculation and Dual-level Association method (BSDA), which considers both users' interests in POIs and POIs' appeal to users. Furthermore, we develop the cross-user and cross-POI association to alleviate the data sparsity by similar users and POIs to enrich the candidates. Experimental results on two public datasets demonstrate that BSDA achieves significant improvements over state-of-the-art methods.

NeurIPS Conference 2021 Conference Paper

Self-Adaptable Point Processes with Nonparametric Time Decays

  • Zhimeng Pan
  • Zheng Wang
  • Jeff M Phillips
  • Shandian Zhe

Many applications involve multi-type event data. Understanding the complex influences of the events on each other is critical to discover useful knowledge and to predict future events and their types. Existing methods either ignore or partially account for these influences. Recent works use recurrent neural networks to model the event rate. While being highly expressive, they couple all the temporal dependencies in a black-box and can hardly extract meaningful knowledge. More important, most methods assume an exponential time decay of the influence strength, which is over-simplified and can miss many important strength varying patterns. To overcome these limitations, we propose SPRITE, a $\underline{S}$elf-adaptable $\underline{P}$oint p$\underline{R}$ocess w$\underline{I}$th nonparametric $\underline{T}$ime d$\underline{E}$cays, which can decouple the influences between every pair of the events and capture various time decays of the influence strengths. Specifically, we use an embedding to represent each event type and model the event influence as an unknown function of the embeddings and time span. We derive a general construction that can cover all possible time decaying functions. By placing Gaussian process (GP) priors over the latent functions and using Gauss-Legendre quadrature to obtain the integral in the construction, we can flexibly estimate all kinds of time-decaying influences, without restricting to any specific form or imposing derivative constraints that bring learning difficulties. We then use weight space augmentation of GPs to develop an efficient stochastic variational learning algorithm. We show the advantages of our approach in both the ablation study and real-world applications.

YNIMG Journal 2021 Journal Article

U-net model for brain extraction: Trained on humans for transfer to non-human primates

  • Xindi Wang
  • Xin-Hui Li
  • Jae Wook Cho
  • Brian E. Russ
  • Nanditha Rajamani
  • Alisa Omelchenko
  • Lei Ai
  • Annachiara Korchmaros

Brain extraction (a.k.a. skull stripping) is a fundamental step in the neuroimaging pipeline as it can affect the accuracy of downstream preprocess such as image registration, tissue classification, etc. Most brain extraction tools have been designed for and applied to human data and are often challenged by non-human primates (NHP) data. Amongst recent attempts to improve performance on NHP data, deep learning models appear to outperform the traditional tools. However, given the minimal sample size of most NHP studies and notable variations in data quality, the deep learning models are very rarely applied to multi-site samples in NHP imaging. To overcome this challenge, we used a transfer-learning framework that leverages a large human imaging dataset to pretrain a convolutional neural network (i.e. U-Net Model), and then transferred this to NHP data using a small NHP training sample. The resulting transfer-learning model converged faster and achieved more accurate performance than a similar U-Net Model trained exclusively on NHP samples. We improved the generalizability of the model by upgrading the transfer-learned model using additional training datasets from multiple research sites in the Primate Data-Exchange (PRIME-DE) consortium. Our final model outperformed brain extraction routines from popular MRI packages (AFNI, FSL, and FreeSurfer) across a heterogeneous sample from multiple sites in the PRIME-DE with less computational cost (20 s~10 min). We also demonstrated the transfer-learning process enables the macaque model to be updated for use with scans from chimpanzees, marmosets, and other mammals (e.g. pig). Our model, code, and the skull-stripped mask repository of 136 macaque monkeys are publicly available for unrestricted use by the neuroimaging community at https://github.com/HumanBrainED/NHP-BrainExtraction.

AAAI Conference 2021 Conference Paper

Very Important Person Localization in Unconstrained Conditions: A New Benchmark

  • Xiao Wang
  • Zheng Wang
  • Toshihiko Yamasaki
  • Wenjun Zeng

This paper presents a new high-quality dataset for Very Important Person Localization (VIPLoc), named Unconstrained-7k. Generally, existing datasets are: 1) limited in scale; 2) built under simple and constrained conditions, where the number of disturbing non-VIPs is not large, the scene is relatively simple, and the face of VIP is always in frontal view and salient. To tackle these problems, the proposed Unconstrained-7k dataset is featured in two aspects. First, it contains over 7, 000 annotated images, making it the largest VIPLoc dataset under unconstrained conditions to date. Second, our dataset is collected freely on the Internet, including multiple scenes, where images are in unconstrained conditions. VIPs in the new dataset are in different settings, e. g. , large view variation, varying sizes, occluded, and complex scenes. Meanwhile, each image has more persons (> 20), making the dataset more challenging. As a minor contribution, motivated by the observation that VIPs are highly related to not only neighbors but also iconic objects, this paper proposes a Joint Social Relation and Individual Interaction Graph Neural Networks (JSRII-GNN) for VIPLoc. Experiments show that the JSRII-GNN yields competitive accuracy on NCAA (National Collegiate Athletic Association), MS (Multi-scene), and Unconstrained-7k datasets. https: //github. com/xiaowang1516/VIPLoc.

NeurIPS Conference 2020 Conference Paper

BAIL: Best-Action Imitation Learning for Batch Deep Reinforcement Learning

  • Xinyue Chen
  • Zijian Zhou
  • Zheng Wang
  • CHE WANG
  • Yanqiu Wu
  • Keith Ross

There has recently been a surge in research in batch Deep Reinforcement Learning (DRL), which aims for learning a high-performing policy from a given dataset without additional interactions with the environment. We propose a new algorithm, Best-Action Imitation Learning (BAIL), which strives for both simplicity and performance. BAIL learns a V function, uses the V function to select actions it believes to be high-performing, and then uses those actions to train a policy network using imitation learning. For the MuJoCo benchmark, we provide a comprehensive experimental study of BAIL, comparing its performance to four other batch Q-learning and imitation-learning schemes for a large variety of batch datasets. Our experiments show that BAIL's performance is much higher than the other schemes, and is also computationally much faster than the batch Q-learning schemes.

IJCAI Conference 2020 Conference Paper

Beyond Intra-modality: A Survey of Heterogeneous Person Re-identification

  • Zheng Wang
  • Zhixiang Wang
  • Yinqiang Zheng
  • Yang Wu
  • Wenjun Zeng
  • Shin'ichi Satoh

An efficient and effective person re-identification (ReID) system relieves the users from painful and boring video watching and accelerates the process of video analysis. Recently, with the explosive demands of practical applications, a lot of research efforts have been dedicated to heterogeneous person re-identification (Hetero-ReID). In this paper, we provide a comprehensive review of state-of-the-art Hetero-ReID methods that address the challenge of inter-modality discrepancies. According to the application scenario, we classify the methods into four categories --- low-resolution, infrared, sketch, and text. We begin with an introduction of ReID, and make a comparison between Homogeneous ReID (Homo-ReID) and Hetero-ReID tasks. Then, we describe and compare existing datasets for performing evaluations, and survey the models that have been widely employed in Hetero-ReID. We also summarize and compare the representative approaches from two perspectives, i. e. , the application scenario and the learning pipeline. We conclude by a discussion of some future research directions. Follow-up updates are available at https: //github. com/lightChaserX/Awesome-Hetero-reID

IJCAI Conference 2020 Conference Paper

Discriminative Feature Selection via A Structured Sparse Subspace Learning Module

  • Zheng Wang
  • Feiping Nie
  • Lai Tian
  • Rong Wang
  • Xuelong Li

In this paper, we first propose a novel Structured Sparse Subspace Learning S^3L module to address the long-standing subspace sparsity issue. Elicited by proposed module, we design a new discriminative feature selection method, named Subspace Sparsity Discriminant Feature Selection S^2DFS which enables the following new functionalities: 1) Proposed S^2DFS method directly joints trace ratio objective and structured sparse subspace constraint via L2, 0-norm to learn a row-sparsity subspace, which improves the discriminability of model and overcomes the parameter-tuning trouble with comparison to the methods used L2, 1-norm regularization; 2) An alternative iterative optimization algorithm based on the proposed S^3L module is presented to explicitly solve the proposed problem with a closed-form solution and strict convergence proof. To our best knowledge, such objective function and solver are first proposed in this paper, which provides a new though for the development of feature selection methods. Extensive experiments conducted on several high-dimensional datasets demonstrate the discriminability of selected features via S^2DFS with comparison to several related SOTA feature selection methods. Source matlab code: https: //github. com/StevenWangNPU/L20-FS.

AAAI Conference 2020 Conference Paper

Learning Cross-Aligned Latent Embeddings for Zero-Shot Cross-Modal Retrieval

  • Kaiyi Lin
  • Xing Xu
  • Lianli Gao
  • Zheng Wang
  • Heng Tao Shen

Zero-Shot Cross-Modal Retrieval (ZS-CMR) is an emerging research hotspot that aims to retrieve data of new classes across different modality data. It is challenging for not only the heterogeneous distributions across different modalities, but also the inconsistent semantics across seen and unseen classes. A handful of recently proposed methods typically borrow the idea from zero-shot learning, i. e. , exploiting word embeddings of class labels (i. e. , class-embeddings) as common semantic space, and using generative adversarial network (GAN) to capture the underlying multimodal data structures, as well as strengthen relations between input data and semantic space to generalize across seen and unseen classes. In this paper, we propose a novel method termed Learning Cross-Aligned Latent Embeddings (LCALE) as an alternative to these GAN based methods for ZS-CMR. Unlike using the class-embeddings as the semantic space, our method seeks for a shared low-dimensional latent space of input multimodal features and class-embeddings by modalityspecific variational autoencoders. Notably, we align the distributions learned from multimodal input features and from class-embeddings to construct latent embeddings that contain the essential cross-modal correlation associated with unseen classes. Effective cross-reconstruction and crossalignment criterions are further developed to preserve classdiscriminative information in latent space, which benefits the efficiency for retrieval and enable the knowledge transfer to unseen classes. We evaluate our model using four benchmark datasets on image-text retrieval tasks and one largescale dataset on image-sketch retrieval tasks. The experimental results show that our method establishes the new state-ofthe-art performance for both tasks on all datasets.

AAAI Conference 2020 Conference Paper

Mining on Heterogeneous Manifolds for Zero-Shot Cross-Modal Image Retrieval

  • Fan Yang
  • Zheng Wang
  • Jing Xiao
  • Shin'ichi Satoh

Most recent approaches for the zero-shot cross-modal image retrieval map images from different modalities into a uniform feature space to exploit their relevance by using a pre-trained model. Based on the observation that manifolds of zero-shot images are usually deformed and incomplete, we argue that the manifolds of unseen classes are inevitably distorted during the training of a two-stream model that simply maps images from different modalities into a uniform space. This issue directly leads to poor cross-modal retrieval performance. We propose a bi-directional random walk scheme to mining more reliable relationships between images by traversing heterogeneous manifolds in the feature space of each modality. Our proposed method benefits from intra-modal distributions to alleviate the interference caused by noisy similarities in the cross-modal feature space. As a result, we achieved great improvement in the performance of the thermal v. s. visible image retrieval task. The code of this paper: https: //github. com/fyang93/cross-modal-retrieval

AAAI Conference 2020 Conference Paper

Multi-Type Self-Attention Guided Degraded Saliency Detection

  • Ziqi Zhou
  • Zheng Wang
  • Huchuan Lu
  • Song Wang
  • Meijun Sun

Existing saliency detection techniques are sensitive to image quality and perform poorly on degraded images. In this paper, we systematically analyze the current status of the research on detecting salient objects from degraded images and then propose a new multi-type self-attention network, namely MSANet, for degraded saliency detection. The main contributions include: 1) Applying attention transfer learning to promote semantic detail perception and internal feature mining of the target network on degraded images; 2) Developing a multi-type self-attention mechanism to achieve the weight recalculation of multi-scale features. By computing global and local attention scores, we obtain the weighted features of different scales, effectively suppress the interference of noise and redundant information, and achieve a more complete boundary extraction. The proposed MSANet converts low-quality inputs to high-quality saliency maps directly in an end-to-end fashion. Experiments on seven widely-used datasets show that our approach produces good performance on both clear and degraded images.

AAAI Conference 2020 Short Paper

Supervised Discovery of Unknown Unknowns through Test Sample Mining (Student Abstract)

  • Zheng Wang
  • Bruno Abrahao
  • Ece Kamar

Given a fixed hypothesis space, defined to model class structure in a particular domain of application, unknown unknowns (u. u. s) are data examples that form classes in the feature space whose structure is not represented in a trained model. Accordingly, this leads to incorrect class prediction with high confidence, which represents one of the major sources of blind spots in machine learning. Our method seeks to reduce the structural mismatch between the training model and that of the target space in a supervised way. We illuminate further structure through cross-validation on a modified training model, set up to mine and trap u. u. s in a marginal training class, created from examples of a random sample of the test set. Contrary to previous approaches, our method simplifies the solution, as it does not rely on budgeted queries to an Oracle whose outcomes inform adjustments to training. In addition, our empirically results exhibit consistent performance improvements over baselines, on both synthetic and real-world data sets.

IJCAI Conference 2020 Conference Paper

When Pedestrian Detection Meets Nighttime Surveillance: A New Benchmark

  • Xiao Wang
  • Jun Chen
  • Zheng Wang
  • Wu Liu
  • Shin'ichi Satoh
  • Chao Liang
  • Chia-Wen Lin

Pedestrian detection at nighttime is a crucial and frontier problem in surveillance, but has not been well explored by the computer vision and artificial intelligence communities. Most of existing methods detect pedestrians under favorable lighting conditions (e. g. daytime) and achieve promising performances. In contrast, they often fail under unstable lighting conditions (e. g. nighttime). Night is a critical time for criminal suspects to act in the field of security. The existing nighttime pedestrian detection dataset is captured by a car camera, specially designed for autonomous driving scenarios. The dataset for nighttime surveillance scenario is still vacant. There are vast differences between autonomous driving and surveillance, including viewpoint and illumination. In this paper, we build a novel pedestrian detection dataset from the nighttime surveillance aspect: NightSurveillance1. As a benchmark dataset for pedestrian detection at nighttime, we compare the performances of state-of-the-art pedestrian detectors and the results reveal that the methods cannot solve all the challenging problems of NightSurveillance. We believe that NightSurveillance can further advance the research of pedestrian detection, especially in the field of surveillance security at nighttime.

AAAI Conference 2019 Conference Paper

Composite Binary Decomposition Networks

  • You Qiaoben
  • Zheng Wang
  • Jianguo Li
  • Yinpeng Dong
  • Yu-Gang Jiang
  • Jun Zhu

Binary neural networks have great resource and computing efficiency, while suffer from long training procedure and non-negligible accuracy drops, when comparing to the fullprecision counterparts. In this paper, we propose the composite binary decomposition networks (CBDNet), which first compose real-valued tensor of each layer with a limited number of binary tensors, and then decompose some conditioned binary tensors into two low-rank binary tensors, so that the number of parameters and operations are greatly reduced comparing to the original ones. Experiments demonstrate the effectiveness of the proposed method, as CBDNet can approximate image classification network ResNet-18 using 5. 25 bits, VGG-16 using 5. 47 bits, DenseNet-121 using 5. 72 bits, object detection networks SSD300 using 4. 38 bits, and semantic segmentation networks SegNet using 5. 18 bits, all with minor accuracy drops. 1

AAAI Conference 2019 Conference Paper

Dual-View Ranking with Hardness Assessment for Zero-Shot Learning

  • Yuchen Guo
  • Guiguang Ding
  • Jungong Han
  • Xiaohan Ding
  • Sicheng Zhao
  • Zheng Wang
  • Chenggang Yan
  • Qionghai Dai

Zero-shot learning (ZSL) is to build recognition models for previously unseen target classes which have no labeled data for training by transferring knowledge from some other related auxiliary source classes with abundant labeled samples to the target ones with class attributes as the bridge. The key is to learn a similarity based ranking function between samples and class labels using the labeled source classes so that the proper (unseen) class label for a test sample can be identified by the function. In order to learn the function, single-view ranking based loss is widely used which aims to rank the true label prior to the other labels for a training sample. However, we argue that the ranking can be performed from the other view, which aims to place the images belonging to a label before the images from the other classes. Motivated by it, we propose a novel DuAl-view RanKing (DARK) loss for zeroshot learning simultaneously ranking labels for an image by point-to-point metric and ranking images for a label by pointto-set metric, which is capable of better modeling the relationship between images and classes. In addition, we also notice that previous ZSL approaches mostly fail to well exploit the hardness of training samples, either using only very hard ones or using all samples indiscriminately. In this work, we also introduce a sample hardness assessment method to ZSL which assigns different weights to training samples based on their hardness, which leads to a more accurate and robust ZSL model. Experiments on benchmarks demonstrate that DARK outperforms the state-of-the-arts for (generalized) ZSL.

IJCAI Conference 2019 Conference Paper

Flexible Multi-View Representation Learning for Subspace Clustering

  • Ruihuang Li
  • Changqing Zhang
  • Qinghua Hu
  • Pengfei Zhu
  • Zheng Wang

In recent years, numerous multi-view subspace clustering methods have been proposed to exploit the complementary information from multiple views. Most of them perform data reconstruction within each single view, which makes the subspace representation unpromising and thus can not well identify the underlying relationships among data. In this paper, we propose to conduct subspace clustering based on Flexible Multi-view Representation (FMR) learning, which avoids using partial information for data reconstruction. The latent representation is flexibly constructed by enforcing it to be close to different views, which implicitly makes it more comprehensive and well-adapted to subspace clustering. With the introduction of kernel dependence measure, the latent representation can flexibly encode complementary information from different views and explore nonlinear, high-order correlations among these views. We employ the Alternating Direction Minimization (ADM) method to solve our problem. Empirical studies on real-world datasets show that our method achieves superior clustering performance over other state-of-the-art methods.

AAAI Conference 2019 Conference Paper

Lattice CNNs for Matching Based Chinese Question Answering

  • Yuxuan Lai
  • Yansong Feng
  • Xiaohan Yu
  • Zheng Wang
  • Kun Xu
  • Dongyan Zhao

Short text matching often faces the challenges that there are great word mismatch and expression diversity between the two texts, which would be further aggravated in languages like Chinese where there is no natural space to segment words explicitly. In this paper, we propose a novel lattice based CNN model (LCNs) to utilize multi-granularity information inherent in the word lattice while maintaining strong ability to deal with the introduced noisy information for matching based question answering in Chinese. We conduct extensive experiments on both document based question answering and knowledge based question answering tasks, and experimental results show that the LCNs models can significantly outperform the state-of-the-art matching models and strong baselines by taking advantages of better ability to distill rich but discriminative information from the word lattice input.

IJCAI Conference 2019 Conference Paper

Relation-Aware Entity Alignment for Heterogeneous Knowledge Graphs

  • Yuting Wu
  • Xiao Liu
  • Yansong Feng
  • Zheng Wang
  • Rui Yan
  • Dongyan Zhao

Entity alignment is the task of linking entities with the same real-world identity from different knowledge graphs (KGs), which has been recently dominated by embedding-based methods. Such approaches work by learning KG representations so that entity alignment can be performed by measuring the similarities between entity embeddings. While promising, prior works in the field often fail to properly capture complex relation information that commonly exists in multi-relational KGs, leaving much room for improvement. In this paper, we propose a novel Relation-aware Dual-Graph Convolutional Network (RDGCN) to incorporate relation information via attentive interactions between the knowledge graph and its dual relation counterpart, and further capture neighboring structures to learn better entity representations. Experiments on three real-world cross-lingual datasets show that our approach delivers better and more robust results over the state-of-the-art alignment methods by learning better KG representations.

IJCAI Conference 2018 Conference Paper

Cascaded SR-GAN for Scale-Adaptive Low Resolution Person Re-identification

  • Zheng Wang
  • Mang Ye
  • Fan Yang
  • Xiang Bai
  • Shin'ichi Satoh

Person re-identification (REID) is an important task in video surveillance and forensics applications. Most of previous approaches are based on a key assumption that all person images have uniform and sufficiently high resolutions. Actually, various low-resolutions and scale mismatching always exist in open world REID. We name this kind of problem as Scale-Adaptive Low Resolution Person Re-identification (SALR-REID). The most intuitive way to address this problem is to increase various low-resolutions (not only low, but also with different scales) to a uniform high-resolution. SR-GAN is one of the most competitive image super-resolution deep networks, designed with a fixed upscaling factor. However, it is still not suitable for SALR-REID task, which requires a network not only synthesizing high-resolution images with different upscaling factors, but also extracting discriminative image feature for judging person’s identity. (1) To promote the ability of scale-adaptive upscaling, we cascade multiple SRGANs in series. (2) To supplement the ability of image feature representation, we plug-in a reidentification network. With a unified formulation, a Cascaded Super-Resolution GAN (CSR-GAN) framework is proposed. Extensive evaluations on two simulated datasets and one public dataset demonstrate the advantages of our method over related state-of-the-art methods.

TCS Journal 2018 Journal Article

Data fusion using Hilbert space multi-dimensional models

  • Jerome Busemeyer
  • Zheng Wang

General procedures for constructing, estimating, and testing Hilbert space multi-dimensional (HSM) models, built from quantum probability theory, are presented. HSM models can be applied to collections of K different contingency tables obtained from a set of p variables that are measured under different contexts. A context is defined by the measurement of a subset of the p variables that are used to form a table. HSM models provide a representation of the collection of K tables in a low dimensional vector space, even when no single joint probability distribution across the p variables exists. HSM models produce parameter estimates that provide a simple and informative interpretation of the complex collection of tables.

IJCAI Conference 2018 Conference Paper

Evaluating Brush Movements for Chinese Calligraphy: A Computer Vision Based Approach

  • Pengfei Xu
  • Lei Wang
  • Ziyu Guan
  • Xia Zheng
  • Xiaojiang Chen
  • Zhanyong Tang
  • Dingyi Fang
  • Xiaoqing Gong

Chinese calligraphy is a popular, highly esteemed art form in the Chinese cultural sphere and worldwide. Ink brushes are the traditional writing tool for Chinese calligraphy and the subtle nuances of brush movements have a great impact on the aesthetics of the written characters. However, mastering the brush movement is a challenging task for many calligraphy learners as it requires many years’ practice and expert supervision. This paper presents a novel approach to help Chinese calligraphy learners to quantify the quality of brush movements without expert involvement. Our approach extracts the brush trajectories from a video stream; it then compares them with example templates of reputed calligraphers to produce a score for the writing quality. We achieve this by first developing a novel neural network to extract the spatial and temporal movement features from the video stream. We then employ methods developed in the computer vision and signal processing domains to track the brush movement trajectory and calculate the score. We conducted extensive experiments and user studies to evaluate our approach. Experimental results show that our approach is highly accurate in identifying brush movements, yielding an average accuracy of 90%, and the generated score is within 3% of errors when compared to the one given by human experts.

AAAI Conference 2018 Conference Paper

RSDNE: Exploring Relaxed Similarity and Dissimilarity from Completely-Imbalanced Labels for Network Embedding

  • Zheng Wang
  • Xiaojun Ye
  • Chaokun Wang
  • Yuexin Wu
  • Changping Wang
  • Kaiwen Liang

Network embedding, aiming to project a network into a lowdimensional space, is increasingly becoming a focus of network research. Semi-supervised network embedding takes advantage of labeled data, and has shown promising performance. However, existing semi-supervised methods would get unappealing results in the completely-imbalanced label setting where some classes have no labeled nodes at all. To alleviate this, we propose a novel semi-supervised network embedding method, termed Relaxed Similarity and Dissimilarity Network Embedding (RSDNE). Specifically, to benefit from the completely-imbalanced labels, RSDNE guarantees both intra-class similarity and inter-class dissimilarity in an approximate way. Experimental results on several real-world datasets demonstrate the superiority of the proposed method.

AAAI Conference 2018 Conference Paper

Scale Up Event Extraction Learning via Automatic Training Data Generation

  • Ying Zeng
  • Yansong Feng
  • Rong Ma
  • Zheng Wang
  • Rui Yan
  • Chongde Shi
  • Dongyan Zhao

The task of event extraction has long been investigated in a supervised learning paradigm, which is bound by the number and the quality of the training instances. Existing training data must be manually generated through a combination of expert domain knowledge and extensive human involvement. However, due to drastic efforts required in annotating text, the resultant datasets are usually small, which severally affects the quality of the learned model, making it hard to generalize. Our work develops an automatic approach for generating training data for event extraction. Our approach allows us to scale up event extraction training instances from thousands to hundreds of thousands, and it does this at a much lower cost than a manual approach. We achieve this by employing distant supervision to automatically create event annotations from unlabelled text using existing structured knowledge bases or tables. We then develop a neural network model with post inference to transfer the knowledge extracted from structured knowledge bases to automatically annotate typed events with corresponding arguments in text. We evaluate our approach by using the knowledge extracted from Freebase to label texts from Wikipedia articles. Experimental results show that our approach can generate a large number of highquality training instances. We show that this large volume of training data not only leads to a better event extractor, but also allows us to detect multiple typed events.

AAAI Conference 2018 Conference Paper

Video-Based Person Re-Identification via Self Paced Weighting

  • Wenjun Huang
  • Chao Liang
  • Yi Yu
  • Zheng Wang
  • Weijian Ruan
  • Ruimin Hu

Person re-identification (re-id) is a fundamental technique to associate various person images, captured by different surveillance cameras, to the same person. Compared to the single image based person re-id methods, video-based person re-id has attracted widespread attentions because extra space-time information and more appearance cues that can be used to greatly improve the matching performance. However, most existing video-based person re-id methods equally treat all video frames, ignoring their quality discrepancy caused by object occlusion and motions, which is a common phenomenon in real surveillance scenario. Based on this finding, we propose a novel video-based person re-id method via self paced weighting (SPW). Firstly, we propose a self paced outlier detection method to evaluate the noise degree of video sub sequences. Thereafter, a weighted multi-pair distance metric learning approach is adopted to measure the distance of two person image sequences. Experimental results on two public datasets demonstrate the superiority of the proposed method over current state-of-the-art work.

IJCAI Conference 2018 Conference Paper

Visible Thermal Person Re-Identification via Dual-Constrained Top-Ranking

  • Mang Ye
  • Zheng Wang
  • Xiangyuan Lan
  • Pong C. Yuen

Cross-modality person re-identification between the thermal and visible domains is extremely important for night-time surveillance applications. Existing works in this filed mainly focus on learning sharable feature representations to handle the cross-modality discrepancies. However, besides the cross-modality discrepancy caused by different camera spectrums, visible thermal person re-identification also suffers from large cross-modality and intra-modality variations caused by different camera views and human poses. In this paper, we propose a dual-path network with a novel bi-directional dual-constrained top-ranking loss to learn discriminative feature representations. It is advantageous in two aspects: 1) end-to-end feature learning directly from the data without extra metric learning steps, 2) it simultaneously handles the cross-modality and intra-modality variations to ensure the discriminability of the learnt representations. Meanwhile, identity loss is further incorporated to model the identity-specific information to handle large intra-class variations. Extensive experiments on two datasets demonstrate the superior performance compared to the state-of-the-arts.

AAAI Conference 2017 Conference Paper

Efficient Delivery Policy to Minimize User Traffic Consumption in Guaranteed Advertising

  • Jia Zhang
  • Zheng Wang
  • Qian Li
  • Jialin Zhang
  • Yanyan Lan
  • Qiang Li
  • Xiaoming Sun

In this work, we study the guaranteed delivery model which is widely used in online advertising. In the guaranteed delivery scenario, ad exposures (which are also called impressions in some works) to users are guaranteed by contracts signed in advance between advertisers and publishers. A crucial problem for the advertising platform is how to fully utilize the valuable user traffic to generate as much as possible revenue. Different from previous works which usually minimize the penalty of unsatisfied contracts and some other cost (e. g. representativeness), we propose the novel consumption minimization model, in which the primary objective is to minimize the user traffic consumed to satisfy all contracts. Under this model, we develop a near optimal method to deliver ads for users. The main advantage of our method lies in that it consumes nearly as least as possible user traffic to satisfy all contracts, therefore more contracts can be accepted to produce more revenue. It also enables the publishers to estimate how much user traffic is redundant or short so that they can sell or buy this part of traffic in bulk in the exchange market. Furthermore, it is robust with regard to priori knowledge of user type distribution. Finally, the simulation shows that our method outperforms the traditional state-of-the-art methods.

AAAI Conference 2017 Conference Paper

Multiple Source Detection without Knowing the Underlying Propagation Model

  • Zheng Wang
  • Chaokun Wang
  • Jisheng Pei
  • Xiaojun Ye

Information source detection, which is the reverse problem of information diffusion, has attracted considerable research effort recently. Most existing approaches assume that the underlying propagation model is fixed and given as input, which may limit their application range. In this paper, we study the multiple source detection problem when the underlying propagation model is unknown. Our basic idea is source prominence, namely the nodes surrounded by larger proportions of infected nodes are more likely to be infection sources. As such, we propose a multiple source detection method called Label Propagation based Source Identification (LPSI). Our method lets infection status iteratively propagate in the network as labels, and finally uses local peaks of the label propagation result as source nodes. In addition, both the convergent and iterative versions of LPSI are given. Extensive experiments are conducted on several real-world datasets to demonstrate the effectiveness of the proposed method.

IJCAI Conference 2016 Conference Paper

Causality Based Propagation History Ranking in Social Networks

  • Zheng Wang
  • Chaokun Wang
  • Jisheng Pei
  • Xiaojun Ye
  • Philip S. Yu

In social network sites (SNS), propagation histories which record the information diffusion process can be used to explain to users what happened in their networks. However, these histories easily grow in size and complexity, limiting their intuitive understanding by users. To reduce this information overload, in this paper, we present the problem of propagation history ranking. The goal is to rank participant edges/nodes by their contribution to the diffusion. Firstly, we discuss and adapt Difference of Causal Effects (DCE) as the ranking criterion. Then, to avoid the complex calculation of DCE, we propose a resp-cap ranking strategy by adopting two indicators. The first is responsibility which captures the necessary face of causal effects. We further give an approximate algorithm for this indicator. The second is capability which is defined to capture the sufficient face of causal effects. Finally, promising experimental results are presented to verify the feasibility of our method.

IJCAI Conference 2016 Conference Paper

Scale-Adaptive Low-Resolution Person Re-Identification via Learning a Discriminating Surface

  • Zheng Wang
  • Ruimin Hu
  • Yi Yu
  • Junjun Jiang
  • Chao Liang
  • Jinqiao Wang

Person re-identification, as an important task in video surveillance and forensics applications, has been widely studied. But most of previous approaches are based on the key assumption that images for comparison have the same resolution and a uniform scale. Some recent works investigate how to match low resolution query images against high resolution gallery images, but still assume that the low-resolution query images have the same scale. In real scenarios, person images may not only be with low-resolution but also have different scales. Through investigating the distance variation behavior by changing image scales, we observe that scale-distance functions, generated by image pairs under different scales from the same person or different persons, are distinguishable and can be classified as feasible (for a pair of images from the same person) or infeasible (for a pair of images from different persons). The scale-distance functions are further represented by parameter vectors in the scale-distance function space. On this basis, we propose to learn a discriminating surface separating these feasible and infeasible functions in the scale-distance function space, and use it for reidentifying persons. Experimental results on two simulated datasets and one public dataset demonstrate the effectiveness of the proposed framework.

IJCAI Conference 2009 Conference Paper

  • Zheng Wang
  • Yangqiu Song
  • Changshui Zhang

In machine learning problems, labeled data are often in short supply. One of the feasible solution for this problem is transfer learning. It can make use of the labeled data from other domain to discriminate those unlabeled data in the target domain. In this paper, we propose a transfer learning framework based on similarity matrix approximation to tackle such problems. Two practical algorithms are proposed, which are the label propagation and the similarity propagation. In these methods, we build a hybrid graph based on all available data. Then the information is transferred cross domains through alternatively constructing the similarity matrix for different part of the graph. Among all related methods, similarity propagation approach can make maximum use of all available similarity information across domains. This leads to more efficient transfer and better learning result. The experiment on real world text mining applications demonstrates the promise and effectiveness of our algorithms.

ICRA Conference 2009 Conference Paper

Regulation control of underactuated mechanical systems based on a new matching equation of port-controlled hamiltonian systems

  • Zheng Wang
  • Peter B. Goldsmith
  • Jason Jianjun Gu

We consider the control of Port-Controlled Hamiltonian (PCH) systems, which are a generalization of Euler-Lagrange Systems. A new matching equation for PCH systems is developed so that interconnection damping assignment passivity-based control (IDA-PBC) can be extended to the regulation of some underactuated PCH systems whosekinetic energy must be modified. A simple underactuated mechanical system (the inertial wheel pendulum) is used to demonstrate the effectiveness of the proposed method.

YNIMG Journal 2006 Journal Article

Linear aspects of transformation from interictal epileptic discharges to BOLD fMRI signals in an animal model of occipital epilepsy

  • Seyed M. Mirsattari
  • Zheng Wang
  • John R. Ives
  • Frank Bihari
  • L. Stan Leung
  • Robert Bartha
  • Ravi S. Menon

Epileptic disorders manifest with seizures and interictal epileptic discharges (IEDs). The hemodynamic changes that accompany IEDs are poorly understood and may be critical for understanding epileptogenesis. Despite a known linear coupling of the neurovascular elements in normal brain tissues, previous simultaneous electroencephalography (EEG)–functional magnetic resonance imaging (fMRI) studies have shown variable correlations between epileptic discharges and blood oxygenation level-dependent (BOLD) response, partly because most previous studies assumed particular hemodynamic properties in normal brain tissue. The occurrence of IEDs in human subjects is unpredictable. Therefore, an animal model with reproducible stereotyped IEDs was developed by the focal injection of penicillin into the right occipital cortex of rats anesthetized with isoflurane. Simultaneous EEG–fMRI was used to study the hemodynamic changes during IEDs. A hybrid of temporal independent component analysis (ICA) of EEG and spatial ICA of fMRI data was used to correlate BOLD fMRI signals with IEDs. A linear autoregression with exogenous input (ARX) model was used to estimate the hemodynamic impulse response function (HIRF) based on the data from simultaneous EEG–fMRI measurement. Changes in the measured BOLD signal from the right primary visual cortex and bilateral visual association cortices were consistently coupled to IEDs. The linear ARX model was applied here to confirm that a linear transform can be used to study the correlation between BOLD signal and its corresponding neural activity in this animal model of occipital epilepsy.

v2026.09.13