Arrow Research search

Author name cluster

Shiqi Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

25 papers
1 author row

Possible papers

25

EAAI Journal 2026 Journal Article

Damage assessment of thermal-humidity-mechanical coupling field of early-age concrete based on adaptive physics informed neural network

  • Shiqi Wang
  • Yue Chen
  • Jinlong Liu
  • Fangzhou Lin
  • Lei Xu

The crack-damage resistance of early-age concrete is affected by multiple factors such as hydration, self-drying, temperature and humidity diffusion, and material properties, which are difficult to be accurately evaluated by traditional theories and numerical models. This paper proposed an adaptive physics-informed back propagation neural network (BPINN) to accurately evaluate the damage of early-age concrete under multi-physics field coupling. The temperature and humidity diffusion and shrinkage models are used as physics loss functions to guide the model in learning the physics laws. Furthermore, time-dependent factor weights are constructed for both the physics and boundary equations to enhance the model's ability to learn the spatiotemporal feature distribution of the sampling points. BPINN effectively simulates the influence of concrete strength grade and boundary conditions on temperature and humidity diffusion, with the average error less than 5 %. The LOSS differences of traditional physics informed neural network (PINN) and BPINN in time step, activation function, hidden layer and neuron number are quantified. Compared with the traditional PINN, the LOSS of BPINN is reduced by 62. 4 %. On this basis, the predictive performance of BPINN and four types of data-driven models is compared to verify the influence of physics constraint, as BPINN has the smallest statistical loss and data discreteness. The model proposed in this paper enhances the learning ability of spatial-temporal features by balancing the weight between boundary and physics equations, providing new insights for the thermo-hygro-mechanical coupling field in early-age concrete.

AAAI Conference 2026 Conference Paper

DeepRAHT: Learning Predictive RAHT for Point Cloud Attribute Compression

  • Chunyang Fu
  • Tai Qin
  • Shiqi Wang
  • Zhu Li

Regional Adaptive Hierarchical Transform (RAHT) is an effective point cloud attribute compression (PCAC) method. However, its application in deep learning lacks research. In this paper, we propose an end-to-end RAHT framework for lossy PCAC based on the sparse tensor, called DeepRAHT. The RAHT transform is performed within the learning reconstruction process, without requiring manual RAHT for pre-processing. We also introduce the predictive RAHT to reduce bitrates and design a learning-based prediction model to enhance the performance. Moreover, we devise a bitrate proxy that applies run-length coding to entropy model, achieving seamless variable-rate coding and improving the robustness. DeepRAHT is a reversible and distortion-controllable framework, ensuring its lower bound performance and offering significant application potential. The experiments demonstrate that DeepRAHT is a high-performance, faster, and more robust solution than the baseline methods.

EAAI Journal 2026 Journal Article

Engineering graphene and carbon nanotube reinforced cement composites for intelligent pavement as noise-resistant wireless monitoring sensors

  • Yucheng Fan
  • Chuang Feng
  • Luming Shen
  • Hongru Xiao
  • Shiqi Wang
  • Jinlong Liu
  • Wengui Li

Hybrid graphene nanoplatelet/carbon nanotube reinforced cement-based sensors (GNP/CNTRCS) offer significant advantages for self-sensing pavements in road infrastructure monitoring. However, real-world engineering applications face challenges such as the necessity of wired electrical resistance measurement and the inherent nonlinear and noisy piezoresistive response, which hinder automated data processing. To address these issues, this study develops a four-electrode wireless pavement monitoring system integrating GNP/CNTRCS with machine learning (ML) algorithms for noise-resistant vehicle classification and speed estimation. Laboratory and field tests are conducted to validate the piezoresistive performance and noisy condition of the GNP/CNTRCS, with fractional change in resistivity (FCR) ranging from −40 % for cars to −3 % for pedestrians and 260 sets of time-series data are collected under different vehicle loads and speeds for ML. Innovatively employing continuous wavelet transform (CWT) for noise-resistant feature extraction, the convolutional neural network (CNN) achieves 96. 1 % classification accuracy, maintaining 91. 8 % under noise. Furthermore, a proposed hierarchical regression strategy establishes state-of-the-art (SOTA) performance for speed estimation with an R2 of 0. 851, sustaining 0. 813 under noise. This work provides a scalable and low-maintenance solution for intelligent transportation infrastructure.

AAAI Conference 2026 Conference Paper

When Privacy Meets Recovery: The Overlooked Half of Surrogate-Driven Privacy Preservation for MLLM Editing

  • Siyuan Xu
  • Yibing Liu
  • Peilin Chen
  • Yung-Hui Li
  • Shiqi Wang
  • Sam Kwong

Privacy leakage in Multimodal Large Language Models (MLLMs) has long been an intractable problem. Existing studies, though effectively obscure private information in MLLMs, often overlook the evaluation of authenticity and recovery quality of user privacy. To this end, this work uniquely focuses on the critical challenge of how to restore surrogate-driven protected data in diverse MLLM scenarios. We first bridge this research gap by contributing the SPPE (Surrogate Privacy Protected Editable) dataset, which includes a wide range of privacy categories and user instructions to simulate real MLLM applications. This dataset offers protected surrogates alongside their various MLLM-edited versions, thus enabling the direct assessment of privacy recovery quality. By formulating privacy recovery as a guided generation task conditioned on complementary multimodal signals, we further introduce a unified approach that reliably reconstructs private content while preserving the fidelity of MLLM-generated edits. The experiments on both SPPE and InstructPix2Pix further show that our approach generalizes well across diverse visual content and editing tasks, achieving a strong balance between privacy protection and MLLM usability.

EAAI Journal 2025 Journal Article

A Bayesian-physical informed conditional tabular generative adversarial network framework for low-carbon concrete data augmentation and hyperparameter optimization

  • Shiqi Wang
  • Peng Xia
  • Fuyuan Gong
  • Yuxi Zhao
  • Peng Lin

Data shortage, unbalanced data distribution and multi-factor coupling mechanism of materials all increase the difficulty of design. This paper proposed a physical constraint-conditional generative adversarial network (PI-CTGAN) to solve the above problems. Firstly, residual layers are added to the generator to enhance model stability. The continuous differentiable function and wasserstein_distance were constructed to embed physical loss functions into the generator, including water-cement ratio, supplementary cementitious materials (SCMs) ratio, and aggregate water absorption ratio. Based on this, Bayesian optimization (BO) was used to optimize the hyperparameters of PI-CTGAN. The results showed that BO effectively optimized the model's hyperparameters, reducing the total of Kolmogorov-Smirnov distribution (K-Stot) of the generated dataset by 27. 9 %. Additionally, applying physical loss to the optimized model can improve the model's data recognition capability, with generation accuracy increasing by 16. 2 %. The influence of physical weight (WPI) and activation functions on data generation quality was compared. Revealing that K-Stot initially decreased and then increased with WPI. The model using the rectified linear unit exhibited the best generation accuracy, with a K-Stot of 0. 37 and anomaly data ratios (water-binder ratio and supplementary cementitious materials/binder ratio) of 11. 67 % and 3. 2 %, respectively. The generated data and the experimental data show statistical similarity and conform to the physical law. By constructing the target dataset and related physical constraints, the proposed PI-CTGAN can effectively solve the issues of multi-source data sets with data shortage and imbalance, thereby providing numerous datasets to guide engineering design.

AAAI Conference 2025 Conference Paper

AI-generated Image Quality Assessment in Visual Communication

  • Yu Tian
  • Yixuan Li
  • Baoliang Chen
  • Hanwei Zhu
  • Shiqi Wang
  • Sam Kwong

Assessing the quality of artificial intelligence-generated images (AIGIs) plays a crucial role in their application in real-world scenarios. However, traditional image quality assessment (IQA) algorithms primarily focus on low-level visual perception, while existing IQA works on AIGIs overemphasize the generated content itself, neglecting its effectiveness in real-world applications. To bridge this gap, we propose AIGI-VC, a quality assessment database for AI-Generated Images in Visual Communication, which studies the communicability of AIGIs in the advertising field from the perspectives of information clarity and emotional interaction. The dataset consists of 2,500 images spanning 14 advertisement topics and 8 emotion types. It provides coarse-grained human preference annotations and fine-grained preference descriptions, benchmarking the abilities of IQA methods in preference prediction, interpretation, and reasoning. We conduct an empirical study of existing representative IQA methods and large multi-modal models on the AIGI-VC dataset, uncovering their strengths and weaknesses.

NeurIPS Conference 2025 Conference Paper

Large Language Models for Lossless Image Compression: Next-Pixel Prediction in Language Space is All You Need

  • Kecheng Chen
  • Pingping Zhang
  • Hui Liu
  • Jie Liu
  • Yibing Liu
  • Jiaxin Huang
  • Shiqi Wang
  • Hong Yan

We have recently witnessed that ''Intelligence" and `''Compression" are the two sides of the same coin, where the language large model (LLM) with unprecedented intelligence is a general-purpose lossless compressor for various data modalities. This attribute is particularly appealing to the lossless image compression community, given the increasing need to compress high-resolution images in the current streaming media era. Consequently, a spontaneous envision emerges: Can the compression performance of the LLM elevate lossless image compression to new heights? However, our findings indicate that the naive application of LLM-based lossless image compressors suffers from a considerable performance gap compared with existing state-of-the-art (SOTA) codecs on common benchmark datasets. In light of this, we are dedicated to fulfilling the unprecedented intelligence (compression) capacity of the LLM for lossless image compression tasks, thereby bridging the gap between theoretical and practical compression performance. Specifically, we propose P -LLM, a next-pixel prediction-based LLM, which integrates various elaborated insights and methodologies, \textit{e. g. ,} pixel-level priors, the in-context ability of LLM, and a pixel-level semantic preservation strategy, to enhance the understanding capacity of pixel sequences for better next-pixel predictions. Extensive experiments on benchmark datasets demonstrate that P-LLM can beat SOTA classical and learned codecs.

JBHI Journal 2025 Journal Article

MPSol: A Multimodal Prompt Learning Framework for Protein Solubility Prediction

  • Yuhang Zhang
  • Peilin Chen
  • Keyan Ding
  • Han Liu
  • Shiqi Wang
  • Qi Song

Protein solubility is a critical determinant of biologic candidates’ developability, stability, and therapeutic efficacy. However, accurate solubility prediction remains a central challenge in computational protein engineering due to the inherent complexity within protein sequences. In this work, we propose a multimodal prompt learning framework, called MPSol, for protein solubility prediction that integrates complementary representations derived from primary sequences, structural proxies, and textual descriptions generated by large language models (LLMs). MPSol is built upon a unified multimodal backbone with a dedicated cross-modal fusion module that captures fine-grained interactions across modalities. In addition, we design label-aware prompts that encode solubility-specific semantic cues associated with each class. These prompts provide semantic supervision, guiding the alignment of fused protein representations to promote semantic consistency. Extensive experiments demonstrate that MPSol achieves state-of-the-art performance, reaching an accuracy of 0. 815, AUC of 0. 867 and MCC of 0. 642 on the standard PDBSol test set, and generalizes well to the external out-of-distribution test dataset with an accuracy of 0. 632, AUC of 0. 653 and MCC of 0. 332. These results underscore the potential of prompt-driven multimodal learning for interpretable and effective protein property prediction.

NeurIPS Conference 2025 Conference Paper

Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal Retrieval

  • Lanyun Zhu
  • Deyi Ji
  • Tianrun Chen
  • Haiyang Wu
  • Shiqi Wang

The success of DeepSeek-R1 demonstrates the immense potential of using reinforcement learning (RL) to enhance LLMs' reasoning capabilities. This paper introduces Retrv-R1, the first R1-style MLLM specifically designed for multimodal universal retrieval, achieving higher performance by employing step-by-step reasoning to produce more accurate retrieval results. We find that directly applying the methods of DeepSeek-R1 to retrieval tasks is not feasible, mainly due to (1) the high computational cost caused by the large token consumption required for multiple candidates with reasoning processes, and (2) the instability and suboptimal results when directly applying RL to train for retrieval tasks. To address these issues, Retrv-R1 introduces an information compression module with a details inspection mechanism, which enhances computational efficiency by reducing the number of tokens while ensuring that critical information for challenging candidates is preserved. Additionally, a new training paradigm is proposed, including an activation stage using a retrieval-tailored synthetic CoT dataset for more effective optimization, followed by RL with a novel curriculum reward to improve both performance and efficiency. Incorporating these novel designs, Retrv-R1 achieves SOTA performance, high efficiency, and strong generalization ability, as demonstrated by extensive experiments across multiple benchmarks and tasks.

AAAI Conference 2025 Conference Paper

Speed Master: Quick or Slow Play to Attack Speaker Recognition

  • Zhe Ye
  • Wenjie Zhang
  • Ying Ren
  • Xiangui Kang
  • Diqun Yan
  • Bin Ma
  • Shiqi Wang

Backdoor attacks pose a significant threat during the model's training phase. Attackers craft pre-defined triggers to break deep neural networks, ensuring the model accurately classifies clean samples during inference yet erroneously classifies samples added with these triggers. Recent studies have shown that speaker recognition systems trained on large-scale data are susceptible to backdoor attacks. Existing attackers employ unnoticed ambient sounds as triggers. However, these sounds are not inherently part of the training samples themselves. In essence, triggers can be designed to maintain an intrinsic connection with the original speech to enhance stealthiness. Our paper presents a novel attack methodology named Speed Master, which undermines deep neural networks by manipulating the speed of speech samples. Specifically, we execute poison-only backdoor attacks using speed or tempo adjustment. Changes in speech rate have become a common occurrence, as seen on platforms that allow users to adjust playback speed. In real-world scenarios, people naturally adjust their speaking rate depending on the context. As a result, changes in a speaker’s speech rate are typically perceived as normal and are unlikely to raise suspicion. Furthermore, detecting such subtle adjustments becomes challenging for users without reference speech. Our comprehensive experiments demonstrate that Speed Master can achieve an ASR over 99% in the digital domain, with only a 0.6% poisoning rate. Additionally, we validate the feasibility of Speed Master in the real world and its resistance to typical defensive measures.

AAAI Conference 2025 Conference Paper

Unified Coding for Both Human Perception and Generalized Machine Analytics with CLIP Supervision

  • Kangsheng Yin
  • Quan Liu
  • Xuelin Shen
  • Yulin He
  • Wenhan Yang
  • Shiqi Wang

The image compression model has long struggled with adaptability and generalization, as the decoded bitstream typically serves only human or machine needs and fails to preserve information for unseen visual tasks. Therefore, this paper innovatively introduces supervision obtained from multimodal pre-training models and incorporates adaptive multi-objective optimization tailored to support both human visual perception and machine vision simultaneously with a single bitstream, denoted as Unified and Generalized Image Coding for Machine (UG-ICM). Specifically, to get rid of the reliance between compression models with downstream task supervision, we introduce Contrastive Language-Image Pre-training (CLIP) models into the training constraint for improved generalization. Global-to-instance-wise CLIP supervision is applied to help obtain hierarchical semantics that make models more generalizable for the tasks relying on the information of different granularity. Furthermore, for supporting both human and machine visions with only a unifying bitstream, we incorporate a conditional decoding strategy that takes as conditions human or machine preferences, enabling the bitstream to be decoded into different versions for corresponding preferences. As such, our proposed UG-ICM is fully trained in a self-supervised manner, i.e., without awareness of any specific downstream models and tasks. The extensive experiments have shown that the proposed UG-ICM is capable of achieving remarkable improvements in various unseen machine analytics tasks, while simultaneously providing perceptually satisfying images.

NeurIPS Conference 2024 Conference Paper

Adaptive Image Quality Assessment via Teaching Large Multimodal Model to Compare

  • Hanwei Zhu
  • Haoning Wu
  • Yixuan Li
  • Zicheng Zhang
  • Baoliang Chen
  • Lingyu Zhu
  • Yuming Fang
  • Guangtao Zhai

While recent advancements in large multimodal models (LMMs) have significantly improved their abilities in image quality assessment (IQA) relying on absolute quality rating, how to transfer reliable relative quality comparison outputs to continuous perceptual quality scores remains largely unexplored. To address this gap, we introduce an all-around LMM-based NR-IQA model, which is capable of producing qualitatively comparative responses and effectively translating these discrete comparison outcomes into a continuous quality score. Specifically, during training, we present to generate scaled-up comparative instructions by comparing images from the same IQA dataset, allowing for more flexible integration of diverse IQA datasets. Utilizing the established large-scale training corpus, we develop a human-like visual quality comparator. During inference, moving beyond binary choices, we propose a soft comparison method that calculates the likelihood of the test image being preferred over multiple predefined anchor images. The quality score is further optimized by maximum a posteriori estimation with the resulting probability matrix. Extensive experiments on nine IQA datasets validate that the Compare2Score effectively bridges text-defined comparative levels during training with converted single image quality scores for inference, surpassing state-of-the-art IQA models across diverse scenarios. Moreover, we verify that the probability-matrix-based inference conversion not only improves the rating accuracy of Compare2Score but also zero-shot general-purpose LMMs, suggesting its intrinsic effectiveness.

NeurIPS Conference 2024 Conference Paper

DDR: Exploiting Deep Degradation Response as Flexible Image Descriptor

  • Juncheng Wu
  • Zhangkai Ni
  • Hanli Wang
  • Wenhan Yang
  • Yuyin Zhou
  • Shiqi Wang

Image deep features extracted by pre-trained networks are known to contain rich and informative representations. In this paper, we present Deep Degradation Response (DDR), a method to quantify changes in image deep features under varying degradation conditions. Specifically, our approach facilitates flexible and adaptive degradation, enabling the controlled synthesis of image degradation through text-driven prompts. Extensive evaluations demonstrate the versatility of DDR as an image descriptor, with strong correlations observed with key image attributes such as complexity, colorfulness, sharpness, and overall quality. Moreover, we demonstrate the efficacy of DDR across a spectrum of applications. It excels as a blind image quality assessment metric, outperforming existing methodologies across multiple datasets. Additionally, DDR serves as an effective unsupervised learning objective in image restoration tasks, yielding notable advancements in image deblurring and single-image super-resolution. Our code is available at: https: //github. com/eezkni/DDR.

NeurIPS Conference 2024 Conference Paper

LeDex: Training LLMs to Better Self-Debug and Explain Code

  • Nan Jiang
  • Xiaopeng Li
  • Shiqi Wang
  • Qiang Zhou
  • Soneya B. Hossain
  • Baishakhi Ray
  • Varun Kumar
  • Xiaofei Ma

In the domain of code generation, self-debugging is crucial. It allows LLMs to refine their generated code based on execution feedback. This is particularly important because generating correct solutions in one attempt proves challenging for complex tasks. Prior works on self-debugging mostly focus on prompting methods by providing LLMs with few-shot examples, which work poorly on small open-sourced LLMs. In this work, we propose LeDex, a training framework that significantly improves the self-debugging capability of LLMs. Intuitively, we observe that a chain of explanations on the wrong code followed by code refinement helps LLMs better analyze the wrong code and do refinement. We thus propose an automated pipeline to collect a high-quality dataset for code explanation and refinement by generating a number of explanations and refinement trajectories from the LLM itself or a larger teacher model and filtering via execution verification. We perform supervised fine-tuning (SFT) and further reinforcement learning (RL) on both success and failure trajectories with a novel reward design considering code explanation and refinement quality. SFT improves the pass@1 by up to 15. 92\% and pass@10 by 9. 30\% over four benchmarks. RL training brings additional up to 3. 54\% improvement on pass@1 and 2. 55\% improvement on pass@10. The trained LLMs show iterative refinement ability and can keep refining code continuously. Lastly, our human evaluation shows that the LLMs trained with our framework generate more useful code explanations and help developers better understand bugs in source code.

AAAI Conference 2024 Conference Paper

Mitigating the Impact of False Negative in Dense Retrieval with Contrastive Confidence Regularization

  • Shiqi Wang
  • Yeqin Zhang
  • Cam-Tu Nguyen

In open-domain Question Answering (QA), dense text retrieval is crucial for finding relevant passages to generate answers. Typically, contrastive learning is used to train a retrieval model, which maps passages and queries to the same semantic space, making similar ones closer and dissimilar ones further apart. However, training such a system is challenging due to the false negative problem, where relevant passages may be missed during data annotation. Hard negative sampling, commonly used to improve contrastive learning, can introduce more noise in training. This is because hard negatives are those close to a given query, and thus more likely to be false negatives. To address this, we propose a novel contrastive confidence regularizer for Noise Contrastive Estimation (NCE) loss, a commonly used contrastive loss. Our analysis shows that the regularizer helps make the dense retrieval model more robust against false negatives with a theoretical guarantee. Additionally, we propose a model-agnostic method to filter out noisy negative passages in the dataset, improving any downstream dense retrieval models. Through experiments on three datasets, we demonstrate that our method achieves better retrieval performance in comparison to existing state-of-the-art dense retrieval systems.

EAAI Journal 2024 Journal Article

Robust asynchronous fuzzy predictive fault-tolerant tracking control for nonlinear multi-phase batch processes with time-varying reference trajectories

  • Hui Li
  • Shiqi Wang
  • Huiyuan Shi
  • Limin Wang
  • Chengli Su
  • Ping Li

For nonlinear multi-phase batch processes with time-varying reference trajectories, actuator faults and nonlinear characteristics, a robust asynchronous fuzzy predictive fault-tolerant tracking control method is proposed. First, considering an asynchronous switching situation between a state and a controller when switching occurs, an extended Takagi-Sugeno fuzzy switching model, including the matched and mismatched cases, is established. Then, a robust asynchronous fuzzy predictive fault-tolerant tracking controller is constructed based on the switching model by taking into account whether the model rules match the controller rules. Second, in order to ensure the stability of the system, the conditions for system stability indicated by the linear matrix inequality are provided using the relevant theories and methodologies. Next, the stability conditions are solved online, which can obtain the gains of control law in each phase, the minimum running time in the matched case, and the maximum running time in the mismatched case. By utilizing the maximum running duration allows the switching signal to be sent out in advance, preventing asynchronous switching situations and ensuring stable system operation. The simulation results finally demonstrate the viability and effectiveness of the developed controller.

IJCAI Conference 2024 Conference Paper

ScreenAgent: A Vision Language Model-driven Computer Control Agent

  • Runliang Niu
  • Jindong Li
  • Shiqi Wang
  • Yali Fu
  • Xiyu Hu
  • Xueyuan Leng
  • He Kong
  • Yi Chang

Large Language Models (LLM) can invoke a variety of tools and APIs to complete complex tasks. The computer, as the most powerful and universal tool, could potentially be controlled by a trained LLM agent. Powered by the computer, we can hopefully build a more generalized agent to assist humans in various daily digital works. In this paper, we construct an environment for a Vision Language Model (VLM) agent to interact with a real computer screen. Within this environment, the agent can observe screenshots and manipulate the Graphical User Interface (GUI) by outputting mouse and keyboard actions. We also design an automated control pipeline that includes planning, acting, and reflecting phases, guiding the agent to continuously interact with the environment and complete multi-step tasks. Additionally, we construct the ScreenAgent Dataset, which collects screenshots and action sequences when completing daily computer tasks. Finally, we train a model, ScreenAgent, which achieves comparable computer control capabilities to GPT-4V and demonstrated more precise UI positioning capabilities. Our attempts could inspire further research on building a generalist LLM agent. The code and more detailed information are at https: //github. com/niuzaisheng/ScreenAgent.

NeurIPS Conference 2022 Conference Paper

General Cutting Planes for Bound-Propagation-Based Neural Network Verification

  • Huan ZHang
  • Shiqi Wang
  • Kaidi Xu
  • Linyi Li
  • Bo Li
  • Suman Jana
  • Cho-Jui Hsieh
  • J. Zico Kolter

Bound propagation methods, when combined with branch and bound, are among the most effective methods to formally verify properties of deep neural networks such as correctness, robustness, and safety. However, existing works cannot handle the general form of cutting plane constraints widely accepted in traditional solvers, which are crucial for strengthening verifiers with tightened convex relaxations. In this paper, we generalize the bound propagation procedure to allow the addition of arbitrary cutting plane constraints, including those involving relaxed integer variables that do not appear in existing bound propagation formulations. Our generalized bound propagation method, GCP-CROWN, opens up the opportunity to apply general cutting plane methods for neural network verification while benefiting from the efficiency and GPU acceleration of bound propagation methods. As a case study, we investigate the use of cutting planes generated by off-the-shelf mixed integer programming (MIP) solver. We find that MIP solvers can generate high-quality cutting planes for strengthening bound-propagation-based verifiers using our new formulation. Since the branching-focused bound propagation procedure and the cutting-plane-focused MIP solver can run in parallel utilizing different types of hardware (GPUs and CPUs), their combination can quickly explore a large number of branches with strong cutting planes, leading to strong verification performance. Experiments demonstrate that our method is the first verifier that can completely solve the oval20 benchmark and verify twice as many instances on the oval21 benchmark compared to the best tool in VNN-COMP 2021, and also noticeably outperforms state-of-the-art verifiers on a wide range of benchmarks. GCP-CROWN is part of the $\alpha, \beta$-CROWN verifier, the VNN-COMP 2022 winner. Code is available at http: //PaperCode. cc/GCP-CROWN.

AAAI Conference 2021 Conference Paper

Adaptive Verifiable Training Using Pairwise Class Similarity

  • Shiqi Wang
  • Kevin Eykholt
  • Taesung Lee
  • Jiyong Jang
  • Ian Molloy

Verifiable training has shown success in creating neural networks that are provably robust to a given amount of noise. However, despite only enforcing a single robustness criterion, its performance scales poorly with dataset complexity. On CIFAR10, a non-robust LeNet model has a 21. 63% error rate, while a model created using verifiable training and a L∞ robustness criterion of 8/255, has an error rate of 57. 10%. Upon examination, we find that when labeling visually similar classes, the model’s error rate is as high as 61. 65%. Thus, we attribute the loss in performance to inter-class similarity. Classes that are similar (i. e. , close in the feature space) increase the difficulty of learning a robust model. While it may be desirable to train a model to be robust for a large robustness region, pairwise class similarities limit the potential gains. Furthermore, consideration must be made regarding the relative cost of mistaking one class for another. In security or safety critical tasks, similar classes are likely to belong to the same group, and thus are equally sensitive. In this work, we propose a new approach that utilizes interclass similarity to improve the performance of verifiable training and create robust models with respect to multiple adversarial criteria. First, we cluster similar classes using agglomerate clustering and assign robustness criteria based on the degree of similarity between clusters. Next, we propose two methods to apply our approach: (1) the Inter-Group Robustness Prioritization method, which uses a custom loss term to create a single model with multiple robustness guarantees and (2) the neural decision tree method, which trains multiple sub-classifiers with different robustness guarantees and combines them in a decision tree architecture. Our experiments on Fashion-MNIST and CIFAR10 demonstrate that by prioritizing the robustness between the most dissimilar groups, we improve clean performance by up to 9. 63% and 30. 89% respectively. Furthermore, on CIFAR100, our approach reduces the clean error rate by 26. 32%.

NeurIPS Conference 2021 Conference Paper

Beta-CROWN: Efficient Bound Propagation with Per-neuron Split Constraints for Neural Network Robustness Verification

  • Shiqi Wang
  • Huan ZHang
  • Kaidi Xu
  • Xue Lin
  • Suman Jana
  • Cho-Jui Hsieh
  • J. Zico Kolter

Bound propagation based incomplete neural network verifiers such as CROWN are very efficient and can significantly accelerate branch-and-bound (BaB) based complete verification of neural networks. However, bound propagation cannot fully handle the neuron split constraints introduced by BaB commonly handled by expensive linear programming (LP) solvers, leading to loose bounds and hurting verification efficiency. In this work, we develop $\beta$-CROWN, a new bound propagation based method that can fully encode neuron splits via optimizable parameters $\beta$ constructed from either primal or dual space. When jointly optimized in intermediate layers, $\beta$-CROWN generally produces better bounds than typical LP verifiers with neuron split constraints, while being as efficient and parallelizable as CROWN on GPUs. Applied to complete robustness verification benchmarks, $\beta$-CROWN with BaB is up to three orders of magnitude faster than LP-based BaB methods, and is notably faster than all existing approaches while producing lower timeout rates. By terminating BaB early, our method can also be used for efficient incomplete verification. We consistently achieve higher verified accuracy in many settings compared to powerful incomplete verifiers, including those based on convex barrier breaking techniques. Compared to the typically tightest but very costly semidefinite programming (SDP) based incomplete verifiers, we obtain higher verified accuracy with three orders of magnitudes less verification time. Our algorithm empowered the $\alpha, \! \beta$-CROWN (alpha-beta-CROWN) verifier, the winning tool in VNN-COMP 2021. Our code is available at http: //PaperCode. cc/BetaCROWN.

NeurIPS Conference 2020 Conference Paper

ARMA Nets: Expanding Receptive Field for Dense Prediction

  • Jiahao Su
  • Shiqi Wang
  • Furong Huang

Global information is essential for dense prediction problems, whose goal is to compute a discrete or continuous label for each pixel in the images. Traditional convolutional layers in neural networks, initially designed for image classification, are restrictive in these problems since the filter size limits their receptive fields. In this work, we propose to replace any traditional convolutional layer with an autoregressive moving-average (ARMA) layer, a novel module with an adjustable receptive field controlled by the learnable autoregressive coefficients. Compared with traditional convolutional layers, our ARMA layer enables explicit interconnections of the output neurons and learns its receptive field by adapting the autoregressive coefficients of the interconnections. ARMA layer is adjustable to different types of tasks: for tasks where global information is crucial, it is capable of learning relatively large autoregressive coefficients to allow for an output neuron's receptive field covering the entire input; for tasks where only local information is required, it can learn small or near zero autoregressive coefficients and automatically reduces to a traditional convolutional layer. We show both theoretically and empirically that the effective receptive field of networks with ARMA layers (named ARMA networks) expands with larger autoregressive coefficients. We also provably solve the instability problem of learning and prediction in the ARMA layer through a re-parameterization mechanism. Additionally, we demonstrate that ARMA networks substantially improve their baselines on challenging dense prediction tasks, including video prediction and semantic segmentation.

NeurIPS Conference 2020 Conference Paper

Domain Generalization for Medical Imaging Classification with Linear-Dependency Regularization

  • Haoliang Li
  • Yufei Wang
  • Renjie Wan
  • Shiqi Wang
  • Tie-Qiang Li
  • Alex Kot

Recently, we have witnessed great progress in the field of medical imaging classification by adopting deep neural networks. However, the recent advanced models still require accessing sufficiently large and representative datasets for training, which is often unfeasible in clinically realistic environments. When trained on limited datasets, the deep neural network is lack of generalization capability, as the trained deep neural network on data within a certain distribution (e. g. the data captured by a certain device vendor or patient population) may not be able to generalize to the data with another distribution. In this paper, we introduce a simple but effective approach to improve the generalization capability of deep neural networks in the field of medical imaging classification. Motivated by the observation that the domain variability of the medical images is to some extent compact, we propose to learn a representative feature space through variational encoding with a novel linear-dependency regularization term to capture the shareable information among medical data collected from different domains. As a result, the trained neural network is expected to equip with better generalization capability to the ``unseen" medical data. Experimental results on two challenging medical imaging classification tasks indicate that our method can achieve better cross-domain generalization capability compared with state-of-the-art baselines.

NeurIPS Conference 2020 Conference Paper

HYDRA: Pruning Adversarially Robust Neural Networks

  • Vikash Sehwag
  • Shiqi Wang
  • Prateek Mittal
  • Suman Jana

In safety-critical but computationally resource-constrained applications, deep learning faces two key challenges: lack of robustness against adversarial attacks and large neural network size (often millions of parameters). While the research community has extensively explored the use of robust training and network pruning \emph{independently} to address one of these challenges, only a few recent works have studied them jointly. However, these works inherit a heuristic pruning strategy that was developed for benign training, which performs poorly when integrated with robust training techniques, including adversarial training and verifiable robust training. To overcome this challenge, we propose to make pruning techniques aware of the robust training objective and let the training objective guide the search for which connections to prune. We realize this insight by formulating the pruning objective as an empirical risk minimization problem which is solved efficiently using SGD. We demonstrate that our approach, titled HYDRA, achieves compressed networks with \textit{state-of-the-art} benign and robust accuracy, \textit{simultaneously}. We demonstrate the success of our approach across CIFAR-10, SVHN, and ImageNet dataset with four robust training techniques: iterative adversarial training, randomized smoothing, MixTrain, and CROWN-IBP. We also demonstrate the existence of highly robust sub-networks within non-robust networks.

AAAI Conference 2020 Conference Paper

Towards Scale-Free Rain Streak Removal via Self-Supervised Fractal Band Learning

  • Wenhan Yang
  • Shiqi Wang
  • Dejia Xu
  • Xiaodong Wang
  • Jiaying Liu

Data-driven rain streak removal methods, which most of rely on synthesized paired data, usually come across the generalization problem when being applied in real cases. In this paper, we propose a novel deep-learning based rain streak removal method injected with self-supervision to improve the ability to remove rain streaks in various scales. To realize this goal, we made efforts in two aspects. First, considering that rain streak removal is highly correlated with texture characteristics, we create a fractal band learning (FBL) network based on frequency band recovery. It integrates commonly seen band feature operations with neural modules and effectively improves the capacity to capture discriminative features for deraining. Second, to further improve the generalization ability of FBL for rain streaks in various scales, we add cross-scale self-supervision to regularize the network training. The constraint forces the extracted features of inputs in different scales to be equivalent after rescaling. Therefore, FBL can offer similar responses based on solely image content without the interleave of scale and is capable to remove rain streaks in various scales. Extensive experiments in quantitative and qualitative evaluations demonstrate the superiority of our FBL for rain streak removal, especially for the real cases where very large rain streaks exist, and prove the effectiveness of its each component. Our code will be public available at: https: //github. com/flyywh/AAAI-2020-FBL-SS.

NeurIPS Conference 2018 Conference Paper

Efficient Formal Safety Analysis of Neural Networks

  • Shiqi Wang
  • Kexin Pei
  • Justin Whitehouse
  • Junfeng Yang
  • Suman Jana

Neural networks are increasingly deployed in real-world safety-critical domains such as autonomous driving, aircraft collision avoidance, and malware detection. However, these networks have been shown to often mispredict on inputs with minor adversarial or even accidental perturbations. Consequences of such errors can be disastrous and even potentially fatal as shown by the recent Tesla autopilot crash. Thus, there is an urgent need for formal analysis systems that can rigorously check neural networks for violations of different safety properties such as robustness against adversarial perturbations within a certain L-norm of a given image. An effective safety analysis system for a neural network must be able to either ensure that a safety property is satisfied by the network or find a counterexample, i. e. , an input for which the network will violate the property. Unfortunately, most existing techniques for performing such analysis struggle to scale beyond very small networks and the ones that can scale to larger networks suffer from high false positives and cannot produce concrete counterexamples in case of a property violation. In this paper, we present a new efficient approach for rigorously checking different safety properties of neural networks that significantly outperforms existing approaches by multiple orders of magnitude. Our approach can check different safety properties and find concrete counterexamples for networks that are 10x larger than the ones supported by existing analysis techniques. We believe that our approach to estimating tight output bounds of a network for a given input range can also help improve the explainability of neural networks and guide the training process of more robust neural networks.

v2026.09.13