Arrow Research search

Author name cluster

Lin Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

55 papers
2 author rows

Possible papers

55

YNICL Journal 2026 Journal Article

Alteration of fronto-thalamic-striatal and visual network activity to positive emotional stimuli in adolescent patients with bipolar disorder during a Go/No-Go task-based functional brain MRI

  • Xueying Wang
  • Jinfan Zhang
  • Feifei Wu
  • Liying Shen
  • Lin Wang
  • Han Wu
  • Qian Xiao
  • Xiaoping Yi

Background Adolescent patients with bipolar disorder (BD) tend to have abnormal neural activity to emotional stimuli. This study aimed to assess alterations of neural activity within the fronto-thalamic-striatal circuit, and the fusiform gyrus during positive stimuli processing in Go/No-Go task-based brain functional MRI (fMRI). Methods This prospective study enrolled 43 adolescent patients with BD and 18 age- sex-matched healthy controls. All study participants underwent a task-based brain fMRI using a happy versus neutral Go/No-Go paradigm. All study participants also completed multiple affective and cognitive assessment questionnaire. Results Enhanced activity was found in the fronto-thalamic-striatal circuit including the inferior frontal gyrus and the caudate, the fusiform gyrus, the left cerebellum crus I and the hippocampus in adolescent patients with BD, during response inhibition to happy versus neutral distractors in an emotional Go/No-Go fMRI task, compared with matched healthy controls (p<0. 05). Moreover, the left inferior frontal gyrus, the right fusiform gyrus and the left cerebellum crus I responses to happy versus neutral distractors were positively associated with the differences in false response errors in the patients with BD (all p<0. 05, FDR corrected). The enhanced activity of the caudate nucleus and that of the right hippocampus were negatively correlated with cognitive function (all p<0. 05, FDR corrected). Conclusion This study found significant brain functional alterations in the limbic system, the visual brain network and cerebellum, especially the fronto-thalamic-striatal track and fusiform gyrus, which was correlated with cognitive dysfunction in adolescent patients with BD. These changes may serve as potential neuroimaging correlates of BD in adolescent patients.

AAAI Conference 2026 Conference Paper

Beyond Single-Point Perturbation: A Hierarchical, Manifold-Aware Approach to Diffusion Attacks

  • Zhijie Wang
  • Lin Wang
  • Zhenyu Wen
  • Cong Wang

Latent Diffusion Models have become a powerful tool for generating high-fidelity unrestricted adversarial examples. However, the existing methods typically perturb only the initial latent or rely on prompt engineering, which is ill-suited to the iterative nature of the diffusion process, plus optimization instability due to external text prompts and cumulative drift that push the adversarial images off the data manifold. In this paper, we propose a hierarchical attack framework that operates in alignment with the model's generative manifold and leverages intermediate denoising states to maximize attack transferability and visual fidelity. Extensive experiments show that the proposed attack improves adversarial transferability by 10-20% against a diverse set of normally-trained models and achieves over 10.5% higher success rate against adversarially-defended models, while simultaneously enhancing visual quality by 1.0-1.2 FID reduction and 16.7% LPIPS improvements.

TMLR Journal 2026 Journal Article

BrightDreamer: Generic 3D Gaussian Generative Framework for Fast Text-to-3D Synthesis

  • Lutao Jiang
  • Xu Zheng
  • Yuanhuiyi Lyu
  • Jiazhou Zhou
  • Lin Wang

Text-to-3D synthesis has recently seen intriguing advances by combining the text-to-image priors with 3D representation methods, e.g., 3D Gaussian Splatting (3D GS), via Score Distillation Sampling (SDS). However, a hurdle of existing methods is the low efficiency, per-prompt optimization for a single 3D object. Therefore, it is imperative for a paradigm shift from per-prompt optimization to feed-forward generation for any unseen text prompts, which yet remains challenging. An obstacle is how to directly generate a set of millions of 3D Gaussians to represent a 3D object. This paper presents BrightDreamer, an end-to-end feed-forward approach that can achieve generalizable and fast (77 ms) text-to-3D generation. Our key idea is to formulate the generation process as estimating the 3D deformation from an anchor shape with predefined positions. For this, we first propose a Text-guided Shape Deformation (TSD) network to predict the deformed shape and its new positions, used as the centers (one attribute) of 3D Gaussians. To estimate the other four attributes (i.e., scaling, rotation, opacity, and SH), we then design a novel Text-guided Triplane Generator (TTG) to generate a triplane representation for a 3D object. The center of each Gaussian enables us to transform the spatial feature into the four attributes. The generated 3D Gaussians can be finally rendered at 705 frames per second. Extensive experiments demonstrate the superiority of our method over existing methods. Also, BrightDreamer possesses a strong semantic understanding capability even for complex text prompts. The project code is available in supplementary materials.

EAAI Journal 2026 Journal Article

Cross-domain generalization in non-stationary data-based Human Activity Recognition

  • Siyang Wang
  • Lin Wang
  • Lunan Duan
  • Zhi Huang
  • Wenyuan Liu

Human Activity Recognition (HAR) is pivotal for context-aware applications, while WiFi-based HAR offers non-line-of-sight sensing with low cost and privacy preservation. However, due to the high sensitivity of WiFi features, existing methods require explicit manual labeling of multiple domain factors (e. g. , user, location, orientation, environment) to address cross-domain generalization, leading to high annotation costs and scalability challenges. To address these challenges, we introduce a novel concept, domain state (DS), that unifies multi-dimensional domain factors into a single label. This innovation eliminates the need for separate annotations of individual domain attributes. A dynamic threshold iterative clipping clustering (DTICC) algorithm is introduced to automatically assign pseudo-DS labels by iteratively filtering outliers and refining clusters, eliminating manual domain annotation entirely. Furthermore, an adversarial learning-based DS Generalization Network is designed to extract domain-invariant features through a minimax game between a feature extractor and a DS discriminator, enabling robust recognition across unseen domains without target domain samples. Extensive experiments on public datasets demonstrate that our method achieves 87% average accuracy, outperforming state-of-the-art methods by 5%–32%. By decoupling domain-specific and activity-specific features, this work provides a scalable solution for cross-domain WiFi-based HAR, significantly reducing annotation costs compared to traditional approaches while maintaining performance under complex domain shifts. The DS concept and DTICC algorithm together enable a paradigm shift in WiFi-based HAR, where domain-aware recognition is achieved with minimal human intervention.

AAAI Conference 2026 Conference Paper

Ev-iCRF: Self-supervised Event-guided iCRF Estimation for HDR Image Reconstruction

  • Xucheng Guo
  • Bing Li
  • Lin Wang
  • Yiran Shen

In this paper, we present Ev-iCRF, a novel self-supervised pipeline for high dynamic range (HDR) image reconstruction from a single-exposure low dynamic range (LDR) image, guided by asynchronous event streams generated by a bio-inspired event camera. The highlight of Ev-iCRF lies in its formulation of the inverse camera response function (iCRF) based on Event-LDR Correspondence. By leveraging the HDR properties of event data, the method enables direct iCRF estimation, offering a new perspective for event-guided HDR imaging. The pipeline is trained in a self-supervised manner using formulation-driven iCRF estimation loss and refinement loss, without the need for synchronized HDR supervision. Ev-iCRF adopts a two-stage coarse-to-fine reconstruction pipeline, allowing effective fusion of features from both LDR image and event data. The event information is used to optimize the iCRF, enabling accurate HDR reconstruction from LDR inputs. We evaluate Ev-iCRF on real-world datasets, and results show that it outperforms state-of-the-art methods in HDR reconstruction accuracy. Moreover, the reconstructed images demonstrate improved texture fidelity and structural detail.

AAAI Conference 2026 Conference Paper

EvDiff3D: Event-Aware Diffusion Repair for High-Fidelity Event-Based 3D Reconstruction

  • Kanghao Chen
  • Zixin Zhang
  • Hangyu Li
  • Lin Wang
  • Zeyu Wang

Event cameras are bio-inspired sensors that capture visual information through asynchronous brightness changes, offering distinct advantages including high temporal resolution and wide dynamic range. While prior research has investigated event-based 3D reconstruction for extreme scenarios, existing methods face inherent limitations and fail to fully exploit the unique characteristics of event data. In this paper, we present EvDiff3D, a novel two-stage 3D reconstruction framework that integrates event-based geometric constraints with an event-aware diffusion prior for appearance refinement. Our key insight lies in bridging the gap between physically grounded event-based reconstruction and data-driven appearance repair through a unified cyclical pipeline. In the first stage, we reconstruct a coarse 3D scene under supervision from event loss and event-based monocular depth constraints to preserve structural fidelity. The second stage fine-tunes an event-aware diffusion model based on a pretrained video diffusion model as a repair prior to enhance the appearance in under-constrained regions. Based on the diffusion model, our pipeline operates within a reconstruction-generation cycle that progressively refines both geometry and appearance using only event data. Extensive experiments on synthetic and real-world datasets demonstrate that EvDiff3D significantly outperforms existing methods in perceptual quality and structural consistency.

EAAI Journal 2026 Journal Article

Fusion-driven graph representation enhancement for predicting interactions of new drugs

  • Jiankang Liu
  • Yuhan Zhao
  • Rui Han
  • Yihan Fu
  • Lin Wang

Accurate prediction of drug–drug interactions (DDIs) for newly synthesized compounds enables early, in-silico safety screening in drug discovery and formulary review. We target the cold-start regime, where (i) new compounds are topologically isolated on external biomedical knowledge graphs (KGs) and on the DDI graph, and (ii) sparse supervision hampers the learning of discriminative representations. We propose an early-fusion method (LINCS-DDI) that inserts shared substructure nodes to connect a molecular-fingerprint knowledge graph with the DDI graph, turning structural similarity into topological links, and providing two-hop connectivity directly from the Simplified Molecular Input Line Entry System (SMILES) without prior inclusion in external KGs. Building on this substrate, we introduce Native Dual-View Contrastive Learning (NDV-CL): within a single pass of a flow-based graph neural network (GNN), forward and reverse message-passing representations of the same drug pair are treated as deterministic positives, while label-guided negatives (screened using only training-split interaction labels) are mined within the induced subgraph, improving representation quality without stochastic augmentations. Under strict cold-start settings on two open-source datasets, LINCS-DDI improves macro-F1 by up to 4. 1% over the best baseline and reduces contrastive overhead by up to 66%. These properties make the approach suitable for routine, large-scale preclinical DDI triage, prioritizing high-risk combinations for wet-lab validation, and informing pharmacovigilance pipelines.

AAAI Conference 2026 Conference Paper

PhysPatch: A Physically Realizable and Transferable Adversarial Patch Attack for Multimodal Large Language Models-based Autonomous Driving Systems

  • Qi Guo
  • Xiaojun Jia
  • Shanmin Pang
  • Simeng Qin
  • Lin Wang
  • Ju Jia
  • Yang Liu
  • Qing Guo

Multimodal Large Language Models (MLLMs) are becoming integral to autonomous driving (AD) systems due to their strong vision-language reasoning capabilities. However, MLLMs are vulnerable to adversarial attacks—particularly adversarial patch attacks—which can pose serious threats in real-world scenarios. Existing patch-based attack methods are primarily designed for object detection models. Due to the more complex architectures and strong reasoning capabilities of MLLMs, these approaches perform poorly when transferred to MLLM-based systems. To address these limitations, we propose PhysPatch, a physically realizable and transferable adversarial patch framework tailored for MLLM-based AD systems. PhysPatch jointly optimizes patch location, shape, and content to enhance attack effectiveness and real-world applicability. It introduces a semantic-based mask initialization strategy for realistic placement, an SVD-based local alignment loss with patch-guided crop-resize to improve transferability, and a potential field-based mask refinement method. Extensive experiments across open-source, commercial, and reasoning-capable MLLMs demonstrate that PhysPatch significantly outperforms state-of-the-art (SOTA) methods in steering MLLM-based AD systems toward target-aligned perception and planning outputs. Moreover, PhysPatch consistently places adversarial patches in physically feasible regions of AD scenes, ensuring strong real-world applicability and deployability.

AAAI Conference 2026 Conference Paper

Temporal and Spatial Representation Learning for Multimodal Low-Beam 3D Object Detection

  • Lin Wang
  • Shiliang Sun
  • Jing Zhao

To facilitate the large-scale deployment of autonomous driving in real-world scenarios, developing low-cost and high-performance 3D object detection systems has become a critical technical challenge. Although high-beam LiDARs provide denser point cloud data, their prohibitive hardware cost and high power consumption limit their practicality. In contrast, low-beam LiDARs offer advantages in terms of affordability and energy efficiency, but often suffer from inadequate perception accuracy due to their sparser point cloud data. This paper focuses on the task of multimodal 3D object detection with low-beam LiDARs, and proposes a novel approach that integrates temporal and spatial representation learning to enhance detection accuracy under sparser sensor conditions. Specifically, our approach comprises: (1) a Temporal Feature Prediction Learning (TFPL) module, which predicts the current BEV representation based on a sequence of historical BEV features; (2) a Spatial Feature Observation Learning (SFOL) module, which aligns BEV (Bird's-Eye-View) features from high-beam and low-beam LiDAR to enforce the low-beam features to approximate high-beam representations; (3) an Uncertainty-Aware Fusion (UAF) strategy, which performs feature-wise weighting between the predicted and observed BEV features by leveraging channel-wise variances, effectively mitigating perturbations in the learned BEV representations. Extensive experiments on the KITTI and nuScenes 3D object detection datasets demonstrate that the proposed approach significantly improves detection performance under low-beam LiDAR configurations.

EAAI Journal 2025 Journal Article

A flexible job shop scheduling method based on heterogeneous disjunctive graph and deep reinforcement learning

  • Xuhui Zhao
  • Haokai Qu
  • Mingchuan Zhang
  • Jiamei Feng
  • Lin Wang
  • Qingtao Wu

Automated Guided Vehicles (AGVs) have been widely applied in discrete manufacturing systems, but their introduction has increased the uncertainty and complexity of the scheduling process. In a flexible job-shop environment, achieving efficient integrated scheduling of machines and AGVs is of significant research value for enhancing production efficiency. However, existing mathematical models typically incorporate transportation time into processing time or assume it to be zero, neglecting the dynamic interaction characteristics between AGVs and machines, which limits the optimization effect of the scheduling scheme. This paper addresses the flexible job-shop scheduling problem with limited AGVs (FJSP-LA), aiming to minimize the makespan, and proposes an end-to-end optimization method based on deep reinforcement learning to achieve the coordinated optimization of processing and transportation resources. Firstly, a Heterogeneous Disjunctive Graph model is constructed, which includes three types of nodes: operations, machines, and AGVs. By integrating the information of operation nodes, machine nodes, and AGV nodes, efficient trilateral cooperative scheduling is achieved while reducing the model complexity. Secondly, a hierarchical feature extraction framework based on multi-layer perceptrons (MLPs) is designed to effectively capture the node attribute features and their dynamic constraint relationships. On this basis, a decision-making model based on the Actor–Critic architecture is constructed, and the Proximal Policy Optimization (PPO) algorithm is used for training to learn the optimal production scheduling strategy. To verify the effectiveness of the algorithm, comparative experiments were conducted on two public datasets with seven advanced scheduling algorithms. The numerical experimental results show that this method not only demonstrates good generalization ability on unknown datasets but also outperforms existing mainstream algorithms in terms of computational efficiency and scheduling performance, proving its application potential in complex flexible job-shop scheduling problems.

EAAI Journal 2025 Journal Article

A three-stage adaptive memetic algorithm for multi-objective optimization of flexible assembly job-shop scheduling problem

  • Chenlu Zhang
  • Jiamei Feng
  • Mingchuan Zhang
  • Lei Yang
  • Lei Zhang
  • Lin Wang
  • Junlong Zhu
  • Qingtao Wu

The flexible assembly job-shop scheduling problem (FAJSP) widely arises in the manufacturing industry. Various approaches have been designed in recent years to address this problem. However, existing methods have rarely considered assembly process constraints and task assembly wait time. For this reason, this paper proposes a three-stage adaptive memetic algorithm (TA-MA) to solve the FAJSP with process route constraints. Specifically, the proposed algorithm combines memetic algorithms and reinforcement learning. The optimization objectives are completion time, equipment load, and assembly operation waiting time. Moreover, a two-layer integer coding method is proposed to encode the problem, and a reinforcement learning method is introduced to assist the solution search of the memetic algorithm. Further, a three-stage search framework is designed to reasonably equilibrium TA-MA’s exploration and mining capabilities as iterations advance. Finally, the effectiveness of the proposed algorithm is assessed through a series of experiments. The outcomes demonstrate that the proposed algorithm is effective and outperforms existing algorithms.

YNIMG Journal 2025 Journal Article

An implemented predictive coding model of lexico-semantic processing explains the dynamics of univariate and multivariate activity within the left ventromedial temporal lobe during reading comprehension

  • Lin Wang
  • Samer Nour Eddine
  • Trevor Brothers
  • Ole Jensen
  • Gina R. Kuperberg

During language comprehension, the larger neural response to unexpected versus expected inputs is often taken as evidence for predictive coding-a specific computational architecture and optimization algorithm proposed to approximate probabilistic inference in the brain. However, other predictive processing frameworks can also account for this effect, leaving the unique claims of predictive coding untested. In this study, we used MEG to examine both univariate and multivariate neural activity in response to expected and unexpected inputs during word-by-word reading comprehension. We further simulated this activity using an implemented predictive coding model that infers the meaning of words from their orthographic form. Consistent with previous findings, the univariate analysis showed that, between 300 and 500 ms, unexpected words produced a larger evoked response than expected words within a left ventromedial temporal region that supports the mapping of orthographic word-forms onto lexical and conceptual representations. Our model explained this larger evoked response as the enhanced lexico-semantic prediction error produced when prior top-down predictions failed to suppress activity within lexical and semantic "error units". Critically, our simulations showed that despite producing minimal prediction error, expected inputs nonetheless reinstated top-down predictions within the model's lexical and semantic "state" units. Two types of multivariate analyses provided evidence for this functional distinction between state and error units within the ventromedial temporal region. First, within each trial, the same individual voxels that produced a larger response to unexpected inputs between 300 and 500 ms produced unique temporal patterns to expected inputs that resembled the patterns produced within a pre-activation time window. Second, across trials, and again within the same 300-500 ms time window and left ventromedial temporal region, pairs of expected words produced spatial patterns that were more similar to one another than the spatial patterns produced by pairs of expected and unexpected words, regardless of specific item. Together, these findings provide compelling evidence that the left ventromedial temporal lobe employs predictive coding to infer the meaning of incoming words from their orthographic form during reading comprehension.

ICRA Conference 2025 Conference Paper

DAP-LED: Learning Degradation-Aware Priors with Clip for Joint Low-Light Enhancement and Deblurring

  • Ling Wang
  • Chen Wu
  • Lin Wang

Autonomous vehicles and robots often struggle with reliable visual perception at night due to the low illumination and motion blur caused by the long exposure time of RGB cameras. Existing methods address this challenge by sequentially connecting the off-the-shelf pretrained lowlight enhancement and deblurring models. Unfortunately, these methods often lead to noticeable artifacts (e. g. , color distortions) in the over-exposed regions or make it hardly possible to learn the motion cues of the dark regions. In this paper, we interestingly find vision-language models, e. g. , Contrastive LanguageImage Pretraining (CLIP), can comprehensively perceive diverse degradation levels at night. In light of this, we propose a novel transformer-based joint learning framework, named DAP-LED, which can jointly achieve low-light enhancement and deblurring, benefiting downstream tasks, such as depth estimation, segmentation, and detection in the dark. The key insight is to leverage CLIP to adaptively learn the degradation levels from images at night. This subtly enables learning rich semantic information and visual representation for optimization of the joint tasks. To achieve this, we first introduce a CLIPguided cross-fusion module to obtain multi-scale patch-wise degradation heatmaps from the image embeddings. Then, the heatmaps are fused via the designed CLIP-enhanced transformer blocks to retain useful degradation information for effective model optimization. Experimental results show that, compared to existing methods, our DAP-LED achieves state-of-the-art performance in the dark. Meanwhile, the enhanced results are demonstrated to be effective for three downstream tasks. For demo and more results, please check the project page: https://vlislab22.github.io/dap-led/.

ICRA Conference 2025 Conference Paper

Foresee and Act Ahead: Task Prediction and Pre-Scheduling Enabled Efficient Robotic Warehousing

  • Bo Cao
  • Zhe Liu 0022
  • Xingyao Han
  • Shunbo Zhou
  • Heng Zhang
  • Lijun Han
  • Lin Wang
  • Hesheng Wang 0001

In warehousing systems, to enhance efficiency amid surging demand volumes, much attention has been placed on how to reasonably allocate tasks of delivery to robots. However, the labor of robots is still inevitably wasted to some extent. In this paper, we propose a pre-scheduling enhanced warehousing framework aiming to foresee and act in advance, which consists of task flow prediction and hybrid task allocation. For task prediction, we design the spatio-temporal representations of the task flow and introduce a periodicity-decoupled mechanism tailored for the generation patterns of aggregated orders, and then further extract spatial features of task distribution with a novel combination of graph structures. In hybrid tasks allocation, we consider the known tasks and predicted future tasks simultaneously to optimize the task allocation. In addition, we consider factors such as predicted task uncertainty and sector-level efficiency to realize more balanced and rational allocations. We validate our task prediction model across datasets derived from factories, achieving SOTA performance. Furthermore, we implement our system in a real-world robotic warehouse, demonstrating more than 30% improvements in efficiency.

JBHI Journal 2025 Journal Article

Frozen Large-Scale Pretrained Vision-Language Models are the Effective Foundational Backbone for Multimodal Breast Cancer Prediction

  • Hung Q. Vo
  • Lin Wang
  • Kelvin K. Wong
  • Chika F. Ezeana
  • Xiaohui Yu
  • Wei Yang
  • Jenny Chang
  • Hien V. Nguyen

Breast cancer is a pervasive global health concern among women. Leveraging multimodal data from enterprise patient databases—including Picture Archiving and Communication Systems (PACS) and Electronic Health Records (EHRs)—holds promise for improving prediction. This study introduces a multimodal deep-learning model leveraging mammogram datasets to evaluate breast cancer prediction. Our approach integrates frozen large-scale pretrained vision-language models, showcasing superior performance and stability compared to traditional image-tabular models across two public breast cancer datasets. The model consistently outperforms conventional full fine-tuning methods by using frozen pretrained vision-language models alongside a lightweight trainable classifier. The observed improvements are significant. In the CBIS-DDSM dataset, the Area Under the Curve (AUC) increases from 0. 867 to 0. 902 during validation and from 0. 803 to 0. 830 for the official test set. Within the EMBED dataset, AUC improves from 0. 780 to 0. 805 during validation. In scenarios with limited data, using Breast Imaging-Reporting and Data System category three (BI-RADS 3) cases, AUC improves from 0. 91 to 0. 96 on the official CBIS-DDSM test set and from 0. 79 to 0. 83 on a challenging validation set. This study underscores the benefits of vision-language models in jointly training diverse image-clinical datasets from multiple healthcare institutions, effectively addressing challenges related to non-aligned tabular features. Combining training data enhances breast cancer prediction on the EMBED dataset, outperforming all other experiments. In summary, our research emphasizes the efficacy of frozen large-scale pretrained vision-language models in multimodal breast cancer prediction, offering superior performance and stability over conventional methods, reinforcing their potential for breast cancer prediction.

EAAI Journal 2025 Journal Article

Hybrid Grid Search and Bayesian optimization-based random forest regression for predicting material compression pressure in manufacturing processes

  • Youcheng Zong
  • Yi Nian
  • Chaojie Zhang
  • Xinyu Tang
  • Lin Wang
  • Liqiang Zhang

In the realm of high-performance material development and industrial manufacturing, accurately detecting key data during the material compression process and predicting pressure are crucial. Nevertheless, extant methodologies have not attained accurate and real-time monitoring of data or the prediction of pressure in high-stress material compression processes for high-stress materials. To address this issue, this study introduces a pressure prediction methodology predicated on an optimized random forest regression model. This model undergoes efficient optimization via a univariate grid search combined with the integration of Bayesian filtering techniques, and engages in an analysis of variable importance through SHapley Additive exPlanations values. The study collects real-time compression deformation data and pressure values during the compression process of high-stress materials using an innovative material compression apparatus. It processes the experimental data and implements a noise-induced data augmentation strategy to address the issue of small sample sizes to construct a material compression process dataset. The findings reveal that this method reduces the training time to 3. 5 s and diminishes the mean absolute percentage error to 2. 65%, surpassing the performance of other regression algorithms. Moreover, simulating the predictive performance of this method in different noise environments indicates that this method has a lower root mean square error in most scenarios compared to the original algorithm, demonstrating excellent robustness. Conclusively, this methodology showcases exceptional accuracy and robustness in material compression experiments using 6061 aluminum alloy as an example and can be extended to pressure prediction of other similar materials, offering a broad application prospect.

NeurIPS Conference 2025 Conference Paper

Leader360V: A Large-scale, Real-world 360 Video Dataset for Multi-task Learning in Diverse Environment

  • Weiming Zhang
  • Dingwen Xiao
  • Aobotao DAI
  • Yexin Liu
  • Tianbo Pan
  • Shiqi Wen
  • Lei Chen
  • Lin Wang

360 video captures the complete surrounding scenes with the ultra-large field of view of 360x180. This makes 360 scene understanding tasks, e. g. , segmentation and tracking, crucial for appications, such as autonomous driving, robotics. With the recent emergence of foundation models, the community is, however, impeded by the lack of large-scale, labelled real-world datasets. This is caused by the inherent spherical properties, e. g. , severe distortion in polar regions, and content discontinuities, rendering the annotation costly yet complex. This paper introduces Leader360V, the first large-scale (10K+), labeled real-world 360 video datasets for instance segmentation and tracking. Our datasets enjoy high scene diversity, ranging from indoor and urban settings to natural and dynamic outdoor scenes. To automate annotation, we design an automatic labeling pipeline, which subtly coordinates pre-trained 2D segmentors and large language models (LLMs) to facilitate the labeling. The pipeline operates in three novel stages. Specifically, in the Initial Annotation Phase, we introduce a Semantic- and Distortion-aware Refinement ( SDR ) module, which combines object mask proposals from multiple 2D segmentors with LLM-verified semantic labels. These are then converted into mask prompts to guide SAM2 in generating distortion-aware masks for subsequent frames. In the Auto-Refine Annotation Phase, missing or incomplete regions are corrected either by applying the SDR again or resolving the discontinuities near the horizontal borders. The Manual Revision Phase finally incorporates LLMs and human annotators to further refine and validate the annotations. Extensive user studies and evaluations demonstrate the effectiveness of our labeling pipeline. Meanwhile, experiments confirm that Leader360V significantly enhances model performance for 360 video segmentation and tracking, paving the way for more scalable 360 scene understanding. We release our dataset and code at {https: //leader360v. github. io/Leader360V_HomePage/} for better understanding.

JBHI Journal 2025 Journal Article

MOSAIC: A Multi-Granularity Cross-Modal Framework for Predicting Synergistic Drug Combinations in Personalized Healthcare

  • Licai Zhang
  • Xiao Kang
  • Xinxing Yang
  • Lin Wang
  • Genke Yang
  • Jian Chu

The personalization of cancer treatment through drug combinations is critical for improving healthcare outcomes, increasing effectiveness, and reducing side effects. Computational methods have become increasingly important to prioritize synergistic drug pairs because of the vast search space of possible chemicals. However, existing approaches typically rely solely on global molecular structures, neglecting information exchange between different modality representations and interactions between molecular and fine-grained fragments, leading to limited understanding of drug synergy mechanisms for personalized treatment. To address these limitations, we propose MOSAIC ( M ulti-granularity cr OS s-mod A l method for synerg I stic drug combinations predi C tion), an AI-driven multi-granularity cross-modal method for personalized synergistic drug combination prediction that considers both molecular and fragment-level features. MOSAIC employs a dual-layer representation system, decomposing molecules into chemically meaningful fragments using the BRICS algorithm, facilitating information exchange between graph and SMILES representations through a bidirectional cross-attention mechanism, and ensuring semantic consistency of different modal representations of the same molecular fragment through a contrastive learning framework. Additionally, we designed a bilinear attention network to capture interactions between fragments of different drugs and dynamically integrate multi-granularity feature relationships through a multi-head attention mechanism. Through extensive experiments on multiple real-world datasets, MOSAIC demonstrates superior performance over state-of-the-art methods. Literature validation confirms its predicted novel drug combinations align with existing clinical evidence, while visualization analyses elucidate its capability to pinpoint key molecular fragments critical for drug synergy, providing valuable insights for personalized treatment planning and remote patient monitoring.

NeurIPS Conference 2025 Conference Paper

PASS: Path-selective State Space Model for Event-based Recognition

  • Jiazhou Zhou
  • Kanghao Chen
  • Lei Zhang
  • Lin Wang

Event cameras are bio-inspired sensors that capture intensity changes asynchronously with distinct advantages, such as high temporal resolution. Existing methods for event-based object/action recognition predominantly sample and convert event representation at every fixed temporal interval (or frequency). However, they are constrained to processing a limited number of event lengths and show poor frequency generalization, thus not fully leveraging the event's high temporal resolution. In this paper, we present our PASS framework, exhibiting superior capacity for spatiotemporal event modeling towards a larger number of event lengths and generalization across varying inference temporal frequencies. Our key insight is to learn adaptively encoded event features via the state space models (SSMs), whose linear complexity and generalization on input frequency make them ideal for processing high temporal resolution events. Specifically, we propose a Path-selective Event Aggregation and Scan (PEAS) module to encode events into features with fixed dimensions by adaptively scanning and selecting aggregated event presentation. On top of it, we introduce a novel Multi-faceted Selection Guiding (MSG) loss to minimize the randomness and redundancy of the encoded features during the PEAS selection process. Our method outperforms prior methods on five public datasets and shows strong generalization across varying inference frequencies with less accuracy drop (ours -8. 62% v. s. -20. 69% for the baseline). Moreover, our model exhibits strong long spatiotemporal modeling for a broader distribution of event length (1-10^9), precise temporal perception, and effective generalization for real-world scenarios. Code and checkpoints will be released upon acceptance.

NeurIPS Conference 2025 Conference Paper

PolypSense3D: A Multi-Source Benchmark Dataset for Depth-Aware Polyp Size Measurement in Endoscopy

  • Ruyu Liu
  • Lin Wang
  • Zhou Mingming
  • Jianhua Zhang
  • ZHANG HAOYU
  • Xiufeng Liu
  • Xu Cheng
  • Sixian Chan

Accurate polyp sizing during endoscopy is crucial for cancer risk assessment but is hindered by subjective methods and inadequate datasets lacking integrated 2D appearance, 3D structure, and real-world size information. We introduce PolypSense3D, the first multi-source benchmark dataset specifically targeting depth-aware polyp size measurement. It uniquely integrates over 43, 000 frames from virtual simulations, physical phantoms, and clinical sequences, providing synchronized RGB, dense/sparse depth, segmentation masks, camera parameters, and millimeter-scale size labels derived via a novel forceps-assisted in-vivo annotation technique. To establish its value, we benchmark state-of-the-art segmentation and depth estimation models. Results quantify significant domain gaps between simulated/phantom and clinical data and reveal substantial error propagation from perception stages to final size estimation, with the best fully automated pipelines achieving an average Mean Absolute Error (MAE) of 0. 95 mm on the clinical data subset. Publicly released under CC BY-SA 4. 0 with code and evaluation protocols, PolypSense3D offers a standardized platform to accelerate research in robust, clinically relevant quantitative endoscopic vision. The benchmark dataset and code are available at: https: //github. com/HNUicda/PolypSense3D and https: //doi. org/10. 7910/DVN/K13H89.

NeurIPS Conference 2025 Conference Paper

ProteinConformers: Benchmark Dataset for Simulating Protein Conformational Landscape Diversity and Plausibility

  • Yihang Zhou
  • Chen Wei
  • Minghao Sun
  • Jin Song
  • Yang Li
  • Lin Wang
  • Yang Zhang

Understanding the conformational landscape of proteins is essential for elucidating protein function and facilitating drug design. However, existing protein conformation benchmarks fail to capture the full energy landscape, limiting their ability to evaluate the diversity and physical plausibility of AI-generated structures. We introduce ProteinConformers, a large-scale benchmark dataset comprising over 381, 000 physically realistic conformations for 87 CASP targets. These were derived from more than 40, 000 structural decoys via extensive all-atom molecular dynamics simulations totaling over 6 million CPU hours. Using this dataset, we propose novel metrics to evaluate conformational diversity and plausibility, and systematically benchmark six protein conformation generative models. Our results highlight that leveraging large-scale protein sequence data can enhance a model’s ability to explore conformational space, potentially reducing reliance on MD-derived data. Additionally, we find that PDB and MD datasets influence model performance differently, current models perform well on inter-atomic distance prediction but struggle with inter-residue orientation generation. Overall, our dataset, evaluation metrics, and benchmarking results provide the first comprehensive foundation for assessing generative models in protein conformational modeling. Dataset and instructions are available at https: //huggingface. co/ datasets/Jim990908/ProteinConformers/tree/main. Codes are stored at https: //github. com/auroua/ProteinConformers. An interactive website locates at https: //zhanggroup. org/ProteinConformers.

EAAI Journal 2025 Journal Article

Rectified self-supervised monocular depth estimation loss for nighttime and dynamic scenes

  • Xiaofei Qin
  • Lin Wang
  • Yongchao Zhu
  • Fan Mao
  • Xuedian Zhang
  • Changxiang He
  • Qiulei Dong

Self-supervised monocular depth estimation has attracted much attention in computer vision recently. However, most existing methods assume that the scenes are static and the photometric is consistent, so that their performances tend to degrade significantly in nighttime and dynamic scenes. To address this issue, this paper proposes a self-supervised monocular depth estimation model to tackle two challenges. One is the drastic photometric changes problem due to underexposure of distant areas in nighttime scenes, the other is the moving objects problem in dynamic scenes. In the proposed model, an Effective Area Photometric Loss function (EAPL) is designed which is gated by the Effective Area Mask (EAM) and Potentially Moving Objects Mask (PMOM). Then, a motion flow network is introduced to estimate the motion of moving objects, and a Motion Flow Loss function (MFL) is proposed based on three facts, i. e. , the motion flow of static objects should be zero, most moving objects in autonomous driving scenarios are approximately rigid objects, and the relative motion flows between consecutive frames should be mutually inverse. Finally, a decoupled training approach is provided to facilitate the optimization process of the model. Experimental results show that our model achieves state-of-the-art or second best performance on the nuTonomy Scenes (nuScenes) and Dense Depth for Autonomous Driving (DDAD) dataset which contains many nighttime or dynamic scenes, and also achieves competitive performance on the Karlsruhe Institute of Technology and Toyota Technological Institute at Chicago (KITTI) dataset which is dominated by daytime and static scenes. Codes are available at https: //github. com/pandaswfas/effdepth.

ICRA Conference 2025 Conference Paper

Robo-GS: A Physics Consistent Spatial-Temporal Model for Robotic Arm with Hybrid Representation

  • Haozhe Lou
  • Yurong Liu
  • Yike Pan
  • Yiran Geng
  • Jianteng Chen
  • Wenlong Ma 0006
  • Chenglong Li
  • Lin Wang

The Real2Sim2Real (R2S2R) paradigm is critical for advancing robotic learning. Existing methods lack a comprehensive solution to accurately reconstruct real-world objects with both spatial representations and their associated physics attributes in the Real2Sim stage. We propose a Real2Sim pipeline to generate digital assets enabling high-fidelity simulation. We design a hybrid repre-sentation model that integrates mesh geometry, 3D Gaussian kernels, and physics attributes to enhance the representation of robotic arms in digital assets. This hybrid representation is implemented through a Gaussian-Mesh-Pixel binding technique, which establishes an isomorphic mapping between mesh vertices and the Gaussian model. This enables a fully differentiable rendering pipeline that can be optimized through numerical solvers, achieves high-fidelity rendering via Gaussian Splatting, and facilitates physically plausible simulation of the robotic arm's interaction with its environment through mesh geometry. With the digital assets, we propose a fully manipulable Real2Sim pipeline that standardizes coordinate systems and scales, ensuring the seamless integration of multiple components. To demonstrate its effectiveness, we include datasets covering various robotic manipulation tasks with their mesh reconstructions. Our model achieves state-of-the-art results in realistic rendering and mesh reconstruction quality for robotic applications. Our code and datasets will be made publicly available at robostudioapp. com.

EAAI Journal 2025 Journal Article

Weakly supervised histopathology tissue semantic segmentation with multi-scale voting and online noise suppression

  • Xipeng Pan
  • Hualong Zhang
  • Huahu Deng
  • Huadeng Wang
  • Lingqiao Li
  • Zhenbing Liu
  • Lin Wang
  • Yajun An

The development of an Artificial Intelligence (AI) assisted tissue segmentation method of digital pathology images is critical for cancer diagnosis and prognosis. Excellent performance has been achieved with the current fully supervised segmentation approach, which relies on a huge number of annotated data. However, drawing dense pixel-level annotations on the giga-pixel whole slide image (WSI) is extremely time-consuming and labor-intensive. To this end, we propose a tissue segmentation method using only patch-level classification labels to reduce such annotation burden and significantly improve the quality of the pseudo-masks. We introduce a framework with two phases of classification and segmentation. In the classification phase, we propose a multi-scale voting method on the Class Activation Map (CAM) based model to obtain more stable pseudo masks. In the segmentation phase, an Online Noise Suppression Strategy (ONSS) is proposed to encourage the model to focus on more reliable signals in the pseudo mask rather than noisy signals. Extensive experiments on two weakly supervised pathology image tissue segmentation datasets Lung Adenocarcinoma (LUAD-HistoSeg) and Breast Cancer Semantic Segmentation (BCSS-WSSS) demonstrate our model outperforms state-of-the-art weakly-supervised semantic segmentation (WSSS) methods using patch-level labels. Furthermore, our method exhibits superior generalization ability compared to other models, and demonstrates promising adaptation performance on unseen domains with only small amounts of data.

AIIM Journal 2024 Journal Article

A joint entity Relation Extraction method for document level Traditional Chinese Medicine texts

  • Wenxuan Xu
  • Lin Wang
  • Mingchuan Zhang
  • Junlong Zhu
  • Junqiang Yan
  • Qingtao Wu

Chinese medicine is a unique and complex medical system with complete and rich scientific theories. The textual data of Traditional Chinese Medicine (TCM) contains a large amount of relevant knowledge in the field of TCM, which can serve as guidance for accurate disease diagnosis as well as efficient disease prevention and treatment. Existing TCM texts are disorganized and lack a uniform standard. For this reason, this paper proposes a joint extraction framework by using graph convolutional networks to extract joint entity relations on document-level TCM texts to achieve TCM entity relation mining. More specifically, we first finetune the pre-trained language model by using the TCM domain knowledge to obtain the task-specific model. Taking the integrity of TCM into account, we extract the complete entities as well as the relations corresponding to diagnosis and treatment from the document-level medical cases by using multiple features such as word fusion coding, TCM lexicon information, and multi-relational graph convolutional networks. The experimental results show that the proposed method outperforms the state-of-the-art methods. It has an F1-score of 90. 7% for Name Entity Recognization and 76. 14% for Relation Extraction on the TCM dataset, which significantly improves the ability to extract entity relations from TCM texts. Code is available at https: //github. com/xxxxwx/TCMERE.

EAAI Journal 2024 Journal Article

An adaptive financial trading strategy based on proximal policy optimization and financial signal representation

  • Lin Wang
  • Xuerui Wang

Trading strategies play a crucial role in financial trading. However, due to the significant amount of noise present in financial signals, traditional trading strategies and those based on price prediction often fail to achieve optimal results in real market conditions. Nowadays, with the pervasive noise in financial signals, developing successful trading strategies to achieve high returns has become one of the most prominent and challenging research areas in modern finance. Therefore, an adaptive financial trading strategy, called proximal policy optimization based on financial signal representation trading strategy (FSRPPO), is proposed. This strategy employs financial signal representation technology (FSR), which combines complete ensemble extreme-point symmetric mode decomposition with adaptive noise (CEEMDAN) and modified rescaled range analysis (MRS), to accurately and robustly represent dynamic market states, enhancing profitability. Additionally, the designed reward function reduces trading frequency to lower costs, while the designed action space increases trading flexibility to reduce risks. The experimental results on real stock data demonstrate the outstanding profitability and good risk avoidance ability of our proposed trading strategy, which means that the proposed model can effectively filter out noise from financial signals, extract valuable information, and provide reliable decision support for investors.

AIIM Journal 2024 Journal Article

Clinical knowledge-guided deep reinforcement learning for sepsis antibiotic dosing recommendations

  • Yuan Wang
  • Anqi Liu
  • Jucheng Yang
  • Lin Wang
  • Ning Xiong
  • Yisong Cheng
  • Qin Wu

Sepsis is the third leading cause of death worldwide. Antibiotics are an important component in the treatment of sepsis. The use of antibiotics is currently facing the challenge of increasing antibiotic resistance (Evans et al. , 2021). Sepsis medication prediction can be modeled as a Markov decision process, but existing methods fail to integrate with medical knowledge, making the decision process potentially deviate from medical common sense and leading to underperformance. (Wang et al. , 2021). In this paper, we use Deep Q-Network (DQN) to construct a Sepsis Anti-infection DQN (SAI-DQN) model to address the challenge of determining the optimal combination and duration of antibiotics in sepsis treatment. By setting sepsis clinical knowledge as reward functions to guide DQN complying with medical guidelines, we formed personalized treatment recommendations for antibiotic combinations. The results showed that our model had a higher average value for decision-making than clinical decisions. For the test set of patients, our model predicts that 79. 07% of patients will achieve a favorable prognosis with the recommended combination of antibiotics. By statistically analyzing decision trajectories and drug action selection, our model was able to provide reasonable medication recommendations that comply with clinical practices. Our model was able to improve patient outcomes by recommending appropriate antibiotic combinations in line with certain clinical knowledge.

AIIM Journal 2024 Journal Article

Diagnosis knowledge constrained network based on first-order logic for syndrome differentiation

  • Meiwen Li
  • Lin Wang
  • Qingtao Wu
  • Junlong Zhu
  • Mingchuan Zhang

Traditional Chinese medicine (TCM) has been recognized worldwide as a valuable asset of human medicine. The procedure of TCM is to treatment based on syndrome differentiation. However, the effect of TCM syndrome differentiation relies heavily on the experience of doctors. The gratifying progress of machine learning research in recent years has brought new ideas for TCM syndrome differentiation. In this paper, we propose a deep network model for TCM syndrome differentiation, which improves network performance by injecting TCM syndrome differentiation knowledge in the form of first-order logic into the deep network. Experimental results show that the accuracy of our proposed model reaches 89%, which is significantly better than the deep learning model MLP and other traditional machine learning models. In addition, we present the collected and formatted TCM syndrome differentiation (TSD) dataset, which contains more than 40, 000 TCM clinical records. Moreover, 45 symptoms (“ ”), 322 patterns(“ ”), and more than 500 symptoms are labeled in TSD respectively. To the best of our knowledge, this is the first TCM syndrome differentiation dataset labeling diseases, syndromes and pattern. Such detailed labeling is helpful to explore the relationship between various elements of syndrome differentiation.

NeurIPS Conference 2024 Conference Paper

LaSe-E2V: Towards Language-guided Semantic-aware Event-to-Video Reconstruction

  • Kanghao Chen
  • Hangyu Li
  • Jiazhou Zhou
  • Zeyu Wang
  • Lin Wang

Event cameras harness advantages such as low latency, high temporal resolution, and high dynamic range (HDR), compared to standard cameras. Due to the distinct imaging paradigm shift, a dominant line of research focuses on event-to-video (E2V) reconstruction to bridge event-based and standard computer vision. However, this task remains challenging due to its inherently ill-posed nature: event cameras only detect the edge and motion information locally. Consequently, the reconstructed videos are often plagued by artifacts and regional blur, primarily caused by the ambiguous semantics of event data. In this paper, we find language naturally conveys abundant semantic information, rendering it stunningly superior in ensuring semantic consistency for E2V reconstruction. Accordingly, we propose a novel framework, called LaSe-E2V, that can achieve semantic-aware high-quality E2V reconstruction from a language-guided perspective, buttressed by the text-conditional diffusion models. However, due to diffusion models' inherent diversity and randomness, it is hardly possible to directly apply them to achieve spatial and temporal consistency for E2V reconstruction. Thus, we first propose an Event-guided Spatiotemporal Attention (ESA) module to condition the event data to the denoising pipeline effectively. We then introduce an event-aware mask loss to ensure temporal coherence and a noise initialization strategy to enhance spatial consistency. Given the absence of event-text-video paired data, we aggregate existing E2V datasets and generate textual descriptions using the tagging models for training and evaluation. Extensive experiments on three datasets covering diverse challenging scenarios (e. g. , fast motion, low light) demonstrate the superiority of our method. Demo videos for the results are attached to the project page.

NeurIPS Conference 2024 Conference Paper

LinNet: Linear Network for Efficient Point Cloud Representation Learning

  • Hao Deng
  • Kunlei Jing
  • Shengmei Cheng
  • Cheng Liu
  • Jiawei Ru
  • Jiang Bo
  • Lin Wang

Point-based methods have made significant progress, but improving their scalability in large-scale 3D scenes is still a challenging problem. In this paper, we delve into the point-based method and develop a simpler, faster, stronger variant model, dubbed as LinNet. In particular, we first propose the disassembled set abstraction (DSA) module, which is more effective than the previous version of set abstraction. It achieves more efficient local aggregation by leveraging spatial anisotropy and channel anisotropy separately. Additionally, by mapping 3D point clouds onto 1D space-filling curves, we enable parallelization of downsampling and neighborhood queries on GPUs with linear complexity. LinNet, as a purely point-based method, outperforms most previous methods in both indoor and outdoor scenes without any extra attention, and sparse convolution but merely relying on a simple MLP. It achieves the mIoU of 73. 7\%, 81. 4\%, and 69. 1\% on the S3DIS Area5, NuScenes, and SemanticKITTI validation benchmarks, respectively, while speeding up almost 10x times over PointNeXt. Our work further reveals both the efficacy and efficiency potential of the vanilla point-based models in large-scale representation learning. Our code will be available upon publication.

ICLR Conference 2024 Conference Paper

Rethinking CNN's Generalization to Backdoor Attack from Frequency Domain

  • Quanrui Rao
  • Lin Wang
  • Wuying Liu

Convolutional neural network (CNN) is easily affected by backdoor injections, whose models perform normally on clean samples but produce specific outputs on poisoned ones. Most of the existing studies have focused on the effect of trigger feature changes of poisoned samples on model generalization in spatial domain. We focus on the mechanism of CNN memorize poisoned samples in frequency domain, and find that CNN generate generalization to poisoned samples by memorizing the frequency domain distribution of trigger changes. We also explore the influence of trigger perturbations in different frequency domain components on the generalization of poisoned models from visible and invisible backdoor attacks, and prove that high-frequency components are more susceptible to perturbations than low-frequency components. Based on the above fundings, we propose a universal invisible strategy for visible triggers, which can achieve trigger invisibility while maintaining raw attack performance. We also design a novel frequency domain backdoor attack method based on low-frequency semantic information, which can achieve 100\% attack accuracy on multiple models and multiple datasets, and can bypass multiple defenses.

NeurIPS Conference 2023 Conference Paper

DELTA: Diverse Client Sampling for Fasting Federated Learning

  • Lin Wang
  • Yongxin Guo
  • Tao Lin
  • Xiaoying Tang

Partial client participation has been widely adopted in Federated Learning (FL) to reduce the communication burden efficiently. However, an inadequate client sampling scheme can lead to the selection of unrepresentative subsets, resulting in significant variance in model updates and slowed convergence. Existing sampling methods are either biased or can be further optimized for faster convergence. In this paper, we present DELTA, an unbiased sampling scheme designed to alleviate these issues. DELTA characterizes the effects of client diversity and local variance, and samples representative clients with valuable information for global model updates. In addition, DELTA is a proven optimal unbiased sampling scheme that minimizes variance caused by partial client participation and outperforms other unbiased sampling schemes in terms of convergence. Furthermore, to address full-client gradient dependence, we provide a practical version of DELTA depending on the available clients' information, and also analyze its convergence. Our results are validated through experiments on both synthetic and real-world datasets.

EAAI Journal 2023 Journal Article

Fault diagnosis based on residual–knowledge–data jointly driven method for chillers

  • Zhanwei Wang
  • Boyang Liang
  • Jingjing Guo
  • Lin Wang
  • Yingying Tan
  • Xiuzhen Li

Fault diagnosis is crucial for energy conversation in building energy systems. There are three different types of fault diagnosis methods: residual-, knowledge-, and data-driven. Each of them has its unique advantages and drawbacks. This study proposes a hybrid Bayesian network (HBN) to merge these methods to make their advantages complementary. The symptom layer of the HBN includes three different types of nodes: residual nodes composed of the feature residuals from the residual-driven part, knowledge nodes composed of the knowledge from the knowledge-driven part, and data nodes composed of the feature measurements from the data-driven part. Based on such fusion mechanisms, residual, knowledge, and data are merged into a single framework, and play roles in a parallel way, meaning that it can tolerate missing any kind of part in the fault diagnosis process. The proposed method can be easily customized by users, and, to a certain extent, overcomes the weaknesses of the individual methods when used separately, thus achieving outstanding field application performance. Meanwhile, a generic framework to develop the residual–knowledge–data jointly driven method is given. Applied to two experimental chillers and compared with existing frequently-used methods, the proposed method is proven to have better overall and individual performance.

ICRA Conference 2023 Conference Paper

Improved Event-Based Dense Depth Estimation via Optical Flow Compensation

  • Dianxi Shi
  • Luoxi Jing
  • Ruihao Li 0001
  • Zhe Liu 0029
  • Lin Wang
  • Huachi Xu
  • Yi Zhang

Event cameras have the potential to overcome the limitations of classical computer vision in real-world applications. Depth estimation is a crucial step for high-level robotics tasks and has attracted much attention from the community. In this paper, we propose an event-based dense depth estimation architecture, Mixed-EF2DNet, which firstly predicts inter-grid optical flow to compensate for lost temporal information, and then estimates multiple contextual depth maps that are fused to generate a robust depth estimation map. To supervise the network training, we further design a smoothing loss function used to smooth local depth estimates and facilitate estimating reasonable depth for pixels without events. In addition, we introduce SE-resblocks in the depth network to enhance the network representation by selecting feature channels. Experimental evaluations on both real-world and synthetic datasets show that our method performs better in terms of accuracy when compared to state-of-the-art algorithms, especially in scene detail estimation. Besides, our method demonstrates excellent generalization in cross-dataset tasks.

JBHI Journal 2023 Journal Article

Interpretable Inference and Classification of Tissue Types in Histological Colorectal Cancer Slides Based on Ensembles Adaptive Boosting Prototype Tree

  • Meiyan Liang
  • Ru Wang
  • Jianan Liang
  • Lin Wang
  • Bo Li
  • Xiaojun Jia
  • Yu Zhang
  • Qinghui Chen

Digital pathology images are treated as the “gold standard” for the diagnosis of colorectal lesions, especially colon cancer. Real-time, objective and accurate inspection results will assist clinicians to choose symptomatic treatment in a timely manner, which is of great significance in clinical medicine. However, Manual methods suffers from long inspection cycle and serious reliance on subjective interpretation. It is also a challenging task for existing computer-aided diagnosis methods to obtain models that are both accurate and interpretable. Models that exhibit high accuracy are always more complex and opaque, while interpretable models may lack the necessary accuracy. Therefore, the framework of ensemble adaptive boosting prototype tree is proposed to predict the colorectal pathology images and provide interpretable inference by visualizing the decision-making process in each base learner. The results showed that the proposed method could effectively address the “accuracy-interpretability trade-off” issue by ensemble of m adaptive boosting neural prototype trees. The superior performance of the framework provides a novel paradigm for interpretable inference and high-precision prediction of pathology image patches in computational pathology.

NeurIPS Conference 2023 Conference Paper

NurViD: A Large Expert-Level Video Database for Nursing Procedure Activity Understanding

  • Ming Hu
  • Lin Wang
  • Siyuan Yan
  • Don Ma
  • Qingli Ren
  • Peng Xia
  • Wei Feng
  • Peibo Duan

The application of deep learning to nursing procedure activity understanding has the potential to greatly enhance the quality and safety of nurse-patient interactions. By utilizing the technique, we can facilitate training and education, improve quality control, and enable operational compliance monitoring. However, the development of automatic recognition systems in this field is currently hindered by the scarcity of appropriately labeled datasets. The existing video datasets pose several limitations: 1) these datasets are small-scale in size to support comprehensive investigations of nursing activity; 2) they primarily focus on single procedures, lacking expert-level annotations for various nursing procedures and action steps; and 3) they lack temporally localized annotations, which prevents the effective localization of targeted actions within longer video sequences. To mitigate these limitations, we propose NurViD, a large video dataset with expert-level annotation for nursing procedure activity understanding. NurViD consists of over 1. 5k videos totaling 144 hours, making it approximately four times longer than the existing largest nursing activity datasets. Notably, it encompasses 51 distinct nursing procedures and 177 action steps, providing a much more comprehensive coverage compared to existing datasets that primarily focus on limited procedures. To evaluate the efficacy of current deep learning methods on nursing activity understanding, we establish three benchmarks on NurViD: procedure recognition on untrimmed videos, procedure and action recognition on trimmed videos, and action detection. Our benchmark and code will be available at https: //github. com/minghu0830/NurViD-benchmark.

AAAI Conference 2023 Conference Paper

OPT-GAN: A Broad-Spectrum Global Optimizer for Black-Box Problems by Learning Distribution

  • Minfang Lu
  • Shuai Ning
  • Shuangrong Liu
  • Fengyang Sun
  • Bo Zhang
  • Bo Yang
  • Lin Wang

Black-box optimization (BBO) algorithms are concerned with finding the best solutions for problems with missing analytical details. Most classical methods for such problems are based on strong and fixed a priori assumptions, such as Gaussianity. However, the complex real-world problems, especially when the global optimum is desired, could be very far from the a priori assumptions because of their diversities, causing unexpected obstacles. In this study, we propose a generative adversarial net-based broad-spectrum global optimizer (OPT-GAN) which estimates the distribution of optimum gradually, with strategies to balance exploration-exploitation trade-off. It has potential to better adapt to the regularity and structure of diversified landscapes than other methods with fixed prior, e.g., Gaussian assumption or separability. Experiments on diverse BBO benchmarks and high dimensional real world applications exhibit that OPT-GAN outperforms other traditional and neural net-based BBO algorithms. The code and Appendix are available at https://github.com/NBICLAB/OPT-GAN

AAAI Conference 2023 Conference Paper

Pixel Is All You Need: Adversarial Trajectory-Ensemble Active Learning for Salient Object Detection

  • Zhenyu Wu
  • Lin Wang
  • Wei Wang
  • Qing Xia
  • Chenglizhao Chen
  • Aimin Hao
  • Shuo Li

Although weakly-supervised techniques can reduce the labeling effort, it is unclear whether a saliency model trained with weakly-supervised data (e.g., point annotation) can achieve the equivalent performance of its fully-supervised version. This paper attempts to answer this unexplored question by proving a hypothesis: there is a point-labeled dataset where saliency models trained on it can achieve equivalent performance when trained on the densely annotated dataset. To prove this conjecture, we proposed a novel yet effective adversarial trajectory-ensemble active learning (ATAL). Our contributions are three-fold: 1) Our proposed adversarial attack triggering uncertainty can conquer the overconfidence of existing active learning methods and accurately locate these uncertain pixels. 2) Our proposed trajectory-ensemble uncertainty estimation method maintains the advantages of the ensemble networks while significantly reducing the computational cost. 3) Our proposed relationship-aware diversity sampling algorithm can conquer oversampling while boosting performance. Experimental results show that our ATAL can find such a point-labeled dataset, where a saliency model trained on it obtained 97%-99% performance of its fully-supervised version with only 10 annotated points per image.

AAAI Conference 2023 Conference Paper

SEPT: Towards Scalable and Efficient Visual Pre-training

  • Yiqi Lin
  • Huabin Zheng
  • Huaping Zhong
  • Jinjing Zhu
  • Weijia Li
  • Conghui He
  • Lin Wang

Recently, the self-supervised pre-training paradigm has shown great potential in leveraging large-scale unlabeled data to improve downstream task performance. However, increasing the scale of unlabeled pre-training data in real-world scenarios requires prohibitive computational costs and faces the challenge of uncurated samples. To address these issues, we build a task-specific self-supervised pre-training framework from a data selection perspective based on a simple hypothesis that pre-training on the unlabeled samples with similar distribution to the target task can bring substantial performance gains. Buttressed by the hypothesis, we propose the first yet novel framework for Scalable and Efficient visual Pre-Training (SEPT) by introducing a retrieval pipeline for data selection. SEPT first leverage a self-supervised pre-trained model to extract the features of the entire unlabeled dataset for retrieval pipeline initialization. Then, for a specific target task, SEPT retrievals the most similar samples from the unlabeled dataset based on feature similarity for each target instance for pre-training. Finally, SEPT pre-trains the target model with the selected unlabeled samples in a self-supervised manner for target data finetuning. By decoupling the scale of pre-training and available upstream data for a target task, SEPT achieves high scalability of the upstream dataset and high efficiency of pre-training, resulting in high model architecture flexibility. Results on various downstream tasks demonstrate that SEPT can achieve competitive or even better performance compared with ImageNet pre-training while reducing the size of training samples by one magnitude without resorting to any extra annotations.

JMLR Journal 2023 Journal Article

SQLFlow: An Extensible Toolkit Integrating DB and AI

  • Jun Zhou
  • Ke Zhang
  • Lin Wang
  • Hua Wu
  • Yi Wang
  • Chaochao Chen

Integrating AI algorithms into databases is an ongoing effort in both academia and industry. We introduce SQLFlow, a toolkit seamlessly combining data manipulations and AI operations that can be run locally or remotely. SQLFlow extends SQL syntax to support typical AI tasks including model training, inference, interpretation, and mathematical optimization. It is compatible with a variety of database management systems (DBMS) and AI engines, including MySQL, TiDB, MaxCompute, and Hive, as well as TensorFlow, scikit-learn, and XGBoost. Documentations and case studies are available at https://sqlflow.org. The source code and additional details can be found at https://github.com/sql-machine-learning/sqlflow. &copy JMLR 2023. ( edit, beta )

IJCAI Conference 2023 Conference Paper

STS-GAN: Can We Synthesize Solid Texture with High Fidelity from Arbitrary 2D Exemplar?

  • Xin Zhao
  • Jifeng Guo
  • Lin Wang
  • Fanqi Li
  • Jiahao Li
  • Junteng Zheng
  • Bo Yang

Solid texture synthesis (STS), an effective way to extend a 2D exemplar to a 3D solid volume, exhibits advantages in computational photography. However, existing methods generally fail to accurately learn arbitrary textures, which may result in the failure to synthesize solid textures with high fidelity. In this paper, we propose a novel generative adversarial nets-based framework (STS-GAN) to extend the given 2D exemplar to arbitrary 3D solid textures. In STS-GAN, multi-scale 2D texture discriminators evaluate the similarity between the given 2D exemplar and slices from the generated 3D texture, promoting the 3D texture generator synthesizing realistic solid textures. Finally, experiments demonstrate that the proposed method can generate high-fidelity solid textures with similar visual characteristics to the 2D exemplar.

AAAI Conference 2023 Conference Paper

Unsupervised Domain Adaptation for Medical Image Segmentation by Selective Entropy Constraints and Adaptive Semantic Alignment

  • Wei Feng
  • Lie Ju
  • Lin Wang
  • Kaimin Song
  • Xin Zhao
  • Zongyuan Ge

Generalizing a deep learning model to new domains is crucial for computer-aided medical diagnosis systems. Most existing unsupervised domain adaptation methods have made significant progress in reducing the domain distribution gap through adversarial training. However, these methods may still produce overconfident but erroneous results on unseen target images. This paper proposes a new unsupervised domain adaptation framework for cross-modality medical image segmentation. Specifically, We first introduce two data augmentation approaches to generate two sets of semantics-preserving augmented images. Based on the model's predictive consistency on these two sets of augmented images, we identify reliable and unreliable pixels. We then perform a selective entropy constraint: we minimize the entropy of reliable pixels to increase their confidence while maximizing the entropy of unreliable pixels to reduce their confidence. Based on the identified reliable and unreliable pixels, we further propose an adaptive semantic alignment module which performs class-level distribution adaptation by minimizing the distance between same class prototypes between domains, where unreliable pixels are removed to derive more accurate prototypes. We have conducted extensive experiments on the cross-modality cardiac structure segmentation task. The experimental results show that the proposed method significantly outperforms the state-of-the-art comparison algorithms. Our code and data are available at https://github.com/fengweie/SE_ASA.

JBHI Journal 2022 Journal Article

A Time-Series Feature-Based Recursive Classification Model to Optimize Treatment Strategies for Improving Outcomes and Resource Allocations of COVID-19 Patients

  • Lin Wang
  • Zheng Yin
  • Mamta Puppala
  • Chika Ezeana
  • Kelvin Wong
  • Tiancheng He
  • Deepa Gotur
  • Stephen Wong

This paper presents a novel Lasso Logistic Regression model based on feature-based time series data to determine disease severity and when to administer drugs or escalate intervention procedures in patients with coronavirus disease 2019 (COVID-19). Advanced features were extracted from highly enriched and time series vital sign data of hospitalized COVID-19 patients, including oxygen saturation readings, and with a combination of patient demographic and comorbidity information, as inputs into the dynamic feature-based classification model. Such dynamic combinations brought deep insights to guide clinical decision-making of complex COVID-19 cases, including prognosis prediction, timing of drug administration, admission to intensive care units, and application of intervention procedures like ventilation and intubation. The COVID-19 patient classification model was developed utilizing 900 hospitalized COVID-19 patients in a leading multi-hospital system in Texas, United States. By providing mortality prediction based on time-series physiologic data, demographics, and clinical records of individual COVID-19 patients, the dynamic feature-based classification model can be used to improve efficacy of the COVID-19 patient treatment, prioritize medical resources, and reduce casualties. The uniqueness of our model is that it is based on just the first 24 hours of vital sign data such that clinical interventions can be decided early and applied effectively. Such a strategy could be extended to prioritize resource allocations and drug treatment for futurepandemic events.

AAAI Conference 2022 Conference Paper

Deconvolutional Density Network: Modeling Free-Form Conditional Distributions

  • Bing Chen
  • Mazharul Islam
  • Jisuo Gao
  • Lin Wang

Conditional density estimation (CDE) is the task of estimating the probability of an event conditioned on some inputs. A neural network (NN) can also be used to compute the output distribution for continuous-domain, which can be viewed as an extension of regression task. Nevertheless, it is difficult to explicitly approximate a distribution without knowing the information of its general form a priori. In order to fit an arbitrary conditional distribution, discretizing the continuous domain into bins is an effective strategy, as long as we have sufficiently narrow bins and very large data. However, collecting enough data is often hard to reach and falls far short of that ideal in many circumstances, especially in multivariate CDE for the curse of dimensionality. In this paper, we demonstrate the benefits of modeling free-form conditional distributions using a deconvolution-based neural net framework, coping with data deficiency problems in discretization. It has the advantage of being flexible but also takes advantage of the hierarchical smoothness offered by the deconvolution layers. We compare our method to a number of other density-estimation approaches and show that our Deconvolutional Density Network (DDN) outperforms the competing methods on many univariate and multivariate tasks. The code of DDN is available at https: //github. com/NBICLAB/DDN

JBHI Journal 2022 Journal Article

Guest Editorial Sensing Psychological Parameters and AI-Enabled Emotion Care for Human Wellness

  • Min Chen
  • Hamid Gharavi
  • Lin Wang
  • Victor C. M. Leung
  • Zhongchun Liu
  • Iztok Humar

The papers in this special section focus on the use of artificial intelligence (AI)-enabled technologies to address human wellness. As the COVID-19 pandemic took hold over the last several years, there was an urgent demand to pay more attention to psychological health for human wellness by providing methods and means of sensing psychological parameters, emotional care and mental disorder patient monitoring, especially during these difficult times. With the aid of wearable computing technology and artificial intelligence, emotion and mental disorder detections are available through sensing and analyzing psychological parameters. Discusses the use of AI-based patient monitoring and the ability to monitor human wellness via remote sensing technologies. The papers in this issue provide a snapshot of some of the latest research advances on the research and application of Small Things and Big Data, knowledge discovery and knowledge representation for the combination towards biomedical and health informatics.

YNICL Journal 2022 Journal Article

Thalamic-insomnia phenotype in E200K Creutzfeldt–Jakob disease: A PET/MRI study

  • Hong Ye
  • Min Chu
  • Zhongyun Chen
  • Kexin Xie
  • Li Liu
  • Haitian Nan
  • Yue Cui
  • Jing Zhang

BACKGROUND: Insomnia and thalamic involvement were frequently reported in patients with genetic Creutzfeldt-Jakob disease (gCJD) with E200K mutations, suggesting E200K might have discrepancy with typical sporadic CJD (sCJD). The study aimed to explore the clinical and neuroimage characteristics of genetic E200K CJD patients by comprehensive neuroimage analysis. METHODS: Six patients with gCJD carried E200K mutation on Prion Protein (PRNP) gene, 13 patients with sporadic CJD, and 22 age- and sex-matched normal controls were enrolled in the study. All participants completed a hybrid positron emission tomography/magnetic resonance imaging (PET/MRI) examination. Signal intensity on diffusion-weighted imaging (DWI) and metabolism on PET were visually rating analyzed, statistical parameter mapping analysis was performed on PET and 3D-T1 images. Clinical and imaging characteristics were compared between the E200K, sCJD, and control groups. RESULTS: There was no group difference in age or gender among the E200K, sCJD, and control groups. Insomnia was a primary complaint in patients with E200K gCJD (4/2 versus 1/12, p = 0.007). Hyperintensity on DWI and hypometabolism on PET of the thalamus were observed during visual rating analysis of images in patients with E200K gCJD. Gray matter atrophy (uncorrected p < 0.001) and hypometabolism (uncorrected p < 0.001) of the thalamus were more pronounced in patients with E200K gCJD. CONCLUSION: The clinical and imaging characteristics of patients with gCJD with PRNP E200K mutations manifested as a thalamic-insomnia phenotype. PET is a sensitive approach to help identify the functional changes in the thalamus in prion disease.

EAAI Journal 2021 Journal Article

Effective electricity load forecasting using enhanced double-reservoir echo state network

  • Lu Peng
  • Sheng-Xiang Lv
  • Lin Wang
  • Zi-Yun Wang

Accurate electricity load prediction is essential to ensure the efficient, reliable, and secure operation of the power system. In this study, a hybrid forecasting method called improved backtracking search optimization algorithm (IBSA)–double-reservoir echo state network (DRESN) (IBSA–DRESN) is proposed on the basis of IBSA and DRESN. Mutual information is utilized to eliminate low-significance input features and retain key input features. The DRESN structure aims to increase the diversity of the network. Roulette strategy, adaptive mutation operator, and niche operator is introduced to improve the standard BSA algorithm. The IBSA is applied to optimize several critical parameters in the DRESN neural network. The proposed IBSA–DRESN method is evaluated using two electricity load datasets, namely, North-America and PJM. Compared with eight popular benchmark models, prediction results show that IBSA–DRESN is more accurate for one-step ahead electricity load forecasting. In one-day ahead forecasting, IBSA–DRESN obtains better prediction performance in most cases.

AAMAS Conference 2021 Conference Paper

Modeling Replicator Dynamics in Stochastic Games Using Markov Chain Method

  • Chuang Deng
  • Zhihai Rong
  • Lin Wang
  • Xiaofan Wang

In stochastic games, individuals need to make decisions in multiple states and transitions between states influence the dynamics of strategies significantly. In this work, by describing the dynamic process in stochastic game as a Markov chain and utilizing the transition matrix, we introduce a new method, named state-transition replicator dynamics, to obtain the replicator dynamics of a stochastic game. Based on our proposed model, we can gain qualitative and detailed insights into the influence of transition probabilities on the dynamics of strategies. We illustrate that a set of unbalanced transition probabilities can help players to overcome the social dilemmas and lead to mutual cooperation in a cooperation back state, even if the stochastic game has the same social dilemmas in each state. Moreover, we also present that a set of specifically designed transition probabilities can fix the expected payoffs of one player and make him lose the motivation to update his strategies in the stochastic game.

EAAI Journal 2020 Journal Article

Discriminative sparse embedding based on adaptive graph for dimension reduction

  • Zhonghua Liu
  • Kaiming Shi
  • Kaibing Zhang
  • Weihua Ou
  • Lin Wang

The traditional manifold learning methods usually utilize the original observed data to directly define the intrinsic structure among data. Because the original samples often contain a deal of redundant information or it is corrupted by noises, it leads to the unreliability of the obtained intrinsic structure. In addition, the intrinsic structure learning and subspace learning are completely separated. For solving above problems, this paper presents a novel dimension reduction method termed discriminative sparse embedding (DSE) based on adaptive graph. By projecting the original samples into a low-dimensional subspace, DSE learns a sparse weight matrix, which can reduce the effects of redundant information and noises of the original data, and uncover essential structural relationship among the data. In DSE, the robust subspace is learned from the original data. Meanwhile, the intrinsic local structure and the optimal subspace can be simultaneously learned, in which they are mutually improved, and the accurate structure can be captured, and the optimal subspace can be obtained. We propose an alternative and iterative method to solve the DSE model. In order to evaluate the performance of DSE, it is compared with some state-of-the-art feature extraction algorithms. Various experiments show that our DSE is effective and feasible.

EAAI Journal 2020 Journal Article

Estimating cement compressive strength using three-dimensional microstructure images and deep belief network

  • Jifeng Guo
  • Meihui Li
  • Lin Wang
  • Bo Yang
  • Liangliang Zhang
  • Zhenxiang Chen
  • Shiyuan Han
  • Laura Garcia-Hernandez

The estimation of cement compressive strength is of great significance in the quality inspections, technological designs, and engineering applications for cement. Compared to destructive methods, the nondestructive estimation approaches save the cost in the manpower and material. However, the existing nondestructive methods have the large error because the used influence factors are difficult to control and the used two-dimensional microstructure images can not reflect the specific spatial structure of the entire cement. In this paper, a novel model is proposed to estimate the cement compressive strength using three-dimensional microstructure images and deep belief network. To reduce the computation consumption induced by three-dimensional images with abundant information, this method extracts image features that reflect the cement hydration state to estimate cement compressive strength. Deep belief network is applied to build the estimation model. Its unique training pattern and flexibility of parameters improve the ability to learn nonlinear relationships between microstructure images and cement compressive strength. Furthermore, the training processes are accelerated on the graphics processing units. The experimental results prove that the proposed method estimates cement compressive strength nondestructively and improves the efficiency.

YNIMG Journal 2019 Journal Article

Cerebral microbleed detection using Susceptibility Weighted Imaging and deep learning

  • Saifeng Liu
  • David Utriainen
  • Chao Chai
  • Yongsheng Chen
  • Lin Wang
  • Sean K. Sethi
  • Shuang Xia
  • E. Mark Haacke

Detecting cerebral microbleeds (CMBs) is important in diagnosing a variety of diseases including dementia, stroke and traumatic brain injury. However, manual detection of CMBs can be time-consuming and prone to errors, whereas the current automatic algorithms for CMB detection are usually limited by large number of false positives. In this study, we present a two-stage CMB detection framework which contains a candidate detection stage based on a 3D fast radial symmetry transform of the composite images from Susceptibility Weighted Imaging (SWI), and a false positive reduction stage based on deep residual neural networks using both the SWI and the high-pass filtered phase images. While the SWI images provide exquisite sensitivity to the presence of blood products, the high-pass filtered phase images enable the differentiation of diamagnetic calcifications from paramagnetic microbleeds. The deep learning model was trained using 154 data sets, and the best models were selected using 25 validation data sets. Finally, the models were tested using 41 cases, including 13 hemodialysis cases, 9 traumatic brain injury cases, 9 stroke cases and 10 healthy controls. Using 3D SWI and high-pass filtered phase images as input, the best model led to a sensitivity of 95. 8%, a precision of 70. 9%, and 1. 6 false positives per case. This model achieved similar performance to the most experienced human rater and outperformed recently reported CMB detection methods. This study demonstrates the potential of applying deep learning techniques to medical imaging for improving efficiency and accuracy in diagnosis.

EAAI Journal 2019 Journal Article

Optimizing echo state network with backtracking search optimization algorithm for time series forecasting

  • Zhigang Wang
  • Yu-Rong Zeng
  • Sirui Wang
  • Lin Wang

The echo state network (ESN) is a state-of-the art reservoir computing approach, which is particularly effective for time series forecasting problems because it is coupled with a time parameter. However, the linear regression algorithm commonly used to compute the output weights of ESN could usually cause the trained network over-fitted and thus obtain unsatisfactory results. To overcome the problem, we present four optimized ESNs that are based on the backtracking search optimization algorithm (BSA) or its variants to improve generalizability. Concretely, we utilize BSA and its variants to determine the most appropriate output weights of ESN given that the optimization problem is complex while BSA is a novel evolutionary algorithm that effectively unscrambles optimal solutions in complex spaces. The three BSA variants, namely, adaptive population selection scheme (APSS)–BSA, adaptive mutation factor strategy (AMFS)–BSA, and APSS&AMFS–BSA, were designed to further improve the performance of BSA. Time series forecasting experiments were performed using two real-life time series. The experimental results of the optimized ESNs were compared with those of the basic ESN without optimization, and the two other comparison approaches, as well as the other existing approaches. Experimental results showed that (a) the results of the optimized ESNs are more accurate than that of basic ESN and (b) APSS&AMFS–BSA–ESN nearly outperforms basic ESN, the three other optimized ESNs, the two comparison approaches, and other existing optimization approaches.

EAAI Journal 2018 Journal Article

Accelerating nearest neighbor partitioning neural network classifier based on CUDA

  • Lin Wang
  • Xuehui Zhu
  • Bo Yang
  • Jifeng Guo
  • Shuangrong Liu
  • Meihui Li
  • Jian Zhu
  • Ajith Abraham

The nearest neighbor partitioning (NNP) method is a high performance approach which is used for improving traditional neural network classifiers. However, the construction process of NNP model is very time-consuming, particularly for large data sets, thus limiting its range of application. In this study, a parallel NNP method is proposed to accelerate NNP based on Compute Unified Device Architecture(CUDA). In this method, blocks and threads are used to evaluate potential neural networks and to perform parallel subtasks, respectively. Experimental results manifest that the proposed parallel method improves performance of NNP neural network classifier. Furthermore, the application of parallel NNP in performance evaluation of cement microstructure indicates that the proposed approach has favorable performance.

ICML Conference 2018 Conference Paper

Adversarial Attack on Graph Structured Data

  • Hanjun Dai
  • Hui Li
  • Tian Tian 0001
  • Xin Huang
  • Lin Wang
  • Jun Zhu 0001
  • Le Song

Deep learning on graph structures has shown exciting results in various applications. However, few attentions have been paid to the robustness of such models, in contrast to numerous research work for image or text adversarial attack and defense. In this paper, we focus on the adversarial attacks that fool deep learning models by modifying the combinatorial structure of data. We first propose a reinforcement learning based attack method that learns the generalizable attack policy, while only requiring prediction labels from the target classifier. We further propose attack methods based on genetic algorithms and gradient descent in the scenario where additional prediction confidence or gradients are available. We use both synthetic and real-world data to show that, a family of Graph Neural Network models are vulnerable to these attacks, in both graph-level and node-level classification tasks. We also show such attacks can be used to diagnose the learned classifiers.

AAAI Conference 2016 Conference Paper

Toward a Better Understanding of Deep Neural Network Based Acoustic Modelling: An Empirical Investigation

  • Xingfu Wang
  • Lin Wang
  • Jing Chen
  • Litao Wu

Recently, deep neural networks (DNNs) have outperformed traditional acoustic models on a variety of speech recognition benchmarks. However, due to system differences across research groups, although a tremendous breadth and depth of related work has been established, it is still not easy to assess the performance improvements of a particular architectural variant from examining the literature when building DNN acoustic models. Our work aims to uncover which variations among baseline systems are most relevant for automatic speech recognition (ASR) performance via a series of systematic tests on the limits of the major architectural choices. By holding all the other components fixed, we are able to explore the design and training decisions without being confounded by the other influencing factors. Our experiment results suggest that a relatively simple DNN architecture and optimization technique produces strong results. These findings, along with previous work, not only help build a better understanding towards why DNN acoustic models perform well or how they might be improved, but also help establish a set of best practices for new speech corpora and language understanding task variants.

v2026.09.13