Arrow Research search

Author name cluster

Bo Wu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

32 papers
2 author rows

Possible papers

32

EAAI Journal 2026 Journal Article

Dual-attention guided one-shot ore detection with classification loss enhancement

  • Guodong Sun
  • Mingxuan Liu
  • Long Wang
  • Shicheng Li
  • Bo Wu

In ore particle size detection, traditional object detectors often suffer from high computational cost and limited classification capability. Few-shot object detection algorithms alleviate these issues by reducing data requirements and enhancing generalization. Building upon this, one-shot object detection offers greater flexibility, enabling rapid adaptation to new ore samples. In this study, we propose a dual-attention guided detector with enhanced classification supervision specifically optimized for ore particle size detection. The detector integrates contextual and channel attention mechanisms to enable multi-dimensional dynamic modeling of the input, significantly improving feature extraction from limited ore samples—particularly in fine-grained detection tasks. Moreover, we introduce an auxiliary classification design that encourages the network to learn more diverse and discriminative representations, thereby improving category discrimination under extremely limited supervision. Notably, our detector achieves an average precision of 46. 8 percent under the one-shot fine-tuning setting, operates in real time at 48 frames per second, and has a compact model size of 16. 1 megabytes. These results indicate that the proposed method provides an effective and lightweight solution for ore detection under scarce-annotation conditions in the tested single-class setting.

AAAI Conference 2026 Conference Paper

From Semantics to Spectrum: A New Lens on Graph Augmentation Strategy

  • Xiangping Zheng
  • Xiuxin Hao
  • Bo Wu
  • Wei Li
  • Bin Ren
  • Bin Tang
  • Yuhui Guo
  • Xun Liang

Graph augmentation is a cornerstone of effective graph contrastive learning, yet existing methods often rely on random designed perturbations, which may distort latent semantics and impair representation quality. In this work, we argue that semantic consistency can be effectively approximated by low-frequency components in the spectral domain, offering a principled proxy for guiding augmentation. Based on this insight, we propose Frequency-Aware Graph Contrastive Learning (FA-GCL), a novel framework that explicitly preserves low-frequency signals while selectively perturbing high-frequency components. By aligning augmentation with frequency-aware decomposition, FA-GCL generates diverse yet semantically coherent views, mitigating semantic drift and enhancing representational discrimination. Extensive experiments across multiple benchmarks demonstrate that FA-GCL consistently outperforms state-of-the-art baselines with statistically significant gains, validating its exclusive merits.

YNIMG Journal 2026 Journal Article

Partial volume correction for quantifying venous oxygen saturation levels using contrast-enhanced MRI

  • Sagar Buch
  • Yifan Lv
  • Mingming Wang
  • Bo Wu
  • Ryan M. Smith
  • Yu Luo
  • E. Mark Haacke

Quantifying brain oxygenation is crucial for diagnosing and managing neurological conditions like stroke. Quantitative susceptibility mapping (QSM), an MRI technique, can measure venous oxygen saturation (Yv) but is hampered by partial volume effects (PVEs) in small vessels, leading to inaccurate measurements. This study aims to develop a robust method to mitigate these PVEs for the QSM-derived Yv and oxygen extraction fraction (OEF) in small cerebral veins. We integrated QSM with high-resolution, contrast-enhanced T1-weighted imaging to generate regional cerebral blood volume (rCBV) maps, which were used to correct for PVEs in QSM data from 30 stroke patients. The corrected QSM images showed a significant increase in venous susceptibility (Δχ) values compared to the uncorrected images (441.46 ± 61.14 ppb vs. 163.26 ± 19.66 ppb; p < 0.001), translating to a physiologically plausible mean Yv of 70.42 ± 4.09%. The method also improved the distinction of asymmetrically prominent cortical veins (APCVs), revealing lower Yv values in these areas for some cases, consistent with reduced oxygenation. Our findings demonstrate that using contrast-enhanced rCBV maps can correct PVEs in QSM, providing more reliable measurements of Yv and OEF in small cerebral veins. This approach offers valuable clinical insight into assessing cerebral hemodynamics in patients with stroke and other neurological conditions.

YNICL Journal 2026 Journal Article

Remote cortical degeneration related to structural connectivity following recent small subcortical infarcts

  • Youjie Wang
  • Jingyu Cui
  • Yuying Yan
  • Tang Yang
  • Yue Yuan
  • Rumei Lei
  • Rongfeng Luo
  • Bo Wu

BACKGROUND: Secondary cortical degeneration caused by the remote effects of subcortical infarcts has been implicated in long-term outcomes after acute ischemic stroke. However, this process remains insufficiently studied in recent small subcortical infarcts (RSSI). We aimed to verify RSSI-induced cortical damage, determine whether it can be captured by neuroimaging markers, and explore its association with clinical outcomes. METHODS: RSSI patients with longitudinal Magnetic Resonance Imaging (MRI) were included. Cortical degeneration was assessed using linear mixed-effects models, incorporating a direct approach based on individual diffusion weighted imaging and an indirect approach using the normative connectome from the Human Connectome Project (HCP). Principal component analysis (PCA) was employed to extract features of cortical alterations. The resulting component scores were used in general linear models to assess associations with neuroimaging markers and clinical outcomes. RESULTS: A total of 76 RSSI patients were analyzed. RSSI was found to induce progressive cortical thinning and volume loss in structurally connected regions. PCA identified a component reflecting parenchymal atrophy associated with diffusion-based markers of white matter integrity, as well as the presence of track/cap signs. Moreover, faster cortical degeneration in lesion-connected regions was significantly associated with a greater increase in Hamilton Anxiety Rating Scale (HAMA) scores (β = -2.38, 95% CI = -4.30 - -0.47, p = 0.017). CONCLUSIONS: RSSI induces secondary cortical damage through structurally connected fiber tracts, which is detectable by neuroimaging markers of white matter integrity. These regional cortical alterations may be relevant to post-stroke outcomes and require validation in larger longitudinal studies.

JBHI Journal 2025 Journal Article

A Dual-Branch Cross-Modality-Attention Network for Thyroid Nodule Diagnosis Based on Ultrasound Images and Contrast-Enhanced Ultrasound Videos

  • Jianning Chi
  • Jia-hui Chen
  • Bo Wu
  • Jin Zhao
  • Kai Wang
  • Xiaosheng Yu
  • Wenjun Zhang
  • Ying Huang

Contrast-enhanced ultrasound (CEUS) has been extensively employed as an imaging modality in thyroid nodule diagnosis due to its capacity to visualise the distribution and circulation of micro-vessels in organs and lesions in a non-invasive manner. However, current CEUS-based thyroid nodule diagnosis methods suffered from: 1) the blurred spatial boundaries between nodules and other anatomies in CEUS videos, and 2) the insufficient representations of the local structural information of nodule tissues by the features extracted only from CEUS videos. In this paper, we propose a novel dual-branch network with a cross-modality-attention mechanism for thyroid nodule diagnosis by integrating the information from tow related modalities, i. e. , CEUS videos and ultrasound image. The mechanism has two parts: US-attention-from-CEUS transformer (UAC-T) and CEUS-attention-from-US transformer (CAU-T). As such, this network imitates the manner of human radiologists by decomposing the diagnosis into two correlated tasks: 1) the spatio-temporal features extracted from CEUS are hierarchically embedded into the spatial features extracted from US with UAC-T for the nodule segmentation; 2) the US spatial features are used to guide the extraction of the CEUS spatio-temporal features with CAU-T for the nodule classification. The two tasks are intertwined in the dual-branch end-to-end network and optimized with the multi-task learning (MTL) strategy. The proposed method is evaluated on our collected thyroid US-CEUS dataset. Experimental results show that our method achieves the classification accuracy of 86. 92%, specificity of 66. 41%, and sensitivity of 97. 01%, outperforming the state-of-the-art methods. As a general contribution in the field of multi-modality diagnosis of diseases, the proposed method has provided an effective way to combine static information with its related dynamic information, improving the quality of deep learning based diagnosis with an additional benefit of explainability.

EAAI Journal 2025 Journal Article

A hybrid Bayesian model updating and non-dominated sorting genetic algorithm framework for intelligent mix design of steel fiber reinforced concrete

  • Yong Yu
  • Jie Su
  • Bo Wu

Steel fiber reinforced concrete (SFRC) improves the strength and toughness of conventional concrete, but the high cost and carbon footprint of fibers challenge the balance among performance, cost and sustainability. To address this, an intelligent mix design framework is proposed to optimize compressive and splitting tensile strengths, cost and emissions. Based on 671 experimental records, posterior models were built using Markov Chain Monte Carlo sampling and Bayesian model updating, enabling accurate strength predictions. Compared to traditional regression methods, R 2 scores improved by 15. 7 % and 12. 4 %, confirming its predictive advantage. Cost-wise, materials dominate, while emissions mainly arise from production, transport and mixing. A non-dominated sorting genetic algorithm identified optimal designs under given constraints. Results show that reducing water-to-cement and aggregate-to-cement ratios, and increasing sand ratio and fiber reinforcement index, enhances SFRC strength. Larger coarse aggregates reduce compressive strength but have limited effect on tensile strength. Optimization suggests potential cost and emission reductions of up to 60 %. Moreover, for compression-prone components, fiber use is inefficient due to high cost and emissions, whereas for crack-resistant or strength-balanced elements, fiber inclusion offers a more sustainable alternative to merely lowering the water-to-cement ratio. The proposed framework enables tailored SFRC mix designs, guiding the efficient use of steel fibers.

AAAI Conference 2025 Conference Paper

Integrating Large Language Models and Möbius Group Transformations for Temporal Knowledge Graph Embedding on the Riemann Sphere

  • Sensen Zhang
  • Xun Liang
  • Simin Niu
  • Zhendong Niu
  • Bo Wu
  • Gengxin Hua
  • Long Wang
  • Zhenyu Guan

The significance of Temporal Knowledge Graphs (TKGs) in Artificial Intelligence (AI) lies in their capacity to incorporate time-dimensional information, support complex reasoning and prediction, optimize decision-making processes, enhance the accuracy of recommendation systems, promote multimodal data integration, and strengthen knowledge management and updates. This provides a robust foundation for various AI applications. To effectively learn and apply both static and dynamic temporal patterns for reasoning, a range of embedding methods and large language models (LLMs) have been proposed in the literature. However, these methods often rely on a single underlying embedding space, whose geometric properties severely limit their ability to model intricate temporal patterns, such as hierarchical and ring structures. To address this limitation, this paper proposes embedding TKGs into projective geometric space and leverages LLMs technology to extract crucial temporal node information, thereby constructing the 5EL model. By embedding TKGs into projective geometric space and utilizing Möbius Group transformations, we effectively model various temporal patterns. Subsequently, LLMs technology is employed to process the trained TKGs. We adopt a parameter-efficient fine-tuning strategy to align LLMs with specific task requirements, thereby enhancing the model's ability to recognize structural information of key nodes in historical chains and enriching the representation of central entities. Experimental results on five advanced TKG datasets demonstrate that our proposed 5EL model significantly outperforms existing models.

YNIMG Journal 2025 Journal Article

Optimizing the visualization of the locus coeruleus using magnetization transfer contrast 3D imaging

  • Haiying Lyu
  • Naying He
  • Bo Wu
  • Paula Trujillo
  • Fuhua Yan
  • Yong Lu
  • E. Mark Haacke

BACKGROUND: The locus coeruleus (LC) is a key noradrenergic nucleus of the brain. Its dysfunction is implicated in neurodegenerative diseases like Alzheimer's disease and Parkinson's disease, as well as in psychiatric disorders. However, imaging the LC with sufficient contrast-to-noise ratio (CNR) is challenging due to its small size and deep location in the brainstem. This study optimizes a 3D gradient echo (GRE) sequence with magnetization transfer contrast (MTC) to enable rapid, high-resolution LC imaging in under five minutes. METHODS: A high-resolution 3D-GRE-MTC sequence was optimized on a 3T scanner in 11 healthy volunteers (6 young and 5 older adults). Tissue properties were measured using in vivo MRI data, and simulations were performed to identify the optimal flip angle. LC visualization was evaluated by two independent raters using relative contrast ratio (rCR) and CNR. The diameter and the length of the LC were also evaluated. Each volunteer underwent MRI sessions over three days to assess test-retest reliability. The intra-class correlation coefficient (ICC) for inter-rater reliability and the mean ± standard deviation of LC rCR across sessions for test-retest reproducibility were calculated. RESULTS: A total of 98 scans were collected. The optimized protocol achieved 0.67 × 0.73 × 2 mm³ resolution with an 18° flip angle, 6.18 ms first echo, 52 ms repetition time, flow compensation, arterial suppression, and strict head immobilization. The LC exhibited a CNR of 8.27 ± 1.03, and rCR of 16.70% ± 1.77% (left) and 13.97% ± 2.19% (right), with good inter-rater reliability (ICC = 88.51%). Contrast stability between scans had a variability of 4%-11%. The bilateral LC was visible across 3-6 slices (6-12 mm). Using the full width at quarter maximum measure, the LC diameter was 1.94 ± 0.40 mm for the left side and 1.67 ± 0.34 mm for the right side. CONCLUSION: The optimized protocol enabled reliable, high-resolution LC imaging in under five minutes, providing a valuable tool for clinical and research applications.

EAAI Journal 2024 Journal Article

Dynamic balanced teacher: A semi-supervised object detection algorithm for train faults

  • Guodong Sun
  • Xingyu Pan
  • Qihang Liang
  • Bo Wu

Efficient visual fault detection in freight trains is crucial for guaranteeing the safety of rail transportation. At present, deep learning-based methods for diagnosing faults in freight trains necessitate substantial amounts of labeled data for training. However, the procurement of this labeled data is both costly and time-consuming, posing significant challenges for the practical implementation of these methods. In this study, we propose a semi-supervised fault detection algorithm that demands fewer labeled data, while maintaining detection accuracy. Initially, we employ a lightweight vision transformer for feature extraction to minimize computational requirements in the detector. Subsequently, we implement Dynamic Label Assignment to replace the static-based Intersection over Union strategy, thereby enhancing the robustness of the teacher–student network against noise pseudo-frames. Lastly, we establish a balanced learning framework that effectively mitigates sample-level and target-level imbalances during training. We conduct extensive experiments on three faulty datasets, and the results demonstrate that our method surpasses alternative approaches at various labeling ratios(i. e. 1%, 2%, 5%, 10%).

EAAI Journal 2024 Journal Article

Efficient segmentation with texture in ore images based on box-supervised approach

  • Guodong Sun
  • Delong Huang
  • Yuting Peng
  • Le Cheng
  • Bo Wu
  • Yang Zhang

Image segmentation methods have been utilized to determine the particle size distribution of crushed ores. Due to the complex working environment, high-powered computing equipment is difficult to deploy. At the same time, the ore distribution is stacked, and it is difficult to identify the complete features. To address this issue, an effective box-supervised technique with texture features is provided for ore image segmentation that can identify complete and independent ores. Firstly, a ghost feature pyramid network (Ghost-FPN) is proposed to process the features obtained from the backbone to reduce redundant semantic information and computation generated by complex networks. Then, an optimized detection head is proposed to obtain the feature to maintain accuracy. Finally, Lab color space (Lab) and local binary patterns (LBP) texture features are combined to form a fusion feature similarity-based loss function to improve accuracy while incurring no loss. Experiments on MS COCO have shown that the proposed fusion features are also worth studying on other types of datasets. Extensive experimental results demonstrate the effectiveness of the proposed method, which achieves over 50 frames per second with a small model size of 21. 6 MB. Meanwhile, the method maintains a high level of accuracy 67. 8 in A P 50 b o x and 47. 7 in A P 50 m a s k compared with the state-of-the-art approaches on ore image dataset, even better than bounding box tightness prior (BBTP) by 10. 4/1. 3 on A P 50 b o x / A P 50 m a s k metrics with the ResNet50 as backbone. The source code is available at https: //github. com/MVME-HBUT/OREINST.

EAAI Journal 2024 Journal Article

FS-OreDet: Feature enhancement and relationship exploration for boosting few-shot object detector of ore images

  • Guodong Sun
  • Le Cheng
  • Jinyu Liu
  • Yuting Peng
  • Chengming Xu
  • Yanwei Fu
  • Bo Wu
  • Yang Zhang

In the ore beneficiation process, large block detection is necessary to ensure production safety. This typically involves identifying oversized ore on the conveyor belt and preventing material blockage accidents in the transfer buffer bin between the ore feeding belt and the ore receiving belt. Methods based on deep learning can learn to construct complex features from a large amount of data, but they also require a large number of hand-made datasets for training. Although the existing few shot detection methods for ore images reduce the cost of manual labeling, the corresponding detection performance is insufficient. This article mainly explores how to improve the performance of the detector under the ore image detection task in the case of few labeled images. First, a shot enhancement block is proposed to enhance the valuable foreground information for higher-quality support features. Subsequently, we present a dual-attention region proposal network that effectively leverages support features to enhance the precision of generating candidate proposals. Finally, we propose a lightweight multi-relational detector to effectively evaluate the relationship between query and support proposals, leading to a substantial enhancement in guidance performance. The proposed few-shot object detector (FS-OreDet) achieves the best detection results with state-of-the-art methods with an average precision ( A P ) of 55. 1, a speed of 57 frames per second ( F P S ), and a model size of only 17 M B. Furthermore, our framework adeptly captures the feature information of ore images with substantial data. The detector’s accuracy achieves a significant improvement of 14% in A P. Compared with general object detectors, the performance of the detector ranks first and meets the requirements for outdoor scene deployment.

EAAI Journal 2024 Journal Article

Graph relationship-driven label coded mapping and compensation for multi-label textile fiber recognition

  • Daxing Fu
  • Hao Zhong
  • Xin Zhang
  • Quan Zhou
  • Chenhui Wan
  • Bo Wu
  • Youmin Hu

Recently, neural network-based methods for multi-label textile fiber recognition have achieved considerable success. However, a significant limitation of most current approaches is their disregard for the valuable dependencies that exist among different fiber categories. The universal multi-label image recognition methods that consider label relationships often fall short when applied to the challenge posed by the mixture of fibers in textile images. And these relationship modeling manners have not yet been used in the field of multi-label textile fiber recognition. In this work, a graph relationship-driven method is proposed for the recognition of multi-label textile fibers. Based on the graph attention network, the proposed method introduces global relation graph, label coded mapping, and label compensation to equip the feature learning backbone with the relationship modeling ability. First, feature learning backbone extracts the semantic features from images, and global relation graph mines shallow representations of global label relationships from label embeddings. Then, label coded mapping combines these semantic features and global features to model joint relationships. The obtained joint representations constitute the label code, which is used to map the fiber class indices. Finally, label compensation further extracts the deep representations of global label relationships and utilizes them to compensate the fiber class indices to obtain final prediction scores. By employing experiments on a textile fiber dataset and a rearranged dataset, the effectiveness and superiority of the proposed method are validated. Further experiments on PASCAL VOC 2007 showcase the potential applications of the proposed method beyond textile fiber recognition.

IROS Conference 2024 Conference Paper

Uncertainty-Aware Deployment of Pre-trained Language-Conditioned Imitation Learning Policies

  • Bo Wu
  • Bruce D. Lee
  • Kostas Daniilidis
  • Bernadette Bucher
  • Nikolai Matni

Large-scale robotic policies trained on data from diverse tasks and robotic platforms hold great promise for enabling general-purpose robots; however, reliable generalization to new environment conditions remains a major challenge. Toward addressing this challenge, we propose a novel approach for uncertainty-aware deployment of pre-trained language-conditioned imitation learning agents. Specifically, we use temperature scaling to calibrate these models and exploit the calibrated model to make uncertainty-aware decisions by aggregating the local information of candidate actions. We implement our approach in simulation using three such pre-trained models, and showcase its potential to significantly enhance task completion rates. The accompanying code is accessible at the link: https://github.com/BobWu1998/uncertainty_quant_all.git

AAMAS Conference 2023 Conference Paper

Benchmarking Robustness and Generalization in Multi-Agent Systems: A Case Study on Neural MMO

  • Yangkun Chen
  • Joseph Suarez
  • Junjie Zhang
  • Chenghui Yu
  • Bo Wu
  • Hanmo Chen
  • Hengman Zhu
  • Rui Du

We present the results of the second Neural MMO challenge, hosted at IJCAI 2022, which received 1600+ submissions. This competition targets robustness and generalization in multi-agent systems: participants train teams of agents to complete a multi-task objective against opponents not seen during training. We summarize the competition design and results and suggest that, considering our work as a case study, competitions are an effective approach to solving hard problems and establishing a solid benchmark for algorithms. We will open-source our benchmark including the environment wrapper, baselines, a visualization tool, and selected policies for further research.

AAAI Conference 2023 Short Paper

Enhancing Dynamic GCN for Node Attribute Forecasting with Meta Spatial-Temporal Learning (Student Abstract)

  • Bo Wu
  • Xun Liang
  • Xiangping Zheng
  • Jun Wang

Node attribute forecasting has recently attracted considerable attention. Recent attempts have thus far utilize dynamic graph convolutional network (GCN) to predict future node attributes. However, few prior works have notice that the complex spatial and temporal interaction between nodes, which will hamper the performance of dynamic GCN. In this paper, we propose a new dynamic GCN model named meta-DGCN, leveraging meta spatial-temporal tasks to enhance the ability of dynamic GCN for better capturing node attributes in the future. Experiments show that meta-DGCN effectively modeling comprehensive spatio-temporal correlations between nodes and outperforms state-of-the-art baselines on various real-world datasets.

AAAI Conference 2023 Short Paper

Exploiting High-Order Interaction Relations to Explore User Intent (Student Abstract)

  • Xiangping Zheng
  • Xun Liang
  • Bo Wu

This paper studies the problem of exploring the user intent for session-based recommendations. Its challenges come from the uncertainty of user behavior and limited information. However, current endeavors cannot fully explore the mutual interactions among sessions and do not explicitly model the complex high-order relations among items. To circumvent these critical issues, we innovatively propose a HyperGraph Convolutional Contrastive framework (termed HGCC) that consists of two crucial tasks: 1) The session-based recommendation (SBR task) that aims to capture the beyond pair-wise relationships between items and sessions. 2) The self-supervised learning (SSL task) acted as the auxiliary task to boost the former task. By jointly optimizing the two tasks, the performance of the recommendation task achieves decent gains. Experiments on multiple real-world datasets demonstrate the superiority of the proposed approach over the state-of-the-art methods.

EAAI Journal 2023 Journal Article

Graph features dynamic fusion learning driven by multi-head attention for large rotating machinery fault diagnosis with multi-sensor data

  • Xin Zhang
  • Xi Zhang
  • Jie Liu
  • Bo Wu
  • Youmin Hu

Recently, rotating machinery fault diagnosis studies based on graph neural networks (GNN) have received some satisfactory achievements. But most of them are based on the analysis of the single sensor signals, which cannot capture the comprehensive fault information, especially aiming at large rotating machineries. A few research using GNN for multi-sensor fault diagnosis only fuse multi-source features in the construction of the input graph, and the fusion effect largely depends on the manual feature selection. Graph attention network (GAT), as an emerging GNN, can give trainable weights to vertices based on the self-attention mechanism to improve the effectiveness of feature learning. And it has not yet been used in the field of multi-sensor fault diagnosis. To fill this gap and utilize GAT’s advantages, this paper presents a multi-sensor multi-head GAT (MMHGAT) model for large rotating machinery fault diagnosis. With the input of several subgraphs, the designed MMHGAT model consisting of two graph attention layers (GAL), a feature fusion process and a Softmax classifier, can dynamically fuse and mine the high-level fault characteristics during the training process. By employing the experiment on the axial flow pump, the effectiveness and superiority of the proposed method are validated.

AAAI Conference 2023 Conference Paper

Personalized Dialogue Generation with Persona-Adaptive Attention

  • Qiushi Huang
  • Yu Zhang
  • Tom Ko
  • Xubo Liu
  • Bo Wu
  • Wenwu Wang
  • H Tang

Persona-based dialogue systems aim to generate consistent responses based on historical context and predefined persona. Unlike conventional dialogue generation, the persona-based dialogue needs to consider both dialogue context and persona, posing a challenge for coherent training. Specifically, this requires a delicate weight balance between context and persona. To achieve that, in this paper, we propose an effective framework with Persona-Adaptive Attention (PAA), which adaptively integrates the weights from the persona and context information via our designed attention. In addition, a dynamic masking mechanism is applied to the PAA to not only drop redundant information in context and persona but also serve as a regularization mechanism to avoid overfitting. Experimental results demonstrate the superiority of the proposed PAA framework compared to the strong baselines in both automatic and human evaluation. Moreover, the proposed PAA approach can perform equivalently well in a low-resource regime compared to models trained in a full-data setting, which achieve a similar result with only 20% to 30% of data compared to the larger models trained in the full-data setting. To fully exploit the effectiveness of our design, we designed several variants for handling the weighted information in different ways, showing the necessity and sufficiency of our weighting and masking designs.

AAAI Conference 2022 Short Paper

Capsule Graph Neural Network for Multi-Label Image Recognition (Student Abstract)

  • Xiangping Zheng
  • Xun Liang
  • Bo Wu

This paper studies the problem of learning complex relationships between multi-labels for image recognition. Its challenges come from the rich and diverse semantic information in images. However, current methods cannot fully explore the mutual interactions among labels and do not explicitly model the label co-occurrence. To overcome these shortcomings, we innovatively propose CGML that consists of two crucial modules: 1) an image representation learning module that aims to complete the feature extraction of an image whose features are expressed in the form of primary capsules; 2) a label adaptive graph convolutional network module that leverages the popular graph convolutional networks with an adaptive label correlation graph to model label dependencies. Experiments show that our approach obviously outperforms the existing state-of-the-art methods.

AAAI Conference 2021 Conference Paper

Advice-Guided Reinforcement Learning in a non-Markovian Environment

  • Daniel Neider
  • Jean-Raphael Gaglione
  • Ivan Gavran
  • Ufuk Topcu
  • Bo Wu
  • Zhe Xu

We study a class of reinforcement learning tasks in which the agent receives its reward for complex, temporally-extended behaviors sparsely. For such tasks, the problem is how to augment the state-space so as to make the reward function Markovian in an efficient way. While some existing solutions assume that the reward function is explicitly provided to the learning algorithm (e. g. , in the form of a reward machine), the others learn the reward function from the interactions with the environment, assuming no prior knowledge provided by the user. In this paper, we generalize both approaches and enable the user to give advice to the agent, representing the user’s best knowledge about the reward function, potentially fragmented, partial, or even incorrect. We formalize advice as a set of DFAs and present a reinforcement learning algorithm that takes advantage of such advice, with optimal convergence guarantee. The experiments show that using wellchosen advice can reduce the number of training steps needed for convergence to optimal policy, and can decrease the computation time to learn the reward function by up to two orders of magnitude.

AAMAS Conference 2021 Conference Paper

Intrinsic Motivated Multi-Agent Communication

  • Chuxiong Sun
  • Bo Wu
  • Rui Wang
  • Xiaohui Hu
  • Xiaoya Yang
  • Cong Cong

Efficient communication is a promising way to achieve cooperation among agents in many real-world scenarios. However, aimless and motiveless information sharing may not work or even degrade the cooperative performance. Typically, the multi-agent communication behaviors are motivated by extrinsic rewards from environment. We conclude the mechanism as ’Communicate what rewards you’. In this work, we present a novel communication mechanism called Intrinsic Motivated Multi-Agent Communication (IMMAC). Our key insight can be summarized as ’Communicate what surprises you’. Concretely, we use an observation-dependent intrinsic value to represent the importance of observed information. Then a gating mechanism and an attentional mechanism based on intrinsic values are designed to control communication. By encouraging agent to communicate and focus on the observations with uncertain and important information, our algorithm achieves superior communication efficiency and cooperative performance. We evaluate IMMAC on a variety of challenging tasks, and demonstrate that intrinsic values are sufficient to drive efficient communication behaviors. Moreover, we found that the combination of intrinsic values and extrinsic values can further improve the communication efficiency. Consequently, intrinsic motivation is a promising way to control communication and it is capable of being a good complement to the existing extrinsic motivated communication methods.

NeurIPS Conference 2021 Conference Paper

NxMTransformer: Semi-Structured Sparsification for Natural Language Understanding via ADMM

  • Connor Holmes
  • Minjia Zhang
  • Yuxiong He
  • Bo Wu

Natural Language Processing (NLP) has recently achieved great success by using huge pre-trained Transformer networks. However, these models often contain hundreds of millions or even billions of parameters, bringing challenges to online deployment due to latency constraints. Recently, hardware manufacturers have introduced dedicated hardware for NxM sparsity to provide the flexibility of unstructured pruning with the runtime efficiency of structured approaches. NxM sparsity permits arbitrarily selecting M parameters to retain from a contiguous group of N in the dense representation. However, due to the extremely high complexity of pre-trained models, the standard sparse fine-tuning techniques often fail to generalize well on downstream tasks, which have limited data resources. To address such an issue in a principled manner, we introduce a new learning framework, called NxMTransformer, to induce NxM semi-structured sparsity on pretrained language models for natural language understanding to obtain better performance. In particular, we propose to formulate the NxM sparsity as a constrained optimization problem and use Alternating Direction Method of Multipliers (ADMM) to optimize the downstream tasks while taking the underlying hardware constraints into consideration. ADMM decomposes the NxM sparsification problem into two sub-problems that can be solved sequentially, generating sparsified Transformer networks that achieve high accuracy while being able to effectively execute on newly released hardware. We apply our approach to a wide range of NLP tasks, and our proposed method is able to achieve 1. 7 points higher accuracy in GLUE score than current best practices. Moreover, we perform detailed analysis on our approach and shed light on how ADMM affects fine-tuning accuracy for downstream tasks. Finally, we illustrate how NxMTransformer achieves additional performance improvement with knowledge distillation based methods.

AAMAS Conference 2021 Conference Paper

Reward Machines for Cooperative Multi-Agent Reinforcement Learning

  • Cyrus Neary
  • Zhe Xu
  • Bo Wu
  • Ufuk Topcu

In cooperative multi-agent reinforcement learning, a collection of agents learns to interact in a shared environment to achieve a common goal. We propose the use of reward machines (RM) — Mealy machines used as structured representations of reward functions — to encode the team’s task. The proposed novel interpretation of RMs in the multi-agent setting explicitly encodes required teammate interdependencies, allowing the team-level task to be decomposed into sub-tasks for individual agents. We define such a notion of RM decomposition and present algorithmically verifiable conditions guaranteeing that distributed completion of the sub-tasks leads to team behavior accomplishing the original task. This framework for task decomposition provides a natural approach to decentralized learning: agents may learn to accomplish their sub-tasks while observing only their local state and abstracted representations of their teammates. We accordingly propose a decentralized q-learning algorithm. Furthermore, in the case of undiscounted rewards, we use local value functions to derive lower and upper bounds for the global value function corresponding to the team task. Experimental results in three discrete settings exemplify the effectiveness of the proposed RM decomposition approach, which converges to a successful team policy an order of magnitude faster than a centralized learner and significantly outperforms hierarchical and independent q-learning approaches.

NeurIPS Conference 2021 Conference Paper

STAR: A Benchmark for Situated Reasoning in Real-World Videos

  • Bo Wu
  • Shoubin Yu
  • Zhenfang Chen
  • Josh Tenenbaum
  • Chuang Gan

Reasoning in the real world is not divorced from situations. How to capture the present knowledge from surrounding situations and perform reasoning accordingly is crucial and challenging for machine intelligence. This paper introduces a new benchmark that evaluates the situated reasoning ability via situation abstraction and logic-grounded question answering for real-world videos, called Situated Reasoning in Real-World Videos (STAR). This benchmark is built upon the real-world videos associated with human actions or interactions, which are naturally dynamic, compositional, and logical. The dataset includes four types of questions, including interaction, sequence, prediction, and feasibility. We represent the situations in real-world videos by hyper-graphs connecting extracted atomic entities and relations (e. g. , actions, persons, objects, and relationships). Besides visual perception, situated reasoning also requires structured situation comprehension and logical reasoning. Questions and answers are procedurally generated. The answering logic of each question is represented by a functional program based on a situation hyper-graph. We compare various existing video reasoning models and find that they all struggle on this challenging situated reasoning task. We further propose a diagnostic neuro-symbolic model that can disentangle visual perception, situation abstraction, language understanding, and functional reasoning to understand the challenges of this benchmark.

AAAI Conference 2021 Conference Paper

Temporal-Logic-Based Reward Shaping for Continuing Reinforcement Learning Tasks

  • Yuqian Jiang
  • Suda Bharadwaj
  • Bo Wu
  • Rishi Shah
  • Ufuk Topcu
  • Peter Stone

In continuing tasks, average-reward reinforcement learning may be a more appropriate problem formulation than the more common discounted reward formulation. As usual, learning an optimal policy in this setting typically requires a large amount of training experiences. Reward shaping is a common approach for incorporating domain knowledge into reinforcement learning in order to speed up convergence to an optimal policy. However, to the best of our knowledge, the theoretical properties of reward shaping have thus far only been established in the discounted setting. This paper presents the first reward shaping framework for averagereward learning and proves that, under standard assumptions, the optimal policy under the original reward function can be recovered. In order to avoid the need for manual construction of the shaping function, we introduce a method for utilizing domain knowledge expressed as a temporal logic formula. The formula is automatically translated to a shaping function that provides additional reward throughout the learning process. We evaluate the proposed method on three continuing tasks. In all cases, shaping speeds up the average-reward learning rate without any reduction in the performance of the learned policy compared to relevant baselines.

AAAI Conference 2020 Conference Paper

General Partial Label Learning via Dual Bipartite Graph Autoencoder

  • Brian Chen
  • Bo Wu
  • Alireza Zareian
  • Hanwang Zhang
  • Shih-Fu Chang

We formulate a practical yet challenging problem: General Partial Label Learning (GPLL). Compared to the traditional Partial Label Learning (PLL) problem, GPLL relaxes the supervision assumption from instance-level — a label set partially labels an instance — to group-level: 1) a label set partially labels a group of instances, where the within-group instance-label link annotations are missing, and 2) crossgroup links are allowed — instances in a group may be partially linked to the label set from another group. Such ambiguous group-level supervision is more practical in real-world scenarios as additional annotation on the instance-level is no longer required, e. g. , face-naming in videos where the group consists of faces in a frame, labeled by a name set in the corresponding caption. In this paper, we propose a novel graph convolutional network (GCN) called Dual Bipartite Graph Autoencoder (DB-GAE) to tackle the label ambiguity challenge of GPLL. First, we exploit the cross-group correlations to represent the instance groups as dual bipartite graphs: within-group and cross-group, which reciprocally complements each other to resolve the linking ambiguities. Second, we design a GCN autoencoder to encode and decode them, where the decodings are considered as the refined results. It is worth noting that DB-GAE is self-supervised and transductive, as it only uses the group-level supervision without a separate offline training stage. Extensive experiments on two real-world datasets demonstrate that DB-GAE significantly outperforms the best baseline over absolute 0. 159 F1-score and 24. 8% accuracy. We further offer analysis on various levels of label ambiguities.

IJCAI Conference 2020 Conference Paper

Learning the Compositional Visual Coherence for Complementary Recommendations

  • Zhi Li
  • Bo Wu
  • Qi Liu
  • Likang Wu
  • Hongke Zhao
  • Tao Mei

Complementary recommendations, which aim at providing users product suggestions that are supplementary and compatible with their obtained items, have become a hot topic in both academia and industry in recent years. Existing work mainly focused on modeling the co-purchased relations between two items, but the compositional associations of item collections are largely unexplored. Actually, when a user chooses the complementary items for the purchased products, it is intuitive that she will consider the visual semantic coherence (such as color collocations, texture compatibilities) in addition to global impressions. Towards this end, in this paper, we propose a novel Content Attentive Neural Network (CANN) to model the comprehensive compositional coherence on both global contents and semantic contents. Specifically, we first propose a Global Coherence Learning (GCL) module based on multi-heads attention to model the global compositional coherence. Then, we generate the semantic-focal representations from different semantic regions and design a Focal Coherence Learning (FCL) module to learn the focal compositional coherence from different semantic-focal representations. Finally, we optimize the CANN in a novel compositional optimization strategy. Extensive experiments on the large-scale real-world data clearly demonstrate the effectiveness of CANN compared with several state-of-the-art methods.

AAAI Conference 2020 Conference Paper

Tensor FISTA-Net for Real-Time Snapshot Compressive Imaging

  • Xiaochen Han
  • Bo Wu
  • Zheng Shou
  • Xiao-Yang Liu
  • Yimeng Zhang
  • Linghe Kong

Snapshot compressive imaging (SCI) cameras capture highspeed videos by compressing multiple video frames into a measurement frame. However, reconstructing video frames from the compressed measurement frame is challenging. The existing state-of-the-art reconstruction algorithms suffer from low reconstruction quality or heavy time consumption, making them not suitable for real-time applications. In this paper, exploiting the powerful learning ability of deep neural networks (DNN), we propose a novel Tensor Fast Iterative Shrinkage-Thresholding Algorithm Net (Tensor FISTA-Net) as a decoder for SCI video cameras. Tensor FISTA-Net not only learns the sparsest representation of the video frames through convolution layers, but also reduces the reconstruction time significantly through tensor calculations. Experimental results on synthetic datasets show that the proposed Tensor FISTA-Net achieves average PSNR improvement of 1. 63∼3. 89dB over the state-of-the-art algorithms. Moreover, Tensor FISTA-Net takes less than 2 seconds running time and 12MB memory footprint, making it practical for real-time IoT applications.

IJCAI Conference 2017 Conference Paper

Sequential Prediction of Social Media Popularity with Deep Temporal Context Networks

  • Bo Wu
  • Wen-Huang Cheng
  • Yongdong Zhang
  • Qiushi Huang
  • Jintao Li
  • Tao Mei

Prediction of popularity has profound impact for social media, since it offers opportunities to reveal individual preference and public attention from evolutionary social systems. Previous research, although achieves promising results, neglects one distinctive characteristic of social data, i. e. , sequentiality. For example, the popularity of online content is generated over time with sequential post streams of social media. To investigate the sequential prediction of popularity, we propose a novel prediction framework called Deep Temporal Context Networks (DTCN) by incorporating both temporal context and temporal attention into account. Our DTCN contains three main components, from embedding, learning to predicting. With a joint embedding network, we obtain a unified deep representation of multi-modal user-post data in a common embedding space. Then, based on the embedded data sequence over time, temporal context learning attempts to recurrently learn two adaptive temporal contexts for sequential popularity. Finally, a novel temporal attention is designed to predict new popularity (the popularity of a new user-post pair) with temporal coherence across multiple time-scales. Experiments on our released image dataset with about 600K Flickr photos demonstrate that DTCN outperforms state-of-the-art deep prediction algorithms, with an average of 21. 51% relative performance improvement in the popularity prediction (Spearman Ranking Correlation).

AAAI Conference 2016 Conference Paper

Unfolding Temporal Dynamics: Predicting Social Media Popularity Using Multi-scale Temporal Decomposition

  • Bo Wu
  • Tao Mei
  • Wen-Huang Cheng
  • Yongdong Zhang

Time information plays a crucial role on social media popularity. Existing research on popularity prediction, effective though, ignores temporal information which is highly related to user-item associations and thus often results in limited success. An essential way is to consider all these factors (user, item, and time), which capture the dynamic nature of photo popularity. In this paper, we present a novel approach to factorize the popularity into user-item context and time-sensitive context for exploring the mechanism of dynamic popularity. The user-item context provides a holistic view of popularity, while the time-sensitive context captures the temporal dynamics nature of popularity. Accordingly, we develop two kinds of time-sensitive features, including user activeness variability and photo prevalence variability. To predict photo popularity, we propose a novel framework named Multi-scale Temporal Decomposition (MTD), which decomposes the popularity matrix in latent spaces based on contextual associations. Specifically, the proposed MTD models time-sensitive context on different time scales, which is beneficial to automatically learn temporal patterns. Based on the experiments conducted on a real-world dataset with 1. 29M photos from Flickr, our proposed MTD can achieve the prediction accuracy of 79. 8% and outperform the best three state-of-the-art methods with a relative improvement of 9. 6% on average.

IJCAI Conference 2015 Conference Paper

An Iterative Approach to Synthesize Data Transformation Programs

  • Bo Wu
  • Craig A. Knoblock

Programming-by-Example approaches allow users to transform data by simply entering the target data. However, current methods do not scale well to complicated examples, where there are many examples or the examples are long. In this paper, we present an approach that exploits the fact that users iteratively provide examples. It reuses the previous subprograms to improve the efficiency in generating new programs. We evaluated the approach with a variety of transformation scenarios. The results show that the approach significantly reduces the time used to generate the transformation programs, especially in complicated scenarios.

v2026.09.13