Arrow Research search

Author name cluster

Hu Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

12 papers
2 author rows

Possible papers

12

TMLR Journal 2026 Journal Article

Deep Multimodal Learning with Missing Modality: A Survey

  • Renjie Wu
  • Hu Wang
  • Hsiang-Ting Chen
  • Gustavo Carneiro

During multimodal model training and testing, certain data modalities may be absent due to sensor limitations, cost constraints, privacy concerns, or data loss, which can degrade performance. Multimodal learning techniques that explicitly account for missing modalities aim to improve robustness by enabling models to perform reliably even when certain inputs are unavailable. This survey presents the first comprehensive review of Multimodal Learning with Missing Modality (MLMM), with a focus on deep learning approaches. We outline the motivations and key distinctions between MLMM and conventional multimodal learning, provide a detailed analysis of existing methods, applications, and datasets, and conclude by highlighting open challenges and future research directions.

EAAI Journal 2026 Journal Article

Discrete physics-informed neural network with enforced interface constraint for domain decomposition

  • Jichao Yin
  • Mingxuan Li
  • Jianguang Fang
  • Chi Wu
  • Hu Wang
  • Guangyao Li

While domain decomposition method (DDM) constitutes an effective strategy for improving the training efficiency of physics-informed neural network (PINN), the approach simultaneously introduces an increased risk of training instability owing to the additional loss terms introduced. To address this issue, the work proposes an energy-based discrete PINN (dPINN) approach incorporating a proposed enforced interface constraint (EIC) mechanism within the context of the DDM. The dPINN builds upon the DDM with the EIC mechanism and will henceforth be referred to as EIC-DDM-dPINN. Within this framework, the dPINN computes the system energy in an element-wise fashion using Gaussian integration, guided by finite element-inspired formulations. Meanwhile, displacement continuity across subdomain interfaces is explicitly enforced through the EIC mechanism. This enforcement obviates the need to incorporate supplementary loss terms into the loss function, thereby substantially mitigating the risk of training instability. The integration of the EIC-based DDM facilitates simpler and more flexible subdomain mesh partitioning within the EIC-DDM-dPINN framework, thereby reducing the strong dependence on sampling strategies typically required in conventional DDM-based PINN. Beyond improving computational efficiency via parallelization, the DDM also helps decouple the weak spatial constraint (WSC) effect, which can otherwise result in spurious displacement continuity across geometrically discontinuous gaps. Comprehensive numerical experiments in both two- and three-dimensional settings are conducted to assess the accuracy and efficiency of the proposed approach, and the results demonstrate its scalability and robustness, highlighting its potential for application to large-scale problems with complex geometries.

EAAI Journal 2026 Journal Article

Multi-exposure high dynamic range reconstruction by incorporating imaging knowledge

  • Hu Wang
  • Mao Ye
  • Dengyan Luo
  • Yan Gan

The existing photographic equipment is not able to capture scenes of the natural world very well. Thus, the problem of reconstructing high dynamic range (HDR) images from multi-exposure low dynamic range (LDR) images arises because these images have different details. The existing methods do not fully leverage imaging knowledge in the LDR image generation pipeline, resulting in design redundancy and inefficient resource utilization. We propose a new Multi-Exposure HDR reconstruction by incorporating Imaging Knowledge (MEIK) for efficient HDR image reconstruction. Our method consists of two parts: fusion of LDR features and reconstruction of HDR feature. Due to object motion and exposure time effects, LDR features with different exposures need to be fused. A Multi-Exposure Information Aggregation (MEIA) module is proposed to fuse LDR features based on Mamba. After that, an Inverse imaging Knowledge-Driven (IKD) cluster is employed to reconstruct the HDR feature, which is a cascade of IKD blocks at different scales. The IKD block consists of three parts: HDR information recovery, imaging parameter adjustment, and noise suppression, used to simulate the mathematical formula for multi-exposure HDR imaging. Experimental results demonstrate that the proposed MEIK model outperforms existing state-of-the-art models and exhibits strong scalability.

EAAI Journal 2025 Journal Article

Adaptive extreme learning framework for nitrogen oxide conversion prediction: a multi-scale approach

  • Chengzheng Tong
  • Yanhua Wang
  • Liehao Wei
  • Hu Wang
  • Qingling Liu
  • Caixia Liu
  • Mingfa Yao

Efficient selective catalytic reduction (SCR) systems for nitrogen oxide (NOx) emission control face significant challenges in catalyst design optimization due to data scarcity and complex parameter relationships. We present an Adaptive Extreme Learning Machine (AELM) framework designed to address these engineering challenges. The framework innovatively integrates automated text mining to significantly expand the dataset from literature, utilizes adaptive mechanisms (e. g. , feature fusion, temperature bias) within a multi-output regression structure to model complex parameter interactions and temperature dependencies, enabling accurate catalyst performance prediction even under data-scarce conditions. This artificial intelligence approach demonstrates exceptional engineering applicability, achieving up to a 93. 3 % improvement in prediction accuracy (measured by Fréchet distance for curve similarity on challenging cases) while reducing computational overhead by approximately 87 % compared to conventional machine learning models like Artificial Neural Networks. Through comprehensive validation across multiple zeolite catalysts, our framework successfully captures the complex relationships between key structural parameters (Copper content: 2–6 wt percent, Silicon/Aluminum ratio: 5–12) and conversion efficiency, even for diverse zeolite catalysts and challenging test conditions. This development provides engineers with a powerful tool for rapid catalyst screening and optimization in real-world emission control applications, offering practical solutions for sustainable environmental protection technologies.

ICML Conference 2025 Conference Paper

High Dynamic Range Novel View Synthesis with Single Exposure

  • Kaixuan Zhang
  • Hu Wang
  • Minxian Li
  • Mingwu Ren
  • Mao Ye 0001
  • Xiatian Zhu

High Dynamic Range Novel View Synthesis (HDR-NVS) aims to establish a 3D scene HDR model from Low Dynamic Range (LDR) imagery. Typically, multiple-exposure LDR images are employed to capture a wider range of brightness levels in a scene, as a single LDR image cannot represent both the brightest and darkest regions simultaneously. While effective, this multiple-exposure HDR-NVS approach has significant limitations, including susceptibility to motion artifacts (e. g. , ghosting and blurring), high capture and storage costs. To overcome these challenges, we introduce, for the first time, the single-exposure HDR-NVS problem, where only single exposure LDR images are available during training. We further introduce a novel approach, Mono-HDR-3D, featuring two dedicated modules formulated by the LDR image formation principles, one for converting LDR colors to HDR counterparts, and the other for transforming HDR images to LDR format so that unsupervised learning is enabled in a closed loop. Designed as a meta-algorithm, our approach can be seamlessly integrated with existing NVS models. Extensive experiments show that Mono-HDR-3D significantly outperforms previous methods. Source code is released at https: //github. com/prinasi/Mono-HDR-3D.

AAAI Conference 2024 Conference Paper

Segment beyond View: Handling Partially Missing Modality for Audio-Visual Semantic Segmentation

  • Renjie Wu
  • Hu Wang
  • Feras Dayoub
  • Hsiang-Ting Chen

Augmented Reality (AR) devices, emerging as prominent mobile interaction platforms, face challenges in user safety, particularly concerning oncoming vehicles. While some solutions leverage onboard camera arrays, these cameras often have limited field-of-view (FoV) with front or downward perspectives. Addressing this, we propose a new out-of-view semantic segmentation task and Segment Beyond View (SBV), a novel audio-visual semantic segmentation method. SBV supplements the visual modality, which miss the information beyond FoV, with the auditory information using a teacher-student distillation model (Omni2Ego). The model consists of a vision teacher utilising panoramic information, an auditory teacher with 8-channel audio, and an audio-visual student that takes views with limited FoV and binaural audio as input and produce semantic segmentation for objects outside FoV. SBV outperforms existing models in comparative evaluations and shows a consistent performance across varying FoV ranges and in monaural audio settings.

EAAI Journal 2023 Journal Article

Single-image HDR reconstruction by dual learning the camera imaging process

  • Lei She
  • Mao Ye
  • Shuai Li
  • Yu Zhao
  • Ce Zhu
  • Hu Wang

It is a very challenging problem to reconstruct a high dynamic range (HDR) image from a single exposure image. There exist three problems, i. e. , the many-to-many mapping problem between low dynamic range (LDR) images and HDR images, the image quality problem caused by the change of dynamic range and the problem of unpaired LDR–HDR training images. These problems can be solved to some extent through a dual learning framework simultaneously to learn the forward and reverse of camera imaging processes. This procedure is divided into a primary module, to reconstruct HDR from LDR, and a secondary module to reversely mapping the HDR to LDR. The secondary module guides the learning of primary module by constraining the outputs of the primary module. After that, the attention mechanism is used to solve the problem of unnatural perception caused by the change of dynamic range. In the end, with the advantage of our dual learning framework, unpaired data is further explored to train our model, which enriches the training samples. Compared with the state-of-the-art methods, a large number of quantitative and qualitative experiments confirm that our method can achieve better performance.

TIST Journal 2023 Journal Article

Toward Balancing the Efficiency and Effectiveness in k-Facility Relocation Problem

  • Hu Wang
  • Hui Li
  • Meng Wang
  • Jiangtao Cui

Facility Relocation (FR), which is an effort to reallocate the placement of facilities to adapt to the changes of urban planning, has remarkable impact on many areas. Existing solutions fail to guarantee the result quality on relocating k > 1 facilities. As k -FR problem is NP-complete and is not submodular or non-decreasing, traditional greedy algorithm cannot be directly applied. We propose to transform k -FR into another facility placement problem, which is submodular and non-decreasing. We prove that the optimal solutions of both problems are equivalent. Accordingly, we present the first approximate solution toward the k -FR, FR2FP. Our extensive comparison over both FR2FP and the state-of-the-art solution shows that FR2FP, although it provides approximation guarantee, cannot necessarily given superior results. The comparison motivates us to present an advanced approximate solution, FR2FP-ex. Moreover, based on Lagrangian relaxation, we develop an algorithm that can adjust the approximation ratio. Extensive experiments verified that, FR2FP-ex demonstrates the best result quality, and it is very close to the optimal solution. In addition, we also unveil the scenarios when the state-of-the-art would fail. We further generalize the k -FR problem, considering the budget for relocation and the cost of each facility. We also present corresponding approximate solutions toward the new problem and prove the approximation ratio.

IJCAI Conference 2022 Conference Paper

KUNet: Imaging Knowledge-Inspired Single HDR Image Reconstruction

  • Hu Wang
  • Mao Ye
  • Xiatian Zhu
  • Shuai Li
  • Ce Zhu
  • Xue Li

Recently, with the rise of high dynamic range (HDR) display devices, there is a great demand to transfer traditional low dynamic range (LDR) images into HDR versions. The key to success is how to solve the many-to-many mapping problem. However, the existing approaches either do not consider constraining solution space or just simply imitate the inverse camera imaging pipeline in stages, without directly formulating the HDR image generation process. In this work, we address this problem by integrating LDR-to-HDR imaging knowledge into an UNet architecture, dubbed as Knowledge-inspired UNet (KUNet). The conversion from LDR-to-HDR image is mathematically formulated, and can be conceptually divided into recovering missing details, adjusting imaging parameters and reducing imaging noise. Accordingly, we develop a basic knowledge-inspired block (KIB) including three subnetworks corresponding to the three procedures in this HDR imaging process. The KIB blocks are cascaded in the similar way to the UNet to construct HDR image with rich global information. In addition, we also propose a knowledge inspired jump-connect structure to fit a dynamic range gap between HDR and LDR images. Experimental results demonstrate that the proposed KUNet achieves superior performance compared with the state-of-the-art methods. The code, dataset and appendix materials are available at https: //github. com/wanghu178/KUNet. git.

NeurIPS Conference 2022 Conference Paper

M$^4$I: Multi-modal Models Membership Inference

  • Pingyi Hu
  • Zihan Wang
  • Ruoxi Sun
  • Hu Wang
  • Minhui Xue

With the development of machine learning techniques, the attention of research has been moved from single-modal learning to multi-modal learning, as real-world data exist in the form of different modalities. However, multi-modal models often carry more information than single-modal models and they are usually applied in sensitive scenarios, such as medical report generation or disease identification. Compared with the existing membership inference against machine learning classifiers, we focus on the problem that the input and output of the multi-modal models are in different modalities, such as image captioning. This work studies the privacy leakage of multi-modal models through the lens of membership inference attack, a process of determining whether a data record involves in the model training process or not. To achieve this, we propose Multi-modal Models Membership Inference (M$^4$I) with two attack methods to infer the membership status, named metric-based (MB) M$^4$I and feature-based (FB) M$^4$I, respectively. More specifically, MB M$^4$I adopts similarity metrics while attacking to infer target data membership. FB M$^4$I uses a pre-trained shadow multi-modal feature extractor to achieve the purpose of data inference attack by comparing the similarities from extracted input and output features. Extensive experimental results show that both attack methods can achieve strong performances. Respectively, 72. 5% and 94. 83% of attack success rates on average can be obtained under unrestricted scenarios. Moreover, we evaluate multiple defense mechanisms against our attacks. The source code of M$^4$I attacks is publicly available at https: //github. com/MultimodalMI/Multimodal-membership-inference. git.

AAAI Conference 2021 Conference Paper

Memory-Gated Recurrent Networks

  • Yaquan Zhang
  • Qi Wu
  • Nanbo Peng
  • Min Dai
  • Jing Zhang
  • Hu Wang

The essence of multivariate sequential learning is all about how to extract dependencies in data. These data sets, such as hourly medical records in intensive care units and multifrequency phonetic time series, often time exhibit not only strong serial dependencies in the individual components (the “marginal” memory) but also non-negligible memories in the cross-sectional dependencies (the “joint” memory). Because of the multivariate complexity in the evolution of the joint distribution that underlies the data generating process, we take a data-driven approach and construct a novel recurrent network architecture, termed Memory-Gated Recurrent Networks (mGRN), with gates explicitly regulating two distinct types of memories: the marginal memory and the joint memory. Through a combination of comprehensive simulation studies and empirical experiments on a range of public datasets, we show that our proposed mGRN architecture consistently outperforms state-of-the-art architectures targeting multivariate time series.

IJCAI Conference 2020 Conference Paper

Unsupervised Representation Learning by Predicting Random Distances

  • Hu Wang
  • Guansong Pang
  • Chunhua Shen
  • Congbo Ma

Deep neural networks have gained great success in a broad range of tasks due to its remarkable capability to learn semantically rich features from high-dimensional data. However, they often require large-scale labelled data to successfully learn such features, which significantly hinders their adaption in unsupervised learning tasks, such as anomaly detection and clustering, and limits their applications to critical domains where obtaining massive labelled data is prohibitively expensive. To enable unsupervised learning on those domains, in this work we propose to learn features without using any labelled data by training neural networks to predict data distances in a randomly projected space. Random mapping is a theoretically proven approach to obtain approximately preserved distances. To well predict these distances, the representation learner is optimised to learn genuine class structures that are implicitly embedded in the randomly projected space. Empirical results on 19 real-world datasets show that our learned representations substantially outperform a few state-of-the-art methods for both anomaly detection and clustering tasks. Code is available at: \url{https: //git. io/RDP}

v2026.09.13