Arrow Research search

Author name cluster

Wenbo Hu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

14 papers
1 author row

Possible papers

14

EAAI Journal 2026 Journal Article

Automated crack measurement in slab tracks using deformable instance segmentation and boundary augmentation with unsupervised style transfer

  • Wenbo Hu
  • Zheng Wu
  • Weidong Wang
  • Xianhua Liu
  • Jun Peng

Crack detection and measurement in slab tracks are critical for maintenance decision-making. Pre-trained deep learning segmentation models often struggle with cracking instances due to domain adaptation and data scarcity. This study proposes an instance segmentation framework incorporating dynamic snake convolution (DSConv) modules, combined with an unsupervised style transfer-based boundary augmentation strategy. The DSConv-enhanced architecture prioritizes linear crack features in cluttered backgrounds, while the augmentation introduces controlled perturbations to global pixels and local crack boundaries, generating structurally consistent diversified training samples. The results demonstrate that the deformable DSConv enhanced architecture achieves optimal mean average precision (mAP), improving segmentation performance by nearly 13 % compared to "fine-tuned Segment Anything Model". Its segmentation capability surpasses eight state-of-the-art models, especially for multiple intermittent microcracks. Furthermore, unsupervised style transfer-generated augmented data enhances crack instance segmentation performance by 10 % compared to non-augmented baselines, surpassing conventional methods including horizontal flipping and color jittering. Quantitative crack width distributions from segmentation-quantification analysis provide more comprehensive structural health insights than manual discrete-point measurements, facilitating precise maintenance decisions for railway infrastructure.

AAAI Conference 2026 Conference Paper

Benchmarking Trustworthiness in Multimodal LLMs for Video Understanding

  • Youze Wang
  • Zijun Chen
  • Ruoyu Chen
  • Shishen Gu
  • Wenbo Hu
  • Jiayang Liu
  • Yinpeng Dong
  • Hang Su

Recent advancements in multimodal large language models for video understanding (videoLLMs) have enhanced their capacity to process complex spatiotemporal data. However, challenges such as factual inaccuracies, harmful content, biases, hallucinations, and privacy risks compromise their reliability. This study introduces Trust-videoLLMs, a first comprehensive benchmark evaluating 23 state-of-the-art videoLLMs (5 commercial, 18 open-source) across five critical dimensions: truthfulness, robustness, safety, fairness, and privacy. Comprising 30 tasks with adapted, synthetic, and annotated videos, the framework assesses spatiotemporal risks, temporal consistency and cross-modal impact. Results reveal significant limitations in dynamic scene comprehension, cross-modal perturbation resilience and real-world risk mitigation. While open-source models occasionally outperform, proprietary models generally exhibit superior credibility, though scaling does not consistently improve performance. These findings underscore the need for enhanced training datat diversity and robust multimodal alignment. Trust-videoLLMs provides a publicly available, extensible toolkit for standardized trustworthiness assessments, addressing the critical gap between accuracy-focused benchmarks and demands for robustness, safety, fairness, and privacy.

AAAI Conference 2026 Conference Paper

Deep Hidden Cognition Facilitates Reliable Chain-of-Thought Reasoning

  • Zijun Chen
  • Wenbo Hu
  • Richang Hong

Chain of Thought (CoT) reasoning has demonstrated remarkable deep reasoning capabilities in both large language models (LLMs) and multimodal large language models (MLLMs). However, its reliability is often undermined by the accumulation of errors in intermediate steps. This paper proposes a novel approach to calibrating CoT reasoning accuracy by leveraging the model’s internal cognition of truthfulness. Our findings suggest that the model implicitly tracks the evolving veracity of intermediate steps throughout the dynamic, progressive reasoning process. We train a confidence predictor to quantify the model’s internal cognition of truthfulness at each reasoning step, enabling dynamic selection of the most plausible reasoning path through beam search. Experimental results demonstrate that our method significantly outperforms the state-of-the-art baselines (e.g., Self-Consistency, and PRM Guided Search) across the mathematical, symbolic, and commonsense reasoning tasks, exhibiting superior accuracy and reliability in both unimodal and multimodal settings. This study proposes a novel path toward improving the reliability of CoT reasoning, demonstrating strong potential for wide-ranging applications.

AAAI Conference 2026 Conference Paper

Sparse-Scale Transformer with Bidirectional Awareness for Time Series Forecasting

  • Ying Liu
  • Bo Liu
  • Sheng Huang
  • Gang Luo
  • Wenbo Hu
  • Meng Wang
  • Richang Hong

Time series forecasting (TSF) plays a crucial role in many real-world applications, such as weather prediction and economic planning. While Transformer-based models have shown strong capabilities in modeling long-range dependencies, effectively capturing the multi-scale temporal dynamics inherent in time series remains a major challenge. Existing methods often adopt time-windows of varying sizes, which may introduce noisy or irrelevant representations when mismatched with the underlying temporal patterns, potentially leading to overfitting. In this paper, we propose Sparse-Scale Transformer (SSformer) with Bidirectional Awareness for Time Series Forecasting to enhance the multi-scale modeling for time series. Specifically, we propose a novel Sparse-Scale Convolution (SSC) block that imposes sparsity on scales to obtain the informative representations by evaluating the intra-scale segment similarity of time series, and utilizes scale-specific convolutions to extract local patterns. Furthermore, we design a Bidirectional-Scale Interaction (BSI) block to explicitly model scale correlations in both coarse-to-fine and fine-to-coarse directions. Finally, scale predictions are ensembled to fully exploit the complementary forecasting capabilities across scales. Extensive experiments on various real-world datasets demonstrate that SSformer achieves state-of-the-art performance with superior efficiency.

NeurIPS Conference 2025 Conference Paper

3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model

  • Wenbo Hu
  • Yining Hong
  • Yanjun Wang
  • Leison Gao
  • Zibu Wei
  • Xingcheng Yao
  • Nanyun Peng
  • Yonatan Bitton

Humans excel at performing complex tasks by leveraging long-term memory across temporal and spatial experiences. In contrast, current Large Language Models (LLMs) struggle to effectively plan and act in dynamic, multi-room 3D environments. We posit that part of this limitation is due to the lack of proper 3D spatial-temporal memory modeling in LLMs. To address this, we first introduce 3DMem-Bench, a comprehensive benchmark comprising over 26, 000 trajectories and 2, 892 embodied tasks, question-answering and captioning, designed to evaluate an agent's ability to reason over long-term memory in 3D environments. Second, we propose 3DLLM-Mem, a novel dynamic memory management and fusion model for embodied spatial-temporal reasoning and actions in LLMs. Our model uses working memory tokens, which represents current observations, as queries to selectively attend to and fuse the most useful spatial and temporal features from episodic memory, which stores past observations and interactions. Our approach allows the agent to focus on task-relevant information while maintaining memory efficiency in complex, long-horizon environments. Experimental results demonstrate that 3DLLM-Mem achieves state-of-the-art performance across various tasks, outperforming the strongest baselines by 16. 5\% in success rate on 3DMem-Bench's most challenging in-the-wild embodied tasks.

EAAI Journal 2024 Journal Article

Automated detection and quantification of pavement cracking around manhole

  • Jun Peng
  • Weidong Wang
  • Wenbo Hu
  • Chengbo Ai
  • Xinyue Xu
  • Youyin Shi
  • Jin Wang
  • Zhifa Ran

Damage detection plays an important role in pavement health monitoring and inspection. Unfortunately, research about damage detection of the special component of pavement structures, such as pavement manhole covers, is relatively few. A new pipeline for the detection and quantification of damage around the pavement manhole covers is proposed in this research. In this pipeline, the Attention-enhanced Manhole Detection Model (AMDM) is proposed to detect manhole covers. AMDM achieves an ideal balance between accuracy and speed by eliminating redundant structures and incorporating an attention mechanism. The BCSM (Boundary-enhanced Crack Segmentation Model) is proposed to segment the damage around the manhole cover, and the boundary loss function is used to enhance the fine segmentation ability of the model on the boundary. The MAP (Mean Average Precision) of the manhole covers detection model is 96. 68%, and the MIOU (Mean Intersection Over Union) of the crack segmentation model is 89. 73%. Attributing to this efficient and accurate pipeline, a reasonable damage evaluation method is proposed in the end, which is based on statistical data and engineering experience. Overall, not only this research will contribute an automatic and cost-effective method to the detection and evaluation of damage around the manhole cover, but also inspire the detection and evaluation of other special components of civil engineering structures.

AAAI Conference 2024 Conference Paper

BLIVA: A Simple Multimodal LLM for Better Handling of Text-Rich Visual Questions

  • Wenbo Hu
  • Yifan Xu
  • Yi Li
  • Weiyue Li
  • Zeyuan Chen
  • Zhuowen Tu

Vision Language Models (VLMs), which extend Large Language Models (LLM) by incorporating visual understanding capability, have demonstrated significant advancements in addressing open-ended visual question-answering (VQA) tasks. However, these models cannot accurately interpret images infused with text, a common occurrence in real-world scenarios. Standard procedures for extracting information from images often involve learning a fixed set of query embeddings. These embeddings are designed to encapsulate image contexts and are later used as soft prompt inputs in LLMs. Yet, this process is limited to the token count, potentially curtailing the recognition of scenes with text-rich context. To improve upon them, the present study introduces BLIVA: an augmented version of InstructBLIP with Visual Assistant. BLIVA incorporates the query embeddings from InstructBLIP and also directly projects encoded patch embeddings into the LLM, a technique inspired by LLaVA. This approach assists the model to capture intricate details potentially missed during the query decoding process. Empirical evidence demonstrates that our model, BLIVA, significantly enhances performance in processing text-rich VQA benchmarks (up to 17.76% in OCR-VQA benchmark) and in undertaking general (not particularly text-rich) VQA benchmarks (up to 7.9% in Visual Spatial Reasoning benchmark), and achieved 17.72% overall improvement in a comprehensive multimodal LLM benchmark (MME), comparing to our baseline InstructBLIP. BLIVA demonstrates significant capability in decoding real-world images, irrespective of text presence. To demonstrate the broad industry applications enabled by BLIVA, we evaluate the model using a new dataset comprising YouTube thumbnails paired with question-answer sets across 11 diverse categories. For researchers interested in further exploration, our code and models are freely accessible at https://github.com/mlpc-ucsd/BLIVA.

NeurIPS Conference 2024 Conference Paper

CV-VAE: A Compatible Video VAE for Latent Generative Video Models

  • Sijie Zhao
  • Yong Zhang
  • Xiaodong Cun
  • Shaoshu Yang
  • Muyao Niu
  • Xiaoyu Li
  • Wenbo Hu
  • Ying Shan

Spatio-temporal compression of videos, utilizing networks such as Variational Autoencoders (VAE), plays a crucial role in OpenAI's SORA and numerous other video generative models. For instance, many LLM-like video models learn the distribution of discrete tokens derived from 3D VAEs within the VQVAE framework, while most diffusion-based video models capture the distribution of continuous latent extracted by 2D VAEs without quantization. The temporal compression is simply realized by uniform frame sampling which results in unsmooth motion between consecutive frames. Currently, there lacks of a commonly used continuous video (3D) VAE for latent diffusion-based video models in the research community. Moreover, since current diffusion-based approaches are often implemented using pre-trained text-to-image (T2I) models, directly training a video VAE without considering the compatibility with existing T2I models will result in a latent space gap between them, which will take huge computational resources for training to bridge the gap even with the T2I models as initialization. To address this issue, we propose a method for training a video VAE of latent video models, namely CV-VAE, whose latent space is compatible with that of a given image VAE, e. g. , image VAE of Stable Diffusion (SD). The compatibility is achieved by the proposed novel latent space regularization, which involves formulating a regularization loss using the image VAE. Benefiting from the latent space compatibility, video models can be trained seamlessly from pre-trained T2I or video models in a truly spatio-temporally compressed latent space, rather than simply sampling video frames at equal intervals. To improve the training efficiency, we also design a novel architecture for the video VAE. With our CV-VAE, existing video models can generate four times more frames with minimal finetuning. Extensive experiments are conducted to demonstrate the effectiveness of the proposed video VAE.

NeurIPS Conference 2024 Conference Paper

Matryoshka Query Transformer for Large Vision-Language Models

  • Wenbo Hu
  • Zi-Yi Dou
  • Liunian H. Li
  • Amita Kamath
  • Nanyun Peng
  • Kai-Wei Chang

Large Vision-Language Models (LVLMs) typically encode an image into a fixed number of visual tokens (e. g. , 576) and process these tokens with a language model. Despite their strong performance, LVLMs face challenges in adapting to varying computational constraints. This raises the question: can we achieve flexibility in the number of visual tokens to suit different tasks and computational resources? We answer this with an emphatic yes. Inspired by Matryoshka Representation Learning, we introduce the Matryoshka Query Transformer (MQT), capable of encoding an image into $m$ visual tokens during inference, where $m$ can be any number up to a predefined maximum. This is achieved by employing a query transformer with $M$ latent query tokens to compress the visual embeddings. During each training step, we randomly select $m \leq M$ latent query tokens and train the model using only these first $m$ tokens, discarding the rest. Combining MQT with LLaVA, we train a single model once, and flexibly and drastically reduce the number of inference-time visual tokens while maintaining similar or better performance compared to training independent models for each number of tokens. Our model, MQT-LLaVA, matches LLaVA-1. 5 performance across 11 benchmarks using a maximum of 256 tokens instead of LLaVA’s fixed 576. Reducing to 16 tokens (8x less TFLOPs) only sacrifices the performance by 2. 4 points on MMBench. On certain tasks such as ScienceQA and MMMU, we can even go down to only 2 visual tokens with performance drops of just 3\% and 6\% each. Our exploration of the trade-off between the accuracy and computational cost brought about by the number of visual tokens facilitates future research to achieve the best of both worlds.

AAAI Conference 2024 Conference Paper

Spotting the Unseen: Reciprocal Consensus Network Guided by Visual Archetypes

  • Wenbo Hu
  • Hongjian Zhan
  • Xinchen Ma
  • Yue Lu
  • Ching Y. Suen

Humans often require only a few visual archetypes to spot novel objects. Based on this observation, we present a strategy rooted in ``spotting the unseen" by establishing dense correspondences between potential query image regions and a visual archetype, and we propose the Consensus Network (CoNet). Our method leverages relational patterns intra and inter images via Auto-Correlation Representation (ACR) and Mutual-Correlation Representation (MCR). Within each image, the ACR module is capable of encoding both local self-similarity and global context simultaneously. Between the query and support images, the MCR module computes the cross-correlation across two image representations and introduces a reciprocal consistency constraint, which can incorporate to exclude outliers and enhance model robustness. To overcome the challenges of low-resource training data, particularly in one-shot learning scenarios, we incorporate an adaptive margin strategy to better handle diverse instances. The experimental results indicate the effectiveness of the proposed method across diverse domains such as object detection in natural scenes, and text spotting in both historical manuscripts and natural scenes, which demonstrates its sparkling generalization ability. Our code is available at: https://github.com/infinite-hwb/conet.

EAAI Journal 2024 Journal Article

TANet: Text region attention learning for vehicle re-identification

  • Wenbo Hu
  • Hongjian Zhan
  • Palaiahnakote Shivakumara
  • Umapada Pal
  • Yue Lu

In recent years, the challenge of distinguishing vehicles of the same model has prompted a shift towards leveraging both global appearances and local features, such as lighting and rearview mirrors, for vehicle re-identification (ReID). Despite advancements, accurately identifying vehicles remains complex, particularly due to the underutilization of highly discriminative text regions. This paper introduces the Text Region Attention Network (TANet), a novel approach that integrates global and local information with a specific focus on text regions for improved feature learning. TANet uniquely captures stable and distinctive features across various vehicle views, demonstrating its effectiveness through rigorous evaluation on the VeRi-776, VehicleID, and VERI-Wild datasets. TANet significantly outperforms existing methods, achieving mAP scores of 83. 6% on VeRi-776, 84. 4% on VehicleID (Large), and 76. 6% on VERI-Wild (Large). Statistical tests further validate the superiority of TANet over the baseline, showcasing notable improvements in mAP and Top-1 through Top-15 accuracy metrics.

IJCAI Conference 2021 Conference Paper

Two Birds with One Stone: Series Saliency for Accurate and Interpretable Multivariate Time Series Forecasting

  • Qingyi Pan
  • Wenbo Hu
  • Ning Chen

It is important yet challenging to perform accurate and interpretable time series forecasting. Though deep learning methods can boost forecasting accuracy, they often sacrifice interpretability. In this paper, we present a new scheme of series saliency to boost both accuracy and interpretability. By extracting series images from sliding windows of the time series, we design series saliency as a mixup strategy with a learnable mask between the series images and their perturbed versions. Series saliency is model agnostic and performs as an adaptive data augmentation method for training deep models. Moreover, by slightly changing the objective, we optimize series saliency to find a mask for interpretable forecasting in both feature and time dimensions. Experimental results on several real datasets demonstrate that series saliency is effective to produce accurate time-series forecasting results as well as generate temporal interpretations.

NeurIPS Conference 2020 Conference Paper

Calibrated Reliable Regression using Maximum Mean Discrepancy

  • Peng Cui
  • Wenbo Hu
  • Jun Zhu

Accurate quantification of uncertainty is crucial for real-world applications of machine learning. However, modern deep neural networks still produce unreliable predictive uncertainty, often yielding over-confident predictions. In this paper, we are concerned with getting well-calibrated predictions in regression tasks. We propose the calibrated regression method using the maximum mean discrepancy by minimizing the kernel embedding measure. Theoretically, the calibration error of our method asymptotically converges to zero when the sample size is large enough. Experiments on non-trivial real datasets show that our method can produce well-calibrated and sharp prediction intervals, which outperforms the related state-of-the-art methods.

IJCAI Conference 2017 Conference Paper

Semi-supervised Max-margin Topic Model with Manifold Posterior Regularization

  • Wenbo Hu
  • Jun Zhu
  • Hang Su
  • Jingwei Zhuo
  • Bo Zhang

Supervised topic models leverage label information to learn discriminative latent topic representations. As collecting a fully labeled dataset is often time-consuming, semi-supervised learning is of high interest. In this paper, we present an effective semi-supervised max-margin topic model by naturally introducing manifold posterior regularization to a regularized Bayesian topic model, named LapMedLDA. The model jointly learns latent topics and a related classifier with only a small fraction of labeled documents. To perform the approximate inference, we derive an efficient stochastic gradient MCMC method. Unlike the previous semi-supervised topic models, our model adopts a tight coupling between the generative topic model and the discriminative classifier. Extensive experiments demonstrate that such tight coupling brings significant benefits in quantitative and qualitative performance.

v2026.09.13