Arrow Research search

Author name cluster

Hongbo Liu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

11 papers
1 author row

Possible papers

11

AAAI Conference 2026 Conference Paper

EchoGen: Cycle-Consistent Learning for Unified Layout-Image Generation and Understanding

  • Kai Zou
  • Hongbo Liu
  • Dian Zheng
  • Jianxiong Gao
  • Zhiwei Zhao
  • Bin Liu

In this work, we present EchoGen, a unified framework for layout-to-image generation and image grounding, capable of generating images with both accurate layout and high fidelity to the text description.(e.g., spatial relationship), and grounding the image robustly at the same time. We believe that image grounding possesses strong text and layout understanding abilities, which can compensate for the corresponding limitations in layout-to-image generation. At the same time, images generated from layouts exhibit high diversity in content, thereby enhancing the robustness of image grounding. Jointly training both tasks within a unified model can promote performance improvements for each. However, we identify that this joint training paradigm encounters several optimization challenges and results in restricted performance. To address these issues, we propose progressive training strategies. First, the Parallel Multi-Task Pre-training (PMTP) stage equips the model with basic abilities for both tasks, leveraging shared tokens to accelerate training. Next, the Dual Joint Optimization (DJO) stage exploits task duality to sequentially integrate the two tasks, enabling unified optimization. Finally, the Cycle RL stage eliminates reliance on visual supervision by using consistency constraints as rewards, significantly enhancing the model’s unified capabilities via the GRPO strategy. Extensive experiments demonstrate state-of-the-art results on both layout-to-image generation and image grounding benchmarks, and reveal clear synergistic gains from optimizing the two tasks together.

AAAI Conference 2026 Conference Paper

From Subtle to Significant: Prompt-Driven Self-Improving Optimization in Test-Time Graph OOD Detection

  • Luzhi Wang
  • Xuanshuo Fu
  • He Zhang
  • Chuang Liu
  • Xiaobao Wang
  • Hongbo Liu

Graph Out-of-Distribution (OOD) detection aims to identify whether a test graph deviates from the distribution of graphs observed during training, which is critical for ensuring the reliability of Graph Neural Networks (GNNs) when deployed in open-world scenarios. Recent advances in graph OOD detection have focused on test-time training techniques that facilitate OOD detection without accessing potential supervisory information (e.g., training data). However, most of these methods employ a one-pass inference paradigm, which prevents them from progressively correcting erroneous predictions to amplify OOD signals. To this end, we propose a Self-Improving Graph Out-of-Distribution detector (SIGOOD), which is an unsupervised framework that integrates continuous self-learning with test-time training for effective graph OOD detection. Specifically, SIGOOD generates a prompt to construct a prompt-enhanced graph that amplifies potential OOD signals. To optimize prompts, SIGOOD introduces an Energy Preference Optimization (EPO) loss, which leverages energy variations between the original test graph and the prompt-enhanced graph. By iteratively optimizing the prompt by involving it into the detection model in a self-improving loop, the resulting optimal prompt-enhanced graph is ultimately used for OOD detection. Comprehensive evaluations on 21 real-world datasets confirm the effectiveness and outperformance of our SIGOOD method.

AAMAS Conference 2026 Conference Paper

RBC: Retroactive Belief State Compensation for Multi-Agent Collaboration Under Information Delay

  • Dongkun Huo
  • Hongbo Liu
  • Shu Yin
  • Yixue Hao
  • Long Hu
  • Rui Wang
  • Min Chen

Real-time information is usually not satisfied in real world due to communication or observation delay. Although existing works address individual delay, they do not fully consider the complex effects of composite delay, denoted as “Information Delay”, which severely reduce the efficiency of these methods. To address information delay, we propose Retroactive Belief state Compensation (RBC), a multi-agent framework with enhanced robustness and collaboration. Specifically, we design a multi-step reconstruction model that retroactively rebuilds agents’ belief states starting from the generation time of the information. This process corrects the accumulated deviation in the current belief state caused by information delay. Moreover, to enhance proactive collaboration, we introduce an intent inference module. This module enables agents to generate intents, which represent short-term action plans, as content of communication. By aggregating intents from teammates, agents will choose more coherent and synchronized joint actions. To evaluate the performance of RBC, we design scenarios with multiple levels of observation, communication, and composite delays. Experimental results demonstrate that RBC outperforms the baselines in all scenarios with delays. Yixue Hao is corresponding author. Email: yixuehao@hust. edu. cn. This work is licensed under a Creative Commons Attribution International 4. 0 License. Proc. of the 25th International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2026), C. Amato, L. Dennis, V. Mascardi, J. Thangarajah (eds.), May 25 – 29, 2026, Paphos, Cyprus. © 2026 International Foundation for Autonomous Agents and Multiagent Systems (www. ifaamas. org). https: //doi. org/10. 65109/MFLP4403

AAAI Conference 2026 Conference Paper

Towards Distance-Invariant Radio Frequency Fingerprinting via Augmented Unsupervised Learning

  • Shiyue Huang
  • Yuchen Su
  • Hongbo Liu
  • Zikang Ding
  • Xuewan He
  • Yanzhi Ren
  • Haitao Jia

Radio Frequency Fingerprinting (RFF) exploits inherent hardware-level imperfections of wireless transmitters as unclonable identifiers for device identification. These unique signatures, concealed in transmitted signals, inevitably experience complex distortions during wireless propagation (i.e., coupled with ambient noise and channel fading), making it extremely challenging for reliable extraction. Despite substantial research efforts dedicated to advancing effective fingerprint extraction techniques, current approaches still struggle in handling fingerprint robustness under distance variations, leading to severe SNR fluctuations and complex multipath effects. To address this gap, we propose the first unsupervised framework for distance-invariant radio frequency fingerprinting, eliminating dependence on labeled target domain data. Specifically, we first preprocess raw RF samples by confining them within a specified variation range and filtering noisy high-frequency components while avoiding aliasing. For source domain data, we then propose a set of physics-inspired data augmentation techniques designed to emulate realistic wireless signal propagation effects. Building on this, we introduce a dual alignment contrastive learning method to explicitly decouple identity-discriminative features, ensuring the model focuses on device-specific traits. Furthermore, we incorporate a pseudo-labeling-based domain adaptation module to refine the model for the unlabeled target domain, enhancing its generalization to unseen distances. Extensive experiments on public datasets show that our method achieves the identification accuracy outperforming state-of-the-art approaches by 40%, while maintaining computational efficiency suitable for edge deployment.

NeurIPS Conference 2025 Conference Paper

ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models

  • Hongbo Liu
  • Jingwen He
  • Yi Jin
  • Dian Zheng
  • Yuhao Dong
  • Fan Zhang
  • Ziqi Huang
  • Yinan He

Recent Vision-Language Models (VLMs) have shown strong performance in general-purpose visual understanding and reasoning, but their ability to comprehend the visual grammar of movie shots remains underexplored and insufficiently evaluated. To bridge this gap, we present \textbf{ShotBench}, a dedicated benchmark for assessing VLMs’ understanding of cinematic language. ShotBench includes 3, 049 still images and 500 video clips drawn from more than 200 films, with each sample annotated by trained annotators or curated from professional cinematography resources, resulting in 3, 608 high-quality question-answer pairs. We conduct a comprehensive evaluation of over 20 state-of-the-art VLMs across eight core cinematography dimensions. Our analysis reveals clear limitations in fine-grained perception and cinematic reasoning of current VLMs. To improve VLMs capability in cinematography understanding, we construct a large-scale multimodal dataset, named ShotQA, which contains about 70k Question-Answer pairs derived from movie shots. Besides, we propose ShotVL and train this VLM model with a two-stage training strategy, integrating both supervised fine-tuning and Group Relative Policy Optimization (GRPO). Experimental results demonstrate that our model achieves substantial improvements, surpassing all existing strongest open-source and proprietary models evaluated on ShotBench, establishing a new state-of-the-art performance.

EAAI Journal 2024 Journal Article

SOFT: Self-supervised sparse Optical Flow Transformer for video stabilization via quaternion

  • Naiyao Wang
  • Changdong Zhou
  • Rongfeng Zhu
  • Bo Zhang
  • Ye Wang
  • Hongbo Liu

Video stabilization is crucial for video representation learning, which suffers from the challenges such as the perception of unstable vision, the stripping and cognition of target motion features in complex scenes, the correction of the jittery camera systems trails. In this paper, we propose a Self-supervised sparse Optical Flow Transformer (SOFT) model, consisting of a self-supervised contrastive learning transformer network, a sparse optical flow perception network and a multimodal cognitive fusion network. The SOFT model takes advantage of optical flow to estimate motion. The sparse optical flow perception network perceiving partially sparse optical flow containing motion features. This serves as the input to the self-supervised contrastive learning transformer network for generating sparse optical flow features, which are fed into the multimodal cognitive fusion network together with the real and virtual camera pose for video frame warping. Experimental comparisons with state-of-the-art models on 4 metrics demonstrate the effectiveness of the SOFT model. It achieves the best performance with an average Stability of 0. 869 and average Distortion of 0. 993 across 6 categories videos, which shows that the SOFT model can effectively perceive the motion in the video and smooth the jitter track of videos.

EAAI Journal 2022 Journal Article

Human trajectory forecasting using a flow-based generative model

  • Bo Zhang
  • Tao Wang
  • Changdong Zhou
  • Nicola Conci
  • Hongbo Liu

In this article, we present a flow-based framework for multi-modal trajectory prediction, which is able to provide an accurate and explicit inference of the latent representations on trajectory data. Differently from other typical generative models (such as GAN, VAE, etc.), the flow-based models aim at learning data distribution explicitly through an invertible network, which can convert a complicated distribution into a tractable form via invertible transformations. The whole framework is built upon the standard encoder–decoder architecture, where the LSTM is exploited as the fundamental block to capture the temporal structure of a trajectory. As a core module, we incorporate an invertible network that can learn the multi-modal distributions of trajectory data and further generate plausible future paths by sampling tricks from the standard Gaussian distribution. Extensive experiments carried out on synthetic and realistic datasets demonstrate the effectiveness of the proposed approach, and show the advantages as compared to the GAN-based and the VAE-based prediction frameworks.

I&C Journal 2022 Journal Article

On enumerating algorithms of novel multiple leaf-distance granular regular α-subtrees of trees

  • Yu Yang
  • Hongbo Liu
  • Hua Wang
  • Xiao-Dong Zhang
  • C.L. Philip Chen

Subtrees and BC-subtrees (subtrees in which the distance between any two leaves is even) are important concepts in the study of complex graphical structures. In this article, we propose a novel generalization called the leaf-distance granular regular α-tree (abbreviated as LDR α-tree for short). This is a tree in which the distance between any two leaves is divisible by α (α is a positive integer). A LDR α-subtree is simply a subtree that is also a LDR α-tree. We present basic properties and generating functions related to the LDR α-subtrees enumeration. Based on those theoretical results we provide efficient algorithms for enumerating various LDR α-subtrees of trees. Our algorithms can serve as multi-distance granularity sifters of a graph to screen all the α-subtree uniformly, and thus provide novel insights into exploring new structural properties from the perspective of multiple leaf-distance granularity.

EAAI Journal 2021 Journal Article

AONet: Active Offset Network for crowd flow prediction

  • Dafeng Wang
  • Qian Ma
  • Naiyao Wang
  • Xuanzhe Fan
  • MingYu Lu
  • Hongbo Liu

Predicting crowd flow is of great importance to public safety and traffic management. The crowd flow is difficult to predict accurately and timely due to the uncertainty of the future positions. In this paper, we propose a novel Active Offset Network (AONet), in which ActiveGRU (Active Gate Recurrent Unit) is designed to predict the variation of pedestrians’ positions in the crowd flow. Its inner location-variant recurrent structure is implemented by utilizing convolution operation on low dimensional spatio-temporal sequences to obtain fractional offset locations. Afterwards, the sampling locations are determined by bilinear interpolation on fractional offset locations. Moreover, a probabilistic sparse strategy is introduced to reduce the links between sampling locations during supervised training. Finally, the experiments over popular benchmarks demonstrate that our method can actively characterize the future positions of pedestrians. Meanwhile, the performance of the proposed AONet is superior over state-of-art baselines with regard to both accuracy and computational savings.

EAAI Journal 2021 Journal Article

STENet: A hybrid spatio-temporal embedding network for human trajectory forecasting

  • Bo Zhang
  • Chengzhi Yuan
  • Tao Wang
  • Hongbo Liu

In this paper, we present a hybrid spatio-temporal embedding network (named as STENet) for human trajectory forecasting, which is built upon a GAN-based hierarchical framework. Differently from traditional approaches that only use LSTM for trajectory modeling, we exploit the 1D Convolutional Neural Network (1D-CNN) to embed position features at multiple temporal scales. Moreover, we propose a two-stage graph attention mechanism, which can better describe mutual interactions among pedestrians in the crowd. Additionally, group influences at every time step are taken into account as well. The overall framework is designed using a hierarchical manner, and trained using the Wasserstein distance. We carry out our experiments on the ETH and the UCY datasets. The corresponding results demonstrate the effectiveness of the proposed framework.

TCS Journal 2015 Journal Article

Enumeration of BC-subtrees of trees

  • Yu Yang
  • Hongbo Liu
  • Hua Wang
  • Scott Makeig

A BC-tree (block-cutpoint-tree) is a tree (with at least two vertices) where the distance between any two leaves is even. Motivated from the study of the “core” of a graph, BC-trees form an interesting class of trees. We provide a comprehensive study of questions related to BC-trees. As the analogue of the study of extremal questions on subtrees of trees, we first characterize the general extremal trees that maximize or minimize the number of BC-subtrees or leaf-containing BC-subtrees. We further discuss the “middle part” of a tree with respect to the number of BC-subtrees, namely the BC-subtree-core that behaves in a rather different way than all previously known “middle parts” of a tree. Last but not least, fast algorithms are proposed (following similar ideas as those of the enumeration of subtrees) for enumerating various classes of BC-subtrees of a tree.

v2026.09.13