Arrow Research search

Author name cluster

Haoran Liu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

12 papers
2 author rows

Possible papers

12

JBHI Journal 2026 Journal Article

KAFSTExp: Kernel Adaptive Filtering With Nyström Approximation for Predicting Spatial Gene Expression From Histology Images

  • Haoran Liu
  • Hossein Farahani
  • Xifeng Li
  • Yongle Xie
  • Ali Bashashati

Spatial transcriptomics (ST), known as an expensive medical examination, plays an important role in analyzing the spatial heterogeneity of tumors. When considering the correlation between tissue morphological patterns and gene profiles, predicting corresponding gene expression from pathology images obtained from affordable biopsies is regarded as an instantaneous and cost-effective alternative. However, accurately modeling the complex and nonlinear relationship between histological features and gene expression remains challenging. Existing deep learning models often struggle to generalize on limited ST datasets due to their large and overparameterized architectures. The primary advantage of kernel adaptive filtering (KAF) lies in its ability to transform a challenging nonlinear problem arising in the original space into a linear regression problem in the higher-dimensional feature space via kernel methods. Therefore, this paper proposes a framework called KAFSTExp, which utilizes the state-of-the-art pathology foundation model UNI to extract image feature vectors, and then introduces the kernel least mean square algorithm with Nyström approximation to predict the normalized transcript counts of specific genes. Extensive experiments show that KAFSTExp significantly improves prediction accuracy while reducing computational cost and training time. KAFSTExp demonstrates consistent performance gains across multiple ST datasets, achieving relative improvements in Pearson correlation coefficient ranging from 1. 24% to 94. 23%, with an average increase of 19. 80% over the best-performing non-KAF methods. External validation and further clinical analysis confirm the generalization performance and clinical application value of the proposed KAFSTExp.

AAAI Conference 2026 Conference Paper

SkyMoE: A Vision-Language Foundation Model for Enhancing Geospatial Interpretation with Mixture of Experts

  • Jiaqi Liu
  • Ronghao Fu
  • Lang Sun
  • Haoran Liu
  • Xiao Yang
  • Weipeng Zhang
  • Xu Na
  • Zhuoran Duan

The emergence of large vision-language models (VLMs) has significantly enhanced the efficiency and flexibility of geospatial interpretation. However, general-purpose VLMs remain suboptimal for remote sensing (RS) tasks. Existing geospatial VLMs typically adopt a unified modeling strategy and struggle to differentiate between task types and interpretation granularities, limiting their ability to balance local detail perception and global contextual understanding. In this paper, we present SkyMoE, a Mixture-of-Experts (MoE) vision-language model tailored for multimodal, multi-task RS interpretation. SkyMoE employs an adaptive router that generates task- and granularity-aware routing instructions, enabling specialized large language model experts to handle diverse sub-tasks. To further promote expert decoupling and granularity sensitivity, we introduce a context-disentangled augmentation strategy that creates contrastive pairs between local and global features, guiding experts toward level-specific representation learning. We also construct MGRS-Bench, a comprehensive benchmark covering multiple RS interpretation tasks and granularity levels, to evaluate generalization in complex scenarios. Extensive experiments on 21 public datasets demonstrate that SkyMoE achieves state-of-the-art performance across tasks, validating its adaptability, scalability, and superior multi-granularity understanding in remote sensing.

AAAI Conference 2025 Conference Paper

Learning Disentangled Equivariant Representation for Explicitly Controllable 3D Molecule Generation

  • Haoran Liu
  • Youzhi Luo
  • Tianxiao Li
  • James Caverlee
  • Martin Renqiang Min

We consider the conditional generation of 3D drug-like molecules with explicit control over molecular properties such as drug-like properties (e.g., Quantitative Estimate of Druglikeness or Synthetic Accessibility score) and effectively binding to specific protein sites. To tackle this problem, we propose an E(3)-equivariant Wasserstein autoencoder and factorize the latent space of our generative model into two disentangled aspects: molecular properties and the remaining structural context of 3D molecules. Our model ensures explicit control over these molecular attributes while maintaining equivariance of coordinate representation and invariance of data likelihood. Furthermore, we introduce a novel alignment-based coordinate loss to adapt equivariant networks for auto-regressive de-novo 3D molecule generation from scratch. Extensive experiments validate our model's effectiveness on property-guided and context-guided molecule generation, both for de-novo 3D molecule design and structure-based drug discovery against protein targets.

ICRA Conference 2025 Conference Paper

Na Vid-4D: Unleashing Spatial Intelligence in Egocentric RGB-D Videos for Vision-and-Language Navigation

  • Haoran Liu
  • Weikang Wan
  • Xiqian Yu
  • Minghan Li
  • Jiazhao Zhang
  • Bo Zhao 0015
  • Zhibo Chen 0001
  • Zhongyuan Wang 0006

Understanding and reasoning about the 4D space-time is crucial for Vision-and-Language Navigation (VLN). However, previous works lack in-depth exploration in this aspect, resulting in bottlenecked spatial perception and action precision of VLN agents. In this work, we introduce NaVid-4D, a Vision Language Model (VLM) based navigation agent taking the lead in explicitly showcasing the capabilities of spatial intelligence in the real world. Given natural language instructions, NaVid-4D requires only egocentric RGB-D video streams as observations to perform spatial understanding and reasoning for generating precise instruction-following robotic actions. NaVid-4D learns navigation policies using the data from simulation environments and is endowed with precise spatial understanding and reasoning capabilities using web data. Without the need to pre-train an RGB-D foundation model, we propose a method capable of directly injecting the depth features into the visual encoder of a VLM. We further compare the use of factually captured depth information with the monocularly estimated one and find NaVid-4D works well with both while using estimated depth offers greater gener-alization capability and better mitigates the sim-to-real gap. Extensive experiments demonstrate that NaVid-4D achieves state-of-the-art performance in simulation environment and makes impressive VLN performance with spatial intelligence happen in the real world.

IROS Conference 2025 Conference Paper

Tele-GS: 3D Gaussian Scene Representation for Low-Bandwidth Teleoperation

  • Chunyang Zhao
  • Zeyu Zhou
  • Haoran Liu
  • Dogan Kircali
  • Huan Yang
  • Chang Boon Low
  • Yuanzhe Wang
  • Danwei Wang

Video streaming based teleoperation often faces a trade-off between bandwidth consumption and the need for high-fidelity telepresence. Higher image resolution or a wider field of view (FOV) substantially increases bandwidth requirements. In this paper, we propose a novel telepresence model for teleoperated vehicles operating in bandwidth-constrained environments. Our approach employs a LiDAR-fused 3D Gaussian Splatting (3DGS) as a compact scene representation to efficiently generate remote views. Initially, a static point cloud map is constructed using LiDAR-based semantic mapping, which serves as the initial Gaussians for optimizing the 3DGS model. During teleoperation, the prebuilt 3DGS is then rendered on the teleoperation platform, while only safety-critical information, such as vehicle pose and dynamic objects, is transmitted from the vehicle to the teleoperator in real-time. The proposed telepresence model significantly reduces data transmission requirements while maintaining photorealistic telepresence, enabling reliable and effective teleoperation even under stringent bandwidth constraints. This capability ensures safe and efficient vehicle teleoperation under challenging environments without relying on traditional high-bandwidth communication, thereby broadening the applicability of teleoperation technology to more demanding and diverse operational scenarios. Real-world experimental results show that the developed system can provide immersive teleoperation experiences at Kbps-level bandwidth consumption.

EAAI Journal 2024 Journal Article

An efficient skeleton learning approach-based hybrid algorithm for identifying Bayesian network structure

  • Niantai Wang
  • Haoran Liu
  • Liyue Zhang
  • Yanbin Cai
  • Qianrui Shi

Bayesian network (BN) structure learning is the basis of BN applications and plays a pivotal role in many machine learning tasks. Whereas remarkable progress in structure learning has been achieved in the past, making further improvements in the efficiency and accuracy of structure learning is a significant challenge. In this paper, we propose an efficient skeleton learning approach-based hybrid algorithm (ESLH), which consists of two phases. In the constraint-based phase, a dynamic threshold (DTH) strategy and a skeleton learning method based on triangle breaking (SLTB) are proposed to learn the skeleton of a BN structure efficiently. DTH designs a dynamic threshold to remove redundant edges in the initial skeleton with little time overhead. By the result of DTH, SLTB first finds, tests and breaks triangles in the initial skeleton to efficiently remove redundant edges and then removes the remaining redundant edges to discover the final skeleton. In the score-and-search phase, ESLH employs the hill-climbing algorithm to find the highest-scored structure. We propose a novel strategy to divide this phase into three steps, both utilizing the learned skeleton to constrain the search space and preventing the errors of the learned skeleton from reducing the quality of the final learned structure. Extensive experiments on benchmark BNs validate the effectiveness of DTH and SLTB and demonstrate that ESLH is more than five times faster than the state-of-the-art structure learning algorithms while maintaining the highest average accuracy.

AAAI Conference 2024 Conference Paper

DiDA: Disambiguated Domain Alignment for Cross-Domain Retrieval with Partial Labels

  • Haoran Liu
  • Ying Ma
  • Ming Yan
  • Yingke Chen
  • Dezhong Peng
  • Xu Wang

Driven by generative AI and the Internet, there is an increasing availability of a wide variety of images, leading to the significant and popular task of cross-domain image retrieval. To reduce annotation costs and increase performance, this paper focuses on an untouched but challenging problem, i.e., cross-domain image retrieval with partial labels (PCIR). Specifically, PCIR faces great challenges due to the ambiguous supervision signal and the domain gap. To address these challenges, we propose a novel method called disambiguated domain alignment (DiDA) for cross-domain retrieval with partial labels. In detail, DiDA elaborates a novel prototype-score unitization learning mechanism (PSUL) to extract common discriminative representations by simultaneously disambiguating the partial labels and narrowing the domain gap. Additionally, DiDA proposes a prototype-based domain alignment mechanism (PBDA) to further bridge the inherent cross-domain discrepancy. Attributed to PSUL and PBDA, our DiDA effectively excavates domain-invariant discrimination for cross-domain image retrieval. We demonstrate the effectiveness of DiDA through comprehensive experiments on three benchmarks, comparing it to existing state-of-the-art methods. Code available: https://github.com/lhrrrrrr/DiDA.

IROS Conference 2024 Conference Paper

Towards Kbps-level Vehicle Teleoperation via Persistent-Transient Environment Modelling

  • Chunyang Zhao
  • Zeyu Zhou
  • Haoran Liu
  • Dogan Kircali
  • Guoyi Chi
  • Hongming Shen
  • Yuanzhe Wang
  • Danwei Wang

Traditional teleoperation technologies based on video streaming are facing several challenges in practical applications, including limited bandwidth, constrained spatial awareness, and sensitivity to illumination. Existing studies have not adequately addressed these issues. This paper presents a novel non-video based teleoperation framework for autonomous vehicles operating in bandwidth-limited environments. To reduce the amount of data being transmitted, a persistent-transient environment model is proposed for telepresence. Initially, a digital twin of the environment is preconstructed, containing only persistent environmental information. Subsequently, transient information captured by onboard sensors, such as vehicle state and dynamic objects, necessitate real-time transmission. Based on this model, a 3D virtual scene is rendered in front of the teleoperator, offering any desired virtual viewpoint to enhance spatial awareness. This telepresence model only requires real-time transmission of minimal data, i. e. , vehicle state and detected objects, and remains unaffected by illumination conditions, enabling teleoperation even in applications with Kbps-level bandwidth constraints. Experimental results showcase the substantial potential of the proposed framework in bandwidth-limited settings.

ICLR Conference 2023 Conference Paper

Gradient-Guided Importance Sampling for Learning Binary Energy-Based Models

  • Meng Liu 0015
  • Haoran Liu
  • Shuiwang Ji

Learning energy-based models (EBMs) is known to be difficult especially on discrete data where gradient-based learning strategies cannot be applied directly. Although ratio matching is a sound method to learn discrete EBMs, it suffers from expensive computation and excessive memory requirements, thereby resulting in difficulties in learning EBMs on high-dimensional data. Motivated by these limitations, in this study, we propose ratio matching with gradient-guided importance sampling (RMwGGIS). Particularly, we use the gradient of the energy function w.r.t. the discrete data space to approximately construct the provably optimal proposal distribution, which is subsequently used by importance sampling to efficiently estimate the original ratio matching objective. We perform experiments on density modeling over synthetic discrete data, graph generation, and training Ising models to evaluate our proposed method. The experimental results demonstrate that our method can significantly alleviate the limitations of ratio matching, perform more effectively in practice, and scale to high-dimensional problems. Our implementation is available at https://github.com/divelab/RMwGGIS.

ICLR Conference 2023 Conference Paper

Learning Hierarchical Protein Representations via Complete 3D Graph Networks

  • Limei Wang
  • Haoran Liu
  • Yi Liu 0059
  • Jerry Kurtin
  • Shuiwang Ji

We consider representation learning for proteins with 3D structures. We build 3D graphs based on protein structures and develop graph networks to learn their representations. Depending on the levels of details that we wish to capture, protein representations can be computed at different levels, \emph{e.g.}, the amino acid, backbone, or all-atom levels. Importantly, there exist hierarchical relations among different levels. In this work, we propose to develop a novel hierarchical graph network, known as ProNet, to capture the relations. Our ProNet is very flexible and can be used to compute protein representations at different levels of granularity. By treating each amino acid as a node in graph modeling as well as harnessing the inherent hierarchies, our ProNet is more effective and efficient than existing methods. We also show that, given a base 3D graph network that is complete, our ProNet representations are also complete at all levels. Experimental results show that ProNet outperforms recent methods on most datasets. In addition, results indicate that different downstream tasks may require representations at different levels. Our code is publicly available as part of the DIG library (\url{https://github.com/divelab/DIG}).

NeurIPS Conference 2022 Conference Paper

ComENet: Towards Complete and Efficient Message Passing for 3D Molecular Graphs

  • Limei Wang
  • Yi Liu
  • Yuchao Lin
  • Haoran Liu
  • Shuiwang Ji

Many real-world data can be modeled as 3D graphs, but learning representations that incorporates 3D information completely and efficiently is challenging. Existing methods either use partial 3D information, or suffer from excessive computational cost. To incorporate 3D information completely and efficiently, we propose a novel message passing scheme that operates within 1-hop neighborhood. Our method guarantees full completeness of 3D information on 3D graphs by achieving global and local completeness. Notably, we propose the important rotation angles to fulfill global completeness. Additionally, we show that our method is orders of magnitude faster than prior methods. We provide rigorous proof of completeness and analysis of time complexity for our methods. As molecules are in essence quantum systems, we build the \underline{com}plete and \underline{e}fficient graph neural network (ComENet) by combing quantum inspired basis functions and the proposed message passing scheme. Experimental results demonstrate the capability and efficiency of ComENet, especially on real-world datasets that are large in both numbers and sizes of graphs. Our code is publicly available as part of the DIG library (\url{https: //github. com/divelab/DIG}).

JMLR Journal 2021 Journal Article

DIG: A Turnkey Library for Diving into Graph Deep Learning Research

  • Meng Liu
  • Youzhi Luo
  • Limei Wang
  • Yaochen Xie
  • Hao Yuan
  • Shurui Gui
  • Haiyang Yu
  • Zhao Xu

Although there exist several libraries for deep learning on graphs, they are aiming at implementing basic operations for graph deep learning. In the research community, implementing and benchmarking various advanced tasks are still painful and time-consuming with existing libraries. To facilitate graph deep learning research, we introduce DIG: Dive into Graphs, a turnkey library that provides a unified testbed for higher level, research-oriented graph deep learning tasks. Currently, we consider graph generation, self-supervised learning on graphs, explainability of graph neural networks, and deep learning on 3D graphs. For each direction, we provide unified implementations of data interfaces, common algorithms, and evaluation metrics. Altogether, DIG is an extensible, open-source, and turnkey library for researchers to develop new methods and effortlessly compare with common baselines using widely used datasets and evaluation metrics. Source code is available at https://github.com/divelab/DIG. [abs] [ pdf ][ bib ] [ code ] &copy JMLR 2021. ( edit, beta )

v2026.09.13