Arrow Research search

Author name cluster

Yanjun Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

6 papers
2 author rows

Possible papers

6

AAAI Conference 2026 Conference Paper

Enhancing Multimodal Misinformation Detection by Replaying the Whole Story from Image Modality Perspective

  • Bing Wang
  • Ximing Li
  • Yanjun Wang
  • Changchun Li
  • Lin Yuanbo Wu
  • Buyu Wang
  • Shengsheng Wang

Multimodal Misinformation Detection (MMD) refers to the task of detecting social media posts involving misinformation, where the post often contains text and image modalities. However, by observing the MMD posts, we hold that the text modality may be much more informative than the image modality because the text generally describes the whole event/story of the current post but the image often presents partial scenes only. Our preliminary empirical results indicate that the image modality exactly contributes less to MMD. Upon this idea, we propose a new MMD method named RETSIMD. Specifically, we suppose that each text can be divided into several segments, and each text segment describes a partial scene that can be presented by an image. Accordingly, we split the text into a sequence of segments, and feed these segments into a pre-trained text-to-image generator to augment a sequence of images. We further incorporate two auxiliary objectives concerning text-image and image-label mutual information, and further post-train the generator over an auxiliary text-to-image generation benchmark dataset. Additionally, we propose a graph structure by defining three heuristic relationships between images, and use a graph neural network to generate the fused features. Extensive empirical results validate the effectiveness of RETSIMD.

NeurIPS Conference 2025 Conference Paper

3DLLM-Mem: Long-Term Spatial-Temporal Memory for Embodied 3D Large Language Model

  • Wenbo Hu
  • Yining Hong
  • Yanjun Wang
  • Leison Gao
  • Zibu Wei
  • Xingcheng Yao
  • Nanyun Peng
  • Yonatan Bitton

Humans excel at performing complex tasks by leveraging long-term memory across temporal and spatial experiences. In contrast, current Large Language Models (LLMs) struggle to effectively plan and act in dynamic, multi-room 3D environments. We posit that part of this limitation is due to the lack of proper 3D spatial-temporal memory modeling in LLMs. To address this, we first introduce 3DMem-Bench, a comprehensive benchmark comprising over 26, 000 trajectories and 2, 892 embodied tasks, question-answering and captioning, designed to evaluate an agent's ability to reason over long-term memory in 3D environments. Second, we propose 3DLLM-Mem, a novel dynamic memory management and fusion model for embodied spatial-temporal reasoning and actions in LLMs. Our model uses working memory tokens, which represents current observations, as queries to selectively attend to and fuse the most useful spatial and temporal features from episodic memory, which stores past observations and interactions. Our approach allows the agent to focus on task-relevant information while maintaining memory efficiency in complex, long-horizon environments. Experimental results demonstrate that 3DLLM-Mem achieves state-of-the-art performance across various tasks, outperforming the strongest baselines by 16. 5\% in success rate on 3DMem-Bench's most challenging in-the-wild embodied tasks.

NeurIPS Conference 2025 Conference Paper

Approximately Aligned Decoding

  • Daniel Melcer
  • Sujan Kumar Gonugondla
  • Pramuditha Perera
  • Haifeng Qian
  • Wen-Hao Chiang
  • Yanjun Wang
  • Nihal Jain
  • Pranav Garg

It is common to reject undesired outputs of Large Language Models (LLMs); however, current methods to do so require an excessive amount of computation to re-sample after a rejection, or distort the distribution of outputs by constraining the output to highly improbable tokens. We present a method, Approximately Aligned Decoding (AprAD), to balance the distortion of the output distribution with computational efficiency, inspired by algorithms from the speculative decoding literature. AprAD allows for the generation of long sequences of text with difficult-to-satisfy constraints, while amplifying low probability outputs much less compared to existing methods. We show through a series of experiments that the task-specific performance of AprAD is comparable to methods that do not distort the output distribution, while being much more computationally efficient.

AAAI Conference 2023 Conference Paper

AudioEar: Single-View Ear Reconstruction for Personalized Spatial Audio

  • Xiaoyang Huang
  • Yanjun Wang
  • Yang Liu
  • Bingbing Ni
  • Wenjun Zhang
  • Jinxian Liu
  • Teng Li

Spatial audio, which focuses on immersive 3D sound rendering, is widely applied in the acoustic industry. One of the key problems of current spatial audio rendering methods is the lack of personalization based on different anatomies of individuals, which is essential to produce accurate sound source positions. In this work, we address this problem from an interdisciplinary perspective. The rendering of spatial audio is strongly correlated with the 3D shape of human bodies, particularly ears. To this end, we propose to achieve personalized spatial audio by reconstructing 3D human ears with single-view images. First, to benchmark the ear reconstruction task, we introduce AudioEar3D, a high-quality 3D ear dataset consisting of 112 point cloud ear scans with RGB images. To self-supervisedly train a reconstruction model, we further collect a 2D ear dataset composed of 2,000 images, each one with manual annotation of occlusion and 55 landmarks, named AudioEar2D. To our knowledge, both datasets have the largest scale and best quality of their kinds for public use. Further, we propose AudioEarM, a reconstruction method guided by a depth estimation network that is trained on synthetic data, with two loss functions tailored for ear data. Lastly, to fill the gap between the vision and acoustics community, we develop a pipeline to integrate the reconstructed ear mesh with an off-the-shelf 3D human body and simulate a personalized Head-Related Transfer Function (HRTF), which is the core of spatial audio rendering. Code and data are publicly available in https://github.com/seanywang0408/AudioEar.

ICLR Conference 2022 Conference Paper

Representation-Agnostic Shape Fields

  • Xiaoyang Huang
  • Jiancheng Yang
  • Yanjun Wang
  • Ziyu Chen
  • Linguo Li
  • Teng Li 0001
  • Bingbing Ni
  • Wenjun Zhang 0001

3D shape analysis has been widely explored in the era of deep learning. Numerous models have been developed for various 3D data representation formats, e.g., MeshCNN for meshes, PointNet for point clouds and VoxNet for voxels. In this study, we present Representation-Agnostic Shape Fields (RASF), a generalizable and computation-efficient shape embedding module for 3D deep learning. RASF is implemented with a learnable 3D grid with multiple channels to store local geometry. Based on RASF, shape embeddings for various 3D shape representations (point clouds, meshes and voxels) are retrieved by coordinate indexing. While there are multiple ways to optimize the learnable parameters of RASF, we provide two effective schemes among all in this paper for RASF pre-training: shape reconstruction and normal estimation. Once trained, RASF becomes a plug-and-play performance booster with negligible cost. Extensive experiments on diverse 3D representation formats, networks and applications, validate the universal effectiveness of the proposed RASF. Code and pre-trained models are publicly available\footnote{\url{https://github.com/seanywang0408/RASF}}.

ICRA Conference 2021 Conference Paper

An MR Safe Rotary Encoder Based on Eccentric Sheave and FBG Sensors

  • Shaoping Huang
  • Anzhu Gao
  • Zicong Wu
  • Chuqian Lou
  • Yanjun Wang
  • Guang-Zhong Yang

MRI-guided robotic systems are emerging platforms for minimally invasive intervention because of high positioning accuracy and excellent tissue contrast. MR safe encoders are critical components for closed-loop robotic control. This paper develops an MR safe absolute rotary encoder based on eccentric sheave and FBG sensors. The eccentric sheave transforms the rotational motion of the shaft to the bending deflection of the beam on which FBG sensors are integrated. A model is built by establishing the relationship of the kinematics of the sheave, the mechanical properties of the beam with unknown length, and the strain model of two Fiber Bragg Grating (FBG) sensors. A Pseudo-Rigid Body (PRB) 3R model is used to solve a set of constrained equations for accurate rotary encoding. A prototype is built to calibrate the parameters and validate the accuracy of the encoder and its MR compatibility. Results show that the maximum angular error is 1. 6°, and the RMS error is 0. 46°. MRI shows that no noticeable artifacts are observed, and the Signal to Noise Ratio (SNR) is not affected. The results demonstrate the potential of the proposed method for it to be integrated with MR safe robots with easy fabrication, compact structures, and continuous measurement.

v2026.09.13