Arrow Research search

Author name cluster

Yifan Xie

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

8 papers
2 author rows

Possible papers

8

EAAI Journal 2026 Journal Article

Leveraging high-frequency intraday pattern for improved deep temporal learning of stock price dynamics

  • Jiaqi He
  • Yifan Xie
  • Xinyuan Song
  • Hanwen Ning

High-frequency intraday data is crucial in financial modeling, but its high dimensionality and noise present significant challenges for pattern recognition and effective use in machine learning. In this paper, motivated by the stylized characteristics of intraday trading, we propose a novel dimension reduction method called price–volume balanced functional principal component analysis (PVB-FPCA) in terms of cumulative intraday return (CIDR) curves. Through PVB-FPCA, we make a key discovery: when crucial volume–price characteristics are incorporated, high-frequency intraday data reveals a remarkably robust and significant intraday trend pattern, described by the first functional principal component. By exploiting this pattern, we develop a network module that generates efficient low-dimensional representations of high-frequency intraday data. This approach addresses the challenges of high dimensionality and noise, enhancing the utility of high-frequency data. Building upon the module, we develop a novel deep temporal forecasting model, termed PVB-IntraFusionNet. We design a dual-channel temporal convolutional network (TCN) to extract low-frequency temporal price–volume features. These features are subsequently fused with high-frequency features via a cross multi-head attention mechanism to capture multi-scale dependencies. Through a carefully designed projection module, the fused features are utilized to generate dynamic forecasting results. Compared to state-of-the-arts, PVB-IntraFusionNet can efficiently leverage both low- and high-frequency patterns simultaneously, resulting in significantly improved forecasting performance. Through PVB-IntraFusionNet, this paper explores an effective approach for developing more accurate deep temporal learning models with high-frequency intraday data, demonstrating the potential of our method in addressing various financial learning problems. The numerical results on prominent NYSE/NASDAQ stocks further illustrate the efficiency and advantages of our approach.

IROS Conference 2025 Conference Paper

GaussianPU: Color Point Cloud Upsampling via 3D Gaussian Splatting

  • Zixuan Guo
  • Yifan Xie
  • Weijing Xie
  • Peng Huang
  • Chenyang Wang 0001
  • Fei Ma 0006
  • F. Richard Yu

Dense colored point clouds enhance visual perception and are of significant value in various robotic applications. However, existing learning-based point cloud upsampling methods are constrained by computational resources and batch processing strategies, which often require subdividing point clouds into smaller patches, leading to distortions that degrade perceptual quality. To address this challenge, we propose a novel 2D-3D hybrid colored point cloud upsampling framework (GaussianPU) based on 3D Gaussian Splatting (3DGS) for robotic perception. This approach leverages 3DGS to bridge 3D point clouds with their 2D rendered images in robot vision systems. A dual scale rendered image restoration network transforms sparse point cloud renderings into dense representations, which are then input into 3DGS along with precise robot camera poses and interpolated sparse point clouds to reconstruct dense 3D point clouds. We have made a series of enhancements to the vanilla 3DGS, enabling precise control over the number of points and significantly boosting the quality of the upsampled point cloud for robotic scene understanding. Our framework supports processing entire point clouds on a single consumer-grade GPU, eliminating the need for segmentation and thus producing high-quality, dense colored point clouds with millions of points for robot navigation and manipulation tasks. Extensive experimental results on generating million-level point cloud data validate the effectiveness of our method, substantially improving the quality of colored point clouds and demonstrating significant potential for applications involving large-scale point clouds in autonomous robotics and human-robot interaction scenarios.

ICRA Conference 2025 Conference Paper

Multiple Rotation Averaging with Constrained Reweighting Deep Matrix Factorization

  • Shiqi Li
  • Jihua Zhu
  • Yifan Xie
  • Naiwen Hu
  • Mingchen Zhu
  • Zhongyu Li 0002
  • Di Wang 0006
  • Huimin Lu 0001

Multiple rotation averaging plays a crucial role in computer vision and robotics domains. The conventional optimization-based methods optimize a nonlinear cost function based on certain noise assumptions, while most previous learning-based methods require ground truth labels in the supervised training process. Recognizing the handcrafted noise assumption may not be reasonable in all real-world scenarios, this paper proposes an effective rotation averaging method for mining data patterns in a learning manner while avoiding the requirement of labels. Specifically, we apply deep matrix factorization to directly solve the multiple rotation averaging problem in free linear space. For deep matrix factorization, we design a neural network model, which is explicitly low-rank and symmetric to better suit the background of multiple rotation averaging. Meanwhile, we utilize a spanning tree-based edge filtering to suppress the influence of rotation outliers. What's more, we also adopt a reweighting scheme and dynamic depth selection strategy to further improve the robustness. Our method synthesizes the merit of both optimization-based and learning-based methods. Experimental results on various datasets validate the effectiveness of our proposed method.

IROS Conference 2025 Conference Paper

Observation-Graph Interaction and Key-Detail Guidance for Vision and Language Navigation

  • Yifan Xie
  • Binkai Ou
  • Fei Ma 0006
  • Yaohua Liu

Vision and Language Navigation (VLN) requires an agent to navigate through environments following natural language instructions. However, existing methods often struggle with effectively integrating visual observations and instruction details during navigation, leading to suboptimal path planning and limited success rates. In this paper, we propose OIKG (Observation-graph Interaction and Key-detail Guidance), a novel framework that addresses these limitations through two key components: (1) an observation-graph interaction module that decouples angular and visual information while strengthening edge representations in the navigation space, and (2) a key-detail guidance module that dynamically extracts and utilizes fine-grained location and object information from instructions. By enabling more precise cross-modal alignment and dynamic instruction interpretation, our approach significantly improves the agent’s ability to follow complex navigation instructions. Extensive experiments on the R2R and RxR datasets demonstrate that OIKG achieves state-of-the-art performance across multiple evaluation metrics, validating the effectiveness of our method in enhancing navigation precision through better observation-instruction alignment.

AAAI Conference 2025 Conference Paper

PointTalk: Audio-Driven Dynamic Lip Point Cloud for 3D Gaussian-based Talking Head Synthesis

  • Yifan Xie
  • Tao Feng
  • Xin Zhang
  • Xiangyang Luo
  • Zixuan Guo
  • Weijiang Yu
  • Heng Chang
  • Fei Ma

Talking head synthesis with arbitrary speech audio is a crucial challenge in the field of digital humans. Recently, methods based on radiance fields have received increasing attention due to their ability to synthesize high-fidelity and identity-consistent talking heads from just a few minutes of training video. However, due to the limited scale of the training data, these methods often exhibit poor performance in audio-lip synchronization and visual quality. In this paper, we propose a novel 3D Gaussian-based method called PointTalk, which constructs a static 3D Gaussian field of the head and deforms it in sync with the audio. It also incorporates an audio-driven dynamic lip point cloud as a critical component of the conditional information, thereby facilitating the effective synthesis of talking heads. Specifically, the initial step involves generating the corresponding lip point cloud from the audio signal and capturing its topological structure. The design of the dynamic difference encoder aims to capture the subtle nuances inherent in dynamic lip movements more effectively. Furthermore, we integrate the audio-point enhancement module, which not only ensures the synchronization of the audio signal with the corresponding lip point cloud within the feature space, but also facilitates a deeper understanding of the interrelations among cross-modal conditional features. Extensive experiments demonstrate that our method achieves superior high-fidelity and audio-lip synchronization in talking head synthesis compared to previous methods.

NeurIPS Conference 2025 Conference Paper

Universal Visuo-Tactile Video Understanding for Embodied Interaction

  • Yifan Xie
  • Mingyang Li
  • Shoujie Li
  • Xingting Li
  • Guangyu Chen
  • Fei Ma
  • Fei Yu
  • Wenbo Ding

Tactile perception is essential for embodied agents to understand the physical attributes of objects that cannot be determined through visual inspection alone. While existing methods have made progress in visual and language modalities for physical understanding, they fail to effectively incorporate tactile information that provides crucial haptic feedback for real-world interaction. In this paper, we present VTV-LLM, the first multi-modal large language model that enables universal Visuo-Tactile Video (VTV) understanding, bridging the gap between tactile perception and natural language. To address the challenges of cross-sensor and cross-modal integration, we contribute VTV150K, a comprehensive dataset comprising 150, 000 video frames from 100 diverse objects captured across three different tactile sensors (GelSight Mini, DIGIT, and Tac3D), annotated with four fundamental tactile attributes (hardness, protrusion, elasticity, and friction). We develop a novel three-stage training paradigm that includes VTV enhancement for robust visuo-tactile representation, VTV-text alignment for cross-modal correspondence, and text prompt finetuning for natural language generation. Our framework enables sophisticated tactile reasoning capabilities including feature assessment, comparative analysis, and scenario-based decision-making. Extensive experimental evaluations demonstrate that VTV-LLM achieves superior performance in tactile reasoning tasks, establishing a foundation for more intuitive human-machine interaction in tactile domains.

IROS Conference 2024 Conference Paper

Hyperbolic Image-and-Pointcloud Contrastive Learning for 3D Classification

  • Naiwen Hu
  • Haozhe Cheng
  • Yifan Xie
  • Pengcheng Shi
  • Jihua Zhu

3D contrastive representation learning has exhibited remarkable efficacy across various downstream tasks. However, existing contrastive learning paradigms based on cosine similarity fail to deeply explore the potential intra-modal hierarchical and cross-modal semantic correlations about multi-modal data in Euclidean space. In response, we seek solutions in hyperbolic space and propose a hyperbolic image-and-pointcloud contrastive learning method (HyperIPC). For the intra-modal branch, we rely on the intrinsic geometric structure to explore the hyperbolic embedding representation of point cloud to capture invariant features. For the cross-modal branch, we leverage images to guide the point cloud in establishing strong semantic hierarchical correlations. Empirical experiments underscore the outstanding classification performance of HyperIPC. Notably, HyperIPC enhances object classification results by 2. 8% and few-shot classification outcomes by 5. 9% on ScanObjectNN compared to the baseline. Furthermore, ablation studies and confirmatory testing validate the rationality of HyperIPC’s parameter settings and the effectiveness of its submodules.

TCS Journal 2023 Journal Article

The search and rescue game on a cycle

  • Thomas Lidbetter
  • Yifan Xie

We consider a search and rescue game introduced recently by the first author. An immobile target or targets (for example, injured hikers) are hidden on a graph. The terrain is assumed to be dangerous, so that when any given vertex of the graph is searched, there is a certain probability that the search will come to an end, otherwise with the complementary success probability the search can continue. A Searcher searches the graph with the aim of finding all the targets with maximum probability. Here, we focus on the game in the case that the graph is a cycle. In the case that there is only one target, we solve the game for equal success probabilities, and for a class of games with unequal success probabilities. For multiple targets and equal success probabilities, we give a solution for an adaptive Searcher and a solution in a special case for a non-adaptive Searcher. We also consider a continuous version of the model, giving a full solution for an adaptive Searcher and approximately optimal solutions in the non-adaptive case.

v2026.09.13