Arrow Research search

Author name cluster

Ying He

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

28 papers
2 author rows

Possible papers

28

AAAI Conference 2026 Conference Paper

AquaSplatting: A Hybrid 3D Representation for Robust Underwater Scene Reconstruction via Dual-Branch Rendering

  • Jiangbei Hu
  • Haobo Wang
  • Baixin Xu
  • Nan Ding
  • Zhimao Lu
  • Na Lei
  • Ying He

While 3D Gaussian Splatting (3DGS) excels at real-time rendering of standard scenes, it struggles to reconstruct underwater environments due to severe challenges such as light scattering, color attenuation, and sparse coverage of Gaussian kernels in far-field aqueous regions. To address this, we introduce AquaSplatting, a hybrid framework that combines explicit and implicit modeling methods for robust underwater scene reconstruction. Our dual-branch architecture employs 3DGS in a geometry-guided branch to model solid surfaces like the seabed, while a medium-aware branch uses a compact, view-dependent MLP to represent volumetric water effects. Furthermore, a neural underwater hybrid rendering mechanism adaptively fuses these two representations based on accumulated opacity. Thanks to this dual-branch framework, our method can also synthesize restored images without water medium. To enhance efficiency, our proposed engagement-based pruning (EBP) strategy quantifies each Gaussian's contribution by accumulating its image-space gradients over multiple frames, enabling the principled removal of primitives with negligible impact. The entire framework is optimized using a comprehensive loss function that integrates photometric, exposure, semantic, and depth priors to maximize visual fidelity. Experiments on challenging underwater datasets demonstrate that AquaSplatting achieves the state-of-the-art in reconstruction quality surpassing prior methods while maintaining real-time performance.

AAAI Conference 2026 Conference Paper

ForeDiffusion: Foresight-Conditioned Diffusion Policy via Future View Construction for Robot Manipulation

  • Weize Xie
  • Yi Ding
  • Ying He
  • Leilei Wang
  • Binwen Bai
  • Zheyi Zhao
  • Chenyang Wang
  • F. Richard Yu

Diffusion strategies have advanced visual motor control by progressively denoising high-dimensional action sequences, providing a promising method for robot manipulation. However, as task complexity increases, the success rate of existing baseline models decreases considerably. Analysis indicates that current diffusion strategies are confronted with two limitations. First, these strategies only rely on short-term observations as conditions. Second, the training objective remains limited to a single denoising loss, which leads to error accumulation and causes grasping deviations. To address these limitations, this paper proposes Foresight-Conditioned Diffusion (ForeDiffusion), by injecting the predicted future view representation into the diffusion process. As a result, the policy is guided to be forward-looking, enabling it to correct trajectory deviations. Following this design, ForeDiffusion employs a dual loss mechanism, combining the traditional denoising loss and the consistency loss of future observations, to achieve the unified optimization. Extensive evaluation on the Adroit suite and the MetaWorld benchmark demonstrates that ForeDiffusion achieves an average success rate of 80% for the overall task, significantly outperforming the existing mainstream diffusion methods by approximately 20% in high difficulty tasks, while maintaining more stable performance across the entire tasks.

AAAI Conference 2026 Conference Paper

Guiding Point Cloud Denoising with Learned Structural Priors

  • Chuchen Guo
  • Zheng Liu
  • Ying He

Recovering precise surface geometry from corrupted point clouds remains a core challenge in 3D vision. Although existing denoising techniques achieve remarkable success, balancing noise removal with preserving intricate geometric details continues to pose difficulties. A critical limitation of current methods is that their adaptive feature aggregation mechanisms rely heavily on intermediate network features that have not been explicitly regularized, resulting in unstable guidance signals. This instability restricts the capability of the network to optimally differentiate true geometric details from noise. To overcome this limitation, we propose a novel deep learning framework that explicitly learns structured representations as robust priors to guide feature refinement. Our approach first derives a set of representative local structural primitives from input features by means of a learned codebook. This learned structured representation then serves as a robust conditional signal, directing a subsequent feature fusion mechanism to dynamically aggregate information in a structure-aware manner, thereby more effectively discerning noise and meticulously reconstructing geometric details. Extensive experiments on several benchmarks have demonstrated the superiority of our framework over existing advanced techniques in terms of detail preservation and noise suppression.

AAAI Conference 2026 Conference Paper

LiteGE: Lightweight Geodesic Embedding for Efficient Geodesics Computation and Non-Isometric Shape Correspondence

  • Yohanes Yudhi Adikusuma
  • Qixing Huang
  • Ying He

Computing geodesic distances on 3D surfaces is fundamental to many tasks in 3D vision and geometry processing, with deep connections to tasks such as shape correspondence. Recent learning-based methods achieve strong performance but rely on large 3D backbones, leading to high memory usage and latency, which limit their use in interactive or resource-constrained settings. We introduce LiteGE, a lightweight approach that constructs compact, category-aware shape descriptors by applying PCA to unsigned distance field (UDFs) samples at informative voxels. This descriptor is efficient to compute and removes the need for high-capacity networks. LiteGE remains robust on sparse point clouds, supporting inputs with as few as 300 points, where prior methods fail. Extensive experiments show that LiteGE reduces memory usage and inference time by up to 300x compared to existing neural approaches. In addition, by exploiting the intrinsic relationship between geodesic distance and shape correspondence, LiteGE enables fast and accurate shape matching. Our method achieves up to 1000x speedup over state-of-the-art mesh-based approaches while maintaining comparable accuracy on non-isometric shape pairs, including evaluations on point-cloud inputs.

AAAI Conference 2026 Conference Paper

MonoCloth: Reconstruction and Animation of Cloth-Decoupled Human Avatars from Monocular Videos

  • Daisheng Jin
  • Ying He

Reconstructing realistic 3D human avatars from monocular videos is a challenging task due to the limited geometric information and complex non-rigid motion involved. We present MonoCloth, a new method for reconstructing and animating clothed human avatars from monocular videos. To overcome the limitations of monocular input, we introduce a part-based decomposition strategy that separates the avatar into body, face, hands, and clothing. This design reflects the varying levels of reconstruction difficulty and deformation complexity across these components. Specifically, we focus on detailed geometry recovery for the face and hands. For clothing, we propose a dedicated cloth simulation module that captures garment deformation using temporal motion cues and geometric constraints. Experimental results demonstrate that MonoCloth improves both visual reconstruction quality and animation realism compared to existing methods. Furthermore, thanks to its part-based design, MonoCloth also supports additional tasks such as clothing transfer, underscoring its versatility and practical utility.

YNIMG Journal 2026 Journal Article

Morphometric dissimilarity in association cortices linked to autism subtype with more severe symptoms

  • Hongxiu Jiang
  • Raul Rodriguez-Cruces
  • Ke Xie
  • Valeria Kebets
  • Yezhou Wang
  • Clara F. Weber
  • Ying He
  • Jonah Kember

Autism spectrum disorder (ASD) is a prevalent and heterogeneous neurodevelopmental condition marked by atypical brain connectivity. Understanding ASD neural subtypes at the network level is critical for clarifying its neuroanatomical heterogeneity. Morphometric similarity networks (MSNs), derived from region-to-region similarity across multiple anatomical features, offer a powerful approach for capturing individual-level neural architecture. In this study, MSNs were estimated from seven anatomical features in 348 individuals with ASD and 452 typically developing (TD) controls. Across all ASD participants, the first principal component of MSN values was negatively correlated with social and communication severity. Three ASD subtypes with distinct MSN patterns were identified. Subtype-1, characterized by weaker morphometric similarity values in frontotemporal association regions compared to TD individuals, exhibited the most severe symptoms in social, communication and repetitive behaviors, and displayed hyperconnectivity between the salience and visual networks, and between language and visual networks. Subtype-2 showed greater values of morphometric similarities than TD and less severe social symptoms compared to subtype-1, along with hyperconnectivity between default and salience networks relative to TD. Subtype-3 displayed morphometric similarity values largely comparable to TD and the least severe symptoms out of the three subtypes. Transcriptomic analysis revealed that GABAergic parvalbumin and glutamatergic intratelencephalic-projecting neurons were key cell types differentiating subtypes. These findings suggest the existence of distinct ASD neuroanatomical subtypes defined by regional morphometric similarity, each linked to unique behavioral, functional, and transcriptomic profiles. Morphometric dissimilarity in association regions may serve as a neural signature for ASD subtypes characterized by more severe clinical manifestations.

AAAI Conference 2026 Conference Paper

PSPO: Prompt-Level Prioritization and Experience-Weighted Smoothing for Efficient Policy Optimization

  • Xinxin Zhu
  • Ying He
  • Haowen Hou
  • Ruichong Zhang
  • Nianbo Zeng
  • Yulin Peng
  • Jiongfeng Fang
  • F. Richard Yu

Reinforcement Fine-tuning (RFT) methods such as Group Relative Policy Optimization (GRPO) have demonstrated strong capabilities in aligning Large Language Models with human preferences. However, these approaches often suffer from limited data efficiency, necessitating extensive on-policy rollouts to maintain competitive performance. We propose PSPO (Prompt-Level Prioritization and Experience-Weighted Smoothing for Efficient Policy Optimization), a lightweight yet effective enhancement to GRPO that improves training stability and sample efficiency through two complementary techniques. First, we introduce an experience-weighted reward smoothing mechanism, which uses exponential moving averages to track group-level reward statistics for each prompt. This enables more stable advantage estimation across training steps without storing entire trajectories, allowing the model to capture historical reward trends in a lightweight and memory-efficient manner. Second, we adopt a prompt-level prioritized sampling strategy, which is an online data selection method inspired by prioritized experience replay. It dynamically emphasizes higher-impact prompts based on their relative advantages, thereby improving data efficiency. Experiments on multiple mathematical reasoning benchmarks and models show that PSPO achieves comparable or better accuracy than GRPO, while significantly accelerating convergence, and maintaining low computational and memory overhead.

EAAI Journal 2026 Journal Article

Restoring neural radiance fields performance under adverse weather conditions

  • Ying He
  • Gan Chen
  • F. Richard Yu
  • Ming Li
  • Fei Ma
  • Guang Zhou

Neural Radiance Fields (NeRFs) have emerged as a powerful paradigm for modeling complex, photorealistic three-dimensional environments, garnering significant attention in the field of scene-based robotic localization. Regrettably, environments encountered in robotic applications are frequently susceptible to adverse weather conditions (e. g. , rain, snow, fog). Under such conditions, the inherent quality degradation introduced by existing image restoration algorithms severely disrupts the spatial consistency reconstruction of NeRFs. To address this challenge, this paper proposes a novel methodology for three-dimensional Scene Reconstruction under adverse weather, termed WeatherNeRF, which seamlessly integrates an image restoration algorithm with the neural radiance field framework. The proposed method effectively leverages the image restoration algorithm to process input image sequences affected by inclement weather. To mitigate the introduction of invalid artifacts by these processed sequences during scene reconstruction, we employ two regularization functions specifically designed to enhance scene compactness. Furthermore, to bolster scene consistency and facilitate effective scene restoration, we incorporate two-dimensional prior knowledge extracted from an image restoration model—WeatherDiffusion—during the reconstruction process, utilizing Score Distillation Sampling (SDS). Comprehensive experimental evaluations demonstrate that the proposed WeatherNeRF framework effectively restores neural radiance fields in everyday scenes degraded by adverse weather conditions and is capable of synthesizing high-fidelity novel view images. The code and data are publicly available at https: //github. com/C2022G/WeatherNeRF.

AAAI Conference 2025 Conference Paper

3DMambaIPF: A State Space Model for Iterative Point Cloud Filtering via Differentiable Rendering

  • Qingyuan Zhou
  • Weidong Yang
  • Ben Fei
  • Jingyi Xu
  • Rui Zhang
  • Keyi Liu
  • Yeqi Luo
  • Ying He

Noise is an inevitable aspect of point cloud acquisition, necessitating filtering as a fundamental task within the realm of 3D vision. Existing learning-based filtering methods have shown promising capabilities on commonly used datasets. Nonetheless, the effectiveness of these methods is constrained when dealing with a substantial quantity of point clouds. This limitation primarily stems from their limited denoising capabilities for dense and large-scale point clouds and their inclination to generate noisy outliers after denoising. To deal with this challenge, we introduce 3DMambaIPF, for the first time, exploiting Selective State Space Models (SSMs) architecture to handle highly-dense and large-scale point clouds, capitalizing on its strengths in selective input processing and large context modeling capabilities. Additionally, we present a robust and fast differentiable rendering loss to constrain the noisy points around the surface. In contrast to previous methodologies, this differentiable rendering loss enhances the visual realism of denoised geometric structures and aligns point cloud boundaries more closely with those observed in real-world objects. Extensive evaluations on commonly used datasets (typically with up to 50K points) demonstrate that 3DMambaIPF achieves state-of-the-art results. Moreover, we showcase the superior scalability and efficiency of 3DMambaIPF on highly dense and large-scale point clouds with up to 500K points compared to off-the-shelf methods.

AAAI Conference 2025 Conference Paper

Details Enhancement in Unsigned Distance Field Learning for High-fidelity 3D Surface Reconstruction

  • Cheng Xu
  • Fei Hou
  • Wencheng Wang
  • Hong Qin
  • Zhebin Zhang
  • Ying He

While Signed Distance Fields (SDF) are well-established for modeling watertight surfaces, Unsigned Distance Fields (UDF) broaden the scope to include open surfaces and models with complex inner structures. Despite their flexibility, UDFs encounter significant challenges in high-fidelity 3D reconstruction, such as non-differentiability at the zero level set, difficulty in achieving the exact zero value, numerous local minima, vanishing gradients, and oscillating gradient directions near the zero level set. To address these challenges, we propose Details Enhanced UDF (DEUDF) learning that integrates normal alignment and the SIREN network for capturing fine geometric details, adaptively weighted Eikonal constraints to address vanishing gradients near the target surface, unconditioned MLP-based UDF representation to relax non-negativity constraints, and DCUDF for extracting the local minimal average distance surface. These strategies collectively stabilize the learning process from unoriented point clouds and enhance the accuracy of UDFs. Our computational results demonstrate that DEUDF outperforms existing UDF learning methods in both accuracy and the quality of reconstructed surfaces.

AAAI Conference 2025 Conference Paper

DMF-Net: Image-Guided Point Cloud Completion with Dual-Channel Modality Fusion and Shape-Aware Upsampling Transformer

  • Aihua Mao
  • Yuxuan Tang
  • Jiangtao Huang
  • Ying He

In this paper we study the task of a single-view image-guided point cloud completion. Existing methods have got promising results by fusing the information of image into point cloud explicitly or implicitly. However, given that the image has global shape information and the partial point cloud has rich local details, We believe that both modalities need to be given equal attention when performing modality fusion. To this end, we propose a novel dual-channel modality fusion network for image-guided point cloud completion(named DMF-Net), in a coarse-to-fine manner. In the first stage, DMF-Net takes a partial point cloud and corresponding image as input to recover a coarse point cloud. In the second stage, the coarse point cloud will be upsampled twice with shape-aware upsampling transformer to get the dense and complete point cloud. Extensive quantitative and qualitative experimental results show that DMF-Net outperforms the state-of-the-art unimodal and multimodal point cloud completion works on ShapeNet-ViPC dataset.

AAAI Conference 2025 Conference Paper

Do Not DeepFake Me: Privacy-Preserving Neural 3D Head Reconstruction Without Sensitive Images

  • Jiayi Kong
  • Xurui Song
  • Shuo Huai
  • Baixin Xu
  • Jun Luo
  • Ying He

While 3D head reconstruction is widely used for modeling, existing neural reconstruction approaches rely on high-resolution multi-view images, posing notable privacy issues. Individuals are particularly sensitive to facial features, and facial image leakage can lead to many malicious activities, such as unauthorized tracking and deepfake. In contrast, geometric data is less susceptible to misuse due to its complex processing requirements, and absence of facial texture features. In this paper, we propose a novel two-stage 3D facial reconstruction method aimed at avoiding exposure to sensitive facial information while preserving detailed geometric accuracy. Our approach first uses non-sensitive rear-head images for initial geometry and then refines this geometry using processed privacy-removed gradient images. Extensive experiments show that the resulting geometry is comparable to methods using full images, while the process is resistant to DeepFake applications and facial recognition (FR) systems, thereby proving its effectiveness in privacy protection.

IS Journal 2025 Journal Article

Facilitating Autonomous Driving Tasks With Large Language Models

  • Mengyao Wu
  • F. Richard Yu
  • Peter Xiaoping Liu
  • Ying He

We explore how large language models (LLMs) can expedite and automate the learning process for autonomous driving tasks. This involves harnessing LLM knowledge to shape a learning framework and utilizing LLMs to guide the learning process. We conduct a case study to demonstrate LLMs’ ability to export driving rules. LLM outputs may not be entirely reliable for the direct handling of driving decisions due to potential inaccuracies and inconsistencies. To address these issues, we propose integrating LLM knowledge with statistical learning. This enables LLMs to export task-specific knowledge as symbolic rules, forming the initial learning structure. Rule weights are calculated based on statistical salience derived from training data, resulting in a set of weighted rules for robust decision making. Furthermore, this set of weighted rules preserves strong semantics, allowing LLMs to comprehend and make modifications based on varying needs. Simulations using a highway driving simulator validate the effectiveness of our approach.

IJCAI Conference 2025 Conference Paper

Inter3D: A Benchmark and Strong Baseline for Human-Interactive 3D Object Reconstruction

  • Gan Chen
  • Ying He
  • Mulin Yu
  • F. Richard Yu
  • Gang Xu
  • Fei Ma
  • Ming Li
  • Guang Zhou

Recent advancements in implicit 3D reconstruction methods, e. g. , neural rendering fields and Gaussian splatting, have primarily focused on novel view synthesis of static or dynamic objects with continuous motion states. However, these approaches struggle to efficiently model a human-interactive object with n movable parts, requiring 2^n separate models to represent all discrete states. To overcome this limitation, we propose Inter3D, a new benchmark and approach for novel state synthesis of human-interactive objects. We introduce a self-collected dataset featuring commonly encountered interactive objects and a new evaluation pipeline, where only individual part states are observed during training, while part combination states remain unseen. We also propose a strong baseline approach that leverages Space Discrepancy Tensors to efficiently modelling all states of an object. To alleviate the impractical constraints on camera trajectories across training states, we propose a Mutual State Regularization mechanism to enhance the spatial density consistency of movable parts. In addition, we explore two occupancy grid sampling strategies to facilitate training efficiency. We conduct extensive experiments on the proposed benchmark, showcasing the challenges of the task and the superiority of our approach. The code and data are publicly available at https: //github. com/Inter3D-ui/Inter3D.

IROS Conference 2025 Conference Paper

JAM: Keypoint-Guided Joint Prediction after Classification-Aware Marginal Proposal for Multi-Agent Interaction

  • Fangze Lin
  • Ying He
  • Fei Yu
  • Hong Zhang

Predicting the future motion of road participants is a critical task in autonomous driving. In this work, we address the challenge of low-quality generation of low-probability modes in multi-agent joint prediction. To tackle this issue, we propose a two-stage multi-agent interactive prediction framework named keypoint-guided joint prediction after classification-aware marginal proposal (JAM). The first stage is modeled as a marginal prediction process, which classifies queries by trajectory type to encourage the model to learn all categories of trajectories, providing comprehensive mode information for the joint prediction module. The second stage is modeled as a joint prediction process, which takes the scene context and the marginal proposals from the first stage as inputs to learn the final joint distribution. We explicitly introduce key waypoints to guide the joint prediction module in better capturing and leveraging the critical information from the initial predicted trajectories. We conduct extensive experiments on the real-world Waymo Open Motion Dataset interactive prediction benchmark. The results show that our approach achieves competitive performance. In particular, in the framework comparison experiments, the proposed JAM outperforms other prediction frameworks and achieves state-of-the-art performance in interactive trajectory prediction. The code is available at https://github.com/LinFunster/JAM to facilitate future research.

NeurIPS Conference 2025 Conference Paper

MIND: Material Interface Generation from UDFs for Non-Manifold Surface Reconstruction

  • Xuhui Chen
  • Fei Hou
  • Wencheng Wang
  • Hong Qin
  • Ying He

Unsigned distance fields (UDFs) are widely used in 3D deep learning due to their ability to represent shapes with arbitrary topology. While prior work has largely focused on learning UDFs from point clouds or multi-view images, extracting meshes from UDFs remains challenging, as the learned fields rarely attain exact zero distances. A common workaround is to reconstruct signed distance fields (SDFs) locally from UDFs to enable surface extraction via Marching Cubes. However, this often introduces topological artifacts such as holes or spurious components. Moreover, local SDFs are inherently incapable of representing non-manifold geometry, leading to complete failure in such cases. To address this gap, we propose MIND ($\mathrm{\underline{M}aterial}$ $\mathrm{\underline{I}nterface}$ $\mathrm{from}$ $\mathrm{\underline{N}on}$-$\mathrm{manifold}$ $\mathrm{\underline{D}istance}$ $\mathrm{fields}$), a novel algorithm for generating material interfaces directly from UDFs, enabling non-manifold mesh extraction from a global perspective. The core of our method lies in deriving a meaningful spatial partitioning from the UDF, where the target surface emerges as the interface between distinct regions. We begin by computing a two-signed local field to distinguish the two sides of manifold patches, and then extend this to a multi-labeled global field capable of separating all sides of a non-manifold structure. By combining this multi-labeled field with the input UDF, we construct material interfaces that support non-manifold mesh extraction via a multi-labeled Marching Cubes algorithm. Extensive experiments on UDFs generated from diverse data sources, including point cloud reconstruction, multi-view reconstruction, and medial axis transforms, demonstrate that our approach robustly handles complex non-manifold surfaces and significantly outperforms existing methods. The source code is available at https: //github. com/jjjkkyz/MIND.

NeurIPS Conference 2025 Conference Paper

ReDit: Reward Dithering for Improved LLM Policy Optimization

  • Chenxing Wei
  • Jiarui Yu
  • Ying He
  • Hande Dong
  • Yao Shu
  • Fei Yu

DeepSeek-R1 has successfully enhanced Large Language Model (LLM) reasoning capabilities through its rule-based reward system. While it's a ''perfect'' reward system that effectively mitigates reward hacking, such reward functions are often discrete. Our experimental observations suggest that discrete rewards can lead to gradient anomaly, unstable optimization, and slow convergence. To address this issue, we propose ReDit (Reward Dithering), a method that dithers the discrete reward signal by adding simple random noise. With this perturbed reward, exploratory gradients are continuously provided throughout the learning process, enabling smoother gradient updates and accelerating convergence. The injected noise also introduces stochasticity into flat reward regions, encouraging the model to explore novel policies and escape local optima. Experiments across diverse tasks demonstrate the effectiveness and efficiency of ReDit. On average, ReDit achieves performance comparable to vanilla GRPO with only approximately 10% the training steps, and furthermore, still exhibits a 4% performance improvement over vanilla GRPO when trained for a similar duration. Visualizations confirm significant mitigation of gradient issues with ReDit. Moreover, theoretical analyses are provided to further validate these advantages.

AAAI Conference 2025 Conference Paper

You Should Learn to Stop Denoising on Point Clouds in Advance

  • Chuchen Guo
  • Weijie Zhou
  • Zheng Liu
  • Ying He

Point clouds have become the preferred data format for a variety of tasks in 3D vision and graphics. However, raw point clouds often contain significant noise. This paper introduces the Adaptive Stop Denoising Network (ASDN), a novel approach aimed at restoring high-quality point clouds from noisy data. Our method is built upon a pivotal observation: during the denoising phase, high-noise points draw more focus from the network, which may suppress the points that have already been effectively denoised. This observation has led us to develop an adaptive strategy that ceases denoising already cleaned points to prevent over-denoising, while continuing to refine points that remain noisy. We employ a U-Net architecture complemented by an adaptive classifier, which utilizes a recoverability factor to assess the completion of denoising and make dynamic decisions about when to halt the process. Our method not only demonstrates superior noise removal efficiency but also preserves geometric details more effectively, reducing over- or under-denoising artifacts. Extensive experiments and evaluations demonstrate that our method outperforms the state-of-the-art both qualitatively and quantitatively.

IJCAI Conference 2024 Conference Paper

ABM: Attention before Manipulation

  • Fan Zhuo
  • Ying He
  • Fei Yu
  • Pengteng Li
  • Zheyi Zhao
  • Xilong Sun

Vision-language models (VLMs) show promising generalization and zero-shot capabilities, offering a potential solution to the impracticality and cost of enabling robots to comprehend diverse human instructions and scene semantics in the real world. Existing approaches most directly integrate the semantic representations from pre-trained VLMs with policy learning. However, these methods are limited to the labeled data learned, resulting in poor generalization ability to unseen instructions and objects. To address the above limitation, we propose a simple method called "Attention before Manipulation" (ABM), which fully leverages the object knowledge encoded in CLIP to extract information about the target object in the image. It constructs an Object Mask Field, serving as a better representation of the target object for the model to separate visual grounding from action prediction and acquire specific manipulation skills effectively. We train ABM for 8 RLBench tasks and 2 real-world tasks via behavior cloning. Extensive experiments show that our method significantly outperforms the baselines in the zero-shot and compositional generalization experiment settings.

NeurIPS Conference 2024 Conference Paper

Flatten Anything: Unsupervised Neural Surface Parameterization

  • Qijian Zhang
  • Junhui Hou
  • Wenping Wang
  • Ying He

Surface parameterization plays an essential role in numerous computer graphics and geometry processing applications. Traditional parameterization approaches are designed for high-quality meshes laboriously created by specialized 3D modelers, thus unable to meet the processing demand for the current explosion of ordinary 3D data. Moreover, their working mechanisms are typically restricted to certain simple topologies, thus relying on cumbersome manual efforts (e. g. , surface cutting, part segmentation) for pre-processing. In this paper, we introduce the Flatten Anything Model (FAM), an unsupervised neural architecture to achieve global free-boundary surface parameterization via learning point-wise mappings between 3D points on the target geometric surface and adaptively-deformed UV coordinates within the 2D parameter domain. To mimic the actual physical procedures, we ingeniously construct geometrically-interpretable sub-networks with specific functionalities of surface cutting, UV deforming, unwrapping, and wrapping, which are assembled into a bi-directional cycle mapping framework. Compared with previous methods, our FAM directly operates on discrete surface points without utilizing connectivity information, thus significantly reducing the strict requirements for mesh quality and even applicable to unstructured point cloud data. More importantly, our FAM is fully-automated without the need for pre-cutting and can deal with highly-complex topologies, since its learning process adaptively finds reasonable cutting seams and UV boundaries. Extensive experiments demonstrate the universality, superiority, and inspiring potential of our proposed neural surface parameterization paradigm. Our code is available at https: //github. com/keeganhk/FlattenAnything.

NeurIPS Conference 2024 Conference Paper

From Transparent to Opaque: Rethinking Neural Implicit Surfaces with $\alpha$-NeuS

  • Haoran Zhang
  • Junkai Deng
  • Xuhui Chen
  • Fei Hou
  • Wencheng Wang
  • Hong Qin
  • Chen Qian
  • Ying He

Traditional 3D shape reconstruction techniques from multi-view images, such as structure from motion and multi-view stereo, face challenges in reconstructing transparent objects. Recent advances in neural radiance fields and its variants primarily address opaque or transparent objects, encountering difficulties to reconstruct both transparent and opaque objects simultaneously. This paper introduces $\alpha$-NeuS$\textemdash$an extension of NeuS$\textemdash$that proves NeuS is unbiased for materials from fully transparent to fully opaque. We find that transparent and opaque surfaces align with the non-negative local minima and the zero iso-surface, respectively, in the learned distance field of NeuS. Traditional iso-surfacing extraction algorithms, such as marching cubes, which rely on fixed iso-values, are ill-suited for such data. We develop a method to extract the transparent and opaque surface simultaneously based on DCUDF. To validate our approach, we construct a benchmark that includes both real-world and synthetic scenes, demonstrating its practical utility and effectiveness. Our data and code are publicly available at https: //github. com/728388808/alpha-NeuS.

AAAI Conference 2024 Conference Paper

O^2-Recon: Completing 3D Reconstruction of Occluded Objects in the Scene with a Pre-trained 2D Diffusion Model

  • Yubin Hu
  • Sheng Ye
  • Wang Zhao
  • Matthieu Lin
  • Yuze He
  • Yu-Hui Wen
  • Ying He
  • Yong-Jin Liu

Occlusion is a common issue in 3D reconstruction from RGB-D videos, often blocking the complete reconstruction of objects and presenting an ongoing problem. In this paper, we propose a novel framework, empowered by a 2D diffusion-based in-painting model, to reconstruct complete surfaces for the hidden parts of objects. Specifically, we utilize a pre-trained diffusion model to fill in the hidden areas of 2D images. Then we use these in-painted images to optimize a neural implicit surface representation for each instance for 3D reconstruction. Since creating the in-painting masks needed for this process is tricky, we adopt a human-in-the-loop strategy that involves very little human engagement to generate high-quality masks. Moreover, some parts of objects can be totally hidden because the videos are usually shot from limited perspectives. To ensure recovering these invisible areas, we develop a cascaded network architecture for predicting signed distance field, making use of different frequency bands of positional encoding and maintaining overall smoothness. Besides the commonly used rendering loss, Eikonal loss, and silhouette loss, we adopt a CLIP-based semantic consistency loss to guide the surface from unseen camera angles. Experiments on ScanNet scenes show that our proposed framework achieves state-of-the-art accuracy and completeness in object-level reconstruction from scene-level RGB-D videos. Code: https://github.com/THU-LYJ-Lab/O2-Recon.

IJCAI Conference 2024 Conference Paper

OTOcc: Optimal Transport for Occupancy Prediction

  • Pengteng Li
  • Ying He
  • F. Richard Yu
  • Pinhao Song
  • Xingchen Zhou
  • Guang Zhou

The autonomous driving community is highly interested in 3D occupancy prediction due to its outstanding geometric perception and object recognition capabilities. However, previous methods are limited to existing semantic conversion mechanisms for solving sparse ground truths problem, causing excessive computational demands and sub-optimal voxels representation. To tackle the above limitations, we propose OTOcc, a novel 3D occupancy prediction framework that models semantic conversion from 2D pixels to 3D voxels as Optimal Transport (OT) problem, offering accurate semantic mapping to adapt to sparse scenarios without attention or depth estimation. Specifically, the unit transportation cost between each demander (voxel) and supplier (pixel) pair is defined as the weighted occupancy prediction loss. Then, we utilize the Sinkhorn-Knopp Iteration to find the best mapping matrices with minimal transportation costs. To reduce the computational cost, we propose a block reading technique with multi-perspective feature representation, which also brings fine-grained scene understanding. Extensive experiments show that OTOcc not only has the competitive prediction performance but also has about more than 4. 58% reduction in computational overhead compared to state-of-the-art methods.

JBHI Journal 2023 Journal Article

DeepTPpred: A Deep Learning Approach With Matrix Factorization for Predicting Therapeutic Peptides by Integrating Length Information

  • Zhen Cui
  • Si-Guo Wang
  • Ying He
  • Zhan-Heng Chen
  • Qin-Hu Zhang

The abuse of traditional antibiotics has led to increased resistance of bacteria and viruses. Efficient therapeutic peptide prediction is critical for peptide drug discovery. However, most of the existing methods only make effective predictions for one class of therapeutic peptides. It is worth noting that currently no predictive method considers sequence length information as a distinct feature of therapeutic peptides. In this article, a novel deep learning approach with matrix factorization for predicting therapeutic peptides (DeepTPpred) by integrating length information are proposed. The matrix factorization layer can learn the potential features of the encoded sequence through the mechanism of first compression and then restoration. And the length features of the sequence of therapeutic peptides are embedded with encoded amino acid sequences. To automatically learn therapeutic peptide predictions, these latent features are input into the neural networks with self-attention mechanism. On eight therapeutic peptide datasets, DeepTPpred achieved excellent prediction results. Based on these datasets, we first integrated eight datasets to obtain a full therapeutic peptide integration dataset. Then, we obtained two functional integration datasets based on the functional similarity of the peptides. Finally, we also conduct experiments on the latest versions of the ACP and CPP datasets. Overall, the experimental results show that our work is effective for the identification of therapeutic peptides.

NeurIPS Conference 2023 Conference Paper

NeuroGF: A Neural Representation for Fast Geodesic Distance and Path Queries

  • Qijian Zhang
  • Junhui Hou
  • Yohanes Adikusuma
  • Wenping Wang
  • Ying He

Geodesics play a critical role in many geometry processing applications. Traditional algorithms for computing geodesics on 3D mesh models are often inefficient and slow, which make them impractical for scenarios requiring extensive querying of arbitrary point-to-point geodesics. Recently, deep implicit functions have gained popularity for 3D geometry representation, yet there is still no research on neural implicit representation of geodesics. To bridge this gap, we make the first attempt to represent geodesics using implicit learning frameworks. Specifically, we propose neural geodesic field (NeuroGF), which can be learned to encode all-pairs geodesics of a given 3D mesh model, enabling to efficiently and accurately answer queries of arbitrary point-to-point geodesic distances and paths. Evaluations on common 3D object models and real-captured scene-level meshes demonstrate our exceptional performances in terms of representation accuracy and querying efficiency. Besides, NeuroGF also provides a convenient way of jointly encoding both 3D geometry and geodesics in a unified representation. Moreover, the working mode of per-model overfitting is further extended to generalizable learning frameworks that can work on various input formats such as unstructured point clouds, which also show satisfactory performances for unseen shapes and categories. Our code and data are available at https: //github. com/keeganhk/NeuroGF.

IJCAI Conference 2023 Conference Paper

RePaint-NeRF: NeRF Editting via Semantic Masks and Diffusion Models

  • Xingchen Zhou
  • Ying He
  • F. Richard Yu
  • Jianqiang Li
  • You Li

The emergence of Neural Radiance Fields (NeRF) has promoted the development of synthesized high-fidelity views of the intricate real world. However, it is still a very demanding task to repaint the content in NeRF. In this paper, we propose a novel framework that can take RGB images as input and alter the 3D content in neural scenes. Our work leverages existing diffusion models to guide changes in the designated 3D content. Specifically, we semantically select the target object and a pre-trained diffusion model will guide the NeRF model to generate new 3D objects, which can improve the editability, diversity, and application range of NeRF. Experiment results show that our algorithm is effective for editing 3D objects in NeRF under different text prompts, including editing appearance, shape, and more. We validate our method on both real-world datasets and synthetic-world datasets for these editing tasks. Please visit https: //repaintnerf. github. io for a better view of our results.

JBHI Journal 2022 Journal Article

FCNGRU: Locating Transcription Factor Binding Sites by Combing Fully Convolutional Neural Network With Gated Recurrent Unit

  • Siguo Wang
  • Ying He
  • Zhanheng Chen
  • Qinhu Zhang

Deciphering the relationship between transcription factors (TFs) and DNA sequences is very helpful for computational inference of gene regulation and a comprehensive understanding of gene regulation mechanisms. Transcription factor binding sites (TFBSs) are specific DNA short sequences that play a pivotal role in controlling gene expression through interaction with TF proteins. Although recently many computational and deep learning methods have been proposed to predict TFBSs aiming to predict sequence specificity of TF-DNA binding, there is still a lack of effective methods to directly locate TFBSs. In order to address this problem, we propose FCNGRU combing a fully convolutional neural network (FCN) with the gated recurrent unit (GRU) to directly locate TFBSs in this paper. Furthermore, we present a two-task framework (FCNGRU-double): one is a classification task at nucleotide level which predicts the probability of each nucleotide and locates TFBSs, and the other is a regression task at sequence level which predicts the intensity of each sequence. A series of experiments are conducted on 45 in-vitro datasets collected from the UniPROBE database derived from universal protein binding microarrays (uPBMs). Compared with competing methods, FCNGRU-double achieves much better results on these datasets. Moreover, FCNGRU-double has an advantage over a single-task framework, FCNGRU-single, which only contains the branch of locating TFBSs. In addition, we combine with in vivo datasets to make a further analysis and discussion.

IJCAI Conference 2022 Conference Paper

Multi-Constraint Deep Reinforcement Learning for Smooth Action Control

  • Guangyuan Zou
  • Ying He
  • F. Richard Yu
  • Longquan Chen
  • Weike Pan
  • Zhong Ming

Deep reinforcement learning (DRL) has been studied in a variety of challenging decision-making tasks, e. g. , autonomous driving. \textcolor{black}{However, DRL typically suffers from the action shaking problem, which means that agents can select actions with big difference even though states only slightly differ. } One of the crucial reasons for this issue is the inappropriate design of the reward in DRL. In this paper, to address this issue, we propose a novel way to incorporate the smoothness of actions in the reward. Specifically, we introduce sub-rewards and add multiple constraints related to these sub-rewards. In addition, we propose a multi-constraint proximal policy optimization (MCPPO) method to solve the multi-constraint DRL problem. Extensive simulation results show that the proposed MCPPO method has better action smoothness compared with the traditional proportional-integral-differential (PID) and mainstream DRL algorithms. The video is available at https: //youtu. be/F2jpaSm7YOg.

v2026.09.13