Arrow Research search

Author name cluster

Huan Yang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

12 papers
2 author rows

Possible papers

12

EAAI Journal 2025 Journal Article

A 6-dimensional pose estimation method combining sparse viewpoint classification initialization and optical flow-guided iterative refinement

  • Huan Yang
  • Yue Wang
  • Xinghang Yin
  • Yongxu Liu
  • Peng Wang

Electronic equipment is typically a complex and high-precision electromechanical system, where the routing and bundling of Radio Frequency (RF) cables are crucial to equipment performance. Traditional assembly methods require workers to assemble according to the assembly process card, which can easily lead to incorrect or missing assembly, poor assembly consistency, and low efficiency. Augmented Reality (AR) assembly guidance can effectively improve efficiency and reduce errors. 6-dimensional (6D) pose estimation is a key technology for AR assembly guidance. In the assembly process of complex electronic products, existing deep learning methods suffer from poor tracking and localization robustness and real-time performance due to factors such as arm occlusion, resulting in slow tracking recovery. This article proposes a two-stage real-time 6D pose estimation method from coarse to fine, which can estimate the pose of target objects in complex backgrounds at a speed of about 20 frames per second and quickly recover after tracking target loss. The real-time and effectiveness were verified through experiments on the red squirrel and electronic chassis.

ECAI Conference 2025 Conference Paper

DialogGraph-LLM: Graph-Informed LLMs for End-to-End Audio Dialogue Intent Recognition

  • HongYu Liu
  • Junxin Li
  • Changxi Guo
  • Hao Chen
  • Yaqian Huang
  • Yifu Guo
  • Huan Yang
  • Lihua Cai

Recognizing speaker intent in long audio dialogues among speakers has a wide range of applications, but is a non-trivial AI task due to complex inter-dependencies in speaker utterances and scarce annotated data. To address these challenges, an end-to-end framework, namely DialogGraph-LLM, is proposed in the current work. DialogGraph-LLM combines a novel Multi-Relational Dialogue Attention Network (MR-DAN) architecture with multimodal foundation models (e. g. , Qwen2. 5-Omni-7B) for direct acoustic-to-intent inference. An adaptive semi-supervised learning strategy is designed using LLM with a confidence-aware pseudo-label generation mechanism based on dual-threshold filtering using both global and class confidences, and an entropy-based sample selection process that prioritizes high-information unlabeled instances. Extensive evaluations on the proprietary MarketCalls corpus and the publicly available MIntRec 2. 0 benchmark demonstrate DialogGraph-LLM’s superiority over strong audio and text-driven baselines. The framework demonstrates strong performance and efficiency in intent recognition in real world scenario audio dialogues, proving its practical value for audio-rich domains with limited supervision. Our code is available at https: //github. com/david188888/DialogGraph-LLM

NeurIPS Conference 2025 Conference Paper

Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference Optimization

  • Tao Zhang
  • Cheng Da
  • Kun Ding
  • Huan Yang
  • Kun Jin
  • Yan Li
  • Tingting Gao
  • Di Zhang

Preference optimization for diffusion models aims to align them with human preferences for images. Previous methods typically use Vision-Language Models (VLMs) as pixel-level reward models to approximate human preferences. However, when used for step-level preference optimization, these models face challenges in handling noisy images of different timesteps and require complex transformations into pixel space. In this work, we show that pre-trained diffusion models are naturally suited for step-level reward modeling in the noisy latent space, as they are explicitly designed to process latent images at various noise levels. Accordingly, we propose the Latent Reward Model (LRM), which repurposes components of the diffusion model to predict preferences of latent images at arbitrary timesteps. Building on LRM, we introduce Latent Preference Optimization (LPO), a step-level preference optimization method conducted directly in the noisy latent space. Experimental results indicate that LPO significantly improves the model's alignment with general, aesthetic, and text-image alignment preferences, while achieving a 2. 5-28x training speedup over existing preference optimization methods.

IROS Conference 2025 Conference Paper

Tele-GS: 3D Gaussian Scene Representation for Low-Bandwidth Teleoperation

  • Chunyang Zhao
  • Zeyu Zhou
  • Haoran Liu
  • Dogan Kircali
  • Huan Yang
  • Chang Boon Low
  • Yuanzhe Wang
  • Danwei Wang

Video streaming based teleoperation often faces a trade-off between bandwidth consumption and the need for high-fidelity telepresence. Higher image resolution or a wider field of view (FOV) substantially increases bandwidth requirements. In this paper, we propose a novel telepresence model for teleoperated vehicles operating in bandwidth-constrained environments. Our approach employs a LiDAR-fused 3D Gaussian Splatting (3DGS) as a compact scene representation to efficiently generate remote views. Initially, a static point cloud map is constructed using LiDAR-based semantic mapping, which serves as the initial Gaussians for optimizing the 3DGS model. During teleoperation, the prebuilt 3DGS is then rendered on the teleoperation platform, while only safety-critical information, such as vehicle pose and dynamic objects, is transmitted from the vehicle to the teleoperator in real-time. The proposed telepresence model significantly reduces data transmission requirements while maintaining photorealistic telepresence, enabling reliable and effective teleoperation even under stringent bandwidth constraints. This capability ensures safe and efficient vehicle teleoperation under challenging environments without relying on traditional high-bandwidth communication, thereby broadening the applicability of teleoperation technology to more demanding and diverse operational scenarios. Real-world experimental results show that the developed system can provide immersive teleoperation experiences at Kbps-level bandwidth consumption.

AIIM Journal 2024 Journal Article

A clinical consensus-compliant deep learning approach to quantitatively evaluate human in vitro fertilization early embryonic development with optical microscope images

  • Zaowen Liao
  • Chaoyu Yan
  • Jianbo Wang
  • Ningfeng Zhang
  • Huan Yang
  • Chenghao Lin
  • Haiyue Zhang
  • Wenjun Wang

The selection of embryos is a key for the success of in vitro fertilization (IVF). However, automatic quality assessment on human IVF embryos with optical microscope images is still challenging. In this study, we developed a clinical consensus-compliant deep learning approach, named Esava (Embryo Segmentation and Viability Assessment), to quantitatively evaluate the development of IVF embryos using optical microscope images. In total 551 optical microscope images of human IVF embryos of day-2 to day-3 were collected, preprocessed, and annotated. Using the Faster R-CNN model as baseline, our Esava model was constructed, refined, trained, and validated for precise and robust blastomere detection. A novel algorithm Crowd-NMS was proposed and employed in Esava to enhance the object detection and to precisely quantify the embryonic cells and their size uniformity. Additionally, an innovative GrabCut-based unsupervised module was integrated for the segmentation of blastomeres and embryos. Independently tested on 94 embryo images for blastomere detection, Esava obtained the high rates of 0. 9940, 0. 9121, and 0. 9531 for precision, recall, and mAP respectively, and gained significant advances compared with previous computational methods. Intraclass correlation coefficients indicated the consistency between Esava and three experienced embryologists. Another test on 51 extra images demonstrated that Esava surpassed other tools significantly, achieving the highest average precision 0. 9025. Moreover, it also accurately identified the borders of blastomeres with mIoU over 0. 88 on the independent testing dataset. Esava is compliant with the Istanbul clinical consensus and compatible to senior embryologists. Taken together, Esava improves the accuracy and efficiency of embryonic development assessment with optical microscope images.

TMLR Journal 2024 Journal Article

AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks

  • Max Ku
  • Cong Wei
  • Weiming Ren
  • Huan Yang
  • Wenhu Chen

In the dynamic field of digital content creation using generative models, state-of-the-art video editing models still do not offer the level of quality and control that users desire. Previous works on video editing either extended from image-based generative models in a zero-shot manner or necessitated extensive fine-tuning, which can hinder the production of fluid video edits. Furthermore, these methods frequently rely on textual input as the editing guidance, leading to ambiguities and limiting the types of edits they can perform. Recognizing these challenges, we introduce AnyV2V, a novel tuning-free paradigm designed to simplify video editing into two primary steps: (1) employing an off-the-shelf image editing model to modify the first frame, (2) utilizing an existing image-to-video generation model to generate the edited video through temporal feature injection. AnyV2V can leverage any existing image editing tools to support an extensive array of video editing tasks, including prompt-based editing, reference-based style transfer, subject-driven editing, and identity manipulation, which were unattainable by previous methods. AnyV2V can also support any video length. Our evaluation shows that AnyV2V achieved CLIP Scores comparable to other baseline methods. Furthermore, AnyV2V significantly outperformed these baselines in human evaluations, demonstrating notable improvements in visual consistency with the source video while producing high-quality edits across all editing tasks.

TMLR Journal 2024 Journal Article

ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation

  • Weiming Ren
  • Huan Yang
  • Ge Zhang
  • Cong Wei
  • Xinrun Du
  • Wenhao Huang
  • Wenhu Chen

Image-to-video (I2V) generation aims to use the initial frame (alongside a text prompt) to create a video sequence. A grand challenge in I2V generation is to maintain visual consistency throughout the video: existing methods often struggle to preserve the integrity of the subject, background, and style from the first frame, as well as ensure a fluid and logical progression within the video narrative. To mitigate these issues, we propose ConsistI2V, a diffusion-based method to enhance visual consistency for I2V generation. Specifically, we introduce (1) spatiotemporal attention over the first frame to maintain spatial and motion consistency, (2) noise initialization from the low-frequency band of the first frame to enhance layout consistency. These two approaches enable ConsistI2V to generate highly consistent videos. We also extend the proposed approaches to show their potential to improve consistency in auto-regressive long video generation and camera motion control. To verify the effectiveness of our method, we propose I2V-Bench, a comprehensive evaluation benchmark for I2V generation. Our automatic and human evaluation results demonstrate the superiority of ConsistI2V over existing methods.

YNIMG Journal 2024 Journal Article

Precise detection of awareness in disorders of consciousness using deep learning framework

  • Huan Yang
  • Hang Wu
  • Lingcong Kong
  • Wen Luo
  • Qiuyou Xie
  • Jiahui Pan
  • Wuxiu Quan
  • Lianting Hu

Diagnosis of disorders of consciousness (DOC) remains a formidable challenge. Deep learning methods have been widely applied in general neurological and psychiatry disorders, while limited in DOC domain. Considering the successful use of resting-state functional MRI (rs-fMRI) for evaluating patients with DOC, this study seeks to explore the conjunction of deep learning techniques and rs-fMRI in precisely detecting awareness in DOC. We initiated our research with a benchmark dataset comprising 140 participants, including 76 unresponsive wakefulness syndrome (UWS), 25 minimally conscious state (MCS), and 39 Controls, from three independent sites. We developed a cascade 3D EfficientNet-B3-based deep learning framework tailored for discriminating MCS from UWS patients, referred to as "DeepDOC", and compared its performance against five state-of-the-art machine learning models. We also included an independent dataset consists of 11 DOC patients to test whether our model could identify patients with cognitive motor dissociation (CMD), in which DOC patients were behaviorally diagnosed unconscious but could be detected conscious by brain computer interface (BCI) method. Our results demonstrate that DeepDOC outperforms the five machine learning models, achieving an area under curve (AUC) value of 0.927 and accuracy of 0.861 for distinguishing MCS from UWS patients. More importantly, DeepDOC excels in CMD identification, achieving an AUC of 1 and accuracy of 0.909. Using gradient-weighted class activation mapping algorithm, we found that the posterior cortex, encompassing the visual cortex, posterior middle temporal gyrus, posterior cingulate cortex, precuneus, and cerebellum, as making a more substantial contribution to classification compared to other brain regions. This research offers a convenient and accurate method for detecting covert awareness in patients with MCS and CMD using rs-fMRI data.

JBHI Journal 2022 Journal Article

Attention Gate Based Dual-Pathway Network for Vertebra Segmentation of X-Ray Spine Images

  • Wenbo Shi
  • Tongshuai Xu
  • Huan Yang
  • Yongming Xi
  • Yukun Du
  • Jinhua Li
  • Jinxu Li

Automatic spine and vertebra segmentation from X-ray spine images is a critical and challenging problem in many computer-aid spinal image analysis and disease diagnosis applications. In this paper, a two-stage automatic segmentation framework for spine X-ray images is proposed, which can firstly locate the spine regions (including backbone, sacrum and ilium) in the coarse stage and then identify eighteen vertebrae (i. e. , cervical vertebra 7, thoracic vertebra 1-12 and lumbar vertebra 1-5) with isolate and clear boundary in the fine stage. A novel Attention Gate based dual-pathway Network (AGNet) composed of context and edge pathways is designed to extract semantic and boundary information for segmentation of both spine and vertebra regions. Multi-scale supervision mechanism is applied to explore comprehensive features and an Edge aware Fusion Mechanism (EFM) is proposed to fuse features extracted from the two pathways. Some other image processing skills, such as centralized backbone clipping, patch cropping and convex hull detection are introduced to further refine the vertebra segmentation results. Experimental validations on spine X-ray images dataset and vertebrae dataset suggest that the proposed AGNet achieves superior performance compared with state-of-the-art segmentation methods, and the coarse-to-fine framework can be implemented in real spinal diagnosis systems.

NeurIPS Conference 2022 Conference Paper

Long-Form Video-Language Pre-Training with Multimodal Temporal Contrastive Learning

  • Yuchong Sun
  • Hongwei Xue
  • Ruihua Song
  • Bei Liu
  • Huan Yang
  • Jianlong Fu

Large-scale video-language pre-training has shown significant improvement in video-language understanding tasks. Previous studies of video-language pretraining mainly focus on short-form videos (i. e. , within 30 seconds) and sentences, leaving long-form video-language pre-training rarely explored. Directly learning representation from long-form videos and language may benefit many long-formvideo-language understanding tasks. However, it is challenging due to the difficulty of modeling long-range relationships and the heavy computational burden caused by more frames. In this paper, we introduce a Long-Form VIdeo-LAnguage pre-training model (LF-VILA) and train it on a large-scale long-form video and paragraph dataset constructed from an existing public dataset. To effectively capturethe rich temporal dynamics and to better align video and language in an efficient end-to-end manner, we introduce two novel designs in our LF-VILA model. We first propose a Multimodal Temporal Contrastive (MTC) loss to learn the temporal relation across different modalities by encouraging fine-grained alignment between long-form videos and paragraphs. Second, we propose a Hierarchical Temporal Window Attention (HTWA) mechanism to effectively capture long-range dependency while reducing computational cost in Transformer. We fine-tune the pre-trained LF-VILA model on seven downstream long-form video-language understanding tasks of paragraph-to-video retrieval and long-form video question-answering, and achieve new state-of-the-art performances. Specifically, our model achieves 16. 1% relative improvement on ActivityNet paragraph-to-video retrieval task and 2. 4% on How2QA task, respectively. We release our code, dataset, and pre-trained models at https: //github. com/microsoft/XPretrain.

NeurIPS Conference 2021 Conference Paper

Improving Visual Quality of Image Synthesis by A Token-based Generator with Transformers

  • Yanhong Zeng
  • Huan Yang
  • Hongyang Chao
  • Jianbo Wang
  • Jianlong Fu

We present a new perspective of achieving image synthesis by viewing this task as a visual token generation problem. Different from existing paradigms that directly synthesize a full image from a single input (e. g. , a latent code), the new formulation enables a flexible local manipulation for different image regions, which makes it possible to learn content-aware and fine-grained style control for image synthesis. Specifically, it takes as input a sequence of latent tokens to predict the visual tokens for synthesizing an image. Under this perspective, we propose a token-based generator (i. e. , TokenGAN). Particularly, the TokenGAN inputs two semantically different visual tokens, i. e. , the learned constant content tokens and the style tokens from the latent space. Given a sequence of style tokens, the TokenGAN is able to control the image synthesis by assigning the styles to the content tokens by attention mechanism with a Transformer. We conduct extensive experiments and show that the proposed TokenGAN has achieved state-of-the-art results on several widely-used image synthesis benchmarks, including FFHQ and LSUN CHURCH with different resolutions. In particular, the generator is able to synthesize high-fidelity images with (1024x1024) size, dispensing with convolutions entirely.

YNICL Journal 2020 Journal Article

Brain GABA+ changes in primary hypothyroidism patients before and after levothyroxine treatment: A longitudinal magnetic resonance spectroscopy study

  • Bo Liu
  • Zhensong Wang
  • Liangjie Lin
  • Huan Yang
  • Fei Gao
  • Tao Gong
  • Richard A.E. Edden
  • Guangbin Wang

OBJECTIVE: Increasing evidence indicates the involvement of the GABAergic system in the pathophysiology of hypothyroidism. We aimed to investigate longitudinal changes of brain GABA in primary hypothyroidism before and after levothyroxine (L-T4) treatment. MATERIAL AND METHODS: In 18 patients with hypothyroidism, we used the MEGA-PRESS (Mescher-Garwood point-resolved spectroscopy) editing sequence to measure brain GABA levels from medial prefrontal cortex (mPFC) and posterior cingulate cortex (PCC) at baseline and after 6-months of L-T4 treatment. Sex- and age-matched healthy controls (n = 18) were scanned at baseline. Thyroid function and neuropsychological tests were also performed. RESULTS: GABA signals were successfully quantified from all participants with fitting errors lower than 15%. GABA signal was labeled as GABA+ due to contamination from co-edited macromoleculars and homocarnosine. In hypothyroid patients, mean GABA+ was significantly lower in the mPFC region compared with controls (p = 0.031), and the mPFC GABA+ measurements were significantly correlated with depressive symptoms and memory function (r = -0.558, p = 0.016; r = 0.522, p = 0.026, respectively). After adequate L-T4 treatment, the mPFC GABA+ in hypothyroid patients increased to normal level, along with relieved neuropsychological impairments. CONCLUSION: The study suggested the decrease of GABA+ may be an important neurobiological factor in the pathophysiology of hypothyroidism. Treatment of L-T4 may reverse the abnormal GABA+ and hypothyroidism-induced neuropsychiatric impairments, indicating the action mode of L-T4 in adjunctive treatment of affective disorders.

v2026.09.13