Arrow Research search

Author name cluster

Yiran Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

10 papers
2 author rows

Possible papers

10

EAAI Journal 2026 Journal Article

An enhanced deep learning framework with Large Separable Kernel Attention and Reparameterized Dual Convolution for real-time cold-crack detection in laser cladding

  • Jinyang Du
  • Ruipeng Gao
  • Yiran Wang
  • Jiabao Zhao
  • Yuechen Meng
  • WEI SHAO

Despite extensive industrial application, laser cladding continues to be challenged by cold-crack formation—a critical defect degrading structural integrity through brittle fracture propagation along grain boundaries. To address conventional Non-Destructive Testing (NDT) limitations in real-time monitoring (<50 ms latency) and micron-scale resolution (5-10 μm (μm)), a novel architecture—You Only Look Once version 9 (YOLOv9) integrated with Large Separable Kernel Attention (LSKA) and Reparameterized Dual Convolution (Rep-DualConv), hereinafter referred to as YOLOv9-LSKA-Rep-DualConv—is proposed for in situ cold-crack detection. This framework integrates three synergistic innovations: LSKA blocks expanding the receptive field by 40% to enhance contextual feature extraction, Rep-DualConv operations enabling hierarchical feature fusion with an 18% parameter reduction relative to standard convolutional blocks, and Dynamic Noise Suppression (DNS) Module elevating the Signal-to-Noise Ratio (SNR) by 12 dB (dB) in thermal environments. Comprehensive evaluation demonstrates that the model reaches 94. 2% mean Average Precision (mAP)@0. 5 on the proprietary dataset. While maintaining real-time performance, it achieves a 2% accuracy improvement compared to the baseline YOLOv9. Crucially, under stringent mAP@0. 5: 0. 95 evaluation, a consistent 1. 8-percentage-point advantage is maintained, enabling deployment in industrial closed-loop systems that achieve American Society of Mechanical Engineers (ASME) B46. 1-compliant 8 μm crack detection at 90% confidence—fulfilling critical quality control requirements for laser processing.

AAAI Conference 2026 Conference Paper

MacVQA: Adaptive Memory Allocation and Global Noise Filtering for Continual Visual Question Answering

  • Zhifei Li
  • Yiran Wang
  • Chenyi Xiong
  • Yujing Xia
  • Xiaoju Hou
  • Yue Zhao
  • Miao Zhang
  • Kui Xiao

Visual Question Answering (VQA) requires models to reason over multimodal information, combining visual and textual data. With the development of continual learning, significant progress has been made in retaining knowledge and adapting to new information in the VQA domain. However, current methods often struggle with balancing knowledge retention, adaptation, and robust feature representation. To address these challenges, we propose a novel framework with adaptive memory allocation and global noise filtering called MacVQA for visual question answering. MacVQA fuses visual and question information while filtering noise to ensure robust representations, and employs prototype-based memory allocation to optimize feature quality and memory usage. These designs enable MacVQA to balance knowledge acquisition, retention, and compositional generalization in continual VQA learning. Experiments on ten continual VQA tasks show that MacVQA outperforms existing baselines, achieving 43.38% average accuracy and 2.32% average forgetting on standard tasks, and 42.53% average accuracy and 3.60% average forgetting on novel composition tasks.

NeurIPS Conference 2025 Conference Paper

FlashMo: Geometric Interpolants and Frequency-Aware Sparsity for Scalable Efficient Motion Generation

  • Zeyu Zhang
  • Yiran Wang
  • Danning Li
  • Dong Gong
  • Ian Reid
  • Richard Hartley

Diffusion models have recently advanced 3D human motion generation by producing smoother and more realistic sequences from natural language. However, existing approaches face two major challenges: high computational cost during training and inference, and limited scalability due to reliance on U-Net inductive bias. To address these challenges, we propose **FlashMo**, a frequency-aware sparse motion diffusion model that prunes low-frequency tokens to enhance efficiency without custom kernel design. We further introduce *MotionSiT*, a scalable diffusion transformer based on a joint-temporal factorized interpolant with Lie group geodesics over $\mathrm{SO}(3)$ manifolds, enabling principled generation of joint rotations. Extensive experiments on the large-scale MotionHub V2 dataset and standard benchmarks including HumanML3D and KIT-ML demonstrate that our method significantly outperforms previous approaches in motion quality, efficiency, and scalability. Compared to the state-of-the-art 1-step distillation baseline, FlashMo reduces **12. 9%** inference time and FID by **34. 1%**. Project website: https: //steve-zeyu-zhang. github. io/FlashMo.

ICML Conference 2025 Conference Paper

Hyper: Hyperparameter Robust Efficient Exploration in Reinforcement Learning

  • Yiran Wang
  • Chenshu Liu
  • Yunfan Li
  • Sanae Amani
  • Bolei Zhou
  • Lin F. Yang

The exploration & exploitation dilemma poses significant challenges in reinforcement learning (RL). Recently, curiosity-based exploration methods achieved great success in tackling hard-exploration problems. However, they necessitate extensive hyperparameter tuning on different environments, which heavily limits the applicability and accessibility of this line of methods. In this paper, we characterize this problem via analysis of the agent behavior, concluding the fundamental difficulty of choosing a proper hyperparameter. We then identify the difficulty and the instability of the optimization when the agent learns with curiosity. We propose our method, hyperparameter robust exploration ( Hyper ), which extensively mitigates the problem by effectively regularizing the visitation of the exploration and decoupling the exploitation to ensure stable training. We theoretically justify that Hyper is provably efficient under function approximation setting and empirically demonstrate its appealing performance and robustness in various environments.

ICLR Conference 2025 Conference Paper

MetaUrban: An Embodied AI Simulation Platform for Urban Micromobility

  • Wayne Wu
  • Honglin He
  • Jack He
  • Yiran Wang
  • Chenda Duan
  • Zhizheng Liu
  • Quanyi Li
  • Bolei Zhou

Public urban spaces such as streetscapes and plazas serve residents and accommodate social life in all its vibrant variations. Recent advances in robotics and embodied AI make public urban spaces no longer exclusive to humans. Food delivery bots and electric wheelchairs have started sharing sidewalks with pedestrians, while robot dogs and humanoids have recently emerged in the street. **Micromobility** enabled by AI for short-distance travel in public urban spaces plays a crucial component in future transportation systems. It is essential to ensure the generalizability and safety of AI models used for maneuvering mobile machines. In this work, we present **MetaUrban**, a *compositional* simulation platform for the AI-driven urban micromobility research. MetaUrban can construct an *infinite* number of interactive urban scenes from compositional elements, covering a vast array of ground plans, object placements, pedestrians, vulnerable road users, and other mobile agents' appearances and dynamics. We design point navigation and social navigation tasks as the pilot study using MetaUrban for urban micromobility research and establish various baselines of Reinforcement Learning and Imitation Learning. We conduct extensive evaluation across mobile machines, demonstrating that heterogeneous mechanical structures significantly influence the learning and execution of AI policies. We perform a thorough ablation study, showing that the compositional nature of the simulated environments can substantially improve the generalizability and safety of the trained mobile agents. MetaUrban will be made publicly available to provide research opportunities and foster safe and trustworthy embodied AI and micromobility in cities. The code and data have been released.

YNIMG Journal 2024 Journal Article

Optimization-derived blood input function using a kernel method and its evaluation with total-body PET for brain parametric imaging

  • Yansong Zhu
  • Quyen Tran
  • Yiran Wang
  • Ramsey D. Badawi
  • Simon R. Cherry
  • Jinyi Qi
  • Shiva Abbaszadeh
  • Guobao Wang

Dynamic PET allows quantification of physiological parameters through tracer kinetic modeling. For dynamic imaging of brain or head and neck cancer on conventional PET scanners with a short axial field of view, the image-derived input function (ID-IF) from intracranial blood vessels such as the carotid artery (CA) suffers from severe partial volume effects. Alternatively, optimization-derived input function (OD-IF) by the simultaneous estimation (SIME) method does not rely on an ID-IF but derives the input function directly from the data. However, the optimization problem is often highly ill-posed. We proposed a new method that combines the ideas of OD-IF and ID-IF together through a kernel framework. While evaluation of such a method is challenging in human subjects, we used the uEXPLORER total-body PET system that covers major blood pools to provide a reference for validation. METHODS: F-fluorodeoxyglucose studies with both computer simulations and 20 human-subject scans acquired on the uEXPLORER scanner. The effect of the number of ROIs on kernel SIME was also explored. RESULTS: The estimated OD-IF by kernel SIME showed a good match with the reference input function and provided more accurate estimation of kinetic parameters for both simulation and human-subject data. The kernel SIME led to the highest correlation coefficient (R = 0.97) and the lowest mean absolute error (MAE = 10.5 %) compared to using the CA ID-IF (R = 0.86, MAE = 108.2 %) and conventional SIME (R = 0.57, MAE = 78.7 %) in the human-subject evaluation. Adding more ROIs improved the overall performance of the kernel SIME method. CONCLUSION: The proposed kernel SIME method shows promise to provide an accurate estimation of the blood input function and kinetic parameters for brain PET parametric imaging.

EAAI Journal 2024 Journal Article

Robust minimum cost consensus models with uncertain asymmetric costs based on linear uncertain-constrained tolerance level

  • Zhongming Wu
  • Pan Gao
  • Yiran Wang
  • Xiaoxia Xu
  • Neng Wan
  • Francisco Javier Cabrerizo

The ability of the minimum cost consensus model (MCCM) to promote consensus reaching in the domain of group decision-making (GDM) has been extensively studied. Recently, the MCCM has been enhanced by introducing the consensus principle and tolerance level to achieve a soft consensus. However, the potential impact of asymmetric and uncertain unit adjustment costs on the effectiveness of the consensus reaching process (CRP) has been overlooked. This paper aims to investigate the implications of uncertain asymmetric costs for achieving consensus with a certain level of tolerance, where new robust MCCMs with uncertain asymmetric costs are constructed under four uncertainty sets for the unit adjustment costs. Considering the linear uncertain-constrained tolerance level and consensus level, we incorporate the insight of an expert with a cost-free threshold into models. Additionally, through a pollutant emission application, the proposed robust MCCMs are able to effectively handle uncertainties arising from costs and improve the quality of the CRP compared to the traditional models. Finally, we conduct simulation experiments and sensitivity analysis to illustrate the effectiveness of the proposed models on achieving a consensus by identifying appropriate parameters.

NeurIPS Conference 2024 Conference Paper

Self-Distilled Depth Refinement with Noisy Poisson Fusion

  • Jiaqi Li
  • Yiran Wang
  • Jinghong Zheng
  • Zihao Huang
  • Ke Xian
  • Zhiguo Cao
  • Jianming Zhang

Depth refinement aims to infer high-resolution depth with fine-grained edges and details, refining low-resolution results of depth estimation models. The prevailing methods adopt tile-based manners by merging numerous patches, which lacks efficiency and produces inconsistency. Besides, prior arts suffer from fuzzy depth boundaries and limited generalizability. Analyzing the fundamental reasons for these limitations, we model depth refinement as a noisy Poisson fusion problem with local inconsistency and edge deformation noises. We propose the Self-distilled Depth Refinement (SDDR) framework to enforce robustness against the noises, which mainly consists of depth edge representation and edge-based guidance. With noisy depth predictions as input, SDDR generates low-noise depth edge representations as pseudo-labels by coarse-to-fine self-distillation. Edge-based guidance with edge-guided gradient loss and edge-based fusion loss serves as the optimization objective equivalent to Poisson fusion. When depth maps are better refined, the labels also become more noise-free. Our model can acquire strong robustness to the noises, achieving significant improvements in accuracy, edge quality, efficiency, and generalizability on five different benchmarks. Moreover, directly training another model with edge labels produced by SDDR brings improvements, suggesting that our method could help with training robust refinement models in future works.

NeurIPS Conference 2024 Conference Paper

Unleashing Region Understanding in Intermediate Layers for MLLM-based Referring Expression Generation

  • Yaoyuan Liang
  • Zhuojun Cai
  • Jian Xu
  • Guanbo Huang
  • Yiran Wang
  • Xiao Liang
  • Jiahao Liu
  • Ziran Li

The Multi-modal Large Language Model (MLLM) based Referring Expression Generation (REG) task has gained increasing popularity, which aims to generate an unambiguous text description that applies to exactly one object or region in the image by leveraging foundation models. We empirically found that there exists a potential trade-off between the detailedness and the correctness of the descriptions for the referring objects. On the one hand, generating sentences with more details is usually required in order to provide more precise object descriptions. On the other hand, complicated sentences could easily increase the probability of hallucinations. To address this issue, we propose a training-free framework, named ``unleash-then-eliminate'', which first elicits the latent information in the intermediate layers, and then adopts a cycle-consistency-based decoding method to alleviate the production of hallucinations. Furthermore, to reduce the computational load of cycle-consistency-based decoding, we devise a Probing-based Importance Estimation method to statistically estimate the importance weights of intermediate layers within a subset. These importance weights are then incorporated into the decoding process over the entire dataset, intervening in the next token prediction from intermediate layers. Extensive experiments conducted on the RefCOCOg and PHD benchmarks show that our proposed framework could outperform existing methods on both semantic and hallucination-related metrics. Code will be made available in https: //github. com/Glupayy/unleash-eliminate.

ICML Conference 2023 Conference Paper

Low-Switching Policy Gradient with Exploration via Online Sensitivity Sampling

  • Yunfan Li
  • Yiran Wang
  • Yu Cheng
  • Lin Yang

Policy optimization methods are powerful algorithms in Reinforcement Learning (RL) for their flexibility to deal with policy parameterization and ability to handle model misspecification. However, these methods usually suffer from slow convergence rates and poor sample complexity. Hence it is important to design provably sample efficient algorithms for policy optimization. Yet, recent advances for this problems have only been successful in tabular and linear setting, whose benign structures cannot be generalized to non-linearly parameterized policies. In this paper, we address this problem by leveraging recent advances in value-based algorithms, including bounded eluder-dimension and online sensitivity sampling, to design a low-switching sample-efficient policy optimization algorithm, LPO, with general non-linear function approximation. We show that, our algorithm obtains an $\varepsilon$-optimal policy with only $\widetilde{O}(\frac{\text{poly}(d)}{\varepsilon^3})$ samples, where $\varepsilon$ is the suboptimality gap and $d$ is a complexity measure of the function class approximating the policy. This drastically improves previously best-known sample bound for policy optimization algorithms, $\widetilde{O}(\frac{\text{poly}(d)}{\varepsilon^8})$. Moreover, we empirically test our theory with deep neural nets to show the benefits of the theoretical inspiration.

v2026.09.13