Arrow Research search

Author name cluster

Yang Tian

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

9 papers
2 author rows

Possible papers

9

AAAI Conference 2026 Conference Paper

Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling

  • Hao Li
  • Shuai Yang
  • Yilun Chen
  • Xinyi Chen
  • Xiaoda Yang
  • Yang Tian
  • Hanqing Wang
  • Tai WANG

Recent vision-language-action (VLA) models built on pretrained vision-language models (VLMs) have demonstrated strong performance in robotic manipulation. However, these models remain constrained by the single-frame image paradigm and fail to fully leverage the temporal information offered by multi-frame histories, as directly feeding multiple frames into VLM backbones incurs substantial computational overhead and inference latency. We propose CronusVLA, a unified framework that extends single-frame VLA models to the multi-frame paradigm. CronusVLA follows a two-stage process: (1) Single-frame pretraining on large-scale embodied datasets with autoregressive prediction of action tokens, establishing an effective embodied vision-language foundation; (2) Multi-frame post-training, which adapts the prediction of the vision-language backbone from discrete tokens to learnable features, and aggregates historical information via feature chunking. CronusVLA effectively addresses the existing challenges of multi-frame modeling while enhancing performance. To evaluate the robustness under temporal and spatial disturbances, we introduce SimplerEnv-OR, a novel benchmark featuring 24 types of observational disturbances and 120 severity levels. Experiments across three embodiments in simulated and real-world environments demonstrate that CronusVLA achieves leading performance and superior robustness, with a 70.9% success rate on SimplerEnv, a 26.8% improvement over OpenVLA on LIBERO, and the highest robustness score on SimplerEnv-OR, showing the promise of efficient multi-frame adaptation for real-world VLA deployment.

EAAI Journal 2025 Journal Article

Coarse-to-fine dual-branch network for ship target recognition in complex environments

  • Yang Tian
  • Hao Meng

Harsh sea conditions and the complex and variable positions of ships significantly impact the capacity of imaging devices to capture high-quality ship images, making ship target recognition challenging in the application of artificial intelligence. Many scholars have recently proposed cascaded recognition models to address this issue. Following this method, in this paper, we propose a novel method for ship target recognition in complex environments called the coarse-to-fine dual-branch (CFDB) network. The CFDB model designs a dual-branch network from coarse to fine to lock the target area fine features and then uses peer-to-peer communication to extract and exchange learning of the target region’s final discriminative contour features, assisting in predicting ship classes in the complex environment. The proposed method is evaluated on the constructed complex in background ships (CIB-ships) dataset and the publicly available Marine Argos Recognition Ships (MAR-ships) and Game-of-Ships datasets. Compared with the suboptimal method, the proposed CFDB network exhibits improvements of 2. 11%, 1. 33%, and 1. 24% accuracy on the CIB-ships, MAR-ships, and Game-of-ships datasets, respectively. The results demonstrate that the proposed method provides useful ideas for the dynamic monitoring of ships in real environments. Our code will be published at https: //github. com/yangt1013/CFDB-master.

EAAI Journal 2025 Journal Article

Fuzzy reinforcement learning prescribed-time algorithm for the rigid–flexible coupled robotic mechanisms with input deadzone

  • Xingyu Zhou
  • Haoping Wang
  • Yang Tian

In the horizontal plane, the dynamic model for rigid–flexible coupled robotic mechanisms under large beam-deformations are determined through the utilization of a comprehensive modeling approach based on the virtual work concept. To track the desired angular positions of such robotic mechanisms with input nonsymmetric deadzone, the fuzzy nonsymmetric deadzone compensation based prescribed time adaptive reinforcement learning control strategy, incorporated with virtual robust linear quadratic state feedback input is proposed. To handle the unknown nonsymmetric input deadzone and uncertain system dynamics, an actor prescribed time fuzzy law is adopted. For further reduce the large vibration modes and tracking errors simultaneously, a virtual input and the proposition of a robust linear quadratic state feedback controller are developed. With the Lyapunov direct strategy, the angular position tracking errors and the flexible vibration of robotic mechanisms are demonstrated to converge to a tiny confined compact set. In numerical scenarios, the proposed fuzzy nonsymmetric deadzone compensation-based prescribed time adaptive reinforcement learning strategy simultaneously reduced mean angular tracking errors in a preset time and flexible vibration when compared respectively to virtual robust state feedback-free and backstepping mode control baselines.

NeurIPS Conference 2025 Conference Paper

GraphChain: Large Language Models for Large-scale Graph Analysis via Tool Chaining

  • Chunyu Wei
  • Wenji Hu
  • Xingjia Hao
  • Xin Wang
  • Yifan Yang
  • Yunhai Wang
  • Yang Tian
  • Yueguo Chen

Large Language Models (LLMs) face significant limitations when applied to large-scale graphs, struggling with context constraints and inflexible reasoning. We introduce GraphChain, a novel framework enabling LLMs to analyze large graphs by orchestrating dynamic sequences of specialized tools, mimicking human exploratory processes. GraphChain incorporates two core technical contributions: (1) Progressive Graph Distillation, a reinforcement learning approach that learns to generate tool sequences balancing task relevance and intermediate state compression, thereby overcoming LLM context limitations. (2) Structure-aware Test-Time Adaptation (STTA), a mechanism using a lightweight, self-supervised adapter conditioned on graph spectral properties to efficiently adapt a frozen LLM policy to diverse graph structures via soft prompts without retraining. Experiments show GraphChain significantly outperforms prior methods, enabling scalable and adaptive LLM-driven graph analysis.

ICLR Conference 2025 Conference Paper

Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation

  • Yang Tian
  • Sizhe Yang
  • Jia Zeng
  • Ping Wang
  • Dahua Lin
  • Hao Dong 0003
  • Jiangmiao Pang

Current efforts to learn scalable policies in robotic manipulation primarily fall into two categories: one focuses on "action," which involves behavior cloning from extensive collections of robotic data, while the other emphasizes "vision," enhancing model generalization by pre-training representations or generative models, also referred to as world models, using large-scale visual datasets. This paper presents an end-to-end paradigm that predicts actions using inverse dynamics models conditioned on the robot's forecasted visual states, named Predictive Inverse Dynamics Models (PIDM). By closing the loop between vision and action, the end-to-end PIDM can be a better scalable action learner. In practice, we use Transformers to process both visual states and actions, naming the model Seer. It is initially pre-trained on large-scale robotic datasets, such as DROID, and can be adapted to real-world scenarios with a little fine-tuning data. Thanks to large-scale, end-to-end training and the continuous synergy between vision and action at each execution step, Seer significantly outperforms state-of-the-art methods across both simulation and real-world experiments. It achieves improvements of 13% on the LIBERO-LONG benchmark, 22% on CALVIN ABC-D, and 43% in real-world tasks. Notably, it demonstrates superior generalization for novel objects, lighting conditions, and environments under high-intensity disturbances. Code and models will be publicly available.

ICRA Conference 2024 Conference Paper

RoboKeyGen: Robot Pose and Joint Angles Estimation via Diffusion-based 3D Keypoint Generation

  • Yang Tian
  • Jiyao Zhang
  • Guowei Huang 0002
  • Bin Wang
  • Ping Wang
  • Jiangmiao Pang
  • Hao Dong 0003

Estimating robot pose and joint angles is significant in advanced robotics, enabling applications like robot collaboration and online hand-eye calibration. However, the introduction of unknown joint angles makes prediction more complex than simple robot pose estimation, due to its higher dimensionality. Previous methods either regress 3D keypoints directly or utilise a render&compare strategy. These approaches often falter in terms of performance or efficiency and grapple with the cross-camera gap problem. This paper presents a novel framework that bifurcates the high-dimensional prediction task into two manageable subtasks: 2D keypoints detection and lifting 2D keypoints to 3D. This separation promises enhanced performance without sacrificing the efficiency innate to keypoint-based techniques. A vital component of our method is the lifting of 2D keypoints to 3D keypoints. Common deterministic regression methods may falter when faced with uncertainties from 2D detection errors or self-occlusions. Leveraging the robust modeling potential of diffusion models, we reframe this issue as a conditional 3D keypoints generation task. To bolster cross-camera adaptability, we introduce the Normalised Camera Coordinate Space (NCCS), ensuring alignment of estimated 2D keypoints across varying camera intrinsics. Experimental results demonstrate that the proposed method outperforms the state-of-the-art render&compare method and achieves higher inference speed. Furthermore, the tests accentuate our method’s robust cross-camera generalisation capabilities. We intend to release both the dataset and code in https://nimolty.github.io/Robokeygen/.

NeurIPS Conference 2023 Conference Paper

GLIME: General, Stable and Local LIME Explanation

  • Zeren Tan
  • Yang Tian
  • Jian Li

As black-box machine learning models become more complex and are applied in high-stakes settings, the need for providing explanations for their predictions becomes crucial. Although Local Interpretable Model-agnostic Explanations (LIME) \cite{ribeiro2016should} is a widely adopted method for understanding model behavior, it suffers from instability with respect to random seeds \cite{zafar2019dlime, shankaranarayana2019alime, bansal2020sam} and exhibits low local fidelity (i. e. , how the explanation explains model's local behaviors) \cite{rahnama2019study, laugel2018defining}. Our study demonstrates that this instability is caused by small sample weights, resulting in the dominance of regularization and slow convergence. Additionally, LIME's sampling approach is non-local and biased towards the reference, leading to diminished local fidelity and instability to references. To address these challenges, we propose \textsc{Glime}, an enhanced framework that extends LIME and unifies several previous methods. Within the \textsc{Glime} framework, we derive an equivalent formulation of LIME that achieves significantly faster convergence and improved stability. By employing a local and unbiased sampling distribution, \textsc{Glime} generates explanations with higher local fidelity compared to LIME, while being independent of the reference choice. Moreover, \textsc{Glime} offers users the flexibility to choose sampling distribution based on their specific scenarios.

ICML Conference 2023 Conference Paper

Robust Explanation for Free or At the Cost of Faithfulness

  • Zeren Tan
  • Yang Tian

Devoted to interpreting the explicit behaviors of machine learning models, explanation methods can identify implicit characteristics of models to improve trustworthiness. However, explanation methods are shown as vulnerable to adversarial perturbations, implying security concerns in high-stakes domains. In this paper, we investigate when robust explanations are necessary and what they cost. We prove that the robustness of explanations is determined by the robustness of the model to be explained. Therefore, we can have robust explanations for free for a robust model. To have robust explanations for a non-robust model, composing the original model with a kernel is proved as an effective way that returns strictly more robust explanations. Nevertheless, we argue that this also incurs a robustness-faithfulness trade-off, i. e. , contrary to common expectations, an explanation method may also become less faithful when it becomes more robust. This argument holds for any model. We are the first to introduce this trade-off and theoretically prove its existence for SmoothGrad. Theoretical findings are verified by empirical evidence on six state-of-the-art explanation methods and four backbones.

ICRA Conference 2020 Conference Paper

An Actuation Fault Tolerance Approach to Reconfiguration Planning of Modular Self-folding Robots

  • Meibao Yao
  • Xueming Xiao
  • Yang Tian
  • Hutao Cui
  • Jamie Paik

This paper presents a novel approach to fault tolerant reconfiguration of modular self-folding robots. Among various types of faults that probably occur in the modular system, we focus on the tolerance of complete actuation failure of active modules that might cause imprecise robotic motion and even reconfiguration failure. Our approach is to utilize the reconfigurability of modular self-folding robots and investigate intra-module connection to determine initial patterns that are inherently fault tolerant. We exploit the redundancy of actuation and distribute active modules in both layout-based and target-based scenarios, such that reconfiguration schemes with user-specified fault tolerant capability can be generated for an arbitrary input initial pattern or 3D configuration. Our methods are demonstrated in computer-aided simulation on the robotic platform of Mori, a modular origami robot. The simulation results validate that the proposed algorithms yield fault tolerant initial patterns and distribution schemes of active modules for several 2D and 3D configurations with Mori, while retaining generalizability for a large number of modular self-folding robots.

v2026.09.13