Arrow Research search

Author name cluster

Yuntao Chen

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

11 papers
2 author rows

Possible papers

11

NeurIPS Conference 2025 Conference Paper

DriveDPO: Policy Learning via Safety DPO For End-to-End Autonomous Driving

  • Shuyao Shang
  • Yuntao Chen
  • Yuqi Wang
  • Yingyan Li
  • ZHAO-XIANG ZHANG

End-to-end autonomous driving has substantially progressed by directly predicting future trajectories from raw perception inputs, which bypasses traditional modular pipelines. However, mainstream methods trained via imitation learning suffer from critical safety limitations, as they fail to distinguish between trajectories that appear human-like but are potentially unsafe. Some recent approaches attempt to address this by regressing multiple rule-driven scores but decoupling supervision from policy optimization, resulting in suboptimal performance. To tackle these challenges, we propose DriveDPO, a Safety Direct Preference Optimization Policy Learning framework. First, we distill a unified policy distribution from human imitation similarity and rule-based safety scores for direct policy optimization. Further, we introduce an iterative Direct Preference Optimization stage formulated as trajectory-level preference alignment. Extensive experiments on the NAVSIM benchmark demonstrate that DriveDPO achieves a new state-of-the-art PDMS of 90. 0. Furthermore, qualitative results across diverse challenging scenarios highlight DriveDPO’s ability to produce safer and more reliable driving behaviors.

ICLR Conference 2025 Conference Paper

Enhancing End-to-End Autonomous Driving with Latent World Model

  • Yingyan Li
  • Lue Fan
  • Jiawei He 0002
  • Yuqi Wang 0001
  • Yuntao Chen
  • Zhaoxiang Zhang 0001
  • Tieniu Tan

In autonomous driving, end-to-end planners directly utilize raw sensor data, enabling them to extract richer scene features and reduce information loss compared to traditional planners. This raises a crucial research question: how can we develop better scene feature representations to fully leverage sensor data in end-to-end driving? Self-supervised learning methods show great success in learning rich feature representations in NLP and computer vision. Inspired by this, we propose a novel self-supervised learning approach using the LAtent World model (LAW) for end-to-end driving. LAW predicts future latent scene features based on current features and ego trajectories. This self-supervised task can be seamlessly integrated into perception-free and perception-based frameworks, improving scene feature learning while optimizing trajectory prediction. LAW achieves state-of-the-art performance across multiple benchmarks, including real-world open-loop benchmark nuScenes, NAVSIM, and simulator-based closed-loop benchmark CARLA. The code will be released.

EAAI Journal 2025 Journal Article

Fabric defect detection via Explicit De-Background

  • Yuntao Chen
  • Hao Liu
  • Jiuzhen Liang

Fabric defect detection under complex backgrounds faces challenges like high interference, false positives, and lack of robustness. To address these issues, a frequency-domain Explicit De-Background method is proposed to separate background from defects and enhance defect focus. The network uses a pixel-level De-background Layer to suppress noise and emphasize defects after extracting multi-scale features with Swin Transformer. This layer includes the Frequency-domain Background Extraction Module (F-BEM) and the Background Suppression Unit (BSU): F-BEM leverages Fourier amplitude to capture global background features, while BSU suppresses them to produce a difference map that accentuates defect regions. The De-Background Attention (DBA) module leverages the difference map as a weighting matrix to enhance spatial focus on defect features while minimizing background interference. To integrate multi-scale information, the Feature Cross-Shrinking Decoder (FCSD) progressively fuses adjacent layers via Cross Aggregation Nodes(CAN), ensuring semantic consistency, reducing redundancy, and mitigating information loss and gradient vanishing for precise defect segmentation. Our method enhances robustness and accuracy in complex backgrounds using explicit background separation and multi-stage feature processing. It surpasses state-of-the-art techniques on fabric defect datasets and shows good generalization in complex and transfer learning scenarios, providing a practical solution for industrial defect detection.

ICLR Conference 2025 Conference Paper

FreeVS: Generative View Synthesis on Free Driving Trajectory

  • Qitai Wang
  • Lue Fan
  • Yuqi Wang 0001
  • Yuntao Chen
  • Zhaoxiang Zhang 0001

Existing reconstruction-based novel view synthesis methods for driving scenes focus on synthesizing camera views along the recorded trajectory of the ego vehicle. Their image rendering performance will severely degrade on viewpoints falling out of the recorded trajectory, where camera rays are untrained. We propose FreeVS, a novel fully generative approach that can synthesize camera views on free new trajectories in real driving scenes. To control the generation results to be 3D consistent with the real scenes and accurate in viewpoint pose, we propose the pseudo-image representation of view priors to control the generation process. Viewpoint translation simulation is applied on pseudo-images to simulate camera movement in each direction. Once trained, FreeVS can be applied to any validation sequences without reconstruction process and synthesis views on novel trajectories. Moreover, we propose two new challenging benchmarks tailored to driving scenes, which are novel camera synthesis and novel trajectory synthesis, emphasizing the freedom of viewpoints. Given that no ground truth images are available on novel trajectories, we also propose to evaluate the consistency of images synthesized on novel trajectories with 3D perception models. Experiments on the Waymo Open Dataset show that FreeVS has a strong image synthesis performance on both the recorded trajectories and novel trajectories. The code is released. Project page: https://freevs24.github.io/.

AILAW Journal 2025 Journal Article

LLMES: an LLMs-based expert system for quality management system audits

  • Yunhan Li
  • Zeyang Shi
  • Yujing Li
  • Yuntao Chen
  • Gengshen Wu
  • Min Yang

Abstract Many organizations still face significant challenges in document audits, such as contract reviews, accounting audits, and compliance checks, which require substantial manpower and time. Current Artificial Intelligence solutions struggle with ambiguous requirements, unclear standards, and complex documents, limiting their scalability and effectiveness. To address these challenges, we propose a novel architecture, LLMES, designed to generate a concise and independent audit checklist. This approach allows Large Language Models to perform audits and produce results without prior fine-tuning. LLMES effectively integrates complex legal documents, expert knowledge, and mandatory regulatory guidelines with large language modeling to enhance artificial intelligence-assisted auditing. In our experiments auditing a medical device quality management system using a 292-item sampled checklist, LLMES achieved an F1 score of 0. 895 on core datasets, significantly outperforming the 0. 696 score of direct auditing. Additionally, LLMES attained an F1 score of 0. 876 ± 0. 034 across full datasets and 0. 905 on a complete 1696-item checklist. The source code and datasets are publicly available on our GitHub repository: https: //github. com/lyxx3rd/LLMES

NeurIPS Conference 2024 Conference Paper

DrivingDojo Dataset: Advancing Interactive and Knowledge-Enriched Driving World Model

  • Yuqi Wang
  • Ke Cheng
  • Jiawei He
  • Qitai Wang
  • Hengchen Dai
  • Yuntao Chen
  • Fei Xia
  • Zhaoxiang Zhang

Driving world models have gained increasing attention due to their ability to model complex physical dynamics. However, their superb modeling capability is yet to be fully unleashed due to the limited video diversity in current driving datasets. We introduce DrivingDojo, the first dataset tailor-made for training interactive world models with complex driving dynamics. Our dataset features video clips with a complete set of driving maneuvers, diverse multi-agent interplay, and rich open-world driving knowledge, laying a stepping stone for future world model development. We further define an action instruction following (AIF) benchmark for world models and demonstrate the superiority of the proposed dataset for generating action-controlled future predictions.

NeurIPS Conference 2024 Conference Paper

OpenSatMap: A Fine-grained High-resolution Satellite Dataset for Large-scale Map Construction

  • Hongbo Zhao
  • Lue Fan
  • Yuntao Chen
  • Haochen Wang
  • Yuran Yang
  • Xiaojuan Jin
  • Yixin Zhang
  • Gaofeng Meng

In this paper, we propose OpenSatMap, a fine-grained, high-resolution satellite dataset for large-scale map construction. Map construction is one of the foundations of the transportation industry, such as navigation and autonomous driving. Extracting road structures from satellite images is an efficient way to construct large-scale maps. However, existing satellite datasets provide only coarse semantic-level labels with a relatively low resolution (up to level 19), impeding the advancement of this field. In contrast, the proposed OpenSatMap (1) has fine-grained instance-level annotations; (2) consists of high-resolution images (level 20); (3) is currently the largest one of its kind; (4) collects data with high diversity. Moreover, OpenSatMap covers and aligns with the popular nuScenes dataset and Argoverse 2 dataset to potentially advance autonomous driving technologies. By publishing and maintaining the dataset, we provide a high-quality benchmark for satellite-based map construction and downstream tasks like autonomous driving.

NeurIPS Conference 2023 Conference Paper

SheetCopilot: Bringing Software Productivity to the Next Level through Large Language Models

  • Hongxin Li
  • Jingran Su
  • Yuntao Chen
  • Qing Li
  • ZHAO-XIANG ZHANG

Computer end users have spent billions of hours completing daily tasks like tabular data processing and project timeline scheduling. Most of these tasks are repetitive and error-prone, yet most end users lack the skill to automate these burdensome works. With the advent of large language models (LLMs), directing software with natural language user requests become a reachable goal. In this work, we propose a SheetCopilot agent that takes natural language task and control spreadsheet to fulfill the requirements. We propose a set of atomic actions as an abstraction of spreadsheet software functionalities. We further design a state machine-based task planning framework for LLMs to robustly interact with spreadsheets. We curate a representative dataset containing 221 spreadsheet control tasks and establish a fully automated evaluation pipeline for rigorously benchmarking the ability of LLMs in software control tasks. Our SheetCopilot correctly completes 44. 3\% of tasks for a single generation, outperforming the strong code generation baseline by a wide margin. Our project page: https: //sheetcopilot. github. io/.

NeurIPS Conference 2022 Conference Paper

4D Unsupervised Object Discovery

  • Yuqi Wang
  • Yuntao Chen
  • ZHAO-XIANG ZHANG

Object discovery is a core task in computer vision. While fast progresses have been made in supervised object detection, its unsupervised counterpart remains largely unexplored. With the growth of data volume, the expensive cost of annotations is the major limitation hindering further study. Therefore, discovering objects without annotations has great significance. However, this task seems impractical on still-image or point cloud alone due to the lack of discriminative information. Previous studies underlook the crucial temporal information and constraints naturally behind multi-modal inputs. In this paper, we propose 4D unsupervised object discovery, jointly discovering objects from 4D data -- 3D point clouds and 2D RGB images with temporal information. We present the first practical approach for this task by proposing a ClusterNet on 3D point clouds, which is jointly iteratively optimized with a 2D localization network. Extensive experiments on the large-scale Waymo Open Dataset suggest that the localization network and ClusterNet achieve competitive performance on both class-agnostic 2D object detection and 3D instance segmentation, bridging the gap between unsupervised methods and full supervised ones. Codes and models will be made available at https: //github. com/Robertwyq/LSMOL.

JMLR Journal 2019 Journal Article

SimpleDet: A Simple and Versatile Distributed Framework for Object Detection and Instance Recognition

  • Yuntao Chen
  • Chenxia Han
  • Yanghao Li
  • Zehao Huang
  • Yi Jiang
  • Naiyan Wang
  • Zhaoxiang Zhang

Object detection and instance recognition play a central role in many AI applications like autonomous driving, video surveillance and medical image analysis. However, training object detection models on large scale datasets remains computationally expensive and time consuming. This paper presents an efficient and open source object detection framework called SimpleDet which enables the training of state-of-the-art detection models on consumer grade hardware at large scale. SimpleDet covers a wide range of models including both high-performance and high-speed ones. SimpleDet is well-optimized for both low precision training and distributed training and achieves 70% higher throughput for the Mask R-CNN detector compared with existing frameworks. Codes, examples and documents of SimpleDet can be found at https://github.com/tusimple/simpledet. [abs] [ pdf ][ bib ] [ code ] &copy JMLR 2019. ( edit, beta )

AAAI Conference 2018 Conference Paper

DarkRank: Accelerating Deep Metric Learning via Cross Sample Similarities Transfer

  • Yuntao Chen
  • Naiyan Wang
  • Zhaoxiang Zhang

We have witnessed rapid evolution of deep neural network architecture design in the past years. These latest progresses greatly facilitate the developments in various areas such as computer vision and natural language processing. However, along with the extraordinary performance, these state-of-theart models also bring in expensive computational cost. Directly deploying these models into applications with real-time requirement is still infeasible. Recently, Hinton et al. (?) have shown that the dark knowledge within a powerful teacher model can significantly help the training of a smaller and faster student network. These knowledge are vastly beneficial to improve the generalization ability of the student model. Inspired by their work, we introduce a new type of knowledge – cross sample similarities for model compression and acceleration. This knowledge can be naturally derived from deep metric learning model. To transfer them, we bring the “learning to rank” technique into deep metric learning formulation. We test our proposed DarkRank method on various metric learning tasks including pedestrian re-identification, image retrieval and image clustering. The results are quite encouraging. Our method can improve over the baseline method by a large margin. Moreover, it is fully compatible with other existing methods. When combined, the performance can be further boosted.

v2026.09.13