Arrow Research search

Author name cluster

Dong An

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

9 papers
2 author rows

Possible papers

9

EAAI Journal 2026 Journal Article

A dual branch fusion network for self-supervised sonar image despeckling

  • Yunhong Duan
  • Shubin Zhang
  • Yaoguang Wei
  • Dong An
  • Jincun Liu
  • Yan Meng

Sonar image despeckling is necessary for downstream scene understanding tasks by providing high signal-to-noise ratio despeckled images. Current despeckling methods often require clean images for supervision, however, clean sonar images are inaccessible. Therefore, we propose a self-supervised method based on blind spot network, only using single noisy image for training. In addition, due to the strong spatial correlation of speckle, speckle noise blends with the fine textures of image, making it challenging to keep fine textures while removing speckles. To address this issue, we design a novel dual branch fusion blind spot network that despeckles flat and textual areas separately to maintain fine textures while removing speckles. During training, the flat area is smoothed by the flat branch utilizing a large blind spot excluding more correlated noisy pixels. In contrast, the details in the textual area are recovered by the texture branch employing a small blind spot which incorporates more adjacent pixels for prediction. The flat and texture branches are supervised by a novel spatially adaptive loss to enhance feature extraction. In addition, a variable blind spot is utilized to further balance the capabilities of speckle suppression and detail preservation. Extensive experiments on the synthetic, sidescan, and forward-looking sonar datasets demonstrate that our method achieves superior performances in balancing the speckle suppression and detail preservation compared to state-of-the-art methods, producing high-quality despeckled images while removing visible speckles. The proposed method may improve underwater perception by providing high quality acoustic images, hence advancing ocean exploration.

EAAI Journal 2026 Journal Article

Joint prediction of state-of-charge and state-of-energy using a multiple channels bidirectional long short-term memory and inverted transformer framework

  • Haichi Huang
  • Chong Bian
  • Dong An
  • Shunkun Yang

The state-of-charge (SOC) and state-of-energy (SOE) are critical parameters for lithium batteries, essential for extending battery life and enhancing system reliability. These parameters are highly correlated, yet there is limited research on their joint estimation. This paper proposes a joint deep learning framework, called Multi-channel Bidirectional Long Short-term Memory and Inverted Transformer (MLIT), for SOC and SOE prediction. MLIT utilizes an inverted transformer to separately extract the features of each physical quantity, allowing for better extraction of independent feature information and avoiding alignment issues caused by different sampling frequencies. Additionally, the unique encoder–decoder structure of MLIT enables the model to effectively integrate historical SOC and SOE data and learn the deep correlation between SOC and SOE. Experimental results demonstrate that MLIT achieves higher prediction accuracy compared to current mainstream models, and its joint estimation approach results in more stable prediction outcomes with smaller maximum errors.

AAAI Conference 2025 Conference Paper

Learning Fine-Grained Alignment for Aerial Vision-Dialog Navigation

  • Yifei Su
  • Dong An
  • Kehan Chen
  • Weichen Yu
  • Baiyang Ning
  • Yonggen Ling
  • Yan Huang
  • Liang Wang

Aerial Vision-Dialog Navigation (AVDN) is a new task that requires drones to navigate to a target location based on human-robot dialog history. This paper focuses on the critical fine-grained cross-modal alignment problem in AVDN, requiring the drone to align language entities with visual landmarks in top-down views. To achieve this, we first construct a Fine-Grained AVDN (FG-AVDN) dataset via a semi-automatic annotation pipeline, providing diverse multimodal annotations at the entity-landmark level. Based on this, a novel Fine-grained Entity-Landmark Alignment (FELA) method is proposed to learn the cross-modal alignment explicitly. Concretely, FELA first boosts the drone's visual understanding with a precise semantic grid representation, which captures the environmental semantics and spatial structure simultaneously. Subsequently, to learn the entity-landmark alignment, we devise cross-modal auxiliary tasks from three perspectives, including grounding, captioning, and contrastive learning. Extensive experiments demonstrate that our explicit entity-landmark alignment learning is beneficial for AVDN. As a result, FELA achieves leading performance with 3.2% SR and 4.9% GP improvements over prior arts. Code and dataset will be publicly available.

NeurIPS Conference 2025 Conference Paper

NavBench: Probing Multimodal Large Language Models for Embodied Navigation

  • Yanyuan Qiao
  • Haodong Hong
  • Wenqi Lyu
  • Dong An
  • Siqi Zhang
  • Yutong Xie
  • Xinyu Wang
  • Qi Wu

Multimodal Large Language Models (MLLMs) have demonstrated strong generalization in vision-language tasks, yet their ability to understand and act within embodied environments remains underexplored. We present NavBench, a benchmark to evaluate the embodied navigation capabilities of MLLMs under zero-shot settings. NavBench consists of two components: (1) navigation comprehension, assessed through three cognitively grounded tasks including global instruction alignment, temporal progress estimation, and local observation-action reasoning, covering 3, 200 question-answer pairs; and (2) step-by-step execution in 432 episodes across 72 indoor scenes, stratified by spatial, cognitive, and execution complexity. To support real-world deployment, we introduce a pipeline that converts MLLMs' outputs into robotic actions. We evaluate both proprietary and open-source models, finding that GPT-4o performs well across tasks, while lighter open-source models succeed in simpler cases. Results also show that models with higher comprehension scores tend to achieve better execution performance. Providing map-based context improves decision accuracy, especially in medium-difficulty scenarios. However, most models struggle with temporal understanding, particularly in estimating progress during navigation, which may pose a key challenge.

ECAI Conference 2024 Conference Paper

Large Language Models Understand Layout

  • Weiming Li 0001
  • Manni Duan
  • Dong An
  • Yan Shao

Large language models (LLMs) demonstrate extraordinary abilities in a wide range of natural language processing (NLP) tasks. In this paper, we show that, beyond text understanding capability, LLMs are capable of processing text layouts that are denoted by spatial markers. They are able to answer questions that require explicit spatial perceiving and reasoning, while a drastic performance drop is observed when the spatial markers from the original data are excluded. We perform a series of experiments with the GPT-3. 5/4, Baichuan2, Llama2, and ChatGLM3 models on various types of layout-sensitive datasets for further analysis. The experimental results reveal that the layout understanding ability of LLMs is mainly introduced by the coding data for pre-training, which is further enhanced at the instruction-tuning stage. In addition, layout understanding can be enhanced by integrating low-cost, auto-generated data approached by a novel text game. Finally, we show that layout understanding ability is beneficial for building efficient visual question-answering (VQA) systems.

EAAI Journal 2024 Journal Article

Maize seed fraud detection based on hyperspectral imaging and one-class learning

  • Liu Zhang
  • Yaoguang Wei
  • Jincun Liu
  • Dong An
  • Jianwei Wu

Premium maize varieties are the focus of attention of farmers, breeders, food manufacturers, and people in other industries. Maize seed fraud causes huge financial losses to these industries and many varieties are difficult to distinguish due to their similar appearance. Hyperspectral imaging, as a powerful tool for rapid non-destructive testing, combined with traditional pattern recognition/classification algorithms, has had many successful reports in seed variety identification. In practice, however, fake varieties are too complex to enumerate and a new fake variety may be developed at any time, posing a serious challenge to existing variety identification models. In view of this, this paper proposes a deep one-class learning (OCL) network for seed fraud detection. Specifically, it trains a hypersphere with minimum-volume that can enclose the real variety and isolate all fake varieties outside the hypersphere. In order to improve the performance and stability of the model, the spectral and spatial information of seeds are fused, and a band attention module is used to amplify the weights of the effective bands to suppress the interference from redundant bands. The experimental results show that our method has a mean accuracy of 93. 70% for receiving real varieties and 94. 28% for rejecting fake varieties, which is superior to several existing state-of-the-art OCL models.

EAAI Journal 2023 Journal Article

Boosting fish counting in sonar images with global attention and point supervision

  • Yunhong Duan
  • Shubin Zhang
  • Yang Liu
  • Jincun Liu
  • Dong An
  • Yaoguang Wei

Automatically counting fish in sonar images has been attracting increasing attention in recent years because extreme efforts are needed in manual counting. Density map regression provides a promising approach in the counting field, but two obstacles are placed in front of fish counting in low resolution sonar images: the difficulty in distinguishing fish from the similar background noise and the inconsistency between the strip-shaped fishes in input images and dot-shaped ground truth density map. To address these issues, we present GPNet, a novel encoder-decoder network with global attention and point supervision, to boost sonar image-based fish counting accuracy. To alleviate the impact of background noise, we incorporate a segmentation module (SM) with global self-attention to the neck of the network to identify the fish region and space out background noise. Furthermore, feature enhancement modules (FEM) with a global receptive field are introduced to the encoder to enhance the feature representation and discrimination. To break down the performance upper bound resulting from target shape inconsistency between input and ground truth, we leverage fish center coordinates instead of the Gaussian density map to supervise the network training directly. Extensive experiments on a challenging public sonar image-based fish counting dataset, the ARIS dataset, demonstrate that GPNet achieves state-of-the-art performance both in counting accuracy and noise removal.

NeurIPS Conference 2021 Conference Paper

Landmark-RxR: Solving Vision-and-Language Navigation with Fine-Grained Alignment Supervision

  • Keji He
  • Yan Huang
  • Qi Wu
  • Jianhua Yang
  • Dong An
  • Shuanglin Sima
  • Liang Wang

In Vision-and-Language Navigation (VLN) task, an agent is asked to navigate inside 3D indoor environments following given instructions. Cross-modal alignment is one of the most critical challenges in VLN because the predicted trajectory needs to match the given instruction accurately. In this paper, we address the cross-modal alignment challenge from the perspective of fine-grain. Firstly, to alleviate weak cross-modal alignment supervision from coarse-grained data, we introduce a human-annotated fine-grained VLN dataset, namely Landmark-RxR. Secondly, to further enhance local cross-modal alignment under fine-grained supervision, we investigate the focal-oriented rewards with soft and hard forms, by focusing on the critical points sampled from fine-grained Landmark-RxR. Moreover, to fully evaluate the navigation process, we also propose a re-initialization mechanism that makes metrics insensitive to difficult points, which can cause the agent to deviate from the correct trajectories. Experimental results show that our agent has superior navigation performance on Landmark-RxR, en-RxR and R2R. Our dataset and code are available at https: //github. com/hekj/Landmark-RxR.

JBHI Journal 2020 Journal Article

Automatic Identification of Breast Ultrasound Image Based on Supervised Block-Based Region Segmentation Algorithm and Features Combination Migration Deep Learning Model

  • Wen-Xuan Liao
  • Ping He
  • Jin Hao
  • Xuan-Yu Wang
  • Ruo-Lin Yang
  • Dong An
  • Li-Gang Cui

Breast cancer is a high-incidence type of cancer for women. Early diagnosis plays a crucial role in the successful treatment of the disease and the effective reduction of deaths. In this paper, deep learning technology combined with ultrasound imaging diagnosis was used to identify and determine whether the tumors were benign or malignant. First, the tumor regions were segmented from the breast ultrasound (BUS) images using the supervised block-based region segmentation algorithm. Then, a VGG-19 network pretrained on the ImageNet dataset was applied to the segmented BUS images to predict whether the breast tumor was benign or malignant. The benchmark data for bio-validation were obtained from 141 patients with 199 breast tumors, including 69 cases of malignancy and 130 cases of benign tumors. The experiment showed that the accuracy of the supervised block-based region segmentation algorithm was almost the same as that of manual segmentation; therefore, it can replace manual work. The diagnostic effect of the combination feature model established based on the depth feature of the B-mode ultrasonic imaging and strain elastography was better than that of the model established based on these two images alone. The correct recognition rate was 92. 95%, and the AUC was 0. 98 for the combination feature model.

v2026.09.13