Arrow Research search

Author name cluster

Miao Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

9 papers
1 author row

Possible papers

9

AIIM Journal 2025 Journal Article

Anatomical prior-based vertebral landmark detection for spinal disorder diagnosis

  • Yukang Yang
  • Yu Wang
  • Tianyu Liu
  • Miao Wang
  • Ming Sun
  • Shiji Song
  • Wenhui Fan
  • Gao Huang

As one of fundamental ways to interpret spine images, detection of vertebral landmarks is an informative prerequisite for further diagnosis and management of spine disorders such as scoliosis and fractures. Most existing machine learning-based methods for automatic vertebral landmark detection suffer from overlapping landmarks or abnormally long distances between nearby landmarks against anatomical priors, and thus lack sufficient reliability and interpretability. To tackle the problem, this paper systematically utilizes anatomical prior knowledge in vertebral landmark detection. We explicitly formulate anatomical priors of the spine, related to distances among vertebrae and spatial order within the spine, and integrate these geometrical constraints within training loss, inference procedure, and evaluation metrics. First, we introduce an anatomy-constraint loss to regularize the training process with the aforementioned contextual priors explicitly. Second, we propose a simple-yet-effective anatomy-aided inference procedure by employing sequential prediction rather than a parallel counterpart. Third, we provide novel anatomy-related metrics to quantitatively evaluate to which extent landmark predictions follow the anatomical priors, as is not reflected within the widely-used landmark localization error metric. We employ the localization framework on 1410 anterior–posterior radiographic images. Compared with competitive baseline models, we achieve superior landmark localization accuracy and comparable Cobb angle estimation for scoliosis assessment. Ablation studies demonstrate the effectiveness of designed components on the decrease of localization error and improvement of anatomical plausibility. Additionally, we exhibit effective generalization performance by transferring our detection method onto sagittal 2-D slices of CT scans and boost the performance of downstream compression fracture classification at vertebra-level.

EAAI Journal 2025 Journal Article

Multi-objective deployment optimization for integrated sensing and communication-enabled unmanned aerial vehicle swarm

  • Hongjuan Li
  • Haiyuan Chen
  • Miao Wang
  • Jiahui Li
  • Hui Kang
  • Yuzhuo Guan
  • Xu Lin

With the convergence of mobile communication, sensing, and computational networks in sixth-generation technology, the integration of sensing and communication with unmanned aerial vehicles (UAVs) is promising. This paper focuses on the contribution of artificial intelligence in optimizing the deployment of UAV swarms for multi-objective target detection applications in sixth-generation networks. Specifically, the artificial intelligence contribution lies in the development of an improved multi-objective particle swarm optimization (IMOPSO) algorithm for solving a complex multi-objective deployment problem. The problem aims to simultaneously optimize communication rate, sensing quality, and energy consumption in the deployment of UAV swarms. To address this, the proposed IMOPSO incorporates chaotic initialization, Lévy flight mutation, dynamic mutation rate, and an elimination mechanism based on opposition-based learning. These innovations are designed to enhance the algorithm’s ability to explore the solution space effectively, overcome premature convergence to local solutions, and improve solution quality. In terms of engineering applications, the IMOPSO is applied to the deployment of UAV swarms for target detection, demonstrating its ability to enhance communication and sensing performance while reducing energy consumption in practical scenarios. Through extensive simulations, we show that the IMOPSO outperforms traditional optimization methods and other baseline algorithms, achieving superior results across all optimization objectives. Specifically, the IMOPSO achieves approximately 5% higher transmission data rate, 9% better sensing quality, and 19% lower energy consumption compared to baseline algorithms across multiple test scenarios. Furthermore, the solutions obtained are not only closer to the optimal front but also more concentrated, indicating higher-quality results.

AAAI Conference 2025 Conference Paper

PAT: Pruning-Aware Tuning for Large Language Models

  • Yijiang Liu
  • Huanrui Yang
  • Youxin Chen
  • Rongyu Zhang
  • Miao Wang
  • Yuan Du
  • Li Du

Large language models (LLMs) excel in language tasks, especially with supervised fine-tuning after pre-training. However, their substantial memory and computational requirements hinder practical applications. Structural pruning, which reduces less significant weight dimensions, is one solution. Yet, traditional post-hoc pruning often leads to significant performance loss, with limited recovery from further fine-tuning due to reduced capacity. Since the model fine-tuning refines the general and chaotic knowledge in pre-trained models, we aim to incorporate structural pruning with the fine-tuning, and propose the Pruning-Aware Tuning (PAT) paradigm to eliminate model redundancy while preserving the model performance to the maximum extend. Specifically, we insert the innovative Hybrid Sparsification Modules (HSMs) between the Attention and FFN components to accordingly sparsify the upstream and downstream linear modules. The HSM comprises a lightweight operator and a globally shared trainable mask. The lightweight operator maintains a training overhead comparable to that of LoRA, while the trainable mask unifies the channels to be sparsified, ensuring structural pruning. Additionally, we propose the Identity Loss which decouples the transformation and scaling properties of the HSMs to enhance training robustness. Extensive experiments demonstrate that PAT excels in both performance and efficiency. For example, our Llama2-7b model with a 25% pruning ratio achieves 1.33x speedup while outperforming the LoRA-finetuned model by up to 1.26% in accuracy with a similar training cost.

AAAI Conference 2024 Conference Paper

A Non-parametric Graph Clustering Framework for Multi-View Data

  • Shengju Yu
  • Siwei Wang
  • Zhibin Dong
  • Wenxuan Tu
  • Suyuan Liu
  • Zhao Lv
  • Pan Li
  • Miao Wang

Multi-view graph clustering (MVGC) derives encouraging grouping results by seamlessly integrating abundant information inside heterogeneous data, and has captured surging focus recently. Nevertheless, the majority of current MVGC works involve at least one hyper-parameter, which not only requires additional efforts for tuning, but also leads to a complicated solving procedure, largely harming the flexibility and scalability of corresponding algorithms. To this end, in the article we are devoted to getting rid of hyper-parameters, and devise a non-parametric graph clustering (NpGC) framework to more practically partition multi-view data. To be specific, we hold that hyper-parameters play a role in balancing error item and regularization item so as to form high-quality clustering representations. Therefore, under without the assistance of hyper-parameters, how to acquire high-quality representations becomes the key. Inspired by this, we adopt two types of anchors, view-related and view-unrelated, to concurrently mine exclusive characteristics and common characteristics among views. Then, all anchors' information is gathered together via a consensus bipartite graph. By such ways, NpGC extracts both complementary and consistent multi-view features, thereby obtaining superior clustering results. Also, linear complexities enable it to handle datasets with over 120000 samples. Numerous experiments reveal NpGC's strong points compared to lots of classical approaches.

JBHI Journal 2024 Journal Article

An Efficient Multi-Task Synergetic Network for Polyp Segmentation and Classification

  • Miao Wang
  • Xingwei An
  • Zhengcun Pei
  • Ning Li
  • Li Zhang
  • Gang Liu
  • Dong Ming

Colonoscopy is considered the best diagnostic tool for early detection and resection of polyps, which can effectively prevent consequential colorectal cancer. In clinical practice, segmenting and classifying polyps from colonoscopic images have a great significance since they provide precious information for diagnosis and treatment. In this study, we propose an efficient multi-task synergetic network (EMTS-Net) for concurrent polyp segmentation and classification, and we introduce a polyp classification benchmark for exploring the potential correlations of the above-mentioned two tasks. This framework is composed of an enhanced multi-scale network (EMS-Net) for coarse-grained polyp segmentation, an EMTS-Net (Class) for accurate polyp classification, and an EMTS-Net (Seg) for fine-grained polyp segmentation. Specifically, we first obtain coarse segmentation masks by using EMS-Net. Then, we concatenate these rough masks with colonoscopic images to assist EMTS-Net (Class) in locating and classifying polyps precisely. To further enhance the segmentation performance of polyps, we propose a random multi-scale (RMS) training strategy to eliminate the interference caused by redundant information. In addition, we design an offline dynamic class activation mapping (OFLD CAM) generated by the combined effect of EMTS-Net (Class) and RMS strategy, which optimizes bottlenecks between multi-task networks efficiently and elegantly and helps EMTS-Net (Seg) to perform more accurate polyp segmentation. We evaluate the proposed EMTS-Net on the polyp segmentation and classification benchmarks, and it achieves an average mDice of 0. 864 in polyp segmentation and an average AUC of 0. 913 with an average accuracy of 0. 924 in polyp classification. Quantitative and qualitative evaluations on the polyp segmentation and classification benchmarks demonstrate that our EMTS-Net achieves the best performance and outperforms previous state-of-the-art methods in terms of both efficiency and generalization.

AAAI Conference 2024 Conference Paper

DVSAI: Diverse View-Shared Anchors Based Incomplete Multi-View Clustering

  • Shengju Yu
  • Siwei Wang
  • Pei Zhang
  • Miao Wang
  • Ziming Wang
  • Zhe Liu
  • Liming Fang
  • En Zhu

In numerous real-world applications, it is quite common that sample information is partially available for some views due to machine breakdown or sensor failure, causing the problem of incomplete multi-view clustering (IMVC). While several IMVC approaches using view-shared anchors have successfully achieved pleasing performance improvement, (1) they generally construct anchors with only one dimension, which could deteriorate the multi-view diversity, bringing about serious information loss; (2) the constructed anchors are typically with a single size, which could not sufficiently characterize the distribution of the whole samples, leading to limited clustering performance. For generating view-shared anchors with multi-dimension and multi-size for IMVC, we design a novel framework called Diverse View-Shared Anchors based Incomplete multi-view clustering (DVSAI). Concretely, we associate each partial view with several potential spaces. In each space, we enable anchors to communicate among views and generate the view-shared anchors with space-specific dimension and size. Consequently, spaces with various scales make the generated view-shared anchors enjoy diverse dimensions and sizes. Subsequently, we devise an integration scheme with linear computational and memory expenditures to integrate the outputted multi-scale unified anchor graphs such that running spectral algorithm generates the spectral embedding. Afterwards, we theoretically demonstrate that DVSAI owns linear time and space costs, thus well-suited for tackling large-size datasets. Finally, comprehensive experiments confirm the effectiveness and advantages of DVSAI.

EAAI Journal 2023 Journal Article

A low-cost framework for the recognition of human motion gait phases and patterns based on multi-source perception fusion

  • Dianbiao Dong
  • Chi Ma
  • Miao Wang
  • Huong Thi Vu
  • Bram Vanderborght
  • Yuanxi Sun

Due to the increasing resolution and usage requirements of sensors and the increasing computational cost of algorithms, the application of existing human motion pattern and phase recognition systems in wearable systems is limited. This paper proposes a recognition system framework based on multi-source information fusion for the recognition of human gait patterns and phases. The framework utilizes multi-channel commercial low-cost sensors to obtain human motion information. By applying data fusion methods and preprocessing algorithms, the framework can uniformly process raw signals of uncertain dimensions from an uncertain number of sensors into a valid recognition vector with fixed dimensions. The obtained recognition vector can be used with commonly applied gait recognition algorithms to reduce the computational cost. Support vector machines, Backpropagation neural networks, AlexNet, and LeNet5 algorithms are used to evaluate the performance of the proposed gait recognition framework in recognizing gait phases and patterns. The experimental results show that the four algorithms using fusion signal can achieve a recognition accuracy of 97. 7% for gait phases and an average recognition accuracy of over 99. 2% for gait patterns, which proves the effectiveness of the proposed framework.

IJCAI Conference 2022 Conference Paper

C3-STISR: Scene Text Image Super-resolution with Triple Clues

  • Minyi Zhao
  • Miao Wang
  • Fan Bai
  • Bingjia Li
  • Jie Wang
  • Shuigeng Zhou

Scene text image super-resolution (STISR) has been regarded as an important pre-processing task for text recognition from low-resolution scene text images. Most recent approaches use the recognizer's feedback as clues to guide super-resolution. However, directly using recognition clue has two problems: 1) Compatibility. It is in the form of probability distribution, has an obvious modal gap with STISR - a pixel-level task; 2) Inaccuracy. it usually contains wrong information, thus will mislead the main task and degrade super-resolution performance. In this paper, we present a novel method C3-STISR that jointly exploits the recognizer's feedback, visual and linguistical information as clues to guide super-resolution. Here, visual clue is from the images of texts predicted by the recognizer, which is informative and more compatible with the STISR task; while linguistical clue is generated by a pre-trained character-level language model, which is able to correct the predicted texts. We design effective extraction and fusion mechanisms for the triple cross-modal clues to generate a comprehensive and unified guidance for super-resolution. Extensive experiments on TextZoom show that C3-STISR outperforms the SOTA methods in fidelity and recognition performance. Code is available in https: //github. com/zhaominyiz/C3-STISR.

AAAI Conference 2022 Conference Paper

Rendering-Aware HDR Environment Map Prediction from a Single Image

  • Jun-Peng Xu
  • Chenyu Zuo
  • Fang-Lue Zhang
  • Miao Wang

High dynamic range (HDR) illumination estimation from a single low dynamic range (LDR) image is a significant task in computer vision, graphics and augmented reality. We present a two-stage deep learning-based method to predict an HDR environment map from a single narrow field-of-view LDR image. We first learn a hybrid parametric representation that sufficiently covers high- and low-frequency illumination components in the environment. Taking the estimated illuminations as the guidance, we build a generative adversarial network to synthesize an HDR environment map that enables realistic rendering effects. We specifically consider the rendering effect by supervising the networks using rendering losses in both stages, on the predicted environment map as well as the hybrid illumination representation. Quantitative and qualitative experiments demonstrate that our approach achieves lower relighting errors for virtual object insertion and is preferred by users compared to state-of-the-art methods.

v2026.09.13