Arrow Research search

Author name cluster

Song Li

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

12 papers
2 author rows

Possible papers

12

EAAI Journal 2026 Journal Article

Application of dunhuang caisson patterns in cultural and creative product design based on artificial intelligence generated content technology

  • Yu Fang
  • Yuyao Zhang
  • Hongjia Tu
  • Wuyang Yang
  • Bin Wu
  • Wenkai Zhu
  • Song Li

The rapid advancement of artificial intelligence generated content (AIGC) provides new technological pathways for digitally preserving and reinventing traditional patterns. This study introduces an integrated framework to explore the creative reproduction of Dunhuang caisson patterns and their application in product design. By coupling the analytic hierarchy process (AHP) with low-rank adaptation (LoRA) fine-tuning, the proposed methodology translates cultural hierarchies into quantitative generative constraints. Firstly, the AHP was employed to establish an evaluation system for the Dunhuang caisson patterns, ranking the key patterns such as lotus and flowers according to their relative importance. Based on these results, representative patterns were collected and preprocessed from published books, academic literature, and open-source datasets to construct a dataset. Subsequently, the weights obtained from the AHP were converted into attention-guided and gradient-guided mechanisms, and the LoRA technique was used to fine-tune three mainstream models (Stable Diffusion 1. 5 (SD 1. 5), Stable Diffusion XL (SDXL), and Flux model). The performance of the fine-tuned models in terms of pattern generation quality and cultural symbol representation was evaluated. Finally, the generated patterns were applied to four representative cultural and creative products (mugs, carpets, handbags, and cushions). The design of the mugs was further examined through Importance-performance analysis (IPA) to verify the practical feasibility and performance of the generated output. This study not only enhances the automated generation quality of Dunhuang caisson patterns but also provides a transferable technical framework for the digital preservation and innovative utilization of traditional patterns in product design.

ICML Conference 2025 Conference Paper

CogMath: Assessing LLMs' Authentic Mathematical Ability from a Human Cognitive Perspective

  • Jiayu Liu 0001
  • Zhenya Huang
  • Wei Dai
  • Cheng Cheng
  • Jinze Wu
  • Jing Sha
  • Song Li
  • Qi Liu 0003

Although large language models (LLMs) show promise in solving complex mathematical tasks, existing evaluation paradigms rely solely on a coarse measure of overall answer accuracy, which are insufficient for assessing their authentic capabilities. In this paper, we propose CogMath, which comprehensively assesses LLMs’ mathematical abilities through the lens of human cognition. Specifically, inspired by psychological theories, CogMath formalizes human reasoning process into 3 stages: problem comprehension, problem solving, and solution summarization. Within these stages, we investigate perspectives such as numerical calculation, knowledge, and counterfactuals, and design a total of 9 fine-grained evaluation dimensions. In each dimension, we develop an “ Inquiry - Judge - Reference ” multi-agent system to generate inquiries that assess LLMs’ mastery from this dimension. An LLM is considered to truly master a problem only when excelling in all inquiries from the 9 dimensions. By applying CogMath on three benchmarks, we reveal that the mathematical capabilities of 7 mainstream LLMs are overestimated by 30%-40%. Moreover, we locate their strengths and weaknesses across specific stages/dimensions, offering in-depth insights to further enhance their reasoning abilities.

ICLR Conference 2025 Conference Paper

Immunogenicity Prediction with Dual Attention Enables Vaccine Target Selection

  • Song Li
  • Yang Tan 0001
  • Song Ke
  • Liang Hong
  • Bingxin Zhou

Immunogenicity prediction is a central topic in reverse vaccinology for finding candidate vaccines that can trigger protective immune responses. Existing approaches typically rely on highly compressed features and simple model architectures, leading to limited prediction accuracy and poor generalizability. To address these challenges, we introduce VenusVaccine, a novel deep learning solution with a dual attention mechanism that integrates pre-trained latent vector representations of protein sequences and structures. We also compile the most comprehensive immunogenicity dataset to date, encompassing over 7000 antigen sequences, structures, and immunogenicity labels from bacteria, viruses, and tumors. Extensive experiments demonstrate that VenusVaccine outperforms existing methods across a wide range of evaluation metrics. Furthermore, we establish a post-hoc validation protocol to assess the practical significance of deep learning models in tackling vaccine design challenges. Our work provides an effective tool for vaccine design and sets valuable benchmarks for future research. The implementation is at \url{https://github.com/songleee/VenusVaccine}.

EAAI Journal 2025 Journal Article

Visual-tactile fusion learning for material recognition based on channel switching and dual cross-attention

  • Song Li
  • Wei Sun
  • Qiaokang Liang
  • Jian Sun
  • Hui Yang
  • YuDong Yang

Visual-tactile multimodal object recognition has attracted increasing attention, as information from different modalities can complement each other and enhance recognition performance. However, the inherent heterogeneity between vision and touch presents a key challenge for effective fusion, limiting accuracy and robustness. To address this, we propose a Multimodal Channel-Switching and Dual Cross-Attention Fusion (MCSDCF) method for integrating multisource data. The channel-switching module adaptively weights and exchanges information across modalities, enabling more discriminative and complementary feature representations. To further capture high-level semantic correlations and heterogeneous cues, we introduce a dual crossattention fusion structure that combines intra-modal self-attention with cross-modal mutual attention, reinforcing the quality of fused representations. Extensive experiments on three public benchmark datasets demonstrate the effectiveness of our MCSDCF framework, with multimodal fusion consistently outperforming unimodal baselines. Ablation studies further confirm the individual contributions of the proposed modules to the overall recognition performance.

ICRA Conference 2024 Conference Paper

A Dragonfly-inspired Flapping Wing Robot Mimicking Force Vector Control Approach

  • Fangyuan Liu
  • Song Li
  • Jinwu Xiang
  • Daochun Li
  • Zhan Tu

Dragonflies show impressive flying skills by achieving both high efficiency and agility. They can perform distinctive flight maneuvers, such as flying backwards, which has proven to be achieved through "force vectoring" mechanism recently. In this paper, to explore the agile flight ability of dragonflies on man-made flapping wing systems, we designed, optimized and fabricated a dragonfly-inspired flapping wing robot (DFWR) with inclinable stroke plane control degrees. The proposed platform employs a four-wing configuration, each of which integrates an extra servo motor to enable the rotation of the flapping plane and imitate the "force vectoring" mechanism. Besides, referring to the flapping kinematics of dragonflies, the installation angle and wing pitch angle of the proposed DFWR are optimized considering the total lift and energy consumption through multiobjective optimization based on NSGA-II method. The "force vector" produced by the proposed platform has been illustrated through both theoretical method and experimental method. Moreover, the feasibility of the design is further verified through a series of operation validation experiments. Such a robot has the potential to provide a highly biomimetic platform to validate the flight mechanism studying of Odonata as well as the relative on-board applications such as bio-inspired vision.

AAAI Conference 2024 Conference Paper

Behavioral Recognition of Skeletal Data Based on Targeted Dual Fusion Strategy

  • Xiao Yun
  • Chenglong Xu
  • Kevin Riou
  • Kaiwen Dong
  • Yanjing Sun
  • Song Li
  • Kevin Subrin
  • Patrick Le Callet

The deployment of multi-stream fusion strategy on behavioral recognition from skeletal data can extract complementary features from different information streams and improve the recognition accuracy, but suffers from high model complexity and a large number of parameters. Besides, existing multi-stream methods using a fixed adjacency matrix homogenizes the model’s discrimination process across diverse actions, causing reduction of the actual lift for the multi-stream model. Finally, attention mechanisms are commonly applied to the multi-dimensional features, including spatial, temporal and channel dimensions. But their attention scores are typically fused in a concatenated manner, leading to the ignorance of the interrelation between joints in complex actions. To alleviate these issues, the Front-Rear dual Fusion Graph Convolutional Network (FRF-GCN) is proposed to provide a lightweight model based on skeletal data. Targeted adjacency matrices are also designed for different front fusion streams, allowing the model to focus on actions of varying magnitudes. Simultaneously, the mechanism of Spatial-Temporal-Channel Parallel Attention (STC-P), which processes attention in parallel and places greater emphasis on useful information, is proposed to further improve model’s performance. FRF-GCN demonstrates significant competitiveness compared to the current state-of-the-art methods on the NTU RGB+D, NTU RGB+D 120 and Kinetics-Skeleton 400 datasets. Our code is available at: https://github.com/sunbeam-kkt/FRF-GCN-master.

NeurIPS Conference 2024 Conference Paper

Virtual Scanning: Unsupervised Non-line-of-sight Imaging from Irregularly Undersampled Transients

  • Xingyu Cui
  • Huanjing Yue
  • Song Li
  • Xiangjun Yin
  • Yusen Hou
  • Yun Meng
  • Kai Zou
  • Xiaolong Hu

Non-line-of-sight (NLOS) imaging allows for seeing hidden scenes around corners through active sensing. Most previous algorithms for NLOS reconstruction require dense transients acquired through regular scans over a large relay surface, which limits their applicability in realistic scenarios with irregular relay surfaces. In this paper, we propose an unsupervised learning-based framework for NLOS imaging from irregularly undersampled transients~(IUT). Our method learns implicit priors from noisy irregularly undersampled transients without requiring paired data, which is difficult and expensive to acquire and align. To overcome the ambiguity of the measurement consistency constraint in inferring the albedo volume, we design a virtual scanning process that enables the network to learn within both range and null spaces for high-quality reconstruction. We devise a physics-guided SURE-based denoiser to enhance robustness to ubiquitous noise in low-photon imaging conditions. Extensive experiments on both simulated and real-world data validate the performance and generalization of our method. Compared with the state-of-the-art (SOTA) method, our method achieves higher fidelity, greater robustness, and remarkably faster inference times by orders of magnitude. The code and model are available at https: //github. com/XingyuCuii/Virtual-Scanning-NLOS.

EAAI Journal 2023 Journal Article

Rotation adaptive grasping estimation network oriented to unknown objects based on novel RGB-D fusion strategy

  • Hongkun Tian
  • Kechen Song
  • Song Li
  • Shuai Ma
  • Yunhui Yan

Accurate grasping estimation is prerequisite and key to achieving accurate robotic grasping. As common data sources, existing RGB and Depth (RGB-D) fusion strategies hardly fully use the advantages and suppress the disadvantages of both modes. In addition, existing methods mainly rely on data augmentation to achieve spatial and rotation adaptation, which cannot fundamentally solve the problem. Therefore, this paper proposes a framework for rotation adaptive grasping estimation based on a novel RGB-D fusion strategy. Specifically, the RGB-D is fused with shared weights in stages based on the proposed Multi-step Weight-learning Fusion (MWF) strategy. The spatial position is encoding learned autonomously based on the proposed Rotation Adaptive Conjoin (RAC) encoder to achieve spatial and rotational adaptiveness oriented to unknown objects with unknown poses. In addition, the Multi-dimensional Interaction-guided Attention (MIA) decoding strategy based on the fused multiscale features is proposed to highlight the practical elements and suppress the invalid ones. The method has been validated on the Cornell and Jacquard grasping datasets with cross-validation accuracies of 99. 3% and 94. 6%. The single-object and multi-object scene grasping success rates on the robot platform are 95. 625% and 87. 5%, respectively. Our performance compares favorably with state-of-the-art methods.

ICRA Conference 2022 Conference Paper

Liftoff of A Motor-Driven Flapping Wing Rotorcraft with Mechanically Decoupled Wings

  • Fangyuan Liu
  • Song Li
  • Ziyu Wang
  • Xin Dong 0020
  • Daochun Li
  • Zhan Tu

Flapping Wing Rotorcraft (FWR) combines flapping and rotating wing motion in one element. Such a hybrid design integrates the high-efficiency characteristics of the rotating wing and the high-lift feature of the flapping wing under low Reynolds number, providing a broader range of simultaneous lift and power efficiency optimization. Nevertheless, the flight performance of the current FWRs is limited by their complex transmission mechanisms. Such mechanical constraints not only induce coupled wing kinematics but also render tedious assembly work and fabrication imperfections. In order to fundamentally address the constraints, we propose a motor-driven FWR with mechanically decoupled wings. The wing of the proposed DFWR is directly actuated by two bi-directional rotating motors instead of using the crank rocker (or alike) transmission. The proposed DFWR flaps within 25Hz to 35Hz, with about 12. 4 grams of system weight and 185mm wingspan. With the direct-drive principle, the wing kinematics can be modulated properly by real-time motor control. In particular, we tuned the flapping frequency, stroke amplitude, and mid-stroke angle of the proposed direct-drive FWR to attain its best lift performance. As a result, it can generate about 16 grams of maximum total lift. In order to validate the proposed design, free flight tests have been conducted. The proposed FWR demonstrates stable liftoff.

IROS Conference 2020 Conference Paper

Crop Height and Plot Estimation for Phenotyping from Unmanned Aerial Vehicles using 3D LiDAR

  • Harnaik Dhami
  • Kevin Yu
  • Tianshu Xu
  • Qian Zhu
  • Kshitiz Dhakal
  • James Friel
  • Song Li
  • Pratap Tokekar

We present techniques to measure crop heights using a 3D Light Detection and Ranging (LiDAR) sensor mounted on an Unmanned Aerial Vehicle (UAV). Knowing the height of plants is crucial to monitor their overall health and growth cycles, especially for high-throughput plant phenotyping. We present a methodology for extracting plant heights from 3D LiDAR point clouds, specifically focusing on plot-based phenotyping environments. We also present a toolchain that can be used to create phenotyping farms for use in Gazebo simulations. The tool creates a randomized farm with realistic 3D plant and terrain models. We conducted a series of simulations and hardware experiments in controlled and natural settings. Our algorithm was able to estimate the plant heights in a field with 112 plots with a root mean square error (RMSE) of 6. 1 cm. This is the first such dataset for 3D LiDAR from an airborne robot over a wheat field. The developed simulation toolchain, algorithmic implementation, and datasets can be found on our GitHub repository. 1

EAAI Journal 2006 Journal Article

Automatic clinical image segmentation using pathological modeling, PCA and SVM

  • Shuo Li
  • Thomas Fevens
  • Adam Krzyżak
  • Song Li

Due to the presence of complicated topological and residual features, the segmentation of medical imagery is a difficult problem. In this paper, an automated approach to clinical image segmentation is presented. The processing of these images in our approach is divided into learning and segmentation stages to facilitate the application of principal component analysis with a support vector machine (SVM) classifier. During the initial learning stage, representative images are chosen to represent typical input images. These images are segmented using a variational level set method driven by a modeled energy functional designed to delineate the pathological characteristics of the images. Then a window-based feature extraction is applied to these segmented images. Principal component analysis is applied to these extracted features and the results are used to train an SVM classifier. After training the SVM, any time a clinical image needs to be segmented, it is simply classified with the trained SVM. By the proposed method, we take the strengths of both machine learning and the variational level set method while limiting their weaknesses to achieve automatic and fast clinical segmentation. To test the proposed system, both chest (thoracic) computed tomography (CT) scans (2D and 3D) and dental X-rays are used. Promising results are demonstrated and analyzed. The proposed method can be used during pre-processing for automatic computer-aided diagnosis.

NeurIPS Conference 1999 Conference Paper

Statistical Dynamics of Batch Learning

  • Song Li
  • K. Y. Michael Wong

An important issue in neural computing concerns the description of learning dynamics with macroscopic dynamical variables. Recen(cid: 173) t progress on on-line learning only addresses the often unrealistic case of an infinite training set. We introduce a new framework to model batch learning of restricted sets of examples, widely applica(cid: 173) ble to any learning cost function, and fully taking into account the temporal correlations introduced by the recycling of the examples. For illustration we analyze the effects of weight decay and early stopping during the learning of teacher-generated examples.

v2026.09.13