Arrow Research search

Author name cluster

Yan Song

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

30 papers
2 author rows

Possible papers

30

EAAI Journal 2025 Journal Article

A global information-guided denoising diffusion probabilistic model for fault diagnosis with imbalanced data

  • Han Yang
  • Yan Song
  • Daichao Wang
  • Jinghao Xing
  • Yibin Li

Data imbalance is a significant challenge in intelligent fault diagnosis. Diffusion models offer potential solutions to data imbalance by generating samples from limited data. However, their limited global inference capability can create discrepancies between the generated and real data. To address this issue, this paper introduces a global information-guided denoising diffusion probabilistic model (GI-DDPM) for intelligent fault diagnosis. GI-DDPM employs a pretrained Vision Transformer (ViT) as a guide for global information. It extracts the global information from the data through ViT and aligns it with features extracted by the diffusion model. The cosine similarity between the two sets of features is calculated and used as a regularization term to guide the diffusion model to pay more attention to the comprehensive features in the data. This approach encourages the model to generate fault data with more comprehensive representations. The proposed method is tested on two datasets with different balance ratios for imbalanced fault diagnosis experiments, achieving average accuracy rates of 96. 95%, 97. 18% and 96. 06%. These results outperform other intelligent methods, demonstrating significant potential in handling imbalanced fault diagnosis.

YNIMG Journal 2025 Journal Article

Atypical visual selective attention in children with dyslexia: evidence from N2pc and PD

  • Hongyu Liu
  • Yulu Sun
  • Zimo Wang
  • Jialiang Guo
  • Yan Song
  • Xiangzhi Meng

Efficient visual attention is fundamental to the development of reading abilities. Previous studies have identified visual attention impairments in individuals with dyslexia, particularly in visual attention span and sluggish attention shifting. However, findings regarding basic visual selective attention remain controversial. To address this issue, the present study provides event-related potential (ERP) evidence to verify whether children with developmental dyslexia (DD) suffered from deficits in visual selective attention. A pop-out visual search task was used to examine visual attentional patterns to a lateral target in 66 children (33 children with DD and 33 typically developing (TD) children) using electroencephalography. Compared to TD children, children with DD showed a larger and prolonged P1 component, as well as a larger target-evoked N2-posterior contralateral (N2pc) component. The larger P1 amplitude was associated with lower reading fluency while the larger N2pc amplitude was associated with poor reading comprehension performance. Interestingly, children with DD were characterised by a larger target-elicited N2pc followed by a diminished distractor positivity (PD) component, as well as a robust correlation between N2pc and PD, compared to TD children. Children with DD show immature early attentional processing and may have an imbalance in the regulation of visual selective attention between attentional selection and suppression. Our findings provide ERP evidence for interpreting the underlying neural mechanism of visual selective attention in children with DD and may offer a new perspective for exploring the mechanism of visual attention deficits in dyslexia.

YNIMG Journal 2025 Journal Article

Development of the relationship between visual selective attention and auditory change detection

  • Yuanjun Kong
  • Xuye Yuan
  • Yiqing Hu
  • Bingkun Li
  • Dongwei Li
  • Jialiang Guo
  • Meirong Sun
  • Yan Song

Understanding the developmental trajectories of the auditory and visual systems is crucial to elucidate cognitive maturation and its associated relationships, which are essential for effectively navigating dynamic environments. Our one recent study has shown a positive correlation between the event-related potential (ERP) amplitudes associated with visual selective attention (posterior contralateral N2) and auditory change detection (mismatch negativity) in adults, suggesting an intimate relationship and potential shared mechanism between visual selective attention and auditory change detection. However, the evolution of these processes and their relationship over time remains unclear. In this study, we recorded electroencephalography signals from 118 participants (42 adults and 76 typically developing children) during separate visual localization and auditory-embedded fixation tasks. Further, we employed both ERP analysis and multivariate pattern machine learning to investigate developmental patterns. ERP amplitude and decoding accuracy provided convergent evidence underlying a linear developmental trajectory for visual selective attention and an inverted U-shaped trajectory for auditory change detection from childhood to adulthood. Importantly, our findings confirmed the established association of an N2 pc-MMN in adults using a larger sample size, and further identified a positive correlation between decoding accuracy for visual target location and decoding accuracy for auditory stimulus type specifically in adults. However, both visual-auditory correlation effects were absent in children. Our study provides neurophysiological insights into the distinct developmental trajectories of visual selective attention and auditory change detection. It highlights that the close relationship between individual differences in the two processes emerges alongside their respective maturation and does not become evident until adulthood.

ICLR Conference 2025 Conference Paper

Efficient Reinforcement Learning with Large Language Model Priors

  • Xue Yan
  • Yan Song
  • Xidong Feng
  • Mengyue Yang
  • Haifeng Zhang
  • Haitham Bou-Ammar
  • Jun Wang

In sequential decision-making (SDM) tasks, methods like reinforcement learning (RL) and heuristic search have made notable advances in specific cases. However, they often require extensive exploration and face challenges in generalizing across diverse environments due to their limited grasp of the underlying decision dynamics. In contrast, large language models (LLMs) have recently emerged as powerful general-purpose tools, due to their capacity to maintain vast amounts of domain-specific knowledge. To harness this rich prior knowledge for efficiently solving complex SDM tasks, we propose treating LLMs as prior action distributions and integrating them into RL frameworks through Bayesian inference methods, making use of variational inference and direct posterior sampling. The proposed approaches facilitate the seamless incorporation of fixed LLM priors into both policy-based and value-based RL frameworks. Our experiments show that incorporating LLM-based action priors significantly reduces exploration and optimization complexity, substantially improving sample efficiency compared to traditional RL techniques, e.g., using LLM priors decreases the number of required samples by over 90\% in offline learning scenarios.

NeurIPS Conference 2025 Conference Paper

ForceVLA: Enhancing VLA Models with a Force-aware MoE for Contact-rich Manipulation

  • Jiawen Yu
  • Hairuo Liu
  • Qiaojun Yu
  • Jieji Ren
  • Ce Hao
  • Haitong Ding
  • Guangyu Huang
  • Guofan Huang

Vision-Language-Action (VLA) models have advanced general-purpose robotic manipulation by leveraging pretrained visual and linguistic representations. However, they struggle with contact-rich tasks that require fine-grained control involving force, especially under visual occlusion or dynamic uncertainty. To address these limitations, we propose \textbf{ForceVLA}, a novel end-to-end manipulation framework that treats external force sensing as a first-class modality within VLA systems. ForceVLA introduces \textbf{FVLMoE}, a force-aware Mixture-of-Experts fusion module that dynamically integrates pretrained visual-language embeddings with real-time 6-axis force feedback during action decoding. This enables context-aware routing across modality-specific experts, enhancing the robot's ability to adapt to subtle contact dynamics. We also introduce \textbf{ForceVLA-Data}, a new dataset comprising synchronized vision, proprioception, and force-torque signals across five contact-rich manipulation tasks. ForceVLA improves average task success by 23. 2\% over strong $\pi_0$-based baselines, achieving up to 80\% success in tasks such as plug insertion. Our approach highlights the importance of multimodal integration for dexterous manipulation and sets a new benchmark for physically intelligent robotic control. Code and data will be released at https: //sites. google. com/view/forcevla2025/.

IROS Conference 2025 Conference Paper

Noise Fusion-based Distillation Learning for Anomaly Detection in Complex Industrial Environments

  • Jiawen Yu
  • Jieji Ren
  • Yang Chang
  • Qiaojun Yu
  • Xuan Tong
  • Boyang Wang 0003
  • Yan Song
  • You Li

Anomaly detection and localization in automated industrial manufacturing can significantly enhance production efficiency and product quality. Existing methods are capable of detecting surface defects in pre-defined or controlled imaging environments. However, accurately detecting workpiece defects in complex and unstructured industrial environments with varying views, poses and illumination remains challenging. We propose a novel anomaly detection and localization method specifically designed to handle inputs with perturbative patterns. Our approach introduces a new framework based on a collaborative distillation heterogeneous teacher network (HetNet), an adaptive local-global feature fusion module, and a local multivariate Gaussian noise generation module. HetNet can learn to model the complex feature distribution of normal patterns using limited information about local disruptive changes. We conducted extensive experiments on mainstream benchmarks. HetNet demonstrates superior performance with approximately 10% improvement across all evaluation metrics on MSC-AD under industrial conditions, while achieving state-of-the-art results on other datasets, validating its resilience to environmental fluctuations and its capability to enhance the reliability of industrial anomaly detection systems across diverse scenarios. Tests in real-world environments further confirm that HetNet can be effectively integrated into production lines to achieve robust and real-time anomaly detection. Codes, images and videos are published on the project website at: https://zihuatanejoyu.github.io/HetNet/

ICLR Conference 2025 Conference Paper

Reinforcement Learning from Imperfect Corrective Actions and Proxy Rewards

  • Zhao-Hui Jiang
  • Xuening Feng
  • Paul Weng
  • Yifei Zhu
  • Yan Song
  • Tianze Zhou
  • Yujing Hu
  • Tangjie Lv

In practice, reinforcement learning (RL) agents are often trained with a possibly imperfect proxy reward function, which may lead to a human-agent alignment issue (i.e., the learned policy either converges to non-optimal performance with low cumulative rewards, or achieves high cumulative rewards but in an undesired manner). To tackle this issue, we consider a framework where a human labeler can provide additional feedback in the form of corrective actions, which expresses the labeler's action preferences although this feedback may possibly be imperfect as well. In this setting, to obtain a better-aligned policy guided by both learning signals, we propose a novel value-based deep RL algorithm called **I**terative learning from **Co**rrective actions and **Pro**xy rewards (ICoPro), which cycles through three phases: (1) Solicit sparse corrective actions from a human labeler on the agent's demonstrated trajectories; (2) Incorporate these corrective actions into the Q-function using a margin loss to enforce adherence to labeler's preferences; (3) Train the agent with standard RL losses regularized with a margin loss to learn from proxy rewards and propagate the Q-values learned from human feedback. Moreover, another novel design in our approach is to integrate pseudo-labels from the target Q-network to reduce human labor and further stabilize training. We experimentally validate our proposition on a variety of tasks (Atari games and autonomous driving on highway). On the one hand, using proxy rewards with different levels of imperfection, our method can better align with human and is more sample-efficient than baseline methods. On the other hand, facing corrective actions with different types of imperfection, our method can overcome the non-optimality of this feedback thanks to the guidance from proxy rewards.

NeurIPS Conference 2025 Conference Paper

ReMA: Learning to Meta-Think for LLMs with Multi-agent Reinforcement Learning

  • Ziyu Wan
  • Yunxiang Li
  • Xiaoyu Wen
  • Yan Song
  • Hanjing Wang
  • Linyi Yang
  • Mark Schmidt
  • Jun Wang

Recent research on Reasoning of Large Language Models (LLMs) has sought to further enhance their performance by integrating meta-thinking—enabling models to monitor, evaluate, and control their reasoning processes for more adaptive and effective problem-solving. However, current single-agent work lacks a specialized design for acquiring meta-thinking, resulting in low efficacy. To address this challenge, we introduce Reinforced Meta-thinking Agents (ReMA), a novel framework that leverages Multi-Agent Reinforcement Learning (MARL) to elicit meta-thinking behaviors, encouraging LLMs to think about thinking. ReMA decouples the reasoning process into two hierarchical agents: a high-level meta-thinking agent responsible for generating strategic oversight and plans, and a low-level reasoning agent for detailed executions. Through iterative reinforcement learning with aligned objectives, these agents explore and learn collaboration, leading to improved generalization and robustness. Empirical results from single-turn experiments demonstrate that ReMA outperforms single-agent RL baselines on complex reasoning tasks, including competitive-level mathematical benchmarks and LLM-as-a-Judge benchmarks. Additionally, we further extend ReMA to multi-turn interaction settings, leveraging turn-level ratio and parameter sharing to improve efficiency. Comprehensive ablation studies further illustrate the evolving dynamics of each distinct agent, providing valuable insights into how the meta-thinking reasoning process enhances the reasoning capabilities of LLMs.

NeurIPS Conference 2025 Conference Paper

ThinkBench: Dynamic Out-of-Distribution Evaluation for Robust LLM Reasoning

  • Shulin Huang
  • Linyi Yang
  • Yan Song
  • Shawn Chen
  • Leyang Cui
  • Ziyu Wan
  • Qingcheng Zeng
  • Ying Wen

Evaluating large language models (LLMs) poses significant challenges, particularly due to issues of data contamination and the leakage of correct answers. To address these challenges, we introduce ThinkBench, a novel evaluation framework designed to robustly evaluate the reasoning capability of LLMs. ThinkBench proposes a dynamic data generation method for constructing out-of-distribution (OOD) datasets and offers an OOD dataset that contains 2, 912 samples drawn from reasoning tasks. ThinkBench unifies the evaluation of reasoning models and non-reasoning models. We evaluate 16 LLMs and 4 PRMs under identical experimental conditions and show that most of the LLMs' performance are far from robust and they face a certain level of data leakage. By dynamically generating OOD datasets, ThinkBench effectively provides a reliable evaluation of LLMs and reduces data contamination impact. Our data and codes are available at https: //github. com/huangshulin123/ThinkBench.

IJCAI Conference 2024 Conference Paper

AI-Olympics: Exploring the Generalization of Agents through Open Competitions

  • Chen Wang
  • Yan Song
  • Shuai Wu
  • Sa Wu
  • Ruizhi Zhang
  • Shu Lin
  • Haifeng Zhang

Between 2021 and 2023, AI-Olympics---a series of online AI competitions, was hosted by the online evaluation platform Jidi in collaboration with the IJCAI committee. In these competitions, an agent is required to accomplish diverse sports tasks in a two-dimensional continuous world, while competing against an opponent. This paper provides a brief overview of the competition series and highlights notable findings. We aim to contribute insights to the field of multi-agent decision-making and explore the generalization of agents through engineering efforts.

EAAI Journal 2024 Journal Article

Attribute reduction algorithms with an anti-noise mechanism for hybrid data based on fuzzy evidence theory

  • Qinli Zhang
  • Yan Song
  • Yichun Peng
  • Zhaowen Li

Attribute reduction can remove data noise and redundancy, thus reducing computational complexity, which is very important for machine learning. Because the difference between nominal attribute values is difficult to measure, attribute reduction for hybrid data faces challenges. In addition, most of the existing methods are sensitive to noise due to the lack of an anti-noise mechanism. Decision attribute contains the most important information of data. This paper proposes some techniques that consider the above problems from the perspective of fuzzy evidence theory. First of all, a new distance incorporating decision attributes is defined, and then a new fuzzy relation with an anti-noise mechanism is defined. Furthermore, fuzzy belief and fuzzy plausibility are defined based on the defined new distance and new fuzzy relation. In this framework, two anti-noise attribute reduction algorithms for hybrid data are proposed. Experiments on 12 data sets of various types show that compared with the other 8 state-of-the-art algorithms, the proposed algorithms improve the classification accuracy by at least 2% and the anti-noise ability by at least 11%. Therefore, it can be concluded that the proposed algorithms have excellent anti-noise ability while maintaining good feature selection ability.

AAMAS Conference 2024 Conference Paper

Boosting Studies of Multi-Agent Reinforcement Learning on Google Research Football Environment: The Past, Present, and Future

  • Yan Song
  • He Jiang
  • Haifeng Zhang
  • Zheng Tian
  • Weinan Zhang
  • Jun Wang

Even though Google Research Football (GRF) was initially benchmarked and studied as a single-agent environment in its original paper [19], recent years have witnessed an increasing focus on its multi-agent nature by researchers utilizing it as a testbed for Multi-Agent Reinforcement Learning (MARL), especially in the cooperative scenarios. However, the absence of standardized environment settings and uni�ed evaluation metrics for multi-agent scenarios hampers the consistent understanding of various studies. Furthermore, the challenging 5 vs 5 and 11 vs 11 full-game scenarios have received limited thorough examination due to their substantial training complexities. To address these gaps, this paper extends the original environment by not only standardizing the environment settings and benchmarking cooperative learning algorithms across di�erent scenarios, including the most challenging full-game scenarios, but also by discussing approaches to enhance football AI from diverse perspectives and introducing related research tools for learning beyond multi-agent cooperation. Speci�cally, we provide a distributed and asynchronous population-based self-play framework with diverse pre-trained policies for faster training, two football-speci�c analytical tools for deeper investigation, and an online leaderboard for broader evaluation. The overall expectation of this work is to advance the study of Multi-Agent Reinforcement Learning both on and with Google Research Football environment, with the ultimate goal of deploying these technologies to real-world applications, such as sports analysis. ∗Equal Contribution †Corresponding Author This work is licensed under a Creative Commons Attribution International 4. 0 License. Proc. of the 23rd International Conference on Autonomous Agents and Multiagent Systems (AAMAS 2024), N. Alechina, V. Dignum, M. Dastani, J. S. Sichman (eds.), May 6 – 10, 2024, Auckland, New Zealand. © 2024 International Foundation for Autonomous Agents and Multiagent Systems (www. ifaamas. org).

AAAI Conference 2024 Conference Paper

Bootstrapping Large Language Models for Radiology Report Generation

  • Chang Liu
  • Yuanhe Tian
  • Weidong Chen
  • Yan Song
  • Yongdong Zhang

Radiology report generation (RRG) aims to automatically generate a free-text description from a specific clinical radiograph, e.g., chest X-Ray images. Existing approaches tend to perform RRG with specific models trained on the public yet limited data from scratch, where they often lead to inferior performance owing to the problem of inefficient capabilities in both aligning visual and textual features and generating informative reports accordingly. Currently, large language models (LLMs) offered a promising solution to text generation with their power in learning from big data, especially for cross-modal scenarios such as RRG. However, most existing LLMs are pre-trained on general data, and suffer from the same problem of conventional approaches caused by knowledge gap between general and medical domain if they are applied to RRG. Therefore in this paper, we propose an approach to bootstrapping LLMs for RRG with a in-domain instance induction and a coarse-to-fine decoding process. Specifically, the in-domain instance induction process learns to align the LLM to radiology reports from general texts through contrastive learning. The coarse-to-fine decoding performs a text elevating process for those reports from the ranker, further enhanced with visual features and refinement prompts. Experimental results on two prevailing RRG datasets, namely, IU X-Ray and MIMIC-CXR, demonstrate the superiority of our approach to previous state-of-the-art solutions. Further analyses illustrate that, for the LLM, the induction process enables it to better align with the medical domain and the coarse-to-fine generation allows it to conduct more precise text generation.

EAAI Journal 2024 Journal Article

Spatial and temporal attention-based and residual-driven long short-term memory networks with implicit features

  • Yameng Zhang
  • Yan Song
  • Guoliang Wei

Long short-term memory (LSTM) network is extensively researched as an effective tool for time series prediction, the addition of spatial attention can portray the spatial relationship between inputs and outputs. However, the input features are not extracted adequately and are weakly representative, resulting in models with inadequate learning ability, and the cumulative errors lead to inaccurate output results. To surmount the above matters, a novel improved LSTM-based model, namely residual-driven and spatial-attentive LSTM networks with implicit features, is proposed to forecast time series. At first, abundant implicit features are extracted by means of the convolutional layer to enhance the learning ability of the model. Then, the preliminary predictive results are derived by utilizing the spatial–temporal attention-based LSTM module. Subsequently, a new kernel ridge regression (KRR)-based residual-driven module is designed to correct the above results to further develop the forecasting performance in a low time-consuming. Finally, the experiments on public time series datasets from different fields depict the effective performance of the designed method in contrast to other tested models.

AAMAS Conference 2024 Conference Paper

TaxAI: A Dynamic Economic Simulator and Benchmark for Multi-agent Reinforcement Learning

  • Qirui Mi
  • Siyu Xia
  • Yan Song
  • Haifeng Zhang
  • Shenghao Zhu
  • Jun Wang

Taxation and government spending are crucial tools for governments to promote economic growth and maintain social equity. However, the difficulty in accurately predicting the dynamic strategies of diverse self-interested households presents a challenge for governments to implement effective tax policies. Given its proficiency in modeling other agents in partially observable environments and adaptively learning to find optimal policies, Multi-Agent Reinforcement Learning (MARL) is highly suitable for solving dynamic games between the government and numerous households. Although MARL shows more potential than traditional methods such as the genetic algorithm and dynamic programming, there is a lack of large-scale multi-agent reinforcement learning economic simulators. Therefore, we propose a MARL environment, named TaxAI, for dynamic games involving 𝑁 households, government, firms, and financial intermediaries based on the Bewley-Aiyagari economic model. Our study benchmarks 2 traditional economic methods with 7 MARL methods on TaxAI, demonstrating the effectiveness and superiority of MARL algorithms. Moreover, TaxAI’s scalability in simulating dynamic interactions between the government and 10, 000 households, coupled with real-data calibration, grants it a substantial improvement in scale and reality over existing simulators. Therefore, TaxAI is the most realistic economic simulator for optimal tax policy, which aims to generate feasible recommendations for governments and individuals.

EAAI Journal 2023 Journal Article

A feature-enhanced long short-term memory network combined with residual-driven ν support vector regression for financial market prediction

  • Yameng Zhang
  • Yan Song
  • Guoliang Wei

In recent years, the long short-term memory (LSTM) network has gained special attention in the investigation of financial market forecasting since it has a good ability to mine crucial information from time series data via network learning. However, most existing LSTM networks cannot perform well on the small number of samples and usually have a weak feature extraction and inadequate use of information. To address the above issues, a novel LSTM network, namely feature-enhanced LSTM network combined with residual-driven ν support vector regression, is put forward. Such a proposed LSTM network has the following two whelming merits: (1) the convolution layers are utilized to extract crucial profitable latent features and then the LSTM is applied to gain the rough prediction by using both long-term and short-term information; (2) a residual-driven ν support vector regression ( ν SVR) model is developed to make an promotion over the rough prediction by taking full consideration of historical information. Finally, extensive experiments in real-world datasets demonstrate the desirable results of the proposed method as opposed to other baseline models.

ICLR Conference 2023 Conference Paper

Neural Episodic Control with State Abstraction

  • Zhuo Li 0021
  • Derui Zhu
  • Yujing Hu
  • Xiaofei Xie
  • Lei Ma 0003
  • Yan Zheng 0002
  • Yan Song
  • Yingfeng Chen

Existing Deep Reinforcement Learning (DRL) algorithms suffer from sample inefficiency. Generally, episodic control-based approaches are solutions that leverage highly rewarded past experiences to improve sample efficiency of DRL algorithms. However, previous episodic control-based approaches fail to utilize the latent information from the historical behaviors (\eg, state transitions, topological similarities, \etc) and lack scalability during DRL training. This work introduces Neural Episodic Control with State Abstraction (NECSA), a simple but effective state abstraction-based episodic control containing a more comprehensive episodic memory, a novel state evaluation, and a multi-step state analysis. We evaluate our approach to the MuJoCo and Atari tasks in OpenAI gym domains. The experimental results indicate that NECSA achieves higher sample efficiency than the state-of-the-art episodic control-based approaches. Our data and code are available at the project website\footnote{\url{https://sites.google.com/view/drl-necsa}}.

YNIMG Journal 2023 Journal Article

Prioritizing flexible working memory representations through retrospective attentional strengthening

  • Dongwei Li
  • Yiqing Hu
  • Mengdi Qi
  • Chenguang Zhao
  • Ole Jensen
  • Jing Huang
  • Yan Song

Previous work has proposed two potential benefits of retrospective attention on working memory (WM): target strengthening and non-target inhibition. It remains unknown which hypothesis contributes to the improved WM performance, yet the neural mechanisms responsible for this attentional benefit are unclear. Here, we recorded electroencephalography (EEG) signals while 33 participants performed a retrospective-cue WM task. Multivariate pattern classification analysis revealed that only representations of target features were enhanced by valid retrospective attention during retention, supporting the target strengthening hypothesis. Further univariate analysis found that mid-frontal theta inter-trial phase coherence (ITPC) and ERP components were modulated by valid retrospective attention and correlated with individual differences and moment-to-moment fluctuations on behavioral outcomes, suggesting that both trait- and state-level variability in attentional preparatory processes influence goal-directed behavior. Furthermore, task-irrelevant target spatial location could be decoded from EEG signals, indicating that enhanced spatial binding of target representation is vital to high WM precision. Importantly, frontoparietal theta-alpha phase-amplitude coupling was increased by valid retrospective attention and predicted the reduced random guessing rates. This long-range connection supported top-down information flow in the engagement of frontoparietal networks, which might organize attentional states to integrate target features. Altogether, these results provide neurophysiological bases that retrospective attention improves WM precision by enhancing flexible target representation and emphasize the critical role of the frontoparietal attentional network in the control of WM representations.

EAAI Journal 2023 Journal Article

Switching synthesizing-incorporated and cluster-based synthetic oversampling for imbalanced binary classification

  • Jun Dou
  • Zihan Gao
  • Guoliang Wei
  • Yan Song
  • Ming Li

Oversampling is a popular yet useful method to fulfill the binary classification of imbalanced data, however many existing results of oversampling are very likely to generate redundant/unsafe/noise samples due primarily to the inadequate consideration of the data distribution. To address this issue, we propose a novel oversampling approach, namely Switching Synthesizing-Incorporated and Cluster-Based Synthetic Oversampling (SSI-CBSO). The core idea of SSI-CBSO is four-fold: (1) noise samples are removed by using K nearest neighbor strategy and Fuzzy C-Means clustering is adopted for the filtered data in the minority class; (2) the number of samples that need to be synthesized is adaptively assigned to each cluster concerning the inter-class distance and the intra-cluster similarity; (3) to better reflect the data distribution, a new method in terms of the concept of the hypersphere is put forward to measure the cluster density in a high dimensional; and (4) a new principle based on the Mahalanobis distance is provided for a better selection of the target sample. Then, a switching synthesizing strategy is established to guarantee the safety of the synthesized samples. Finally, experiments on 13 binary imbalanced data sets by using five evaluation metrics with four classifiers verify that our proposed SSI-CBSO approach can obtain desirable results.

YNIMG Journal 2021 Journal Article

Failure of resting-state frontal–occipital connectivity in linking visual perception with reading fluency in Chinese children with developmental dyslexia

  • Xiujie Yang
  • Jia Zhang
  • Yaping Lv
  • Fang Wang
  • Guosheng Ding
  • Manli Zhang
  • Xiangzhi Meng
  • Yan Song

It is widely accepted that impairment in visual perception impedes children's reading development, and further studies have demonstrated significant enhancement in reading fluency after visual perceptual training. However, the mechanism of the neural linkage between visual perception and reading is unclear. The purpose of this study was to examine the intrinsic functional relationship between visual perception (indexed by the texture discrimination task,TDT) and reading ability (character reading and reading fluency) in Chinese children with developmental dyslexia (DD) and those with typical development (TD). The resting-state functional connectivity (RSFC) between the primary visual cortex (V1, BA17) and the entire brain was analyzed. In addition, how RSFC maps are associated with TDT performance and reading ability in the DD and TD groups was examined. The results demonstrated that the strength of the RSFC between V1 and the left middle frontal gyrus (LMFG, BA9/BA46) was significantly correlated with both the threshold (SOA) of the TDT and reading fluency in TD children but not in DD children. Moreover, LMFG-V1 resting-state connectivity played a mediating role in the association of visual texture discrimination and reading fluency, but not in character reading, in TD children. In contrast, this mediation was absent in DD children, albeit their strengths of RSFC between V1 and the left middle frontal gyrus (LMFG) were comparable to those for the TD group. These findings indicate that typically developing children use the linkage of the RSFC between the V1 and LMFG for visual perception skills, which in turn promote fluent reading; in contrast, children with dyslexia, who had higher TDT thresholds than TD children, could not take advantage of their frontal-occipital connectivity to improve reading fluency abilities. These findings suggest that visual perception plays an important role in reading skills and that children with developmental dyslexia lack the ability to use their frontal-occipital connectivity to link visual perception with reading fluency.

YNICL Journal 2020 Journal Article

Abnormal modulation of theta oscillations in children with attention-deficit/hyperactivity disorder

  • Jialiang Guo
  • Xiangsheng Luo
  • Bingkun Li
  • Qinyuan Chang
  • Li Sun
  • Yan Song

Previous studies have found that theta activities exhibit posterior lateralized modulation as well as midfrontal event-related synchronization (ERS) during covert visual attention in adults. The present study investigated whether these theta modulations existed in children and whether they were associated with attentional problems in attention-deficit/hyperactivity disorder (ADHD). Electroencephalography signals were recorded from typically developing (TD) children and children with ADHD (TD: n = 24; ADHD: n = 22) while they performed a cued covert visual attention task. The participants responded to a target following a cue designed as human eyes that gazed to the left or right visual field (70% validity). Compared with the TD children, the children with ADHD showed increased midfrontal theta ERS and significant posterior theta lateralization in response to the cues. More importantly, we found that the stronger posterior theta lateralization in the right hemisphere exhibited a positive trial-based correlation with the larger midfrontal theta ERS and predicted lower RT variability at the trial level in the children with ADHD. We suggest that ADHD may be associated with some enhanced systems in the frontal and posterior areas via theta oscillations, which may be involved in the compensatory maturation for their attention deficits in childhood, thereby promoting the stability of behavioral responses.

AAAI Conference 2020 Conference Paper

Coordinated Reasoning for Cross-Lingual Knowledge Graph Alignment

  • Kun Xu
  • Linfeng Song
  • Yansong Feng
  • Yan Song
  • Dong Yu

Existing entity alignment methods mainly vary on the choices of encoding the knowledge graph, but they typically use the same decoding method, which independently chooses the local optimal match for each source entity. This decoding method may not only cause the “many-to-one” problem but also neglect the coordinated nature of this task, that is, each alignment decision may highly correlate to the other decisions. In this paper, we introduce two coordinated reasoning methods, i. e. , the Easy-to-Hard decoding strategy and joint entity alignment algorithm. Specifically, the Easy-to- Hard strategy first retrieves the model-confident alignments from the predicted results and then incorporates them as additional knowledge to resolve the remaining model-uncertain alignments. To achieve this, we further propose an enhanced alignment model that is built on the current state-of-the-art baseline. In addition, to address the many-to-one problem, we propose to jointly predict entity alignments so that the oneto-one constraint can be naturally incorporated into the alignment prediction. Experimental results show that our model achieves the state-of-the-art performance and our reasoning methods can also significantly improve existing baselines.

YNICL Journal 2019 Journal Article

The neural correlations of spatial attention and working memory deficits in adults with ADHD

  • Xiangsheng Luo
  • Jialiang Guo
  • Lu Liu
  • Xixi Zhao
  • Dongwei Li
  • Hui Li
  • Qihua Zhao
  • Yanfei Wang

Working memory impairment is a typical cognitive abnormality in patients with attention-deficit/hyperactivity disorder (ADHD) and is closely related to attention. Exploring the interaction between working memory and attention in patients with ADHD is of great significance for studying the pathological mechanism of this disease. In this study, electrophysiological markers of attention, posterior contralateral N2 (N2pc), and working memory, contralateral delay activity (CDA), were used to explore the relationship between these two cognitive abilities in patients with ADHD. EEG data were collected from adults with ADHD and age-, sex-, and IQ-matched normal controls while performing a classical visuospatial working memory task that consisted of low-load and high-load memory conditions. In different memory load conditions, the memory array elicited a smaller N2pc (220-260 ms) and a smaller CDA (400-800 ms) in adults with ADHD than in normal controls. Further analysis revealed that the reduced CDA amplitude could be significantly predicted by the earlier and reduced N2pc amplitude in adults with ADHD. Moreover, when the number of memory items increased, the increase in N2pc highly predicted the increases in CDA. Our findings illustrate the relationship between spatial working memory and attention ability in ADHD adults from the neurophysiological aspect that reduced working memory is closely related to insufficient attention ability and provide a potential physiological basis for the pathological mechanism of ADHD.

IJCAI Conference 2019 Conference Paper

Unsupervised Neural Aspect Extraction with Sememes

  • Ling Luo
  • Xiang Ao
  • Yan Song
  • Jinyao Li
  • Xiaopeng Yang
  • Qing He
  • Dong Yu

Aspect extraction relies on identifying aspects by discovering coherence among words, which is challenging when word meanings are diversified and processing on short texts. To enhance the performance on aspect extraction, leveraging lexical semantic resources is a possible solution to such challenge. In this paper, we present an unsupervised neural framework that leverages sememes to enhance lexical semantics. The overall framework is analogous to an autoenoder which reconstructs sentence representations and learns aspects by latent variables. Two models that form sentence representations are proposed by exploiting sememes via (1) a hierarchical attention; (2) a context-enhanced attention. Experiments on two real-world datasets demonstrate the validity and the effectiveness of our models, which significantly outperforms existing baselines.

IJCAI Conference 2018 Conference Paper

Complementary Learning of Word Embeddings

  • Yan Song
  • Shuming Shi

Continuous bag-of-words (CB) and skip-gram (SG) models are popular approaches to training word embeddings. Conventionally they are two standing-alone techniques used individually. However, with the same goal of building embeddings by leveraging surrounding words, they are in fact a pair of complementary tasks where the output of one model can be used as input of the other, and vice versa. In this paper, we propose complementary learning of word embeddings based on the CB and SG model. Specifically, one round of learning first integrates the predicted output of a SG model with existing context, then forms an enlarged context as input to the CB model. Final models are obtained through several rounds of parameter updating. Experimental results indicate that our approach can effectively improve the quality of initial embeddings, in terms of intrinsic and extrinsic evaluations.

IJCAI Conference 2018 Conference Paper

Joint Learning Embeddings for Chinese Words and their Components via Ladder Structured Networks

  • Yan Song
  • Shuming Shi
  • Jing Li

The components, such as characters and radicals, of a Chinese word are important sources to help in capturing semantic information of the word. In this paper, we propose a novel framework, namely, ladder structured networks (LSN), which contains three layers representing word, character and radical and learns their embeddings synchronously. LSN captures not only the relations among words, but also the relations among their component characters and radicals, as well as the relations across layers. Each layer in LSN is pluggable so that any particular type of unit (word, character, radical) can be removed and the LSN is thus adjusted for particular types of inputs. In evaluating our framework, we use word similarity as the intrinsic evaluation and part-of-speech tagging and document classification as extrinsic evaluations. Experimental results confirm the validity of our approach and show superiority of our approach over previous work.

YNIMG Journal 2016 Journal Article

Predicting perceptual learning from higher-order cortical processing

  • Fang Wang
  • Jing Huang
  • Yaping Lv
  • Xiaoli Ma
  • Bin Yang
  • Encong Wang
  • Boqi Du
  • Wu Li

Visual perceptual learning has been shown to be highly specific to the retinotopic location and attributes of the trained stimulus. Recent psychophysical studies suggest that these specificities, which have been associated with early retinotopic visual cortex, may in fact not be inherent in perceptual learning and could be related to higher-order brain functions. Here we provide direct electrophysiological evidence in support of this proposition. In a series of event-related potential (ERP) experiments, we recorded high-density electroencephalography (EEG) from human adults over the course of learning in a texture discrimination task (TDT). The results consistently showed that the earliest C1 component (68–84ms), known to reflect V1 activity driven by feedforward inputs, was not modulated by learning regardless of whether the behavioral improvement is location specific or not. In contrast, two later posterior ERP components (posterior P1 and P160–350) over the occipital cortex and one anterior ERP component (anterior P160–350) over the prefrontal cortex were progressively modified day by day. Moreover, the change of the anterior component was closely correlated with improved behavioral performance on a daily basis. Consistent with recent psychophysical and imaging observations, our results indicate that perceptual learning can mainly involve changes in higher-level visual cortex as well as in the neural networks responsible for cognitive functions such as attention and decision making.

TIST Journal 2015 Journal Article

Local Structure-Based Sparse Representation for Face Recognition

  • Fan Liu
  • Jinhui Tang
  • Yan Song
  • Liyan Zhang
  • Zhenmin Tang

This article presents a simple yet effective face recognition method, called local structure-based sparse representation classification (LS_SRC). Motivated by the “divide-and-conquer” strategy, we first divide the face into local blocks and classify each local block, then integrate all the classification results to make the final decision. To classify each local block, we further divide each block into several overlapped local patches and assume that these local patches lie in a linear subspace. This subspace assumption reflects the local structure relationship of the overlapped patches, making sparse representation-based classification (SRC) feasible even when encountering the single-sample-per-person (SSPP) problem. To lighten the computing burden of LS_SRC, we further propose the local structure-based collaborative representation classification (LS_CRC). Moreover, the performance of LS_SRC and LS_CRC can be further improved by using the confusion matrix of the classifier. Experimental results on four public face databases show that our methods not only generalize well to SSPP problem but also have strong robustness to occlusion; little pose variation; and the variations of expression, illumination, and time.

YNIMG Journal 2015 Journal Article

Predicting N2pc from anticipatory HbO activity during sustained visuospatial attention: A concurrent fNIRS–ERP study

  • Jing Huang
  • Fang Wang
  • Yulong Ding
  • Haijing Niu
  • Fenghua Tian
  • Hanli Liu
  • Yan Song

Understanding the properties of attentional control, along with the neural mechanisms subserving them, has long invited intense scrutiny in research groups. However, it has not been demonstrated how the top-down anticipatory hemodynamic activation influences the subsequent attentional processing of targets and distractors. Here, with concurrent fNIRS–ERP recording, we explored the potential contribution of anticipatory oxygenated hemoglobin (HbO) based brain activity to attentional control by examining how HbO influences the subsequent ERP N2pc components assumed to reflect attentional selection. We found that expecting a target led to a larger increase of preparatory HbO response over the visual cortex contralateral to the upcoming target, which was positively correlated with the subsequent target-evoked N2pc amplitude. Further, anticipation concerning the presence of a competing distractor resulted in large and prolonged preparatory HbO signals in the visual cortex contralateral to the distractor, indicating that the salient distractor might be actively suppressed by preparatory top-down attentional control. However, the pre-suppressed distractor still captured part of the attention in the subsequent visual search as revealed by a decrease in the N2pc amplitude, and such a distraction effect on N2pc was negatively correlated with preparatory HbO enhancement contralateral to the anticipated distractor. Overall, each individuals attentional shift to the target and resistance to the distractor measured by ERP is predictable in advance via anticipatory hemodynamic activity in the visual cortex measured by fNIRS.

ICRA Conference 2014 Conference Paper

A synchronous and multi-domain feature extraction method of EEG and sEMG in power-assist rehabilitation robot

  • Yan Song
  • Yihao Du
  • Xiaoguang Wu
  • Xiaoling Chen
  • Ping Xie

To propose a synchronous and multi-domain feature extraction method of electroencephalogram (EEG) and surface electromyogram (sEMG) signals is of great significance to power-assist rehabilitation robot control with humancomputer interface (HCI). In this paper, nonnegative Tucker decomposition which is one model of nonnegative tensor factorization (NTF) is used to fuse two kinds of bioelectricity signals (EEG and sEMG) and extract multi-domain features of EEG and sEMG signals for classification which contain time, frequency, and space domains. In the first step the EEG and sEMG data are transformed into multidimensional information using continuous wavelet transform and the 4-D EEG-sEMG tensor is established. Then the tensor is decomposed into four components (spatial components, spectral components, temporal components and category components) and the core tensor is the feature extracted. The feature after being eliminated and compressed are fed into KNN, LDA and SVM classifiers for pattern recognition, and a comparison is done in single EEG analysis, single sEMG analysis and both EEG and sEMG analysis. An experiment about 10 healthy participants' upper limb movements was carried out to verify the validity of this algorithm. The result implied that NTF is a meaningful and valuable synchronous and multi-domain feature extraction method which may be promising in power-assist rehabilitation robot control.

v2026.09.13