Arrow Research search

Author name cluster

Tao Shen

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

22 papers
1 author row

Possible papers

22

JBHI Journal 2026 Journal Article

An OBS-WA Algorithm for Pose Optimization of a Tooth Preparation Robot End-Effector in Confined Spaces

  • Jingang Jiang
  • Zhonghao Xue
  • Jianpeng Sun
  • Chunrui Wang
  • Jingchao Wang
  • Jie Pan
  • Tao Shen

Most traditional instrument pose planning algorithms focus on optimizing the pose of vertical instruments in open spaces. However, there is a lack of research on pose planning for cantilevered instruments in confined environments. In this paper, we propose an innovative method to optimizing instrument pose under multi-objective constraints. The method introduces the concept of a personalized outer outer bounding sphere and defines the safe feasible region for intraoperative instruments based on the Euclidean distance. For optimizing handle orientation during surgery, we propose an algorithm based on a center search strategy, which ensures that the handle orientation solution set avoids interference with adjacent teeth. Additionally, we introduce an improved scheme based on the outer bounding sphere weighted average (OBS-WA) algorithm to optimize robotic arm joint angles, considering multi-objective constraints. One contribution of this study is the development of an improved skeleton-based instrument collision detection method that addresses the limitations of traditional triangular mesh detection in real-time performance. Another innovation lies in solving the multi-objective optimization problem within the oral cavity. By establishing a test system on an experimental platform, this study demonstrates compliance control and safety planning during tooth preparation.

JBHI Journal 2026 Journal Article

Efficient Collaborative Model Training Mechanism With Privacy-Preserving Data for the IoMT

  • Chi Zhang
  • Tao Shen
  • Fenhua Bai
  • Xiaohui Zhang
  • Ziyuan Zhao

As time-series data from the Internet of Medical Things (IoMT) increasingly permeates various aspects of medical research, public governance, and clinical treatment, its sensitivity raises significant privacy concerns, hindering the potential of deep learning applications for cross-institutional data integration. Previous practices focused on deep learning methods based on centralized data storage and processing, which are often unsuitable for decentralized and privacy-sensitive IoMT data scenarios. Most existing methods rely on mechanisms such as trusted coordinators, which face challenges in addressing potential passive data leakage and side-channel attacks, failing to effectively protect the privacy of sensitive data during collaborative training. To address these issues, we propose a privacy-preserving collaborative training model, Secure Long Sequence Time-Series Forecasting (SecLSTF), for IoMT time-series data and design a mapping strategy between model components and Multi-Party Computation (MPC) protocols. Building on this foundation, we propose a novel secret sharing protocol, Pleione, which focuses on optimizing the computational efficiency of the low-level secret-sharing protocol. The protocol centers on a hyper-invertible matrix and adopts a paired double random expansion mechanism, significantly reducing the communication rounds required for random number generation. This optimization enhances the overall training speed of SecLSTF. Subsequently, we replace the original computational support protocol with Pleione. Experimental results show that SecLSTF-Pleione significantly reduces computational time while maintaining computational accuracy, outperforming other protocols in component efficiency. This study offers a potential pathway for cross-institutional IoMT data sharing.

JBHI Journal 2025 Journal Article

Advancing Chinese Conversation-based Patient Guidance with a Benchmark and Knowledge-Evolvable Assistant

  • Wenpeng Lu
  • Kangjun Liu
  • Jianlei Wang
  • Xueping Peng
  • Tao Shen
  • Fa Zhu
  • Weiyu Zhang
  • Jiabing Zhu

Chinese Conversation-based Patient Guidance (CCPG) helps patients reach the correct hospital department through natural-language exchanges with medical staff. Despite the rapid success of large language models (LLMs) in other healthcare tasks, CCPG remains under-explored and lacks dedicated benchmarks. We address this gap with PG-Bench, the first comprehensive CC PG bench mark, spanning five subsets, 19, 814 annotated dialogues, and 98 clinical departments. We evaluate 25 representative LLMs on PG-Bench and observe uniformly poor performance, even the latest models such as GPT-4 and DeepSeek-V3 fail to meet practical requirements. To close this gap, we introduce the Knowledge-Evolvable Assistant (KEA), a novel framework that augments any LLM with (i) an experience bank of validated, successful CCPG cases for analogy-based reasoning; (ii) a reflection bank that records previously misclassified cases together with their corrections and self-summarized error analyses; and (iii) an external medical knowledge base. KEA employs retrieval-augmented generation to evolve its guidance knowledge iteratively. Experiments show that KEA consistently and significantly boosts the CCPG performance of all tested LLMs on PG-Bench. However, current best results still fall short of clinical expectations, underscoring the difficulty of CCPG and the need for further research. PG-Bench and KEA together establish a rigorous foundation and strong baseline for future work on conversation-driven patient guidance in Chinese healthcare settings.

JBHI Journal 2025 Journal Article

BianCang: A Traditional Chinese Medicine Large Language Model

  • Sibo Wei
  • Xueping Peng
  • Yifei Wang
  • Tao Shen
  • Jiasheng Si
  • Weiyu Zhang
  • Fa Zhu
  • Athanasios V. Vasilakos

The surge of large language models (LLMs) has driven significant progress in medical applications, including traditional Chinese medicine (TCM). However, current medical LLMs struggle with TCM diagnosis and syndrome differentiation due to substantial differences between TCM and modern medical theory, and the scarcity of specialized, high-quality corpora. To this end, in this paper we propose BianCang (扁仓) 1, a TCM-specific LLM, using a two-stage training process that first injects domainspecific knowledge and then aligns it through targeted stimulation to enhance diagnostic and differentiation capabilities. Specifically, we constructed pre-training corpora, instruction-aligned datasets based on real hospital records, and the ChP-TCM dataset derived from the Pharmacopoeia of the People's Republic of China. We compiled extensive TCM and medical corpora for continual pre-training and supervised fine-tuning, building a comprehensive dataset to refine the model's understanding of TCM. Evaluations across 11 test sets involving 31 models and 4 tasks demonstrate the effectiveness of BianCang, offering valuable insights for future research. Code, datasets, and models are available on GitHub.

JBHI Journal 2025 Journal Article

Boundary-Enhanced $U^{2}$-Net for Simultaneous Four-Chamber Segmentation in Transthoracic Echocardiography

  • Yuanqin Meng
  • Shengjie Chai
  • Haoyu Xiao
  • Zhaohui Meng
  • Qingwang Wang
  • Tao Shen

The heart, responsible for circulating blood throughout our body, contains four chambers. Existing analysis methods primarily focus on one single ventricle. Transthoracic echocardiography provides real-time estimations of cardiac function and enables comprehensive observations of the entire heart, especially through the apical 4-chamber view. However, no current clinical indices evaluate cardiac function considering all four chambers simultaneously. Manual estimation of the four chambers is laborious, inefficient, and complicated by anatomical complexity and variable image quality, including motion artifacts and unclear borders. There is a significant need for a high-performance segmentation tool that can assess all four chambers concurrently. To address this, we collected a clinically representative dataset of 2D apical 4-chamber view echocardiograms, with annotated 4-chamber regions serving as the basis for automatic 4-chamber synergy analysis. We then proposed a boundary-enhanced network, denoted as $BeU^{2}$ -Net, tailored for transthoracic echocardiography 4-chamber segmentation using our private dataset. Specifically, our network employs a two-level nested encoder-decoder architecture, utilizing a segmentation-specific residual U-block with a mixture of receptive fields at each stage to capture multi-level and multi-scale features. A dedicated boundary prediction branch, incorporating edge details, is integrated to enhance boundary segmentation performance. Experiments on both private and public datasets demonstrate that our $BeU^{2}$ -Net possesses superior boundary detection capabilities and achieves high segmentation performance for echocardiographic images.

AIIM Journal 2025 Journal Article

Enhancing diagnosis prediction with adaptive disease representation learning

  • Hengliang Cheng
  • Shibo Li
  • Tao Shen
  • Weihua Li

Diagnosis prediction predicts which diseases a patient is most likely to suffer from in the future based on their historical electronic health records. The time series model can better capture the temporal progression relationship of patient diseases, but ignores the semantic correlation between all diseases; in fact, multiple diseases that are often diagnosed at the same time reflect hidden patterns that are conducive to diagnosis, so predefined global disease co-occurrence graph can help the model understand disease relationships. But it may contain a lot of noise and ignore the semantic adaptation of the disease under the diagnosis target. To this end, we propose a graph-driven end-to-end framework, named A daptive D isease R epresentation L earning (ADRL), obtain disease representation after learning complex disease relationships, and then use it to improve diagnosis prediction performance. This model introduces an adaptive mechanism to dynamically adjust and optimize disease relationships by performing self-supervised perturbations on a predefined global disease co-occurrence graph, thereby learning a global disease relationship graph that contains complex semantic association information between diseases. The computational burden of adaptive global disease graph can be further alleviated by the proposed SVD-based accelerator. Finally, experimental results on two real-world EHR datasets show that the proposed model outperforms existing models in diagnosis prediction.

EAAI Journal 2025 Journal Article

Enhancing long-term load forecasting with convolutional informer-based hybrid model

  • Bin Sun
  • Xudong Chen
  • Tao Shen
  • Liyao Ma

Long-term load forecasting (LTLF) is essential for energy management but challenged by the complexity of non-stationary time series. The Informer model struggles to capture localized peak–valley patterns, while Variational Mode Decomposition (VMD) faces issues with feature complexity. This study proposes a hybrid framework integrating VMD, Informer, and a Convolutional Long Short-Term Memory (CNN-LSTM) module for accurate LTLF. VMD decomposes non-stationary load data into multi-scale intrinsic mode functions, refined through spectral and autocorrelation analyses to ensure robust feature extraction. The Informer employs sparse self-attention for efficient long-sequence modeling, with CNN-LSTM enhancing the decoder to capture localized temporal dynamics. Experiments on non-stationary load time series across multiple prediction horizons demonstrate that the proposed framework significantly improves forecasting accuracy and robustness compared to baseline models, including Informer and its derivatives. By excelling in complex load pattern prediction, the framework supports efficient grid scheduling and resource optimization in energy systems.

AAAI Conference 2025 Conference Paper

FedCFA: Alleviating Simpson’s Paradox in Model Aggregation with Counterfactual Federated Learning

  • Zhonghua Jiang
  • Jimin Xu
  • Shengyu Zhang
  • Tao Shen
  • Jiwei Li
  • Kun Kuang
  • Haibin Cai
  • Fei Wu

Federated learning (FL) is a promising technology for data privacy and distributed optimization, but it suffers from data imbalance and heterogeneity among clients. Existing FL methods try to solve the problems by aligning client with server model or by correcting client model with control variables. These methods excel on IID and general Non-IID data but perform mediocrely in Simpson's Paradox scenarios. Simpson's Paradox refers to the phenomenon that the trend observed on the global dataset disappears or reverses on a subset, which may lead to the fact that global model obtained through aggregation in FL does not accurately reflect the distribution of global data. Thus, we propose FedCFA, an novel FL framework employing counterfactual learning to generate counterfactual samples by replacing local data critical factors with global average data, aligning local data distributions with the global and mitigating Simpson's Paradox effects. In addition, to improve the counterfactual samples quality, we introduce factor decorrelation (FDC) loss to reduce the correlation among features and thus improve the independence of extracted factors. We conduct extensive experiments on six datasets and verify that our method outperforms other FL methods in terms of efficiency and global model accuracy under limited communication rounds.

NeurIPS Conference 2025 Conference Paper

FlowerTune: A Cross-Domain Benchmark for Federated Fine-Tuning of Large Language Models

  • Yan Gao
  • Massimo R. Scamarcia
  • Javier Fernandez-Marques
  • Mohammad Naseri
  • Chong Ng
  • Dimitris Stripelis
  • Zexi Li
  • Tao Shen

Large Language Models (LLMs) have achieved state-of-the-art results across diverse domains, yet their development remains reliant on vast amounts of publicly available data, raising concerns about data scarcity and the lack of access to domain-specific, sensitive information. Federated Learning (FL) presents a compelling framework to address these challenges by enabling decentralized fine-tuning on pre-trained LLMs without sharing raw data. However, the compatibility and performance of pre-trained LLMs in FL settings remain largely under explored. We introduce the FlowerTune LLM Leaderboard, a first-of-its-kind benchmarking suite designed to evaluate federated fine-tuning of LLMs across four diverse domains: general NLP, finance, medical, and coding. Each domain includes federated instruction-tuning datasets and domain-specific evaluation metrics. Our results, obtained through a collaborative, open-source and community-driven approach, provide the first comprehensive comparison across 26 pre-trained LLMs with different aggregation and fine-tuning strategies under federated settings, offering actionable insights into model performance, resource constraints, and domain adaptation. This work lays the foundation for developing privacy-preserving, domain-specialized LLMs for real-world applications.

AAAI Conference 2025 Conference Paper

PriFold: Biological Priors Improve RNA Secondary Structure Predictions

  • Chenchen Yang
  • Hao Wu
  • Tao Shen
  • Kai Zou
  • Siqi Sun

Predicting RNA secondary structures is crucial for understanding RNA function, designing RNA-based therapeutics, and studying molecular interactions within cells. Existing deep-learning-based methods for RNA secondary structure prediction have mainly focused on local structural properties, often overlooking the global characteristics and evolutionary features of RNA sequences. Guided by biological priors, we propose PriFold, incorporating two key innovations: 1) improving attention mechanism with pairing probabilities to utilize global pairing characteristics, and 2) implementing data augmentation based on RNA covariation to leverage evolutionary information. Our structured enhanced pretraining and finetuning strategy significantly optimizes model performance. Extensive experiments demonstrate that PriFold achieves state-of-the-art (SOTA) results in RNA secondary structure prediction on benchmark datasets such as bpRNA, RNAStrAlign and ArchiveII. These results not only validate our prediction approach but also highlight the potential of integrating biological priors, such as global characteristics and evolutionary information, into RNA structure prediction tasks, opening new avenues for research in RNA biology and bioinformatics.

AAAI Conference 2024 Conference Paper

CrossBind: Collaborative Cross-Modal Identification of Protein Nucleic-Acid-Binding Residues

  • Linglin Jing
  • Sheng Xu
  • Yifan Wang
  • Yuzhe Zhou
  • Tao Shen
  • Zhigang Ji
  • Hui Fang
  • Zhen Li

Accurate identification of protein nucleic acid binding residues poses a significant challenge with important implications for various biological processes and drug design. Many typical computational methods for protein analysis rely on a single model that could ignore either the semantic context of the protein or the global 3D geometric information. Consequently, these approaches may result in incomplete or inaccurate protein analysis. To address the above issue, in this paper, we present CrossBind, a novel collaborative cross modal approach for identifying binding residues by exploiting both protein geometric structure and its sequence prior knowledge extracted from a large scale protein language model. Specifically, our multi modal approach leverages a contrastive learning technique and atom wise attention to capture the positional relationships between atoms and residues, thereby incorporating fine grained local geometric knowledge, for better binding residue prediction. Extensive experimental results demonstrate that our approach outperforms the next best state of the art methods, GraphSite and GraphBind, on DNA and RNA datasets by 10.8/17.3% in terms of the harmonic mean of precision and recall (F1 Score) and 11.9/24.8% in Matthews correlation coefficient (MCC), respectively. We release the code at https://github.com/BEAM-Labs/CrossBind.

NeurIPS Conference 2024 Conference Paper

Dual-Personalizing Adapter for Federated Foundation Models

  • Yiyuan Yang
  • Guodong Long
  • Tao Shen
  • Jing Jiang
  • Michael Blumenstein

Recently, foundation models, particularly large language models (LLMs), have demonstrated an impressive ability to adapt to various tasks by fine-tuning diverse instruction data. Notably, federated foundation models (FedFM) emerge as a privacy preservation method to fine-tune models collaboratively under federated learning (FL) settings by leveraging many distributed datasets with non-IID data. To alleviate communication and computation overhead, parameter-efficient methods are introduced for efficiency, and some research adapted personalization methods to FedFM for better user preferences alignment. However, a critical gap in existing research is the neglect of test-time distribution shifts in real-world applications, and conventional methods for test-time distribution shifts in personalized FL are less effective for FedFM due to their failure to adapt to complex distribution shift scenarios and the requirement to train all parameters. To bridge this gap, we refine the setting in FedFM, termed test-time personalization, which aims to learn personalized federated foundation models on clients while effectively handling test-time distribution shifts simultaneously. To address challenges in this setting, we explore a simple yet effective solution, a Federated Dual-Personalizing Adapter (FedDPA) architecture. By co-working with a foundation model, a global adapter and a local adapter jointly tackle the test-time distribution shifts and client-specific personalization. Additionally, we introduce an instance-wise dynamic weighting mechanism that dynamically integrates the global and local adapters for each test instance during inference, facilitating effective test-time personalization. The effectiveness of the proposed method has been evaluated on benchmark datasets across different NLP tasks.

IJCAI Conference 2024 Conference Paper

Federated Prompt Learning for Weather Foundation Models on Devices

  • Shengchao Chen
  • Guodong Long
  • Tao Shen
  • Jing Jiang
  • Chengqi Zhang

On-device intelligence for weather forecasting uses local deep learning models to analyze weather patterns without centralized cloud computing, holds significance for supporting human activates. Federated Learning is a promising solution for such forecasting by enabling collaborative model training without sharing raw data. However, it faces three main challenges that hinder its reliability: (1) data heterogeneity among devices due to geographic differences; (2) data homogeneity within individual devices and (3) communication overload from sending large model parameters for collaboration. To address these challenges, this paper propose Federated Prompt learning for Weather Foundation Models on Devices (FedPoD), which enables devices to obtain highly customized models while maintaining communication efficiency. Concretely, our Adaptive Prompt Tuning leverages lightweight prompts guide frozen foundation model to generate more precise predictions, also conducts prompt-based multi-level communication to encourage multi-source knowledge fusion and regulate optimization. Additionally, Dynamic Graph Modeling constructs graphs from prompts, prioritizing collaborative training among devices with similar data distributions to against heterogeneity. Extensive experiments demonstrates FedPoD leads the performance among state-of-the-art baselines across various setting in real-world on-device weather forecasting datasets.

AAAI Conference 2024 Conference Paper

Fine-Grained Distillation for Long Document Retrieval

  • Yucheng Zhou
  • Tao Shen
  • Xiubo Geng
  • Chongyang Tao
  • Jianbing Shen
  • Guodong Long
  • Can Xu
  • Daxin Jiang

Long document retrieval aims to fetch query-relevant documents from a large-scale collection, where knowledge distillation has become de facto to improve a retriever by mimicking a heterogeneous yet powerful cross-encoder. However, in contrast to passages or sentences, retrieval on long documents suffers from the \textit{scope hypothesis} that a long document may cover multiple topics. This maximizes their structure heterogeneity and poses a granular-mismatch issue, leading to an inferior distillation efficacy. In this work, we propose a new learning framework, fine-grained distillation (FGD), for long-document retrievers. While preserving the conventional dense retrieval paradigm, it first produces global-consistent representations crossing different fine granularity and then applies multi-granular aligned distillation merely during training. In experiments, we evaluate our framework on two long-document retrieval benchmarks, which show state-of-the-art performance.

NeurIPS Conference 2024 Conference Paper

MSA Generation with Seqs2Seqs Pretraining: Advancing Protein Structure Predictions

  • Le Zhang
  • Jiayang Chen
  • Tao Shen
  • Yu Li
  • Siqi Sun

Deep learning models like AlphaFold2 have revolutionized protein structure prediction, achieving unprecedented accuracy. However, the dependence on robust multiple sequence alignments (MSAs) continues to pose a challenge, especially for proteins that lack a wealth of homologous sequences. To overcome this limitation, we introduce MSA-Generator, a self-supervised generative protein language model. Trained on a sequence-to-sequence task using an automatically constructed dataset, MSA-Generator employs protein-specific attention mechanisms to harness large-scale protein databases, generating virtual MSAs that enrich existing ones and boost prediction accuracy. Our experiments on CASP14 and CASP15 benchmarks reveal significant improvements in LDDT scores, particularly for complex and challenging sequences, enhancing the performance of both AlphaFold2 and RoseTTAFold. The code is released at \url{https: //github. com/lezhang7/MSAGen}.

NeurIPS Conference 2023 Conference Paper

Optimal Treatment Regimes for Proximal Causal Learning

  • Tao Shen
  • Yifan Cui

A common concern when a policymaker draws causal inferences from and makes decisions based on observational data is that the measured covariates are insufficiently rich to account for all sources of confounding, i. e. , the standard no confoundedness assumption fails to hold. The recently proposed proximal causal inference framework shows that proxy variables that abound in real-life scenarios can be leveraged to identify causal effects and therefore facilitate decision-making. Building upon this line of work, we propose a novel optimal individualized treatment regime based on so-called outcome and treatment confounding bridges. We then show that the value function of this new optimal treatment regime is superior to that of existing ones in the literature. Theoretical guarantees, including identification, superiority, excess value bound, and consistency of the estimated regime, are established. Furthermore, we demonstrate the proposed optimal regime via numerical experiments and a real data application.

IJCAI Conference 2023 Conference Paper

Prompt Federated Learning for Weather Forecasting: Toward Foundation Models on Meteorological Data

  • Shengchao Chen
  • Guodong Long
  • Tao Shen
  • Jing Jiang

To tackle the global climate challenge, it urgently needs to develop a collaborative platform for comprehensive weather forecasting on large-scale meteorological data. Despite urgency, heterogeneous meteorological sensors across countries and regions, inevitably causing multivariate heterogeneity and data exposure, become the main barrier. This paper develops a foundation model across regions capable of understanding complex meteorological data and providing weather forecasting. To relieve the data exposure concern across regions, a novel federated learning approach has been proposed to collaboratively learn a brand-new spatio-temporal Transformer-based foundation model across participants with heterogeneous meteorological data. Moreover, a novel prompt learning mechanism has been adopted to satisfy low-resourced sensors' communication and computational constraints. The effectiveness of the proposed method has been demonstrated on classical weather forecasting tasks using three meteorological datasets with multivariate time series.

YNICL Journal 2021 Journal Article

A deep learning algorithm for automatic detection and classification of acute intracranial hemorrhages in head CT scans

  • Xiyue Wang
  • Tao Shen
  • Sen Yang
  • Jun Lan
  • Yanming Xu
  • Minghui Wang
  • Jing Zhang
  • Xiao Han

Acute Intracranial hemorrhage (ICH) is a life-threatening disease that requires emergency medical attention, which is routinely diagnosed using non-contrast head CT imaging. The diagnostic accuracy of acute ICH on CT varies greatly among radiologists due to the difficulty of interpreting subtle findings and the time pressure associated with the ever-increasing workload. The use of artificial intelligence technology may help automate the process and assist radiologists for more prompt and better decision-making. In this work, we design a deep learning approach that mimics the interpretation process of radiologists, and combines a 2D CNN model and two sequence models to achieve accurate acute ICH detection and subtype classification. Being developed using the extensive 2019-RSNA Brain CT Hemorrhage Challenge dataset with over 25000 CT scans, our deep learning algorithm can accurately classify the acute ICH and its five subtypes with AUCs of 0.988 (ICH), 0.984 (EDH), 0.992 (IPH), 0.996 (IVH), 0.985 (SAH), and 0.983 (SDH), respectively, reaching the accuracy level of expert radiologists. Our method won 1st place among 1345 teams from 75 countries in the RSNA challenge. We have further evaluated our algorithm on two independent external validation datasets with 75 and 491 CT scans, respectively, and our method maintained high AUCs of 0.964 and 0.949 for acute ICH detection. These results have demonstrated the high performance and robust generalization ability of our proposed method, which makes it a useful second-read or triage tool that can facilitate routine clinical applications.

IJCAI Conference 2020 Conference Paper

Effective Search of Logical Forms for Weakly Supervised Knowledge-Based Question Answering

  • Tao Shen
  • Xiubo Geng
  • Guodong Long
  • Jing Jiang
  • Chengqi Zhang
  • Daxin Jiang

Many algorithms for Knowledge-Based Question Answering (KBQA) depend on semantic parsing, which translates a question to its logical form. When only weak supervision is provided, it is usually necessary to search valid logical forms for model training. However, a complex question typically involves a huge search space, which creates two main problems: 1) the solutions limited by computation time and memory usually reduce the success rate of the search, and 2) spurious logical forms in the search results degrade the quality of training data. These two problems lead to a poorly-trained semantic parsing model. In this work, we propose an effective search method for weakly supervised KBQA based on operator prediction for questions. With search space constrained by predicted operators, sufficient search paths can be explored, more valid logical forms can be derived, and operators possibly causing spurious logical forms can be avoided. As a result, a larger proportion of questions in a weakly supervised training set are equipped with logical forms, and fewer spurious logical forms are generated. Such high-quality training data directly contributes to a better semantic parsing model. Experimental results on one of the largest KBQA datasets (i. e. , CSQA) verify the effectiveness of our approach and deliver a new state-of-the-art performance.

AAAI Conference 2020 Conference Paper

Self-Attention Enhanced Selective Gate with Entity-Aware Embedding for Distantly Supervised Relation Extraction

  • Yang Li
  • Guodong Long
  • Tao Shen
  • Tianyi Zhou
  • Lina Yao
  • Huan Huo
  • Jing Jiang

Distantly supervised relation extraction intrinsically suffers from noisy labels due to the strong assumption of distant supervision. Most prior works adopt a selective attention mechanism over sentences in a bag to denoise from wrongly labeled data, which however could be incompetent when there is only one sentence in a bag. In this paper, we propose a brand-new light-weight neural framework to address the distantly supervised relation extraction problem and alleviate the defects in previous selective attention framework. Specifically, in the proposed framework, 1) we use an entity-aware word embedding method to integrate both relative position information and head/tail entity embeddings, aiming to highlight the essence of entities for this task; 2) we develop a self-attention mechanism to capture the rich contextual dependencies as a complement for local dependencies captured by piecewise CNN; and 3) instead of using selective attention, we design a pooling-equipped gate, which is based on rich contextual representations, as an aggregator to generate baglevel representation for final relation classification. Compared to selective attention, one major advantage of the proposed gating mechanism is that, it performs stably and promisingly even if only one sentence appears in a bag and thus keeps the consistency across all training examples. The experiments on NYT dataset demonstrate that our approach achieves a new state-of-the-art performance in terms of both AUC and top-n precision metrics.

AAAI Conference 2018 Conference Paper

DiSAN: Directional Self-Attention Network for RNN/CNN-Free Language Understanding

  • Tao Shen
  • Tianyi Zhou
  • Guodong Long
  • Jing Jiang
  • Shirui Pan
  • Chengqi Zhang

Recurrent neural nets (RNN) and convolutional neural nets (CNN) are widely used on NLP tasks to capture the long-term and local dependencies, respectively. Attention mechanisms have recently attracted enormous interest due to their highly parallelizable computation, significantly less training time, and flexibility in modeling dependencies. We propose a novel attention mechanism in which the attention between elements from input sequence(s) is directional and multi-dimensional (i. e. , feature-wise). A light-weight neural net, “Directional Self-Attention Network (DiSAN)”, is then proposed to learn sentence embedding, based solely on the proposed attention without any RNN/CNN structure. DiSAN is only composed of a directional self-attention with temporal order encoded, followed by a multi-dimensional attention that compresses the sequence into a vector representation. Despite its simple form, DiSAN outperforms complicated RNN models on both prediction quality and time efficiency. It achieves the best test accuracy among all sentence encoding methods and improves the most recent best result by 1. 02% on the Stanford Natural Language Inference (SNLI) dataset, and shows stateof-the-art test accuracy on the Stanford Sentiment Treebank (SST), Multi-Genre natural language inference (MultiNLI), Sentences Involving Compositional Knowledge (SICK), Customer Review, MPQA, TREC question-type classification and Subjectivity (SUBJ) datasets.

IJCAI Conference 2018 Conference Paper

Reinforced Self-Attention Network: a Hybrid of Hard and Soft Attention for Sequence Modeling

  • Tao Shen
  • Tianyi Zhou
  • Guodong Long
  • Jing Jiang
  • Sen Wang
  • Chengqi Zhang

Many natural language processing tasks solely rely on sparse dependencies between a few tokens in a sentence. Soft attention mechanisms show promising performance in modeling local/global dependencies by soft probabilities between every two tokens, but they are not effective and efficient when applied to long sentences. By contrast, hard attention mechanisms directly select a subset of tokens but are difficult and inefficient to train due to their combinatorial nature. In this paper, we integrate both soft and hard attention into one context fusion model, "reinforced self-attention (ReSA)", for the mutual benefit of each other. In ReSA, a hard attention trims a sequence for a soft self-attention to process, while the soft attention feeds reward signals back to facilitate the training of the hard one. For this purpose, we develop a novel hard attention called "reinforced sequence sampling (RSS)", selecting tokens in parallel and trained via policy gradient. Using two RSS modules, ReSA efficiently extracts the sparse dependencies between each pair of selected tokens. We finally propose an RNN/CNN-free sentence-encoding model, "reinforced self-attention network (ReSAN)", solely based on ReSA. It achieves state-of-the-art performance on both the Stanford Natural Language Inference (SNLI) and the Sentences Involving Compositional Knowledge (SICK) datasets.

v2026.09.13