Arrow Research search

Author name cluster

Qiong Wu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

27 papers
1 author row

Possible papers

27

EAAI Journal 2026 Journal Article

Automatic stem phenotyping in soybean using keypoint detection

  • Fei Liu
  • Qiong Wu
  • Zhongzhi Han
  • Longgang Zhao
  • Shanchen Pang
  • Shudong Wang

Soybean breeding critically relies on stem phenotypes, as they directly impact yield and lodging resistance. Traditional measurement methods are labor-intensive, prone to human error, and often require destructive sampling. Although artificial intelligence (AI) has emerged as a transformative alternative, existing studies on soybean stem phenotyping remain limited and imprecise. An AI implementation integrating keypoint detection and localization provides a promising solution. This study proposes Soybean-pose, a novel approach that models soybean plants as structural “bodily forms” via keypoint detection. By integrating a semi-supervised iterative self-training paradigm with a hybrid Convolutional Neural Network-Swin Vision Transformer (CNN-SViT) architecture, Soybean-pose achieves high-precision detection of soybean stem nodes via limited labeled data and pseudo-label iterative optimization strategies, and integrates phenotypic quantification algorithms to accomplish automated parsing of stem-related phenotypes. To support this research, the Soybean Stem Keypoint (SSK) dataset is constructed and publicly released. Soybean-pose achieves an average precision at 50 % intersection-over-union (AP50) of 91. 8 % on the validation set and 93. 2 % on the test-dev set. The Pearson correlation coefficients (R) for pitch number, internode length, and main stem length are 0. 986, 0. 989, and 0. 978, respectively. This AI application enables accurate measurement of soybean stem phenotypes, reduces labor costs, and minimizes measurement errors, demonstrating its potential to accelerate breeding processes.

AAAI Conference 2026 Conference Paper

MARS: Multimodal Adaptive Reasoning Model for Avoiding Overthinking

  • Tan Yue
  • Qiong Wu
  • Dongyan Zhao

Multimodal Large Language Models (MLLMs) have shown advanced performance in vision-language tasks. However, existing multimodal reasoning models often suffer from excessive reasoning steps, leading to high computational costs and inefficiency. In this paper, we propose the Multimodal Adaptive Reasoning Model (MARS), which enables adaptive adjustment of the reasoning strategy based on question difficulty. Specifically, MARS adopts a three-stage training framework based on our constructed training dataset (MART): 1) CoT Masking Learning to enhance reasoning logicality by predicting masked reasoning steps. 2) Adaptive Reasoning Instruction Learning to train the model to skip or keep reasoning steps according to difficulty levels. 3) CoT Lightweight Reinforcement Learning with the Information Bottleneck Principle based GRPO algorithm to reduce CoT length while maintaining performance and generalizability. Results on both in-domain and out-of-domain datasets show that MARS significantly reduces the CoT length (90.2% decrease) while improving accuracy (0.54%), outperforming existing SOTA open-source and proprietary MLLMs.

EAAI Journal 2025 Journal Article

A novel local enhanced channel self-attention based on Transformer for industrial remaining useful life prediction

  • Zhizheng Zhang
  • Wen Song
  • Qiong Wu
  • Wenxu Sun
  • Qiqiang Li
  • Lei Jia

Remaining useful life (RUL) prediction is a foundational technique for predictive maintenance (PdM) and is critical to ensuring the reliability and safety of complex industrial machines. Recently, while advanced deep learning architectures like recurrent neural network (RNN), convolutional neural network (CNN) and self-attention (SA) have been widely used for RUL prediction, existing methods still face difficulties in simultaneously processing global long-term dependencies and local contextual information of sequence units as well as the spatial correlations of industrial multi-sensors. In this article, we propose local enhanced channel self-attention based on Transformer (LECformer), a novel deep RUL prediction method to overcome these issues. LECformer can more effectively capture the long-term dependencies by Transformer architecture compared with RNN/CNN-based methods. Moreover, LECformer proposes a novel local enhanced channel self-attention (LECSA) mechanism to replace the traditional SA of vanilla Transformer, which can adaptively extract both long-term dependencies and local contextual information, while dynamically weighting the importance of different channels to improve predictive performance. Two widely used turbofan engine datasets and a bearing dataset are applied to validate the effectiveness of the proposed method. Experimental results show that the LECformer significantly outperforms the state-of-the-art RUL prediction methods.

NeurIPS Conference 2025 Conference Paper

Accelerating Multimodal Large Language Models via Dynamic Visual-Token Exit and the Empirical Findings

  • Qiong Wu
  • Wenhao Lin
  • Yiyi Zhou
  • Weihao Ye
  • Zhanpeng Zeng
  • Xiaoshuai Sun
  • Rongrong Ji

In this paper, we study the visual redundancy problem of multimodal large language models (MLLMs) from the perspective of attention behaviors. Via extensive empirical experiments, we observe and conclude three main inference stages of MLLMs: (i) Early fusion between tokens is first accomplished quickly. (ii) Intra-modality modeling then comes to play. (iii) Multimodal reasoning resumes and lasts until the end of inference. In particular, we reveal that visual tokens will stop contributing to reasoning when the text tokens receive enough image information. Based on this observation, we propose an effective method to improve the efficiency of MLLMs, termed dynamic visual-token exit (DyVTE), which is orthogonal but collaborative to previous token-wise visual compression methods. To validate the efficacy of DyVTE, we apply it to a set of MLLMs, including LLaVA, VILA, EAGLE and InternVL. The experimental results not only show the effectiveness of our DyVTE in improving MLLMs' efficiency, e. g. , DyVTE reduces the computation overhead of LLaVA-1. 5 by up to 45. 7% without performance drop, but also reveal a general pattern across multiple MLLMs, well facilitating the in-depth analysis of MLLMs. Our code is anonymously released at https: //anonymous. 4open. science/r/AnonymousDyVTE-26AB/.

YNICL Journal 2025 Journal Article

Boostering diagnosis of frontotemporal lobar degeneration with AI-driven neuroimaging – A systematic review and meta-analysis

  • Qiong Wu
  • Dimitra Kiakou
  • Karsten Mueller
  • Wolfgang Köhler
  • Matthias L. Schroeter

BACKGROUND AND OBJECTIVES: Frontotemporal lobar degeneration (FTLD) as the second most common dementia encompasses a range of syndromes and often shows overlapping symptoms with other subtypes or neurodegenerative diseases, which poses a significant clinical diagnostic challenge. Recent advancements in artificial intelligence (AI), specifically the application of machine learning (ML) algorithms to neuroimaging, have significantly progressed in addressing this challenge. This study aims to assess the diagnostic and predictive efficacy of neuroimaging feature-based AI algorithms for FTLD. METHODS: We conducted a systematic review and meta-analysis following PRISMA guidelines. We searched Pubmed, Scopus, and Web of Science for English-language, peer-reviewed studies using the following three umbrella terms: artificial intelligence, frontotemporal lobar degeneration, and neuroimaging modality. Our survey focused on computer-aided diagnosis for FTLD, employing machine/deep learning with neuroimaging radiomic features. RESULTS: The meta-analysis includes 75 articles with 20,601 subjects, including 8,051 FTLD patients. The results reveal that FTLD can be automatically classified against healthy controls (HC) with pooled sensitivity and specificity of 86% and 89%, respectively. Likewise, FTLD versus Alzheimer's disease (AD) classification exhibits pooled sensitivity and specificity of 84% and 81%, while FTLD versus Parkinson's disease (PD) demonstrates pooled sensitivity and specificity of 84% and 75%, respectively. Classification performance distinguishing FTLD from atypical Parkinsonian syndromes (APS) showed pooled sensitivity and specificity of 84% and 79%, respectively. Multiclass classification sensitivity ranges from 42% to 100%, with lower sensitivity occurring in higher class distinctions (e.g., 5-class and 11-class). DISCUSSION: Our study demonstrates the effectiveness of utilizing neuroimaging features to distinguish FTLD from HC, AD, APS, and PD in binary classification. Utilizing deep learning with multimodal neuroimaging data to differentiate FTLD subtypes and perform multiclassification among FTLD and other neurodegenerative disease holds promise for expediting diagnosis. In sum, the meta-analysis supports translation of machine learning tools in combination with imaging to clinical routine paving the way to precision medicine.

NeurIPS Conference 2025 Conference Paper

Diagnosing and Addressing Pitfalls in KG-RAG Datasets: Toward More Reliable Benchmarking

  • Liangliang Zhang
  • Zhuorui Jiang
  • Hongliang Chi
  • Haoyang Chen
  • Mohammed ElKoumy
  • Fali Wang
  • Qiong Wu
  • Zhengyi Zhou

Knowledge Graph Question Answering (KGQA) systems rely on high-quality benchmarks to evaluate complex multi-hop reasoning. However, despite their widespread use, popular datasets such as WebQSP and CWQ suffer from critical quality issues, including inaccurate or incomplete ground-truth annotations, poorly constructed questions that are ambiguous, trivial, or unanswerable, and outdated or inconsistent knowledge. Through a manual audit of 16 popular KGQA datasets—including WebQSP and CWQ—we find that the average factual correctness rate is only 57%. To address these issues, we introduce KGQAGen, an LLM-in-the-loop framework that systematically resolves these pitfalls. KGQAGen combines structured knowledge grounding, LLM-guided generation, and symbolic verification to produce challenging and verifiable QA instances. Using KGQAGen, we construct KGQAGen-10k, a 10K-scale benchmark grounded in Wikidata, and evaluate a diverse set of KG-RAG models. Experimental results demonstrate that even state-of-the-art systems struggle on this benchmark, highlighting its ability to expose limitations of existing models. Our findings advocate for more rigorous benchmark construction and position KGQAGen as a scalable framework for advancing KGQA evaluation.

AAAI Conference 2025 Conference Paper

Fit and Prune: Fast and Training-free Visual Token Pruning for Multi-modal Large Language Models

  • Weihao Ye
  • Qiong Wu
  • Wenhao Lin
  • Yiyi Zhou

Recent progress in Multimodal Large Language Models (MLLMs) often use large image tokens to compensate the visual shortcoming of MLLMs, which not only exhibits obvious redundancy but also greatly exacerbates the already high computation. Token pruning is an effective solution for speeding up MLLMs, but when and how to drop tokens still remains a challenge. In this paper, we propose a novel and training-free approach for the effective visual token pruning of MLLMs, termed FitPrune, which can quickly produce a complete pruning recipe for MLLMs according to a pre-defined budget. Specifically, FitPrune considers token pruning as a statistical problem of MLLM and its objective is to find out an optimal pruning scheme that can minimize the divergence of the attention distributions before and after pruning. In practice, FitPrune can be quickly accomplished based on the attention statistics from a small batch of inference data, avoiding the expensive trials of MLLMs. According to the pruning recipe, an MLLM can directly remove the redundant visual tokens of different examples during inference. To validate FitPrune, we apply it to a set of recent MLLMs, including LLaVA-1.5, LLaVA-HR and LLaVA-NEXT, and conduct extensive experiments on a set of benchmarks. The experimental results show that our FitPrune can not only reduce the computational complexity to a large extent, while retaining high performance, e.g., -54.9% FLOPs for LLaVA-NEXT with only 0.5% accuracy drop. Notably, the pruning recipe can be obtained in about 5 minutes.

AAAI Conference 2025 Conference Paper

What Kind of Visual Tokens Do We Need? Training-Free Visual Token Pruning for Multi-Modal Large Language Models from the Perspective of Graph

  • Yutao Jiang
  • Qiong Wu
  • Wenhao Lin
  • Wei Yu
  • Yiyi Zhou

Recent Multimodal Large Language Models(MLLMs) often use a large number of visual tokens to compensate their visual shortcoming, leading to excessive computation and obvious visual redundancy. In this paper, we investigate what kind of visual tokens are needed for MLLMs, and reveal that both foreground and background tokens are critical for MLLMs given the varying difficulties of examples. Based on this observation, we propose a graph-based method towards training-free visual token pruning, termed G-Prune. In particular, G-Prune regards visual tokens as nodes, and construct their connections based on their semantic similarities. Afterwards, the information flow is propagated via weighted links, and the most important tokens after iterations are kept for MLLMs, which can be front or background. To validate G-Prune, we apply it to a recent MLLM called LLaVA-NeXT, and conduct extensive experiments on a set of benchmarks. The experiment results show that G-Prune can greatly reduce computation overhead while retaining high performance on both coarse- and fine-grained tasks. For instance, G-Prune can reduce 63.57% FLOPs of LLaVA-NeXT on VQA2.0 and TextVQA with only 0.95% and 2.34% accuracy drops, respectively.

TMLR Journal 2024 Journal Article

A General-Purpose Multi-Modal OOD Detection Framework

  • Viet Quoc Duong
  • Qiong Wu
  • Zhengyi Zhou
  • Eric Zavesky
  • WenLing Hsu
  • Han Zhao
  • Huajie Shao

Out-of-distribution (OOD) detection seeks to identify test samples that deviate from the training data, which is critical to ensuring the safety and reliability of machine learning (ML) systems. While a plethora of methods have been developed to detect uni-modal OOD samples, only a few have focused on multi-modal OOD detection. Current contrastive learning-based methods primarily address multi-modal OOD detection in a scenario where an image is not related to the class labels in training data. However, ML systems in the real-world applications may encounter a broader spectrum of anomalies caused by different factors like systematic errors in labeling, environmental changes, and sensor malfunctions. Hence, we propose a new method to be able to simultaneously detect anomalies from multiple different OOD scenarios, arising from fine-grained image features and textual descriptions, instead of large categorical information. To achieve this goal, we propose a general-purpose weakly-supervised OOD detection framework, called WOOD, that combines a binary classifier and a contrastive learning module to reap the benefits of both. In order to better distinguish in-distribution (ID) samples from OOD ones, we employ the Hinge loss to constrain the similarity of their latent representations. Moreover, we devise a new scoring metric that fuses predictions from both the binary classifier and contrastive learning to enhance OOD detection. Extensive experimental results on multiple benchmarks demonstrate that the proposed WOOD significantly outperforms the state-of-the-art methods for multi-modal OOD detection. Importantly, our approach can achieve superior detection performance in a variety of OOD scenarios.

AAAI Conference 2024 Conference Paper

Efficient Toxic Content Detection by Bootstrapping and Distilling Large Language Models

  • Jiang Zhang
  • Qiong Wu
  • Yiming Xu
  • Cheng Cao
  • Zheng Du
  • Konstantinos Psounis

Toxic content detection is crucial for online services to remove inappropriate content that violates community standards. To automate the detection process, prior works have proposed varieties of machine learning (ML) approaches to train Language Models (LMs) for toxic content detection. However, both their accuracy and transferability across datasets are limited. Recently, Large Language Models (LLMs) have shown promise in toxic content detection due to their superior zero-shot and few-shot in-context learning ability as well as broad transferability on ML tasks. However, efficiently designing prompts for LLMs remains challenging. Moreover, the high run-time cost of LLMs may hinder their deployments in production. To address these challenges, in this work, we propose BD-LLM, a novel and efficient approach to bootstrapping and distilling LLMs for toxic content detection. Specifically, we design a novel prompting method named Decision-Tree-of-Thought (DToT) to bootstrap LLMs' detection performance and extract high-quality rationales. DToT can automatically select more fine-grained context to re-prompt LLMs when their responses lack confidence. Additionally, we use the rationales extracted via DToT to fine-tune student LMs. Our experimental results on various datasets demonstrate that DToT can improve the accuracy of LLMs by up to 4.6%. Furthermore, student LMs fine-tuned with rationales extracted via DToT outperform baselines on all datasets with up to 16.9% accuracy improvement, while being more than 60x smaller than conventional LLMs. Finally, we observe that student LMs fine-tuned with rationales exhibit better cross-dataset transferability.

EAAI Journal 2024 Journal Article

Expanding the defect image dataset of composite material coating with enhanced image-to-image translation

  • Xinrui Tao
  • Hanjun Gao
  • Kai Yang
  • Qiong Wu

In the defect detection based on machine vision, the limited datasets have consistently impeded the efficacy of deep learning in defect identification. Restricted datasets lead to inadequate model generalization and reduced detection accuracy. Creating simulated images is one proposed solution to address this challenge. However, existing defect image generation methods have limitations in diversity and realism. An enhanced Image-to-image translation method, named E-Pix2Pix, is proposed for enhancing defect image datasets of composite material coating on solid motors to expand the current dataset and enhance the precision of classification models. This method not only improves the quality of generated images, but also has strong generalization ability. This study examines how different convolution kernel sizes and the coefficient of the L1 loss function affect the generation performance. The results of the experiments show that the defect images generated by E-Pix2Pix have realistic visual effects and better capture the essential features of defects. The structure similarity index measurement (SSIM) value reaches approximately 0. 9 under appropriate parameters. Moreover, this study intends to investigate how the content of generated images impacts the classification model. By doubling the dataset size using the generated images, the accuracy of ResNet50 model recognition can be improved by 7%. The accuracy of ResNet101 model recognition can be improved by 5% when the dataset is expanded to 1. 5 times. The generated images have harmful effects on the weaker models. Our method has also been validated on two public datasets, and satisfactory results have been obtained.

AAMAS Conference 2023 Conference Paper

A Deep Reinforcement Learning Approach for Online Parcel Assignment

  • Hao Zeng
  • Qiong Wu
  • Kunpeng Han
  • Junying He
  • Haoyuan Hu

In this paper, we investigate the online parcel assignment (OPA) problem, in which each stochastically generated parcel order needs to be assigned to a candidate route for delivery with the objective to minimize the total delivery cost under certain business constraints. The OPA problem is challenging due to its stochastic nature: each parcel’s candidate routes, which depend on the parcel’s attributes, are unknown until its order is placed, and the total parcel volume to be assigned is uncertain in advance. To tackle this problem, we propose an algorithm based on deep reinforcement learning, namely PPO-OPA, that shows competitive performance. More specifically, we introduce a novel Markov Decision Process (MDP) to model the decision-making process in the OPA problem, and develop a policy gradient algorithm that adopts attention networks for policy evaluation. By designing a dedicated reward function, our proposed algorithm can achieve a lower total cost with a smaller violation of constraints, compared to the traditional method used in the industry that assigns parcels to candidate routes proportionally. In addition, the performances of our proposed algorithm and the Primal-Dual algorithm are comparable, while the later assumes a known total parcel volume in advance, which is unrealistic in practice.

JBHI Journal 2023 Journal Article

Immunotherapy Efficacy Prediction for Non-Small Cell Lung Cancer Using Multi-View Adaptive Weighted Graph Convolutional Networks

  • Qiong Wu
  • Jun Wang
  • Zongqiong Sun
  • Lei Xiao
  • Wenhao Ying
  • Jun Shi

Immunotherapy is an effective way to treat non-small cell lung cancer (NSCLC). The efficacy of immunotherapy differs from person to person and may cause side effects, making it important to predict the efficacy of immunotherapy before surgery. Radiomics based on machine learning has been successfully used to predict the efficacy of NSCLC immunotherapy. However, most studies only considered the radiomic features of the individual patient, ignoring the inter-patient correlations. Besides, they usually concatenated different features as the input of a single-view model, failing to consider the complex correlation among features of multiple types. To this end, we propose a multi-view adaptive weighted graph convolutional network (MVAW-GCN) for the prediction of NSCLC immunotherapy efficacy. Specifically, we group the radiomic features into several views according to the type of the fitered images they extracted from. We construct a graph in each view based on the radiomic features and phenotypic information. An attention mechanism is introduced to automatically assign weights to each view. Considering the view-shared and view-specific knowledge of radiomic features, we propose separable graph convolution that decomposes the output of the last convolution layer into two components, i. e. , the view-shared and view-specific outputs. We maximize the consistency and enhance the diversity among different views in the learning procedure. The proposed MVAW-GCN is evaluated on 107 NSCLC patients, including 52 patients with valid efficacy and 55 patients with invalid efficacy. Our method achieved an accuracy of 77. 27% and an area under the curve (AUC) of 0. 7780, indicating its effectiveness in NSCLC immunotherapy efficacy prediction.

NeurIPS Conference 2023 Conference Paper

Parameter and Computation Efficient Transfer Learning for Vision-Language Pre-trained Models

  • Qiong Wu
  • Wei Yu
  • Yiyi Zhou
  • Shubin Huang
  • Xiaoshuai Sun
  • Rongrong Ji

With ever increasing parameters and computation, vision-language pre-trained (VLP) models exhibit prohibitive expenditure in downstream task adaption. Recent endeavors mainly focus on parameter efficient transfer learning (PETL) for VLP models by only updating a small number of parameters. However, excessive computational overhead still plagues the application of VLPs. In this paper, we aim at parameter and computation efficient transfer learning (PCETL) for VLP models. In particular, PCETL not only needs to limit the number of trainable parameters in VLP models, but also to reduce the computational redundancy during inference, thus enabling a more efficient transfer. To approach this target, we propose a novel dynamic architecture skipping (DAS) approach towards effective PCETL. Instead of directly optimizing the intrinsic architectures of VLP models, DAS first observes the significances of their modules to downstream tasks via a reinforcement learning (RL) based process, and then skips the redundant ones with lightweight networks, i. e. adapters, according to the obtained rewards. In this case, the VLP model can well maintain the scale of trainable parameters while speeding up its inference on downstream tasks. To validate DAS, we apply it to two representative VLP models, namely ViLT and METER, and conduct extensive experiments on a bunch of VL tasks. The experimental results not only show the great advantages of DAS in reducing computational complexity, e. g. -11. 97% FLOPs of METER on VQA2. 0, but also confirm its competitiveness against existing PETL methods in terms of parameter scale and performance. Our source code is given in our appendix.

AAAI Conference 2023 Conference Paper

Symphony in the Latent Space: Provably Integrating High-Dimensional Techniques with Non-linear Machine Learning Models

  • Qiong Wu
  • Jian Li
  • Zhenming Liu
  • Yanhua Li
  • Mihai Cucuringu

This paper revisits building machine learning algorithms that involve interactions between entities, such as those between financial assets in an actively managed portfolio, or interactions between users in a social network. Our goal is to forecast the future evolution of ensembles of multivariate time series in such applications (e.g., the future return of a financial asset or the future popularity of a Twitter account). Designing ML algorithms for such systems requires addressing the challenges of high-dimensional interactions and non-linearity. Existing approaches usually adopt an ad-hoc approach to integrating high-dimensional techniques into non-linear models and recent studies have shown these approaches have questionable efficacy in time-evolving interacting systems. To this end, we propose a novel framework, which we dub as the additive influence model. Under our modeling assumption, we show that it is possible to decouple the learning of high-dimensional interactions from the learning of non-linear feature interactions. To learn the high-dimensional interactions, we leverage kernel-based techniques, with provable guarantees, to embed the entities in a low-dimensional latent space. To learn the non-linear feature-response interactions, we generalize prominent machine learning techniques, including designing a new statistically sound non-parametric method and an ensemble learning algorithm optimized for vector regressions. Extensive experiments on two common applications demonstrate that our new algorithms deliver significantly stronger forecasting power compared to standard and recently proposed methods.

YNIMG Journal 2022 Journal Article

Neuronal efficiency following n-back training task is accompanied by a higher cerebral glucose metabolism

  • Isabelle Ripp
  • Qiong Wu
  • Lara Wallenwein
  • Mónica Emch
  • Igor Yakushev
  • Kathrin Koch

Recent functional magnetic resonance imaging (fMRI) studies revealed lower neural activation during processing of an n-back task following working memory training, indicating a training-related increase in neural efficiency. In the present study, we asked if the training induced regional neural activation is accompanied by changes in glucose consumption. An active control and an experimental group of healthy middle-aged volunteers conducted 32 sessions of visual and verbal n-back trainings over 8 weeks. We analyzed data of 52 subjects (25 experimental and 27 control group) for practice effects underlying verbal working memory task and 50 subjects (24 experimental and 26 control group) for practice effects underlying visual WM task. The samples of these two tasks were nearly identical (data of 47 subjects were available for both verbal and visual tasks). Both groups completed neuroimaging sessions at a hybrid PET/MR system before and after training. Each session included criterion task fMRI and resting state positron emission tomography with FDG (FDG-PET). As reported previously, lower neural activation following n-back training was found in regions of the fronto-parieto-cerebellar circuitry during a verbal n-back task. Notably, these changes co-occurred spatially with a higher relative FDG-uptake. Decreased neural activation within regions of the fronto-parietal network during visual n-back task did not show co-occurring changes in relative FDG-uptake. There was no direct association between neuroimaging and behavioral measures, which could be due to the inter-subjects' variability in reaching capacity limits. Our findings provide new details for working memory training induced neural efficiency on a molecular level by integrating FDG-PET and fMRI measures.

TIST Journal 2021 Journal Article

BATS: A Spectral Biclustering Approach to Single Document Topic Modeling and Segmentation

  • Qiong Wu
  • Adam Hare
  • Sirui Wang
  • Yuwei Tu
  • Zhenming Liu
  • Christopher G. Brinton
  • Yanhua Li

Existing topic modeling and text segmentation methodologies generally require large datasets for training, limiting their capabilities when only small collections of text are available. In this work, we reexamine the inter-related problems of “topic identification” and “text segmentation” for sparse document learning, when there is a single new text of interest. In developing a methodology to handle single documents, we face two major challenges. First is sparse information: with access to only one document, we cannot train traditional topic models or deep learning algorithms. Second is significant noise: a considerable portion of words in any single document will produce only noise and not help discern topics or segments. To tackle these issues, we design an unsupervised, computationally efficient methodology called Biclustering Approach to Topic modeling and Segmentation (BATS). BATS leverages three key ideas to simultaneously identify topics and segment text: (i) a new mechanism that uses word order information to reduce sample complexity, (ii) a statistically sound graph-based biclustering technique that identifies latent structures of words and sentences, and (iii) a collection of effective heuristics that remove noise words and award important words to further improve performance. Experiments on six datasets show that our approach outperforms several state-of-the-art baselines when considering topic coherence, topic diversity, segmentation, and runtime comparison metrics.

YNICL Journal 2021 Journal Article

Temporal-thalamic and cingulo-opercular connectivity in people with schizophrenia

  • Adam J. Culbreth
  • Qiong Wu
  • Shuo Chen
  • Bhim M. Adhikari
  • L. Elliot Hong
  • James M. Gold
  • James A. Waltz

A growing body of research has suggested that people with schizophrenia (SZ) exhibit altered patterns of functional and anatomical brain connectivity. For example, many previous resting state functional connectivity (rsFC) studies have shown that, compared to healthy controls (HC), people with SZ demonstrate hyperconnectivity between subregions of the thalamus and sensory cortices, as well as hypoconnectivity between subregions of the thalamus and prefrontal cortex. In addition to thalamic findings, hypoconnectivity between cingulo-opercular brain regions thought to be involved in salience detection has also been commonly reported in people with SZ. However, previous studies have largely relied on seed-based analyses. Seed-based approaches require researchers to define a single a priori brain region, which is then used to create a rsFC map across the entire brain. While useful for testing specific hypotheses, these analyses are limited in that only a subset of connections across the brain are explored. In the current manuscript, we leverage novel network statistical techniques in order to detect latent functional connectivity networks with organized topology that successfully differentiate people with SZ from HCs. Importantly, these techniques do not require a priori seed selection and allow for whole brain investigation, representing a comprehensive, data-driven approach to determining differential connectivity between diagnostic groups. Across two samples, (Sample 1: 35 SZ, 44 HC; Sample 2: 65 SZ, 79 HC), we found evidence for differential rsFC within a network including temporal and thalamic regions. Connectivity in this network was greater for people with SZ compared to HCs. In the second sample, we also found evidence for hypoconnectivity within a cingulo-opercular network of brain regions in people with SZ compared to HCs. In summary, our results replicate and extend previous studies suggesting hyperconnectivity between the thalamus and sensory cortices and hypoconnectivity between cingulo-opercular regions in people with SZ using data-driven statistical and graph theoretical techniques.

YNIMG Journal 2021 Journal Article

Working memory task induced neural activation: A simultaneous PET/fMRI study

  • Isabelle Ripp
  • Lara A Wallenwein
  • Qiong Wu
  • Monica Emch
  • Kathrin Koch
  • Paul Cumming
  • Igor Yakushev

PURPOSE: F]fluorodeoxyglucose (FDG) is a powerful method for mapping cerebral glucose metabolism as a proxy of neural activity, assuming a steady-state during the recording interval. We asked if a clinical FDG-PET imaging protocol might also capture changes in neural activity associated with performance of a working memory (WM) task. METHODS: To test this concept, we examined hybrid PET/MR data for FDG-PET and simultaneous functional magnetic resonance imaging (fMRI) in a sample of healthy volunteers. The PET image acquisition started 30 min after a bolus injection of approximately 100 MBq FDG, and the WM task was undertaken starting at approximately 60 min post-injection. We reconstructed FDG-PET sum images corresponding to baseline (44-60 min p.i.) and WM tasks (63- 71 min p.i.), each with intensity scaling to the corresponding global mean. RESULTS: Compared to the baseline resting condition, relative FDG uptake increased during WM task performance in brain regions previously associated with WM. Furthermore, these metabolically active regions partly overlapped with the regions showing task-dependent increases in BOLD signal in simultaneous fMRI. CONCLUSION: We find evidence for WM task-induced neural activation using a clinical FDG-PET imaging protocol. These findings encourage the development of dedicated protocols for tracking neural correlates of cognitive function.

NeurIPS Conference 2020 Conference Paper

Adaptive Reduced Rank Regression

  • Qiong Wu
  • Felix M. Wong
  • Yanhua Li
  • Zhenming Liu
  • Varun Kanade

We study the low rank regression problem y = Mx + ε, where x and y are d1 and d2 dimensional vectors respectively. We consider the extreme high-dimensional setting where the number of observations n is less than d1 + d2. Existing algorithms are designed for settings where n is typically as large as rank(M)(d1+d2). This work provides an efficient algorithm which only involves two SVD, and establishes statistical guarantees on its performance. The algorithm decouples the problem by first estimating the precision matrix of the features, and then solving the matrix denoising problem. To complement the upper bound, we introduce new techniques for establishing lower bounds on the performance of any algorithm for this problem. Our preliminary experiments confirm that our algorithm often out-performs existing baseline, and is always at least competitive.

AAAI Conference 2020 Conference Paper

Diversified Interactive Recommendation with Implicit Feedback

  • Yong Liu
  • Yingtai Xiao
  • Qiong Wu
  • Chunyan Miao
  • Juyong Zhang
  • Binqiang Zhao
  • Haihong Tang

Interactive recommender systems that enable the interactions between users and the recommender system have attracted increasing research attention. Previous methods mainly focus on optimizing recommendation accuracy. However, they usually ignore the diversity of the recommendation results, thus usually results in unsatisfying user experiences. In this paper, we propose a novel diversified recommendation model, named Diversified Contextual Combinatorial Bandit (DC2 B), for interactive recommendation with users’ implicit feedback. Specifically, DC2 B employs determinantal point process in the recommendation procedure to promote diversity of the recommendation results. To learn the model parameters, a Thompson sampling-type algorithm based on variational Bayesian inference is proposed. In addition, theoretical regret analysis is also provided to guarantee the performance of DC2 B. Extensive experiments on real datasets are performed to demonstrate the effectiveness of the proposed method in balancing the recommendation accuracy and diversity.

YNIMG Journal 2019 Journal Article

Anterior insular cortex is a bottleneck of cognitive control

  • Tingting Wu
  • Xingchao Wang
  • Qiong Wu
  • Alfredo Spagna
  • Jiaqi Yang
  • Changhe Yuan
  • Yanhong Wu
  • Zhixian Gao

Cognitive control, with a limited capacity, is a core process in human cognition for the coordination of thoughts and actions. Although the regions involved in cognitive control have been identified as the cognitive control network (CCN), it is still unclear whether a specific region of the CCN serves as a bottleneck limiting the capacity of cognitive control (CCC). Here, we used a perceptual decision-making task with conditions of high cognitive load to challenge the CCN and to assess the CCC in a functional magnetic resonance imaging study. We found that the activation of the right anterior insular cortex (AIC) of the CCN increased monotonically as a function of cognitive load, reached its plateau early, and showed a significant correlation to the CCC. In a subsequent study of patients with unilateral lesions of the AIC, we found that lesions of the AIC were associated with a significant impairment of the CCC. Simulated lesions of the AIC resulted in a reduction of the global efficiency of the CCN in a network analysis. These findings suggest that the AIC, as a critical hub in the CCN, is a bottleneck of cognitive control.

AAAI Conference 2019 Conference Paper

Near-Neighbor Methods in Random Preference Completion

  • Ao Liu
  • Qiong Wu
  • Zhenming Liu
  • Lirong Xia

This paper studies a stylized, yet natural, learning-to-rank problem and points out the critical incorrectness of a widely used nearest neighbor algorithm. We consider a model with n agents (users) {xi}i∈[n] and m alternatives (items) {yl}l∈[m], each of which is associated with a latent feature vector. Agents rank items nondeterministically according to the Plackett-Luce model, where the higher the utility of an item to the agent, the more likely this item will be ranked high by the agent. Our goal is to identify near neighbors of an arbitrary agent in the latent space for prediction. We first show that the Kendall-tau distance based kNN produces incorrect results in our model. Next, we propose a new anchor-based algorithm to find neighbors of an agent. A salient feature of our algorithm is that it leverages the rankings of many other agents (the so-called “anchors”) to determine the closeness/similarities of two agents. We provide a rigorous analysis for one-dimensional latent space, and complement the theoretical results with experiments on synthetic and real datasets. The experiments confirm that the new algorithm is robust and practical.

IJCAI Conference 2019 Conference Paper

PD-GAN: Adversarial Learning for Personalized Diversity-Promoting Recommendation

  • Qiong Wu
  • Yong Liu
  • Chunyan Miao
  • Binqiang Zhao
  • Yin Zhao
  • Lu Guan

This paper proposes Personalized Diversity-promoting GAN (PD-GAN), a novel recommendation model to generate diverse, yet relevant recommendations. Specifically, for each user, a generator recommends a set of diverse and relevant items by sequentially sampling from a personalized Determinantal Point Process (DPP) kernel matrix. This kernel matrix is constructed by two learnable components: the general co-occurrence of diverse items and the user's personal preference to items. To learn the first component, we propose a novel pairwise learning paradigm using training pairs, and each training pair consists of a set of diverse items and a set of similar items randomly sampled from the observed data of all users. The second component is learnt through adversarial training against a discriminator which strives to distinguish between recommended items and the ground-truth sets randomly sampled from the observed data of the target user. Experimental results show that PD-GAN is superior to generate recommendations that are both diverse and relevant.

IJCAI Conference 2018 Conference Paper

Cross-Modality Person Re-Identification with Generative Adversarial Training

  • Pingyang Dai
  • Rongrong Ji
  • Haibin Wang
  • Qiong Wu
  • Yuyu Huang

Person re-identification (Re-ID) is an important task in video surveillance which automatically searches and identifies people across different cameras. Despite the extensive Re-ID progress in RGB cameras, few works have studied the Re-ID between infrared and RGB images, which is essentially a cross-modality problem and widely encountered in real-world scenarios. The key challenge lies in two folds, i. e. , the lack of discriminative information to re-identify the same person between RGB and infrared modalities, and the difficulty to learn a robust metric towards such a large-scale cross-modality retrieval. In this paper, we tackle the above two challenges by proposing a novel cross-modality generative adversarial network (termed cmGAN). To handle the issue of insufficient discriminative information, we leverage the cutting-edge generative adversarial training to design our own discriminator to learn discriminative feature representation from different modalities. To handle the issue of large-scale cross-modality metric learning, we integrates both identification loss and cross-modality triplet loss, which minimize inter-class ambiguity while maximizing cross-modality similarity among instances. The entire cmGAN can be trained in an end-to-end manner by using standard deep neural network framework. We have quantized the performance of our work in the newly-released SYSU RGB-IR Re-ID benchmark, and have reported superior performance, i. e. , Cumulative Match Characteristic curve (CMC) and Mean Average Precision (MAP), over the state-of-the-art works [Wu et al. , 2017], respectively.

AAMAS Conference 2013 Conference Paper

The Innovative Application of Learning Companions in Virtual Singapura

  • Qiong Wu
  • Xiaogang Han
  • Han Yu
  • Zhiqi Shen
  • Chunyan Miao

Virtual Singapura (VS) is a virtual world based learning environment designed to facilitate learning of the plant transport system. During field studies of VS, we observed that students in virtual world tend to be attracted by visual and auditorial stimuli and be distracted from learning objectives. Also, intensive cognitive load can affect students’ learning experience. To address these issues, we propose two types of companion agent, namely curious companion and remembrance companion. Results collected from the field studies indicate advantages of learning companion augmented virtual world in enhancing students’ learning experience.

EAAI Journal 2007 Journal Article

Nearly optimal neural network stabilization of bipedal standing using genetic algorithm

  • Reza Ghorbani
  • Qiong Wu
  • G. Gary Wang

In this work, stability control of bipedal standing is investigated. The biped is simplified as an inverted pendulum with a foot-link. The controller consists of a general regression neural network (GRNN) feedback control, which stabilizes the inverted pendulum in a region around the upright position, and a PID feedback control, which keeps the pendulum at the upright position. The GRNN controller is also designed to minimize an energy-related cost function while satisfying the constraints between the foot-link and the ground. The optimization has been carried out using the genetic algorithm (GA) and the GRNN is directly trained during optimization iteration process to provide the closed loop feedback optimal controller. The stability of the controlled system is analyzed using the concept of Lyapunov exponents, and a stability region is determined. Simulation results show that the controller can keep the inverted pendulum at the upright position while nearly minimizing an energy-related cost function and keeping the foot-link stationary on the ground. The work contributes to bipedal balancing control, which is important to the development of bipedal robots.

v2026.09.13