Arrow Research search

Author name cluster

Qi Yu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

28 papers
1 author row

Possible papers

28

EAAI Journal 2026 Journal Article

A dual-response colorimetric sensor array integrated with deep learning for mobile intelligent freshness assessment of aquatic products

  • Qi Yu
  • Min Zhang
  • Dayuan Wang
  • Chenlin Wu
  • Chung Lim Law

The growing demand for rapid, on-site monitoring of aquatic product freshness has highlighted the limitations of conventional methods, which often rely on sophisticated instruments and time-consuming procedures. This study presents an engineering-integrated system for rapid freshness assessment of aquatic products (salmon and shrimp), combining a dual-response colorimetric sensor array with systematically evaluated off-the-shelf deep learning models. In terms of engineering, the sensor array adopts a triple-channel design incorporating pH-responsive anthocyanin, alizarin red S, and indole-specific p-dimethylaminobenzaldehyde, enabling simultaneous detection of total volatile basic nitrogen and indole metabolites during spoilage. On the artificial intelligence front, six representative deep learning architectures were benchmarked on this task, with GhostNet and Xception achieving classification accuracy exceeding 98. 00%. By integrating dual-response signals, the system captures complementary spoilage pathways, improving average accuracy from 95. 90% (single-channel) to 97. 13%, demonstrating the algorithmic of dual-response data fusion. Furthermore, for engineering application, the optimized MobileNet_v1 model was successfully deployed in a mobile application, enabling real-time freshness detection with an accuracy of 97. 20% and an inference time of only 12 ms. This work establishes a reliable framework for on-site food quality monitoring, offering a cost-effective and promising alternative to conventional methods while enhancing supply chain transparency.

EAAI Journal 2026 Journal Article

Integrated process planning and scheduling considering automated guided vehicles with an improved deep Q network method

  • Minghai Yuan
  • Qi Yu
  • Liang Zheng
  • Songwei Lu
  • Fengque Pei
  • Wenbin Gu

Process planning and shop scheduling are two essential pillars of industrial production, closely interlinked in determining overall manufacturing efficiency. Integrated process planning and scheduling (IPPS) provides a coordinated framework that enhances resource utilization and minimizes production costs. When transportation operations and automated guided vehicles (AGVs) are also considered, the problem becomes integrated process planning and scheduling with transportation time (IPPS-T). However, most existing IPPS-T studies do not explicitly model AGV task assignment or support reusable, near real-time shop-floor schedules. In particular, a unified learning framework that can handle both IPPS-T with predefined alternative routes (IPPS-T-I) and IPPS-T with more general routing flexibility (IPPS-T-II) remains underexplored. Motivated by real industrial deployment (application in engineering), we propose an improved deep Q network (DQN) method for integrated process planning and scheduling with transportation time (IPPS-T) with explicit AGV allocation, by constructing a unified state and action space for the two types of IPPS-T problems, the same algorithm can be applied to both, thereby extending its applicability and demonstrating a contribution to the field of artificial intelligence (contribution in AI). Finally, five sets of experiments were conducted. In particular, the method was validated using a real-world industrial case involving small-batch gear shaft production with AGV-based logistics. The results show that the proposed algorithm significantly outperforms both combined dispatching rules and traditional DQN, demonstrating its effectiveness and strong generalization capability.

AIIM Journal 2025 Journal Article

ECGEFNet: A two-branch deep learning model for calculating left ventricular ejection fraction using electrocardiogram

  • Yiqiu Qi
  • Guangyuan Li
  • Jinzhu Yang
  • Honghe Li
  • Qi Yu
  • Mingjun Qu
  • Hongxia Ning
  • Yonghuai Wang

Left ventricular systolic dysfunction (LVSD) and its severity are correlated with the prognosis of cardiovascular diseases. Early detection and monitoring of LVSD are of utmost importance. Left ventricular ejection fraction (LVEF) is an essential indicator for evaluating left ventricular function in clinical practice, the current echocardiography-based evaluation method is not avaliable in primary care and difficult to achieve real-time monitoring capabilities for cardiac dysfunction. We propose a two-branch deep learning model (ECGEFNet) for calculating LVEF using electrocardiogram (ECG), which holds the potential to serve as a primary medical screening tool and facilitate long-term dynamic monitoring of cardiac functional impairments. It integrates original numerical signal and waveform plots derived from the signals in an innovative manner, enabling joint calculation of LVEF by incorporating diverse information encompassing temporal, spatial and phase aspects. To address the inadequate information interaction between the two branches and the lack of efficiency in feature fusion, we propose the fusion attention mechanism (FAT) and the two-branch feature fusion module (BFF) to guide the learning, alignment and fusion of features from both branches. We assemble a large internal dataset and perform experimental validation on it. The accuracy of cardiac dysfunction screening is 92. 3%, the mean absolute error (MAE) in LVEF calculation is 4. 57%. The proposed model performs well and outperforms existing basic models, and is of great significance for real-time monitoring of the degree of cardiac dysfunction.

AAAI Conference 2025 Conference Paper

GLEN: Generalized Focal Loss Ensemble of Low-Rank Networks for Calibrated Visual Question Answering

  • Mahsa Mozaffari
  • Hitesh Sapkota
  • Qi Yu

Deep learning models with large-scale backbones have been increasingly adopted to tackle complex visual question answering (VQA) problems in real settings. While providing powerful learning capacities to handle the high-dimensional and multimodal VQA data, these models tend to suffer from the memorization effect leading to overconfident predictions. This can significantly limit their applicability in critical domains (e.g., medicine, cyber-security, and public safety), where confidently wrong predictions may lead to severe consequences. In this work, we propose to perform novel low-rank network factorization, resulting in much better-calibrated networks. These low-rank factorized networks are then aggregated into an ensemble guided by a generalized focal loss to further improve the overall performance and calibration. The overall framework, referred to as the Generalized focal Loss Ensemble of low-rank Networks (GLEN), is an important step toward developing well-calibrated VQA models. We theoretically demonstrate that the generalized focal loss provides a more balanced bias-variance trade-off, which guarantees to lower the confidence of the incorrect predictions, resulting in improved calibration. Extensive experimentation conducted on benchmark datasets and comparison on various VQA models shows that GLEN leads to much better calibration over both in-distribution and out-of-distribution data without sacrificing the VQA accuracy.

AAAI Conference 2025 Conference Paper

Hierarchical Multi-Source Uncertainty Aggregation for Interactive Video Captioning

  • Ervine Zheng
  • Qi Yu

Video captioning automatically generates natural language phrases to explain the contents in video frames. When deploying captioning models in specialized domains, active learning can help reduce the high annotation cost. However, the generative nature of the captioning process is more complex than standard supervised learning tasks and introduces several challenges for active learning in video captioning. Entropy-based uncertainty estimation, which is widely used in active learning, may be inflated in captioning tasks and mislead active sampling. Another challenge arises from the rich content of videos, as each video could be described in multiple ways. A single uncertainty score obtained from one possible caption does not capture the diversity induced by the rich content. To fill out this gap, we propose identifying multiple sources of uncertainty and performing hierarchical aggregation to integrate uncertainty from distinct sources. This innovates a holistic uncertainty metric to quantify the overall informativeness of video content for active sampling. The overall uncertainty is built upon conditional vacuity, an extension of the second-order uncertainty introduced along with the evidential learning framework to the captioning setting, leading to more robust uncertainty estimation without inflation. Both theoretical analysis and experimental evaluation are conducted to demonstrate the effectiveness of the proposed framework for complex uncertainty estimation and interactive learning.

NeurIPS Conference 2024 Conference Paper

Adaptive Important Region Selection with Reinforced Hierarchical Search for Dense Object Detection

  • Dingrong Wang
  • Hitesh Sapkota
  • Qi Yu

Existing state-of-the-art dense object detection techniques tend to produce a large number of false positive detections on difficult images with complex scenes because they focus on ensuring a high recall. To improve the detection accuracy, we propose an Adaptive Important Region Selection (AIRS) framework guided by Evidential Q-learning coupled with a uniquely designed reward function. Inspired by human visual attention, our detection model conducts object search in a top-down, hierarchical fashion. It starts from the top of the hierarchy with the coarsest granularity and then identifies the potential patches likely to contain objects of interest. It then discards non-informative patches and progressively moves downward on the selected ones for a fine-grained search. The proposed evidential Q-learning systematically encodes epistemic uncertainty in its evidential-Q value to encourage the exploration of unknown patches, especially in the early phase of model training. In this way, the proposed model dynamically balances exploration-exploitation to cover both highly valuable and informative patches. Theoretical analysis and extensive experiments on multiple datasets demonstrate that our proposed framework outperforms the SOTA models.

NeurIPS Conference 2024 Conference Paper

Be Confident in What You Know: Bayesian Parameter Efficient Fine-Tuning of Vision Foundation Models

  • Deep S. Pandey
  • Spandan Pyakurel
  • Qi Yu

Large transformer-based foundation models have been commonly used as pre-trained models that can be adapted to different challenging datasets and settings with state-of-the-art generalization performance. Parameter efficient fine-tuning ($\texttt{PEFT}$) provides promising generalization performance in adaptation while incurring minimum computational overhead. However, adaptation of these foundation models through $\texttt{PEFT}$ leads to accurate but severely underconfident models, especially in few-shot learning settings. Moreover, the adapted models lack accurate fine-grained uncertainty quantification capabilities limiting their broader applicability in critical domains. To fill out this critical gap, we develop a novel lightweight {Bayesian Parameter Efficient Fine-Tuning} (referred to as $\texttt{Bayesian-PEFT}$) framework for large transformer-based foundation models. The framework integrates state-of-the-art $\texttt{PEFT}$ techniques with two Bayesian components to address the under-confidence issue while ensuring reliable prediction under challenging few-shot settings. The first component performs base rate adjustment to strengthen the prior belief corresponding to the knowledge gained through pre-training, making the model more confident in its predictions; the second component builds an evidential ensemble that leverages belief regularization to ensure diversity among different ensemble components. Our thorough theoretical analysis justifies that the Bayesian components can ensure reliable and accurate few-shot adaptations with well-calibrated uncertainty quantification. Extensive experiments across diverse datasets, few-shot learning scenarios, and multiple $\texttt{PEFT}$ techniques demonstrate the outstanding prediction and calibration performance by $\texttt{Bayesian-PEFT}$.

AAAI Conference 2024 Conference Paper

Dual-Level Curriculum Meta-Learning for Noisy Few-Shot Learning Tasks

  • Xiaofan Que
  • Qi Yu

Few-shot learning (FSL) is essential in many practical applications. However, the limited training examples make the models more vulnerable to label noise, which can lead to poor generalization capability. To address this critical challenge, we propose a curriculum meta-learning model that employs a novel dual-level class-example sampling strategy to create a robust curriculum for adaptive task distribution formulation and robust model training. The dual-level framework proposes a heuristic class sampling criterion that measures pairwise class boundary complexity to form a class curriculum; it uses effective example sampling through an under-trained proxy model to form an example curriculum. By utilizing both class-level and example-level information, our approach is more robust to handle limited training data and noisy labels that commonly occur in few-shot learning tasks. The model has efficient convergence behavior, which is verified through rigorous convergence analysis. Additionally, we establish a novel error bound through a hierarchical PAC-Bayesian analysis for curriculum meta-learning under noise. We conduct extensive experiments that demonstrate the effectiveness of our framework in outperforming existing noisy few-shot learning methods under various few-shot classification benchmarks. Our code is available at https://github.com/ritmininglab/DCML.

NeurIPS Conference 2024 Conference Paper

Evidential Mixture Machines: Deciphering Multi-Label Correlations for Active Learning Sensitivity

  • Dayou Yu
  • Minghao Li
  • Weishi Shi
  • Qi Yu

Multi-label active learning is a crucial yet challenging area in contemporary machine learning, often complicated by a large and sparse label space. This challenge is further exacerbated in active learning scenarios where labeling resources are constrained. Drawing inspiration from existing mixture of Bernoulli models, which efficiently compress the label space into a more manageable weight coefficient space by learning correlated Bernoulli components, we propose a novel model called Evidential Mixture Machines (EMM). Our model leverages mixture components derived from unsupervised learning in the label space and improves prediction accuracy by predicting weight coefficients following the evidential learning paradigm. These coefficients are aggregated as proxy pseudo counts to enhance component offset predictions. The evidential learning approach provides an uncertainty-aware connection between input features and the predicted coefficients and components. Additionally, our method combines evidential uncertainty with predicted label embedding covariances for active sample selection, creating a richer, multi-source uncertainty metric beyond traditional uncertainty scores. Experiments on synthetic datasets show the effectiveness of evidential uncertainty prediction and EMM's capability to capture label correlations through predicted components. Further testing on real-world datasets demonstrates improved performance compared to existing multi-label active learning methods.

NeurIPS Conference 2024 Conference Paper

Evidential Stochastic Differential Equations for Time-Aware Sequential Recommendation

  • Krishna P. Neupane
  • Ervine Zheng
  • Qi Yu

Sequential recommender systems are designed to capture users' evolving interests over time. Existing methods typically assume a uniform time interval among consecutive user interactions and may not capture users' continuously evolving behavior in the short and long term. In reality, the actual time intervals of user interactions vary dramatically. Consequently, as the time interval between interactions increases, so does the uncertainty in user behavior. Intuitively, it is beneficial to establish a correlation between the interaction time interval and the model uncertainty to provide effective recommendations. To this end, we formulate a novel Evidential Neural Stochastic Differential Equation ( E-NSDE ) to seamlessly integrate NSDE and evidential learning for effective time-aware sequential recommendations. The NSDE enables the model to learn users' fine-grained time-evolving behavior by capturing continuous user representation while evidential learning quantifies both aleatoric and epistemic uncertainties considering interaction time interval to provide model confidence during prediction. Furthermore, we derive a mathematical relationship between the interaction time interval and model uncertainty to guide the learning process. Experiments on real-world data demonstrate the effectiveness of the proposed method compared to the SOTA methods.

NeurIPS Conference 2023 Conference Paper

Actively Testing Your Model While It Learns: Realizing Label-Efficient Learning in Practice

  • Dayou Yu
  • Weishi Shi
  • Qi Yu

In active learning (AL), we focus on reducing the data annotation cost from the model training perspective. However, "testing'', which often refers to the model evaluation process of using empirical risk to estimate the intractable true generalization risk, also requires data annotations. The annotation cost for "testing'' (model evaluation) is under-explored. Even in works that study active model evaluation or active testing (AT), the learning and testing ends are disconnected. In this paper, we propose a novel active testing while learning (ATL) framework that integrates active learning with active testing. ATL provides an unbiased sample-efficient estimation of the model risk during active learning. It leverages test samples annotated from different periods of a dynamic active learning process to achieve fair model evaluations based on a theoretically guaranteed optimal integration of different test samples. Periodic testing also enables effective early-stopping to further save the total annotation cost. ATL further integrates an "active feedback'' mechanism, which is inspired by human learning, where the teacher (active tester) provides immediate guidance given by the prior performance of the student (active learner). Our theoretical result reveals that active feedback maintains the label complexity of the integrated learning-testing objective, while improving the model's generalization capability. We study the realistic setting where we maximize the performance gain from choosing "testing'' samples for feedback without sacrificing the risk estimation accuracy. An agnostic-style analysis and empirical evaluations on real-world datasets demonstrate that the ATL framework can effectively improve the annotation efficiency of both active learning and evaluation tasks.

NeurIPS Conference 2023 Conference Paper

Distributionally Robust Ensemble of Lottery Tickets Towards Calibrated Sparse Network Training

  • Hitesh Sapkota
  • Dingrong Wang
  • Zhiqiang Tao
  • Qi Yu

The recently developed sparse network training methods, such as Lottery Ticket Hypothesis (LTH) and its variants, have shown impressive learning capacity by finding sparse sub-networks from a dense one. While these methods could largely sparsify deep networks, they generally focus more on realizing comparable accuracy to dense counterparts yet neglect network calibration. However, how to achieve calibrated network predictions lies at the core of improving model reliability, especially when it comes to addressing the overconfident issue and out-of-distribution cases. In this study, we propose a novel Distributionally Robust Optimization (DRO) framework to achieve an ensemble of lottery tickets towards calibrated network sparsification. Specifically, the proposed DRO ensemble aims to learn multiple diverse and complementary sparse sub-networks (tickets) with the guidance of uncertainty sets, which encourage tickets to gradually capture different data distributions from easy to hard and naturally complement each other. We theoretically justify the strong calibration performance by showing how the proposed robust training process guarantees to lower the confidence of incorrect predictions. Extensive experimental results on several benchmarks show that our proposed lottery ticket ensemble leads to a clear calibration improvement without sacrificing accuracy and burdening inference costs. Furthermore, experiments on OOD datasets demonstrate the robustness of our approach in the open-set environment.

AAAI Conference 2023 Conference Paper

Evidential Conditional Neural Processes

  • Deep Shankar Pandey
  • Qi Yu

The Conditional Neural Process (CNP) family of models offer a promising direction to tackle few-shot problems by achieving better scalability and competitive predictive performance. However, the current CNP models only capture the overall uncertainty for the prediction made on a target data point. They lack a systematic fine-grained quantification on the distinct sources of uncertainty that are essential for model training and decision-making under the few-shot setting. We propose Evidential Conditional Neural Processes (ECNP), which replace the standard Gaussian distribution used by CNP with a much richer hierarchical Bayesian structure through evidential learning to achieve epistemic-aleatoric uncertainty decomposition. The evidential hierarchical structure also leads to a theoretically justified robustness over noisy training tasks. Theoretical analysis on the proposed ECNP establishes the relationship with CNP while offering deeper insights on the roles of the evidential parameters. Extensive experiments conducted on both synthetic and real-world data demonstrate the effectiveness of our proposed model in various few-shot settings.

AAAI Conference 2023 Conference Paper

Scaling Up Dynamic Graph Representation Learning via Spiking Neural Networks

  • Jintang Li
  • Zhouxin Yu
  • Zulun Zhu
  • Liang Chen
  • Qi Yu
  • Zibin Zheng
  • Sheng Tian
  • Ruofan Wu

Recent years have seen a surge in research on dynamic graph representation learning, which aims to model temporal graphs that are dynamic and evolving constantly over time. However, current work typically models graph dynamics with recurrent neural networks (RNNs), making them suffer seriously from computation and memory overheads on large temporal graphs. So far, scalability of dynamic graph representation learning on large temporal graphs remains one of the major challenges. In this paper, we present a scalable framework, namely SpikeNet, to efficiently capture the temporal and structural patterns of temporal graphs. We explore a new direction in that we can capture the evolving dynamics of temporal graphs with spiking neural networks (SNNs) instead of RNNs. As a low-power alternative to RNNs, SNNs explicitly model graph dynamics as spike trains of neuron populations and enable spike-based propagation in an efficient way. Experiments on three large real-world temporal graph datasets demonstrate that SpikeNet outperforms strong baselines on the temporal node classification task with lower computational costs. Particularly, SpikeNet generalizes to a large temporal graph (2.7M nodes and 13.9M edges) with significantly fewer parameters and computation overheads.

AAAI Conference 2023 Conference Paper

Sparse Maximum Margin Learning from Multimodal Human Behavioral Patterns

  • Ervine Zheng
  • Qi Yu
  • Zhi Zheng

We propose a multimodal data fusion framework to systematically analyze human behavioral data from specialized domains that are inherently dynamic, sparse, and heterogeneous. We develop a two-tier architecture of probabilistic mixtures, where the lower tier leverages parametric distributions from the exponential family to extract significant behavioral patterns from each data modality. These patterns are then organized into a dynamic latent state space at the higher tier to fuse patterns from different modalities. In addition, our framework jointly performs pattern discovery and maximum-margin learning for downstream classification tasks by using a group-wise sparse prior that regularizes the coefficients of the maximum-margin classifier. Therefore, the discovered patterns are highly interpretable and discriminative to support downstream classification tasks. Experiments on real-world behavioral data from medical and psychological domains demonstrate that our framework discovers meaningful multimodal behavioral patterns with improved interpretability and prediction performance.

AAAI Conference 2023 Conference Paper

STARS: Spatial-Temporal Active Re-sampling for Label-Efficient Learning from Noisy Annotations

  • Dayou Yu
  • Weishi Shi
  • Qi Yu

Active learning (AL) aims to sample the most informative data instances for labeling, which makes the model fitting data efficient while significantly reducing the annotation cost. However, most existing AL models make a strong assumption that the annotated data instances are always assigned correct labels, which may not hold true in many practical settings. In this paper, we develop a theoretical framework to formally analyze the impact of noisy annotations and show that systematically re-sampling guarantees to reduce the noise rate, which can lead to improved generalization capability. More importantly, the theoretical framework demonstrates the key benefit of conducting active re-sampling on label-efficient learning, which is critical for AL. The theoretical results also suggest essential properties of an active re-sampling function with a fast convergence speed and guaranteed error reduction. This inspires us to design a novel spatial-temporal active re-sampling function by leveraging the important spatial and temporal properties of maximum-margin classifiers. Extensive experiments conducted on both synthetic and real-world data clearly demonstrate the effectiveness of the proposed active re-sampling function.

AAAI Conference 2022 Conference Paper

A Dynamic Meta-Learning Model for Time-Sensitive Cold-Start Recommendations

  • Krishna Prasad Neupane
  • Ervine Zheng
  • Yu Kong
  • Qi Yu

We present a novel dynamic recommendation model that focuses on users who have interactions in the past but turn relatively inactive recently. Making effective recommendations to these time-sensitive cold-start users is critical to maintain the user base of a recommender system. Due to the sparse recent interactions, it is challenging to capture these users’ current preferences precisely. Solely relying on their historical interactions may also lead to outdated recommendations misaligned with their recent interests. The proposed model leverages historical and current user-item interactions and dynamically factorizes a user’s (latent) preference into time-specific and time-evolving representations that jointly affect user behaviors. These latent factors further interact with an optimized item embedding to achieve accurate and timely recommendations. Experiments over real-world data help demonstrate the effectiveness of the proposed time-sensitive coldstart recommendation model.

IJCAI Conference 2022 Conference Paper

Spiking Graph Convolutional Networks

  • Zulun Zhu
  • Jiaying Peng
  • Jintang Li
  • Liang Chen
  • Qi Yu
  • Siqiang Luo

Graph Convolutional Networks (GCNs) achieve an impressive performance due to the remarkable representation ability in learning the graph information. However, GCNs, when implemented on a deep network, require expensive computation power, making them difficult to be deployed on battery-powered devices. In contrast, Spiking Neural Networks (SNNs), which perform a bio-fidelity inference process, offer an energy-efficient neural architecture. In this work, we propose SpikingGCN, an end-to-end framework that aims to integrate the embedding of GCNs with the biofidelity characteristics of SNNs. The original graph data are encoded into spike trains based on the incorporation of graph convolution. We further model biological information processing by utilizing a fully connected layer combined with neuron nodes. In a wide range of scenarios (e. g. , citation networks, image graph classification, and recommender systems), our experimental results show that the proposed method could gain competitive performance against state-of-the-art approaches. Furthermore, we show that SpikingGCN on a neuromorphic chip can bring a clear advantage of energy efficiency into graph data analysis, which demonstrates its great potential to construct environment-friendly machine learning models.

AAAI Conference 2021 Conference Paper

A Continual Learning Framework for Uncertainty-Aware Interactive Image Segmentation

  • Ervine Zheng
  • Qi Yu
  • Rui Li
  • Pengcheng Shi
  • Anne Haake

Deep learning models have achieved state-of-the-art performance in semantic image segmentation, but the results provided by fully automatic algorithms are not always guaranteed satisfactory to users. Interactive segmentation offers a solution by accepting user annotations on selective areas of the images to refine the segmentation results. However, most existing models only focus on correcting the current image’s misclassified pixels, with no knowledge carried over to other images. In this work, we formulate interactive image segmentation as a continual learning problem and propose a framework to effectively learn from user annotations, aiming to improve the segmentation on both the current image and unseen images in future tasks while avoiding deteriorated performance on previously-seen images. It employs a probabilistic mask to control the neural network’s kernel activation and extract the most suitable features for segmenting images in each task. We also apply a task-aware embedding to automatically infer the optimal kernel activation for initial segmentation and subsequent refinement. Interactions with users are guided through multi-source uncertainty estimation so that users can focus on the most important areas to minimize the overall manual annotation effort. Experiments are performed on both medical and natural image datasets to illustrate the proposed framework’s effectiveness on basic segmentation performance, forward knowledge transfer, and backward knowledge transfer.

NeurIPS Conference 2021 Conference Paper

A Gaussian Process-Bayesian Bernoulli Mixture Model for Multi-Label Active Learning

  • Weishi Shi
  • Dayou Yu
  • Qi Yu

Multi-label classification (MLC) allows complex dependencies among labels, making it more suitable to model many real-world problems. However, data annotation for training MLC models becomes much more labor-intensive due to the correlated (hence non-exclusive) labels and a potential large and sparse label space. We propose to conduct multi-label active learning (ML-AL) through a novel integrated Gaussian Process-Bayesian Bernoulli Mixture model (GP-B$^2$M) to accurately quantify a data sample's overall contribution to a correlated label space and choose the most informative samples for cost-effective annotation. In particular, the B$^2$M encodes label correlations using a Bayesian Bernoulli mixture of label clusters, where each mixture component corresponds to a global pattern of label correlations. To tackle highly sparse labels under AL, the B$^2$M is further integrated with a predictive GP to connect data features as an effective inductive bias and achieve a feature-component-label mapping. The GP predicts coefficients of mixture components that help to recover the final set of labels of a data sample. A novel auxiliary variable based variational inference algorithm is developed to tackle the non-conjugacy introduced along with the mapping process for efficient end-to-end posterior inference. The model also outputs a predictive distribution that provides both the label prediction and their correlations in the form of a label covariance matrix. A principled sampling function is designed accordingly to naturally capture both the feature uncertainty (through GP) and label covariance (through B$^2$M) for effective data sampling. Experiments on real-world multi-label datasets demonstrate the state-of-the-art AL performance of the proposed GP-B$^2$M model.

NeurIPS Conference 2020 Conference Paper

Dynamic Fusion of Eye Movement Data and Verbal Narrations in Knowledge-rich Domains

  • Ervine Zheng
  • Qi Yu
  • Rui Li
  • Pengcheng Shi
  • Anne Haake

We propose to jointly analyze experts' eye movements and verbal narrations to discover important and interpretable knowledge patterns to better understand their decision-making processes. The discovered patterns can further enhance data-driven statistical models by fusing experts' domain knowledge to support complex human-machine collaborative decision-making. Our key contribution is a novel dynamic Bayesian nonparametric model that assigns latent knowledge patterns into key phases involved in complex decision-making. Each phase is characterized by a unique distribution of word topics discovered from verbal narrations and their dynamic interactions with eye movement patterns, indicating experts' special perceptual behavior within a given decision-making stage. A new split-merge-switch sampler is developed to efficiently explore the posterior state space with an improved mixing rate. Case studies on diagnostic error prediction and disease morphology categorization help demonstrate the effectiveness of the proposed model and discovered knowledge patterns.

NeurIPS Conference 2020 Conference Paper

Multifaceted Uncertainty Estimation for Label-Efficient Deep Learning

  • Weishi Shi
  • Xujiang Zhao
  • Feng Chen
  • Qi Yu

We present a novel multi-source uncertainty prediction approach that enables deep learning (DL) models to be actively trained with much less labeled data. By leveraging the second-order uncertainty representation provided by subjective logic (SL), we conduct evidence-based theoretical analysis and formally decompose the predicted entropy over multiple classes into two distinct sources of uncertainty: vacuity and dissonance, caused by lack of evidence and conflict of strong evidence, respectively. The evidence based entropy decomposition provides deeper insights on the nature of uncertainty, which can help effectively explore a large and high-dimensional unlabeled data space. We develop a novel loss function that augments DL based evidence prediction with uncertainty anchor sample identification. The accurately estimated multiple sources of uncertainty are systematically integrated and dynamically balanced using a data sampling function for label-efficient active deep learning (ADL). Experiments conducted over both synthetic and real data and comparison with competitive AL methods demonstrate the effectiveness of the proposed ADL model.

NeurIPS Conference 2019 Conference Paper

Integrating Bayesian and Discriminative Sparse Kernel Machines for Multi-class Active Learning

  • Weishi Shi
  • Qi Yu

We propose a novel active learning (AL) model that integrates Bayesian and discriminative kernel machines for fast and accurate multi-class data sampling. By joining a sparse Bayesian model and a maximum margin machine under a unified kernel machine committee (KMC), the proposed model is able to identify a small number of data samples that best represent the overall data space while accurately capturing the decision boundaries. The integration is conducted using the maximum entropy discrimination framework, resulting in a joint objective function that contains generalized entropy as a regularizer. Such a property allows the proposed AL model to choose data samples that more effectively handle non-separable classification problems. Parameter learning is achieved through a principled optimization framework that leverages convex duality and sparse structure of KMC to efficiently optimize the joint objective function. Key model parameters are used to design a novel sampling function to choose data samples that can simultaneously improve multiple decision boundaries, making it an effective sampler for problems with a large number of classes. Experiments conducted over both synthetic and real data and comparison with competitive AL methods demonstrate the effectiveness of the proposed model.

TAAS Journal 2017 Journal Article

Integrating Reinforcement Learning with Multi-Agent Techniques for Adaptive Service Composition

  • Hongbign Wang
  • Xin Chen
  • Qin Wu
  • Qi Yu
  • Xingguo Hu
  • Zibin Zheng
  • Athman Bouguettaya

Service-oriented architecture is a widely used software engineering paradigm to cope with complexity and dynamics in enterprise applications. Service composition, which provides a cost-effective way to implement software systems, has attracted significant attention from both industry and research communities. As online services may keep evolving over time and thus lead to a highly dynamic environment, service composition must be self-adaptive to tackle uninformed behavior during the evolution of services. In addition, service composition should also maintain high efficiency for large-scale services, which are common for enterprise applications. This article presents a new model for large-scale adaptive service composition based on multi-agent reinforcement learning. The model integrates reinforcement learning and game theory, where the former is to achieve adaptation in a highly dynamic environment and the latter is to enable agents to work for a common task (i.e., composition). In particular, we propose a multi-agent Q-learning algorithm for service composition, which is expected to achieve better performance when compared with the single-agent Q-learning method and multi-agent SARSA (State-Action-Reward-State-Action) method. Our experimental results demonstrate the effectiveness and efficiency of our approach.

AIIM Journal 2017 Journal Article

Knowledge graph for TCM health preservation: Design, construction, and applications

  • Tong Yu
  • Jinghua Li
  • Qi Yu
  • Ye Tian
  • Xiaofeng Shun
  • Lili Xu
  • Ling Zhu
  • Hongjie Gao

Traditional Chinese Medicine (TCM) is one of the important non-material cultural heritages of the Chinese nation. It is an important development strategy of Chinese medicine to collect, analyzes, and manages the knowledge assets of TCM health care. As a novel and massive knowledge management technology, knowledge graph provides an ideal technical means to solve the problem of “Knowledge Island” in the field of traditional Chinese medicine. In this study, we construct a large-scale knowledge graph, which integrates terms, documents, databases and other knowledge resources. This knowledge graph can facilitate various knowledge services such as knowledge visualization, knowledge retrieval, and knowledge recommendation, and helps the sharing, interpretation, and utilization of TCM health care knowledge.

IJCAI Conference 2017 Conference Paper

Modeling Physicians' Utterances to Explore Diagnostic Decision-making

  • Xuan Guo
  • Rui Li
  • Qi Yu
  • Anne Haake

Diagnostic error prevention is a long-established but specialized topic in clinical and psychological research. In this paper, we contribute to the field by exploring diagnostic decision-making via modeling physicians' utterances of medical concepts during image-based diagnoses. We conduct experiments to collect verbal narratives from dermatologists while they are examining and describing dermatology images towards diagnoses. We propose a hierarchical probabilistic framework to learn domain-specific patterns from the medical concepts in these narratives. The discovered patterns match the diagnostic units of thought identified by domain experts. These meaningful patterns uncover physicians' diagnostic decision-making processes while parsing the image content. Our evaluation shows that these patterns provide key information to classify narratives by diagnostic correctness levels.

AIIM Journal 2014 Journal Article

From spoken narratives to domain knowledge: Mining linguistic data for medical image understanding

  • Xuan Guo
  • Qi Yu
  • Cecilia Ovesdotter Alm
  • Cara Calvelli
  • Jeff B. Pelz
  • Pengcheng Shi
  • Anne R. Haake

Objectives Extracting useful visual clues from medical images allowing accurate diagnoses requires physicians’ domain knowledge acquired through years of systematic study and clinical training. This is especially true in the dermatology domain, a medical specialty that requires physicians to have image inspection experience. Automating or at least aiding such efforts requires understanding physicians’ reasoning processes and their use of domain knowledge. Mining physicians’ references to medical concepts in narratives during image-based diagnosis of a disease is an interesting research topic that can help reveal experts’ reasoning processes. It can also be a useful resource to assist with design of information technologies for image use and for image case-based medical education systems. Methods and materials We collected data for analyzing physicians’ diagnostic reasoning processes by conducting an experiment that recorded their spoken descriptions during inspection of dermatology images. In this paper we focus on the benefit of physicians’ spoken descriptions and provide a general workflow for mining medical domain knowledge based on linguistic data from these narratives. The challenge of a medical image case can influence the accuracy of the diagnosis as well as how physicians pursue the diagnostic process. Accordingly, we define two lexical metrics for physicians’ narratives—lexical consensus score and top N relatedness score—and evaluate their usefulness by assessing the diagnostic challenge levels of corresponding medical images. We also report on clustering medical images based on anchor concepts obtained from physicians’ medical term usage. These analyses are based on physicians’ spoken narratives that have been preprocessed by incorporating the Unified Medical Language System for detecting medical concepts. Results The image rankings based on lexical consensus score and on top 1 relatedness score are well correlated with those based on challenge levels (Spearman correlation >0. 5 and Kendall correlation >0. 4). Clustering results are largely improved based on our anchor concept method (accuracy >70% and mutual information >80%). Conclusions Physicians’ spoken narratives are valuable for the purpose of mining the domain knowledge that physicians use in medical image inspections. We also show that the semantic metrics introduced in the paper can be successfully applied to medical image understanding and allow discussion of additional uses of these metrics.

AAAI Conference 2011 Conference Paper

A Feasible Nonconvex Relaxation Approach to Feature Selection

  • Cuixia Gao
  • Naiyan Wang
  • Qi Yu
  • Zhihua Zhang

Variable selection problems are typically addressed under a penalized optimization framework. Nonconvex penalties such as the minimax concave plus (MCP) and smoothly clipped absolute deviation (SCAD), have been demonstrated to have the properties of sparsity practically and theoretically. In this paper we propose a new nonconvex penalty that we call exponential-type penalty. The exponential-type penalty is characterized by a positive parameter, which establishes a connection with the 0 and 1 penalties. We apply this new penalty to sparse supervised learning problems. To solve to resulting optimization problem, we resort to a reweighted 1 minimization method. Moreover, we devise an ef- ficient method for the adaptive update of the tuning parameter. Our experimental results are encouraging. They show that the exponential-type penalty is competitive with MCP and SCAD.

v2026.09.13