Arrow Research search

Author name cluster

Kun Yang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

19 papers
2 author rows

Possible papers

19

JBHI Journal 2026 Journal Article

FourierMask: Explain EEG-Based End-to-End Deep Learning Models in the Frequency Domain

  • Hanqi Wang
  • Jingyu Zhang
  • Kun Yang
  • Jichuan Xiong
  • Xuefeng Liu
  • Tao Chen
  • Liang Song

The rise of EEG-based end-to-end deep learning models has underscored the need to elucidate how these models process time-series raw EEG signals to generate predictions. The frequency domain provides a more suitable perspective for this task due to two key advantages: the strong correlation with cognitive states and the inherent capacity to model long-range temporal dependencies. However, this perspective remains underexplored in existing research. To bridge this gap, we propose FourierMask, the first mask perturbation framework specifically designed for frequency-domain explanation of EEG-based end-to-end models. Our method introduces three key innovations. First, the Fourier-based domain transformation enables direct manipulation of spectral components. Second, A learnable mask mechanism jointly models the spectral-spatial couplings relationship for EEG explanation. Third, a perturbation generator constrained by a target alignment loss ensures natural perturbations by minimizing distribution shift via cluster-aware regularization. We validate our method through experiments on an EEG benchmark dataset across EEGNet, TSCeption, and DeepConvNet models. Our method reaches a 36. 0% average accuracy drop gap (vs. 8. 6% for LIME and 6. 6% for easyPEASI) at the group-level. And, it reaches a 17. 8% average accuracy drop gap (vs. 8. 9% for LIME and 9. 9% for easyPEASI) at the instance-level. Our model-agnostic framework provides a plug-and-play solution for enhancing transparency of EEG-based end-to-end deep learning models. It links model decisions to frequency biomarkers, with potential applications in neuromedicine and brain-computer interfaces.

AAAI Conference 2026 Conference Paper

MEML-GRPO: Heterogeneous Multi-Expert Mutual Learning for RLVR Advancement

  • Weitao Jia
  • Jinghui Lu
  • Haiyang Yu
  • Siqi Wang
  • Guozhi Tang
  • An-Lan Wang
  • Weijie Yin
  • Dingkang Yang

Recent advances demonstrate that reinforcement learning with verifiable rewards (RLVR) significantly enhances the reasoning capabilities of large language models (LLMs). However, standard RLVR faces challenges with reward sparsity, where zero rewards from consistently incorrect candidate answers provide no learning signal, particularly in challenging tasks. To address this,we propose Multi-Expert Mutual Learning GRPO (MEML-GRPO), an innovative framework that utilizes diverse expert prompts as system prompts to generate a broader range of responses, substantially increasing the likelihood of identifying correct solutions. Additionally, we introduce an inter-expert mutual learning mechanism that facilitates knowledge sharing and transfer among experts, further boosting the model’s performance through RLVR. Extensive experiments across multiple reasoning benchmarks show that MEML-GRPO delivers significant improvements, achieving an average performance gain of 4.89% with Qwen and 11.33% with Llama, effectively overcoming the core limitations of traditional RLVR methods.

IROS Conference 2025 Conference Paper

Autonomous Suturing Method for Robot-Assisted Minimally Invasive Surgery

  • Mei Feng
  • Haoju Li
  • Yao Li
  • Kun Yang
  • Dong He
  • Xiuquan Lu

Robot-assisted minimally invasive surgery is widely used because of its superior postoperative recovery outcomes. However, the workload for surgeons remains high. The development of autonomous suturing capabilities in surgical robots is poised to significantly reduce surgeon workload. In this study, we present a novel method or autonomous suturing using a minimally invasive surgical robot. We quantify the surgical suturing requirements and propose corresponding metrics for evaluating the suturing effect. We also use the dynamic adjustment of stitch position to optimize the surgical robot autonomous suturing scheme. Furthermore, we employ particle swarm algorithms to enhance the grasping posture of surgical instruments, enabling the robot to achieve optimal suture needle clamping. Our method maintains the same level of expert operator in the suturing parametric index of suturing when suturing two types of wounds: gauze and egg membrane. The autonomous suturing method proposed in this study is currently deployed on our own surgical robot, and it can be generalized to other surgical robots. This will lay the foundation for surgical robots to achieve fully autonomous surgery. The experimental results show that the stitching effect of our proposed autonomous robot stitching method is already close to that of surgeons using the same robot, and it maintains good consistency in multiple sets of experiments. The method proposed in this study can be generalized to various other surgical robots, laying the foundation for surgical robots to achieve fully autonomous surgery.

EAAI Journal 2025 Journal Article

Machine learning methods comparison for maritime wireless signal strength prediction

  • Lisha Peng
  • Kun Yang
  • Jianming Wu
  • Chengyuan Wen
  • Tong Peng
  • Xin Li

The 5th generation mobile networks (5G) signal strength prediction techniques based on artificial intelligence (AI) demonstrate numerous advantages, such as easy implementation and high accuracy. However, what kind of machine learning model can provide superior performance in maritime applications is still unknown. In this paper, we developed twelve predictive models, encompassing eight machine learning and four linear regression approaches, for estimating reference signal receiving power (RSRP) and received signal strength indicator (RSSI) in maritime land-to-ship (L2S) scenarios. Our models are rigorously evaluated using nine metrics, trained on 5G band data collected with our self-designed equipment, and refined with carefully selected features to maximize prediction accuracy. The results show that the machine learning models are generally superior to the linear regression models in the fitting and prediction of RSRP and RSSI according to our evaluation and prediction experiments. Especially, rule-based models like decision tree regression (DTR) can accurately learn the impact of both large-scale and small-scale fading on the prediction objects from the data, and also have strong model interpretability. Our work has provided good reference value for future machine learning-based channel modeling and model evaluations.

AAAI Conference 2024 Conference Paper

A Unified Self-Distillation Framework for Multimodal Sentiment Analysis with Uncertain Missing Modalities

  • Mingcheng Li
  • Dingkang Yang
  • Yuxuan Lei
  • Shunli Wang
  • Shuaibing Wang
  • Liuzhen Su
  • Kun Yang
  • Yuzheng Wang

Multimodal Sentiment Analysis (MSA) has attracted widespread research attention recently. Most MSA studies are based on the assumption of modality completeness. However, many inevitable factors in real-world scenarios lead to uncertain missing modalities, which invalidate the fixed multimodal fusion approaches. To this end, we propose a Unified multimodal Missing modality self-Distillation Framework (UMDF) to handle the problem of uncertain missing modalities in MSA. Specifically, a unified self-distillation mechanism in UMDF drives a single network to automatically learn robust inherent representations from the consistent distribution of multimodal data. Moreover, we present a multi-grained crossmodal interaction module to deeply mine the complementary semantics among modalities through coarse- and fine-grained crossmodal attention. Eventually, a dynamic feature integration module is introduced to enhance the beneficial semantics in incomplete modalities while filtering the redundant information therein to obtain a refined and robust multimodal representation. Comprehensive experiments on three datasets demonstrate that our framework significantly improves MSA performance under both uncertain missing-modality and complete-modality testing conditions.

YNICL Journal 2024 Journal Article

A whole-brain neuromark resting-state fMRI analysis of first-episode and early psychosis: Evidence of aberrant cortical-subcortical-cerebellar functional circuitry

  • Kyle M. Jensen
  • Vince D. Calhoun
  • Zening Fu
  • Kun Yang
  • Andreia V. Faria
  • Koko Ishizuka
  • Akira Sawa
  • Pablo Andrés-Camazón

Psychosis (including symptoms of delusions, hallucinations, and disorganized conduct/speech) is a main feature of schizophrenia and is frequently present in other major psychiatric illnesses. Studies in individuals with first-episode (FEP) and early psychosis (EP) have the potential to interpret aberrant connectivity associated with psychosis during a period with minimal influence from medication and other confounds. The current study uses a data-driven whole-brain approach to examine patterns of aberrant functional network connectivity (FNC) in a multi-site dataset comprising resting-state functional magnetic resonance images (rs-fMRI) from 117 individuals with FEP or EP and 130 individuals without a psychiatric disorder, as controls. Accounting for age, sex, race, head motion, and multiple imaging sites, differences in FNC were identified between psychosis and control participants in cortical (namely the inferior frontal gyrus, superior medial frontal gyrus, postcentral gyrus, supplementary motor area, posterior cingulate cortex, and superior and middle temporal gyri), subcortical (the caudate, thalamus, subthalamus, and hippocampus), and cerebellar regions. The prominent pattern of reduced cerebellar connectivity in psychosis is especially noteworthy, as most studies focus on cortical and subcortical regions, neglecting the cerebellum. The dysconnectivity reported here may indicate disruptions in cortical-subcortical-cerebellar circuitry involved in rudimentary cognitive functions which may serve as reliable correlates of psychosis.

JBHI Journal 2024 Journal Article

Automatically Extracting and Utilizing EEG Channel Importance Based on Graph Convolutional Network for Emotion Recognition

  • Kun Yang
  • Zhenning Yao
  • Keze Zhang
  • Jing Xu
  • Li Zhu
  • Shichao Cheng
  • Jianhai Zhang

Graph convolutional network (GCN) based on the brain network has been widely used for EEG emotion recognition. However, most studies train their models directly without considering network dimensionality reduction beforehand. In fact, some nodes and edges are invalid information or even interference information for the current task. It is necessary to reduce the network dimension and extract the core network. To address the problem of extracting and utilizing the core network, a core network extraction model (CWGCN) based on channel weighting and graph convolutional network and a graph convolutional network model (CCSR-GCN) based on channel convolution and style-based recalibration for emotion recognition have been proposed. The CWGCN model automatically extracts the core network and the channel importance parameter in a data-driven manner. The CCSR-GCN model innovatively uses the output information of the CWGCN model to identify the emotion state. The experimental results on SEED show that: 1) the core network extraction can help improve the performance of the GCN model; 2) the models of CWGCN and CCSR-GCN achieve better results than the currently popular methods. The idea and its implementation in this paper provide a novel and successful perspective for the application of GCN in brain network analysis of other specific tasks.

JBHI Journal 2024 Journal Article

DSFE: Decoding EEG-Based Finger Motor Imagery Using Feature-Dependent Frequency, Feature Fusion and Ensemble Learning

  • Kun Yang
  • Ruochen Li
  • Jing Xu
  • Li Zhu
  • Wanzeng Kong
  • Jianhai Zhang

Accurate decoding finger motor imagery is essential for fine motor control using EEG signals. However, decoding finger motor imagery is particularly challenging compared with ordinary motor imagery. This paper proposed a novel EEG decoding method of feature-dependent frequency band selection, feature fusion, and ensemble learning (DSFE) for finger motor imagery. First, a feature-dependent frequency band selection method based on correlation coefficient (FDCC) was proposed to select feature-specific effective bands. Second, a feature fusion method was proposed to fuse different types of candidate features to produce multiple refined sets of decoding features. Finally, an ensemble model using the weighted voting strategy was proposed to make full use of these diverse sets of final features. The results on a public EEG dataset of five fingers motor imagery showed that the DSFE method is effective and achieves the highest decoding accuracy of 50. 64%, which is 7. 64% higher than existing studies using exactly the same data. The experiments further revealed that both the effective frequency bands of different subjects and the effective frequency bands of different types of features are different in finger motor imagery. Furthermore, compared with two-hand motor imagery, the effective decoding information of finger motor imagery is transferred to the lower frequency. The idea and findings in this paper provide a valuable perspective for understanding fine motor imagery in-depth.

NeurIPS Conference 2024 Conference Paper

Efficient Prompt Optimization Through the Lens of Best Arm Identification

  • Chengshuai Shi
  • Kun Yang
  • Zihan Chen
  • Jundong Li
  • Jing Yang
  • Cong Shen

The remarkable instruction-following capability of large language models (LLMs) has sparked a growing interest in automatically finding good prompts, i. e. , prompt optimization. Most existing works follow the scheme of selecting from a pre-generated pool of candidate prompts. However, these designs mainly focus on the generation strategy, while limited attention has been paid to the selection method. Especially, the cost incurred during the selection (e. g. , accessing LLM and evaluating the responses) is rarely explicitly considered. To overcome this limitation, this work provides a principled framework, TRIPLE, to efficiently perform prompt selection under an explicit budget constraint. TRIPLE is built on a novel connection established between prompt optimization and fixed-budget best arm identification (BAI-FB) in multi-armed bandits (MAB); thus, it is capable of leveraging the rich toolbox from BAI-FB systematically and also incorporating unique characteristics of prompt optimization. Extensive experiments on multiple well-adopted tasks using various LLMs demonstrate the remarkable performance improvement of TRIPLE over baselines while satisfying the limited budget constraints. As an extension, variants of TRIPLE are proposed to efficiently select examples for few-shot prompts, also achieving superior empirical performance.

TMLR Journal 2024 Journal Article

Harnessing the Power of Federated Learning in Federated Contextual Bandits

  • Chengshuai Shi
  • Ruida Zhou
  • Kun Yang
  • Cong Shen

Federated learning (FL) has demonstrated great potential in revolutionizing distributed machine learning, and tremendous efforts have been made to extend it beyond the original focus on supervised learning. Among many directions, federated contextual bandits (FCB), a pivotal integration of FL and sequential decision-making, has garnered significant attention in recent years. Despite substantial progress, existing FCB approaches have largely employed their tailored FL components, often deviating from the canonical FL framework. Consequently, even renowned algorithms like FedAvg remain under-utilized in FCB, let alone other FL advancements. Motivated by this disconnection, this work takes one step towards building a tighter relationship between the canonical FL study and the investigations on FCB. In particular, a novel FCB design, termed FedIGW, is proposed to leverage a regression-based CB algorithm, i.e., inverse gap weighting. Compared with existing FCB approaches, the proposed FedIGW design can better harness the entire spectrum of FL innovations, which is concretely reflected as (1) flexible incorporation of (both existing and forthcoming) FL protocols; (2) modularized plug-in of FL analyses in performance guarantees; (3) seamless integration of FL appendages (such as personalization, robustness, and privacy). We substantiate these claims through rigorous theoretical analyses and empirical evaluations.

NeurIPS Conference 2024 Conference Paper

Transformers as Game Players: Provable In-context Game-playing Capabilities of Pre-trained Models

  • Chengshuai Shi
  • Kun Yang
  • Jing Yang
  • Cong Shen

The in-context learning (ICL) capability of pre-trained models based on the transformer architecture has received growing interest in recent years. While theoretical understanding has been obtained for ICL in reinforcement learning (RL), the previous results are largely confined to the single-agent setting. This work proposes to further explore the in-context learning capabilities of pre-trained transformer models in competitive multi-agent games, i. e. , in-context game-playing (ICGP). Focusing on the classical two-player zero-sum games, theoretical guarantees are provided to demonstrate that pre-trained transformers can provably learn to approximate Nash equilibrium in an in-context manner for both decentralized and centralized learning settings. As a key part of the proof, constructional results are established to demonstrate that the transformer architecture is sufficiently rich to realize celebrated multi-agent game-playing algorithms, in particular, decentralized V-learning and centralized VI-ULCB.

NeurIPS Conference 2023 Conference Paper

How2comm: Communication-Efficient and Collaboration-Pragmatic Multi-Agent Perception

  • Dingkang Yang
  • Kun Yang
  • Yuzheng Wang
  • Jing Liu
  • Zhi Xu
  • Rongbin Yin
  • Peng Zhai
  • Lihua Zhang

Multi-agent collaborative perception has recently received widespread attention as an emerging application in driving scenarios. Despite the advancements in previous efforts, challenges remain due to various noises in the perception procedure, including communication redundancy, transmission delay, and collaboration heterogeneity. To tackle these issues, we propose \textit{How2comm}, a collaborative perception framework that seeks a trade-off between perception performance and communication bandwidth. Our novelties lie in three aspects. First, we devise a mutual information-aware communication mechanism to maximally sustain the informative features shared by collaborators. The spatial-channel filtering is adopted to perform effective feature sparsification for efficient communication. Second, we present a flow-guided delay compensation strategy to predict future characteristics from collaborators and eliminate feature misalignment due to temporal asynchrony. Ultimately, a pragmatic collaboration transformer is introduced to integrate holistic spatial semantics and temporal context clues among agents. Our framework is thoroughly evaluated on several LiDAR-based collaborative detection datasets in real-world and simulated scenarios. Comprehensive experiments demonstrate the superiority of How2comm and the effectiveness of all its vital components. The code will be released at https: //github. com/ydk122024/How2comm.

EAAI Journal 2023 Journal Article

Interpretable knowledge-guided framework for modeling minimum miscible pressure of CO2-oil system in CO2-EOR projects

  • Bin Shen
  • Shenglai Yang
  • Xinyuan Gao
  • Shuai Li
  • Kun Yang
  • Jiangtao Hu
  • Hao Chen

Carbon dioxide enhanced oil recovery (CO2-EOR) is a promising application for carbon capture, utilization and storage (CCUS). Accurate modeling of CO2-oil minimum miscible pressure (MMP) is crucial for CO2-EOR projects. In this study, a knowledge-guided framework for an extreme gradient boosting machine (XGBoost) and interpretable tabular learning architecture (TabNet), called KXGB and KTabNet, respectively, are developed to model the MMP. The proposed models are strengthened using a large MMP database of 421 samples collected from literature. Domain knowledge is integrated into intelligent models to prevent data-driven models from producing predictions that violated the domain knowledge. The Shapley Additive Explanations (SHAP) method is used to explain the proposed model to ensure the credibility of petroleum engineers. To further verify the model’s effectiveness, the same experimental strategy is employed to compare the proposed models with existing machine-learning (ML) methods. The results show that KXGB is the most recommended superior solution for modeling MMP owing to its outstanding performance and simplicity of optimization. The correlation coefficient, root mean square error, and mean absolute error are 0. 9833, 0. 7637 and 0. 55, respectively. However, KTabNet has great potential. Although its accuracy is slightly lower than that of the former, its strong representation ability and decision transparency may be favorable for future research. This study also demonstrate that the proposed framework conforms to certain theoretical rules and has a reasonable domain of applicability. To the best of our knowledge, this is the first study on integration of domain knowledge into MMP modeling methods. The experience and insights obtained from this study can guide CO2-EOR projects and other tabular data modeling in the oil and gas industry.

AIIM Journal 2023 Journal Article

Osteoporosis prediction in lumbar spine X-ray images using the multi-scale weighted fusion contextual transformer network

  • Linyan Xue
  • Geng Qin
  • Shilong Chang
  • Cheng Luo
  • Ya Hou
  • Zhiyin Xia
  • Jiacheng Yuan
  • Yucheng Wang

Osteoporosis is a bone-related disease characterized by decreased bone density and mass, leading to brittle fractures. Osteoporosis assessment from radiographs using a deep learning algorithm has proven a low-cost alternative to the golden standard DXA. Due to the considerable noise and low contrast, automated diagnosis of osteoporosis in X-ray images still poses a significant challenge for traditional diagnostic methods. In this paper, an end-to-end transformer-style network was proposed, termed FCoTNet, to overcome the shortcoming of insufficient fusion of texture information and local features in the traditional CoTNet. To extract complementary geometric representations at each scale of the transformer module, we integrated parallel multi-scale feature extraction architectures in each unit layer of FCoTNet to utilize convolution to aggregate features from different receptive fields. Moreover, in order to extract small-scale texture features which were more critical to the diagnosis of osteoporosis in radiographs, larger fusion weights were assigned to the feature maps with small-size receptive fields. Afterward, the multi-scale global modeling was conducted by self-attention mechanism. The proposed model was first investigated on a private lumbar spine X-ray dataset with the 5-fold cross-validation strategy, obtaining an average accuracy of 78. 29 ± 0. 93 %, an average sensitivity of 69. 72 ± 2. 35 %, and an average specificity of 88. 92 ± 0. 67 % for the multi-classification of normal, osteopenia, and osteoporosis categories. We then conducted a controlled trial with five orthopedic clinicians to evaluate the clinical value of the model. The average clinician's accuracy improved from 61. 50 ± 10. 79 % unaided to 80. 00 ± 5. 92 % aided (18. 50 % improvement), sensitivity improved from 64. 38 ± 8. 07 % unaided to 83. 31 ± 5. 43 % aided (18. 93 % improvement), and specificity improved from 80. 11 ± 4. 72 % unaided to 89. 94 ± 3. 82 % aided (9. 83 % improvement). Meanwhile, the prediction consistency among clinicians significantly improved with the assistance of FCoTNet. Furthermore, the proposed model showed good robustness on an external test dataset. These investigations indicate that the proposed deep learning model achieves state-of-the-art performance for osteoporosis prediction, which substantially improves osteoporosis screening and reduced osteoporosis fractures.

TMLR Journal 2022 Journal Article

An Efficient One-Class SVM for Novelty Detection in IoT

  • Kun Yang
  • Samory Kpotufe
  • Nick Feamster

One-Class Support Vector Machines (OCSVM) are a common approach for novelty detection, due to their flexibility in fitting complex nonlinear boundaries between {normal} and {novel} data. Novelty detection is important in the Internet of Things (``IoT'') due to the threats these devices can present, and OCSVM often performs well in these environments due to the variety of devices, traffic patterns, and anomalies that IoT devices present. Unfortunately, conventional OCSVMs can introduce prohibitive memory and computational overhead at detection time. This work designs, implements and evaluates an efficient OCSVM for such practical settings. We extend Nystr\"om and (Gaussian) Sketching approaches to OCSVM, combining these methods with clustering and Gaussian mixture models to achieve 15-30x speedup in prediction time and 30-40x reduction in memory requirements without sacrificing detection accuracy. Here, the very nature of IoT devices is crucial: they tend to admit few modes of \emph{normal} operation, allowing for efficient pattern compression.

ICRA Conference 2020 Conference Paper

An Autonomous Intercept Drone with Image-based Visual Servo

  • Kun Yang
  • Quan Quan

For most people on the ground, facing an unwanted drone buzzing around overhead, there is not a lot that we can do, especially if it is out of gun (radio wave gun or shotgun) range. A solution to this is to use intercept drones that seek out and bring down other drones. In order to make the interception autonomous, an image-based visual servo algorithm is designed with a forward-looking monocular camera. The control command, namely the angular velocity and thrust, is generated for intercept drones to implement accurate and fast interception. The proposed method is demonstrated in both hardware-in-the-loop simulation and demonstrative flight experiments.

YNIMG Journal 2020 Journal Article

Increased power by harmonizing structural MRI site differences with the ComBat batch adjustment method in ENIGMA

  • Joaquim Radua
  • Eduard Vieta
  • Russell Shinohara
  • Peter Kochunov
  • Yann Quidé
  • Melissa J. Green
  • Cynthia S. Weickert
  • Thomas Weickert

A common limitation of neuroimaging studies is their small sample sizes. To overcome this hurdle, the Enhancing Neuro Imaging Genetics through Meta-Analysis (ENIGMA) Consortium combines neuroimaging data from many institutions worldwide. However, this introduces heterogeneity due to different scanning devices and sequences. ENIGMA projects commonly address this heterogeneity with random-effects meta-analysis or mixed-effects mega-analysis. Here we tested whether the batch adjustment method, ComBat, can further reduce site-related heterogeneity and thus increase statistical power. We conducted random-effects meta-analyses, mixed-effects mega-analyses and ComBat mega-analyses to compare cortical thickness, surface area and subcortical volumes between 2897 individuals with a diagnosis of schizophrenia and 3141 healthy controls from 33 sites. Specifically, we compared the imaging data between individuals with schizophrenia and healthy controls, covarying for age and sex. The use of ComBat substantially increased the statistical significance of the findings as compared to random-effects meta-analyses. The findings were more similar when comparing ComBat with mixed-effects mega-analysis, although ComBat still slightly increased the statistical significance. ComBat also showed increased statistical power when we repeated the analyses with fewer sites. Results were nearly identical when we applied the ComBat harmonization separately for cortical thickness, cortical surface area and subcortical volumes. Therefore, we recommend applying the ComBat function to attenuate potential effects of site in ENIGMA projects and other multi-site structural imaging work. We provide easy-to-use functions in R that work even if imaging data are partially missing in some brain regions, and they can be trained with one data set and then applied to another (a requirement for some analyses such as machine learning).

NeurIPS Conference 2016 Conference Paper

Density Estimation via Discrepancy Based Adaptive Sequential Partition

  • Dangna Li
  • Kun Yang
  • Wing Hung Wong

Given $iid$ observations from an unknown continuous distribution defined on some domain $\Omega$, we propose a nonparametric method to learn a piecewise constant function to approximate the underlying probability density function. Our density estimate is a piecewise constant function defined on a binary partition of $\Omega$. The key ingredient of the algorithm is to use discrepancy, a concept originates from Quasi Monte Carlo analysis, to control the partition process. The resulting algorithm is simple, efficient, and has provable convergence rate. We demonstrate empirically its efficiency as a density estimation method. We also show how it can be utilized to find good initializations for k-means.

EAAI Journal 2012 Journal Article

Freely-drawn sketches interpretation using SVMs-chain modeling

  • Kun Yang
  • Zhijun Li
  • Jingwei Ye

The growing popularity of tablet PCs and intelligent pen-centric computing has increased the importance of freehand sketch recognition algorithms. In this paper, the proposed method integrates the temporal, spatial and geometric constraint information to improve the recognition accuracy. To interpret the sketch as an incremental process, the paper investigates the use of the information fusion technique with Support Vector Machines (SVMs) chain for modeling and understanding the spatial and temporal information of sketch sequences. Online sketch recognition is achieved through the use of the SVMs-chain for systematically modeling the dynamic and stochastic behaviors of the sketch. To validate its efficiency, the experimental results in various domains and the comparison with traditional Hidden Markov Models have been presented.

v2026.09.13