Arrow Research search

Author name cluster

Ke Yang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

15 papers
1 author row

Possible papers

15

EAAI Journal 2026 Journal Article

A domain knowledge and cognitive law driven approach to anti-vibration hammer defect detection

  • Hang Niu
  • Xinyu Ge
  • Xiaoyu Zhao
  • Ke Yang
  • Qianming Wang
  • Yongjie Zhai
  • Zhedong Hu

The intelligent detection of anti-vibration hammer defects in transmission lines via computer vision is confronted with challenges due to the limited number of defect samples and the high similarity between defect classes. To this end, a domain knowledge and cognitive law driven approach to anti-vibration hammer defect detection is proposed, which integrates a Structural Knowledge and Geometric Feature-driven image generation method (SKGF) with a Cognitive Law-guided Multilevel Progressive target Detection framework (CLMP-Det). The imposition of morphological and tilt angle constraints is incorporated into the SKGF, based on prior knowledge of the anti-vibration hammer’s structure and its tilt angle distribution characteristics. These constraints can guide the generation of artificial anti-vibration hammer samples semantically consistent with the real physical structure and solve the problem of insufficient defective samples. Secondly, CLMP-Det is designed to simulate the human visual cognitive law through a progressive strategy, progressing from ease to difficulty. This strategy includes two sequential phases: preliminary perception and in-depth discrimination, which enhance the model’s capacity to distinguish between the challenging normal and tilt defect categories. The results of the experiment demonstrate that the proposed method significantly improves the overall detection performance of several widely-used detectors. Compared to the baseline model, our approach achieves a 7. 1% improvement in mean average precision. Thus, the method’s robust generalization capability and potential for engineering applications are fully validated.

NeurIPS Conference 2025 Conference Paper

FRBNet: Revisiting Low-Light Vision through Frequency-Domain Radial Basis Network

  • Fangtong Sun
  • Congyu Li
  • Ke Yang
  • Yuchen Pan
  • Hanwen Yu
  • Xichuan Zhang
  • Yiying Li

Low-light vision remains a fundamental challenge in computer vision due to severe illumination degradation, which significantly affects the performance of downstream tasks such as detection and segmentation. While recent state-of-the-art methods have improved performance through invariant feature learning modules, they still fall short due to incomplete modeling of low-light conditions. Therefore, we revisit low-light image formation and extend the classical Lambertian model to better characterize low-light conditions. By shifting our analysis to the frequency domain, we theoretically prove that the frequency-domain channel ratio can be leveraged to extract illumination-invariant features via a structured filtering process. We then propose a novel and end-to-end trainable module named \textbf{F}requency-domain \textbf{R}adial \textbf{B}asis \textbf{Net}work (\textbf{FRBNet}), which integrates the frequency-domain channel ratio operation with a learnable frequency domain filter for the overall illumination-invariant feature enhancement. As a plug-and-play module, FRBNet can be integrated into existing networks for low-light downstream tasks without modifying loss functions. Extensive experiments across various downstream tasks demonstrate that FRBNet achieves superior performance, including +2. 2 mAP for dark object detection and +2. 9 mIoU for nighttime segmentation. Code is available at: \url{https: //github. com/Sing-Forevet/FRBNet}.

YNIMG Journal 2025 Journal Article

Neural signatures of acute stress on the intention and outcome in third-party punishment: Evidence from univariate and multivariate analysis

  • Jingjing Chang
  • Di Song
  • Ke Yang
  • Rongjun Yu

Third-party punishment, a crucial element of prosocial behavior, involves individuals penalizing wrongdoers who harm the interests of others, even when their own interests are unaffected. Considering that third-party punishment behavior frequently arises in acute stress situations, understanding how stress influences such behavior is important. By using a modified economic game paradigm, this study investigates the impact of acute stress (induced through the Trier Social Stress Test) on the intention and outcome factors in third-party punishment, encompassing both behavioral and neural responses. Moreover, in addition to the conventional univariate activation analysis utilized in previous research, we also implemented multivariate pattern analysis (MVPA). On a behavioral level, participants displayed an increased inclination to allocate more tokens for punishing the dictator in scenarios involving unfair intentions or outcomes, and acute stress heightened the participants' sensitivity to the fairness of both intention and outcome. At the neural level, both univariate and multivariate analyses highlighted the crucial role of Theory of Mind (ToM)-related brain regions and the dACC in processing information related to intention and outcome. The MVPA further revealed distinctive neural activation patterns influenced by acute stress, particularly in the processing of intention. Specifically, brain regions within the ToM-related network showed an enhanced ability to differentiate between fair and unfair intentions in the stress group. Our findings suggest that stress has the potential to sensitize individuals to moral awareness during interpersonal interactions by facilitating perspective-taking and intentional attribution.

EAAI Journal 2025 Journal Article

One for all: Geometric structural-guided single domain generalization for vibration damper defect detection

  • Ke Yang
  • Zhedong Hu
  • Xiaoyu Zhao
  • Qianming Wang
  • Dongyang Hu
  • Yongjie Zhai

Improving the generalization capabilities of deep learning-based vibration damper defect detection models is crucial for their applicability in complex power scenarios. Existing models struggle to cope with scenario data not encountered during training, including diverse geographical regions, weather conditions, and lighting scenarios. To address these challenges, this paper introduces a novel single domain generalization(DG) method for vibration damper defect detection, aimed at enhancing the generalization capabilities of current models. Under analyzing the structural and morphological characteristics of vibration damper defects, this study selects their geometric structural features as domain invariant features(DIF), guiding model training with artificial samples generated via three dimensional(3D) modeling. To ensure comprehensive learning of the DIF of vibration damper defects, this study employs multi-scale contrastive learning(MCL) loss and multi-morphology similarity metric learning(MSM) loss during the feature extraction and candidate box processing stages, respectively. These approaches guide the model in extracting geometric features and achieving consistent expression of defect characteristics. The experimental results demonstrate that the proposed method outperforms advanced object detection models in DG for vibration damper defect detection, achieving an average DG accuracy of 68. 6% across multiple target domains and a sensitivity of only 0. 219, thereby realizing the “one-time training, everywhere application” concept in vibration damper defect detection models.

NeurIPS Conference 2024 Conference Paper

Bias and Volatility: A Statistical Framework for Evaluating Large Language Model's Stereotypes and the Associated Generation Inconsistency

  • Yiran Liu
  • Ke Yang
  • Zehan Qi
  • Xiao Liu
  • Yang Yu
  • ChengXiang Zhai

We present a novel statistical framework for analyzing stereotypes in large language models (LLMs) by systematically estimating the bias and variation in their generation. Current evaluation metrics in the alignment literature often overlook the randomness of stereotypes caused by the inconsistent generative behavior of LLMs. For example, this inconsistency can result in LLMs displaying contradictory stereotypes, including those related to gender or race, for identical professions across varied contexts. Neglecting such inconsistency could lead to misleading conclusions in alignment evaluations and hinder the accurate assessment of the risk of LLM applications perpetuating or amplifying social stereotypes and unfairness. This work proposes a Bias-Volatility Framework (BVF) that estimates the probability distribution function of LLM stereotypes. Specifically, since the stereotype distribution fully captures an LLM's generation variation, BVF enables the assessment of both the likelihood and extent to which its outputs are against vulnerable groups, thereby allowing for the quantification of the LLM's aggregated discrimination risk. Furthermore, we introduce a mathematical framework to decompose an LLM’s aggregated discrimination risk into two components: bias risk and volatility risk, originating from the mean and variation of LLM’s stereotype distribution, respectively. We apply BVF to assess 12 commonly adopted LLMs and compare their risk levels. Our findings reveal that: i) Bias risk is the primary cause of discrimination risk in LLMs; ii) Most LLMs exhibit significant pro-male stereotypes for nearly all careers; iii) Alignment with reinforcement learning from human feedback lowers discrimination by reducing bias, but increases volatility; iv) Discrimination risk in LLMs correlates with key sociol-economic factors like professional salaries. Finally, we emphasize that BVF can also be used to assess other dimensions of generation inconsistency's impact on LLM behavior beyond stereotypes, such as knowledge mastery.

EAAI Journal 2024 Journal Article

Buffer-text: Detecting arbitrary shaped text in natural scene image

  • Ke Yang
  • Jizheng Yi
  • Aibin Chen
  • Ze Jin

Scene text detection has always been a research hotspot in computer vision and image understanding. With the development of deep learning, segmentation-based methods have achieved an exceptional effect in regular or curved text detection, but they cannot separate adjacent word-level texts. In this paper, we proposed a detector called Buffer-Text for the detection of irregular text in the natural scene image. First, the buffer region is proposed for bending text detection which widens the spatial distance between word-level texts. Then, a centerline-based polygon expansion algorithm is developed for the acquisition of the buffer region. After that, the scene text image is divided into different regions which are predicted by adopting the idea of multiclass semantic segmentation. To obtain effective segmentation results and solve the category imbalance problem, a Fully Convolutional Networks (FCN) with Spatial and Channel Squeeze & Excitation Block module is designed, and a loss function with adaptive weight updating is defined for the network. Ultimately, the post-processing including the total erosion and the single expansion is applied to eliminate the areas of noise in the segmented image and to separate the weak junctions in the word-level text. To verify the validity of the proposed method, several experiments were conducted on two curved text datasets, namely Total-Text and CTW1500, and the results indicated that the proposed method achieved significant accuracy in three statistical indicators (precision, recall, and F-score), particularly for the images with natural scenes and various text shapes.

AAAI Conference 2023 Conference Paper

ADEPT: A DEbiasing PrompT Framework

  • Ke Yang
  • Charles Yu
  • Yi R. Fung
  • Manling Li
  • Heng Ji

Several works have proven that finetuning is an applicable approach for debiasing contextualized word embeddings. Similarly, discrete prompts with semantic meanings have shown to be effective in debiasing tasks. With unfixed mathematical representation at the token level, continuous prompts usually surpass discrete ones at providing a pre-trained language model (PLM) with additional task-specific information. Despite this, relatively few efforts have been made to debias PLMs by prompt tuning with continuous prompts compared to its discrete counterpart. Furthermore, for most debiasing methods that alter a PLM's original parameters, a major problem is the need to not only decrease the bias in the PLM but also to ensure that the PLM does not lose its representation ability. Finetuning methods typically have a hard time maintaining this balance, as they tend to violently remove meanings of attribute words (like the words developing our concepts of "male" and "female" for gender), which also leads to an unstable and unpredictable training process. In this paper, we propose ADEPT, a method to debias PLMs using prompt tuning while maintaining the delicate balance between removing biases and ensuring representation ability. To achieve this, we propose a new training criterion inspired by manifold learning and equip it with an explicit debiasing term to optimize prompt tuning. In addition, we conduct several experiments with regard to the reliability, quality, and quantity of a previously proposed attribute training corpus in order to obtain a clearer prototype of a certain attribute, which indicates the attribute's position and relative distances to other words on the manifold. We evaluate ADEPT on several widely acknowledged debiasing benchmarks and downstream tasks, and find that it achieves competitive results while maintaining (and in some cases even improving) the PLM's representation ability. We further visualize words' correlation before and after debiasing a PLM, and give some possible explanations for the visible effects.

EAAI Journal 2022 Journal Article

ConvPatchTrans: A script identification network with global and local semantics deeply integrated

  • Ke Yang
  • Jizheng Yi
  • Aibin Chen
  • Jiaqi Liu
  • Wenjie Chen
  • Ze Jin

Optical Character Recognition (OCR) system serves the need of reading text from images. Script identification that identifies the language of the text in the image is an important part of OCR technology and an indispensable role in the stability and accuracy of the OCR system. The most challenging for script identification is the interference caused by similarities between texts in different languages. In this paper, a two-branch network named ConvPatchTrans is designed to process global and local semantic features separately, focusing on the text and each word in a picture. The ConvPatchTrans extracts feature from different stages of the Visual Geometry Group network (VGGNet) as global and local semantics. For the global branch, the linear classifier is recommended. For the local branch, text image data is converted to image sequence data. Then, multi-layers convolution-enhanced Transformer (MCET) is proposed to bring about the deep fusion of sequence. Finally, the global and local branches are fused by an adaptive weighted fusion method to get the best result. In order to verify the effectiveness of our proposed method, four public script identification datasets are used for comparative experiments. Our method has obtained the highest values among currently published methods on the CVSI2015 and MLE2E datasets, which are 98. 90% and 97. 50%, respectively. At the same time, satisfactory results are also obtained on the other two datasets.

EAAI Journal 2022 Journal Article

Geometric characteristic learning R-CNN for shockproof hammer defect detection

  • Yongjie Zhai
  • Ke Yang
  • Zhenyuan Zhao
  • Qianming Wang
  • Kang Bai

Shockproof hammers are key components of power transmission lines. Aiming at the problems regarding the same kind of defect having various manifestations, different defects having similarities, and defect data being difficult to obtain, a geometric characteristic learning region-based convolutional neural network (GCL R-CNN) is proposed to detect shockproof hammer defects in aerial images of transmission lines. First, a GCL module is proposed for the first time and introduced in the Faster R-CNN, and artificial samples are generated via 3D modeling. Second, to make the model pay more attention to the geometric characteristics of the observed shockproof hammer, artificial samples with monochromatic backgrounds are used to guide the neural network training process. In this way, the model can better learn the salient features of the shockproof hammer defects and the model’s ability to distinguish normal samples from defective samples is enhanced. Finally, in the case of few samples of shockproof hammer defects, artificial samples with real backgrounds are used to expand the training set and improve the accuracy of shockproof hammer defect detection. Experimental results on real aerial images of transmission lines show that the proposed model can accurately detect normal shockproof hammers, missing shockproof hammer heads and tilted shockproof hammers, with detection accuracies of 93. 8%, 89. 94% and 66. 22%, respectively. The above results show that the proposed model can realize the detection of two types of defects in the shockproof hammer, that is, the fault diagnosis of the shockproof hammer.

IJCAI Conference 2019 Conference Paper

Balanced Ranking with Diversity Constraints

  • Ke Yang
  • Vasilis Gkatzelis
  • Julia Stoyanovich

Many set selection and ranking algorithms have recently been enhanced with diversity constraints that aim to explicitly increase representation of historically disadvantaged populations, or to improve the over-all representativeness of the selected set. An unintended consequence of these constraints, however, is reduced in-group fairness: the selected candidates from a given group may not be the best ones, and this unfairness may not be well-balanced across groups. In this paper we study this phenomenon using datasets that comprise multiple sensitive attributes. We then introduce additional constraints, aimed at balancing the in-group fairness across groups, and formalize the induced optimization problems as integer linear programs. Using these programs, we conduct an experimental evaluation with real datasets, and quantify the feasible trade-offs between balance and overall performance in the presence of diversity constraints.

AAAI Conference 2018 Conference Paper

Exploring Temporal Preservation Networks for Precise Temporal Action Localization

  • Ke Yang
  • Peng Qiao
  • Dongsheng Li
  • Shaohe Lv
  • Yong Dou

Temporal action localization is an important task of computer vision. Though a variety of methods have been proposed, it still remains an open question how to predict the temporal boundaries of action segments precisely. Most works use segment-level classifiers to select video segments pre-determined by action proposal or dense sliding windows. However, in order to achieve more precise action boundaries, a temporal localization system should make dense predictions at a fine granularity. A newly proposed work exploits Convolutional-Deconvolutional-Convolutional (CDC) filters to upsample the predictions of 3D ConvNets, making it possible to perform per-frame action predictions and achieving promising performance in terms of temporal action localization. However, CDC network loses temporal information partially due to the temporal downsampling operation. In this paper, we propose an elegant and powerful Temporal Preservation Convolutional (TPC) Network that equips 3D ConvNets with TPC filters. TPC network can fully preserve temporal resolution and downsample the spatial resolution simultaneously, enabling frame-level granularity action localization with minimal loss of time information. TPC network can be trained in an end-to-end manner. Experiment results on public datasets show that TPC network achieves significant improvement in both per-frame action prediction and segment-level temporal action localization.

NeurIPS Conference 2012 Conference Paper

Large Scale Distributed Deep Networks

  • Jeffrey Dean
  • Greg Corrado
  • Rajat Monga
  • Kai Chen
  • Matthieu Devin
  • Mark Mao
  • Marc'Aurelio Ranzato
  • Andrew Senior

Recent work in unsupervised feature learning and deep learning has shown that being able to train large models can dramatically improve performance. In this paper, we consider the problem of training a deep network with billions of parameters using tens of thousands of CPU cores. We have developed a software framework called DistBelief that can utilize computing clusters with thousands of machines to train large models. Within this framework, we have developed two algorithms for large-scale distributed training: (i) Downpour SGD, an asynchronous stochastic gradient descent procedure supporting a large number of model replicas, and (ii) Sandblaster, a framework that supports for a variety of distributed batch optimization procedures, including a distributed implementation of L-BFGS. Downpour SGD and Sandblaster L-BFGS both increase the scale and speed of deep network training. We have successfully used our system to train a deep network 100x larger than previously reported in the literature, and achieves state-of-the-art performance on ImageNet, a visual object recognition task with 16 million images and 21k categories. We show that these same techniques dramatically accelerate the training of a more modestly sized deep network for a commercial speech recognition service. Although we focus on and report performance of these methods as applied to training large neural networks, the underlying algorithms are applicable to any gradient-based machine learning algorithm.

TCS Journal 2007 Journal Article

On the (im)possibility of non-interactive correlation distillation

  • Ke Yang

We study the problem of non-interactive correlation distillation (NICD). Suppose that Alice and Bob each have a string, denoted by A = a 0 a 1 ⋯ a n − 1 and B = b 0 b 1 ⋯ b n − 1, respectively. Furthermore, for every k = 0, 1, …, n − 1, ( a k, b k ) is drawn independently from a distribution N, known as the ‘noise model’. Alice and Bob wish to ‘distill’ the correlation non-interactively, i. e. , they wish to each apply a function to their strings, and output one random bit, denoted by X and Y, such that Pr [ X = Y ] can be made as close to 1 as possible. The problem is, for what noise models can they succeed? This problem is related to various topics in computer science, including information reconciliation and random beacons. In fact, if NICD is indeed possible for some general class of noise models, then some of these topics would, in some sense, become straightforward corollaries. We prove two negative results on NICD for various noise models. We prove that, for these models, it is impossible to distill the correlation to be arbitrarily close to 1. We also give an example where Alice and Bob can increase their correlation with one bit of communication (in this case they need to each output two bits). This example, which may be of interest on its own, demonstrates that even the smallest amount of communication is provably more powerful than no communication.

NeurIPS Conference 2004 Conference Paper

An Investigation of Practical Approximate Nearest Neighbor Algorithms

  • Ting Liu
  • Andrew Moore
  • Ke Yang
  • Alexander Gray

This paper concerns approximate nearest neighbor searching algorithms, which have become increasingly important, especially in high dimen- sional perception areas such as computer vision, with dozens of publica- tions in recent years. Much of this enthusiasm is due to a successful new approximate nearest neighbor approach called Locality Sensitive Hash- ing (LSH). In this paper we ask the question: can earlier spatial data structure approaches to exact nearest neighbor, such as metric trees, be altered to provide approximate answers to proximity queries and if so, how? We introduce a new kind of metric tree that allows overlap: certain datapoints may appear in both the children of a parent. We also intro- duce new approximate k-NN search algorithms on this structure. We show why these structures should be able to exploit the same random- projection-based approximations that LSH enjoys, but with a simpler al- gorithm and perhaps with greater efficiency. We then provide a detailed empirical evaluation on five large, high dimensional datasets which show up to 31-fold accelerations over LSH. This result holds true throughout the spectrum of approximation levels.

v2026.09.13