Arrow Research search

Author name cluster

Kai Zheng

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

28 papers
2 author rows

Possible papers

28

EAAI Journal 2026 Journal Article

A frequency-guided denoising framework based on convolutional transformer for electrocardiogram signals

  • Mingyue Cui
  • Yewei Gan
  • Jiepeng Chen
  • Kai Zheng
  • Yanchong Xie
  • Daosong Hu
  • Yuning Cui
  • Kai Huang

As a non-invasive diagnostic tool, the electrocardiogram (ECG) is easily affected by various noises, which poses difficulties in diagnosing heart diseases accurately. However, traditional denoising methods filter out specific frequencies through time-frequency analysis and are limited by threshold setup, while existing deep-learning methods fail to fully exploit the frequency characteristics of ECG signals. To address this problem, we introduce a frequency-guided framework based on convolutional transformer for ECG denoising, called FGCT. We innovatively combine traditional filtering/decomposition techniques and attention-based networks, and propose a frequency-guided multi-head self-attention (FG-MSA) model and a global channel and spatial enhanced convolution (GCSC) network. For the FG-MSA model, we embed frequency domain priors directly to guide the time domain attention process for extracting intra-band dependencies. For the GCSC network, we employ a global channel-spatial attention to capture inter-band dependencies, distinguish signal from noise to reduce the spectrum overlap noise. Besides, to further enhance the correlation between feature maps across different channels, we use the convolutional layer to perform the down-sampling operation instead of a regular pooling layer. We comprehensively compare our FGCT with the state-of-the-art methods, including the traditional rule-based and learning-based methods. Experimental results demonstrate that our method outperforms these baselines on two widely used ECG benchmarks under four representative noise types (baseline wander, electrode motion, muscle artifact, and their mixture).

EAAI Journal 2025 Journal Article

A new deep sparse unfolding network with variational Bayesian enhancement for high-resolution localization of acoustic sources

  • Siguo Wen
  • Kai Zheng
  • Pu Li
  • Yuying Fang
  • Yin Bai
  • Dewei Yang

The Deep sparse unfolding network has been proven to be interpretable and efficient, which has broad prospects for high-resolution acoustic beamforming. However, the performance of the network model depends on the accurate establishment of high-quality datasets and constraint equations. Therefore, the performance and stability of the network model may decline under the presence of complex source excitation and uncertain interference conditions. To address these issues, we propose a high-resolution acoustic beamforming sparse unfolding network enhanced by variational Bayesian. Firstly, we introduce an adaptive proxy model to generate a high-quality target map unaffected by frequency influence and construct a complete dataset. Then, we establish a deconvolution optimization problem and threshold iteration solution method. On this basis, we construct the Learned Iterative Shrinkage Thresholding Algorithm with Nesterov to enhance solution rates and unroll it into a network model. Subsequently, we propose a variational Bayesian enhancement mechanism to enhance the sparsity of the iterative solution results, improve the diversity of model training data, and characterize imaging position deviation loss. By combining energy deviations, we construct a new loss function to ensure a balance between imaging accuracy and energy concentration. We have developed an interpretable high-resolution imaging network, which we refer to as the Leaning-Nesterov-Variational-Bayesian-Enhancements Network (LNVBENet). Finally, we conduct research and analysis on two available datasets and experimental data to assess the feasibility and superiority of LNVBENet. The results show that LNVBENet not only enhances imaging accuracy and stability but also offers faster imaging speeds and robust anti-interference capabilities.

NeurIPS Conference 2025 Conference Paper

Graph-Theoretic Insights into Bayesian Personalized Ranking for Recommendation

  • Kai Zheng
  • Jianxin Wang
  • Jinhui Xu

Graph self-supervised learning (GSL) is essential for processing graph-structured data, reducing the need for manual labeling. Traditionally, this paradigm has extensively utilized Bayesian Personalized Ranking (BPR) as its primary loss function. Despite its widespread application, the theoretical analysis of its node relations evaluation have remained largely unexplored. This paper employs recent advancements in latent hyperbolic geometry to deepen our understanding of node relationships from a graph-theoretical perspective. We analyze BPR’s limitations, particularly its reliance on local connectivity through 2-hop paths, which overlooks global connectivity and the broader topological structure. To address these shortcomings, we purpose a novel loss function, BPR+, designed to encompass even-hop paths and better capture global connectivity and topological nuances. This approach facilitates a more detailed measurement of user-item relationships and improves the granularity of relationship assessments. We validate BPR+ through extensive empirical testing across five real-world datasets and demonstrate its efficacy in refining graph self-supervised learning frameworks. Additionally, we explore the application of BPR+ in drug repositioning, highlighting its potential to support pharmaceutical research and development. Our findings not only illuminate the success factors of previous methodologies but also offer new theoretical insights into this learning paradigm.

ICML Conference 2025 Conference Paper

NEAR: Neural Electromagnetic Array Response

  • Yinyan Bu
  • Jiajie Yu
  • Kai Zheng
  • Xinyu Zhang
  • Piya Pal

We address the challenge of achieving angular super-resolution in multi-antenna radar systems that are widely used for localization, navigation, and automotive perception. A multi-antenna radar achieves very high resolution by computationally creating a large virtual sensing system using very few physical antennas. However, practical constraints imposed by hardware, noise, and a limited number of antennas can impede its performance. Conventional supervised learning models that rely on extensive pre-training with large datasets, often exhibit poor generalization in unseen environments. To overcome these limitations, we propose NEAR, an untrained implicit neural representation (INR) framework that predicts radar responses at unseen locations from sparse measurements, by leveraging latent harmonic structures inherent in radar wave propagation. We establish new theoretical results linking antenna array response to expressive power of INR architectures, and develop a novel physics-informed and latent geometry-aware regularizer. Our approach integrates classical signal representation with modern implicit neural learning, enabling high-resolution radar sensing that is both interpretable and generalizable. Extensive simulations and real-world experiments using radar platforms demonstrate NEAR’s effectiveness and its ability to adapt to unseen environments.

AAAI Conference 2025 Conference Paper

Trigger3:Refining Query Correction via Adaptive Model Selector

  • Kepu Zhang
  • Zhongxiang Sun
  • Xiao Zhang
  • Xiaoxue Zang
  • Kai Zheng
  • Yang Song
  • Jun Xu

In search scenarios, user experience can be hindered by erroneous queries due to typos, voice errors, or knowledge gaps. Therefore, query correction is crucial for search engines. Current correction models, usually small models trained on specific data, often struggle with queries beyond their training scope or those requiring contextual understanding. While the advent of Large Language Models (LLMs) offers a potential solution, they are still limited by their pre-training data and inference cost, particularly for complex queries, making them not always effective for query correction. To tackle these, we propose Trigger3, a large-small model collaboration framework that integrates the traditional correction model and LLM for query correction, capable of adaptively choosing the appropriate correction method based on the query and the correction results from the traditional correction model and LLM. Trigger3 first employs a correction trigger to filter out correct queries. Incorrect queries are then corrected by the traditional correction model. If this fails, an LLM trigger is activated to call the LLM for correction. Finally, for queries that no model can correct, a fallback trigger decides to return the original query. Extensive experiments demonstrate Trigger3 outperforms correction baselines while maintaining efficiency.

SODA Conference 2023 Conference Paper

Approaching the Soundness Barrier: A Near Optimal Analysis of the Cube versus Cube Test

  • Dor Minzer
  • Kai Zheng

The Cube versus Cube test is a variant of the well-known Plane versus Plane test of Raz and Safra [10], in which to each 3-dimensional affine subspace C of 𝔽 n q, a polynomial of degree at most d, T ( C ), is assigned in a somewhat locally consistent manner: taking two cubes C 1, C 2 that intersect in a plane uniformly at random, the probability that T ( C 1 ) and T ( C 2 ) agree on C 1 ∩ C 2 is at least some ε. An element of interest is the soundness threshold of this test, i. e. the smallest value of ε, such that this amount of local consistency implies a global structure; namely, that there is a global degree d function g such that g| C = T (C) for at least Ω(ε) fraction of the cubes. We show that the cube versus cube low degree test has soundness poly( d )/ q. This result achieves the optimal dependence on q for soundness in low degree testing and improves upon previous soundness results of poly( d )/ q 1/2 due to Bhangale, Dinur and Navon [4].

AAAI Conference 2022 Conference Paper

Boosting Contrastive Learning with Relation Knowledge Distillation

  • Kai Zheng
  • Yuanjiang Wang
  • Ye Yuan

While self-supervised representation learning (SSL) has proved to be effective in the large model, there is still a huge gap between the SSL and supervised method in the lightweight model when following the same solution. We delve into this problem and find that the lightweight model is prone to collapse in semantic space when simply performing instance-wise contrast. To address this issue, we propose a relation-wise contrastive paradigm with Relation Knowledge Distillation (ReKD). We introduce a heterogeneous teacher to explicitly mine the semantic information and transferring a novel relation knowledge to the student (lightweight model). The theoretical analysis supports our main concern about instance-wise contrast and verify the effectiveness of our relation-wise contrastive learning. Extensive experimental results also demonstrate that our method achieves significant improvements on multiple lightweight models. Particularly, the linear evaluation on AlexNet obviously improves the current state-of-art from 44. 7% to 50. 1%, which is the first work to get close to the supervised (50. 5%). Code will be made available.

NeurIPS Conference 2022 Conference Paper

Cache-Augmented Inbatch Importance Resampling for Training Recommender Retriever

  • Jin Chen
  • Defu Lian
  • Yucheng Li
  • Baoyun Wang
  • Kai Zheng
  • Enhong Chen

Recommender retrievers aim to rapidly retrieve a fraction of items from the entire item corpus when a user query requests, with the representative two-tower model trained with the log softmax loss. For efficiently training recommender retrievers on modern hardwares, inbatch sampling, where the items in the mini-batch are shared as negatives to estimate the softmax function, has attained growing interest. However, existing inbatch sampling based strategies just correct the sampling bias of inbatch items with item frequency, being unable to distinguish the user queries within the mini-batch and still incurring significant bias from the softmax. In this paper, we propose a Cache-Augmented Inbatch Importance Resampling (XIR) for training recommender retrievers, which not only offers different negatives to user queries with inbatch items, but also adaptively achieves a more accurate estimation of the softmax distribution. Specifically, XIR resamples items from the given mini-batch training pairs based on certain probabilities, where a cache with more frequently sampled items is adopted to augment the candidate item set, with the purpose of reusing the historical informative samples. XIR enables to sample query-dependent negatives based on inbatch items and to capture dynamic changes of model training, which leads to a better approximation of the softmax and further contributes to better convergence. Finally, we conduct experiments to validate the superior performance of the proposed XIR compared with competitive approaches.

IJCAI Conference 2022 Conference Paper

MetaER-TTE: An Adaptive Meta-learning Model for En Route Travel Time Estimation

  • Yu Fan
  • Jiajie Xu
  • Rui Zhou
  • Jianxin Li
  • Kai Zheng
  • Lu Chen
  • Chengfei Liu

En route travel time estimation (ER-TTE) aims to predict the travel time on the remaining route. Since the traveled and remaining parts of a trip usually have some common characteristics like driving speed, it is desirable to explore these characteristics for improved performance via effective adaptation. This yet faces the severe problem of data sparsity due to the few sampled points in a traveled partial trajectory. Since trajectories with different contextual information tend to have different characteristics, the existing meta-learning method for ER-TTE cannot fit each trajectory well because it uses the same model for all trajectories. To this end, we propose a novel adaptive meta-learning model called MetaER-TTE. Particularly, we utilize soft-clustering and derive cluster-aware initialized parameters to better transfer the shared knowledge across trajectories with similar contextual information. In addition, we adopt a distribution-aware approach for adaptive learning rate optimization, so as to avoid task-overfitting which will occur when guiding the initial parameters with a fixed learning rate for tasks under imbalanced distribution. Finally, we conduct comprehensive experiments to demonstrate the superiority of MetaER-TTE.

UAI Conference 2021 Conference Paper

Combinatorial semi-bandit in the non-stationary environment

  • Wei Chen 0041
  • Liwei Wang
  • Haoyu Zhao
  • Kai Zheng

In this paper, we investigate the non-stationary combinatorial semi-bandit problem, both in the switching case and in the dynamic case. In the general case where (a) the reward function is non-linear, (b) arms may be probabilistically triggered, and (c) only approximate offline oracle exists (Wang and Chen, NIPS 2017), our algorithm achieves $\tilde{O}(m\sqrt{N T}/\Delta_{\min})$ distribution-dependent regret in the switching case, and $\tilde{O}({V}^{1/3}T^{2/3})$ distribution-independent regret in the dynamic case, where ${N}$ is the number of switchings and ${V}$ is the sum of the total “distribution changes”, $m$ is the total number of arms, and $\Delta_{\min}$ is a gap variable dependent on the distributions of arm outcomes. The regret bounds in both scenarios are nearly optimal, but our algorithm needs to know the parameter ${N}$ or ${V}$ in advance. We further show that by employing another technique, our algorithm no longer needs to know the parameters ${N}$ or ${V}$ but the regret bounds could become suboptimal. In a special case where the reward function is linear and we have an exact oracle, we apply a new technique to design a parameter-free algorithm that achieves nearly optimal regret both in the switching case and in the dynamic case without knowing the parameters in advance.

AAAI Conference 2021 Conference Paper

Efficient Optimal Selection for Composited Advertising Creatives with Tree Structure

  • Jin Chen
  • Tiezheng Ge
  • Gangwei Jiang
  • Zhiqiang Zhang
  • Defu Lian
  • Kai Zheng

Ad creatives are one of the prominent mediums for online e-commerce advertisements. Ad creatives with enjoyable visual appearance may increase the click-through rate (CTR) of products. Ad creatives are typically handcrafted by advertisers and then delivered to the advertising platforms for advertisement. In recent years, advertising platforms are capable of instantly compositing ad creatives with arbitrarily designated elements of each ingredient, so advertisers are only required to provide basic materials. While facilitating the advertisers, a great number of potential ad creatives can be composited, making it difficult to accurately estimate CTR for them given limited real-time feedback. To this end, we propose an Adaptive and Efficient ad creative Selection (AES) framework based on a tree structure. The tree structure on compositing ingredients enables dynamic programming for efficient ad creative selection on the basis of CTR. Due to limited feedback, the CTR estimator is usually of high variance. Exploration techniques based on Thompson sampling are widely used for reducing variances of the CTR estimator, alleviating feedback sparsity. Based on the tree structure, Thompson sampling is adapted with dynamic programming, leading to efficient exploration for potential ad creatives with the largest CTR. We finally evaluate the proposed algorithm on the synthetic dataset and the real-world dataset. The results show that our approach can outperform competing baselines in terms of convergence rate and overall CTR.

IJCAI Conference 2021 Conference Paper

MFNP: A Meta-optimized Model for Few-shot Next POI Recommendation

  • Huimin Sun
  • Jiajie Xu
  • Kai Zheng
  • Pengpeng Zhao
  • Pingfu Chao
  • Xiaofang Zhou

Next Point-of-Interest (POI) recommendation is of great value for location-based services. Existing solutions mainly rely on extensive observed data and are brittle to users with few interactions. Unfortunately, the problem of few-shot next POI recommendation has not been well studied yet. In this paper, we propose a novel meta-optimized model MFNP, which can rapidly adapt to users with few check-in records. Towards the cold-start problem, it seamlessly integrates carefully designed user-specific and region-specific tasks in meta-learning, such that region-aware user preferences can be captured via a rational fusion of region-independent personal preferences and region-dependent crowd preferences. In modelling region-dependent crowd preferences, a cluster-based adaptive network is adopted to capture shared preferences from similar users for knowledge transfer. Experimental results on two real-world datasets show that our model outperforms the state-of-the-art methods on next POI recommendation for cold-start users.

TIST Journal 2021 Journal Article

Predicting Human Mobility with Reinforcement-Learning-Based Long-Term Periodicity Modeling

  • Shuo Tao
  • Jingang Jiang
  • Defu Lian
  • Kai Zheng
  • Enhong Chen

Mobility prediction plays an important role in a wide range of location-based applications and services. However, there are three problems in the existing literature: (1) explicit high-order interactions of spatio-temporal features are not systemically modeled; (2) most existing algorithms place attention mechanisms on top of recurrent network, so they can not allow for full parallelism and are inferior to self-attention for capturing long-range dependence; (3) most literature does not make good use of long-term historical information and do not effectively model the long-term periodicity of users. To this end, we propose MoveNet and RLMoveNet. MoveNet is a self-attention-based sequential model, predicting each user’s next destination based on her most recent visits and historical trajectory. MoveNet first introduces a cross-based learning framework for modeling feature interactions. With self-attention on both the most recent visits and historical trajectory, MoveNet can use an attention mechanism to capture the user’s long-term regularity in a more efficient way. Based on MoveNet, to model long-term periodicity more effectively, we add the reinforcement learning layer and named RLMoveNet. RLMoveNet regards the human mobility prediction as a reinforcement learning problem, using the reinforcement learning layer as the regularization part to drive the model to pay attention to the behavior with periodic actions, which can help us make the algorithm more effective. We evaluate both of them with three real-world mobility datasets. MoveNet outperforms the state-of-the-art mobility predictor by around 10% in terms of accuracy, and simultaneously achieves faster convergence and over 4x training speedup. Moreover, RLMoveNet achieves higher prediction accuracy than MoveNet, which proves that modeling periodicity explicitly from the perspective of reinforcement learning is more effective.

IJCAI Conference 2020 Conference Paper

Bilinear Graph Neural Network with Neighbor Interactions

  • Hongmin Zhu
  • Fuli Feng
  • Xiangnan He
  • Xiang Wang
  • Yan Li
  • Kai Zheng
  • Yongdong Zhang

Graph Neural Network (GNN) is a powerful model to learn representations and make predictions on graph data. Existing efforts on GNN have largely defined the graph convolution as a weighted sum of the features of the connected nodes to form the representation of the target node. Nevertheless, the operation of weighted sum assumes the neighbor nodes are independent of each other, and ignores the possible interactions between them. When such interactions exist, such as the co-occurrence of two neighbor nodes is a strong signal of the target node's characteristics, existing GNN models may fail to capture the signal. In this work, we argue the importance of modeling the interactions between neighbor nodes in GNN. We propose a new graph convolution operator, which augments the weighted sum with pairwise interactions of the representations of neighbor nodes. We term this framework as Bilinear Graph Neural Network (BGNN), which improves GNN representation ability with bilinear interactions between neighbor nodes. In particular, we specify two BGNN models named BGCN and BGAT, based on the well-known GCN and GAT, respectively. Empirical results on three public benchmarks of semi-supervised node classification verify the effectiveness of BGNN --- BGCN (BGAT) outperforms GCN (GAT) by 1. 6% (1. 5%) in classification accuracy. Codes are available at: https: //github. com/zhuhm1996/bgnn.

IS Journal 2020 Journal Article

Collaborative Filtering With Ranking-Based Priors on Unknown Ratings

  • Jin Chen
  • Defu Lian
  • Kai Zheng

Advanced collaborative filtering methods based on explicit feedback assume that unknown ratings are missing not at random. The state-of-the-art algorithm hypothesizes that unknown items are weakly rated and sets an explicit prior to unknown ratings. However, the prior assuming unknown ratings be close to zero may be questionable and it is challenging to set appropriate prior ratings for unknown items. In this article, to avert the use of prior ratings, we propose a ranking-based prior by hypothesizing that each user's unknown ratings are close to each other. This prior essentially acts as a regularizer to penalize the discrepancy of predicted ratings between any two unknown items. With the ranking-based prior, we design a generic collaborative filtering framework for explicit feedback and develop an efficient optimization algorithm for parameter learning. We finally evaluate the proposed algorithms on four real-world rating datasets. The results show that the proposed algorithms consistently outperform the state-of-the-art baselines and that the ranking-based prior leads to superior recommendation accuracy.

IJCAI Conference 2020 Conference Paper

Discovering Subsequence Patterns for Next POI Recommendation

  • Kangzhi Zhao
  • Yong Zhang
  • Hongzhi Yin
  • Jin Wang
  • Kai Zheng
  • Xiaofang Zhou
  • Chunxiao Xing

Next Point-of-Interest (POI) recommendation plays an important role in location-based services. State-of-the-art methods learn the POI-level sequential patterns in the user's check-in sequence but ignore the subsequence patterns that often represent the socio-economic activities or coherence of preference of the users. However, it is challenging to integrate the semantic subsequences due to the difficulty to predefine the granularity of the complex but meaningful subsequences. In this paper, we propose Adaptive Sequence Partitioner with Power-law Attention (ASPPA) to automatically identify each semantic subsequence of POIs and discover their sequential patterns. Our model adopts a state-based stacked recurrent neural network to hierarchically learn the latent structures of the user's check-in sequence. We also design a power-law attention mechanism to integrate the domain knowledge in spatial and temporal contexts. Extensive experiments on two real-world datasets demonstrate the effectiveness of our model.

NeurIPS Conference 2020 Conference Paper

Locally Differentially Private (Contextual) Bandits Learning

  • Kai Zheng
  • Tianle Cai
  • Weiran Huang
  • Zhenguo Li
  • Liwei Wang

We study locally differentially private (LDP) bandits learning in this paper. First, we propose simple black-box reduction frameworks that can solve a large family of context-free bandits learning problems with LDP guarantee. Based on our frameworks, we can improve previous best results for private bandits learning with one-point feedback, such as private Bandits Convex Optimization etc, and obtain the first results for Bandits Convex Optimization (BCO) with multi-point feedback under LDP. LDP guarantee and black-box nature make our frameworks more attractive in real applications compared with previous specifically designed and relatively weaker differentially private (DP) algorithms. Further, we also extend our algorithm to Generalized Linear Bandits with regret bound $\tilde{\mc{O}}(T^{3/4}/\varepsilon)$ under $(\varepsilon, \delta)$-LDP and it is conjectured to be optimal. Note given existing $\Omega(T)$ lower bound for DP contextual linear bandits (Shariff & Sheffet, NeurIPS 2018), our result shows a fundamental difference between LDP and DP for contextual bandits.

IJCAI Conference 2019 Conference Paper

DMRAN: A Hierarchical Fine-Grained Attention-Based Network for Recommendation

  • Huizhao Wang
  • Guanfeng Liu
  • An Liu
  • Zhixu Li
  • Kai Zheng

The conventional methods for the next-item recommendation are generally based on RNN or one- dimensional attention with time encoding. They are either hard to preserve the long-term dependencies between different interactions, or hard to capture fine-grained user preferences. In this paper, we propose a Double Most Relevant Attention Network (DMRAN) that contains two layers, i. e. , Item level Attention and Feature Level Self- attention, which are to pick out the most relevant items from the sequence of user’s historical behaviors, and extract the most relevant aspects of relevant items, respectively. Then, we can capture the fine-grained user preferences to better support the next-item recommendation. Extensive experiments on two real-world datasets illustrate that DMRAN can improve the efficiency and effectiveness of the recommendation compared with the state-of-the-art methods.

NeurIPS Conference 2019 Conference Paper

Equipping Experts/Bandits with Long-term Memory

  • Kai Zheng
  • Haipeng Luo
  • Ilias Diakonikolas
  • Liwei Wang

We propose the first black-box approach to obtaining long-term memory guarantees for online learning in the sense of Bousquet and Warmuth, 2002, by reducing the problem to achieving typical switching regret. Specifically, for the classical expert problem with $K$ actions and $T$ rounds, using our general framework we develop various algorithms with a regret bound of order $\order(\sqrt{T(S\ln T + n \ln K)})$ compared to any sequence of experts with $S-1$ switches among $n \leq \min\{S, K\}$ distinct experts. In addition, by plugging specific adaptive algorithms into our framework we also achieve the best of both stochastic and adversarial environments simultaneously, which resolves an open problem of Warmuth and Koolen 2014. Furthermore, we extend our results to the sparse multi-armed bandit setting and show both negative and positive results for long-term memory guarantees. As a side result, our lower bound also implies that sparse losses do not help improve the worst-case regret for contextual bandit, a sharp contrast with the non-contextual case.

AAAI Conference 2019 Conference Paper

Improving One-Class Collaborative Filtering via Ranking-Based Implicit Regularizer

  • Jin Chen
  • Defu Lian
  • Kai Zheng

One-class collaborative filtering (OCCF) problems are vital in many applications of recommender systems, such as news and music recommendation, but suffers from sparsity issues and lacks negative examples. To address this problem, the state-of-the-arts assigned smaller weights to unobserved samples and performed low-rank approximation. However, the ground-truth ratings of unobserved samples are usually set to zero but ill-defined. In this paper, we propose a ranking-based implicit regularizer and provide a new general framework for OCCF, to avert the ground-truth ratings of unobserved samples. We then exploit it to regularize a ranking-based loss function and design efficient optimization algorithms to learn model parameters. Finally, we evaluate them on three realworld datasets. The results show that the proposed regularizer significantly improves ranking-based algorithms and that the proposed framework outperforms the state-of-the-art OCCF algorithms.

AAAI Conference 2019 Conference Paper

Learning Transferable Self-Attentive Representations for Action Recognition in Untrimmed Videos with Weak Supervision

  • Xiao-Yu Zhang
  • Haichao Shi
  • Changsheng Li
  • Kai Zheng
  • Xiaobin Zhu
  • Lixin Duan

Action recognition in videos has attracted a lot of attention in the past decade. In order to learn robust models, previous methods usually assume videos are trimmed as short sequences and require ground-truth annotations of each video frame/sequence, which is quite costly and time-consuming. In this paper, given only video-level annotations, we propose a novel weakly supervised framework to simultaneously locate action frames as well as recognize actions in untrimmed videos. Our proposed framework consists of two major components. First, for action frame localization, we take advantage of the self-attention mechanism to weight each frame, such that the influence of background frames can be effectively eliminated. Second, considering that there are trimmed videos publicly available and also they contain useful information to leverage, we present an additional module to transfer the knowledge from trimmed videos for improving the classification performance in untrimmed ones. Extensive experiments are conducted on two benchmark datasets (i. e. , THUMOS14 and ActivityNet1. 3), and experimental results clearly corroborate the efficacy of our method.

AAAI Conference 2019 Conference Paper

Preference-Aware Task Assignment in Spatial Crowdsourcing

  • Yan Zhao
  • Jinfu Xia
  • Guanfeng Liu
  • Han Su
  • Defu Lian
  • Shuo Shang
  • Kai Zheng

With the ubiquity of smart devices, Spatial Crowdsourcing (SC) has emerged as a new transformative platform that engages mobile users to perform spatio-temporal tasks by physically traveling to specified locations. Thus, various SC techniques have been studied for performance optimization, among which one of the major challenges is how to assign workers the tasks that they are really interested in and willing to perform. In this paper, we propose a novel preference-aware spatial task assignment system based on workers’ temporal preferences, which consists of two components: History-based Context-aware Tensor Decomposition (HCTD) for workers’ temporal preferences modeling and preference-aware task assignment. We model worker preferences with a three-dimension tensor (worker-task-time). Supplementing the missing entries of the tensor through HCTD with the assistant of historical data and other two context matrices, we recover worker preferences for different categories of tasks in different time slots. Several preference-aware task assignment algorithms are then devised, aiming to maximize the total number of task assignments at every time instance, in which we give higher priorities to the workers who are more interested in the tasks. We conduct extensive experiments using a real dataset, verifying the practicability of our proposed methods.

IJCAI Conference 2019 Conference Paper

Profit-driven Task Assignment in Spatial Crowdsourcing

  • Jinfu Xia
  • Yan Zhao
  • Guanfeng Liu
  • Jiajie Xu
  • Min Zhang
  • Kai Zheng

In Spatial Crowdsourcing (SC) systems, mobile users are enabled to perform spatio-temporal tasks by physically traveling to specified locations with the SC platforms. SC platforms manage the systems and recruit mobile users to contribute to the SC systems, whose commercial success depends on the profit attained from the task requesters. In order to maximize its profit, an SC platform needs an online management mechanism to assign the tasks to suitable workers. How to assign the tasks to workers more cost-effectively with the spatio-temporal constraints is one of the most difficult problems in SC. To deal with this challenge, we propose a novel Profit-driven Task Assignment (PTA) problem, which aims to maximize the profit of the platform. Specifically, we first establish a task reward pricing model with tasks' temporal constraints (i. e. , expected completion time and deadline). Then we adopt an optimal algorithm based on tree decomposition to achieve the optimal task assignment and propose greedy algorithms to improve the computational efficiency. Finally, we conduct extensive experiments using real and synthetic datasets, verifying the practicability of our proposed methods.

NeurIPS Conference 2018 Conference Paper

Efficient Online Portfolio with Logarithmic Regret

  • Haipeng Luo
  • Chen-Yu Wei
  • Kai Zheng

We study the decades-old problem of online portfolio management and propose the first algorithm with logarithmic regret that is not based on Cover's Universal Portfolio algorithm and admits much faster implementation. Specifically Universal Portfolio enjoys optimal regret $\mathcal{O}(N\ln T)$ for $N$ financial instruments over $T$ rounds, but requires log-concave sampling and has a large polynomial running time. Our algorithm, on the other hand, ensures a slightly larger but still logarithmic regret of $\mathcal{O}(N^2(\ln T)^4)$, and is based on the well-studied Online Mirror Descent framework with a novel regularizer that can be implemented via standard optimization methods in time $\mathcal{O}(TN^{2. 5})$ per round. The regret of all other existing works is either polynomial in $T$ or has a potentially unbounded factor such as the inverse of the smallest price relative.

IJCAI Conference 2018 Conference Paper

LC-RNN: A Deep Learning Model for Traffic Speed Prediction

  • Zhongjian Lv
  • Jiajie Xu
  • Kai Zheng
  • Hongzhi Yin
  • Pengpeng Zhao
  • Xiaofang Zhou

Traffic speed prediction is known as an important but challenging problem. In this paper, we propose a novel model, called LC-RNN, to achieve more accurate traffic speed prediction than existing solutions. It takes advantage of both RNN and CNN models by a rational integration of them, so as to learn more meaningful time-series patterns that can adapt to the traffic dynamics of surrounding areas. Furthermore, since traffic evolution is restricted by the underlying road network, a network embedded convolution structure is proposed to capture topology aware features. The fusion with other information, including periodicity and context factors, is also considered to further improve accuracy. Extensive experiments on two real datasets demonstrate that our proposed LC-RNN outperforms six well-known existing methods.

IJCAI Conference 2017 Conference Paper

Efficient Private ERM for Smooth Objectives

  • Jiaqi Zhang
  • Kai Zheng
  • Wenlong Mou
  • Liwei Wang

In this paper, we consider efficient differentially private empirical risk minimization from the viewpoint of optimization algorithms. For strongly convex and smooth objectives, we prove that gradient descent with output perturbation not only achieves nearly optimal utility, but also significantly improves the running time of previous state-of-the-art private optimization algorithms, for both $\epsilon$-DP and $(\epsilon, \delta)$-DP. For non-convex but smooth objectives, we propose an RRPSGD (Random Round Private Stochastic Gradient Descent) algorithm, which provably converges to a stationary point with privacy guarantee. Besides the expected utility bounds, we also provide guarantees in high probability form. Experiments demonstrate that our algorithm consistently outperforms existing method in both utility and running time.

v2026.09.13