Arrow Research search

Author name cluster

Wenkai Xu

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

10 papers
2 author rows

Possible papers

10

EAAI Journal 2024 Journal Article

Behavioral response of fish under ammonia nitrogen stress based on machine vision

  • Wenkai Xu
  • Chang Liu
  • Guangxu Wang
  • Yue Zhao
  • Jiaxuan Yu
  • Akhter Muhammad
  • Daoliang Li

The long-term accumulation of ammonia nitrogen in aquaculture seriously affects the life of fish and even causes large-scale death. Moreover, when the concentration of ammonia nitrogen starts to accumulate, it is a judgment standard to provide early warning through the changes in fish behavior to prevent excessive ammonia nitrogen in water. Therefore, this paper proposes a novel approach to monitoring water quality for aquaculture based on deep learning and three-dimensional movement trajectory. The improved YOLOv8 model was used as the object detection approach to obtain three-dimensional position information of fish by combining Kalman filter, Kuhn Munkres (KM) algorithm, and Kernelized Correlation Filters (KCF) algorithm. The proposed approach was evaluated in the recovery experiment of acute ammonia nitrogen stress of sturgeon, bass, and crucian. The experimental results show that the precision, recall, mAP@0. 5, and mAP@0. 5: 0. 95 of the improved YOLOv8 model are 0. 964, 0. 914, 0. 979, and 0. 602, respectively. In addition, the proposed three-dimensional positioning approach can qualitatively and quantitatively analyze the fish behavior in different stages and further explores the fish behavior changes through behavior trajectories, volumes of exercise, spatial distribution, and movement velocity. This research provides a new method and idea for studying the abnormal behavior of aquatic animals under ammonia nitrogen stress and has theoretical and practical significance.

UAI Conference 2023 Conference Paper

Learning Nonlinear Causal Effect via Kernel Anchor Regression

  • Wenqi Shi
  • Wenkai Xu

Learning causal effects is a fundamental problem in science. Anchor regression has been developed to address this problem for a large class of causal graphical models, though the relationships between the variables are assumed to be linear. In this work, we tackle the nonlinear setting by proposing kernel anchor regression (KAR). Beyond a classic two-stage least square (2SLS) estimator, we also study an improved variant that involves nonparametric kernel regression in three separate stages. We provide convergence results for the proposed KAR estimators and the identifiability conditions for KAR to learn the nonlinear structural equation models (SEM). Experimental results demonstrate the superior performances of the proposed KAR estimators over existing baselines.

NeurIPS Conference 2022 Conference Paper

A Kernelised Stein Statistic for Assessing Implicit Generative Models

  • Wenkai Xu
  • Gesine D Reinert

Synthetic data generation has become a key ingredient for training machine learning procedures, addressing tasks such as data augmentation, analysing privacy-sensitive data, or visualising representative samples. Assessing the quality of such synthetic data generators hence has to be addressed. As (deep) generative models for synthetic data often do not admit explicit probability distributions, classical statistical procedures for assessing model goodness-of-fit may not be applicable. In this paper, we propose a principled procedure to assess the quality of a synthetic data generator. The procedure is a Kernelised Stein Discrepancy-type test which is based on a non-parametric Stein operator for the synthetic data generator of interest. This operator is estimated from samples which are obtained from the synthetic data generator and hence can be applied even when the model is only implicit. In contrast to classical testing, the sample size from the synthetic data generator can be as large as desired, while the size of the observed data that the generator aims to emulate is fixed. Experimental results on synthetic distributions and trained generative models on synthetic and real datasets illustrate that the method shows improved power performance compared to existing approaches.

NeurIPS Conference 2022 Conference Paper

AgraSSt: Approximate Graph Stein Statistics for Interpretable Assessment of Implicit Graph Generators

  • Wenkai Xu
  • Gesine D Reinert

We propose and analyse a novel statistical procedure, coined AgraSSt, to assess the quality of graph generators which may not be available in explicit forms. In particular, AgraSSt can be used to determine whether a learned graph generating process is capable of generating graphs which resemble a given input graph. Inspired by Stein operators for random graphs, the key idea of AgraSSt is the construction of a kernel discrepancy based on an operator obtained from the graph generator. AgraSSt can provide interpretable criticisms for a graph generator training procedure and help identify reliable sample batches for downstream tasks. We give theoretical guarantees for a broad class of random graph models. Moreover, we provide empirical results on both synthetic input graphs with known graph generation procedures, and real-world input graphs that the state-of-the-art (deep) generative models for graphs are trained on.

ICML Conference 2021 Conference Paper

Interpretable Stein Goodness-of-fit Tests on Riemannian Manifold

  • Wenkai Xu
  • Takeru Matsuda

In many applications, we encounter data on Riemannian manifolds such as torus and rotation groups. Standard statistical procedures for multivariate data are not applicable to such data. In this study, we develop goodness-of-fit testing and interpretable model criticism methods for general distributions on Riemannian manifolds, including those with an intractable normalization constant. The proposed methods are based on extensions of kernel Stein discrepancy, which are derived from Stein operators on Riemannian manifolds. We discuss the connections between the proposed tests with existing ones and provide a theoretical analysis of their asymptotic Bahadur efficiency. Simulation results and real data applications show the validity and usefulness of the proposed methods.

NeurIPS Conference 2021 Conference Paper

Meta Two-Sample Testing: Learning Kernels for Testing with Limited Data

  • Feng Liu
  • Wenkai Xu
  • Jie Lu
  • Danica J. Sutherland

Modern kernel-based two-sample tests have shown great success in distinguishing complex, high-dimensional distributions by learning appropriate kernels (or, as a special case, classifiers). Previous work, however, has assumed that many samples are observed from both of the distributions being distinguished. In realistic scenarios with very limited numbers of data samples, it can be challenging to identify a kernel powerful enough to distinguish complex distributions. We address this issue by introducing the problem of meta two-sample testing (M2ST), which aims to exploit (abundant) auxiliary data on related tasks to find an algorithm that can quickly identify a powerful test on new target tasks. We propose two specific algorithms for this task: a generic scheme which improves over baselines, and a more tailored approach which performs even better. We provide both theoretical justification and empirical evidence that our proposed meta-testing schemes outperform learning kernel-based tests directly from scarce observations, and identify when such schemes will be successful.

NeurIPS Conference 2020 Conference Paper

A kernel test for quasi-independence

  • Tamara Fernandez
  • Wenkai Xu
  • Marc Ditzhaus
  • Arthur Gretton

We consider settings in which the data of interest correspond to pairs of ordered times, e. g, the birth times of the first and second child, the times at which a new user creates an account and makes the first purchase on a website, and the entry and survival times of patients in a clinical trial. In these settings, the two times are not independent (the second occurs after the first), yet it is still of interest to determine whether there exists significant dependence "beyond" their ordering in time. We refer to this notion as "quasi-(in)dependence. " For instance, in a clinical trial, to avoid biased selection, we might wish to verify that recruitment times are quasi-independent of survival times, where dependencies might arise due to seasonal effects. In this paper, we propose a nonparametric statistical test of quasi-independence. Our test considers a potentially infinite space of alternatives, making it suitable for complex data where the nature of the possible quasi-dependence is not known in advance. Standard parametric approaches are recovered as special cases, such as the classical conditional Kendall's tau, and log-rank tests. The tests apply in the right-censored setting: an essential feature in clinical trials, where patients can withdraw from the study. We provide an asymptotic analysis of our test-statistic, and demonstrate in experiments that our test obtains better power than existing approaches, while being more computationally efficient.

ICML Conference 2020 Conference Paper

Kernelized Stein Discrepancy Tests of Goodness-of-fit for Time-to-Event Data

  • Tamara Fernandez
  • Nicolás Rivera
  • Wenkai Xu
  • Arthur Gretton

Survival Analysis and Reliability Theory are concerned with the analysis of time-to-event data, in which observations correspond to waiting times until an event of interest such as death from a particular disease or failure of a component in a mechanical system. This type of data is unique due to the presence of censoring, a type of missing data that occurs when we do not observe the actual time of the event of interest but, instead, we have access to an approximation for it given by random interval in which the observation is known to belong. Most traditional methods are not designed to deal with censoring, and thus we need to adapt them to censored time-to-event data. In this paper, we focus on non-parametric goodness-of-fit testing procedures based on combining the Stein’s method and kernelized discrepancies. While for uncensored data, there is a natural way of implementing a kernelized Stein discrepancy test, for censored data there are several options, each of them with different advantages and disadvantages. In this paper, we propose a collection of kernelized Stein discrepancy tests for time-to-event data, and we study each of them theoretically and empirically; our experimental results show that our proposed methods perform better than existing tests, including previous tests based on a kernelized maximum mean discrepancy.

ICML Conference 2020 Conference Paper

Learning Deep Kernels for Non-Parametric Two-Sample Tests

  • Feng Liu 0003
  • Wenkai Xu
  • Jie Lu 0001
  • Guangquan Zhang 0001
  • Arthur Gretton
  • Danica J. Sutherland

We propose a class of kernel-based two-sample tests, which aim to determine whether two sets of samples are drawn from the same distribution. Our tests are constructed from kernels parameterized by deep neural nets, trained to maximize test power. These tests adapt to variations in distribution smoothness and shape over space, and are especially suited to high dimensions and complex data. By contrast, the simpler kernels used in prior kernel testing work are spatially homogeneous, and adaptive only in lengthscale. We explain how this scheme includes popular classifier-based two-sample tests as a special case, but improves on them in general. We provide the first proof of consistency for the proposed adaptation method, which applies both to kernels on deep features and to simpler radial basis kernels or multiple kernel learning. In experiments, we establish the superior performance of our deep kernels in hypothesis testing on benchmark and real-world data. The code of our deep-kernel-based two-sample tests is available at github. com/fengliu90/DK-for-TST.

NeurIPS Conference 2017 Conference Paper

A Linear-Time Kernel Goodness-of-Fit Test

  • Wittawat Jitkrittum
  • Wenkai Xu
  • Zoltan Szabo
  • Kenji Fukumizu
  • Arthur Gretton

We propose a novel adaptive test of goodness-of-fit, with computational cost linear in the number of samples. We learn the test features that best indicate the differences between observed samples and a reference model, by minimizing the false negative rate. These features are constructed via Stein's method, meaning that it is not necessary to compute the normalising constant of the model. We analyse the asymptotic Bahadur efficiency of the new test, and prove that under a mean-shift alternative, our test always has greater relative efficiency than a previous linear-time kernel test, regardless of the choice of parameters for that test. In experiments, the performance of our method exceeds that of the earlier linear-time test, and matches or exceeds the power of a quadratic-time kernel test. In high dimensions and where model structure may be exploited, our goodness of fit test performs far better than a quadratic-time two-sample test based on the Maximum Mean Discrepancy, with samples drawn from the model.

v2026.09.13