Arrow Research search

Author name cluster

Haiying Wang

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

7 papers
1 author row

Possible papers

7

JMLR Journal 2025 Journal Article

Optimal subsampling for high-dimensional partially linear models via machine learning methods

  • Yujing Shao
  • Lei Wang
  • Heng Lian
  • Haiying Wang

In this paper, we explore optimal subsampling strategies for estimating the parametric regression coefficients in partially linear models with unknown nuisance functions involving high-dimensional and potentially endogenous covariates. To address model misspecifications and the curse of dimensionality, we leverage flexible machine learning (ML) techniques to estimate the unknown nuisance functions. By constructing an unbiased subsampling Neyman-orthogonal score function, we eliminate regularization bias. A two-step algorithm is then used to obtain appropriate ML estimators of the nuisance functions, mitigating the risk of over-fitting. Using martingale techniques, we establish the unconditional consistency and asymptotic normality of the subsample estimators. Furthermore, we derive optimal subsampling probabilities, including A-optimal and L-optimal probabilities as special cases. The proposed optimal subsampling approach is extended to partially linear instrumental variable models to account for potential endogeneity through instrumental variables. Simulation studies and an empirical analysis of the Physicochemical Properties of Protein Tertiary Structure dataset demonstrate the superior performance of our subsample estimators. [abs] [ pdf ][ bib ] &copy JMLR 2025. ( edit, beta )

NeurIPS Conference 2024 Conference Paper

Scale-invariant Optimal Sampling for Rare-events Data and Sparse Models

  • Jing Wang
  • Haiying Wang
  • Hao H. Zhang

Subsampling is effective in tackling computational challenges for massive data with rare events. Overly aggressive subsampling may adversely affect estimation efficiency, and optimal subsampling is essential to mitigate the information loss. However, existing optimal subsampling probabilities depends on data scales, and some scaling transformations may result in inefficient subsamples. This problem is more significant when there are inactive features, because their influence on the subsampling probabilities can be arbitrarily magnified by inappropriate scaling transformations. We tackle this challenge and introduce a scale-invariant optimal subsampling function in the context of sparse models, where inactive features are commonly assumed. Instead of focusing on estimating model parameters, we define an optimal subsampling function to minimize the prediction error, using adaptive lasso as an example to outline the estimation procedure and study its theoretical guarantee. We first introduce the adaptive lasso estimator for rare-events data and establish its oracle properties, thereby validating the use of subsampling. Then we derive a scale-invariant optimal subsampling function that minimizes the prediction error of the inverse probability weighted (IPW) adaptive lasso. Finally, we present an estimator based on the maximum sampled conditional likelihood (MSCL) to further improve the estimation efficiency. We conduct numerical experiments using both simulated and real-world data sets to demonstrate the performance of the proposed methods.

JMLR Journal 2022 Journal Article

Maximum sampled conditional likelihood for informative subsampling

  • Haiying Wang
  • Jae Kwang Kim

Subsampling is a computationally effective approach to extract information from massive data sets when computing resources are limited. After a subsample is taken from the full data, most available methods use an inverse probability weighted (IPW) objective function to estimate the model parameters. The IPW estimator does not fully utilize the information in the selected subsample. In this paper, we propose to use the maximum sampled conditional likelihood estimator (MSCLE) based on the sampled data. We established the asymptotic normality of the MSCLE and prove that its asymptotic variance covariance matrix is the smallest among a class of asymptotically unbiased estimators, including the IPW estimator. We further discuss the asymptotic results with the L-optimal subsampling probabilities and illustrate the estimation procedure with generalized linear models. Numerical experiments are provided to evaluate the practical performance of the proposed method. [abs] [ pdf ][ bib ] &copy JMLR 2022. ( edit, beta )

NeurIPS Conference 2021 Conference Paper

Nonuniform Negative Sampling and Log Odds Correction with Rare Events Data

  • Haiying Wang
  • Aonan Zhang
  • Chong Wang

We investigate the issue of parameter estimation with nonuniform negative sampling for imbalanced data. We first prove that, with imbalanced data, the available information about unknown parameters is only tied to the relatively small number of positive instances, which justifies the usage of negative sampling. However, if the negative instances are subsampled to the same level of the positive cases, there is information loss. To maintain more information, we derive the asymptotic distribution of a general inverse probability weighted (IPW) estimator and obtain the optimal sampling probability that minimizes its variance. To further improve the estimation efficiency over the IPW method, we propose a likelihood-based estimator by correcting log odds for the sampled data and prove that the improved estimator has the smallest asymptotic variance among a large class of estimators. It is also more robust to pilot misspecification. We validate our approach on simulated data as well as a real click-through rate dataset with more than 0. 3 trillion instances, collected over a period of a month. Both theoretical and empirical results demonstrate the effectiveness of our method.

JBHI Journal 2021 Journal Article

Triple Up-Sampling Segmentation Network With Distribution Consistency Loss for Pathological Diagnosis of Cervical Precancerous Lesions

  • Zhu Meng
  • Zhicheng Zhao
  • Bingyang Li
  • Fei Su
  • Limei Guo
  • Haiying Wang

Objective: Cervical cancer, as one of the most frequently diagnosed cancers in women, is curable when detected early. However, automated algorithms for cervical pathology precancerous diagnosis are limited. Methods: In this paper, instead of popular patch-wise classification, an end-to-end patch-wise segmentation algorithm is proposed to focus on the spatial structure changes of pathological tissues. Specifically, a triple up-sampling segmentation network (TriUpSegNet) is constructed to aggregate spatial information. Second, a distribution consistency loss (DC-loss) is designed to constrain the model to fit the inter-class relationship of the cervix. Third, the Gauss-like weighted post-processing is employed to reduce patch stitching deviation and noise. Results: The algorithm is evaluated on three challenging and public datasets: 1) MTCHI for cervical precancerous diagnosis, 2) DigestPath for colon cancer, and 3) PAIP for liver cancer. The Dice coefficient is 0. 7413 on the MTCHI dataset, which is significantly higher than the published state-of-the-art results. Conclusion: Experiments on the public dataset MTCHI indicate the superiority of the proposed algorithm on cervical pathology precancerous diagnosis. In addition, the experiments on two other pathological datasets, i. e. , DigestPath and PAIP, demonstrate the effectiveness and generalization ability of the TriUpSegNet and weighted post-processing on colon and liver cancers. Significance: The end-to-end TriUpSegNet with DC-loss and weighted post-processing leads to improved segmentation in pathology of various cancers.

JMLR Journal 2019 Journal Article

More Efficient Estimation for Logistic Regression with Optimal Subsamples

  • Haiying Wang

In this paper, we propose improved estimation method for logistic regression based on subsamples taken according the optimal subsampling probabilities developed in Wang et al. (2018). Both asymptotic results and numerical results show that the new estimator has a higher estimation efficiency. We also develop a new algorithm based on Poisson subsampling, which does not require to approximate the optimal subsampling probabilities all at once. This is computationally advantageous when available random-access memory is not enough to hold the full data. Interestingly, asymptotic distributions also show that Poisson subsampling produces a more efficient estimator if the sampling ratio, the ratio of the subsample size to the full data sample size, does not converge to zero. We also obtain the unconditional asymptotic distribution for the estimator based on Poisson subsampling. Pilot estimators are required to calculate subsampling probabilities and to correct biases in un-weighted estimators; interestingly, even if pilot estimators are inconsistent, the proposed method still produce consistent and asymptotically normal estimators. [abs] [ pdf ][ bib ] &copy JMLR 2019. ( edit, beta )

IJCAI Conference 2007 Conference Paper

  • Haiying Wang
  • Huiru Zheng
  • FRANCISCO AZUAJE

The ability to learn from data and to improve its performance through incremental learning makes self-adaptive neural networks (SANNs) a powerful tool to support knowledge discovery. However, the development of SANNs has traditionally focused on data domains that are assumed to be modeled by a Gaussian distribution. The analysis of data governed by other statistical models, such as the Poisson distribution, has received less attention from the data mining community. Based on special considerations of the statistical nature of data following a Poisson distribution, this paper introduces a SANN, Poisson-based Self-Organizing Tree Algorithm (PSOTA), which implements novel similarity matching criteria and neuron weight adaptation schemes. It was tested on synthetic and real world data (serial analysis of gene expression data). PSOTA-based data analysis supported the automated identification of more meaningful clusters. By visualizing the dendrograms generated by PSOTA, complex inter- and intra-cluster relationships encoded in the data were also highlighted and readily understood. This study indicate that, in comparison to the traditional Self-Organizing Tree Algorithm (SOTA), PSOTA offers significant improvements in pattern discovery and visualization in data modeled by the Poisson distribution, such as serial analysis of gene expression data.

v2026.09.13