Arrow Research search

Author name cluster

Jacques Wainer

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

12 papers
2 author rows

Possible papers

12

JMLR Journal 2023 Journal Article

A Bayesian Bradley-Terry model to compare multiple ML algorithms on multiple data sets

  • Jacques Wainer

his paper presents a Bayesian model, called the Bayesian Bradley Terry (BBT) model, for comparing multiple algorithms on multiple data sets based on any metric. The model is an extension of the Bradley Terry model, which tracks the number of wins each algorithm has on different data sets. Unlike frequentist methods such as Demsar tests on mean rank or multiple pairwise Wilcoxon tests, the Bayesian approach provides a more nuanced understanding of the algorithms’ performance and allows for the definition of the “region of practical equivalence” (ROPE) for two algorithms. Additionally, the paper introduces the concept of “local ROPE,” which assesses the significance of the difference in mean measure between two algorithms using effect sizes, and can be applied in frequentist approaches as well. Both an R package and a Python program implementing the BBT are available for use. [abs] [ pdf ][ bib ] [ code ] &copy JMLR 2023. ( edit, beta )

AIIM Journal 2019 Journal Article

A data-driven approach to referable diabetic retinopathy detection

  • Ramon Pires
  • Sandra Avila
  • Jacques Wainer
  • Eduardo Valle
  • Michael D. Abramoff
  • Anderson Rocha

Prior art on automated screening of diabetic retinopathy and direct referral decision shows promising performance; yet most methods build upon complex hand-crafted features whose performance often fails to generalize. Objective We investigate data-driven approaches that extract powerful abstract representations directly from retinal images to provide a reliable referable diabetic retinopathy detector. Methods We gradually build the solution based on convolutional neural networks, adding data augmentation, multi-resolution training, robust feature-extraction augmentation, and a patient-basis analysis, testing the effectiveness of each improvement. Results The proposed method achieved an area under the ROC curve of 98. 2% (95% CI: 97. 4–98. 9%) under a strict cross-dataset protocol designed to test the ability to generalize — training on the Kaggle competition dataset and testing using the Messidor-2 dataset. With a 5 × 2-fold cross-validation protocol, similar results are achieved for Messidor-2 and DR2 datasets, reducing the classification error by over 44% when compared to most published studies in existing literature. Conclusion Additional boost strategies can improve performance substantially, but it is important to evaluate whether the additional (computation- and implementation-) complexity of each improvement is worth its benefits. We also corroborate that novel families of data-driven methods are the state of the art for diabetic retinopathy screening. Significance: By learning powerful discriminative patterns directly from available training retinal images, it is possible to perform referral diagnostics without detecting individual lesions.

JBHI Journal 2017 Journal Article

Beyond Lesion-Based Diabetic Retinopathy: A Direct Approach for Referral

  • Ramon Pires
  • Sandra Avila
  • Herbert F. Jelinek
  • Jacques Wainer
  • Eduardo Valle
  • Anderson Rocha

Diabetic retinopathy (DR) is the leading cause of blindness in adults, but can be managed if detected early. Automated DR screening helps by indicating which patients should be referred to the doctor. However, current techniques of automated screening still depend too much on the detection of individual lesions. In this study, we bypass lesion detection, and directly train a classifier for DR referral. Additional novelties are the use of state-of-the-art mid-level features for the retinal images: BossaNova and Fisher Vector. Those features extend the classical Bags of Visual Words and greatly improve the accuracy of complex classification tasks. The proposed technique for direct referral is promising, achieving an area under the curve of 96. 4%, thus, reducing the classification error by almost 40% over the current state of the art, held by lesion-based techniques.

JMLR Journal 2017 Journal Article

Empirical Evaluation of Resampling Procedures for Optimising SVM Hyperparameters

  • Jacques Wainer
  • Gavin Cawley

Tuning the regularisation and kernel hyperparameters is a vital step in optimising the generalisation performance of kernel methods, such as the support vector machine (SVM). This is most often performed by minimising a resampling/cross-validation based model selection criterion, however there seems little practical guidance on the most suitable form of resampling. This paper presents the results of an extensive empirical evaluation of resampling procedures for SVM hyperparameter selection, designed to address this gap in the machine learning literature. We tested 15 different resampling procedures on 121 binary classification data sets in order to select the best SVM hyperparameters. We used three very different statistical procedures to analyse the results: the standard multi- classifier/multi-data set procedure proposed by Dem\v{s}ar, the confidence intervals on the excess loss of each procedure in relation to 5-fold cross validation, and the Bayes factor analysis proposed by Barber. We conclude that a 2-fold procedure is appropriate to select the hyperparameters of an SVM for data sets for 1000 or more datapoints, while a 3-fold procedure is appropriate for smaller data sets. [abs] [ pdf ][ bib ] &copy JMLR 2017. ( edit, beta )

JMLR Journal 2008 Journal Article

HPB: A Model for Handling BN Nodes with High Cardinality Parents

  • Jorge Jambeiro Filho
  • Jacques Wainer

We replaced the conditional probability tables of Bayesian network nodes whose parents have high cardinality with a multilevel empirical hierarchical Bayesian model called hierarchical pattern Bayes (HPB). The resulting Bayesian networks achieved significant performance improvements over Bayesian networks with the same structure and traditional conditional probability tables, over Bayesian networks with simpler structures like naïve Bayes and tree augmented naïve Bayes, over Bayesian networks where traditional conditional probability tables were substituted by noisy-OR gates, default tables, decision trees and decision graphs and over Bayesian networks constructed after a cardinality reduction preprocessing phase using the agglomerative information bottleneck method. Our main tests took place in important fraud detection domains, which are characterized by the presence of high cardinality attributes and by the existence of relevant interactions among them. Other tests, over UCI data sets, show that HPB may have a quite wide applicability. [abs] [ pdf ][ bib ] &copy JMLR 2008. ( edit, beta )

IJCAI Conference 2007 Conference Paper

  • Jorge Jambeiro Filho
  • Jacques Wainer

We employed a multilevel hierarchical Bayesian model in the task of exploiting relevant interactions among high cardinality attributes in a classification problem without overfitting. With this model, we calculate posterior class probabilities for a pattern W combining the observations of W in the training set with prior class probabilities that are obtained recursively from the observations of patterns that are strictly more generic than W. The model achieved performance improvements over standard Bayesian network methods like Naive Bayes and Tree Augmented Naive Bayes, over Bayesian Networks where traditional conditional probability tables were substituted by Noisy-or gates, Default Tables, Decision Trees and Decision Graphs, and over Bayesian Networks constructed after a cardinality reduction preprocessing phase using the Agglomerative Information Bottleneck method.

AIIM Journal 1997 Journal Article

A temporal extension to the parsimonious covering theory

  • Jacques Wainer
  • Alexandre de Melo Rezende

In this paper, parsimonious covering theory is extended in such a way that temporal knowledge can be accommodated. In addition to causally associating possible manifestations with disorders, temporal relationships about duration and the time elapsed before a manifestation comes into existence can be represented by a graph. Precise definitions of the solution of a temporal diagnostic problem, as well as algorithms to compute the solutions are provided. The medical suitability of the extended parsimonious cover theory is studied in the domain of food-borne disease.

AAAI Conference 1992 Conference Paper

Combining Circumscription and Modal Logic

  • Jacques Wainer

This paper discusses the logic LKM which extends circumscription into an epistemic domain. This extension will allow us to define circumscription of predicates that appear within the context of a modal operator. In fact, LKM can be seen as a method of extending any first-order nonmonotonic logic whose semantic definition is based on a partial-order among models, into a new nonmonotonic logic defined for a modal language, whose modal operator (K) follows an underlying S5 or weak-S5 semantics. One interesting use of this nonmonotonic logic is to model nonmonotonic aspects of the communication between agents.

v2026.09.13