Arrow Research search

Author name cluster

Vipin Kumar

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

18 papers
1 author row

Possible papers

18

AAAI Conference 2026 Conference Paper

Knowledge-Guided Machine Learning: A Paradigm Shift in AI for Science

  • Anuj Karpatne
  • Xiaowei Jia
  • Vipin Kumar

As advances in artificial intelligence (AI) and machine learning (ML) continue to transform commercial applications, the scientific community is increasingly eager to harness AI/ML’s power to accelerate modeling and discovery. However, purely data-driven AI methods often lack interpretability, generalizability, and consistency with established scientific principles. Conversely, traditional process-based models embody deep scientific knowledge but suffer from limited scalability or incomplete representation of complex systems. Knowledge-guided machine learning (KGML) offers a promising path forward by integrating scientific knowledge with data-driven approaches to produce AI models that are robust, trustworthy, and capable of advancing both AI and science. This talk summarizes the foundations of KGML, outlines a taxonomy for organizing research efforts, and highlights emerging opportunities for broad scientific impact.

IJCAI Conference 2023 Conference Paper

Bayesian Federated Learning: A Survey

  • Longbing Cao
  • Hui Chen
  • Xuhui Fan
  • Joao Gama
  • Yew-Soon Ong
  • Vipin Kumar

Federated learning (FL) demonstrates its advantages in integrating distributed infrastructure, communication, computing and learning in a privacy-preserving manner. However, the robustness and capabilities of existing FL methods are challenged by limited and dynamic data and conditions, complexities including heterogeneities and uncertainties, and analytical explainability. Bayesian federated learning (BFL) has emerged as a promising approach to address these issues. This survey presents a critical overview of BFL, including its basic concepts, its relations to Bayesian learning in the context of FL, and a taxonomy of BFL from both Bayesian and federated perspectives. We categorize and discuss client- and server-side and FL-based BFL methods and their pros and cons. The limitations of the existing BFL methods and the future directions of BFL research further address the intricate requirements of real-life FL applications.

JBHI Journal 2021 Journal Article

A Computational Method for Learning Disease Trajectories From Partially Observable EHR Data

  • Wonsuk Oh
  • Michael S. Steinbach
  • M. Regina Castro
  • Kevin A. Peterson
  • Vipin Kumar
  • Pedro J. Caraballo
  • Gyorgy J. Simon

Diseases can show different courses of progression even when patients share the same risk factors. Recent studies have revealed that the use of trajectories, the order in which diseases manifest throughout life, can be predictive of the course of progression. In this study, we propose a novel computational method for learning disease trajectories from EHR data. The proposed method consists of three parts: first, we propose an algorithm for extracting trajectories from EHR data; second, three criteria for filtering trajectories; and third, a likelihood function for assessing the risk of developing a set of outcomes given a trajectory set. We applied our methods to extract a set of disease trajectories from Mayo Clinic EHR data and evaluated it internally based on log-likelihood, which can be interpreted as the trajectories' ability to explain the observed (partial) disease progressions. We then externally evaluated the trajectories on EHR data from an independent health system, M Health Fairview. The proposed algorithm extracted a comprehensive set of disease trajectories that can explain the observed outcomes substantially better than competing methods and the proposed filtering criteria selected a small subset of disease trajectories that are highly interpretable and suffered only a minimal (relative 5%) loss of the ability to explain disease progression in both the internal and external validation.

IJCAI Conference 2019 Conference Paper

Recurrent Generative Networks for Multi-Resolution Satellite Data: An Application in Cropland Monitoring

  • Xiaowei Jia
  • Mengdie Wang
  • Ankush Khandelwal
  • Anuj Karpatne
  • Vipin Kumar

Effective and timely monitoring of croplands is critical for managing food supply. While remote sensing data from earth-observing satellites can be used to monitor croplands over large regions, this task is challenging for small-scale croplands as they cannot be captured precisely using coarse-resolution data. On the other hand, the remote sensing data in higher resolution are collected less frequently and contain missing or disturbed data. Hence, traditional sequential models cannot be directly applied on high-resolution data to extract temporal patterns, which are essential to identify crops. In this work, we propose a generative model to combine multi-scale remote sensing data to detect croplands at high resolution. During the learning process, we leverage the temporal patterns learned from coarse-resolution data to generate missing high-resolution data. Additionally, the proposed model can track classification confidence in real time and potentially lead to an early detection. The evaluation in an intensively cultivated region demonstrates the effectiveness of the proposed method in cropland detection.

AAAI Conference 2016 Conference Paper

Understanding Dominant Factors for Precipitation over the Great Lakes Region

  • Soumyadeep Chatterjee
  • Stefan Liess
  • Arindam Banerjee
  • Vipin Kumar

Statistical modeling of local precipitation involves understanding local, regional and global factors informative of precipitation variability in a region. Modern machine learning methods for feature selection can potentially be explored for identifying statistically significant features from pool of potential predictors of precipitation. In this work, we consider sparse regression, which simultaneously performs feature selection and regression, followed by random permutation tests for selecting dominant factors. We consider average winter precipitation over Great Lakes Region in order to identify its dominant influencing factors. Experiments show that global climate indices, computed at different temporal lags, offer predictive information for winter precipitation. Further, among the dominant factors identified using randomized permutation tests, multiple climate indices indicate the influence of geopotential height patterns on winter precipitation. Using composite analysis, we illustrate that certain patterns are indeed typical in high and low precipitation years, and offer plausible scientific reasons for variations in precipitation. Thus, feature selection methods can be useful in identifying influential climate processes and variables, and thereby provide useful hypotheses over physical mechanisms affecting local precipitation.

IJCAI Conference 2015 Conference Paper

Clustering Dynamic Spatio-Temporal Patterns in The Presence of Noise and Missing Data

  • Xi Chen
  • James H. Faghmous
  • Ankush Khandelwal
  • Vipin Kumar

Clustering has gained widespread use, especially for static data. However, the rapid growth of spatio-temporal data from numerous instruments, such as earth-orbiting satellites, has created a need for spatio-temporal clustering methods to extract and monitor dynamic clusters. Dynamic spatiotemporal clustering faces two major challenges: First, the clusters are dynamic and may change in size, shape, and statistical properties over time. Second, numerous spatio-temporal data are incomplete, noisy, heterogeneous, and highly variable (over space and time). We propose a new spatiotemporal data mining paradigm, to autonomously identify dynamic spatio-temporal clusters in the presence of noise and missing data. Our proposed approach is more robust than traditional clustering and image segmentation techniques in the case of dynamic patterns, non-stationary, heterogeneity, and missing data. We demonstrate our method’s performance on a real-world application of monitoring in-land water bodies on a global scale.

IS Journal 2015 Journal Article

Nonoccurring Behavior Analytics: A New Area

  • Longbing Cao
  • Philip S. Yu
  • Vipin Kumar

Nonoccurring behaviors (NOBs) refer to those behaviors that should happen but do not take place for some reasons. They are widely seen in online, business, government, health, medical, scientific, and social data and applications. Very limited research has been undertaken to explore such behaviors owing to their invisibility and the significant challenges in analyzing NOBs. This article outlines the concept, intrinsic characteristics, significant challenges, main issues, research directions, and state-of-the-art work related to NOBs, followed by the prospects for and applications of NOB analytics.

AAAI Conference 2014 Conference Paper

Spatio-Temporal Consistency as a Means to Identify Unlabeled Objects in a Continuous Data Field

  • James Faghmous
  • Hung Nguyen
  • Matthew Le
  • Vipin Kumar

Mesoscale ocean eddies are a critical component of the Earth System as they dominate the ocean’s kinetic energy and impact the global distribution of oceanic heat, salinity, momentum, and nutrients. Therefore, accurately representing these dynamic features is critical for our planet’s sustainability. The majority of methods that identify eddies from satellite observations analyze the data in a frame-by-frame basis despite the fact that eddies are dynamic objects that propagate across space and time. We introduce the notion of spatio-temporal consistency to identify eddies in a continuous spatiotemporal field, to simultaneously ensure that the features detected are both spatially and temporally consistent. Our spatio-temporal consistency approach allows us to remove most of the expert criteria used in traditional methods to reduce false negatives. The removal of arbitrary heuristics enables us to render more complete eddy dynamics by identifying smaller and longer lived eddies compared to existing methods.

YNICL Journal 2013 Journal Article

Complex biomarker discovery in neuroimaging data: Finding a needle in a haystack

  • Gowtham Atluri
  • Kanchana Padmanabhan
  • Gang Fang
  • Michael Steinbach
  • Jeffrey R. Petrella
  • Kelvin Lim
  • Angus MacDonald
  • Nagiza F. Samatova

Neuropsychiatric disorders such as schizophrenia, bipolar disorder and Alzheimer's disease are major public health problems. However, despite decades of research, we currently have no validated prognostic or diagnostic tests that can be applied at an individual patient level. Many neuropsychiatric diseases are due to a combination of alterations that occur in a human brain rather than the result of localized lesions. While there is hope that newer imaging technologies such as functional and anatomic connectivity MRI or molecular imaging may offer breakthroughs, the single biomarkers that are discovered using these datasets are limited by their inability to capture the heterogeneity and complexity of most multifactorial brain disorders. Recently, complex biomarkers have been explored to address this limitation using neuroimaging data. In this manuscript we consider the nature of complex biomarkers being investigated in the recent literature and present techniques to find such biomarkers that have been developed in related areas of data mining, statistics, machine learning and bioinformatics.

AAAI Conference 2013 Conference Paper

Multiple Hypothesis Object Tracking For Unsupervised Self-Learning: An Ocean Eddy Tracking Application

  • James Faghmous
  • Muhammed Uluyol
  • Luke Styles
  • Matthew Le
  • Varun Mithal
  • Shyam Boriah
  • Vipin Kumar

Mesoscale ocean eddies transport heat, salt, energy, and nutrients across oceans. As a result, accurately identifying and tracking such phenomena are crucial for understanding ocean dynamics and marine ecosystem sustainability. Traditionally, ocean eddies are monitored through two phases: identification and tracking. A major challenge for such an approach is that the tracking phase is dependent on the performance of the identification scheme, which can be susceptible to noise and sampling errors. In this paper, we focus on tracking, and introduce the concept of multiple hypothesis assignment (MHA), which extends traditional multiple hypothesis tracking for cases where the features tracked are noisy or uncertain. Under this scheme, features are assigned to multiple potential tracks, and the final assignment is deferred until more data are available to make a relatively unambiguous decision. Unlike the most widely used methods in the eddy tracking literature, MHA uses contextual spatio-temporal information to take corrective measures autonomously on the detection step a posteriori and performs significantly better in the presence of noise. This study is also the first to empirically analyze the relative robustness of eddy tracking algorithms.

AAAI Conference 2012 Conference Paper

A Novel and Scalable Spatio-Temporal Technique for Ocean Eddy Monitoring

  • James Faghmous
  • Yashu Chamber
  • Shyam Boriah
  • Frode Vikebø
  • Stefan Liess
  • Michel dos Santos Mesquita
  • Vipin Kumar

Swirls of ocean currents known as ocean eddies are a crucial component of the ocean’s dynamics. In addition to dominating the ocean’s kinetic energy, eddies play a significant role in the transport of water, salt, heat, and nutrients. Therefore, understanding current and future eddy patterns is a central climate challenge to address future sustainability of marine ecosystems. The emergence of sea surface height observations from satellite radar altimeter has recently enabled researchers to track eddies at a global scale. The majority of studies that identify eddies from observational data employ highly parametrized connected component algorithms using expert filtered data, effectively making reproducibility and scalability challenging. In this paper, we frame the challenge of monitoring ocean eddies as an unsupervised learning problem. We present a novel change detection algorithm that automatically identifies and monitors eddies in sea surface height data based on heuristics derived from basic eddy properties. Our method is accurate, efficient, and scalable. To demonstrate its performance we analyze eddy activity in the Nordic Sea (60 − 80◦ N and 20◦ W − 20◦ E), an area that has received limited attention and has proven to be difficult to analyze using other methods.

IJCAI Conference 2011 Conference Paper

Classification of Emerging Extreme Event Tracks in Multivariate Spatio-Temporal Physical Systems Using Dynamic Network Structures: Application to Hurricane Track Prediction

  • Huseyin Sencan
  • Zhengzhang Chen
  • William Hendrix
  • Tatdow Pansombut
  • Frederick Semazzi
  • Alok Choudhary
  • Vipin Kumar
  • Anatoli V. Melechko

Understanding extreme events, such as hurricanes or forest fires, is of paramount importance because of their adverse impacts on human beings. Such events often propagate in space and time. Predicting-even a few days in advance-what locations will get affected by the event tracks could benefit our society in many ways. Arguably, simulations from "first principles, " where underlying physics-based models are described by a system of equations, provide least reliable predictions for variables characterizing the dynamics of these extreme events. Data-driven model building has been recently emerging as a complementary approach that could learn the relationships between historically observed or simulated multiple, spatio-temporal ancillary variables and the dynamic behavior of extreme events of interest. While promising, the methodology for predictive learning from such complex data is still in its infancy. In this paper, we propose a dynamic networks-based methodology for in-advance prediction of the dynamic tracks of emerging extreme events. By associating a network model of the system with the known tracks, our method is capable of learning the recurrent network motifs that could be used as discriminatory signatures for the event's behavioral class. When applied to classifying the behavior of the hurricane tracks at their early formation stages in Western Africa region, our method is able to predict whether hurricane tracks will hit the land of the North Atlantic region at least 10-15 days lead lag time in advance with more than 90% accuracy using 10-fold cross-validation. To the best of our knowledge, no comparable methodology exists for solving this problem using data-driven models.

TIST Journal 2011 Journal Article

Monitoring global forest cover using data mining

  • Varun Mithal
  • Ashish Garg
  • Shyam Boriah
  • Michael Steinbach
  • Vipin Kumar
  • Christopher Potter
  • Steven Klooster
  • Juan Carlos Castilla-Rubio

Forests are a critical component of the planet's ecosystem. Unfortunately, there has been significant degradation in forest cover over recent decades as a result of logging, conversion to crop, plantation, and pasture land, or disasters (natural or man made) such as forest fires, floods, and hurricanes. As a result, significant attention is being given to the sustainable use of forests. A key to effective forest management is quantifiable knowledge about changes in forest cover. This requires identification and characterization of changes and the discovery of the relationship between these changes and natural and anthropogenic variables. In this article, we present our preliminary efforts and achievements in addressing some of these tasks along with the challenges and opportunities that need to be addressed in the future. At a higher level, our goal is to provide an overview of the exciting opportunities and challenges in developing and applying data mining approaches to provide critical information for forest and land use management.

AAAI Conference 1988 Conference Paper

Parallel Best-First Search of State-Space Graphs: A Summary of Results

  • Vipin Kumar

This paper presents many different parallel formulations of the A*/Branch-and-Bound search algorithm. The parallel formulations primarily differ in the data structures used. Some formulations are suited only for shared-memory architectures, whereas others are suited for distributedmemory architectures as well. These parallel formulations have been implemented to solve the vertex cover problem and the TSP problem on the BBN Butterfly parallel processor. Using appropriate data structures, we are able to obtain fairly linear speedups for as many as 100 processors. We also discovered problem characteristics that make certain formulations more (or less) suitable for some search problems. Since the bestfirst search paradigm of A*/Branch-and-Bound is very commonly used, we expect these parallel formulations to be effective for a variety of problems. Concurrent and distributed priority queues used in these parallel formulations can be used in many parallel algorithms other than parallel A*/branch-and-bound.

AAAI Conference 1984 Conference Paper

A General Bottom-up Procedure for Searching And/Or Graphs

  • Vipin Kumar

This paper summarizes work on a general bottom- up procedure for searching AND/OR graphs which includes a number of procedures for searching AND/OR graphs, state-space graphs, and dynamic programming procedures as its special cases. The paper concludes with comments on the significance of this work in the context of the author’s unified approach to search procedures.

AIJ Journal 1984 Journal Article

General Branch and Bound, and its relation to A∗ and AO∗

  • Dana S. Nau
  • Vipin Kumar
  • Laveen Kanal

Branch and Bound (B&B) is a problem-solving technique which is widely used for various problems encountered in operations research and combinatorial mathematics. Various heuristic search procedures used in artificial intelligence (AI) are considered to be related to B&B procedures. However, in the absence of any generally accepted terminology for B&B procedures, there have been widely differing opinions regarding the relationships between these procedures and B&B. This paper presents a formulation of B&B general enough to include previous formulations as special cases, and shows how two well-known AI search procedures (A∗ and AO∗) are special cases of this general formulation.

AIJ Journal 1983 Journal Article

A general branch and bound formulation for understanding and synthesizing and/or tree search procedures

  • Vipin Kumar
  • Laveen N. Kanal

Citing the confusing statements in the AI literature concerning the relationship between branch and bound (B&B) and heuristic search procedures were present a simple and general formulation of B&B which should help dispel much of the confusion. We illustrate the utility of the formulation by showing that through it some apparently very different algorithms for searching And/Or trees reveal the specific nature of their similarities and differences. In addition to giving new insights into the relationships among some AI search algorithms, the general formulation also provides suggestions on how existing search procedures may be varied to obtain new algorithms.

v2026.09.13