Arrow Research search

Author name cluster

Jiaming Cui

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

9 papers
2 author rows

Possible papers

9

AAAI Conference 2026 Conference Paper

Bridging Public Health with Clinical Decisions from a Data Centric Perspective

  • Jiaming Cui

Public health and clinical decisions are intertwined. Public health crises place a high burden on healthcare facilities, forcing them to make decisions such as maintaining quality verses treating more people. Meanwhile, sub-optimal clinical decisions also cause downstream effects on communities. For ex- ample, discharging patients too early may increase disease spread. Motivated by this, we bring a data-centric perspective to bridge clinical decisions within the context of infectious diseases for public health. This work addresses multiple challenges arising from effectively utilizing rich clinical datasets and issues stemming from the complexity of disease spread dynamics in healthcare facilities. We will cover methods developed to address these challenges with better designed models to optimize disease surveillance and control policies and new techniques for end-to-end learning with mechanistic models. We will conclude by discussing emerging challenges and opportunities at the intersection of machine learning, scientific modeling, and clinical decision-making for computer scientists, epidemiologists, and computational biologists.

UAI Conference 2025 Conference Paper

DF 2: Distribution-Free Decision-Focused Learning

  • Lingkai Kong
  • Wenhao Mu
  • Jiaming Cui
  • Yuchen Zhuang
  • B. Aditya Prakash
  • Bo Dai 0001
  • Chao Zhang 0014

Decision-focused learning (DFL), which differentiates through the KKT conditions, has recently emerged as a powerful approach for predict-then-optimize problems. However, under probabilistic settings, DFL faces three major bottlenecks: model mismatch error, sample average approximation error, and gradient approximation error. Model mismatch error stems from the misalignment between the model’s parameterized predictive distribution and the true probability distribution. Sample average approximation error arises when using finite samples to approximate the expected optimization objective. Gradient approximation error occurs when the objectives are non-convex and KKT conditions cannot be directly applied. In this paper, we present DF$^2$-the first \textit{distribution-free} decision-focused learning method designed to mitigate these three bottlenecks. Rather than depending on a task-specific forecaster that requires precise model assumptions, our method directly learns the expected optimization function during training. To efficiently learn the function in a data-driven manner, we devise an attention-based model architecture inspired by the distribution-based parameterization of the expected objective. We evaluate DF$^2$ on two synthetic problems and three real-world problems, demonstrating the effectiveness of DF$^2$. Our code can be found at: https: //github. com/Lingkai-Kong/DF2.

ICRA Conference 2024 Conference Paper

SAGE-ICP: Semantic Information-Assisted ICP

  • Jiaming Cui
  • Jiming Chen
  • Liang Li

Robust and accurate pose estimation in unknown environments is an essential part of robotic applications. We focus on LiDAR-based point-to-point ICP combined with effective semantic information. This paper proposes a novel semantic information-assisted ICP method named SAGE-ICP, which leverages semantics in odometry. The semantic information for the whole scan is timely and efficiently extracted by a 3D convolution network, and these point-wise labels are deeply involved in every part of the registration, including semantic voxel downsampling, data association, adaptive local map, and dynamic vehicle removal. Unlike previous semantic-aided approaches, the proposed method can improve localization accuracy in large-scale scenes even if the semantic information has certain errors. Experimental evaluations on KITTI and KITTI-360 show that our method outperforms the baseline methods, and improves accuracy while maintaining real-time performance, i. e. , runs faster than the sensor frame rate.

NeurIPS Conference 2024 Conference Paper

Time-MMD: Multi-Domain Multimodal Dataset for Time Series Analysis

  • Haoxin Liu
  • Shangqing Xu
  • Zhiyuan Zhao
  • Lingkai Kong
  • Harshavardhan Kamarthi
  • Aditya B. Sasanur
  • Megha Sharma
  • Jiaming Cui

Time series data are ubiquitous across a wide range of real-world domains. Whilereal-world time series analysis (TSA) requires human experts to integrate numerical series data with multimodal domain-specific knowledge, most existing TSAmodels rely solely on numerical data, overlooking the significance of information beyond numerical series. This oversight is due to the untapped potentialof textual series data and the absence of a comprehensive, high-quality multimodal dataset. To overcome this obstacle, we introduce Time-MMD, the firstmulti-domain, multimodal time series dataset covering 9 primary data domains. Time-MMD ensures fine-grained modality alignment, eliminates data contamination, and provides high usability. Additionally, we develop MM-TSFlib, thefirst-cut multimodal time-series forecasting (TSF) library, seamlessly pipeliningmultimodal TSF evaluations based on Time-MMD for in-depth analyses. Extensiveexperiments conducted on Time-MMD through MM-TSFlib demonstrate significant performance enhancements by extending unimodal TSF to multimodality, evidenced by over 15% mean squared error reduction in general, and up to 40%in domains with rich textual data. More importantly, our datasets and libraryrevolutionize broader applications, impacts, research topics to advance TSA. Thedataset is available at https: //github. com/AdityaLab/Time-MMD.

ICML Conference 2023 Conference Paper

Autoregressive Diffusion Model for Graph Generation

  • Lingkai Kong
  • Jiaming Cui
  • Haotian Sun
  • Yuchen Zhuang
  • B. Aditya Prakash
  • Chao Zhang 0014

Diffusion-based graph generative models have recently obtained promising results for graph generation. However, existing diffusion-based graph generative models are mostly one-shot generative models that apply Gaussian diffusion in the dequantized adjacency matrix space. Such a strategy can suffer from difficulty in model training, slow sampling speed, and incapability of incorporating constraints. We propose an autoregressive diffusion model for graph generation. Unlike existing methods, we define a node-absorbing diffusion process that operates directly in the discrete graph space. For forward diffusion, we design a diffusion ordering network, which learns a data-dependent node absorbing ordering from graph topology. For reverse generation, we design a denoising network that uses the reverse node ordering to efficiently reconstruct the graph by predicting the node type of the new node and its edges with previously denoised nodes at a time. Based on the permutation invariance of graph, we show that the two networks can be jointly trained by optimizing a simple lower bound of data likelihood. Our experiments on six diverse generic graph datasets and two molecule datasets show that our model achieves better or comparable generation performance with previous state-of-the-art, and meanwhile enjoys fast generation speed.

AAAI Conference 2023 Conference Paper

Detecting Sources of Healthcare Associated Infections

  • Hankyu Jang
  • Andrew Fu
  • Jiaming Cui
  • Methun Kamruzzaman
  • B. Aditya Prakash
  • Anil Vullikanti
  • Bijaya Adhikari
  • Sriram V. Pemmaraju

Healthcare acquired infections (HAIs) (e.g., Methicillin-resistant Staphylococcus aureus infection) have complex transmission pathways, spreading not just via direct person-to-person contacts, but also via contaminated surfaces. Prior work in mathematical epidemiology has led to a class of models – which we call load sharing models – that provide a discrete-time, stochastic formalization of HAI-spread on temporal contact networks. The focus of this paper is the source detection problem for the load sharing model. The source detection problem has been studied extensively in SEIR type models, but this prior work does not apply to load sharing models. We show that a natural formulation of the source detection problem for the load sharing model is computationally hard, even to approximate. We then present two alternate formulations that are much more tractable. The tractability of our problems depends crucially on the submodularity of the expected number of infections as a function of the source set. Prior techniques for showing submodularity, such as the "live graph" technique are not applicable for the load sharing model and our key technical contribution is to use a more sophisticated "coupling" technique to show the submodularity result. We propose algorithms for our two problem formulations by extending existing algorithmic results from submodular optimization and combining these with an expectation propagation heuristic for the load sharing model that leads to orders-of-magnitude speedup. We present experimental results on temporal contact networks based on fine-grained EMR data from three different hospitals. Our results on synthetic outbreaks on these networks show that our algorithms outperform baselines by up to 5.97 times. Furthermore, case studies based on hospital outbreaks of Clostridioides difficile infection show that our algorithms identify clinically meaningful sources.

AAAI Conference 2023 Conference Paper

EINNs: Epidemiologically-Informed Neural Networks

  • Alexander Rodríguez
  • Jiaming Cui
  • Naren Ramakrishnan
  • Bijaya Adhikari
  • B. Aditya Prakash

We introduce EINNs, a framework crafted for epidemic forecasting that builds upon the theoretical grounds provided by mechanistic models as well as the data-driven expressibility afforded by AI models, and their capabilities to ingest heterogeneous information. Although neural forecasting models have been successful in multiple tasks, predictions well-correlated with epidemic trends and long-term predictions remain open challenges. Epidemiological ODE models contain mechanisms that can guide us in these two tasks; however, they have limited capability of ingesting data sources and modeling composite signals. Thus, we propose to leverage work in physics-informed neural networks to learn latent epidemic dynamics and transfer relevant knowledge to another neural network which ingests multiple data sources and has more appropriate inductive bias. In contrast with previous work, we do not assume the observability of complete dynamics and do not need to numerically solve the ODE equations during training. Our thorough experiments on all US states and HHS regions for COVID-19 and influenza forecasting showcase the clear benefits of our approach in both short-term and long-term forecasting as well as in learning the mechanistic dynamics over other non-trivial alternatives.

NeurIPS Conference 2022 Conference Paper

End-to-end Stochastic Optimization with Energy-based Model

  • Lingkai Kong
  • Jiaming Cui
  • Yuchen Zhuang
  • Rui Feng
  • B. Aditya Prakash
  • Chao Zhang

Decision-focused learning (DFL) was recently proposed for stochastic optimization problems that involve unknown parameters. By integrating predictive modeling with an implicitly differentiable optimization layer, DFL has shown superior performance to the standard two-stage predict-then-optimize pipeline. However, most existing DFL methods are only applicable to convex problems or a subset of nonconvex problems that can be easily relaxed to convex ones. Further, they can be inefficient in training due to the requirement of solving and differentiating through the optimization problem in every training iteration. We propose SO-EBM, a general and efficient DFL method for stochastic optimization using energy-based models. Instead of relying on KKT conditions to induce an implicit optimization layer, SO-EBM explicitly parameterizes the original optimization problem using a differentiable optimization layer based on energy functions. To better approximate the optimization landscape, we propose a coupled training objective that uses a maximum likelihood loss to capture the optimum location and a distribution-based regularizer to capture the overall energy landscape. Finally, we propose an efficient training procedure for SO-EBM with a self-normalized importance sampler based on a Gaussian mixture proposal. We evaluate SO-EBM in three applications: power scheduling, COVID-19 resource allocation, and non-convex adversarial security game, demonstrating the effectiveness and efficiency of SO-EBM.

AAAI Conference 2022 Conference Paper

Provable Sensor Sets for Epidemic Detection over Networks with Minimum Delay

  • Jack Heavey
  • Jiaming Cui
  • Chen Chen
  • B. Aditya Prakash
  • Anil Vullikanti

The efficient detection of outbreaks and other cascading phenomena is a fundamental problem in a number of domains, including disease spread, social networks, and infrastructure networks. In such settings, monitoring and testing a small group of pre-selected nodes from the susceptible population (i. e. , a sensor set) is often the preferred testing regime. We study the problem of selecting a sensor set that minimizes the delay in detection—we refer to this as the MinDelSS problem. Prior methods for minimizing the detection time rely on greedy algorithms using submodularity. We show that this approach can sometimes lead to a worse approximation for minimizing the detection time than desired. We also show that MinDelSS is hard to approximate within an O(n1−1/γ )factor for any constant γ ≥ 2 for a graph with n nodes. This instead motivates seeking a bicriteria approximations. We present the algorithm ROUNDSENSOR, which gives a rigorous worst case O(log n)-factor for the detection time, while violating the budget by a factor of O(log2 n). Our algorithm is based on the sample average approximation technique from stochastic optimization, combined with linear programming and rounding. We evaluate our algorithm on several networks, including hospital contact networks, which validates its effectiveness in real settings.

v2026.09.13