Arrow Research search

Author name cluster

James S. Duncan

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

6 papers
2 author rows

Possible papers

6

ICLR Conference 2022 Conference Paper

Surrogate Gap Minimization Improves Sharpness-Aware Training

  • Juntang Zhuang
  • Boqing Gong
  • Liangzhe Yuan
  • Yin Cui
  • Hartwig Adam
  • Nicha C. Dvornek
  • Sekhar Tatikonda
  • James S. Duncan

The recently proposed Sharpness-Aware Minimization (SAM) improves generalization by minimizing a perturbed loss defined as the maximum loss within a neighborhood in the parameter space. However, we show that both sharp and flat minima can have a low perturbed loss, implying that SAM does not always prefer flat minima. Instead, we define a surrogate gap, a measure equivalent to the dominant eigenvalue of Hessian at a local minimum when the radius of neighborhood (to derive the perturbed loss) is small. The surrogate gap is easy to compute and feasible for direct minimization during training. Based on the above observations, we propose Surrogate Gap Guided Sharpness-Aware Minimization (GSAM), a novel improvement over SAM with negligible computation overhead. Conceptually, GSAM consists of two steps: 1) a gradient descent like SAM to minimize the perturbed loss, and 2) an ascent step in the orthogonal direction (after gradient decomposition) to minimize the surrogate gap and yet not affect the perturbed loss. GSAM seeks a region with both small loss (by step 1) and low sharpness (by step 2), giving rise to a model with high generalization capabilities. Theoretically, we show the convergence of GSAM and provably better generalization than SAM.Empirically, GSAM consistently improves generalization (e.g., +3.2% over SAM and +5.4% over AdamW on ImageNet top-1 accuracy for ViT-B/32). Code is released at https://sites.google.com/view/gsam-iclr22/home

ICLR Conference 2021 Conference Paper

MALI: A memory efficient and reverse accurate integrator for Neural ODEs

  • Juntang Zhuang
  • Nicha C. Dvornek
  • Sekhar Tatikonda
  • James S. Duncan

Neural ordinary differential equations (Neural ODEs) are a new family of deep-learning models with continuous depth. However, the numerical estimation of the gradient in the continuous case is not well solved: existing implementations of the adjoint method suffer from inaccuracy in reverse-time trajectory, while the naive method and the adaptive checkpoint adjoint method (ACA) have a memory cost that grows with integration time. In this project, based on the asynchronous leapfrog (ALF) solver, we propose the Memory-efficient ALF Integrator (MALI), which has a constant memory cost $w.r.t$ integration time similar to the adjoint method, and guarantees accuracy in reverse-time trajectory (hence accuracy in gradient estimation). We validate MALI in various tasks: on image recognition tasks, to our knowledge, MALI is the first to enable feasible training of a Neural ODE on ImageNet and outperform a well-tuned ResNet, while existing methods fail due to either heavy memory burden or inaccuracy; for time series modeling, MALI significantly outperforms the adjoint method; and for continuous generative models, MALI achieves new state-of-the-art performance. We provide a pypi package: https://jzkay12.github.io/TorchDiffEqPack

ICML Conference 2020 Conference Paper

Adaptive Checkpoint Adjoint Method for Gradient Estimation in Neural ODE

  • Juntang Zhuang
  • Nicha C. Dvornek
  • Xiaoxiao Li
  • Sekhar Tatikonda
  • Xenophon Papademetris
  • James S. Duncan

The empirical performance of neural ordinary differential equations (NODEs) is significantly inferior to discrete-layer models on benchmark tasks (e. g. image classification). We demonstrate an explanation is the inaccuracy of existing gradient estimation methods: the adjoint method has numerical errors in reverse-mode integration; the naive method suffers from a redundantly deep computation graph. We propose the Adaptive Checkpoint Adjoint (ACA) method: ACA applies a trajectory checkpoint strategy which records the forward- mode trajectory as the reverse-mode trajectory to guarantee accuracy; ACA deletes redundant components for shallow computation graphs; and ACA supports adaptive solvers. On image classification tasks, compared with the adjoint and naive method, ACA achieves half the error rate in half the training time; NODE trained with ACA outperforms ResNet in both accuracy and test-retest reliability. On time-series modeling, ACA outperforms competing methods. Furthermore, NODE with ACA can incorporate physical knowledge to achieve better accuracy.

YNICL Journal 2015 Journal Article

An unbiased Bayesian approach to functional connectomics implicates social-communication networks in autism

  • Archana Venkataraman
  • James S. Duncan
  • Daniel Y.-J. Yang
  • Kevin A. Pelphrey

Resting-state functional magnetic resonance imaging (rsfMRI) studies reveal a complex pattern of hyper- and hypo-connectivity in children with autism spectrum disorder (ASD). Whereas rsfMRI findings tend to implicate the default mode network and subcortical areas in ASD, task fMRI and behavioral experiments point to social dysfunction as a unifying impairment of the disorder. Here, we leverage a novel Bayesian framework for whole-brain functional connectomics that aggregates population differences in connectivity to localize a subset of foci that are most affected by ASD. Our approach is entirely data-driven and does not impose spatial constraints on the region foci or dictate the trajectory of altered functional pathways. We apply our method to data from the openly shared Autism Brain Imaging Data Exchange (ABIDE) and pinpoint two intrinsic functional networks that distinguish ASD patients from typically developing controls. One network involves foci in the right temporal pole, left posterior cingulate cortex, left supramarginal gyrus, and left middle temporal gyrus. Automated decoding of this network by the Neurosynth meta-analytic database suggests high-level concepts of "language" and "comprehension" as the likely functional correlates. The second network consists of the left banks of the superior temporal sulcus, right posterior superior temporal sulcus extending into temporo-parietal junction, and right middle temporal gyrus. Associated functionality of these regions includes "social" and "person". The abnormal pathways emanating from the above foci indicate that ASD patients simultaneously exhibit reduced long-range or inter-hemispheric connectivity and increased short-range or intra-hemispheric connectivity. Our findings reveal new insights into ASD and highlight possible neural mechanisms of the disorder.

YNIMG Journal 2004 Journal Article

Geometric strategies for neuroanatomic analysis from MRI

  • James S. Duncan
  • Xenophon Papademetris
  • Jing Yang
  • Marcel Jackowski
  • Xiaolan Zeng
  • Lawrence H. Staib

In this paper, we describe ongoing work in the Image Processing and Analysis Group (IPAG) at Yale University specifically aimed at the analysis of structural information as represented within magnetic resonance images (MRI) of the human brain. Specifically, we will describe our applied mathematical approaches to the segmentation of cortical and subcortical structure, the analysis of white matter fiber tracks using diffusion tensor imaging (DTI), and the intersubject registration of neuroanatomical (aMRI) data sets. Many of our methods rally around the use of geometric constraints, statistical (MAP) estimation, and the use of level set evolution strategies. The analysis of gray matter structure and connecting white matter paths combined with the ability to bring all information into a common space via intersubject registration should provide us with a rich set of data to investigate structure and variation in the human brain in neuropsychiatric disorders, as well as provide a basis for current work in the development of integrated brain function–structure analysis.

IROS Conference 1991 Conference Paper

Noncooperative games for decentralized integration architectures in modular systems

  • H. Isil Bozma
  • James S. Duncan

Describes a novel integration architecture for decentralized decision making in modular intelligent sensor systems. In contrast to previous approaches, this framework preserves the decentralized and coexisting natures of the objectives and is analytically and computationally tractable. The starting point for this approach is noncooperative N-player game theory. >

v2026.09.13