Arrow Research search

Author name cluster

James Demmel

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

8 papers
2 author rows

Possible papers

8

NeurIPS Conference 2025 Conference Paper

StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model Training

  • Ziming Liu
  • Shaoyu Wang
  • Shenggan Cheng
  • Zhongkai Zhao
  • Kai Wang
  • Xuanlei Zhao
  • James Demmel
  • Yang You

Training Transformer models on long sequences in a distributed setting poses significant challenges in terms of efficiency and scalability. Current methods are either constrained by the number of attention heads or excessive communication overheads. To address this problem, we propose StarTrail, a multi-dimensional concentric distributed training system for long sequences, fostering an efficient communication paradigm and providing additional tuning flexibility for communication arrangements. Specifically, StarTrail introduces an extra parallel dimension and divides the peer-to-peer communication into sub-rings to substantially reduce communication volume and avoid bandwidth bottlenecks. Through comprehensive experiments across diverse hardware environments and on both Natural Language Processing (NLP) and Computer Vision (CV) tasks, we demonstrate that our approach significantly surpasses state-of-the-art methods that support Long sequence lengths, achieving performance improvements of up to 77. 12% on GPT-style models and up to 114. 33% on DiT (Diffusion Transformer) models without affecting the computations results.

NeurIPS Conference 2023 Conference Paper

Fast Exact Leverage Score Sampling from Khatri-Rao Products with Applications to Tensor Decomposition

  • Vivek Bharadwaj
  • Osman Asif Malik
  • Riley Murray
  • Laura Grigori
  • Aydin Buluc
  • James Demmel

We present a data structure to randomly sample rows from the Khatri-Rao product of several matrices according to the exact distribution of its leverage scores. Our proposed sampler draws each row in time logarithmic in the height of the Khatri-Rao product and quadratic in its column count, with persistent space overhead at most the size of the input matrices. As a result, it tractably draws samples even when the matrices forming the Khatri-Rao product have tens of millions of rows each. When used to sketch the linear least-squares problems arising in Candecomp / PARAFAC decomposition, our method achieves lower asymptotic complexity per solve than recent state-of-the-art methods. Experiments on billion-scale sparse tensors and synthetic data validate our theoretical claims, with our algorithm achieving higher accuracy than competing methods as the decomposition rank grows.

ICLR Conference 2020 Conference Paper

Large Batch Optimization for Deep Learning: Training BERT in 76 minutes

  • Yang You 0001
  • Jing Li
  • Sashank J. Reddi
  • Jonathan Hseu
  • Sanjiv Kumar
  • Srinadh Bhojanapalli
  • Xiaodan Song
  • James Demmel

Training large deep neural networks on massive datasets is computationally very challenging. There has been recent surge in interest in using large batch stochastic optimization methods to tackle this issue. The most prominent algorithm in this line of research is LARS, which by employing layerwise adaptive learning rates trains ResNet on ImageNet in a few minutes. However, LARS performs poorly for attention models like BERT, indicating that its performance gains are not consistent across tasks. In this paper, we first study a principled layerwise adaptation strategy to accelerate training of deep neural networks using large mini-batches. Using this strategy, we develop a new layerwise adaptive large batch optimization technique called LAMB; we then provide convergence analysis of LAMB as well as LARS, showing convergence to a stationary point in general nonconvex settings. Our empirical results demonstrate the superior performance of LAMB across various tasks such as BERT and ResNet-50 training with very little hyperparameter tuning. In particular, for BERT training, our optimizer enables use of very large batch sizes of 32868 without any degradation of performance. By increasing the batch size to the memory limit of a TPUv3 Pod, BERT training time can be reduced from 3 days to just 76 minutes.

NeurIPS Conference 2016 Conference Paper

Asynchronous Parallel Greedy Coordinate Descent

  • Yang You
  • Xiangru Lian
  • Ji Liu
  • Hsiang-Fu Yu
  • Inderjit Dhillon
  • James Demmel
  • Cho-Jui Hsieh

n this paper, we propose and study an Asynchronous parallel Greedy Coordinate Descent (Asy-GCD) algorithm for minimizing a smooth function with bounded constraints. At each iteration, workers asynchronously conduct greedy coordinate descent updates on a block of variables. In the first part of the paper, we analyze the theoretical behavior of Asy-GCD and prove a linear convergence rate. In the second part, we develop an efficient kernel SVM solver based on Asy-GCD in the shared memory multi-core setting. Since our algorithm is fully asynchronous---each core does not need to idle and wait for the other cores---the resulting algorithm enjoys good speedup and outperforms existing multi-core kernel SVM solvers including asynchronous stochastic coordinate descent and multi-core LIBSVM.

NeurIPS Conference 2003 Conference Paper

Iterative Scaled Trust-Region Learning in Krylov Subspaces via Pearlmutter's Implicit Sparse Hessian

  • Eiji Mizutani
  • James Demmel

The online incremental gradient (or backpropagation) algorithm is widely considered to be the fastest method for solving large-scale neural-network (NN) learning problems. In contrast, we show that an appropriately implemented iterative batch-mode (or block-mode) learning method can be much faster. For example, it is three times faster in the UCI letter classification problem (26 outputs, 16, 000 data items, 6, 066 parameters with a two-hidden-layer multilayer perceptron) and 353 times faster in a nonlinear regression problem arising in color recipe prediction (10 outputs, 1, 000 data items, 2, 210 parameters with a neuro-fuzzy modular network). The three principal innovative ingredients in our algorithm are the following: First, we use scaled trust-region regularization with inner-outer it- eration to solve the associated “overdetermined” nonlinear least squares problem, where the inner iteration performs a truncated (or inexact) Newton method. Second, we employ Pearlmutter’s implicit sparse Hessian matrix-vector multiply algorithm to con- struct the Krylov subspaces used to solve for the truncated New- ton update. Third, we exploit sparsity (for preconditioning) in the matrices resulting from the NNs having many outputs.

NeurIPS Conference 2000 Conference Paper

On Iterative Krylov-Dogleg Trust-Region Steps for Solving Neural Networks Nonlinear Least Squares Problems

  • Eiji Mizutani
  • James Demmel

This paper describes a method of dogleg trust-region steps, or re(cid: 173) stricted Levenberg-Marquardt steps, based on a projection pro(cid: 173) cess onto the Krylov subspaces for neural networks nonlinear least squares problems. In particular, the linear conjugate gradient (CG) method works as the inner iterative algorithm for solving the lin(cid: 173) earized Gauss-Newton normal equation, whereas the outer nonlin(cid: 173) ear algorithm repeatedly takes so-called "Krylov-dogleg" steps, re(cid: 173) lying only on matrix-vector multiplication without explicitly form(cid: 173) ing the Jacobian matrix or the Gauss-Newton model Hessian. That is, our iterative dogleg algorithm can reduce both operational counts and memory space by a factor of O(n) (the number of pa(cid: 173) rameters) in comparison with a direct linear-equation solver. This memory-less property is useful for large-scale problems.

ICRA Conference 1989 Conference Paper

Optimal three finger grasps

  • James Demmel
  • Gerardo Lafferriere

The authors address the problem of optimal force distribution among three point fingers holding a planar object. A scheme that reduces the nonlinear optimization problem to an easily solved generalized eigenvalue problem is proposed. This scheme generalizes and simplifies results of Z. Ji and B. Roth (1988). The generalizations include all possible geometric arrangements and extensions to three dimensions and to the case of variable coefficients of friction. For the two-dimensional case with constant coefficients of friction, it is proved that except for some special cases, the optimal grasping forces, in the sense of minimizing the dependence on friction, are those for which the angles with the corresponding normals are all equal (in absolute value). >

ICRA Conference 1988 Conference Paper

Theoretical and experimental studies using a multifinger planar manipulator

  • James Demmel
  • Gerardo Lafferriere
  • Jacob T. Schwartz
  • Micha Sharir

An approach to manipulation tasks involving dextrous hands is presented. This approach treats the gripped object as a virtual finger and hence reduces the description of the task objectives to target forces and torques at reference points on the object. Computations are carried out for the case of three fingers holding a planar object. The authors implemented these ideas on the Four Finger Manipulator at New York University. Experimental results are presented for door opening and wall following with a gripped tool. >

v2026.09.13