Arrow Research search

Author name cluster

Joshua Batson

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

5 papers
2 author rows

Possible papers

5

TMLR Journal 2025 Journal Article

Open Problems in Mechanistic Interpretability

  • Lee Sharkey
  • Bilal Chughtai
  • Joshua Batson
  • Jack Lindsey
  • Jeffrey Wu
  • Lucius Bushnaq
  • Nicholas Goldowsky-Dill
  • Stefan Heimersheim

Mechanistic interpretability aims to understand the computational mechanisms underlying neural networks' capabilities in order to accomplish concrete scientific and engineering goals. Progress in this field thus promises to provide greater assurance over AI system behavior and shed light on exciting scientific questions about the nature of intelligence. Despite recent progress toward these goals, there are many open problems in the field that require solutions before many scientific and practical benefits can be realized: Our methods require both conceptual and practical improvements to reveal deeper insights; we must figure out how best to apply our methods in pursuit of specific goals; and the field must grapple with socio-technical challenges that influence and are influenced by our work. This forward-facing review discusses the current frontier of mechanistic interpretability and the open problems that the field may benefit from prioritizing.

NeurIPS Conference 2024 Conference Paper

Many-shot Jailbreaking

  • Cem Anil
  • Esin Durmus
  • Nina Panickssery
  • Mrinank Sharma
  • Joe Benton
  • Sandipan Kundu
  • Joshua Batson
  • Meg Tong

We investigate a family of simple long-context attacks on large language models: prompting with hundreds of demonstrations of undesirable behavior. This attack is newly feasible with the larger context windows recently deployed by language model providers like Google DeepMind, OpenAI and Anthropic. We find that in diverse, realistic circumstances, the effectiveness of this attack follows a power law, up to hundreds of shots. We demonstrate the success of this attack on the most widely used state-of-the-art closed-weight models, and across various tasks. Our results suggest very long contexts present a rich new attack surface for LLMs.

YNIMG Journal 2023 Journal Article

Denoising of diffusion MRI in the cervical spinal cord – effects of denoising strategy and acquisition on intra-cord contrast, signal modeling, and feature conspicuity

  • Kurt G. Schilling
  • Shreyas Fadnavis
  • Joshua Batson
  • Mereze Visagie
  • Anna J.E. Combes
  • Samantha By
  • Colin D. McKnight
  • Francesca Bagnato

Quantitative diffusion MRI (dMRI) is a promising technique for evaluating the spinal cord in health and disease. However, low signal-to-noise ratio (SNR) can impede interpretation and quantification of these images. The purpose of this study is to evaluate several dMRI denoising approaches on their ability to improve the quality, reliability, and accuracy of quantitative diffusion MRI of the spinal cord. We evaluate three denoising approaches (Non-Local Means, Marchenko-Pastur PCA, and a newly proposed Patch2Self algorithm) and conduct five experiments to validate the denoising performance on clinical-quality and commonly-acquired dMRI acquisitions: 1) a phantom experiment to assess denoising error and bias; 2) a multi-vendor, multi-acquisition open experiment for both qualitative and quantitative evaluation of noise residuals; 3) a bootstrapping experiment to estimate uncertainty of parametric maps; 4) an assessment of spinal cord lesion conspicuity in a multiple sclerosis group; and 5) an evaluation of denoising for advanced parametric multi-compartment modeling. We find that all methods improve signal-to-noise ratio and conspicuity of MS lesions in individual diffusion weighted images (DWIs), but MPPCA and Patch2Self excel at improving the quality and intra-cord contrast of diffusion weighted images - removing signal fluctuations due to thermal noise while improving precision of estimation of diffusion parameters even with very few DWIs (i.e., 16-32) typical of clinical acquisitions. These denoising approaches hold promise for facilitating reliable diffusion observations and measurements in the spinal cord to investigate biological and pathological processes.

NeurIPS Conference 2020 Conference Paper

Patch2Self: Denoising Diffusion MRI with Self-Supervised Learning​

  • Shreyas Fadnavis
  • Joshua Batson
  • Eleftherios Garyfallidis

Diffusion-weighted magnetic resonance imaging (DWI) is the only non-invasive method for quantifying microstructure and reconstructing white-matter pathways in the living human brain. Fluctuations from multiple sources create significant noise in DWI data which must be suppressed before subsequent microstructure analysis. We introduce a self-supervised learning method for denoising DWI data, Patch2Self, which uses the entire volume to learn a full-rank locally linear denoiser for that volume. By taking advantage of the oversampled q-space of DWI data, Patch2Self can separate structure from noise without requiring an explicit model for either. We demonstrate the effectiveness of Patch2Self via quantitative and qualitative improvements in microstructure modeling, tracking (via fiber bundle coherency) and model estimation relative to other unsupervised methods on real and simulated data.

ICML Conference 2019 Conference Paper

Noise2Self: Blind Denoising by Self-Supervision

  • Joshua Batson
  • Loïc Royer

We propose a general framework for denoising high-dimensional measurements which requires no prior on the signal, no estimate of the noise, and no clean training data. The only assumption is that the noise exhibits statistical independence across different dimensions of the measurement, while the true signal exhibits some correlation. For a broad class of functions (“$\mathcal{J}$-invariant”), it is then possible to estimate the performance of a denoiser from noisy data alone. This allows us to calibrate $\mathcal{J}$-invariant versions of any parameterised denoising algorithm, from the single hyperparameter of a median filter to the millions of weights of a deep neural network. We demonstrate this on natural image and microscopy data, where we exploit noise independence between pixels, and on single-cell gene expression data, where we exploit independence between detections of individual molecules. This framework generalizes recent work on training neural nets from noisy images and on cross-validation for matrix factorization.

v2026.09.13