Arrow Research search

Author name cluster

Jonathan D Cohen

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

4 papers
1 author row

Possible papers

4

NeurIPS Conference 2025 Conference Paper

Causal Head Gating: A Framework for Interpreting Roles of Attention Heads in Transformers

  • Andrew Nam
  • Henry Conklin
  • Yukang Yang
  • Tom Griffiths
  • Jonathan D Cohen
  • Sarah-Jane Leslie

We present causal head gating (CHG), a scalable method for interpreting the functional roles of attention heads in transformer models. CHG learns soft gates over heads and assigns them a causal taxonomy—facilitating, interfering, or irrelevant—based on their impact on task performance. Unlike prior approaches in mechanistic interpretability, which are hypothesis-driven and require prompt templates or target labels, CHG applies directly to any dataset using standard next-token prediction. We evaluate CHG across multiple large language models (LLMs) in the Llama 3 model family and diverse tasks, including syntax, commonsense, and mathematical reasoning, and show that CHG scores yield causal, not merely correlational, insight validated via ablation and causal mediation analyses. We also introduce contrastive CHG, a variant that isolates sub-circuits for specific task components. Our findings reveal that LLMs contain multiple sparse task-sufficient sub-circuits, that individual head roles depend on interactions with others (low modularity), and that instruction following and in-context learning rely on separable mechanisms.

NeurIPS Conference 2023 Conference Paper

Systematic Visual Reasoning through Object-Centric Relational Abstraction

  • Taylor Webb
  • Shanka Subhra Mondal
  • Jonathan D Cohen

Human visual reasoning is characterized by an ability to identify abstract patterns from only a small number of examples, and to systematically generalize those patterns to novel inputs. This capacity depends in large part on our ability to represent complex visual inputs in terms of both objects and relations. Recent work in computer vision has introduced models with the capacity to extract object-centric representations, leading to the ability to process multi-object visual inputs, but falling short of the systematic generalization displayed by human reasoning. Other recent models have employed inductive biases for relational abstraction to achieve systematic generalization of learned abstract rules, but have generally assumed the presence of object-focused inputs. Here, we combine these two approaches, introducing Object-Centric Relational Abstraction (OCRA), a model that extracts explicit representations of both objects and abstract relations, and achieves strong systematic generalization in tasks (including a novel dataset, CLEVR-ART, with greater visual complexity) involving complex visual displays.

NeurIPS Conference 2022 Conference Paper

Using natural language and program abstractions to instill human inductive biases in machines

  • Sreejan Kumar
  • Carlos G. Correa
  • Ishita Dasgupta
  • Raja Marjieh
  • Michael Y Hu
  • Robert Hawkins
  • Jonathan D Cohen
  • Nathaniel Daw

Strong inductive biases give humans the ability to quickly learn to perform a variety of tasks. Although meta-learning is a method to endow neural networks with useful inductive biases, agents trained by meta-learning may sometimes acquire very different strategies from humans. We show that co-training these agents on predicting representations from natural language task descriptions and programs induced to generate such tasks guides them toward more human-like inductive biases. Human-generated language descriptions and program induction models that add new learned primitives both contain abstract concepts that can compress description length. Co-training on these representations result in more human-like behavior in downstream meta-reinforcement learning agents than less abstract controls (synthetic language descriptions, program induction without learned primitives), suggesting that the abstraction supported by these representations is key.

YNIMG Journal 2004 Journal Article

The neural correlates of theory of mind within interpersonal interactions

  • James K Rilling
  • Alan G Sanfey
  • Jessica A Aronson
  • Leigh E Nystrom
  • Jonathan D Cohen

Tasks that engage a theory of mind seem to activate a consistent set of brain areas. In this study, we sought to determine whether two different interactive tasks, both of which involve receiving consequential feedback from social partners that can be used to infer intent, similarly engaged the putative theory of mind neural network. Participants were scanned using fMRI as they played the Ultimatum Game (UG) and the Prisoner's Dilemma Game (PDG) with both alleged human and computer partners who were outside the scanner. We observed a remarkable degree of overlap in brain areas that activated to partner decisions in the two games, including commonly observed theory of mind areas, as well as several brain areas that have not been reported previously and may relate to immersion of participants in real social interactions that have personally meaningful consequences. Although computer partners elicited activation in some of the same areas activated by human partners, most of these activations were stronger for human partners.

v2026.09.13