Arrow Research search

Author name cluster

Sophia Han

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

3 papers
1 author row

Possible papers

3

NeurIPS Conference 2025 Conference Paper

Creativity or Brute Force? Using Brainteasers as a Window into the Problem-Solving Abilities of Large Language Models

  • Sophia Han
  • Howard Dai
  • Stephen Xia
  • Grant Zhang
  • Chen Liu
  • Lichang Chen
  • Hoang H Nguyen
  • Hongyuan Mei

Accuracy remains a standard metric for evaluating AI systems, but it offers limited insight into how models arrive at their solutions. In this work, we introduce a benchmark based on brainteasers written in long narrative form to probe more deeply into the types of reasoning strategies that models use. Brainteasers are well-suited for this goal because they can be solved with multiple approaches, such as a few-step solution that uses a creative insight or a longer solution that uses more brute force. We investigate large language models (LLMs) across multiple layers of reasoning, focusing not only on correctness but also on the quality and creativity of their solutions. We investigate many aspects of the reasoning process: (1) semantic parsing of the brainteasers into precise mathematical competition style formats; (2) self-correcting solutions based on gold solutions; (3) producing step-by-step sketches of solutions; and (4) making use of hints. We find that LLMs are in many cases able to find creative, insightful solutions to brainteasers, suggesting that they capture some of the capacities needed to solve novel problems in creative ways. Nonetheless, there also remain situations where they rely on brute force despite the availability of more efficient, creative solutions, highlighting a potential direction for improvement in the reasoning abilities of LLMs.

NeurIPS Conference 2025 Conference Paper

Measuring what Matters: Construct Validity in Large Language Model Benchmarks

  • Andrew M. Bean
  • Ryan Othniel Kearns
  • Angelika Romanou
  • Franziska Sofia Hafner
  • Harry Mayne
  • Jan Batzner
  • Negar Foroutan Eghlidi
  • Chris Schmitz

Evaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstract and complex phenomena such as safety' and robustness' requires strong construct validity, that is, having measures that represent what matters to the phenomenon. With a team of 29 expert reviewers, we conduct a systematic review of 445 LLM benchmarks from leading conferences in natural language processing and machine learning. Across the reviewed articles, we find patterns related to the measured phenomena, tasks, and scoring metrics which undermine the validity of the resulting claims. To address these shortcomings, we provide eight key recommendations and detailed actionable guidance to researchers and practitioners in developing LLM benchmarks.

YNIMG Journal 2020 Journal Article

Isometric exercise facilitates attention to salient events in women via the noradrenergic system

  • Mara Mather
  • Ringo Huang
  • David Clewett
  • Shawn E. Nielsen
  • Ricardo Velasco
  • Kristie Tu
  • Sophia Han
  • Briana L. Kennedy

The locus coeruleus (LC) regulates attention via the release of norepinephrine (NE), with levels of tonic LC activity constraining the intensity of phasic LC responses. In the current fMRI study, we used isometric handgrip to modulate tonic LC-NE activity in older women and in young women with different hormone statuses during the time period immediately after the handgrip. During this post-handgrip time, an oddball detection task was used to probe how changes in tonic arousal influenced functional coordination between the LC and a right frontoparietal network that supports attentional selectivity. As expected, the frontoparietal network responded more to infrequent target and novel sounds than to frequent sounds. Across participants, greater LC-frontoparietal functional connectivity, pupil dilation, and faster oddball detection were all positively associated with LC MRI structural contrast from a neuromelanin-sensitive scan. Thus, LC structure was related to LC functional dynamics and attentional performance during the oddball task. We also found that handgrip influenced pupil and attentional processing during a subsequent oddball task. Handgrip decreased subsequent tonic pupil size, increased phasic pupil responses to oddball sounds, speeded oddball detection speed, and increased frontoparietal network activation, suggesting that inducing strong LC activity benefits attentional performance in the next few minutes, potentially due to reduced tonic LC activity. In addition, older women showed a similar benefit of handgrip on frontoparietal network activation as younger women, despite showing lower frontoparietal network activation overall. Together these findings suggest that a simple exercise may improve selective attention in healthy aging, at least for several minutes afterwards.

v2026.09.13