Arrow Research search

Author name cluster

Julia Schnabel

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

3 papers
1 author row

Possible papers

3

NeurIPS Conference 2025 Conference Paper

NOVA: A Benchmark for Rare Anomaly Localization and Clinical Reasoning in Brain MRI

  • Cosmin Bercea
  • Jun Li
  • Philipp Raffler
  • Evamaria O. Riedel
  • Lena Schmitzer
  • Angela Kurz
  • Felix Bitzer
  • Paula Roßmüller

In many real-world applications, deployed models encounter inputs that differ from the data seen during training. Open-world recognition ensures that such systems remain robust as ever-emerging, previously _unknown_ categories appear and must be addressed without retraining. Foundation and vision-language models are pre-trained on large and diverse datasets with the expectation of broad generalization across domains, including medical imaging. However, benchmarking these models on test sets with only a few common outlier types silently collapses the evaluation back to a closed-set problem, masking failures on rare or truly novel conditions encountered in clinical use. We therefore present NOVA, a challenging, real-life _evaluation-only_ benchmark of $\sim$900 brain MRI scans that span 281 rare pathologies and heterogeneous acquisition protocols. Each case includes rich clinical narratives and double-blinded expert bounding-box annotations. Together, these enable joint assessment of anomaly localisation, visual captioning, and diagnostic reasoning. Because NOVA is never used for training, it serves as an _extreme_ stress-test of out-of-distribution generalisation: models must bridge a distribution gap both in sample appearance and in semantic space. Baseline results with leading vision-language models (GPT-4o, Gemini 2. 0 Flash, and Qwen2. 5-VL-72B) reveal substantial performance drops, with approximately a 65\% gap in localisation compared to natural-image benchmarks and 40\% and 20\% gaps in captioning and reasoning, respectively, compared to resident radiologists. Therefore, NOVA establishes a testbed for advancing models that can detect, localize, and reason about truly unknown anomalies.

YNIMG Journal 2009 Journal Article

An evaluation of four automatic methods of segmenting the subcortical structures in the brain

  • Kolawole Oluwole Babalola
  • Brian Patenaude
  • Paul Aljabar
  • Julia Schnabel
  • David Kennedy
  • William Crum
  • Stephen Smith
  • Tim Cootes

The automation of segmentation of subcortical structures in the brain is an active research area. We have comprehensively evaluated four novel methods of fully automated segmentation of subcortical structures using volumetric, spatial overlap and distance-based measures. Two methods are atlas-based — classifier fusion and labelling (CFL) and expectation–maximisation segmentation using a brain atlas (EMS), and two incorporate statistical models of shape and appearance — profile active appearance models (PAM) and Bayesian appearance models (BAM). Each method was applied to the segmentation of 18 subcortical structures in 270 subjects from a diverse pool varying in age, disease, sex and image acquisition parameters. Our results showed that all four methods perform on par with recently published methods. CFL performed better than the others according to all three classes of metrics. In summary over all structures, the ranking by the Dice coefficient was CFL, BAM, joint EMS and PAM. The Hausdorff distance ranked the methods as CFL, joint PAM and BAM, EMS, whilst percentage absolute volumetric difference ranked them as joint CFL and PAM, joint BAM and EMS. Furthermore, as we had four methods of performing segmentation, we investigated whether the results obtained by each method were more similar to each other than to the manual segmentations using Williams' Index. Reassuringly, the Williams' Index was close to 1 for most subjects (mean=1. 02, sd=0. 05), indicating better agreement of each method with the gold standard than with the other methods. However, 2% of cases (mainly amygdala and nucleus accumbens) had values outside 3 standard deviations of the mean.

v2026.09.13