Arrow Research search

Author name cluster

Dilyara Bareeva

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

3 papers
1 author row

Possible papers

3

NeurIPS Conference 2025 Conference Paper

Beyond Scalars: Concept-Based Alignment Analysis in Vision Transformers

  • Johanna Vielhaben
  • Dilyara Bareeva
  • Jim Berend
  • Wojciech Samek
  • Nils Strodthoff

Measuring the alignment between representations lets us understand similarities between the feature spaces of different models, such as Vision Transformers trained under diverse paradigms. However, traditional measures for representational alignment yield only scalar values that obscure how these spaces agree in terms of learned features. To address this, we combine alignment analysis with concept discovery, allowing a fine-grained breakdown of alignment into individual concepts. This approach reveals both universal concepts across models and each representation’s internal concept structure. We introduce a new definition of concepts as non-linear manifolds, hypothesizing they better capture the geometry of the feature space. A sanity check demonstrates the advantage of this manifold-based definition over linear baselines for concept-based alignment. Finally, our alignment analysis of four different ViTs shows that increased supervision tends to reduce semantic organization in learned representations.

NeurIPS Conference 2025 Conference Paper

Manipulating Feature Visualizations with Gradient Slingshots

  • Dilyara Bareeva
  • Marina Höhne
  • Alexander Warnecke
  • Lukas Pirch
  • Klaus-Robert Müller
  • Konrad Rieck
  • Sebastian Lapuschkin
  • Kirill Bykov

Feature Visualization (FV) is a widely used technique for interpreting concepts learned by Deep Neural Networks (DNNs), which synthesizes input patterns that maximally activate a given feature. Despite its popularity, the trustworthiness of FV explanations has received limited attention. We introduce Gradient Slingshots, a novel method that enables FV manipulation without modifying model architecture or significantly degrading performance. By shaping new trajectories in off-distribution regions of a feature's activation landscape, we coerce the optimization process to converge to a predefined visualization. We evaluate our approach on several DNN architectures, demonstrating its ability to replace faithful FVs with arbitrary targets. These results expose a critical vulnerability: auditors relying solely on FV may accept entirely fabricated explanations. To mitigate this risk, we propose a straightforward defense and quantitatively demonstrate its effectiveness.

JMLR Journal 2023 Journal Article

Quantus: An Explainable AI Toolkit for Responsible Evaluation of Neural Network Explanations and Beyond

  • Anna Hedström
  • Leander Weber
  • Daniel Krakowczyk
  • Dilyara Bareeva
  • Franz Motzkus
  • Wojciech Samek
  • Sebastian Lapuschkin
  • Marina M.-C. Höhne

The evaluation of explanation methods is a research topic that has not yet been explored deeply, however, since explainability is supposed to strengthen trust in artificial intelligence, it is necessary to systematically review and compare explanation methods in order to confirm their correctness. Until now, no tool with focus on XAI evaluation exists that exhaustively and speedily allows researchers to evaluate the performance of explanations of neural network predictions. To increase transparency and reproducibility in the field, we therefore built Quantus—a comprehensive, evaluation toolkit in Python that includes a growing, well-organised collection of evaluation metrics and tutorials for evaluating explainable methods. The toolkit has been thoroughly tested and is available under an open-source license on PyPi (or on https://github.com/understandable-machine-intelligence-lab/Quantus/). [abs] [ pdf ][ bib ] [ code ] &copy JMLR 2023. ( edit, beta )

v2026.09.13