Arrow Research search
Back to NeSy

NeSy 2025

Explainable Zero-Shot Visual Question Answering via Logic-Based Reasoning

Conference Paper Accepted Paper Artificial Intelligence · Logic in Computer Science · Neurosymbolic Artificial Intelligence

Abstract

Visual Question Answering (VQA) is the task of answering natural language questions about images, which is a challenge for AI systems. To enhance adaptability and reduce training overhead, we address VQA in a zero-shot setting by leveraging pre-trained neural modules without additional fine-tuning. Our proposed hybrid neurosymbolic framework, whose capabilities are demonstrated on the challenging GQA dataset, integrates neural and symbolic components through logic-based reasoning via Answer-Set Programming. Specifically, our pipeline employs large language models for semantic parsing of input questions, followed by the generation of a scene graph that captures relevant visual content. Interpretable rules then operate on the symbolic representations of both the question and the scene graph to derive an answer. Our framework provides a key advantage: it enables full transparency into the reasoning process. Using an existing explanation tool, we illustrate how our method fosters trust by making decisions interpretable and facilitates error analysis when predictions are incorrect. Beyond explaining its own reasoning, our framework can also explain answers from more opaque models by integrating their answers into our system, enabling broader interpretability in VQA.

Authors

Keywords

No keywords are indexed for this paper.

Context

Venue
International Conference on Neurosymbolic Learning and Reasoning
Archive span
2007-2025
Indexed papers
258
Paper id
136359219089904878
v2026.09.13