Arrow Research search
Back to AAAI

AAAI 2017

Sherlock: Scalable Fact Learning in Images

Conference Paper AAAI Technical Track: Vision Artificial Intelligence

Abstract

The human visual system is capable of learning an unbounded number of facts from images including not only objects but also their attributes, actions and interactions. Such uniform understanding of visual facts has not received enough attention. Existing visual recognition systems are typically modeled differently for each fact type such as objects, actions, and interactions. We propose a setting where all these facts can be modeled simultaneously with a capacity to understand an unbounded number of facts in a structured way. The training data comes as structured facts in images, including (1) objects (e. g. , ), (2) attributes (e. g. , ), (3) actions (e. g. , ), and (4) interactions (e. g. , ). Each fact has a language view (e. g. , ) and a visual view (an image). We show that learning visual facts in a structured way enables not only a uniform but also generalizable visual understanding. We propose and investigate recent and strong approaches from the multiview learning literature and also introduce a structured embedding model. We applied the investigated methods on several datasets that we augmented with structured facts and a large scale dataset of > 202, 000 facts and 814, 000 images. Our results show the advantage of relating facts by the structure by the proposed model compared to the baselines.

Authors

Keywords

No keywords are indexed for this paper.

Context

Venue
AAAI Conference on Artificial Intelligence
Archive span
1980-2026
Indexed papers
28718
Paper id
41071485656815903
v2026.09.13