Arrow Research search

Author name cluster

Guillem Collell

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

2 papers
1 author row

Possible papers

2

AAAI Conference 2018 Conference Paper

Acquiring Common Sense Spatial Knowledge Through Implicit Spatial Templates

  • Guillem Collell
  • Luc Van Gool
  • Marie-Francine Moens

Spatial understanding is a fundamental problem with widereaching real-world applications. The representation of spatial knowledge is often modeled with spatial templates, i. e. , regions of acceptability of two objects under an explicit spatial relationship (e. g. , “on”, “below”, etc.). In contrast with prior work that restricts spatial templates to explicit spatial prepositions (e. g. , “glass on table”), here we extend this concept to implicit spatial language, i. e. , those relationships (generally actions) for which the spatial arrangement of the objects is only implicitly implied (e. g. , “man riding horse”). In contrast with explicit relationships, predicting spatial arrangements from implicit spatial language requires significant common sense spatial understanding. Here, we introduce the task of predicting spatial templates for two objects under a relationship, which can be seen as a spatial question-answering task with a (2D) continuous output (“where is the man w. r. t. a horse when the man is walking the horse? ”). We present two simple neural-based models that leverage annotated images and structured text to learn this task. The good performance of these models reveals that spatial locations are to a large extent predictable from implicit spatial language. Crucially, the models attain similar performance in a challenging generalized setting, where the object-relation-object combinations (e. g. ,“man walking dog”) have never been seen before. Next, we go one step further by presenting the models with unseen objects (e. g. , “dog”). In this scenario, we show that leveraging word embeddings enables the models to output accurate spatial predictions, proving that the models acquire solid common sense spatial knowledge allowing for such generalization.

AAAI Conference 2017 Conference Paper

Imagined Visual Representations as Multimodal Embeddings

  • Guillem Collell
  • Ted Zhang
  • Marie-Francine Moens

Language and vision provide complementary information. Integrating both modalities in a single multimodal representation is an unsolved problem with wide-reaching applications to both natural language processing and computer vision. In this paper, we present a simple and effective method that learns a language-to-vision mapping and uses its output visual predictions to build multimodal representations. In this sense, our method provides a cognitively plausible way of building representations, consistent with the inherently reconstructive and associative nature of human memory. Using seven benchmark concept similarity tests we show that the mapped (or imagined) vectors not only help to fuse multimodal information, but also outperform strong unimodal baselines and state-of-the-art multimodal methods, thus exhibiting more human-like judgments. Ultimately, the present work sheds light on fundamental questions of natural language understanding concerning the fusion of vision and language such as the plausibility of more associative and reconstructive approaches.

v2026.09.13