Arrow Research search

Author name cluster

Elena Simperl

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

16 papers
2 author rows

Possible papers

16

TIST Journal 2026 Journal Article

OntoChat Assistant for User Story Generation in Ontology Engineering

  • Yihang Zhao
  • Anelia Kurteva
  • Albert Meroño Peñuela
  • Elena Simperl

An ontology is a formal, explicit specification of a shared conceptualisation, which can be combined with problem-solving methods and reasoning functionality to develop high-quality technology and application systems efficiently. Ontology engineering (OE) typically involves extensive manual effort to elicit intended use cases (user stories) from users for the target ontology-based systems. Recent studies have demonstrated the positive potential of Large Language Model (LLM)-based conversational agents in supporting user story generation in OE. However, we argue that we are not leveraging LLM to its fullest potential by not supporting users in formulating effective prompts. To address this, we identify the prompt guidance users need during user story generation workflows by conducting a formative study (N = 10) using participatory prompting. We demonstrate its usefulness through the design and development of the OntoChat LLM-based system for OE, as well as a user evaluation with knowledge engineers (N = 24). To our knowledge, this is the first work to design and validate a prompt guidance framework that helps users leverage LLM to its fullest potential to generate effective requirements for ontology development. This advances how we interact with LLM for requirements elicitation.

NAI Journal 2025 Journal Article

Towards Interpretable Embeddings: Aligning Representations With Semantic Aspects

  • Nitisha Jain
  • Antoine Domingues
  • Adwait Baokar
  • Albert Meroño Peñuela
  • Elena Simperl

Knowledge graph embedding models (KGEMs) project entities and relations from knowledge graphs (KGs) into dense vector spaces, enabling tasks such as link prediction and recommendation systems. However, these embeddings typically suffer from a lack of interpretability and struggle to represent entity similarities in a way that is meaningful to humans. To address these challenges, we introduce InterpretE, a neuro-symbolic approach that generates interpretable vector spaces aligned with human-understandable entity aspects. By explicitly linking entity representations to their desired semantic aspects, InterpretE not only improves interpretability but also enhances the clustering of similar entities based on these aspects. Our experiments demonstrate that InterpretE effectively produces embeddings that are interpretable and improve the evaluation of semantic similarities, making it a valuable tool in explainable AI research by supporting transparent decision-making. By offering insights into how embeddings represent entities, InterpretE enables KGEMs to be used for semantic tasks in a more trustworthy and reliable manner.

NeSy Conference 2024 Conference Paper

Bringing Back Semantics to Knowledge Graph Embeddings: An Interpretability Approach

  • Antoine Domingues
  • Nitisha Jain
  • Albert Meroño-Peñuela
  • Elena Simperl

Abstract Knowledge Graph Embeddings Models project entities and relations from Knowledge Graphs into a vector space. Despite their widespread application, concerns persist about the ability of these models to capture entity similarity effectively. To address this, we introduce InterpretE, a novel neuro-symbolic approach to derive interpretable vector spaces with human-understandable dimensions in terms of the features of the entities. We demonstrate the efficacy of InterpretE in encapsulating desired semantic features, presenting evaluations both in the vector space as well as in terms of semantic similarity measurements.

NeurIPS Conference 2024 Conference Paper

Croissant: A Metadata Format for ML-Ready Datasets

  • Mubashara Akhtar
  • Omar Benjelloun
  • Costanza Conforti
  • Luca Foschini
  • Pieter Gijsbers
  • Joan Giner-Miguelez
  • Sujata Goswami
  • Nitisha Jain

Data is a critical resource for machine learning (ML), yet working with data remains a key friction point. This paper introduces Croissant, a metadata format for datasets that creates a shared representation across ML tools, frameworks, and platforms. Croissant makes datasets more discoverable, portable, and interoperable, thereby addressing significant challenges in ML data management. Croissant is already supported by several popular dataset repositories, spanning hundreds of thousands of datasets, enabling easy loading into the most commonly-used ML frameworks, regardless of where the data is stored. Our initial evaluation by human raters shows that Croissant metadata is readable, understandable, complete, yet concise.

TIST Journal 2023 Journal Article

Qrowdsmith: Enhancing Paid Microtask Crowdsourcing with Gamification and Furtherance Incentives

  • Eddy Maddalena
  • Luis-Daniel Ibáñez
  • Neal Reeves
  • Elena Simperl

Microtask crowdsourcing platforms are social intelligence systems in which volunteers, called crowdworkers, complete small, repetitive tasks in return for a small fee. Beyond payments, task requesters are considering non-monetary incentives such as points, badges, and other gamified elements to increase performance and improve crowdworker experience. In this article, we present Qrowdsmith, a platform for gamifying microtask crowdsourcing. To design the system, we explore empirically a range of gamified and financial incentives and analyse their impact on how efficient, effective, and reliable the results are. To maintain participation over time and save costs, we propose furtherance incentives, which are offered to crowdworkers to encourage additional contributions in addition to the fee agreed upfront. In a series of controlled experiments, we find that while gamification can work as furtherance incentives, it impacts negatively on crowdworkers’ performance, both in terms of the quantity and quality of work, as compared to a baseline where they can continue to contribute voluntarily. Gamified incentives are also less effective than paid bonus equivalents. Our results contribute to the understanding of how best to encourage engagement in microtask crowdsourcing activities and design better crowd intelligence systems.

TIST Journal 2020 Journal Article

Mapping Points of Interest Through Street View Imagery and Paid Crowdsourcing

  • Eddy Maddalena
  • Luis-Daniel Ibáñez
  • Elena Simperl

We present the Virtual City Explorer (VCE), an online crowdsourcing platform for the collection of rich geotagged information in urban environments. Compared to other volunteered geographic information approaches, which are constrained by the number and availability of mapping enthusiasts on the ground, the VCE uses digital street imagery to allow people to virtually explore a city from anywhere in the world, using a browser or a mobile phone. In addition, contributions in VCE are designed as paid microtasks—small jobs that can be carried out without any specific knowledge of the local area or previous mapping expertise in exchange for a fee. We tested the VCE in two cities to map points of interest (PoIs) in transport and mobility, using FigureEight to recruit participants. We were able to show that our platform enables crowdworkers to submit PoI location seamlessly, cover almost all of the tested areas, and discover several PoIs not reported by other approaches. This allows the VCE to complement existing approaches that leverage experts or grassroot communities.

JAIR Journal 2020 Journal Article

Point at the Triple: Generation of Text Summaries from Knowledge Base Triples

  • Pavlos Vougiouklis
  • Eddy Maddalena
  • Jonathon Hare
  • Elena Simperl

We investigate the problem of generating natural language summaries from knowledge base triples. Our approach is based on a pointer-generator network, which, in addition to generating regular words from a fixed target vocabulary, is able to verbalise triples in several ways. We undertake an automatic and a human evaluation on single and open-domain summaries generation tasks. Both show that our approach significantly outperforms other data-driven baselines.

IJCAI Conference 2020 Conference Paper

Point at the Triple: Generation of Text Summaries from Knowledge Base Triples (Extended Abstract)

  • Pavlos Vougiouklis
  • Eddy Maddalena
  • Jonathon Hare
  • Elena Simperl

We investigate the problem of generating natural language summaries from knowledge base triples. Our approach is based on a pointer-generator network, which, in addition to generating regular words from a fixed target vocabulary, is able to verbalise triples in several ways. We undertake an automatic and a human evaluation on single and open-domain summaries generation tasks. Both show that our approach significantly outperforms other data-driven baselines.

IS Journal 2017 Journal Article

Redecentralizing the Web with Distributed Ledgers

  • Luis-Daniel Ibanez
  • Elena Simperl
  • Fabien Gandon
  • Henry Story

The web was originally conceived as decentralized and universal, but during its popularization, its big value was built on centralized servers and nonuniversal access. A key element to redecentralize the web is to be able to generate trustable, secure, and accountable updates among autonomous participants without a central server. The authors believe that the marriage between distributed ledgers and linked data can provide this functionality and unlock the web's true potential. As a first step toward it, the authors propose a minimal vocabulary to describe and link distributed ledgers.

TIST Journal 2017 Journal Article

Social Incentives in Paid Collaborative Crowdsourcing

  • Oluwaseyi Feyisetan
  • Elena Simperl

Paid microtask crowdsourcing has traditionally been approached as an individual activity, with units of work created and completed independently by the members of the crowd. Other forms of crowdsourcing have, however, embraced more varied models, which allow for a greater level of participant interaction and collaboration. This article studies the feasibility and uptake of such an approach in the context of paid microtasks. Specifically, we compare engagement, task output, and task accuracy in a paired-worker model with the traditional, single-worker version. Our experiments indicate that collaboration leads to better accuracy and more output, which, in turn, translates into lower costs. We then explore the role of the social flow and social pressure generated by collaborating partners as sources of incentives for improved performance. We utilise a Bayesian method in conjunction with interface interaction behaviours to detect when one of the workers in a pair tries to exit the task. Upon this realisation, the other worker is presented with the opportunity to contact the exiting partner to stay: either for personal financial reasons (i.e., they have not completed enough tasks to qualify for a payment) or for fun (i.e., they are enjoying the task). The findings reveal that: (1) these socially motivated incentives can act as furtherance mechanisms to help workers attain and exceed their task requirements and produce better results than baseline collaborations; (2) microtask crowd workers are empathic (as opposed to selfish) agents, willing to go the extra mile to help their partners get paid; and, (3) social furtherance incentives create a win-win scenario for the requester and for the workers by helping more workers get paid by re-engaging them before they drop out.

IS Journal 2016 Journal Article

The Role of Data Science in Web Science

  • Christopher Phethean
  • Elena Simperl
  • Thanassis Tiropanis
  • Ramine Tinati
  • Wendy Hall

Web science relies on an interdisciplinary approach that seeks to go beyond what any one subject can say about the World Wide Web. By incorporating numerous disciplinary perspectives and relying heavily on domain knowledge and expertise, data science has emerged as an important new area that integrates statistics with computational knowledge, data collection, cleaning and processing, analysis methods, and visualization to produce actionable insights from big data. As a discipline to use within Web science research, data science offers significant opportunities for uncovering trends in large Web-based datasets. A Web science observatory exemplifies this relationship by offering an online platform of tools for carrying out Web science research, allowing users to carry out data science techniques to produce insights into Web science issues such as community development, online behavior, and information propagation. The authors outline the similarities and differences of these two growing subject areas to demonstrate the important relationship developing between them.

AAAI Conference 2015 Conference Paper

Towards Knowledge-Driven Annotation

  • Yassine Mrabet
  • Claire Gardent
  • Muriel Foulonneau
  • Elena Simperl
  • Eric Ras

While the Web of data is attracting increasing interest and rapidly growing in size, the major support of information on the surface Web are still multimedia documents. Semantic annotation of texts is one of the main processes that are intended to facilitate meaning-based information exchange between computational agents. However, such annotation faces several challenges such as the heterogeneity of natural language expressions, the heterogeneity of documents structure and context dependencies. While a broad range of annotation approaches rely mainly or partly on the target textual context to disambiguate the extracted entities, in this paper we present an approach that relies mainly on formalized-knowledge expressed in RDF datasets to categorize and disambiguate noun phrases. In the proposed method, we represent the reference knowledge bases as co-occurrence matrices and the disambiguation problem as a 0-1 Integer Linear Programming (ILP) problem. The proposed approach is unsupervised and can be ported to any RDF knowledge base. The system implementing this approach, called KODA, shows very promising results w. r. t. state-of-the-art annotation tools in cross-domain experimentations.

IS Journal 2014 Journal Article

The Web of Data: Bridging the Skills Gap

  • John Domingue
  • Mathieu d'Aquin
  • Elena Simperl
  • Alexander Mikroyannidis

With a projected six-figure skills gap looming in the US alone, here the authors share strategies and lessons learned regarding how to bridge the gap in training competent data scientists in the near future.

KER Journal 2013 Journal Article

Collaborative ontology engineering: a survey

  • Elena Simperl
  • Markus Luczak-Rösch

Abstract Building ontologies in a collaborative and increasingly community-driven fashion has become a central paradigm of modern ontology engineering. This understanding of ontologies and ontology engineering processes is the result of intensive theoretical and empirical research within the Semantic Web community, supported by technology developments such as Web 2.0. Over 6 years after the publication of the first methodology for collaborative ontology engineering, it is generally acknowledged that, in order to be useful, but also economically feasible, ontologies should be developed and maintained in a community-driven manner, with the help of fully-fledged environments providing dedicated support for collaboration and user participation. Wikis, and similar communication and collaboration platforms enabling ontology stakeholders to exchange ideas and discuss modeling decisions are probably the most important technological components of such environments. In addition, process-driven methodologies assist the ontology engineering team throughout the ontology life cycle, and provide empirically grounded best practices and guidelines for optimizing ontology development results in real-world projects. The goal of this article is to analyze the state of the art in the field of collaborative ontology engineering. We will survey several of the most outstanding methodologies, methods and techniques that have emerged in the last years, and present the most popular development environments, which can be utilized to carry out, or facilitate specific activities within the methodologies. A discussion of the open issues identified concludes the survey and provides a roadmap for future research and development in this lively and promising field.

KER Journal 2008 Journal Article

Tuplespace-based computing for the Semantic Web: a survey of the state-of-the-art

  • LYNDON J. B. NIXON
  • Elena Simperl
  • RETO KRUMMENACHER
  • FRANCISCO MARTIN-RECUERDA

Abstract Semantic technologies promise to solve many challenging problems of the present Web applications. As they achieve a feasible level of maturity, they become increasingly accepted in various business settings at enterprise level. By contrast, their usability in open environments such as the Web—with respect to issues such as scalability, dynamism and openness —still requires additional investigation. In particular, Semantic Web services have inherited the Web service communication model, which is primarily based on synchronous message exchange technology such as remote procedure call (RPC), thus being incompatible with the REST (REpresentational State Transfer) architectural model of the Web. Recent advances in the field of middleware propose ‘semantic tuplespace computing’ as an instrument for coping with this situation. Arguing that truly Web-compliant Web service communication should be based, analogously to the conventional Web, on shared access to persistently published data instead of message passing, space-based middleware introduces a coordination infrastructure by means of which services can exchange information in a time- and reference-decoupled manner. In this article, we introduce the most important approaches in this newly emerging field. Our objective is to analyze and compare the solutions proposed so far, thus giving an account of the current state-of-the-art, and identifying new directions of research and development.

IS Journal 2007 Journal Article

Argumentation-Based Ontology Engineering

  • Christoph Tempich
  • Elena Simperl
  • Markus Luczak
  • Rudi Studer
  • H. Sofia Pinto

This article applies the theory of argumentation to ontology engineering. Recent research in ontology engineering has highlighted the importance of controlled discussions for creating commonly agreed-on and widely accepted ontologies. The article analyzes how agreement is reached in the context of ontology development using rhetorical structure theory and identifies the most frequently used argument types. Case study-based investigations have shown that restricting the set of arguments participants use to express their positions can significantly facilitate reaching an agreement. The DILIGENT argumentation framework, consisting of a process, a formal model and a support tool, was built on the basis of these empirical findings. The formal model complies to the IBIS methodology, which was adapted to the ontology-specific requirements. It helps capture and record the design deliberations in ontology-engineering discussions, makes consensus building tasks more efficient, and provides detailed guidance for nonexperts. The authors successfully evaluated the framework in several case studies. This article is part of a special issue on argumentation technology.

v2026.09.13