Arrow Research search

Author name cluster

Lora Aroyo

Possible papers associated with this exact author name in Arrow. This page groups case-insensitive exact name matches and is not a full identity disambiguation profile.

8 papers
1 author row

Possible papers

8

TMLR Journal 2026 Journal Article

Decoding Safety Feedback from Diverse Raters: A Data-driven Lens on Responsiveness to Severity

  • Pushkar Mishra
  • Charvi Rastogi
  • Stephen R Pfohl
  • Alicia Parrish
  • Tian Huey Teh
  • Roma Patel
  • Mark Diaz
  • Ding Wang

Ensuring the safety of Generative AI requires a nuanced understanding of pluralistic viewpoints. In this paper, we introduce a novel data-driven approach for analyzing ordinal safety ratings in pluralistic settings. Specifically, we address the challenge of interpreting nuanced differences in safety feedback from a diverse population expressed via ordinal scales (e.g., a Likert scale). We define non-parametric responsiveness metrics that quantify how raters convey broader distinctions and granular variations in the severity of safety violations. Leveraging publicly available datasets of pluralistic safety feedback as our case studies, we investigate how raters from different demographic groups use an ordinal scale to express their perceptions of the severity of violations. We apply our metrics across violation types, demonstrating their utility in extracting nuanced insights that are crucial for aligning AI systems reliably in multi-cultural contexts. We show that our approach can inform rater selection and feedback interpretation by capturing nuanced viewpoints across different demographic groups, hence improving the quality of pluralistic data collection and in turn contributing to more robust AI alignment.

NeurIPS Conference 2025 Conference Paper

Whose View of Safety? A Deep DIVE Dataset for Pluralistic Alignment of Text-to-Image Models

  • Charvi Rastogi
  • Tian Huey Teh
  • Pushkar Mishra
  • Roma Patel
  • Ding Wang
  • Mark Díaz
  • Alicia Parrish
  • Aida Mostafazadeh Davani

Current text-to-image (T2I) models often fail to account for diverse human experiences, leading to misaligned systems. We advocate for pluralism in AI alignment, where an AI understands and is steerable towards diverse, and often conflicting, human values. Our work provides three core contributions to achieve this in T2I models. First, we introduce a novel dataset for Diverse Intersectional Visual Evaluation (DIVE) -- the first multimodal dataset for pluralistic alignment. It enables deep alignment to diverse safety perspectives through a large pool of demographically intersectional human raters who provided extensive feedback across 1000 prompts, with high replication, capturing nuanced safety perceptions. Second, we empirically confirm demographics as a crucial proxy for diverse viewpoints in this domain, revealing significant, context-dependent differences in harm perception that diverge from conventional evaluations. Finally, we discuss implications for building aligned T2I models, including efficient data collection strategies, LLM judgment capabilities, and model steerability towards diverse perspectives. This research offers foundational tools for more equitable and aligned T2I systems. Content Warning: The paper includes sensitive content that may be harmful.

NeurIPS Conference 2023 Conference Paper

DataPerf: Benchmarks for Data-Centric AI Development

  • Mark Mazumder
  • Colby Banbury
  • Xiaozhe Yao
  • Bojan Karlaš
  • William Gaviria Rojas
  • Sudnya Diamos
  • Greg Diamos
  • Lynn He

Machine learning research has long focused on models rather than datasets, and prominent datasets are used for common ML tasks without regard to the breadth, difficulty, and faithfulness of the underlying problems. Neglecting the fundamental importance of data has given rise to inaccuracy, bias, and fragility in real-world applications, and research is hindered by saturation across existing dataset benchmarks. In response, we present DataPerf, a community-led benchmark suite for evaluating ML datasets and data-centric algorithms. We aim to foster innovation in data-centric AI through competition, comparability, and reproducibility. We enable the ML community to iterate on datasets, instead of just architectures, and we provide an open, online platform with multiple rounds of challenges to support this iterative development. The first iteration of DataPerf contains five benchmarks covering a wide spectrum of data-centric techniques, tasks, and modalities in vision, speech, acquisition, debugging, and diffusion prompting, and we support hosting new contributed benchmarks from the community. The benchmarks, online evaluation platform, and baseline implementations are open source, and the MLCommons Association will maintain DataPerf to ensure long-term benefits to academia and industry.

NeurIPS Conference 2023 Conference Paper

DICES Dataset: Diversity in Conversational AI Evaluation for Safety

  • Lora Aroyo
  • Alex Taylor
  • Mark Díaz
  • Christopher Homan
  • Alicia Parrish
  • Gregory Serapio-García
  • Vinodkumar Prabhakaran
  • Ding Wang

Machine learning approaches often require training and evaluation datasets with a clear separation between positive and negative examples. This requirement overly simplifies the natural subjectivity present in many tasks, and obscures the inherent diversity in human perceptions and opinions about many content items. Preserving the variance in content and diversity in human perceptions in datasets is often quite expensive and laborious. This is especially troubling when building safety datasets for conversational AI systems, as safety is socio-culturally situated in this context. To demonstrate this crucial aspect of conversational AI safety, and to facilitate in-depth model performance analyses, we introduce the DICES (Diversity In Conversational AI Evaluation for Safety) dataset that contains fine-grained demographics information about raters, high replication of ratings per item to ensure statistical power for analyses, and encodes rater votes as distributions across different demographics to allow for in-depth explorations of different aggregation strategies. The DICES dataset enables the observation and measurement of variance, ambiguity, and diversity in the context of safety for conversational AI. We further describe a set of metrics that show how rater diversity influences safety perception across different geographic regions, ethnicity groups, age groups, and genders. The goal of the DICES dataset is to be used as a shared resource and benchmark that respects diverse perspectives during safety evaluation of conversational AI systems.

TIST Journal 2013 Journal Article

Analyzing user behavior across social sharing environments

  • Pasquale De Meo
  • Emilio Ferrara
  • Fabian Abel
  • Lora Aroyo
  • Geert-Jan Houben

In this work we present an in-depth analysis of the user behaviors on different Social Sharing systems. We consider three popular platforms, Flickr, Delicious and StumbleUpon, and, by combining techniques from social network analysis with techniques from semantic analysis, we characterize the tagging behavior as well as the tendency to create friendship relationships of the users of these platforms. The aim of our investigation is to see if (and how) the features and goals of a given Social Sharing system reflect on the behavior of its users and, moreover, if there exists a correlation between the social and tagging behavior of the users. We report our findings in terms of the characteristics of user profiles according to three different dimensions: (i) intensity of user activities, (ii) tag-based characteristics of user profiles, and (iii) semantic characteristics of user profiles.

IS Journal 2009 Journal Article

Using AI to Access and Experience Cultural Heritage

  • Lynda Hardman
  • Lora Aroyo
  • J. van Ossenbruggen
  • Eero Hyvönen

Cultural heritage involves rich and highly heterogeneous collections of different people, organizations and collections. Preserved mainly by professionals it is challenging to convey this diversity of perspectives and information to the general public. Professionals also experience a great deal of obstacles archiving digital collections. This special issue presents current trends in employing AI and Web technologies to overcome such problems.

IS Journal 2007 Journal Article

Guest Editors' Introduction: Intelligent Educational Systems of the Present and Future

  • Lora Aroyo
  • Arthur Graesser
  • Lewis Johnson

Education in the 21st century faces an increasing number of challenges that require intelligent systems. The learning environments of today and tomorrow must handle distributed and dynamically changing content, a geographical dispersion of students and teachers, and generations of learners who spend hours a day interacting with multimillion-dollar multimedia environments. The new learning environments are moving beyond kindergarten through college classrooms to distance learning, lifelong education, and on-the-job training; they're entering the multifaceted universe of virtual worlds and ambient spaces. The articles in this special issue report novel, cutting-edge, intelligent learning environments. The systems not only have intelligent computational architectures but also are empirically tested on humans. Empirical research is of course essential to test system fidelity in delivering the intended pedagogy and to assess learning gains, usability, and learner satisfaction. This article is part of a special issue on intelligent educational systems.

v2026.09.13