Patrick Lange Papers

AAAI Conference 2023 Conference Paper

Towards Credible Human Evaluation of Open-Domain Dialog Systems Using Interactive Setup

Sijia Liu
Patrick Lange
Behnam Hedayatnia
Alexandros Papangelis
Di Jin
Andrew Wirth
Yang Liu
Dilek Hakkani-Tur

Evaluating open-domain conversation models has been an open challenge due to the open-ended nature of conversations. In addition to static evaluations, recent work has started to explore a variety of per-turn and per-dialog interactive evaluation mechanisms and provide advice on the best setup. In this work, we adopt the interactive evaluation framework and further apply to multiple models with a focus on per-turn evaluation techniques. Apart from the widely used setting where participants select the best response among different candidates at each turn, one more novel per-turn evaluation setting is adopted, where participants can select all appropriate responses with different fallback strategies to continue the conversation when no response is selected. We evaluate these settings based on sensitivity and consistency using four GPT2-based models that differ in model sizes or fine-tuning data. To better generalize to any model groups with no prior assumptions on their rankings and control evaluation costs for all setups, we also propose a methodology to estimate the required sample size given a minimum performance gap of interest before running most experiments. Our comprehensive human evaluation results shed light on how to conduct credible human evaluations of open domain dialog systems using the interactive setup, and suggest additional future directions.

PDF Details DOI

AAAI Conference 2022 Conference Paper

TEACh: Task-Driven Embodied Agents That Chat

Aishwarya Padmakumar
Jesse Thomason
Ayush Shrivastava
Patrick Lange
Anjali Narayan-Chen
Spandana Gella
Robinson Piramuthu
Gokhan Tur

Robots operating in human spaces must be able to engage in natural language interaction, both understanding and executing instructions, and using conversation to resolve ambiguity and correct mistakes. To study this, we introduce TEACh, a dataset of over 3, 000 human–human, interactive dialogues to complete household tasks in simulation. A Commander with access to oracle information about a task communicates in natural language with a Follower. The Follower navigates through and interacts with the environment to complete tasks varying in complexity from MAKE COFFEE to PREPARE BREAKFAST, asking questions and getting additional information from the Commander. We propose three benchmarks using TEACh to study embodied intelligence challenges, and we evaluate initial models’ abilities in dialogue understanding, language grounding, and task execution.

PDF Details

Possible papers

Towards Credible Human Evaluation of Open-Domain Dialog Systems Using Interactive Setup

TEACh: Task-Driven Embodied Agents That Chat