AAMAS Conference 2026 Conference Paper
Exploring Cognitive Bias Impact, Detection and Mitigation in Large Language Models
- Ana Gutiérrez-Mandingorra
- Stella Heras
- Javier Palanca
- Vicent Botti
Large Language Models have revolutionized a wide range of domains—including education, healthcare, law, and industry—by enabling the automation of complex tasks through advanced natural language understanding and text generation. However, theirwidespreaddeploymenthasraisedsignificantethical and practical concerns, particularly regarding the biases embedded in their outputs. While social biases in LLMs have been extensively examined across the literature, cognitive biases—systematic patterns of deviation from normative reasoning rooted in human cognition—remain comparatively underexplored. These biases pose a unique challenge, as they can be subtly introduced, for instance, through prompt design or inherited from training data. Therefore, the study of cognitive biases in LLMs represents an emerging and increasingly critical area of research. This work presents a structured investigation into the presence, detection, and mitigation of cognitive biases in LLMs. We propose a three-stage experimental strategy: (1) evaluating the influence of prompt-induced cognitive biases on model outputs, (2) exploring bias detection strategies based on Retrieval-Augmented Generation systems enhanced with cognitive theory knowledge and incorporating agent-based reasoning elements, and (3) mitigating bias effects through warning-based interventions. Our findings aim to contribute towards a better understanding of LLMs’ alignment with human cognition and offer a foundation for safer and more trustworthy AI systems.