- Main
When more precision is worse: Do people recognize inadequate scene representations in concept-based explainable AI?
Abstract
Explainable artificial intelligence (XAI) aims to uncover flaws in an AI model's internal representations. Can it help people recognize an AI's inability to distinguish between relevant and irrelevant features? In this study, a simulated AI classified images of railway trespassers as dangerous or not. To explain its decisions, similar images from the dataset were shown. These concept images varied in three relevant features (i.e., distance to tracks, direction, and action) and in an irrelevant feature (i.e., scene background). A feature used by the AI is retained in the concept images, otherwise the images randomize over it (e.g., same distance, varied backgrounds). Participants rated the AI more favourably when it retained relevant features. For the irrelevant feature, they did not mind in general, and sometimes even preferred it to be retained. This suggests that people may not recognize it when an AI model relies on irrelevant features to make its decisions.