- Main
The Persistence of Self-Preference Bias in LLM Evaluations of Creativity
Abstract
Large language models (LLMs) are increasingly used to evaluate human performance in high-stakes decisions. Here, we investigate LLMs' ability to evaluate creativity, a hallmark of human cognition. Creativity is a robust predictor of academic and professional success. It is also shown that prioritizing creativity in high-stakes decisions such as college admissions mitigates traditional biases. We examined how GPT and Llama models evaluate creativity in authentic admission essays and AI-generated counterparts. LLM ratings were compared against an independent semantic divergence index that closely approximates human judgments of creativity. Results revealed strong self-preference bias in LLMs, inflating ratings for their own outputs and penalizing human-authored essays. We applied different mitigation strategies including zero- and few-shot prompting, and fine-tuning models on expert creativity ratings. Mitigation strategies reduced bias, with fine-tuning proving most effective, but could not fully eliminate it. These findings raise important concerns about using LLMs as judges of human creativity.