Skip to main content
eScholarship
Open Access Publications from the University of California

JUDICIOUS: Evaluating Robustness of Large Language Models in the Legal Realm

Creative Commons 'BY' version 4.0 license
Abstract

In recent years, the remarkable performance of large language models (LLMs) in tasks such as legal judgment prediction (LJP) has garnered widespread attention. An increasing number of LLMs have been successfully implemented to assist judges in performing various legal tasks. However, their robustness and reliability in complex judicial scenarios remain a subject of debate, particularly when confronted with real-world legal cases. Existing research often overlooks the systematic evaluation of these LLMs in terms of judicial fairness, robustness and other ethical considerations. To fill this gap, we propose a novel benchmark that integrates authentic legal cases to evaluate the robustness of LLMs in the legal judgment prediction (LJP) task. Our work establishes foundational safety standards for applying LLMs in the legal domain.