- Main
The Turing Test Is Here to Stay
Abstract
The Turing test, first proposed by Alan Turing in 1950, has historically served as a benchmark for evaluating artificial intelligence (AI). However, since the release of ELIZA in 1966, and particularly with recent advancements in large language models (LLMs), AI has been claimed to pass the Turing test. Furthermore, criticism argues that the Turing test primarily assesses deceptive mimicry rather than genuine intelligence, prompting the continuous emergence of alternative benchmarks. This study argues against discarding the Turing test, proposing instead using more refined versions of it, for example, by interacting simultaneously with both an AI and human candidate to determine who is who, allowing a longer interaction duration, motivating participants with a bonus payment, "training'' the participants on past conversations, access to the Internet and other AIs, using experienced people as evaluators, etc. Through systematic experimentation using a web-based platform, we demonstrate that richer, contextually structured testing environments significantly enhance participants' ability to differentiate between AI and human interactions. Namely, we show that, while an off-the-shelf LLM can pass some Turing test environments, it fails to do so when faced with a more robust one. We further demonstrate that GPT-4.5, which previous work has shown to trick 70% of the human participants, fools less than 20% of the human participants in the most advanced Turing test we propose. Our findings highlight that the Turing test remains an important and effective method for evaluating AI, provided it continues to adapt as AI technology advances.