Thera-Turing Test: A Framework for Evaluating Mental Health Artificial Intelligence-Based Chatbots
preprint
OA: closed
Abstract
AI-based chatbots have emerged as popular tools for addressing mental health concerns. While research has primarily focused on usability or outcome measures, there is a scarcity of studies examining the quality of chatbot-delivered conversations. This article emphasizes the importance of evaluating chatbot content by drawing parallels with the rigorous training and evaluation process for human therapists. The Thera-Turing Test (TTT) proposes a comprehensive evaluation model that scrutinizes chatbot conversations independently, reducing potential bias. The TTT recommends that judges rate user conversations with the chatbot without knowing that AI-chatbots delivered those conversations. The components of the TTT include timing, conversation creation, specific vs. full assessment, selection criteria for conversations and the judges, and the assessment scale. The proposed TTT scale encompasses safety, text-based therapy specifics, confidentiality, ethical standards, common factors, specific factors, cultural awareness, referrals, and readiness for “practice” (or full-scale launch). Alternative approaches to chatbot evaluation are also explored, and potential critiques are addressed, highlighting the unique challenges of text-based therapy interactions. The Thera-Turing Test offers a promising avenue for ensuring the quality and effectiveness of mental health chatbots, contributing to the delivery of high-quality mental health care at scale.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.
Source provenance
- europepmc
- last seen: 2026-05-20T01:45:00.602351+00:00