Safety Evaluation of a Generative AI Agent for Anxiety and Depression Symptoms

preprint OA: closed CC-BY-4.0
🔓 Open OA copy View at publisher

Abstract

Background: Generative AI could radically improve engagement with digital mental health interventions. However, its probabilistic nature and rapid model iteration raise safety challenges that conventional, one-off evaluation approaches cannot fully address. Objective: To evaluate the safety of a clinician-designed generative AI care agent for anxiety and depression symptoms using a scalable dual-methods evaluation paradigm: high-throughput in silico simulation paired with continuous real-world monitoring. An exploratory aim assessed preliminary within-person symptom reduction. Methods: The digital program delivered a structured cognitive behavioral skills training program through a constrained generative AI architecture governed by a multi-agent safety system combining synthetic high-risk scenario testing, automated harm detection, and clinician oversight. Safety was assessed through: (1) evaluation of 43,325 generated responses to synthetic patients with high- and low-risk profiles (Study One), and (2) a two-week prospective, single-arm study in US adults with moderate-to-severe anxiety and depression symptoms evaluating 12,040 responses to real-user interactions (Study Two, N=111). Results: In Study One (simulations), the prevalence of clinician-verified AI content safety risk was 0.01% (95% CI: 0.01%, 0.03%); no responses encouraged self-harm, used offensive language, or actively invalidated users. In Study Two (observational study), no AI-generated responses were identified as containing AI content safety risk (95% CI: 0.00%, 0.04%), and no serious adverse events occurred. More than 50% of users met responder criteria after a median of approximately 90 minutes of program use. Conclusions: Rapid-cycle in silico evaluation paired with continuous real-world monitoring offers a scalable safety paradigm that can keep pace with the stochastic nature of LLMs and ongoing model iteration. Applied to a constrained, clinician-designed generative AI care agent, this paradigm delivered safe, high-quality mental health support across more than 55,000 evaluated responses, with early signals of symptom reduction. Controlled trials are needed to establish effectiveness.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-20T11:00:21.680559+00:00
License: CC-BY-4.0