# InterviewPlayground pilot shows LLM recovers 88% clinical items but underreports safety

Digest AI · Research · published 2026-09-21T04:00:00Z

Canonical: https://digestai.news/story/interviewplayground-pilot-shows-llm-recovers-88-clinical-items-but-und

## Summary

Health systems that plan to deploy AI‑assisted psychiatric intake tools need a way to verify that those systems meet clinical quality standards. The authors propose a clinician‑grounded evaluation platform called InterviewPlayground, built around a memory‑augmented patient simulator that can generate open‑ended interview scenarios from expert‑authored vignettes.

In a pilot involving six clinicians who each performed a 25‑minute assessment using the simulator, the authors compared the clinicians’ performance to that of a GPT‑based large language model acting as an intake interviewer. The LLM extracted 88.0% of the clinically relevant items embedded in the vignettes, while clinicians captured 38.9%. However, the model also generated more clinical inferences that were not grounded in the interview (56.8% vs. 27.8%) and identified safety concerns less frequently (33.3% vs. 66.7%).

The results illustrate that while AI interviewers can retrieve a higher proportion of relevant clinical information, they may also miss or misinterpret safety signals, underscoring the need for systematic quality‑assurance processes before wide deployment.

## Key points

- InterviewPlayground is a memory‑augmented patient simulator for open‑ended AI psychiatric interviews.
- In a pilot with six clinicians, the GPT‑based LLM recovered 88.0% of clinically relevant items versus 38.9% for clinicians.
- The LLM made more unfounded clinical inferences (56.8% vs 27.8%) and identified safety concerns less often (33.3% vs 66.7%).

## Why it matters

Provides a systematic way for health systems to assess AI psychiatric intake tools, highlighting gaps in safety detection that could affect patient care.

## Sources

1. [Clinician-Grounded Quality Assurance for AI-Assisted Psychiatric Intake](https://arxiv.org/abs/2609.21149) (arXiv cs.AI, 2026-09-21, primary source)

## Cite

Digest AI, "InterviewPlayground pilot shows LLM recovers 88% clinical items but underreports safety", 21 September 2026, https://digestai.news/story/interviewplayground-pilot-shows-llm-recovers-88-clinical-items-but-und

---

Written by Digest AI's editorial model from the linked sources; the sources are the record. Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse
JSON: https://digestai.news/story/interviewplayground-pilot-shows-llm-recovers-88-clinical-items-but-und.json
