Researchers introduce DischargeBench to evaluate LLMs as hospital discharge educators
A team of researchers presents DischargeBench, a persona‑grounded simulation that pits a candidate LLM educator against a Virtual Patient in multi‑turn discharge education dialogues. An Education Monitor Agent oversees patient realism without altering the educator, preserving the evaluation signal.
Key points
- DischargeBench simulates multi‑turn discharge education dialogues between LLMs and a Virtual Patient.
- The benchmark includes 477 cases spanning 24 ICD chapters with persona dimensions like literacy and personality.
- Evaluation across LLMs shows performance varies by ICD chapter and patient persona, revealing comprehension gaps.
The benchmark draws on 477 cases extracted from MIMIC‑IV, covering 24 ICD chapters and varying along four persona axes: personality, education level, health literacy, and past‑medical‑history recall. Each simulated session is scored on Conversation Quality, Topic Checklist adherence, Comprehension, and Factual Consistency by an LLM‑as‑a‑Judge that is aligned with physician annotations.
Results across closed‑ and open‑source LLMs show that aggregate scores mask clinically relevant differences. Certain ICD chapters and difficult patient personas reveal coverage failures, comprehension gaps, and lower source‑answer agreement. The authors argue that LLM evaluation for discharge education should prioritize patient understanding rather than solely text quality or answer accuracy.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
More in Research
All →- TatBLiMP benchmarks Tatar linguistic minimal pairs · 1 src
- Recursive language models generalize out of domain, study shows · 1 src
- Embed‑TTT improves rule induction in ARC‑like tasks, authors say · 1 src
- Researchers introduce SpecOpt for agentic molecule specificity optimization · 1 src
- researchers report agent-based hls with rtl refinement speeds chip design 2.6× · 1 src
Comments
via GitHub Discussions