Researchers release German court benchmark for LLM legal reasoning
A new arXiv paper introduces a sentence-level benchmark to test how well large language models can classify interpretive canons used by the German Federal Constitutional Court. The dataset consists of court decisions annotated at the sentence level, based on a legal theory from Larenz in the Savigny tradition. The authors evaluated four LLMs from three model families using both expert-written…
Key points
- New sentence-level benchmark uses German Federal Constitutional Court decisions
- Four LLMs from three families scored 70.4-79.2 mean F1 across seven canons
- GEPA-optimized prompts did not systematically beat expert hand-written prompts
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Anthropic's Claude discovers new enzyme system ART resembling crispr · 10 src
- MIT researchers publish book on visual AI for urban studies · 1 src
- Research shows evidence quality boosts fact‑verification scores · 1 src
- AI transforms radiotherapy workflows while clinical adoption lags · 1 src
- SemiAnalysis releases ClusterMAX 3.0 rating 77 GPU cloud providers · 1 src
Comments
via GitHub Discussions