DigestAI news desk

Cut through the AI noise.

Research

Researchers release German court benchmark for LLM legal reasoning

A new arXiv paper introduces a sentence-level benchmark to test how well large language models can classify interpretive canons used by the German Federal Constitutional Court. The dataset consists of court decisions annotated at the sentence level, based on a legal theory from Larenz in the Savigny tradition. The authors evaluated four LLMs from three model families using both expert-written…

1 source primary source

Key points

  • New sentence-level benchmark uses German Federal Constitutional Court decisions
  • Four LLMs from three families scored 70.4-79.2 mean F1 across seven canons
  • GEPA-optimized prompts did not systematically beat expert hand-written prompts
Read the original at arXiv cs.CL · by Felix Ringe primary sourceOpen source ↗
Topics · follow one to build your own front page
LarenzSavigny

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories