Neo-Classic benchmark evaluates linguistic-aesthetic reasoning in Classical Chinese poetry
Researchers present Neo-Classic, a new evaluation benchmark that uses out‑of‑sample, contemporary metrical poems written by experts to probe linguistic‑aesthetic reasoning in large language models. Unlike prior tests that rely on historical corpora, Neo-Classic includes reverse‑understanding probes that reduce the chance of simple retrieval.
Key points
- Neo-Classic benchmark uses out‑of‑sample contemporary metrical poetry to test linguistic‑aesthetic reasoning.
- State‑of‑the‑art LLMs drop 20‑50% performance on new texts and score 0‑13% on discourse ordering.
- Expert guidance lifts reasoning‑enhanced models to 36% but still trails human experts.
The study assesses three state‑of‑the‑art models—Qwen3-Max, Gemini-3-Pro, and DeepSeek-V3.2—across five behavioral probes. Results show a 20 to 50 percent performance gap when models move from historical to contemporary texts, and discourse‑level ordering tasks yield very low accuracy, ranging from 0 to 13 percent. Providing expert‑level guidance improves reasoning‑enhanced models to 36 percent, yet a sizable gap remains compared with human experts. The findings suggest current LLMs capture local formal patterns but struggle with global hierarchical planning required for robust linguistic‑aesthetic reasoning.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
More in Research
All →- VRR lets LLMs verify, repair and generate candidates, improving code and reasoning results · 1 src
- QVAC Genesis III: 191.43B-token synthetic STEM corpus improves small model performance · 1 src
- Blindspot Benchmark Tests Long-Horizon Safety of Tool-Using LLM Agents · 3 src
- Study finds decomposed-to-composed asymmetry in RL‑post‑trained language models · 1 src
- MAGS framework enables multi‑agent LLM coders to generate formally verified programs · 1 src
Comments
via GitHub Discussions