DigestAI news desk

AI news, digested. Every story with its sources, every 30 minutes.

Researchupdated

Neo-Classic benchmark evaluates linguistic-aesthetic reasoning in Classical Chinese poetry

Researchers present Neo-Classic, a new evaluation benchmark that uses out‑of‑sample, contemporary metrical poems written by experts to probe linguistic‑aesthetic reasoning in large language models. Unlike prior tests that rely on historical corpora, Neo-Classic includes reverse‑understanding probes that reduce the chance of simple retrieval.

1 source primary source

Key points

  • Neo-Classic benchmark uses out‑of‑sample contemporary metrical poetry to test linguistic‑aesthetic reasoning.
  • State‑of‑the‑art LLMs drop 20‑50% performance on new texts and score 0‑13% on discourse ordering.
  • Expert guidance lifts reasoning‑enhanced models to 36% but still trails human experts.

The study assesses three state‑of‑the‑art models—Qwen3-Max, Gemini-3-Pro, and DeepSeek-V3.2—across five behavioral probes. Results show a 20 to 50 percent performance gap when models move from historical to contemporary texts, and discourse‑level ordering tasks yield very low accuracy, ranging from 0 to 13 percent. Providing expert‑level guidance improves reasoning‑enhanced models to 36 percent, yet a sizable gap remains compared with human experts. The findings suggest current LLMs capture local formal patterns but struggle with global hierarchical planning required for robust linguistic‑aesthetic reasoning.

Read the original atarXiv cs.CL · by Han Zhang, Zihan Gu, Zhiyuan Wang, Tianyi Ma, Jiacheng Lu, Xinyan Zhang, Yuhao Wei, Cheng Hua primary sourceOpen source ↗
Topics · follow one to build your own front page
Qwen3-MaxGemini-3-ProDeepSeek-V3.2

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.

Comments

via GitHub Discussions

More in Research

All →

Related stories