GPT OSS120B selects correct kinship terms 90.67% of the time in Hindi, Tamil, Korean
The paper introduces a generation‑based benchmark for culturally specific kinship terms in Hindi, Tamil, and Korean, evaluating five open‑weight large language models (LLMs). It contrasts generation with a matched option‑supported selection baseline.
Key points
- GPT OSS120B selects correct kinship terms 90.67% of 75 valid cells.
- Llama 3.370B selects correct terms 77.92% of 75 valid cells.
- GLM-5.1 accuracy 72.29% on explicitly specified L3 prompts.
Results show GPT OSS120B selects the correct term in 90.67% of 75 valid cells but produces an accepted term in only 36.00% of attempts. Llama 3.370B achieves 77.92% correct selections and 24.24% accepted terms. On explicitly specified L3 prompts, GLM‑5.1 attains 72.29% accuracy, while Llama 3.370B drops to 24.24%. The study notes a paternal‑lineage advantage that varies by language.
These findings suggest that generation tasks expose gaps in LLMs’ cultural knowledge that multiple‑choice tests may mask, motivating the use of generation‑based evaluation alongside traditional selection methods.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Anthropic's Claude discovers new enzyme system ART resembling crispr · 10 src
- MIT researchers publish book on visual AI for urban studies · 1 src
- Research shows evidence quality boosts fact‑verification scores · 1 src
- EduBehaviors framework provides auditable labeling of dialogues, reaching 0.673 macro-F1 · 1 src
- Researchers propose LEGO framework for legal reasoning · 1 src
Comments
via GitHub Discussions