{"version":1,"type":"story","url":"https://digestai.news/story/gpt-oss120b-selects-correct-kinship-terms-90-67-of-the-time-in-hindi-t","json":"https://digestai.news/story/gpt-oss120b-selects-correct-kinship-terms-90-67-of-the-time-in-hindi-t.json","markdown":"https://digestai.news/story/gpt-oss120b-selects-correct-kinship-terms-90-67-of-the-time-in-hindi-t.md","slug":"gpt-oss120b-selects-correct-kinship-terms-90-67-of-the-time-in-hindi-t","headline":"GPT OSS120B selects correct kinship terms 90.67% of the time in Hindi, Tamil, Korean","summary":"The paper introduces a generation‑based benchmark for culturally specific kinship terms in Hindi, Tamil, and Korean, evaluating five open‑weight large language models (LLMs). It contrasts generation with a matched option‑supported selection baseline.\n\nResults show GPT OSS120B selects the correct term in 90.67% of 75 valid cells but produces an accepted term in only 36.00% of attempts. Llama 3.370B achieves 77.92% correct selections and 24.24% accepted terms. On explicitly specified L3 prompts, GLM‑5.1 attains 72.29% accuracy, while Llama 3.370B drops to 24.24%. The study notes a paternal‑lineage advantage that varies by language.\n\nThese findings suggest that generation tasks expose gaps in LLMs’ cultural knowledge that multiple‑choice tests may mask, motivating the use of generation‑based evaluation alongside traditional selection methods.","keyPoints":["GPT OSS120B selects correct kinship terms 90.67% of 75 valid cells.","Llama 3.370B selects correct terms 77.92% of 75 valid cells.","GLM-5.1 accuracy 72.29% on explicitly specified L3 prompts."],"whyItMatters":"The benchmark reveals that large language models struggle with culturally specific kinship generation, highlighting gaps in their lexical knowledge and the need for generation‑based evaluation in multilingual AI.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":["OpenAI","Meta","Alibaba"],"models":["GPT OSS120B","Llama 3.370B","GLM-5.1"],"people":[]},"firstPublishedAt":"2026-09-24T04:00:00Z","updatedAt":"2026-09-24T04:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"arXiv cs.CL","title":"Recognized but Not Produced: A Generation Benchmark for Culturally Specific Kinship Terms","url":"https://arxiv.org/abs/2609.26942","publishedAt":"2026-09-24T04:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"GPT OSS120B selects correct kinship terms 90.67% of the time in Hindi, Tamil, Korean\", 24 September 2026, https://digestai.news/story/gpt-oss120b-selects-correct-kinship-terms-90-67-of-the-time-in-hindi-t","publisher":"Digest AI","title":"GPT OSS120B selects correct kinship terms 90.67% of the time in Hindi, Tamil, Korean","datePublished":"2026-09-24T04:00:00Z","url":"https://digestai.news/story/gpt-oss120b-selects-correct-kinship-terms-90-67-of-the-time-in-hindi-t"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}