Study reveals severe linguistic and cultural errors in LLM-generated Urdu stories
A new research paper from arXiv examines the reliability of multilingual large language models when generating content in low-resource languages, using Urdu as the primary case study. The researchers created a corpus of 93 stories using three contemporary models: GPT-5.1, Qwen-3-Max, and DeepSeek-3.1. These outputs were manually annotated using a nine-label taxonomy covering linguistic,…
Key points
- Researchers analyzed 93 Urdu stories generated by GPT-5.1, Qwen-3-Max, and DeepSeek-3.1.
- Models exhibited basic grammar errors, incoherence, and pervasive cultural shallowness in outputs.
- Few-shot prompting failed to resolve the identified cultural and contextual errors in the stories.
The analysis found that the models frequently committed basic grammatical and semantic errors. Beyond technical flaws, the generated stories exhibited a lack of coherence, unnatural repetition, and significant cultural shallowness. The study further tested whether few-shot prompting could mitigate these issues, finding that cultural and contextual errors largely persisted despite this intervention.
These results suggest that current LLMs are not yet reliable for open-ended content generation or information retrieval in low-resource languages. The findings highlight a critical gap between the multilingual capabilities advertised by AI developers and the actual quality of output for non-English, low-resource contexts.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
More in Research
All →- Guide: How AI Embeddings Encode Meaning into Numerical Vectors · 1 src
- LLM-Anchored Paralinguistic Boost for Alzheimer's Detection · 1 src
- New probability-wave framework links trader behavior to AGI architecture design · 1 src
- Cognitive Digital Twins: Self-Evolving Architectures · 1 src
- Linguistic Structure Enrichment Fails to Improve Text Coherence · 1 src
Comments
via GitHub Discussions