{"version":1,"type":"story","url":"https://digestai.news/story/researchers-test-72-000-rag-combos-on-indian-government-documents","json":"https://digestai.news/story/researchers-test-72-000-rag-combos-on-indian-government-documents.json","markdown":"https://digestai.news/story/researchers-test-72-000-rag-combos-on-indian-government-documents.md","slug":"researchers-test-72-000-rag-combos-on-indian-government-documents","headline":"Researchers test 72,000 RAG combos on Indian government documents","summary":"The work evaluated eight hundred question instances across four distinct Indian central-government regulatory documents, with evidence strings validated against source text and manually reviewed for 10 percent of the sample.\n\nThe study found no single retriever family performed best across all documents, highlighting significant interaction effects between parser and chunker choices. MPNet-base emerged as a consistent underperformer, particularly for table-derived questions. The corpus reached a near-saturated evidence-preservation ceiling above 98 percent, suggesting retrieval differences stem primarily from ranking quality rather than information loss during ingestion. The researchers also released their evaluation harness, corpus manifest, and benchmark dataset for further study.","keyPoints":["Tested 3 parsers, 3 chunking strategies, and 5 dense embedding models plus a sparse BM25 baseline","Evaluated 800 question instances across four Indian government regulatory documents with validated evidence","Found no dominant retriever family, MPNet-base underperformed, and retrieval quality drives results over ingestion"],"whyItMatters":"The findings challenge assumptions about one-size-fits-all RAG setups, offering actionable insights for building systems that handle complex, structured documents like regulations. The released benchmark and tools could accelerate research in retrieval efficiency and ranking quality.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":[],"models":["MPNet-base"],"people":[]},"firstPublishedAt":"2026-09-29T04:00:00Z","updatedAt":"2026-09-29T04:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"arXiv cs.CL","title":"Parser, Chunking, and Embedding Interactions in Retrieval-Augmented Generation over Indian Government Regulatory Documents","url":"https://arxiv.org/abs/2609.31660","publishedAt":"2026-09-29T04:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"Researchers test 72,000 RAG combos on Indian government documents\", 29 September 2026, https://digestai.news/story/researchers-test-72-000-rag-combos-on-indian-government-documents","publisher":"Digest AI","title":"Researchers test 72,000 RAG combos on Indian government documents","datePublished":"2026-09-29T04:00:00Z","url":"https://digestai.news/story/researchers-test-72-000-rag-combos-on-indian-government-documents"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}