DigestAI news desk

Cut through the AI noise.

Research

Researchers test 72,000 RAG combos on Indian government documents

The work evaluated eight hundred question instances across four distinct Indian central-government regulatory documents, with evidence strings validated against source text and manually reviewed for 10 percent of the sample.

1 source primary source

Key points

  • Tested 3 parsers, 3 chunking strategies, and 5 dense embedding models plus a sparse BM25 baseline
  • Evaluated 800 question instances across four Indian government regulatory documents with validated evidence
  • Found no dominant retriever family, MPNet-base underperformed, and retrieval quality drives results over ingestion

The study found no single retriever family performed best across all documents, highlighting significant interaction effects between parser and chunker choices. MPNet-base emerged as a consistent underperformer, particularly for table-derived questions. The corpus reached a near-saturated evidence-preservation ceiling above 98 percent, suggesting retrieval differences stem primarily from ranking quality rather than information loss during ingestion. The researchers also released their evaluation harness, corpus manifest, and benchmark dataset for further study.

Read the original at arXiv cs.CL · by Shubham Kumar Singh primary sourceOpen source ↗
Topics · follow one to build your own front page
MPNet-base

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories