Study shows Conformal factuality control improves claim support in Multi-Hop RAG
Researchers tested claim-level conformal factuality control in multi-hop retrieval-augmented generation using Llama 3.1 8B and GPT-4o-mini on HotpotQA, Natural Questions, and TriviaQA. At a 95% conformal target, the fraction of responses with fully supported claims rose to between 95.80% and 97.20%, up from 55.60%-76.03% without filtering. However, this improvement came with low claim retention:…
Key points
- Conformal factuality control increased fully supported claims to 95.80%-97.20% at 95% target
- Without filtering, supported claims ranged from 55.60% to 76.03%
- Claim retention was low: only 4.41%-31.09% of claims retained at 95% target
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- A Holistic Assessment of the Carbon Footprint of Noor, a Very Large Arabic Language Model · 1 src
- Korean legal study finds KLUE-BERT outperforms GPT models in sexual offense text classification · 1 src
- Researchers question human-derived bias measures for LLM evaluation · 1 src
- arXiv study finds reading LLM judges from first token overstates position bias · 1 src
- Study compares On-Device NER models for speed, cost and accuracy · 1 src
Comments
via GitHub Discussions