Weakly Supervised Framework Uses LLMs to Extract Dataset Mentions in Displacement Documents
A new weakly supervised pipeline lets researchers pull dataset references from forced‑displacement and fragile‑state documents without a large hand‑labeled corpus. First, a lightweight model trained on generic research text proposes candidate mentions in unlabeled humanitarian reports. Those candidates are then vetted by a frontier large‑language model, which validates, rejects, or adjusts…
Key points
- Lightweight model proposes dataset mentions; frontier LLM refines them in context
- Fine‑tuned system attains 74.1% precision and 70.5% recall on 1,706‑passage benchmark
- Framework works without any manually labeled training data for the domain
On an independent benchmark of 1,706 passages spanning academic papers, project briefs, and operational reports, the system reaches 74.1% precision and 70.5% recall at the mention level, with precision climbing to 89.5% on passages that actually contain dataset references. Passage‑level classification hits 88.2% accuracy and 88.6% specificity. The results show that domain‑specific supervision can be built cheaply, enabling broader analysis of data usage and gaps in the displacement data ecosystem.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
More in Research
All →- R2VC Boosts Fact‑Checking Accuracy by 13.74% on FEVER Using Modular Retrieval and Calibration · 1 src
- Study Finds Deictic Ambiguity Can Undermine Draft‑Verify‑Revise LLM Pipelines · 1 src
- New GLARE model improves meeting continuation forecasting on MDFB benchmark · 1 src
- Chopthin-Consensus Power Sampling Boosts LLM Reasoning Accuracy Without Retraining · 1 src
- Context-Augmented KG Training Boosts Multi-Hop QA Accuracy on Disease Graphs · 1 src
Comments
via GitHub Discussions