DigestAI news desk
Research updated

Weakly Supervised Framework Uses LLMs to Extract Dataset Mentions in Displacement Documents

A new weakly supervised pipeline lets researchers pull dataset references from forced‑displacement and fragile‑state documents without a large hand‑labeled corpus. First, a lightweight model trained on generic research text proposes candidate mentions in unlabeled humanitarian reports. Those candidates are then vetted by a frontier large‑language model, which validates, rejects, or adjusts…

1 source primary source

Key points

  • Lightweight model proposes dataset mentions; frontier LLM refines them in context
  • Fine‑tuned system attains 74.1% precision and 70.5% recall on 1,706‑passage benchmark
  • Framework works without any manually labeled training data for the domain

On an independent benchmark of 1,706 passages spanning academic papers, project briefs, and operational reports, the system reaches 74.1% precision and 70.5% recall at the mention level, with precision climbing to 89.5% on passages that actually contain dataset references. Passage‑level classification hits 88.2% accuracy and 88.6% specificity. The results show that domain‑specific supervision can be built cheaply, enabling broader analysis of data usage and gaps in the displacement data ecosystem.

Read the original at arXiv cs.CL · by Rafael Macalaba, Aivin V. Solatorio, Patrick Michael Brock, Olivier Dupriez primary source Open source ↗

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.

Comments

via GitHub Discussions

More in Research

All →

Related stories