Researchers release benchmark for AI in systematic review screening
A new paper on arXiv introduces a benchmark dataset and framework for evaluating large language models in systematic review screening. The dataset contains labeled entries to test how well LLMs classify article relevance, addressing the imbalance between included and excluded articles. The authors propose SRBench, a tool for prompt experimentation and result analysis, alongside a use case…
Key points
- New benchmark dataset and framework for AI-assisted systematic review screening published on arXiv
- Tool called PromptSR supports prompt experimentation and result analysis for LLMs in screening tasks
- Authors argue existing evaluation metrics may not account for class imbalance in screening datasets
The work aims to improve evaluation methods for AI-assisted screening, which is often slow and labor-intensive. Existing metrics may not accurately reflect performance on imbalanced datasets, the authors argue. The paper does not specify the number of entries or studies but highlights the need for better tools to support AI-driven research workflows.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Researchers propose attention-free model mixing with autoencoders for masked language tasks · 1 src
- Study identifies under 1% of BERT neurons driving AI-Text detection · 1 src
- Researchers release Atelier for cryoEM map analysis via hypernetworks · 1 src
- Researchers find intuitive prompts improve LLM social media simulation · 1 src
- Researchers introduce Benchy, a standardized language for AI task benchmarks · 1 src
Comments
via GitHub Discussions