DigestAI news desk

Cut through the AI noise.

Research

Researchers release benchmark for AI in systematic review screening

A new paper on arXiv introduces a benchmark dataset and framework for evaluating large language models in systematic review screening. The dataset contains labeled entries to test how well LLMs classify article relevance, addressing the imbalance between included and excluded articles. The authors propose SRBench, a tool for prompt experimentation and result analysis, alongside a use case…

1 source primary source

Key points

  • New benchmark dataset and framework for AI-assisted systematic review screening published on arXiv
  • Tool called PromptSR supports prompt experimentation and result analysis for LLMs in screening tasks
  • Authors argue existing evaluation metrics may not account for class imbalance in screening datasets

The work aims to improve evaluation methods for AI-assisted screening, which is often slow and labor-intensive. Existing metrics may not accurately reflect performance on imbalanced datasets, the authors argue. The paper does not specify the number of entries or studies but highlights the need for better tools to support AI-driven research workflows.

Read the original at arXiv cs.CL · by Gauransh Kumar, Luciano Marchezan, Guillaume Genois, K\'evin Delcourt, Eugene Syriani primary sourceOpen source ↗

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories