Datalab releases OmniExtractBench to audit PDF extraction accuracy
Datalab launched OmniExtractBench, an open benchmark for structured document extraction from PDFs into JSON. It pools 620 documents from four existing datasets, including regulatory filings and synthetic tests, to measure how well systems fill JSON schemas. A deterministic scorer grades each value with one of six auditable verdicts—matched, misread, unfound, fabricated, invented item, or…
Key points
- OmniExtractBench pools 620 PDFs from four benchmarks, including 88 regulatory filings and 33 multi-page documents
- Scorer assigns six verdicts per value and uses content-based row pairing to avoid table-scoring errors
- Datalab’s accurate mode scores 93.85% accuracy; GPT-5.6-sol and LlamaExtract lag in precision or recall
The benchmark addresses flaws in current leaderboards, which Datalab claims favor vendors or obscure failures like broken harnesses. OmniExtractBench’s scorer uses content-based row pairing (via the Hungarian algorithm) to avoid penalizing tables for a single missed row. It also blocks score inflation by ignoring null fields, a tactic some systems exploit. Datalab tested 10 system configurations, with its own accurate mode leading at 93.85% accuracy, while others like GPT-5.6-sol and LlamaExtract struggled with unfound or fabricated values. The scorer is open-source under Apache 2.0 (PyPI) and the dataset under CC BY 4.0 (Hugging Face). Vendors must provide their own API keys and credits to rerun tests.
Datalab Introduces OmniExtractBench to Fix Bias and Opacity in Extraction Benchmarks
MarkTechPost · 2 October 2026Loading the full article…
This text was published by MarkTechPost and written by Asif Razzaq. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Trillium Labs launches open AI research nonprofit with $30M training budget · 1 src
- GitHub releases KernelOPT for agentic GPU kernel optimization · 1 src
- Karpathy shares how to make AI outputs like aircraft manuals · 1 src
- Allen Institute for AI open-sources AstaBrief 8B for faster scientific reports · 1 src
- What are world models and how do they predict environments? · 1 src
Comments
via GitHub Discussions