# Datalab releases OmniExtractBench to audit PDF extraction accuracy

Digest AI · Research · published 2026-10-02T15:23:35Z

Canonical: https://digestai.news/story/datalab-releases-omniextractbench-to-audit-pdf-extraction-accuracy

## Summary

Datalab launched **OmniExtractBench**, an open benchmark for structured document extraction from PDFs into JSON. It pools **620 documents** from four existing datasets, including regulatory filings and synthetic tests, to measure how well systems fill JSON schemas. A deterministic scorer grades each value with one of six auditable verdicts—matched, misread, unfound, fabricated, invented item, or invented field—and normalizes formats like dates to avoid false mismatches.

The benchmark addresses flaws in current leaderboards, which Datalab claims favor vendors or obscure failures like broken harnesses. OmniExtractBench’s scorer uses content-based row pairing (via the Hungarian algorithm) to avoid penalizing tables for a single missed row. It also blocks score inflation by ignoring null fields, a tactic some systems exploit. Datalab tested **10 system configurations**, with its own **accurate mode** leading at **93.85% accuracy**, while others like **GPT-5.6-sol** and **LlamaExtract** struggled with unfound or fabricated values. The scorer is open-source under **Apache 2.0** (PyPI) and the dataset under **CC BY 4.0** (Hugging Face). Vendors must provide their own API keys and credits to rerun tests.

## Key points

- OmniExtractBench pools 620 PDFs from four benchmarks, including 88 regulatory filings and 33 multi-page documents
- Scorer assigns six verdicts per value and uses content-based row pairing to avoid table-scoring errors
- Datalab’s accurate mode scores 93.85% accuracy; GPT-5.6-sol and LlamaExtract lag in precision or recall

## Why it matters

A standardized, auditable benchmark could force extraction vendors to improve transparency and reduce bias in their claims. Current leaderboards let companies game scores with padding or opaque harnesses, but OmniExtractBench’s open rules may become a reference for developers and buyers.

## Sources

1. [Datalab Introduces OmniExtractBench to Fix Bias and Opacity in Extraction Benchmarks](https://marktechpost.com/2026/10/02/datalab-introduces-omniextractbench-to-fix-bias-and-opacity-in-extraction-benchmarks) (MarkTechPost, 2026-10-02)

## Cite

Digest AI, "Datalab releases OmniExtractBench to audit PDF extraction accuracy", 2 October 2026, https://digestai.news/story/datalab-releases-omniextractbench-to-audit-pdf-extraction-accuracy

---

Written by Digest AI's editorial model from the linked sources; the sources are the record. Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse
JSON: https://digestai.news/story/datalab-releases-omniextractbench-to-audit-pdf-extraction-accuracy.json
