DigestAI news desk

Cut through the AI noise.

Research4 min read

Datalab releases OmniExtractBench to audit PDF extraction accuracy

Datalab launched OmniExtractBench, an open benchmark for structured document extraction from PDFs into JSON. It pools 620 documents from four existing datasets, including regulatory filings and synthetic tests, to measure how well systems fill JSON schemas. A deterministic scorer grades each value with one of six auditable verdicts—matched, misread, unfound, fabricated, invented item, or…

1 source

Key points

  • OmniExtractBench pools 620 PDFs from four benchmarks, including 88 regulatory filings and 33 multi-page documents
  • Scorer assigns six verdicts per value and uses content-based row pairing to avoid table-scoring errors
  • Datalab’s accurate mode scores 93.85% accuracy; GPT-5.6-sol and LlamaExtract lag in precision or recall

The benchmark addresses flaws in current leaderboards, which Datalab claims favor vendors or obscure failures like broken harnesses. OmniExtractBench’s scorer uses content-based row pairing (via the Hungarian algorithm) to avoid penalizing tables for a single missed row. It also blocks score inflation by ignoring null fields, a tactic some systems exploit. Datalab tested 10 system configurations, with its own accurate mode leading at 93.85% accuracy, while others like GPT-5.6-sol and LlamaExtract struggled with unfound or fabricated values. The scorer is open-source under Apache 2.0 (PyPI) and the dataset under CC BY 4.0 (Hugging Face). Vendors must provide their own API keys and credits to rerun tests.

Full story from MarkTechPost · by Asif RazzaqOpen source ↗

Datalab Introduces OmniExtractBench to Fix Bias and Opacity in Extraction Benchmarks

MarkTechPost · 2 October 2026

Loading the full article…

This text was published by MarkTechPost and written by Asif Razzaq. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗

Topics · follow one to build your own front page
DatalabMarkTechPostGPT-5.6-solLlamaExtractGeminiClaudeReducto deepextract v2Mistral OCR 4.1Asif Razzaq

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories