{"version":1,"type":"story","url":"https://digestai.news/story/datalab-releases-omniextractbench-to-audit-pdf-extraction-accuracy","json":"https://digestai.news/story/datalab-releases-omniextractbench-to-audit-pdf-extraction-accuracy.json","markdown":"https://digestai.news/story/datalab-releases-omniextractbench-to-audit-pdf-extraction-accuracy.md","slug":"datalab-releases-omniextractbench-to-audit-pdf-extraction-accuracy","headline":"Datalab releases OmniExtractBench to audit PDF extraction accuracy","summary":"Datalab launched **OmniExtractBench**, an open benchmark for structured document extraction from PDFs into JSON. It pools **620 documents** from four existing datasets, including regulatory filings and synthetic tests, to measure how well systems fill JSON schemas. A deterministic scorer grades each value with one of six auditable verdicts—matched, misread, unfound, fabricated, invented item, or invented field—and normalizes formats like dates to avoid false mismatches.\n\nThe benchmark addresses flaws in current leaderboards, which Datalab claims favor vendors or obscure failures like broken harnesses. OmniExtractBench’s scorer uses content-based row pairing (via the Hungarian algorithm) to avoid penalizing tables for a single missed row. It also blocks score inflation by ignoring null fields, a tactic some systems exploit. Datalab tested **10 system configurations**, with its own **accurate mode** leading at **93.85% accuracy**, while others like **GPT-5.6-sol** and **LlamaExtract** struggled with unfound or fabricated values. The scorer is open-source under **Apache 2.0** (PyPI) and the dataset under **CC BY 4.0** (Hugging Face). Vendors must provide their own API keys and credits to rerun tests.","keyPoints":["OmniExtractBench pools 620 PDFs from four benchmarks, including 88 regulatory filings and 33 multi-page documents","Scorer assigns six verdicts per value and uses content-based row pairing to avoid table-scoring errors","Datalab’s accurate mode scores 93.85% accuracy; GPT-5.6-sol and LlamaExtract lag in precision or recall"],"whyItMatters":"A standardized, auditable benchmark could force extraction vendors to improve transparency and reduce bias in their claims. Current leaderboards let companies game scores with padding or opaque harnesses, but OmniExtractBench’s open rules may become a reference for developers and buyers.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":["Datalab","MarkTechPost"],"models":["GPT-5.6-sol","LlamaExtract","Gemini","Claude","Reducto deepextract v2","Mistral OCR 4.1"],"people":["Asif Razzaq"]},"firstPublishedAt":"2026-10-02T15:23:35Z","updatedAt":"2026-10-02T15:23:35Z","sourceCount":1,"hasPrimarySource":false,"sources":[{"outlet":"MarkTechPost","title":"Datalab Introduces OmniExtractBench to Fix Bias and Opacity in Extraction Benchmarks","url":"https://marktechpost.com/2026/10/02/datalab-introduces-omniextractbench-to-fix-bias-and-opacity-in-extraction-benchmarks","publishedAt":"2026-10-02T15:23:35Z","type":"press","primary":false,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"Datalab releases OmniExtractBench to audit PDF extraction accuracy\", 2 October 2026, https://digestai.news/story/datalab-releases-omniextractbench-to-audit-pdf-extraction-accuracy","publisher":"Digest AI","title":"Datalab releases OmniExtractBench to audit PDF extraction accuracy","datePublished":"2026-10-02T15:23:35Z","url":"https://digestai.news/story/datalab-releases-omniextractbench-to-audit-pdf-extraction-accuracy"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}