{"version":1,"type":"story","url":"https://digestai.news/story/study-compares-on-device-ner-models-for-speed-cost-and-accuracy","json":"https://digestai.news/story/study-compares-on-device-ner-models-for-speed-cost-and-accuracy.json","markdown":"https://digestai.news/story/study-compares-on-device-ner-models-for-speed-cost-and-accuracy.md","slug":"study-compares-on-device-ner-models-for-speed-cost-and-accuracy","headline":"Study compares On-Device NER models for speed, cost and accuracy","summary":"A new arXiv paper tests nine named-entity recognition systems for on-device use, where latency and data privacy matter more than leaderboard scores. They evaluate accuracy, latency (milliseconds to seconds), and output validity on three datasets, including RSS-News, which lacked gold-standard labels. The team created silver-standard labels via an LLM judge panel and validated them against human-annotated gold data (strict F1 score of 0.95). Results show that while a 4B generative LLM leads in accuracy on clean newswire text, bidirectional encoders like GLiNER offer comparable performance at 1/9th to 1/24th the size, with faster inference and zero malformed outputs. Generative models, however, produce up to 27% invalid outputs on long inputs, a flaw that scales with model size rather than output budget. The study also examines GLiNER’s confidence calibration, finding it ranks correctness well (AUROC 0.76–0.86) but is overconfident (ECE 0.24–0.47), with temperature scaling mitigating this. Thresholding confidence yields modest accuracy gains, and cascading small-to-large models locally improves performance modestly but depends on the dataset. Confidence correlates with correctness but not novelty.","keyPoints":["Bidirectional encoders like GLiNER match generative LLMs in accuracy but run faster and produce no invalid outputs","Generative models (e.g., Qwen3-4B) lead in accuracy on clean text but fail 27% of long-input tasks","Confidence calibration in GLiNER improves correctness ranking but remains overconfident without adjustments"],"whyItMatters":"For developers deploying NER locally, this study clarifies trade-offs: smaller, faster encoders may suffice for most use cases, while generative models require larger budgets to avoid errors. The findings help prioritize models based on latency, cost, and output reliability—not just accuracy.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":[],"models":["spaCy","GLiNER","Qwen3","DeepSeek-R1"],"people":[]},"firstPublishedAt":"2026-10-02T04:00:00Z","updatedAt":"2026-10-02T04:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"arXiv cs.CL","title":"On-Device Named-Entity Recognition: A Deployability Study of Accuracy, Cost, Reliability, and Confidence","url":"https://arxiv.org/abs/2610.00007","publishedAt":"2026-10-02T04:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":{"title":"Advancing Legal NER Through On-Device Optimization","url":"https://digestai.news/thread/researchers-release-intlawner-dataset-for-international-law-ner","storyCount":2},"cite":{"text":"Digest AI, \"Study compares On-Device NER models for speed, cost and accuracy\", 2 October 2026, https://digestai.news/story/study-compares-on-device-ner-models-for-speed-cost-and-accuracy","publisher":"Digest AI","title":"Study compares On-Device NER models for speed, cost and accuracy","datePublished":"2026-10-02T04:00:00Z","url":"https://digestai.news/story/study-compares-on-device-ner-models-for-speed-cost-and-accuracy"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}