Study compares On-Device NER models for speed, cost and accuracy
A new arXiv paper tests nine named-entity recognition systems for on-device use, where latency and data privacy matter more than leaderboard scores. They evaluate accuracy, latency (milliseconds to seconds), and output validity on three datasets, including RSS-News, which lacked gold-standard labels. The team created silver-standard labels via an LLM judge panel and validated them against…
Key points
- Bidirectional encoders like GLiNER match generative LLMs in accuracy but run faster and produce no invalid outputs
- Generative models (e.g., Qwen3-4B) lead in accuracy on clean text but fail 27% of long-input tasks
- Confidence calibration in GLiNER improves correctness ranking but remains overconfident without adjustments
The story so far
2 episodes →- Study compares On-Device NER models for speed, cost and accuracythis story
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- A Holistic Assessment of the Carbon Footprint of Noor, a Very Large Arabic Language Model · 1 src
- Korean legal study finds KLUE-BERT outperforms GPT models in sexual offense text classification · 1 src
- Researchers question human-derived bias measures for LLM evaluation · 1 src
- Researchers introduce Context language models that manage their own Context · 3 src
- Researchers release Build2SPARQL benchmark for text-to-SPARQL in building KGs · 1 src
Comments
via GitHub Discussions