Study finds token-level entropy blind in small language models, semantic entropy helps
A new arXiv paper investigates whether entropy‑based confidence signals can improve the accuracy of small language models (SLMs) with fewer than 3 billion parameters that run on consumer hardware. The authors evaluate seven approaches—including token‑level entropy early stopping, semantic entropy estimation, and uncertainty‑aware routing—to larger expert models across seven model pairs and five…
Key points
- Token-level entropy is near zero in 91% of dataset‑model combos, making it ineffective for confidence in SLMs under 3 B parameters.
- Semantic entropy, derived from multiple sampled outputs and clustering, provides a usable confidence signal.
- Routing uncertain queries to larger expert models improves accuracy up to +50 pp, with cross‑family routing averaging +22.0% gain.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
More in Research
All →- TatBLiMP benchmarks Tatar linguistic minimal pairs · 1 src
- Recursive language models generalize out of domain, study shows · 1 src
- reviser proposes cursor-based text generation · 1 src
- SAGE system raises grant review agreement to kappa 0.58, beating baseline · 1 src
- Researchers introduce SpecOpt for agentic molecule specificity optimization · 1 src
Comments
via GitHub Discussions