Researchers test fixed token codes for language models at 100B-token scale
A new paper on arXiv explores whether language models need trainable input embeddings to function well. The authors trained three decoder-only models from scratch with identical tokenizers, backbones, and training recipes, each using a different input interface: a learned embedding table, canonical 16-bit token-ID codes, and a fixed invertible recoding over GF(2). All models were trained on a…
Key points
- Three models trained on 100B tokens each, using learned embeddings, fixed 16-bit codes, and GF(2) recoding
- Fixed-code models scored 52.40% on HellaSwag, 70.51% on PIQA, and 42.75% on LAMBADA
- Learned embeddings outperformed fixed codes but proved they are not strictly required for model capability
The fixed-code models achieved strong performance: 52.40% on HellaSwag, 70.51% on PIQA, and 42.75% on LAMBADA. While the learned-input model outperformed the fixed ones on some benchmarks, the results suggest fixed token codes can be viable without sacrificing capability. The fixed interfaces also reduced trainable parameters by 100.7 million, yielding models with 1.7B parameters—though the paper emphasizes this is not the main finding. The study aims to clarify whether token-specific input vectors are architecturally necessary or empirically useful.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Research
All →- Anthropic study: task understanding beats job title for AI success · 1 src
- Researchers propose FEM-ASM to separate storage, execution, and coordination in language models · 1 src
- Researchers propose surrogate log-probabilities to audit LLM agents · 3 src
- Behavioral history outperforms descriptions for LLM synthetic personas · 1 src
- Researchers introduce ROAR to unify AI-driven research system runs · 1 src
Comments
via GitHub Discussions