{"version":1,"type":"story","url":"https://digestai.news/story/researchers-test-fixed-token-codes-for-language-models-at-100b-token-s","json":"https://digestai.news/story/researchers-test-fixed-token-codes-for-language-models-at-100b-token-s.json","markdown":"https://digestai.news/story/researchers-test-fixed-token-codes-for-language-models-at-100b-token-s.md","slug":"researchers-test-fixed-token-codes-for-language-models-at-100b-token-s","headline":"Researchers test fixed token codes for language models at 100B-token scale","summary":"A new paper on arXiv explores whether language models need trainable input embeddings to function well. The authors trained three decoder-only models from scratch with identical tokenizers, backbones, and training recipes, each using a different input interface: a learned embedding table, canonical 16-bit token-ID codes, and a fixed invertible recoding over GF(2). All models were trained on a target budget of **100 billion prediction tokens** each.\n\nThe fixed-code models achieved strong performance: **52.40%** on HellaSwag, **70.51%** on PIQA, and **42.75%** on LAMBADA. While the learned-input model outperformed the fixed ones on some benchmarks, the results suggest fixed token codes can be viable without sacrificing capability. The fixed interfaces also reduced trainable parameters by **100.7 million**, yielding models with **1.7B parameters**—though the paper emphasizes this is not the main finding. The study aims to clarify whether token-specific input vectors are architecturally necessary or empirically useful.","keyPoints":["Three models trained on 100B tokens each, using learned embeddings, fixed 16-bit codes, and GF(2) recoding","Fixed-code models scored 52.40% on HellaSwag, 70.51% on PIQA, and 42.75% on LAMBADA","Learned embeddings outperformed fixed codes but proved they are not strictly required for model capability"],"whyItMatters":"If fixed token codes work as well as learned embeddings, it could simplify model architecture and reduce training costs without losing performance.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":[],"models":[],"people":[]},"firstPublishedAt":"2026-10-06T04:00:00Z","updatedAt":"2026-10-06T04:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"arXiv cs.CL","title":"Do Language Models Need a Trainable Input Embedding Table? Fixed Minimal Token Codes at 1.7B-Class Scale","url":"https://arxiv.org/abs/2610.04002","publishedAt":"2026-10-06T04:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"Researchers test fixed token codes for language models at 100B-token scale\", 6 October 2026, https://digestai.news/story/researchers-test-fixed-token-codes-for-language-models-at-100b-token-s","publisher":"Digest AI","title":"Researchers test fixed token codes for language models at 100B-token scale","datePublished":"2026-10-06T04:00:00Z","url":"https://digestai.news/story/researchers-test-fixed-token-codes-for-language-models-at-100b-token-s"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}