{"version":1,"type":"story","url":"https://digestai.news/story/mini-agi-is-a-continual-learning-byte-level-model-that-trains-on-a-sin","json":"https://digestai.news/story/mini-agi-is-a-continual-learning-byte-level-model-that-trains-on-a-sin.json","markdown":"https://digestai.news/story/mini-agi-is-a-continual-learning-byte-level-model-that-trains-on-a-sin.md","slug":"mini-agi-is-a-continual-learning-byte-level-model-that-trains-on-a-sin","headline":"mini-AGI is a continual learning byte-level model that trains on a single 8 GB GPU","summary":"mini-AGI is a byte‑level language model that builds its own architecture and trains from scratch on a single 8 GB VRAM GPU. It pages weight files from disk so the parameter count is limited by free disk space rather than VRAM, allowing the model to grow beyond the 8 GB card. The project targets PCs or laptops with at least an 8 GB GPU and currently runs as a toy‑level experiment, not a frontier system.\n\nAt the time of writing the model has about 530 M parameters (total 467.7 M reported) and uses a 256‑byte alphabet with a 4,096‑token context window. Only 32 experts are resident in VRAM at once – roughly 101 M parameters of the 468 M total – while the rest reside on disk. Training costs about 2.4 GFLOPs per byte, comparable to a dense 400 M byte‑level transformer. The run is still on its first pass over the corpus and is weeks away from completing, with early tests reading 524,000 characters of chess showing minimal forgetting.\n\nThe reference machine is an RTX 3070 Laptop GPU with 8 GB. The codebase includes support for adaptive depth up to 24 rows per character, PonderNet halting, and a routing system that selects top‑8 experts per block application. The project was assisted by the Claude Opus 5 model, and the repository is hosted at https://github.com/volotat/mini-AGI.","keyPoints":["mini-AGI runs on a single 8 GB VRAM GPU by paging weights from disk, allowing parameter count to exceed VRAM limits","Current size is about 530 M parameters (total 467.7 M) with a 256‑byte alphabet and 4,096‑token context window","Training cost is ~2.4 GFLOPs per byte, comparable to a dense 400 M byte‑level transformer, and the first corpus pass is weeks away"],"whyItMatters":"Shows continual learning and scalable model growth are possible on consumer‑grade hardware, opening personal AI development.","category":{"slug":"models","name":"Generative AI & Models","url":"https://digestai.news/category/models"},"entities":{"companies":["NVIDIA"],"models":["mini-AGI","Claude Opus 5"],"people":["Alexey Borsky"]},"firstPublishedAt":"2026-09-21T03:26:40Z","updatedAt":"2026-09-21T03:26:40Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"github.com","title":"mini-AGI - dynamically grown (530M params currently and growing) continual learning model trained from scratch on 8GB VRAM laptop from batch-1 stream of data.","url":"https://github.com/volotat/mini-AGI","publishedAt":"2026-09-21T03:26:40Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[{"site":"Reddit","url":"https://www.reddit.com/r/LocalLLaMA/comments/1wm1gab/miniagi_dynamically_grown_530m_params_currently/","points":null}],"thread":null,"cite":{"text":"Digest AI, \"mini-AGI is a continual learning byte-level model that trains on a single 8 GB GPU\", 21 September 2026, https://digestai.news/story/mini-agi-is-a-continual-learning-byte-level-model-that-trains-on-a-sin","publisher":"Digest AI","title":"mini-AGI is a continual learning byte-level model that trains on a single 8 GB GPU","datePublished":"2026-09-21T03:26:40Z","url":"https://digestai.news/story/mini-agi-is-a-continual-learning-byte-level-model-that-trains-on-a-sin"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}