{"version":1,"type":"story","url":"https://digestai.news/story/researchers-compose-frozen-rwkv-and-pythia-models-via-shared-latent-ad","json":"https://digestai.news/story/researchers-compose-frozen-rwkv-and-pythia-models-via-shared-latent-ad.json","markdown":"https://digestai.news/story/researchers-compose-frozen-rwkv-and-pythia-models-via-shared-latent-ad.md","slug":"researchers-compose-frozen-rwkv-and-pythia-models-via-shared-latent-ad","headline":"Researchers compose frozen RWKV and Pythia models via shared latent adapter","summary":"A new arXiv paper introduces NinaXander, a method that connects layers of two frozen language models from different architecture families using a single trained adapter. The study combines RWKV-4-Raven-7B, a recurrent model, with Tulu-Pythia-6.9b, a Transformer-based model. Once the adapter is trained, multiple composed models can be created by connecting at different layer boundaries without retraining.\n\nThe best configuration uses the first 5 layers of Pythia and the remaining 27 layers of RWKV. This setup reduces the Transformer key-value cache by 84.4% while maintaining accuracy comparable to RWKV alone on multiple-choice tasks. However, no composed model matches Pythia's accuracy, and language modeling performance drops sharply on WikiText, an out-of-domain corpus. The authors note that representation alignment was only achieved under favorable conditions — shared tokenizer, depth, and hidden width — and does not prove a general shared semantic space.","keyPoints":["NinaXander connects frozen RWKV and Pythia models via a shared latent adapter","Best composition cuts KV cache 84.4% with RWKV-level multiple-choice accuracy","No composed model matches Pythia accuracy; WikiText perplexity degrades sharply"],"whyItMatters":"Shows post-hoc model composition across architectures is feasible for inference efficiency, but reveals limits in preserving capabilities and generalization.","category":{"slug":"research","name":"Research","url":"https://digestai.news/category/research"},"entities":{"companies":[],"models":["RWKV-4-Raven-7B","Tulu-Pythia-6.9b"],"people":[]},"firstPublishedAt":"2026-10-01T04:00:00Z","updatedAt":"2026-10-01T04:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"arXiv cs.CL","title":"NinaXander: Feasibility and Limits of Composing Frozen Language Models Across Architecture Families via a Shared Latent Space","url":"https://arxiv.org/abs/2609.38261","publishedAt":"2026-10-01T04:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"Researchers compose frozen RWKV and Pythia models via shared latent adapter\", 1 October 2026, https://digestai.news/story/researchers-compose-frozen-rwkv-and-pythia-models-via-shared-latent-ad","publisher":"Digest AI","title":"Researchers compose frozen RWKV and Pythia models via shared latent adapter","datePublished":"2026-10-01T04:00:00Z","url":"https://digestai.news/story/researchers-compose-frozen-rwkv-and-pythia-models-via-shared-latent-ad"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}