DigestAI news desk

Cut through the AI noise.

Research

Researchers compose frozen RWKV and Pythia models via shared latent adapter

A new arXiv paper introduces NinaXander, a method that connects layers of two frozen language models from different architecture families using a single trained adapter. The study combines RWKV-4-Raven-7B, a recurrent model, with Tulu-Pythia-6.9b, a Transformer-based model. Once the adapter is trained, multiple composed models can be created by connecting at different layer boundaries without…

1 source primary source

Key points

  • NinaXander connects frozen RWKV and Pythia models via a shared latent adapter
  • Best composition cuts KV cache 84.4% with RWKV-level multiple-choice accuracy
  • No composed model matches Pythia accuracy; WikiText perplexity degrades sharply

The best configuration uses the first 5 layers of Pythia and the remaining 27 layers of RWKV. This setup reduces the Transformer key-value cache by 84.4% while maintaining accuracy comparable to RWKV alone on multiple-choice tasks. However, no composed model matches Pythia's accuracy, and language modeling performance drops sharply on WikiText, an out-of-domain corpus. The authors note that representation alignment was only achieved under favorable conditions — shared tokenizer, depth, and hidden width — and does not prove a general shared semantic space.

Read the original at arXiv cs.CL · by Takanori Kotama, Shun-ichiro Hayashi, Daichi Mukunoki, Tetsuya Hoshino, Takahiro Katagiri primary sourceOpen source ↗
Topics · follow one to build your own front page
RWKV-4-Raven-7BTulu-Pythia-6.9b

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Research

All →

Related stories