QVAC Genesis III: 191.43B-token synthetic STEM corpus improves small model performance
Researchers introduce QVAC Genesis III, a 191.43 billion‑token synthetic corpus focused on STEM education. The dataset spans 19 domains, multiple difficulty levels, and varied educational styles. It is generated with a dual‑generation strategy: a weak edge‑scale student model’s failures are turned into corrective explanations, while its successes are expanded into contrastive, option‑level…
Key points
- QVAC Genesis III contains 191.43 billion tokens across 19 STEM domains and multiple difficulty levels.
- 1.7 B‑parameter models trained on it beat Cosmopedia‑v2 and Cosmo‑1B, up to +28.57 % on ARC‑E.
- Evaluation uses an LLM‑as‑parser protocol, achieving up to 99.45 % valid answer rate.
Controlled from‑scratch ablations using 1.7 B‑parameter language models show consistent gains over models trained on the open‑source synthetic corpus Cosmopedia‑v2 and the publicly released Cosmo‑1B model. On the ARC, GPQA Diamond, and MMLU STEM benchmarks, the QVAC‑trained models achieve up to +28.57 % improvement on ARC‑E, +21.35 % on ARC‑C, and a Valid Answer Rate reaching 99.45 %. The results suggest the corpus delivers high per‑token learning value for edge‑AI and on‑device deployments where token budgets are tight.
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
More in Research
All →- New framework optimizes LLM inference costs via adaptive model activation · 4 src
- Qwen3.5-4B outperforms larger LLMs on new user-side conflict benchmark · 1 src
- Neo-Classic benchmark evaluates linguistic-aesthetic reasoning in Classical Chinese poetry · 1 src
- Study finds trust and friction issues in major generative AI app reviews · 1 src
- Study finds PCA can detect stylistic axes in LLM activations without training · 1 src
Comments
via GitHub Discussions