DigestAI news desk

AI news, digested. Every story with its sources, every 30 minutes.

Researchupdated

QVAC Genesis III: 191.43B-token synthetic STEM corpus improves small model performance

Researchers introduce QVAC Genesis III, a 191.43 billion‑token synthetic corpus focused on STEM education. The dataset spans 19 domains, multiple difficulty levels, and varied educational styles. It is generated with a dual‑generation strategy: a weak edge‑scale student model’s failures are turned into corrective explanations, while its successes are expanded into contrastive, option‑level…

1 source primary source

Key points

  • QVAC Genesis III contains 191.43 billion tokens across 19 STEM domains and multiple difficulty levels.
  • 1.7 B‑parameter models trained on it beat Cosmopedia‑v2 and Cosmo‑1B, up to +28.57 % on ARC‑E.
  • Evaluation uses an LLM‑as‑parser protocol, achieving up to 99.45 % valid answer rate.

Controlled from‑scratch ablations using 1.7 B‑parameter language models show consistent gains over models trained on the open‑source synthetic corpus Cosmopedia‑v2 and the publicly released Cosmo‑1B model. On the ARC, GPQA Diamond, and MMLU STEM benchmarks, the QVAC‑trained models achieve up to +28.57 % improvement on ARC‑E, +21.35 % on ARC‑C, and a Valid Answer Rate reaching 99.45 %. The results suggest the corpus delivers high per‑token learning value for edge‑AI and on‑device deployments where token budgets are tight.

Read the original atarXiv cs.AI · by Davide Vitabile, N. Ranjan, Akshay Nambiar, Kamal K. Gupta, Amril Nazir primary sourceOpen source ↗
Topics · follow one to build your own front page
QVAC Genesis IIICosmopedia-v2Cosmo-1B

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.

Comments

via GitHub Discussions

More in Research

All →

Related stories