DigestAI news desk

Cut through the AI noise.

Generative AI & Models2 min read

DW-dev-UE releases Apex-2, a 3.87B MoE LLM trained on 86.5B tokens

DW-dev-UE published the Apex-2 language model, a mixture‑of‑experts system with 3.87 billion total parameters and 1.45 billion active ones. The model was pretrained from scratch on 86.5 billion tokens drawn from web, code, math and curated sources, then instruction‑tuned with a two‑stage SFT pipeline. The release includes full code, architecture write‑up and a training diary on GitHub, and the…

1 source primary source

Key points

  • Apex-2 has 3.87 B total, 1.45 B active parameters and was trained on 86.5 B tokens.
  • HumanEval+ score matches Qwen2.5‑1.5B (32.9) despite using far less pretraining data.
  • Model released under Apache 2.0, with code and checkpoint on GitHub (DW-dev-UE).

In benchmark tests Apex-2’s base model matches the performance of Qwen2.5‑1.5B (which used 18 trillion tokens) on HumanEval+ (32.9 vs 32.9) despite using only a fraction of the data. Gaps remain in knowledge (MMLU ~29%) and math, and the model is English‑centric with limited multilingual ability. A DPO stage was abandoned after it reduced scores on code and math tasks. The model supports a 4096‑token context window, uses the ChatML template, and can be loaded directly with vLLM.

Model page: Apex-2 →

Full story from huggingface.co · via Reddit AI communities primary sourceOpen source ↗

I trained a 3.87B MoE (1.45B active) from scratch on only 86.5B tokens

huggingface.co · 28 September 2026

Loading the full article…

This text was published by huggingface.co. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗

Coverage and discussion

1source
Topics · follow one to build your own front page
DW-dev-UEApex-2

The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.

Comments

via GitHub Discussions

More in Generative AI & Models

All →

Related stories