# DW-dev-UE releases Apex-2, a 3.87B MoE LLM trained on 86.5B tokens

Digest AI · Generative AI & Models · published 2026-09-28T00:00:00Z

Canonical: https://digestai.news/story/dw-dev-ue-releases-apex-2-a-3-87b-moe-llm-trained-on-86-5b-tokens

## Summary

DW-dev-UE published the Apex-2 language model, a mixture‑of‑experts system with 3.87 billion total parameters and 1.45 billion active ones. The model was pretrained from scratch on 86.5 billion tokens drawn from web, code, math and curated sources, then instruction‑tuned with a two‑stage SFT pipeline. The release includes full code, architecture write‑up and a training diary on GitHub, and the checkpoint is available under an Apache 2.0 license.

In benchmark tests Apex-2’s base model matches the performance of Qwen2.5‑1.5B (which used 18 trillion tokens) on HumanEval+ (32.9 vs 32.9) despite using only a fraction of the data. Gaps remain in knowledge (MMLU ~29%) and math, and the model is English‑centric with limited multilingual ability. A DPO stage was abandoned after it reduced scores on code and math tasks. The model supports a 4096‑token context window, uses the ChatML template, and can be loaded directly with vLLM.

## Key points

- Apex-2 has 3.87 B total, 1.45 B active parameters and was trained on 86.5 B tokens.
- HumanEval+ score matches Qwen2.5‑1.5B (32.9) despite using far less pretraining data.
- Model released under Apache 2.0, with code and checkpoint on GitHub (DW-dev-UE).

## Why it matters

Apex-2 shows that high‑quality code and instruction performance can be achieved with modest data, offering an open‑weight alternative for researchers and developers seeking a small‑scale MoE LLM.

## Sources

1. [I trained a 3.87B MoE (1.45B active) from scratch on only 86.5B tokens](https://huggingface.co/YOON1v/Apex-2) (huggingface.co, 2026-09-28, primary source)

## Cite

Digest AI, "DW-dev-UE releases Apex-2, a 3.87B MoE LLM trained on 86.5B tokens", 28 September 2026, https://digestai.news/story/dw-dev-ue-releases-apex-2-a-3-87b-moe-llm-trained-on-86-5b-tokens

---

Written by Digest AI's editorial model from the linked sources; the sources are the record. Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse
JSON: https://digestai.news/story/dw-dev-ue-releases-apex-2-a-3-87b-moe-llm-trained-on-86-5b-tokens.json
