{"version":1,"type":"story","url":"https://digestai.news/story/dw-dev-ue-releases-apex-2-a-3-87b-moe-llm-trained-on-86-5b-tokens","json":"https://digestai.news/story/dw-dev-ue-releases-apex-2-a-3-87b-moe-llm-trained-on-86-5b-tokens.json","markdown":"https://digestai.news/story/dw-dev-ue-releases-apex-2-a-3-87b-moe-llm-trained-on-86-5b-tokens.md","slug":"dw-dev-ue-releases-apex-2-a-3-87b-moe-llm-trained-on-86-5b-tokens","headline":"DW-dev-UE releases Apex-2, a 3.87B MoE LLM trained on 86.5B tokens","summary":"DW-dev-UE published the Apex-2 language model, a mixture‑of‑experts system with 3.87 billion total parameters and 1.45 billion active ones. The model was pretrained from scratch on 86.5 billion tokens drawn from web, code, math and curated sources, then instruction‑tuned with a two‑stage SFT pipeline. The release includes full code, architecture write‑up and a training diary on GitHub, and the checkpoint is available under an Apache 2.0 license.\n\nIn benchmark tests Apex-2’s base model matches the performance of Qwen2.5‑1.5B (which used 18 trillion tokens) on HumanEval+ (32.9 vs 32.9) despite using only a fraction of the data. Gaps remain in knowledge (MMLU ~29%) and math, and the model is English‑centric with limited multilingual ability. A DPO stage was abandoned after it reduced scores on code and math tasks. The model supports a 4096‑token context window, uses the ChatML template, and can be loaded directly with vLLM.","keyPoints":["Apex-2 has 3.87 B total, 1.45 B active parameters and was trained on 86.5 B tokens.","HumanEval+ score matches Qwen2.5‑1.5B (32.9) despite using far less pretraining data.","Model released under Apache 2.0, with code and checkpoint on GitHub (DW-dev-UE)."],"whyItMatters":"Apex-2 shows that high‑quality code and instruction performance can be achieved with modest data, offering an open‑weight alternative for researchers and developers seeking a small‑scale MoE LLM.","category":{"slug":"models","name":"Generative AI & Models","url":"https://digestai.news/category/models"},"entities":{"companies":["DW-dev-UE"],"models":["Apex-2"],"people":[]},"firstPublishedAt":"2026-09-28T00:00:00Z","updatedAt":"2026-09-28T00:00:00Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"huggingface.co","title":"I trained a 3.87B MoE (1.45B active) from scratch on only 86.5B tokens","url":"https://huggingface.co/YOON1v/Apex-2","publishedAt":"2026-09-28T00:00:00Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[{"site":"Reddit","url":"https://www.reddit.com/r/LocalLLaMA/comments/1wxiy8y/i_trained_a_387b_moe_145b_active_from_scratch_on/","points":null}],"thread":null,"cite":{"text":"Digest AI, \"DW-dev-UE releases Apex-2, a 3.87B MoE LLM trained on 86.5B tokens\", 28 September 2026, https://digestai.news/story/dw-dev-ue-releases-apex-2-a-3-87b-moe-llm-trained-on-86-5b-tokens","publisher":"Digest AI","title":"DW-dev-UE releases Apex-2, a 3.87B MoE LLM trained on 86.5B tokens","datePublished":"2026-09-28T00:00:00Z","url":"https://digestai.news/story/dw-dev-ue-releases-apex-2-a-3-87b-moe-llm-trained-on-86-5b-tokens"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}