{"version":1,"type":"story","url":"https://digestai.news/story/bottlecap-ai-cuts-qwen3-8-27bs-reasoning-tokens-by-37-2-with-thinkingc","json":"https://digestai.news/story/bottlecap-ai-cuts-qwen3-8-27bs-reasoning-tokens-by-37-2-with-thinkingc.json","markdown":"https://digestai.news/story/bottlecap-ai-cuts-qwen3-8-27bs-reasoning-tokens-by-37-2-with-thinkingc.md","slug":"bottlecap-ai-cuts-qwen3-8-27bs-reasoning-tokens-by-37-2-with-thinkingc","headline":"BottleCap AI cuts Qwen3.8-27B’s reasoning tokens by 37.2% with ThinkingCap-Qwen3.8-27B","summary":"BottleCap AI released **ThinkingCap-Qwen3.8-27B**, a fine-tuned version of Qwen’s **Qwen3.8-27B** designed to reduce reasoning tokens by **37.2%** across 12 benchmarks. The model maintains near-identical accuracy (a **0.86pp drop** to **85.79%**), with some benchmarks—like **MMMLU**—saving up to **65.5%** fewer tokens. Long-context retrieval accuracy improved by **2.25pp**, while math-heavy tasks like **AIME 2026** saw a **3.85pp accuracy drop** for **30.2% fewer tokens**. BottleCap claims the trade-off favors efficiency without sacrificing core reasoning capabilities.\n\nThe model drops into existing Qwen3.8-27B setups via **vLLM or SGLang**, with quantized builds (FP8, NVFP4, GGUF, MLX) for deployment. Weights are gated under **PolyForm Small Business 1.0.0**, requiring a BottleCap agreement for commercial use. Benchmarking used identical settings (NVIDIA H200, vLLM 0.29.0) and highlights **xhigh effort mode** as optimal for accuracy-to-token balance. Future work will explore individual reasoning modes.","keyPoints":["ThinkingCap-Qwen3.8-27B cuts reasoning tokens by 37.2% across 12 benchmarks, with accuracy dropping 0.86pp to 85.79%","Long-context retrieval accuracy rises 2.25pp, but math-heavy tasks like AIME 2026 lose 3.85pp accuracy for 30.2% fewer tokens","Model drops into Qwen3.8-27B setups via vLLM/SGLang, with gated weights under PolyForm Small Business license"],"whyItMatters":"For AI developers, this model offers a **37% token reduction** without sacrificing core reasoning, cutting inference costs for applications needing long traces. The trade-off—smaller accuracy drops for some tasks—may appeal to cost-sensitive deployments, though math-heavy workloads require caution.","category":{"slug":"models","name":"Generative AI & Models","url":"https://digestai.news/category/models"},"entities":{"companies":["BottleCap AI","Qwen"],"models":["ThinkingCap-Qwen3.8-27B","Qwen3.8-27B","Qwen3.6-27B"],"people":["Michal Sutter"]},"firstPublishedAt":"2026-09-24T18:58:56Z","updatedAt":"2026-09-24T18:58:56Z","sourceCount":1,"hasPrimarySource":false,"sources":[{"outlet":"MarkTechPost","title":"BottleCap AI Releases ThinkingCap-Qwen3.8-27B: 37.2% Fewer Thinking Tokens at a 0.86pp Accuracy Cost","url":"https://marktechpost.com/2026/09/24/bottlecap-ai-releases-thinkingcap-qwen3-8-27b-37-2-fewer-thinking-tokens-at-a-0-86pp-accuracy-cost","publishedAt":"2026-09-24T18:58:56Z","type":"press","primary":false,"lead":true}],"sourceNotes":null,"discussions":[],"thread":null,"cite":{"text":"Digest AI, \"BottleCap AI cuts Qwen3.8-27B’s reasoning tokens by 37.2% with ThinkingCap-Qwen3.8-27B\", 24 September 2026, https://digestai.news/story/bottlecap-ai-cuts-qwen3-8-27bs-reasoning-tokens-by-37-2-with-thinkingc","publisher":"Digest AI","title":"BottleCap AI cuts Qwen3.8-27B’s reasoning tokens by 37.2% with ThinkingCap-Qwen3.8-27B","datePublished":"2026-09-24T18:58:56Z","url":"https://digestai.news/story/bottlecap-ai-cuts-qwen3-8-27bs-reasoning-tokens-by-37-2-with-thinkingc"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}