BottleCap AI cuts Qwen3.8-27B’s reasoning tokens by 37.2% with ThinkingCap-Qwen3.8-27B
BottleCap AI released ThinkingCap-Qwen3.8-27B, a fine-tuned version of Qwen’s Qwen3.8-27B designed to reduce reasoning tokens by 37.2% across 12 benchmarks. The model maintains near-identical accuracy (a 0.86pp drop to 85.79%), with some benchmarks—like MMMLU—saving up to 65.5% fewer tokens. Long-context retrieval accuracy improved by 2.25pp, while math-heavy tasks like AIME 2026 saw a 3.85pp…
Key points
- ThinkingCap-Qwen3.8-27B cuts reasoning tokens by 37.2% across 12 benchmarks, with accuracy dropping 0.86pp to 85.79%
- Long-context retrieval accuracy rises 2.25pp, but math-heavy tasks like AIME 2026 lose 3.85pp accuracy for 30.2% fewer tokens
- Model drops into Qwen3.8-27B setups via vLLM/SGLang, with gated weights under PolyForm Small Business license
The model drops into existing Qwen3.8-27B setups via vLLM or SGLang, with quantized builds (FP8, NVFP4, GGUF, MLX) for deployment. Weights are gated under PolyForm Small Business 1.0.0, requiring a BottleCap agreement for commercial use. Benchmarking used identical settings (NVIDIA H200, vLLM 0.29.0) and highlights xhigh effort mode as optimal for accuracy-to-token balance. Future work will explore individual reasoning modes.
Model page: ThinkingCap-Qwen3.8-27B →
BottleCap AI Releases ThinkingCap-Qwen3.8-27B: 37.2% Fewer Thinking Tokens at a 0.86pp Accuracy Cost
MarkTechPost · 24 September 2026
Loading the full article…
This text was published by MarkTechPost and written by Michal Sutter. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Generative AI & Models
All →- Anthropic releases Claude Opus 5.5; OpenAI halves prices for GPT-6 Sol and Luna · 14 src
- OpenAI releases ChatGPT Images 2.5 with sketch, point, and template tools · 4 src
- PrismML launches 1-bit Bonsai LLM for Qualcomm smart glasses · 1 src
- What has changed with GPT-6 Astra? Basic specifications of the new model that supports 'thinking, researching, and creating' in ChatGPT · 2 src
- Anthropic launches Claude Opus 5.5, claims 40% lower running costs than Opus 5 · 44 src
Comments
via GitHub Discussions