Qwen3.8-27B-pi fixes coding agent effort levels with RL fine-tuning
Qwen3.8-27B-pi is a fine-tuned version of the Qwen3.8-27B model designed for agentic coding tasks within the Pi agent harness. The model addresses a key flaw in the base model: its effort levels (low, medium, xhigh) were not consistently ordered, with low often spending more reasoning tokens than medium while achieving lower pass rates. The fine-tuning process involved two stages: supervised…
Key points
- Qwen3.8-27B-pi fixes inconsistent effort-level reasoning in coding agents via RL fine-tuning, ensuring lower levels spend less reasoning than higher ones
- On Terminal-Bench 2.1, Pi at medium effort matches base xhigh performance with 41% fewer output tokens and higher pass rates across all benchmarks
- Open-source release includes BF16/FP8 weights (~30.4 GB) and 17 GGUF variants (9–29 GB), with no official deployment support beyond local use
The results show significant improvements. On Terminal-Bench 2.1, Pi at medium effort matches the base model’s xhigh performance with 41% fewer output tokens. On GPQA Diamond, Pi achieves a pass rate of 86.4%, outperforming the base model across all effort levels. The model is released as open-source under the MIT license, with BF16 and FP8 weights (~30.4 GB) and 17 GGUF variants (9–29 GB). The author, Mario Zechner, notes this is a proof-of-concept release constrained by budget and calls for further exploration, such as scaling RL or testing smaller models through the same pipeline.
Model page: Qwen3.8-27B-pi →
The story so far
2 episodes →- Qwen3.8-27B-pi fixes coding agent effort levels with RL fine-tuningthis story
Qwen3.8-27B-pi: Effort-Ordered Reasoning for Agentic Coding
huggingface.co · 30 September 2026
Loading the full article…
This text was published by huggingface.co. It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
Coverage and discussion
1source- Reddit discussionreddit.com
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us. Published by Martin K., who runs Digest AI and handles corrections.
More in Agents & Tools
All →- OpenAI adds Ultrafast API tier, Decisions API and security scans at DevDay · 2 src
- OpenAI and Meta launch cute AI agents with potential risks · 1 src
- OpenAI launches Dots, always-on AI agents powered by GPT-6 Astra · 49 src
- OpenAI’s Agents API lets AI handle tasks without human input · 1 src
- Meta’s Muse tops app charts with 1.1M installs in 10 days · 5 src
Comments
via GitHub Discussions