{"version":1,"type":"story","url":"https://digestai.news/story/qwen3-8-27b-pi-fixes-coding-agent-effort-levels-with-rl-fine-tuning","json":"https://digestai.news/story/qwen3-8-27b-pi-fixes-coding-agent-effort-levels-with-rl-fine-tuning.json","markdown":"https://digestai.news/story/qwen3-8-27b-pi-fixes-coding-agent-effort-levels-with-rl-fine-tuning.md","slug":"qwen3-8-27b-pi-fixes-coding-agent-effort-levels-with-rl-fine-tuning","headline":"Qwen3.8-27B-pi fixes coding agent effort levels with RL fine-tuning","summary":"Qwen3.8-27B-pi is a fine-tuned version of the Qwen3.8-27B model designed for agentic coding tasks within the Pi agent harness. The model addresses a key flaw in the base model: its effort levels (low, medium, xhigh) were not consistently ordered, with low often spending more reasoning tokens than medium while achieving lower pass rates. The fine-tuning process involved two stages: supervised fine-tuning (SFT) on curated Pi sessions to teach complete workflows, followed by reinforcement learning (RL) with a success-conditioned reward to enforce effort ordering—ensuring lower levels do not reason more than higher levels for the same task.\n\nThe results show significant improvements. On Terminal-Bench 2.1, Pi at medium effort matches the base model’s xhigh performance with 41% fewer output tokens. On GPQA Diamond, Pi achieves a pass rate of 86.4%, outperforming the base model across all effort levels. The model is released as open-source under the MIT license, with BF16 and FP8 weights (~30.4 GB) and 17 GGUF variants (9–29 GB). The author, Mario Zechner, notes this is a proof-of-concept release constrained by budget and calls for further exploration, such as scaling RL or testing smaller models through the same pipeline.","keyPoints":["Qwen3.8-27B-pi fixes inconsistent effort-level reasoning in coding agents via RL fine-tuning, ensuring lower levels spend less reasoning than higher ones","On Terminal-Bench 2.1, Pi at medium effort matches base xhigh performance with 41% fewer output tokens and higher pass rates across all benchmarks","Open-source release includes BF16/FP8 weights (~30.4 GB) and 17 GGUF variants (9–29 GB), with no official deployment support beyond local use"],"whyItMatters":"Fixing effort-level inconsistencies in coding agents improves efficiency and reliability for real-world tasks like repository edits and terminal work, lowering costs for developers and enterprises using AI-driven workflows.","category":{"slug":"agents","name":"Agents & Tools","url":"https://digestai.news/category/agents"},"entities":{"companies":["Tongyi Lab"],"models":["Qwen3.8-27B","Qwen3.8-27B-pi"],"people":["Mario Zechner"]},"firstPublishedAt":"2026-09-30T20:32:41Z","updatedAt":"2026-09-30T20:32:41Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"huggingface.co","title":"Qwen3.8-27B-pi: Effort-Ordered Reasoning for Agentic Coding","url":"https://huggingface.co/blog/bytkim/qwen38-27b-pi","publishedAt":"2026-09-30T20:32:41Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[{"site":"Reddit","url":"https://www.reddit.com/r/LocalLLaMA/comments/1wufzrh/qwen3827bpi_effortordered_reasoning_for_agentic/","points":null}],"thread":{"title":"Qwen3.8-27B Evolution Token Cuts and RL Tuning","url":"https://digestai.news/thread/bottlecap-ai-cuts-qwen3-8-27bs-reasoning-tokens-by-37-2-with-thinkingcap-qwen3","storyCount":2},"cite":{"text":"Digest AI, \"Qwen3.8-27B-pi fixes coding agent effort levels with RL fine-tuning\", 30 September 2026, https://digestai.news/story/qwen3-8-27b-pi-fixes-coding-agent-effort-levels-with-rl-fine-tuning","publisher":"Digest AI","title":"Qwen3.8-27B-pi fixes coding agent effort levels with RL fine-tuning","datePublished":"2026-09-30T20:32:41Z","url":"https://digestai.news/story/qwen3-8-27b-pi-fixes-coding-agent-effort-levels-with-rl-fine-tuning"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}