# Qwen3.8-27B-pi fixes coding agent effort levels with RL fine-tuning

Digest AI · Agents & Tools · published 2026-09-30T20:32:41Z

Canonical: https://digestai.news/story/qwen3-8-27b-pi-fixes-coding-agent-effort-levels-with-rl-fine-tuning

## Summary

Qwen3.8-27B-pi is a fine-tuned version of the Qwen3.8-27B model designed for agentic coding tasks within the Pi agent harness. The model addresses a key flaw in the base model: its effort levels (low, medium, xhigh) were not consistently ordered, with low often spending more reasoning tokens than medium while achieving lower pass rates. The fine-tuning process involved two stages: supervised fine-tuning (SFT) on curated Pi sessions to teach complete workflows, followed by reinforcement learning (RL) with a success-conditioned reward to enforce effort ordering—ensuring lower levels do not reason more than higher levels for the same task.

The results show significant improvements. On Terminal-Bench 2.1, Pi at medium effort matches the base model’s xhigh performance with 41% fewer output tokens. On GPQA Diamond, Pi achieves a pass rate of 86.4%, outperforming the base model across all effort levels. The model is released as open-source under the MIT license, with BF16 and FP8 weights (~30.4 GB) and 17 GGUF variants (9–29 GB). The author, Mario Zechner, notes this is a proof-of-concept release constrained by budget and calls for further exploration, such as scaling RL or testing smaller models through the same pipeline.

## Key points

- Qwen3.8-27B-pi fixes inconsistent effort-level reasoning in coding agents via RL fine-tuning, ensuring lower levels spend less reasoning than higher ones
- On Terminal-Bench 2.1, Pi at medium effort matches base xhigh performance with 41% fewer output tokens and higher pass rates across all benchmarks
- Open-source release includes BF16/FP8 weights (~30.4 GB) and 17 GGUF variants (9–29 GB), with no official deployment support beyond local use

## Why it matters

Fixing effort-level inconsistencies in coding agents improves efficiency and reliability for real-world tasks like repository edits and terminal work, lowering costs for developers and enterprises using AI-driven workflows.

## Sources

1. [Qwen3.8-27B-pi: Effort-Ordered Reasoning for Agentic Coding](https://huggingface.co/blog/bytkim/qwen38-27b-pi) (huggingface.co, 2026-09-30, primary source)

Part of the developing story: [Qwen3.8-27B Evolution Token Cuts and RL Tuning](https://digestai.news/thread/bottlecap-ai-cuts-qwen3-8-27bs-reasoning-tokens-by-37-2-with-thinkingcap-qwen3) (2 stories)

## Cite

Digest AI, "Qwen3.8-27B-pi fixes coding agent effort levels with RL fine-tuning", 30 September 2026, https://digestai.news/story/qwen3-8-27b-pi-fixes-coding-agent-effort-levels-with-rl-fine-tuning

---

Written by Digest AI's editorial model from the linked sources; the sources are the record. Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse
JSON: https://digestai.news/story/qwen3-8-27b-pi-fixes-coding-agent-effort-levels-with-rl-fine-tuning.json
