{"version":1,"type":"story","url":"https://digestai.news/story/aws-shows-multi-turn-rl-fine-tuning-for-search-agents-on-sagemaker-ai","json":"https://digestai.news/story/aws-shows-multi-turn-rl-fine-tuning-for-search-agents-on-sagemaker-ai.json","markdown":"https://digestai.news/story/aws-shows-multi-turn-rl-fine-tuning-for-search-agents-on-sagemaker-ai.md","slug":"aws-shows-multi-turn-rl-fine-tuning-for-search-agents-on-sagemaker-ai","headline":"AWS shows multi-turn RL fine-tuning for search agents on SageMaker AI","summary":"AWS published a tutorial on using Amazon SageMaker AI's multi-turn reinforcement learning (MTRL) to fine-tune a Qwen3.6-27B model into a search agent. The agent uses BM25 and vector search tools across multiple interaction turns to retrieve information. Training used nDCG@10 as a trajectory-level reward, with a -1 penalty for hitting turn or token limits.\n\nThe fine-tuned model improved nDCG@10 on three of four benchmarks: +23.7% on BrowseComp-Plus, +18.4% on WixQA, and +6% on Wands, with a slight regression on FreshStack. Failure rates dropped sharply, from 22.89% to 0.68% on BrowseComp-Plus. Training ran with maxepochs=1, globalbatchsize=128, and rolloutmaxconcurrency=32 on serverless infrastructure in the US West (Oregon) region.","keyPoints":["Fine-tuned Qwen3.6-27B with SageMaker AI MTRL for multi-turn search agent","nDCG@10 improved up to 23.7% on BrowseComp-Plus, failure rate fell to 0.68%","Serverless training with default PPO/CISPO algorithms, per-token pricing"],"whyItMatters":"Shows enterprises can build reliable, low-cost search agents by fine-tuning smaller models with multi-turn RL instead of relying on expensive frontier models.","category":{"slug":"agents","name":"Agents & Tools","url":"https://digestai.news/category/agents"},"entities":{"companies":["AWS","Amazon"],"models":["Qwen3.6-27B"],"people":[]},"firstPublishedAt":"2026-10-02T15:44:20Z","updatedAt":"2026-10-02T15:44:20Z","sourceCount":1,"hasPrimarySource":true,"sources":[{"outlet":"AWS Machine Learning Blog","title":"Fine-tune a search agent with multi-turn RL on Amazon SageMaker AI","url":"https://aws.amazon.com/blogs/machine-learning/fine-tune-a-search-agent-with-multi-turn-rl-on-amazon-sagemaker-ai","publishedAt":"2026-10-02T15:44:20Z","type":"primary","primary":true,"lead":true}],"sourceNotes":null,"discussions":[],"thread":{"title":"AWS Expands AI Tools For Search And Conversion","url":"https://digestai.news/thread/aws-shows-how-contextual-bandits-lift-conversions-in-acquisition-funnels","storyCount":2},"cite":{"text":"Digest AI, \"AWS shows multi-turn RL fine-tuning for search agents on SageMaker AI\", 2 October 2026, https://digestai.news/story/aws-shows-multi-turn-rl-fine-tuning-for-search-agents-on-sagemaker-ai","publisher":"Digest AI","title":"AWS shows multi-turn RL fine-tuning for search agents on SageMaker AI","datePublished":"2026-10-02T15:44:20Z","url":"https://digestai.news/story/aws-shows-multi-turn-rl-fine-tuning-for-search-agents-on-sagemaker-ai"},"generatedBy":"Written by Digest AI's editorial model from the linked sources; the sources are the record.","license":"Headlines, digests and key points are written by Digest AI and may be quoted with a link to the story page. Linked articles belong to their publishers. Terms: https://digestai.news/terms#reuse"}