MikhbarMIKHBAR
Artificial Intelligence

Fine-Tune Search Agents with Multi-Turn RL on SageMaker AI

AWS explores how fine-tuning an LLM-powered search agent with multi-turn reinforcement learning on Amazon SageMaker AI can achieve frontier-model reliability at a lower latency and cost.

Fine-Tune Search Agents with Multi-Turn RL on SageMaker AI

Introduction to Multi-Turn Search Agents

Search agents powered by large language models are transforming how enterprises retrieve information by autonomously deciding search strategies across multiple rounds of interaction. However, achieving reliable multi-step behavior can be challenging. While small models lack awareness of custom environments, frontier models often introduce higher latency and costs. Fine-tuning offers a balanced solution, yielding a smaller model's speed and cost efficiency with enhanced reliability. According to the [AWS Machine Learning Blog](https://aws.amazon.com/blogs/machine-learning/), traditional approaches such as supervised fine-tuning and single-turn reinforcement learning fall short because they fail to capture interdependent decisions across a full sequence of steps.

Understanding Amazon SageMaker AI MTRL

To address optimization across full multi-turn trajectories, developers can leverage [Amazon SageMaker AI MTRL](https://docs.aws.amazon.com/sagemaker/latest/dg/model-customize-mtrl.html) to fine-tune LLMs using reinforcement learning in multi-turn interaction settings. The platform frames agentic tasks as sequences of decisions, uses multi-turn rollouts to generate training data, and optimizes models with policy gradient algorithms. Key architectural features include a modular agent-environment interface, serverless execution at per-token pricing, asynchronous rollout and trajectory collection, and a native algorithm library featuring algorithms like PPO and CISPO.

Enterprise Search Setting and Tool Integration

In practical implementations on Amazon SageMaker AI, engineers configure an enterprise search setting where an LLM-powered system autonomously utilizes search tools to answer questions and complete tasks. The agent integrates two distinct search mechanisms: lexical search utilizing BM25 for exact keyword matches, and vector search employing embedding vectors for semantic and conceptual queries. Limiting the total number of turns helps prevent excessively long responses and encourages efficient search behavior during rollouts.

Line chart of nDCG@10 reward rising over training steps for both the training and validation sets
Image related to the report from AWS Machine Learning Blog · Source: AWS Machine Learning Blog

Datasets and Reward Configuration

The training methodology relies on preprocessed datasets formatted specifically for the service, with five percent of instances reserved for validation. The primary evaluation and reward metric used is nDCG@10, which measures how well the top retrieved documents match ideal rankings. When agents reach maximum turn limits or token thresholds, a negative reward penalty is applied to discourage undesirable behaviors without requiring complex intermediate rewards.

Training Job Configuration and SDK Usage

Configuring training jobs requires minimal adjustments to default settings using the MultiTurnRLTrainer SDK. Developers define hyperparameters such as max_epochs to control passes over the training data and global_batch_size to specify prompts per training step. This streamlined approach enables efficient fine-tuning of models like Qwen3.6-27B directly within supported AWS regions.

Sources

Continue chronologically

Related entity coverage