Google Research introduces ToolGrad for efficient AI tool-use dataset generation
Google Research has unveiled ToolGrad, a new framework designed to streamline the creation of datasets for training large language models in tool use. Unlike traditional methods that generate user queries first and then search for solutions via depth-first search, ToolGrad reverses this process by generating valid tool-use chains before creating the corresponding user prompts. This…
Key points
- ToolGrad generates tool-use datasets by creating solutions before queries, improving efficiency over traditional search methods.
- Fine-tuned Gemma-3 models using ToolGrad-500 data matched or surpassed proprietary LLMs on the BFCL benchmark.
- The framework achieves near 100% pass rates in data generation, enabling compact models to handle complex, long-horizon tasks.
The team evaluated this method using the ToolBench API database, which contains over 16,000 real-world APIs. They generated a small-scale dataset called ToolGrad-500 and used it to fine-tune Gemma-3 models ranging from 1B to 12B parameters. The resulting models, designated ToolGrad-1B, 4B, and 12B, were tested on the Berkeley Function Calling Leaderboard (BFCL). The results showed that these compact, fine-tuned models achieved near-perfect pass rates in data generation and outperformed their base versions. Notably, they matched or surpassed state-of-the-art proprietary models like Gemini, GPT, and Claude on out-of-distribution tasks involving unseen tools.
Presented at ACL 2026, this research addresses a key bottleneck in agentic AI development: the high cost and inefficiency of manually annotating or searching for complex tool-use trajectories. By enabling smaller models to perform at the level of larger proprietary systems, ToolGrad offers a scalable path for deploying capable digital agents in enterprise and consumer applications.
The story so far
32 episodes →- Google Research introduces ToolGrad for efficient AI tool-use dataset generation this story
ToolGrad: Efficient tool-use dataset generation with textual "gradients"
Google Research · 10 September 2026
September 10, 2026
Zhongyi Zhou, Research Scientist, and Ruofei Du, Interactive Perception & Graphics Lead, Google XR
ToolGrad is a data generation framework that reverses the traditional paradigm by first generating tool-use answers before user queries. We show this design enables LLMs to achieve better tool-use performance.
AI agents have shown great potential in automating real-world tasks, such as conducting a Google Search, reading local computer files, or executing generated Python scripts. To achieve such agentic workflows, LLMs need to learn how to use tools correctly and efficiently. To teach large language models tool uses, we need datasets of tool-use chains and their corresponding user queries. In our prior work introduced in InstructPipe, we manually annotated our evaluation data, but it is impractical to scale up the human annotation for advanced LLM fine-tuning workstreams. To streamline the data workstream, prior work, e.g., ToolBench and ToolACE, explored using an agent to automatically search a tool-use path with trial and error. This representative annotation approach involves two steps: (1) generate a hypothetical user instruction from a sampled API pool, and (2) use a depth-first search (DFS) agent to find its tool-use solution. This approach is inherently inefficient because its core concept is to distill valuable trajectories from a complex agent exploration for training an LLM.
In “ToolGrad: Efficient Tool-use Dataset Generation with Textual ‘Gradients’”, presented at ACL 2026, we introduce an alternative solution paradigm. ToolGrad first generates a ground-truth tool-use chain and then annotates its corresponding user prompt. Intuitively, an explicit tool-use solution provides more unambiguous information than a prompt, making the annotation, from tool usage to the use query, much easier and requiring only one LLM step. Our result shows that our answer-first approach can generate more complex (long-horizon) tool-use data with lower cost. LLMs trained on our generated data also outperform those trained on baseline methods, and even match SoTA proprietary LLMs on out-of-distribution (OOD) datasets with unseen tools.
Standard machine learning (ML) systems improve by computing numerical loss gradients across mini-batches of training samples, which are then used by an optimization algorithm to update model weights. Recently, TextGrad adapted this paradigm for prompt engineering using an LLM critic to provide rich, descriptive feedback in plain text — feedback called “textual gradients”. These textual gradients then guide the refinements of a given prompt into a new draft that can better resolve the target task.
ToolGrad adapts the concept of textual gradients from prompt optimization to synthetic dataset generation. Rather than optimizing a static text prompt, ToolGrad uses these gradients to iteratively construct complex, valid API workflows from large tool libraries.
ToolGrad features four core modules that sequentially propose, execute, select, and update.
Repeating this iterative process results in a data sample consisting of a user query, a verified API workflow, and the final AI response.
We first evaluate the cost and quality of the data generation. We use ToolBench as our API database, consisting of 16k+ real-world APIs, to generate our tool-use dataset. We compare the original query-first data generation approach on ToolBench, using depth-first search (DFS), with our answer-first approach, ToolGrad. The results demonstrate that ToolGrad can generate more complex tool-use data with higher pass rate, using lower generation cost.
We generated small-scale tool-use datasets called ToolGrad-500, using API databases from ToolBench. We then fine-tuned Gemma-3 models (1B, 4B and 12B) using ToolGrad-500, and we called these fine-tuned models ToolGrad-1B, ToolGrad-4B and ToolGrad-12B. We evaluated these models' tool-use performance on Berkeley Function Calling Leaderboard (BFCL), a tool-use benchmark with a different tool set from ToolBench. We compare our fine-tuned models against (1) base models without fine-tuning, (2) SoTA proprietary models (Gemini, GPT and Claude), and (3) SoTA tool-use specialized models (ToolACE, Hammer-2.1-7B).
The following summarizes our findings.
ToolGrad demonstrates that high-quality tool-use datasets can be generated more efficiently and reliably through an answer-first paradigm. By designing an agentic framework that iteratively chains APIs via textual gradients, ToolGrad addresses the longstanding cost and scalability bottlenecks in producing ground-truth data. Our design achieves almost 100% pass rate in data generation, enables relatively compact models to perform exceptionally well, and shows that student LLMs can even surpass their teachers.
Looking ahead, this research can be expanded to broader, real-world applications by scaling the framework to handle increasingly dynamic and vast API ecosystems. Future work will also explore extending this self-evolving capability to support continuous, on-the-fly learning for personalization over time. As agentic workflows become increasingly embedded in enterprise and everyday tasks, frameworks like ToolGrad lay the essential groundwork for training digital agents that are both highly capable and economically scalable to deploy.
This research was primarily conducted by Zhongyi Zhou during his Visiting Researcher tenure at Google. We extend our sincere gratitude to key contributors, Kohei Uehara, Haoyu Zhang, Jingtao Zhou, Lin Gu, Zheng Xu, Tatsuya Harada, for their support, and to Adarsh Kowdle and Shahram Izadi for their strategic guidance and thoughtful reviews.
This text was published by Google Research . It is reproduced here with attribution so you can read it in full; the rights remain with the publisher. Read it at the source ↗
The headline, key points and digest above were generated by Digest AI's editorial model from the linked sources. Automated summaries can contain errors: the sources are the record. Spotted a mistake? Tell us.
Comments
via GitHub Discussions