@hooshaaii: LLMs shouldn't guess facts; they should use tools. "Toolformer" (2023) trains LMs to self-teach how to use external API…

X AI KOLs Timeline Papers

Summary

Toolformer trains language models to self-teach how to use external APIs like calculators and search engines via self-supervised learning, significantly improving zero-shot performance across tasks.

LLMs shouldn't guess facts; they should use tools. "Toolformer" (2023) trains LMs to self-teach how to use external APIs (calculators, search engines) via simple text tags. A huge step toward real-world LLM utility! 🛠️ 👇 https://t.co/bFMeGE08Mc https://t.co/sD455FAGqj
Original Article
View Cached Full Text

Cached at: 09/20/26, 03:25 PM

LLMs shouldn’t guess facts; they should use tools.

“Toolformer” (2023) trains LMs to self-teach how to use external APIs (calculators, search engines) via simple text tags.

A huge step toward real-world LLM utility! 🛠️

👇 https://t.co/bFMeGE08Mc https://t.co/sD455FAGqj


Toolformer: Language Models Can Teach Themselves to Use Tools

Source: https://arxiv.org/html/2302.04761 Jane Dwivedi-YuRoberto Dessì†Roberta RaileanuAffiliation:Maria LomeliLuke ZettlemoyerNicola CanceddaThomas ScialomAffiliation:Meta AI Research†Universitat Pompeu Fabra

Abstract

Language models (LMs) exhibit remarkable abilities to solve new tasks from just a few examples or textual instructions, especially at scale. They also, paradoxically, struggle with basic functionality, such as arithmetic or factual lookup, where much simpler and smaller models excel. In this paper, we show that LMs can teach themselves touse external toolsvia simple APIs and achieve the best of both worlds. We introduceToolformer, a model trained to decide which APIs to call, when to call them, what arguments to pass, and how to best incorporate the results into future token prediction. This is done in a self-supervised way, requiring nothing more than a handful of demonstrations for each API. We incorporate a range of tools, including a calculator, a Q&A system, a search engine, a translation system, and a calendar. Toolformer achieves substantially improved zero-shot performance across a variety of downstream tasks, often competitive with much larger models, without sacrificing its core language modeling abilities.

1Introduction

Large language models achieve impressive zero- and few-shot results on a variety of natural language processing tasks(Brown et al., 2020;Chowdhery et al., 2022, i.a.)and show several emergent capabilities(Wei et al., 2022). However, all of these models have several inherent limitations that can at best be partially addressed by further scaling. These limitations include an inability to access up-to-date information on recent events(Komeili et al., 2022)and the related tendency to hallucinate facts(Maynez et al., 2020;Ji et al., 2022), difficulties in understanding low-resource languages(Lin et al., 2021), a lack of mathematical skills to perform precise calculations(Patel et al., 2021)and an unawareness of the progression of time(Dhingra et al., 2022).

Figure 1:Exemplary predictions of Toolformer. The model autonomously decides to call different APIs (from top to bottom: a question answering system, a calculator, a machine translation system, and a Wikipedia search engine) to obtain information that is useful for completing a piece of text.Figure 2:Key steps in our approach, illustrated for aquestion answeringtool: Given an input text𝐱\mathbf{x}, we first sample a positioniiand corresponding API call candidatesci1,ci2,…,cikc_{i}^{1},c_{i}^{2},\ldots,c_{i}^{k}. We then execute these API calls and filter out all calls which do not reduce the lossLiL_{i}over the next tokens. All remaining API calls are interleaved with the original text, resulting in a new text𝐱∗\mathbf{x}^{*}.A simple way to overcome these limitations of today’s language models is to give them the ability touse external toolssuch as search engines, calculators, or calendars. However, existing approaches either rely on large amounts of human annotations(Komeili et al., 2022;Thoppilan et al., 2022)or limit tool use to task-specific settings only(Gao et al., 2022;Parisi et al., 2022, e.g.,), hindering a more widespread adoption of tool use in LMs. Therefore, we proposeToolformer, a model that learns to use tools in a novel way, which fulfills the following desiderata:

  • •The use of tools should be learned in a self-supervised way without requiring large amounts ofhuman annotations. This is important not only because of the costs associated with such annotations, but also because what humans find useful may be different from what a model finds useful.
  • •The LM should not lose any of itsgeneralityand should be able to decide for itselfwhenandhowto use which tool. In contrast to existing approaches, this enables a much more comprehensive use of tools that is not tied to specific tasks.

Our approach for achieving these goals is based on the recent idea of using large LMs within-context learning(Brown et al., 2020)to generate entire datasets from scratch(Schick and Schütze, 2021b;Honovich et al., 2022;Wang et al., 2022): Given just a handful of human-written examples of how an API can be used, we let a LM annotate a huge language modeling dataset with potential API calls. We then use a self-supervised loss to determine which of these API calls actually help the model in predicting future tokens. Finally, we finetune the LM itself on the API calls that it considers useful. As illustrated in Figure1, through this simple approach, LMs can learn to control a variety of tools, and to choose for themselves which tool to use when and how.

As our approach is agnostic of the dataset being used, we can apply it to the exact same dataset that was used to pretrain a model in the first place. This ensures that the model does not lose any of its generality and language modeling abilities. We conduct experiments on a variety of different downstream tasks, demonstrating that after learning to use tools, Toolformer, which is based on a pretrained GPT-J model(Wang and Komatsuzaki, 2021)with 6.7B parameters, achieves much stronger zero-shot results, clearly outperforming a much larger GPT-3 model(Brown et al., 2020)and several other baselines on various tasks.

2Approach

Our aim is to equip a language modelMMwith the ability to use different tools by means of API calls. We require that inputs and outputs for each API can be represented as text sequences. This allows seamless insertion of API calls into any given text, using special tokens to mark the start and end of each such call.

We represent each API call as a tuplec=(ac,ic)c=({a}_{c},{i}_{c})whereaca_{c}is the name of the API andici_{c}is the corresponding input. Given an API callccwith a corresponding resultrr, we denote the linearized sequences of the API call not including and including its result, respectively, as:

e​(c)\displaystyle\text{e}(c)=<API>​ac​(​ic​)​</API>\displaystyle=\texttt{<API>}\,a_{c}\texttt{(}i_{c}\texttt{)}\,\texttt{</API>}e​(c,r)\displaystyle\text{e}(c,r)=<API>​ac​(​ic​)→r​</API>\displaystyle=\texttt{<API>}\,a_{c}\texttt{(}i_{c}\texttt{)}\rightarrow r\,\texttt{</API>}where “<API>”, “</API>” and “→\rightarrow’’ are special tokens.111In practice, we use the token sequences “[”, “]” and “->” to represent “<API>”, “</API>” and “→\rightarrow”, respectively. This enables our approach to work without modifying the existing LM’s vocabulary. For reasons of readability, we still refer to them as “<API>”, “</API>” and “→\rightarrow” throughout this section.Some examples of linearized API calls inserted into text sequences are shown in Figure1.

Given a dataset𝒞={𝐱1,…,𝐱|𝒞|}\mathcal{C}=\{\mathbf{x}^{1},\ldots,\mathbf{x}^{|\mathcal{C}|}\}of plain texts, we first convert this dataset into a dataset𝒞∗\mathcal{C}^{*}augmented with API calls. This is done in three steps, illustrated in Figure2: First, we exploit the in-context learning ability ofMMto sample a large number of potential API calls. We then execute these API calls and finally check whether the obtained responses are helpful for predicting future tokens; this is used as a filtering criterion. After filtering, we merge API calls for different tools, resulting in the augmented dataset𝒞∗\mathcal{C}^{*}, and finetuneMMitself on this dataset. Each of these steps is described in more detail below.

Sampling API Calls

For each API, we write a promptP⁡(𝐱)P(\mathbf{x})that encourages the LM to annotate an example𝐱=x1,…,xn\mathbf{x}=x_{1},\ldots,x_{n}with API calls. An example of such a prompt for a question answering tool is shown in Figure3; all prompts used are shown in AppendixA.2. LetpM​(zn+1∣z1,…,zn)p_{M}(z_{n+1}\mid z_{1},\ldots,z_{n})be the probability thatMMassigns to tokenzn+1z_{n+1}as a continuation for the sequencez1,…,znz_{1},\ldots,z_{n}. We first sample up tokkcandidatepositionsfor doing API calls by computing, for eachi∈{1,…,n}i\in\{1,\ldots,n\}, the probability

pi=pM(<API>∣P(𝐱),x1:i−1)p_{i}=p_{M}(\texttt{<API>}\mid P(\mathbf{x}),x_{1:i-1})thatMMassigns to starting an API call at positionii. Given a sampling thresholdτs\tau_{s}, we keep all positionsI={i∣pi>τs}I=\{i\mid p_{i}>\tau_{s}\}; if there are more thankksuch positions, we only keep the topkk.

For each positioni∈Ii\in I, we then obtain up tommAPI callsci1,…,cimc_{i}^{1},\ldots,c_{i}^{m}by sampling fromMMgiven the sequence[P⁡(𝐱),x1,…,xi−1,<API>][P(\mathbf{x}),x_{1},\ldots,x_{i-1},\texttt{<API>}]as a prefix and</API>as an end-of-sequence token.222We discard all examples whereMMdoes not generate the</API>token.

Figure 3:An exemplary promptP⁡(𝐱)P(\mathbf{x})used to generate API calls for the question answering tool.

Executing API Calls

As a next step, we execute all API calls generated byMMto obtain the corresponding results. How this is done depends entirely on the API itself – for example, it can involve calling another neural network, executing a Python script or using a retrieval system to perform search over a large corpus. The response for each API callcic_{i}needs to be a single text sequencerir_{i}.

Filtering API Calls

Letiibe the position of the API callcic_{i}in the sequence𝐱=x1,…,xn\mathbf{x}=x_{1},\ldots,x_{n}, and letrir_{i}be the response from the API. Further, given a sequence(wi∣i∈ℕ)(w_{i}\mid i\in\mathbb{N})ofweights, let

Li(𝐳)=−∑j=inwj−i⋅logpM(xj∣𝐳,x1:j−1)L_{i}(\mathbf{z})=-\sum_{j=i}^{n}w_{j-i}\cdot\log{p_{M}(x_{j}\mid\mathbf{z},x_{1:j-1})}be the weighted cross entropy loss forMMover the tokensxi,…,xnx_{i},\ldots,x_{n}if the model is prefixed with𝐳\mathbf{z}. We compare two different instantiations of this loss:

Li+\displaystyle L_{i}^{+}=Li​(e​(ci,ri))\displaystyle=L_{i}(\text{e}(c_{i},r_{i}))Li−\displaystyle L_{i}^{-}=min⁡(Li​(ε),Li​(e​(ci,ε)))\displaystyle=\min\left(L_{i}(\varepsilon),L_{i}(\text{e}(c_{i},\varepsilon))\right)whereε\varepsilondenotes an empty sequence. The former is the weighted loss over all tokensxi,…,xnx_{i},\ldots,x_{n}if the API call and its result are given toMMas a prefix;333We providee​(ci,ri)\text{e}(c_{i},r_{i})as a prefix instead of inserting it at positioniibecauseMMis not yet finetuned on any examples containing API calls, so inserting it in the middle of𝐱\mathbf{x}would interrupt the flow and not align with patterns in the pretraining corpus, thus hurting perplexity.the latter is the minimum of the losses obtained from (i) doing no API call at all and (ii) doing an API call, but not providing the response. Intuitively, an API call is helpful toMMif providing it with both the inputandthe output of this call makes it easier for the model to predict future tokens, compared to not receiving the API call at all, or receiving only its input. Given a filtering thresholdτf\tau_{f}, we thus only keep API calls for which

Li−−Li+≥τfL_{i}^{-}-L_{i}^{+}\geq\tau_{f}holds, i.e., adding the API call and its resultreducesthe loss by at leastτf\tau_{f}, compared to not doing any API call or obtaining no result from it.

Model Finetuning

After sampling and filtering calls for all APIs, we finally merge the remaining API calls and interleave them with the original inputs. That is, for an input text𝐱=x1,…,xn\mathbf{x}=x_{1},\ldots,x_{n}with a corresponding API call and result(ci,ri)(c_{i},r_{i})at positionii, we construct the new sequence𝐱∗=x1:i−1,e(ci,ri),xi:n\mathbf{x}^{*}=x_{1:{i-1}},\text{e}(c_{i},r_{i}),x_{i:n}; we proceed analogously for texts with multiple API calls. Doing this for all𝐱∈𝒞\mathbf{x}\in\mathcal{C}results in the new dataset𝒞∗\mathcal{C}^{*}augmented with API calls. We use this new dataset to finetuneMM, using a standard language modeling objective. Crucially, apart from inserted API calls the augmented dataset𝒞∗\mathcal{C}^{*}contains the exact same texts as𝒞\mathcal{C}, the original dataset. As a consequence, finetuningMMon𝒞∗\mathcal{C}^{*}exposes it to the same content as finetuning on𝒞\mathcal{C}. Moreover, as API calls are inserted in exactly those positions and with exactly those inputs that helpMMpredict future tokens, finetuning on𝒞∗\mathcal{C}^{*}enables the language model to decide when and how to use which tool, based purely on its own feedback.

Inference

When generating text withMMafter finetuning with our approach, we perform regular decoding untilMMproduces the “→\rightarrow” token, indicating that it next expects the response for an API call. At this point, we interrupt the decoding process, call the appropriate API to get a response, and continue the decoding process after inserting both the response and the</API>token.

3Tools

We explore a variety of tools to address different shortcomings of regular LMs. The only constraints we impose on these tools is that (i) both their inputs and outputs can be represented as text sequences, and (ii) we can obtain a few demonstrations of their intended use. Concretely, we explore the following five tools: a question answering system, a Wikipedia search engine, a calculator, a calendar, and a machine translation system. Some examples of potential calls and return strings for the APIs associated with each of these tools are shown in Table1. We briefly discuss all tools below; further details can be found in AppendixA.

Table 1:Examples of inputs and outputs for all APIs used.##### Question Answering

Our first tool is a question answering system based on another LM that can answer simple factoid questions. Specifically, we useAtlas(Izacard et al., 2022), a retrieval-augmented LM finetuned on Natural Questions(Kwiatkowski et al., 2019).

Calculator

As a second tool, we use a calculator that can perform simple numeric calculations; we only support the four basic arithmetic operations. Results are always rounded to two decimal places.

Wikipedia Search

Our third tool is a search engine that, given a search term, returns short text snippets from Wikipedia. Compared to our question answering tool, this search enables a model to get more comprehensive information on a subject, but requires it to extract the relevant parts by itself. As our search engine, we use a BM25 retriever(Robertson et al., 1995;Baeza-Yates et al., 1999)that indexes the Wikipedia dump from KILT(Petroni et al., 2021).

Machine Translation System

Our fourth tool is a machine translation system based on a LM that can translate a phrase from any language into English. More concretely, we use the 600M parameter NLLB(Costa-jussà et al., 2022)as our multilingual machine translation model that works for 200 languages (including low-resource ones). The source language is automatically detected using thefastTextclassifier(Joulin et al., 2016), while the target language is always set to English.

Calendar

Our final tool is a calendar API that, when queried, returns the current date without taking any input. This provides temporal context for predictions that require some awareness of time.

4Experiments

We investigate whether our approach enables a model to use tools without any further supervision and to decide for itself when and how to call which of the available tools. To test this, we select a variety of downstream tasks where we assume at least one of the considered tools to be useful, and evaluate performance in zero-shot settings (Section4.2). Beyond that, we also ensure that our approach does not hurt the model’s core language modeling abilities; we verify this by looking at perplexity on two language modeling datasets (Section4.3). Finally, we investigate how the ability to learn using tools is affected by model size (Section4.4).

4.1Experimental Setup

Dataset Generation

Throughout all of our experiments, we use a subset of CCNet(Wenzek et al., 2020)as our language modeling dataset𝒞\mathcal{C}and GPT-J(Wang and Komatsuzaki, 2021)as our language modelMM. To reduce the computational cost of annotating𝒞\mathcal{C}with API calls, we define heuristics for some APIs to get a subset of𝒞\mathcal{C}for which API calls are more likely to be helpful than for an average text. For example, we only consider texts for the calculator tool if they contain at least three numbers. Details of the heuristics used are given in AppendixA. For obtaining𝒞∗\mathcal{C}^{*}from𝒞\mathcal{C}, we perform all steps described in Section2and additionally filter out all examples for which all API calls were eliminated in the filtering step.444While this filtering alters the distribution of training examples, we assume that the remaining examples are close enough to the original distribution so thatMM’s language modeling abilities remain unaffected. This assumption is empirically validated in Section4.3.For the weighting function, we use

wt=w~t∑s∈ℕw~s​with​w~t=max⁡(0,1−0.2⋅t)w_{t}=\frac{\tilde{w}_{t}}{\sum_{s\in\mathbb{N}}\tilde{w}_{s}}\text{ with }\tilde{w}_{t}=\max(0,1-0.2\cdot t)to make sure that API calls happen close to where the information provided by the API is actually helpful for the model. The thresholdsτs\tau_{s}andτf\tau_{f}are chosen individually for each tool to ensure a sufficiently larger number of examples; see AppendixAfor details. Table2shows relevant statistics of our final dataset augmented with API calls.

Table 2:Number of examples with API calls in𝒞∗\mathcal{C}^{*}for different values of our filtering thresholdτf\tau_{f}.

Model Finetuning

We finetuneMMon𝒞∗\mathcal{C}^{*}using a batch size of 128 and a learning rate of1⋅10−51\cdot 10^{-5}with linear warmup for the first 10% of training. Details of our finetuning procedure are given in AppendixB.

Baseline Models

Throughout the remainder of this section, we mainly compare the following models:

  • •GPT-J: A regular GPT-J model without any finetuning.
  • •GPT-J + CC: GPT-J finetuned on𝒞\mathcal{C}, our subset of CCNetwithoutany API calls.
  • •Toolformer: GPT-J finetuned on𝒞∗\mathcal{C}^{*}, our subset of CCNet augmented with API calls.
  • •Toolformer (disabled): The same model as Toolformer, but API calls are disabled during decoding.555This is achieved by manually setting the probability of the<API>token to 0.

For most tasks, we additionally compare to OPT (66B)(Zhang et al., 2022)and GPT-3666We use the originaldavincivariant that is not finetuned on any instructions.(175B)(Brown et al., 2020), two models that are about 10 and 25 times larger than our other baseline models, respectively.

4.2Downstream Tasks

We evaluate all models on a variety of downstream tasks. In all cases, we consider a prompted zero-shot setup – i.e., models are instructed to solve each task in natural language, but we do not provide any in-context examples. This is in contrast to prior work on tool use(Gao et al., 2022;Parisi et al., 2022, e.g.,), where models are provided with dataset-specific examples of how a tool can be used to solve a concrete task. We choose the more challenging zero-shot setup as we are interested in seeing whether Toolformer works in precisely those cases where a user does not specify in advance which tools should be used in which way for solving a specific problem.

We use standard greedy decoding, but with one modification for Toolformer: We let the model start an API call not just when<API>is the most likely token, but whenever it is one of thekkmost likely tokens. Fork=1k=1, this corresponds to regular greedy decoding; we instead usek=10k=10to increase the disposition of our model to make use of the APIs that it has access to. At the same time, we only at most one API call per input to make sure the model does not get stuck in a loop where it constantly calls APIs without producing any actual output. The effect of these modifications is explored in Section5.

4.2.1LAMA

We evaluate our models on the SQuAD, Google-RE and T-REx subsets of the LAMA benchmark(Petroni et al., 2019). For each of these subsets, the task is to complete a short statement with a missing fact (e.g., a date or a place). As LAMA was originally designed to evaluatemaskedlanguage models(Devlin et al., 2019, e.g.,), we filter out examples where the mask token is not the final token, so that the remaining examples can be processed in a left-to-right fashion. To account for different tokenizations and added complexity from not informing the model that a single word is required, we use a slightly more lenient evaluation criterion than exact match and simply check whether the correct word is within the first five words predicted by the model. As LAMA is based on statements obtained directly from Wikipedia, we prevent Toolformer from using the Wikipedia Search API to avoid giving it an unfair advantage.

Results for all models can be seen in Table3. All GPT-J models without tool use achieve similar performance. Crucially, Toolformer clearly outperforms these baseline models, improving upon the best baseline by 11.7, 5.2 and 18.6 points, respectively. It also clearly outperforms OPT (66B) and GPT-3 (175B), despite both models being much larger. This is achieved because the model independently decides to ask the question answering tool for the required information in almost all cases (98.1%); for only very few examples, it uses a different tool (0.7%) or no tool at all (1.2%).

ModelSQuADGoogle-RET-RExGPT-J17.84.931.9GPT-J + CC19.25.633.2Toolformer (disabled)22.16.334.9Toolformer33.811.553.5OPT (66B)21.62.930.1GPT-3 (175B)26.87.039.8Table 3:Results on subsets of LAMA. Toolformer uses the question answering tool for most examples, clearly outperforming all baselines of the same size and achieving results competitive with GPT-3 (175B).

4.2.2Math Datasets

We test mathematical reasoning abilities on ASDiv(Miao et al., 2020), SVAMP(Patel et al., 2021)and the MAWPS benchmark(Koncel-Kedziorski et al., 2016). We again account for the fact that we test all models in a zero-shot setup by using a more lenient evaluation criterion: As the required output is always a number, we simply check for the first number predicted by the model.777An exception to this is if the model’s prediction contains an equation (e.g., “The correct answer is 5+3=8”), in which case we consider the first number after the “=” sign to be its prediction.

Table4shows results for all benchmarks. While GPT-J and GPT-J + CC perform about the same, Toolformer achieves stronger results even when API calls are disabled. We surmise that this is because the model is finetuned on many examples of API calls and their results, improving its own mathematical capabilities. Nonetheless, allowing the model to make API calls more than doubles performance for all tasks, and also clearly outperforms the much larger OPT and GPT-3 models. This is because across all benchmarks, for 97.9% of all examples the model decides to ask the calculator tool for help.

Table 4:Results for various benchmarks requiring mathematical reasoning. Toolformer makes use of the calculator tool for most examples, clearly outperforming even OPT (66B) and GPT-3 (175B).

4.2.3Question Answering

We look at Web Questions(Berant et al., 2013), Natural Questions(Kwiatkowski et al., 2019)and TriviaQA(Joshi et al., 2017), the three question answering datasets considered byBrown et al. (2020). For evaluation, we check whether the first 20 words predicted by a model contain the correct answer instead of requiring an exact match. For Toolformer, we disable the question answering tool as this would make solving the tasks trivial, especially given that the underlying QA system was finetuned on Natural Questions.

Results are shown in Table5. Once again, Toolformer clearly outperforms all other models based on GPT-J, this time mostly relying on the Wikipedia search API (99.3%) to find relevant information. However, Toolformer still lags behind the much larger GPT-3 (175B) model. This is likely due to both the simplicity of our search engine (in many cases, it returns results that are clearly not a good match for a given query) and the inability of Toolformer tointeractwith it, e.g., by reformulating its query if results are not helpful or by browsing through multiple of the top results. We believe that adding this functionality is an exciting direction for future work.

Table 5:Results for various question answering dataset. Using the Wikipedia search tool for most examples, Toolformer clearly outperforms baselines of the same size, but falls short of GPT-3 (175B).

4.2.4Multilingual Question Answering

We evaluate Toolformer and all baseline models on MLQA(Lewis et al., 2019), a multilingual question-answering benchmark. A context paragraph for each question is provided in English, while the question can be in Arabic, German, Spanish, Hindi, Vietnamese, or Simplified Chinese. In order to solve the task, the model needs to be able to understand both the paragraph and the question, so it may benefit from translating the question into English. Our evaluation metric is the percentage of times the model’s generation, capped at 10 words, contains the correct answer.

Results are shown in Table6. Using API calls consistently improves Toolformer’s performance for all languages, suggesting that it has learned to make use of the machine translation tool. Depending on the language, this tool is used for 63.8% to 94.9% of all examples; the only exception to this is Hindi, for which the machine translation tool is used in only 7.3% of cases. However, Toolformer does not consistently outperform vanilla GPT-J. This is mainly because for some languages, finetuning on CCNet deteriorates performance; this might be due to a distribution shift compared to GPT-J’s original pretraining data.

OPT and GPT-3 perform surprisingly weak across all languages, mostly because they fail to provide an answer in English despite being instructed to do so. A potential reason for GPT-J not suffering from this problem is that it was trained on more multilingual data than both OPT and GPT-3, including the EuroParl corpus(Koehn, 2005;Gao et al., 2020). As an upper bound, we also evaluate GPT-J and GPT-3 on a variant of MLQA where both the context and the question are provided in English. In this setup, GPT-3 performs better than all other models, supporting our hypothesis that its subpar performance on MLQA is due to the multilingual aspect of the task.

ModelEsDeHiViZhArGPT-J15.216.51.38.218.28.2GPT-J + CC15.714.90.58.313.74.6Toolformer (disabled)19.811.91.210.115.03.1Toolformer20.613.51.410.616.83.7OPT (66B)0.30.11.10.20.70.1GPT-3 (175B)3.41.10.11.717.70.1GPT-J (All En)24.327.023.923.323.123.6GPT-3 (All En)24.727.226.124.923.624.0Table 6:Results on MLQA for Spanish (Es), German (De), Hindi (Hi), Vietnamese (Vi), Chinese (Zh) and Arabic (Ar). While using the machine translation tool to translate questions is helpful across all languages, further pretraining on CCNet deteriorates performance; consequently, Toolformer does not consistently outperform GPT-J. The final two rows correspond to models that are given contexts and questions in English.

4.2.5Temporal Datasets

To investigate the calendar API’s utility, we evaluate all models onTempLAMA(Dhingra et al., 2022)and a new dataset that we callDateset.TempLAMAis a dataset built from Wikidata that contains cloze queries about facts that change with time (e.g., “Cristiano Ronaldo plays for ___”) as well as the correct answer for the years between 2010 and 2020.Dateset, described in AppendixD, is also generated through a series of templates, but populated using a combination of random dates/durations (e.g., “What day of the week was it 30 days ago?”). Critically, knowing the current date is required to answer these questions. For both tasks, we use the same evaluation as for the original LAMA dataset.

Results shown in Table7illustrate that Toolformer outperforms all baselines for bothTempLAMAandDateset. However, closer inspection shows that improvements onTempLAMAcan not be attributed to the calendar tool, which is only used for 0.2% of all examples, but mostly to the Wikipedia search and question answering tools, which Toolformer calls the most. This makes sense given that named entities inTempLamaare often so specific and rare that even knowing the exact date alone would be of little help. The best course of action for this dataset – first querying the calendar API to get the current date, and then querying the question answering system with this date – is not only prohibited by our restriction of using at most one API call per example, but also hard to learn for Toolformer given that all API calls in its training data are sampled independently.

ForDateset, on the other hand, the considerable improvement of Toolformer compared to other models can be fully accredited to the calendar tool, which it makes use of for 54.8% of all examples.

ModelTempLAMADatesetGPT-J13.73.9GPT-J + CC12.92.9Toolformer (disabled)12.75.9Toolformer16.327.3OPT (66B)14.51.3GPT-3 (175B)15.50.8Table 7:Results for the temporal datasets. Toolformer outperforms all baselines, but does not make use of the calendar tool forTempLAMA.

4.3Language Modeling

In addition to verifying improved performance on various downstream tasks, we also want to ensure that language modeling performance of Toolformer does not degrade through our finetuning with API calls. To this end, we evaluate our models on two language modeling datasets: WikiText(Merity et al., 2017)and a subset of 10,000 randomly selected documents from CCNet(Wenzek et al., 2020)that were not used during training. Perplexities of various models are shown in Table8. As one would expect, finetuning on CCNet leads to slightly improved performance on a different CCNet subset, but it slightly deteriorates performance on WikiText, presumably because the original pretraining data for GPT-J is more similar to WikiText than our randomly selected subset of CCNet. Most importantly, however, training on𝒞∗\mathcal{C}^{*}(our dataset annotated with API calls) does not lead to an increase in perplexity compared to training on𝒞\mathcal{C}when API calls are disabled at inference time.888We do not evaluate the perplexity of Toolformer with API calls enabled as computing the probabilitypM​(xt∣x1,…,xt−1)p_{M}(x_{t}\mid x_{1},\ldots,x_{t-1})of tokenxtx_{t}givenx1,…,xt−1x_{1},\ldots,x_{t-1}would require marginalizing over all potential API calls that the model could make at positiontt, which is intractable.

Table 8:Perplexities of different models on WikiText and our validation subset of CCNet. Adding API calls comes without a cost in terms of perplexity for language modeling without any API calls.

4.4Scaling Laws

Figure 4:Average performance on LAMA, our math benchmarks and our QA benchmarks for GPT-2 models of different sizes and GPT-J finetuned with our approach, both with and without API calls. While API calls are not helpful to the smallest models, larger models learn how to make good use of them. Even for bigger models, the gap between model predictions with and without API calls remains high.We investigate how the ability to ask external tools for help affects performance as we vary the size of our LM. To this end, we apply our approach not just to GPT-J, but also to four smaller models from the GPT-2 family(Radford et al., 2019), with 124M, 355M, 775M and 1.6B parameters, respectively. We do so using only a subset of three tools: the question answering system, the calculator, and the Wikipedia search engine. Apart from this, we follow the experimental setup described in Section4.1.

Figure4shows that the ability to leverage the provided tools only emerges at around 775M parameters: smaller models achieve similar performance both with and without tools. An exception to this is the Wikipedia search engine used mostly for QA benchmarks; we hypothesize that this is because the API is comparably easy to use. While models become better at solving taskswithoutAPI calls as they grow in size, their ability to make good use of the provided API improves at the same time. As a consequence, there remains a large gap between predictions with and without API calls even for our biggest model.

5Analysis

Decoding Strategy

We investigate the effect of our modified decoding strategy introduced in Section4.2, where instead of always generating the most likely token, we generate the<API>token if it is one of thekkmost likely tokens. Table9shows performance on the T-REx subset of LAMA and on WebQS for different values ofkk. As expected, increasingkkleads to the model doing API calls for more examples – from 40.3% and 8.5% withk=1k=1(i.e., regular greedy decoding) to 98.1% and 100% fork=10k=10. While for T-REx, there is already a clear improvement in performance with greedy decoding, on WebQS our model only starts to make a substantial number of API calls as we slightly increasekk. Interestingly, fork=1k=1the model is calibrated to some extent: It decides to call APIs for examples that it would perform particularly badly on without making API calls. This can be seen from the fact that performance on examples where it decidesnotto make an API call (44.3 and 19.9) is higher than average performance if no API calls are made at all (34.9 and 18.9). However, this calibration is lost for higher values ofkk.

Table 9:Toolformer results on the T-REx subset of LAMA and on WebQS for different values ofkkused during decoding. Numbers shown are overall performance (All), performance on the subset where the model decides to make an API call (AC) and all remaining examples (NC), as well as the percentage of examples for which the model decides to call an API (%).Table 10:Examples of API calls for different tools, sorted by the value ofLi−−Li+L_{i}^{-}\,{-}\,L_{i}^{+}that is used as a filtering criterion. High values typically correspond to API calls that are intuitively useful for predicting future tokens.

Data Quality

We qualitatively analyze some API calls generated with our approach for different APIs. Table10shows some examples of texts from CCNet augmented with API calls, as well as the corresponding scoreLi−−Li+L_{i}^{-}-L_{i}^{+}that is used as a filtering criterion, and whether the API calls made by the model are intuitively useful in the given context. As can be seen, high values ofLi−−Li+L_{i}^{-}-L_{i}^{+}typically correspond to useful API calls, whereas low values correspond to API calls that do not provide any information that is useful for predicting future tokens. There are some exceptions, e.g., an API call for “Fast train success” in the fourth example that does not give any relevant information but still reduces perplexity. However, some amount of noise in the API calls that are not filtered can actually be useful as it forces the model finetuned on𝒞∗\mathcal{C}^{*}to not always blindly follow the results of each call it makes.

6Related Work

Language Model Pretraining

There are various approaches that augment language models with some form of additional textual information during pretraining, including various forms of metadata(Keskar et al., 2019), HTML tags(Aghajanyan et al., 2021), Wikipedia markup(Schick et al., 2022), or related texts obtained from an information retrieval system(Guu et al., 2020;Borgeaud et al., 2021;Izacard et al., 2022). For all of these approaches, additional information isalwaysprovided, regardless of whether it is helpful or not. In contrast, Toolformer learns for itself to explicitly asks for the right information.

Tool Use

Several approaches aim to equip LMs with the ability to use external tools such as search engines(Komeili et al., 2022;Thoppilan et al., 2022;Lazaridou et al., 2022;Shuster et al., 2022;Yao et al., 2022), web browsers(Nakano et al., 2021), calculators(Cobbe et al., 2021;Thoppilan et al., 2022), translation systems(Thoppilan et al., 2022)and Python interpreters(Gao et al., 2022). The way these models learn to use tools can roughly be divided into two approaches: Either they rely on large amounts of human supervision(Komeili et al., 2022;Nakano et al., 2021;Thoppilan et al., 2022)or they work by prompting the language model in a few-shot setup tailored towards a specific task where it is known a priori which tools needs to be used(Gao et al., 2022;Lazaridou et al., 2022;Yao et al., 2022). In contrast, the self-supervised nature of Toolformer enables it to learn how and when to use tools without requiring a specific prompt that shows task-specific examples of how a tool could be used. Perhaps most closely related to our work is TALM(Parisi et al., 2022), an approach that uses a similar self-supervised objective for teaching a model to use a calculator and a search engine, but explores this only in settings where a model is finetuned for downstream tasks.

Bootstrapping

The idea of using self-training and bootstrapping techniques to improve models has been investigated in various contexts, ranging from word sense disambiguation(Yarowsky, 1995), relation extraction(Brin, 1999;Agichtein and Gravano, 2000), parsing(McClosky et al., 2006;Reichart and Rappoport, 2007), sequence generation(He et al., 2020), few-shot text classification(Schick and Schütze, 2021a)and retrieval(Izacard and Grave, 2021)to reasoning(Zelikman et al., 2022). In a similar spirit to these approaches, Toolformer is trained on its own predictions after applying a perplexity-based filtering step.

7Limitations

While our approach enables LMs to learn how to use a variety of tools in a self-supervised way, there are some clear limitations to what can be achieved with our method in its current form. One such limitation is the inability of Toolformer to use tools in achain(i.e., using the output of one tool as an input for another tool). This is due to the fact that API calls for each tool are generated independently; as a consequence, there are no examples of chained tool use in the finetuning dataset. Our current approach also does not allow the LM to use a tool in aninteractiveway – especially for tools such as search engines, that could potentially return hundreds of different results, enabling a LM to browse through these results or to refine its search query in a similar spirit toNakano et al. (2021)can be crucial for certain applications. Beyond this, we found models trained with Toolformer to often be sensitive to the exact wording of their input when deciding whether or not to call an API; this is perhaps unsurprising given that LMs are known to be very sensitive to the prompt they are provided with in both zero-and few-shot settings(Jiang et al., 2020;Schick and Schütze, 2021a). Depending on the tool, our method is also very sample-inefficient; for example, processing more than a million documents results in only a few thousand examples of useful calls to the calculator API. A potential solution to this problem might be to iteratively apply our approach, similar to how this is done in related bootstrapping approaches(Schick and Schütze, 2021a;Izacard and Grave, 2021;Parisi et al., 2022). Finally, when deciding whether or not to make an API call, Toolformer currently does not take into account the tool-dependent, computational cost incurred from making an API call.

8Conclusion

We have introduced Toolformer, a language model that learns in a self-supervised way how to use different tools such as search engines, calculators, and translation systems via simple API calls. This is done by finetuning on a large number of sampled API calls that are filtered based on whether they reduce perplexity on future tokens. Toolformer considerably improves zero-shot performance of a 6.7B parameter GPT-J model, enabling it to even outperform a much larger GPT-3 model on a range of different downstream tasks.

References

  • Aghajanyan et al. (2021)Armen Aghajanyan, Dmytro Okhonko, Mike Lewis, Mandar Joshi, Hu Xu, Gargi Ghosh, and Luke Zettlemoyer. 2021.Htlm: Hyper-text pre-training and prompting of language models.
  • Agichtein and Gravano (2000)Eugene Agichtein and Luis Gravano. 2000.Snowball: Extracting relations from large plain-text collections.InProceedings of the Fifth ACM Conference on Digital Libraries, DL ’00, page 85–94, New York, NY, USA. Association for Computing Machinery.
  • Baeza-Yates et al. (1999)Ricardo Baeza-Yates, Berthier Ribeiro-Neto, et al. 1999.Modern information retrieval, volume 463.ACM press New York.
  • Berant et al. (2013)Jonathan Berant, Andrew Chou, Roy Frostig, and Percy Liang. 2013.Semantic parsing on Freebase from question-answer pairs.InProceedings of the 2013 Conference on Empirical Methods in Natural Language Processing, pages 1533–1544, Seattle, Washington, USA. Association for Computational Linguistics.
  • Borgeaud et al. (2021)Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George van den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, Diego de Las Casas, Aurelia Guy, Jacob Menick, Roman Ring, Tom Hennigan, Saffron Huang, Loren Maggiore, Chris Jones, Albin Cassirer, Andy Brock, Michela Paganini, Geoffrey Irving, Oriol Vinyals, Simon Osindero, Karen Simonyan, Jack W. Rae, Erich Elsen, and Laurent Sifre. 2021.Improving language models by retrieving from trillions of tokens.
  • Brin (1999)Sergey Brin. 1999.Extracting patterns and relations from the world wide web.InThe World Wide Web and Databases, pages 172–183, Berlin, Heidelberg. Springer Berlin Heidelberg.
  • Brown et al. (2020)Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, Chris Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gray, Benjamin Chess, Jack Clark, Christopher Berner, Sam McCandlish, Alec Radford, Ilya Sutskever, and Dario Amodei. 2020.Language models are few-shot learners.InAdvances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.
  • Chowdhery et al. (2022)Aakanksha Chowdhery, Sharan Narang, Jacob Devlin, Maarten Bosma, Gaurav Mishra, Adam Roberts, Paul Barham, Hyung Won Chung, Charles Sutton, Sebastian Gehrmann, Parker Schuh, Kensen Shi, Sasha Tsvyashchenko, Joshua Maynez, Abhishek Rao, Parker Barnes, Yi Tay, Noam Shazeer, Vinodkumar Prabhakaran, Emily Reif, Nan Du, Ben Hutchinson, Reiner Pope, James Bradbury, Jacob Austin, Michael Isard, Guy Gur-Ari, Pengcheng Yin, Toju Duke, Anselm Levskaya, Sanjay Ghemawat, Sunipa Dev, Henryk Michalewski, Xavier Garcia, Vedant Misra, Kevin Robinson, Liam Fedus, Denny Zhou, Daphne Ippolito, David Luan, Hyeontaek Lim, Barret Zoph, Alexander Spiridonov, Ryan Sepassi, David Dohan, Shivani Agrawal, Mark Omernick, Andrew M. Dai, Thanumalayan Sankaranarayana Pillai, Marie Pellat, Aitor Lewkowycz, Erica Moreira, Rewon Child, Oleksandr Polozov, Katherine Lee, Zongwei Zhou, Xuezhi Wang, Brennan Saeta, Mark Diaz, Orhan Firat, Michele Catasta, Jason Wei, Kathy Meier-Hellstern, Douglas Eck, Jeff Dean, Slav Petrov, and Noah Fiedel. 2022.Palm: Scaling language modeling with pathways.
  • Cobbe et al. (2021)Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. 2021.Training verifiers to solve math word problems.arXiv preprint arXiv:2110.14168.
  • Costa-jussà et al. (2022)Marta R Costa-jussà, James Cross, Onur Çelebi, Maha Elbayad, Kenneth Heafield, Kevin Heffernan, Elahe Kalbassi, Janice Lam, Daniel Licht, Jean Maillard, et al. 2022.No language left behind: Scaling human-centered machine translation.arXiv preprint arXiv:2207.04672.
  • Devlin et al. (2019)Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019.BERT: Pre-training of deep bidirectional transformers for language understanding.InProceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota. Association for Computational Linguistics.
  • Dhingra et al. (2022)Bhuwan Dhingra, Jeremy R. Cole, Julian Martin Eisenschlos, Daniel Gillick, Jacob Eisenstein, and William W. Cohen. 2022.Time-aware language models as temporal knowledge bases.Transactions of the Association for Computational Linguistics, 10:257–273.
  • Gao et al. (2020)Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, et al. 2020.The pile: An 800gb dataset of diverse text for language modeling.arXiv preprint arXiv:2101.00027.
  • Gao et al. (2022)Luyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon, Pengfei Liu, Yiming Yang, Jamie Callan, and Graham Neubig. 2022.Pal: Program-aided language models.
  • Guu et al. (2020)Kelvin Guu, Kenton Lee, Zora Tung, Panupong Pasupat, and Ming-Wei Chang. 2020.Realm: Retrieval-augmented language model pre-training.
  • He et al. (2020)Junxian He, Jiatao Gu, Jiajun Shen, and Marc’Aurelio Ranzato. 2020.Revisiting self-training for neural sequence generation.InInternational Conference on Learning Representations.
  • Honovich et al. (2022)Or Honovich, Thomas Scialom, Omer Levy, and Timo Schick. 2022.Unnatural instructions: Tuning language models with (almost) no human labor.
  • Izacard and Grave (2021)Gautier Izacard and Edouard Grave. 2021.Distilling knowledge from reader to retriever for question answering.InInternational Conference on Learning Representations.
  • Izacard et al. (2022)Gautier Izacard, Patrick Lewis, Maria Lomeli, Lucas Hosseini, Fabio Petroni, Timo Schick, Jane Dwivedi-Yu, Armand Joulin, Sebastian Riedel, and Edouard Grave. 2022.Atlas: Few-shot learning with retrieval augmented language models.
  • Ji et al. (2022)Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu, Etsuko Ishii, Yejin Bang, Andrea Madotto, and Pascale Fung. 2022.Survey of hallucination in natural language generation.ACM Computing Surveys.
  • Jiang et al. (2020)Zhengbao Jiang, Frank F. Xu, Jun Araki, and Graham Neubig. 2020.How can we know what language models know?Transactions of the Association for Computational Linguistics, 8:423–438.
  • Joshi et al. (2017)Mandar Joshi, Eunsol Choi, Daniel Weld, and Luke Zettlemoyer. 2017.TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension.InProceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1601–1611, Vancouver, Canada. Association for Computational Linguistics.
  • Joulin et al. (2016)Armand Joulin, Edouard Grave, Piotr Bojanowski, Matthijs Douze, Hérve Jégou, and Tomas Mikolov. 2016.Fasttext. zip: Compressing text classification models.arXiv preprint arXiv:1612.03651.
  • Keskar et al. (2019)Nitish Shirish Keskar, Bryan McCann, Lav R. Varshney, Caiming Xiong, and Richard Socher. 2019.Ctrl: A conditional transformer language model for controllable generation.
  • Koehn (2005)Philipp Koehn. 2005.Europarl: A parallel corpus for statistical machine translation.InProceedings of machine translation summit x: papers, pages 79–86.
  • Komeili et al. (2022)Mojtaba Komeili, Kurt Shuster, and Jason Weston. 2022.Internet-augmented dialogue generation.InProceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 8460–8478, Dublin, Ireland. Association for Computational Linguistics.
  • Koncel-Kedziorski et al. (2016)Rik Koncel-Kedziorski, Subhro Roy, Aida Amini, Nate Kushman, and Hannaneh Hajishirzi. 2016.MAWPS: A math word problem repository.InProceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 1152–1157, San Diego, California. Association for Computational Linguistics.
  • Kwiatkowski et al. (2019)Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming-Wei Chang, Andrew M. Dai, Jakob Uszkoreit, Quoc Le, and Slav Petrov. 2019.Natural questions: A benchmark for question answering research.Transactions of the Association for Computational Linguistics, 7:452–466.
  • Lazaridou et al. (2022)Angeliki Lazaridou, Elena Gribovskaya, Wojciech Stokowiec, and Nikolai Grigorev. 2022.Internet-augmented language models through few-shot prompting for open-domain question answering.arXiv preprint arXiv:2203.05115.
  • Lewis et al. (2019)Patrick Lewis, Barlas Oğuz, Ruty Rinott, Sebastian Riedel, and Holger Schwenk. 2019.Mlqa: Evaluating cross-lingual extractive question answering.arXiv preprint arXiv:1910.07475.
  • Lin et al. (2021)Xi Victoria Lin, Todor Mihaylov, Mikel Artetxe, Tianlu Wang, Shuohui Chen, Daniel Simig, Myle Ott, Naman Goyal, Shruti Bhosale, Jingfei Du, Ramakanth Pasunuru, Sam Shleifer, Punit Singh Koura, Vishrav Chaudhary, Brian O’Horo, Jeff Wang, Luke Zettlemoyer, Zornitsa Kozareva, Mona Diab, Veselin Stoyanov, and Xian Li. 2021.Few-shot learning with multilingual language models.
  • Maynez et al. (2020)Joshua Maynez, Shashi Narayan, Bernd Bohnet, and Ryan McDonald. 2020.On faithfulness and factuality in abstractive summarization.
  • McClosky et al. (2006)David McClosky, Eugene Charniak, and Mark Johnson. 2006.Effective self-training for parsing.InProceedings of the Human Language Technology Conference of the NAACL, Main Conference, pages 152–159, New York City, USA. Association for Computational Linguistics.
  • Merity et al. (2017)Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017.Pointer sentinel mixture models.InInternational Conference on Learning Representations.
  • Miao et al. (2020)Shen-yun Miao, Chao-Chun Liang, and Keh-Yih Su. 2020.A diverse corpus for evaluating and developing English math word problem solvers.InProceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 975–984, Online. Association for Computational Linguistics.
  • Nakano et al. (2021)Reiichiro Nakano, Jacob Hilton, Suchir Balaji, Jeff Wu, Long Ouyang, Christina Kim, Christopher Hesse, Shantanu Jain, Vineet Kosaraju, William Saunders, Xu Jiang, Karl Cobbe, Tyna Eloundou, Gretchen Krueger, Kevin Button, Matthew Knight, Benjamin Chess, and John Schulman. 2021.Webgpt: Browser-assisted question-answering with human feedback.
  • Parisi et al. (2022)Aaron Parisi, Yao Zhao, and Noah Fiedel. 2022.Talm: Tool augmented language models.
  • Patel et al. (2021)Arkil Patel, Satwik Bhattamishra, and Navin Goyal. 2021.Are NLP models really able to solve simple math word problems?InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2080–2094, Online. Association for Computational Linguistics.
  • Petroni et al. (2021)Fabio Petroni, Aleksandra Piktus, Angela Fan, Patrick Lewis, Majid Yazdani, Nicola De Cao, James Thorne, Yacine Jernite, Vladimir Karpukhin, Jean Maillard, Vassilis Plachouras, Tim Rocktäschel, and Sebastian Riedel. 2021.KILT: a benchmark for knowledge intensive language tasks.InProceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 2523–2544, Online. Association for Computational Linguistics.
  • Petroni et al. (2019)Fabio Petroni, Tim Rocktäschel, Sebastian Riedel, Patrick Lewis, Anton Bakhtin, Yuxiang Wu, and Alexander Miller. 2019.Language models as knowledge bases?InProceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2463–2473, Hong Kong, China. Association for Computational Linguistics.
  • Radford et al. (2019)Alec Radford, Jeffrey Wu, Rewon Child, David Luan, Dario Amodei, Ilya Sutskever, et al. 2019.Language models are unsupervised multitask learners.OpenAI blog, 1(8):9.
  • Reichart and Rappoport (2007)Roi Reichart and Ari Rappoport. 2007.Self-training for enhancement and domain adaptation of statistical parsers trained on small datasets.InProceedings of the 45th Annual Meeting of the Association of Computational Linguistics, pages 616–623, Prague, Czech Republic. Association for Computational Linguistics.
  • Robertson et al. (1995)Stephen E Robertson, Steve Walker, Susan Jones, Micheline M Hancock-Beaulieu, Mike Gatford, et al. 1995.Okapi at trec-3.Nist Special Publication Sp, 109:109.
  • Schick et al. (2022)Timo Schick, Jane Dwivedi-Yu, Zhengbao Jiang, Fabio Petroni, Patrick Lewis, Gautier Izacard, Qingfei You, Christoforos Nalmpantis, Edouard Grave, and Sebastian Riedel. 2022.Peer: A collaborative language model.
  • Schick and Schütze (2021a)Timo Schick and Hinrich Schütze. 2021a.Exploiting cloze-questions for few-shot text classification and natural language inference.InProceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 255–269, Online. Association for Computational Linguistics.
  • Schick and Schütze (2021b)Timo Schick and Hinrich Schütze. 2021b.Generating datasets with pretrained language models.InProceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6943–6951, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
  • Shuster et al. (2022)Kurt Shuster, Jing Xu, Mojtaba Komeili, Da Ju, Eric Michael Smith, Stephen Roller, Megan Ung, Moya Chen, Kushal Arora, Joshua Lane, Morteza Behrooz, William Ngan, Spencer Poff, Naman Goyal, Arthur Szlam, Y-Lan Boureau, Melanie Kambadur, and Jason Weston. 2022.Blenderbot 3: a deployed conversational agent that continually learns to responsibly engage.
  • Thoppilan et al. (2022)Romal Thoppilan, Daniel De Freitas, Jamie Hall, Noam Shazeer, Apoorv Kulshreshtha, Heng-Tze Cheng, Alicia Jin, Taylor Bos, Leslie Baker, Yu Du, YaGuang Li, Hongrae Lee, Huaixiu Steven Zheng, Amin Ghafouri, Marcelo Menegali, Yanping Huang, Maxim Krikun, Dmitry Lepikhin, James Qin, Dehao Chen, Yuanzhong Xu, Zhifeng Chen, Adam Roberts, Maarten Bosma, Vincent Zhao, Yanqi Zhou, Chung-Ching Chang, Igor Krivokon, Will Rusch, Marc Pickett, Pranesh Srinivasan, Laichee Man, Kathleen Meier-Hellstern, Meredith Ringel Morris, Tulsee Doshi, Renelito Delos Santos, Toju Duke, Johnny Soraker, Ben Zevenbergen, Vinodkumar Prabhakaran, Mark Diaz, Ben Hutchinson, Kristen Olson, Alejandra Molina, Erin Hoffman-John, Josh Lee, Lora Aroyo, Ravi Rajakumar, Alena Butryna, Matthew Lamm, Viktoriya Kuzmina, Joe Fenton, Aaron Cohen, Rachel Bernstein, Ray Kurzweil, Blaise Aguera-Arcas, Claire Cui, Marian Croak, Ed Chi, and Quoc Le. 2022.Lamda: Language models for dialog applications.
  • Wang and Komatsuzaki (2021)Ben Wang and Aran Komatsuzaki. 2021.GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model.https://github.com/kingoflolz/mesh-transformer-jax.
  • Wang et al. (2022)Yizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu, Noah A. Smith, Daniel Khashabi, and Hannaneh Hajishirzi. 2022.Self-instruct: Aligning language model with self generated instructions.
  • Wei et al. (2022)Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean, and William Fedus. 2022.Emergent abilities of large language models.
  • Wenzek et al. (2020)Guillaume Wenzek, Marie-Anne Lachaux, Alexis Conneau, Vishrav Chaudhary, Francisco Guzmán, Armand Joulin, and Edouard Grave. 2020.CCNet: Extracting high quality monolingual datasets from web crawl data.InProceedings of the Twelfth Language Resources and Evaluation Conference, pages 4003–4012, Marseille, France. European Language Resources Association.
  • Yao et al. (2022)Shunyu Yao, Jeffrey Zhao, Dian Yu, Nan Du, Izhak Shafran, Karthik Narasimhan, and Yuan Cao. 2022.React: Synergizing reasoning and acting in language models.
  • Yarowsky (1995)David Yarowsky. 1995.Unsupervised word sense disambiguation rivaling supervised methods.In33rd Annual Meeting of the Association for Computational Linguistics, pages 189–196, Cambridge, Massachusetts, USA. Association for Computational Linguistics.
  • Zelikman et al. (2022)Eric Zelikman, Yuhuai Wu, Jesse Mu, and Noah D. Goodman. 2022.Star: Bootstrapping reasoning with reasoning.
  • Zhang et al. (2022)Susan Zhang, Stephen Roller, Naman Goyal, Mikel Artetxe, Moya Chen, Shuohui Chen, Christopher Dewan, Mona Diab, Xian Li, Xi Victoria Lin, Todor Mihaylov, Myle Ott, Sam Shleifer, Kurt Shuster, Daniel Simig, Punit Singh Koura, Anjali Sridhar, Tianlu Wang, and Luke Zettlemoyer. 2022.Opt: Open pre-trained transformer language models.

Appendix AAPI Details

When sampling and filtering API calls, by default we use values ofτs=0.05\tau_{s}=0.05andτf=1.0\tau_{f}=1.0– i.e., we only make API calls at positions where the probability of the<API>token is at least 5%, and we keep API calls if they reduce the loss by at least 1.0. We only keep the topk=5k=5such positions and sample up tom=5m=5API calls for each position identified in a piece of text. Due to the heuristic filtering described below, we generate API calls for the calculator and machine translation system on only a small subset of𝒞\mathcal{C}; to compensate for this, we setτs=0.0\tau_{s}=0.0,k=20k=20andm=10m=10for these tools. As the resulting sets of API calls are still comparably small, we additionally setτf=0.5\tau_{f}=0.5.

A.1Implementation

Question Answering

We use the Atlas model ofIzacard et al. (2022)finetuned on Natural Questions(Kwiatkowski et al., 2019)as our question answering system. For creating𝒞∗\mathcal{C}^{*}we use Atlas-large, enabling us to efficiently process millions of API calls; during inference, we use the larger Atlas-xxl model.

Calculator

Our calculator is based on a simple Python script and only supports the operators “++”, “−-”, “∗*”, and “//”. It does not return any result for syntactically invalid equations. For sampling API calls, we apply heuristic filters to our subset of CCNet and only process documents that either (i) contain at least three numbers within a window of 100 tokens, where one of these numbers is the result of applying a mathematical operation to the other two, (ii) contain one of the sequences “=”, “equals”, “equal to”, “total of”, “average of” followed by a number, or (iii) contain at least three numbers; for texts that only match the last criterion, we only keep a random subset of 1%.

Calendar

For creating our dataset𝒞∗\mathcal{C}^{*}, we operate under the assumption that the calendar date in such cases should be the date that the document was created. We approximate this by extracting the date from the URL, if it is present. We filter out texts for which a date cannot be extracted, leaving around 18% of the documents.

Machine Translation

For both training and inference, we use the 600M parameter NLLB(Costa-jussà et al., 2022)as our machine translation (MT) model. The source language is automatically detected using the fastText classifier(Joulin et al., 2016), while the target language is always set to English. Since most of the CCNet dataset is in English, we filter out the parts that contain only English text before generating API calls. More specifically, we only keep those paragraphs which contain text chunks in a language other than English preceded and followed by English text. We use text chunks of size 10 tokens. To determine whether the middle text chunk is in a language different than English we again use the fastText classifier with a confidence greater than 0.8. We also filter out any text chunks that contain only numbers or special symbols. This filtering mechanism allows us to generate data more efficiently by focusing our API call generations in places where the MT tool is likely to be helpful. After generating the MT API calls, we additionally remove from our training set those where the input to the MT tool appears after the API call but not before it. While during data generation the model can look ahead to generate API calls, this is not possible at inference time, so we want to dissuade the model from calling the API in such cases.

A.2Prompts

Below, we list the prompts used to sample API calls for each tool considered.

Question Answering

We use the following prompt for the question answering tool:

Your task is to add calls to a Question Answering API to a piece of text. The questions should help you get information required to complete the text. You can call the API by writing "[QA(question)]" where "question" is the question you want to ask. Here are some examples of API calls:
Input: Joe Biden was born in Scranton, Pennsylvania.
Output: Joe Biden was born in [QA("Where was Joe Biden born?")] Scranton, [QA("In which state is Scranton?")] Pennsylvania.
Input: Coca-Cola, or Coke, is a carbonated soft drink manufactured by the Coca-Cola Company.
Output: Coca-Cola, or [QA("What other name is Coca-Cola known by?")] Coke, is a carbonated soft drink manufactured by [QA("Who manufactures Coca-Cola?")] the Coca-Cola Company.
Input: x
Output:
Calculator

We use the following prompt for the calculator:

Your task is to add calls to a Calculator API to a piece of text. The calls should help you get information required to complete the text. You can call the API by writing "[Calculator(expression)]" where "expression" is the expression to be computed. Here are some examples of API calls:
Input: The number in the next term is 18 + 12 x 3 = 54.
Output: The number in the next term is 18 + 12 x 3 = [Calculator(18 + 12 * 3)] 54.
Input: The population is 658,893 people. This is 11.4% of the national average of 5,763,868 people.
Output: The population is 658,893 people. This is 11.4% of the national average of [Calculator(658,893 / 11.4%)] 5,763,868 people.
Input: A total of 252 qualifying matches were played, and 723 goals were scored (an average of 2.87 per match). This is three times less than the 2169 goals last year.
Output: A total of 252 qualifying matches were played, and 723 goals were scored (an average of [Calculator(723 / 252)] 2.87 per match). This is twenty goals more than the [Calculator(723 - 20)] 703 goals last year.
Input: I went to Paris in 1994 and stayed there until 2011, so in total, it was 17 years.
Output: I went to Paris in 1994 and stayed there until 2011, so in total, it was [Calculator(2011 - 1994)] 17 years.
Input: From this, we have 4 * 30 minutes = 120 minutes.
Output: From this, we have 4 * 30 minutes = [Calculator(4 * 30)] 120 minutes.
Input: x
Output:
Wikipedia Search

We use the following prompt for the Wikipedia search tool:

Your task is to complete a given piece of text. You can use a Wikipedia Search API to look up information. You can do so by writing "[WikiSearch(term)]" where "term" is the search term you want to look up. Here are some examples of API calls:
Input: The colors on the flag of Ghana have the following meanings: red is for the blood of martyrs, green for forests, and gold for mineral wealth.
Output: The colors on the flag of Ghana have the following meanings: red is for [WikiSearch("Ghana flag red meaning")] the blood of martyrs, green for forests, and gold for mineral wealth.
Input: But what are the risks during production of nanomaterials? Some nanomaterials may give rise to various kinds of lung damage.
Output: But what are the risks during production of nanomaterials? [WikiSearch("nanomaterial production risks")] Some nanomaterials may give rise to various kinds of lung damage.
Input: Metformin is the first-line drug for patients with type 2 diabetes and obesity.
Output: Metformin is the first-line drug for [WikiSearch("Metformin first-line drug")] patients with type 2 diabetes and obesity.
Input: x
Output:
Machine Translation

We use the following prompt for the machine translation tool:

Your task is to complete a given piece of text by using a Machine Translation API.
You can do so by writing "[MT(text)]" where text is the text to be translated into English.
Here are some examples:
Input: He has published one book: O homem suprimido (“The Supressed Man”)
Output: He has published one book: O homem suprimido [MT(O homem suprimido)] (“The Supressed Man”)
Input: In Morris de Jonge’s Jeschuah, der klassische jüdische Mann, there is a description of a Jewish writer
Output: In Morris de Jonge’s Jeschuah, der klassische jüdische Mann [MT(der klassische jüdische Mann)], there is a description of a Jewish writer
Input: 南京高淳县住房和城乡建设局 城市新区设计 a plane of reference Gaochun is one of seven districts of the provincial capital Nanjing
Output: [MT(南京高淳县住房和城乡建设局 城市新区设计)] a plane of reference Gaochun is one of seven districts of the provincial capital Nanjing
Input: x
Output:
Calendar

We use the following prompt for the calendar tool:

Your task is to add calls to a Calendar API to a piece of text. The API calls should help you get information required to complete the text. You can call the API by writing "[Calendar()]" Here are some examples of API calls:
Input: Today is the first Friday of the year.
Output: Today is the first [Calendar()] Friday of the year.
Input: The president of the United States is Joe Biden.
Output: The president of the United States is [Calendar()] Joe Biden.
Input: The current day of the week is Wednesday.
Output: The current day of the week is [Calendar()] Wednesday.
Input: The number of days from now until Christmas is 30.
Output: The number of days from now until Christmas is [Calendar()] 30.
Input: The store is never open on the weekend, so today it is closed.
Output: The store is never open on the weekend, so today [Calendar()] it is closed.
Input: x
Output:

Appendix BToolformer Training

We use up to 25k examples per API. Max sequence length 1,024. Effective batch size of 128. All models are trained using DeepSpeed’s ZeRO-3 (Rasley et al., 2020). We used 8 NVIDIA A100 40GB GPUs with BF16. Training up to 2k steps, where we evaluate PPL on a small development set from CCNet containing 1,000 examples every 500 steps. We pick the checkpoint that performs best.

Appendix CZero-Shot Prompts

C.1LAMA andTempLAMA

For both LAMA andTempLAMA, given an input text𝐱\mathbf{x}, we use the following prompt:Please complete the following text so that it is factually correct:𝐱\mathbf{x}.

C.2Math Benchmarks

For all math benchmarks, given a context𝐱\mathbf{x}and a question𝐪\mathbf{q}, our prompt is:𝐱​𝐪\mathbf{x}\ \mathbf{q}The answer is.

C.3Question Answering

For all question answering datasets, includingDateset, we simply prefix the question withAnswer the following question:. We append a question mark if the question does not already end with one.

C.4Multilingual Question Answering

For MLQA, given a context𝐱\mathbf{x}and a question𝐪\mathbf{q}, our prompt is:Your task is to answer a question based on the following paragraph:𝐱\mathbf{x}Now answer the following question in English:𝐪\mathbf{q}.

Appendix DDateset

Datesetis created by first randomly selecting 500 “current dates”. For each current date, another relatively past/future date is randomly selected within a four-year range, and the two dates are used to fill the query templates in Table11. An example of one such query using the first template would be, “How many days ago was August 14, 2020?” If called, the Calendar tool would return the presumed current date (e.g., “Today is Sunday, November 20, 2020”).

Table 11:Templates used to createDatesetwhere acurrent_dateis randomly selected. For eachcurrent_date, a randompast_dateandfuture_dateis generated and used to fill each template, if relevant. The federal holidays in the United States (e.g., Thanksgiving) were used in the templates involving holidays.

Similar Articles

LLM Agents Already Know When to Call Tools -- Even Without Reasoning

Hugging Face Daily Papers

This paper introduces When2Tool, a benchmark to study when LLM agents actually need to call tools, and reveals that models already know tool necessity from hidden states but fail to act. The proposed Probe&Prefill method reduces unnecessary tool calls by 48% with minimal accuracy loss.

Beyond APIs: Probing the Limits of MLLMs in Physical Tool Use

arXiv cs.CL

This paper introduces PhysTool-Bench, a benchmark for evaluating multimodal large language models' ability to recognize and plan the use of physical tools in real-world scenes. The authors find that even the best model identifies only 58.7% of tools and completes just 21.0% of queries end-to-end, revealing a two-level deficit in perception and functional commonsense.