LAVOIR: Teaching a Single-Pass Decision Encoder When and What to Ask with Amortized Value of Information
Summary
LAVOIR introduces a method for single-pass decision models to ask questions based on amortized value of information, improving accuracy and efficiency in handling underspecified inputs.
View Cached Full Text
Cached at: 09/28/26, 09:42 AM
# Teaching a Single-Pass Decision EncoderWhen and What to Ask with Amortized Value of Information
Source: [https://arxiv.org/html/2609.30706](https://arxiv.org/html/2609.30706)
September 2026
###### Abstract
“System One” decision models such as TypeSafe’s Jev and its open counterpart Laya answer typed questions about a text in a single forward pass with calibrated probabilities, but they cannot ask for missing information: when a first message does not say what separates two departments, they guess\. We present LAVOIR \(Laya withValue\-Of\-InformationRouting\), which places the candidate pieces of missing information \(*slots*\) in the input next to the answer options, so that one forward pass returns both the decision distribution and, for every slot, the expected gain in the probability of the correct decision if the user were asked about it\. VOI targets need no human labels: gold decisions come from schema rules, an LLM only verbalizes messages and answers, a model from another family checks every text, and pairing each message with several profiles makes regression on realized gains estimate the expected gain\. A Gini\-impurity cap bounds the predicted value by what a calibrated model can still gain\. In a controlled study, decisions on seen schemas are statistically indistinguishable from the Bayes ceiling\. The final model’s question policy matches a greedy oracle VOI policy on seen schemas \(AUC 0\.799 vs\. 0\.797\), and with at most 0\.5 questions per conversation it is 14\.1 points more accurate than never asking\. On real ABCD conversations, one real exchange raises accuracy by 8\.3 points where LAVOIR asks and leaves it unchanged where it does not; on SGD the cap lowers the asking rate from 93% to 8\.6%\. On Laya’s twelve benchmarks LAVOIR is above Laya’s reported scores on seven, and it answers a question in 31 ms \(median, GH200\)\.
## 1Introduction
Many production decisions over text are small, typed and frequent: which team should handle a ticket, whether an e\-mail is phishing, which tool to call\. LLMs can make them, but they generate text token by token and do not expose calibrated probabilities\. TypeSafe introduced Jev\([Almeida, 2026](https://arxiv.org/html/2609.30706#bib.bib2)\)as the first*System One*model, named after the fast mode of thinking in[Kahneman \(2011\)](https://arxiv.org/html/2609.30706#bib.bib38): it returns typed decisions with probabilities from one parallel query, and its weights are not public\. Laya\([Convai Innovations, 2026](https://arxiv.org/html/2609.30706#bib.bib20)\), released three days later as an open model with a Jev\-compatible interface, places every answer option behind its own\[MASK\]marker in the input of a bidirectional encoder\([Warner et al\., 2024](https://arxiv.org/html/2609.30706#bib.bib73)\)and scores all options in one pass; the question schema is given at request time, so a new decision needs no retraining\.
These models must decide from the text they are given, and real first messages are often underspecified\.*“My information was shared without permission, what can I do?”*may belong to the data\-protection or to the legal team, depending on what was shared and what the customer wants\. A human agent asks one question; a single\-pass model guesses or escalates every uncertain case \(Figure[1](https://arxiv.org/html/2609.30706#S1.F1)\)\.
Laya: decide or hand off“hi, i need help with my recent purchase\. i would like a replacement for the item\. can you help me with that? thanks”p\(team∣message\)p\(\\text\{team\}\\mid\\text\{message\}\)refunds\.36returns\_desk\.29marketplace\_support\.22warranty\_service\.07logistics\.06confidence0\.12<0\.850\.12<0\.85⇒\\Rightarrowhand off to a human
best guessrefunds\(36%\): wrongLAVOIR: ask, then decide“hi, i need help with my recent purchase\. i would like a replacement for the item\. can you help me with that? thanks”VOI per slot \(loop 0\),c=0\.05c=0\.05seller\.197problem\.171delivery\_age\.036askseller:“Was the item sold by us directly or by a seller on our marketplace?”
“The item was sold by our store\.”
nowreturns\_desk\.45,logistics\.43: still unsureaskproblem\(VOI \.385\): “What is the problem with your order?”
“The item arrived damaged\.”max VOI≤0\.05\\leq 0\.05⇒\\Rightarrowdecidelogistics\(p=1\.00p=1\.00\)✓\\checkmark
without asking it would have pickedreturns\_desk\(\.35\): wrongFigure 1:The same underspecified first message \(a case of theecommerce\_returnsschema\) given to Laya and to LAVOIR\. Laya’s confidence \(one minus the normalized entropy, as Laya computes it\) stays at 0\.12, so its recommended usage \(hand off below 0\.85\) sends the conversation to a human, and its best guess is wrong anyway; Laya never saw this schema in training\. LAVOIR asks about the seller and then about the problem, skips the three slots whose answers would not change the decision, and routes to the correct team\.LAVOIR adds a second marker block for slots after the answer options\. In the same forward pass that produces the decision distribution, a small*VOI head*reads each slot marker and predicts how much asking about that slot would raise the probability of the correct decision; the model asks about the best slot if its value exceeds a question cost and otherwise decides or hands off \(Figure[2](https://arxiv.org/html/2609.30706#S3.F2)\)\. The question*content*is chosen by the model; its*wording*is a template attached to the slot\. The training signal needs no human labels: schema rules map complete user profiles to gold units, an LLM only verbalizes partial profiles, and because every message is paired with several profiles that differ in their hidden values, least\-squares regression on the realized gainpk\(y⋆\)−p0\(y⋆\)p\_\{k\}\(y^\{\\star\}\)\-p\_\{0\}\(y^\{\\star\}\)estimates the expected gain, an amortized value of information\([Howard, 1966](https://arxiv.org/html/2609.30706#bib.bib35);[Rao and Daumé III, 2018](https://arxiv.org/html/2609.30706#bib.bib62)\)\. Our contributions are:
- •Amortized VOI for typed decisions: a slot block, a zero\-initialized segment embedding \(with no slots the model reproduces Laya’s logits bit for bit\), a VOI head, and a Gini cap that bounds predicted value by the maximum expected gain of a calibrated model\.
- •Label\-free VOI targetsfrom rule\-defined gold with an exact posterior, an answer simulator, and two\-way cross\-family checking, together with the schema design lessons that made generation work\.
- •An evaluation against exact references: the decisions match the Bayes ceiling and the question policy matches an oracle policy; on real conversations the questions go where answers help\.
- •Findings about Laya’s recipe: its policy\-gradient term brings no gain and slows the VOI head, and one temperature per question type miscalibrates heterogeneous mixtures\.
## 2Related Work
#### System One models and decision encoders\.
Jev\([Almeida, 2026](https://arxiv.org/html/2609.30706#bib.bib2)\)returns typed answers \(categories, scores, a choice among up to 255 options\) with probabilities, decodes in parallel and is trained with what TypeSafe calls Reinforcement Learning for Calibrated Decisions \(RLCD\); it is API\-only, with 236–276 ms median latency in independent measurements\([AbdelStark, 2026](https://arxiv.org/html/2609.30706#bib.bib1);[nibzard, 2026](https://arxiv.org/html/2609.30706#bib.bib59)\)\. Laya\([Convai Innovations, 2026](https://arxiv.org/html/2609.30706#bib.bib20)\)implements this interface on ModernBERT\-large\([Warner et al\., 2024](https://arxiv.org/html/2609.30706#bib.bib73)\)or mmBERT\([Marone et al\., 2025](https://arxiv.org/html/2609.30706#bib.bib53)\), trains with strictly proper scoring rules\([Gneiting and Raftery, 2007](https://arxiv.org/html/2609.30706#bib.bib31)\)plus a group\-baseline policy\-gradient term\([Williams, 1992](https://arxiv.org/html/2609.30706#bib.bib77);[Shao et al\., 2024](https://arxiv.org/html/2609.30706#bib.bib67)\), and calibrates with per\-type temperatures\([Guo et al\., 2017](https://arxiv.org/html/2609.30706#bib.bib32)\)\. GLiNER and GLiClass\([Zaratiana et al\., 2024](https://arxiv.org/html/2609.30706#bib.bib81);[Stepanov et al\., 2025](https://arxiv.org/html/2609.30706#bib.bib70)\)encode labels jointly with the text, and SCX Router\([Stepanov et al\., 2026](https://arxiv.org/html/2609.30706#bib.bib69)\)scores label tokens against a decoder cache for model routing\([Ong et al\., 2024](https://arxiv.org/html/2609.30706#bib.bib60)\)\. None of these models can ask for missing information\.
#### Clarifying questions\.
[Rao and Daumé III \(2018\)](https://arxiv.org/html/2609.30706#bib.bib62)rank clarification questions by the expected value of perfect information, generating candidate answers\. Qulac and ClariQ\([Aliannejadi et al\., 2019](https://arxiv.org/html/2609.30706#bib.bib4);[Aliannejadi et al\., 2021](https://arxiv.org/html/2609.30706#bib.bib3)\)study clarification in search,[Yu et al\. \(2020\)](https://arxiv.org/html/2609.30706#bib.bib80)ask binary questions chosen by information gain, and LLM\-based work decides when to clarify\([Kuhn et al\., 2022](https://arxiv.org/html/2609.30706#bib.bib42);[Zhang and Choi, 2023](https://arxiv.org/html/2609.30706#bib.bib82)\), trains models to ask\([Andukuri et al\., 2024](https://arxiv.org/html/2609.30706#bib.bib5)\), or simulates answers to plan questions\([Hu et al\., 2024](https://arxiv.org/html/2609.30706#bib.bib36)\)\. LAVOIR instead predicts the value of every candidate question in one encoder pass, without generating or simulating answers at inference\.
#### Value of information, deferral and conformal prediction\.
VOI\([Howard, 1966](https://arxiv.org/html/2609.30706#bib.bib35)\)and expected information\([Lindley, 1956](https://arxiv.org/html/2609.30706#bib.bib49)\)are the classical criteria for choosing what to observe; deep adaptive design amortizes sequential design into a policy network\([Foster et al\., 2021](https://arxiv.org/html/2609.30706#bib.bib27)\), and POMDP dialogue managers trade off asking and acting with simulated users\([Williams and Young, 2007](https://arxiv.org/html/2609.30706#bib.bib76);[Schatzmann et al\., 2007](https://arxiv.org/html/2609.30706#bib.bib66)\)\. Selective classification\([Geifman and El\-Yaniv, 2017](https://arxiv.org/html/2609.30706#bib.bib29)\)and learning to defer\([Mozannar and Sontag, 2020](https://arxiv.org/html/2609.30706#bib.bib55)\)decide whether to answer at all; KnowNo\([Ren et al\., 2023](https://arxiv.org/html/2609.30706#bib.bib64)\)asks for help when a conformal set\([Vovk et al\., 2005](https://arxiv.org/html/2609.30706#bib.bib72);[Angelopoulos and Bates, 2021](https://arxiv.org/html/2609.30706#bib.bib6)\)has more than one option; our baseline B5 uses the same trigger with an uncalibrated prediction set\. Because LLM\-generated data carries exploitable artifacts\([Li et al\., 2023](https://arxiv.org/html/2609.30706#bib.bib47);[Gururangan et al\., 2018](https://arxiv.org/html/2609.30706#bib.bib33);[Geirhos et al\., 2020](https://arxiv.org/html/2609.30706#bib.bib30)\), we keep labels out of the LLM and check texts with a model from another family\.
## 3Problem Formulation
We follow Laya’s interface: a*state*ss\(a message or JSON object\) and a*question*of typechoice,scoreornoul\(yes/no\) with optionso1,…,omo\_\{1\},\\dots,o\_\{m\}; the model returnsp\(⋅∣s,q\)p\(\\cdot\\mid s,q\)\.
#### Schemas and the exact posterior\.
A*schema*has a set of units𝒰\\mathcal\{U\}\(the options\), decisive and non\-decisive slots, a finite value setVkV\_\{k\}for every decisive slot, a conditional priorP\(vk∣vπ\(k\)\)P\(v\_\{k\}\\mid v\_\{\\pi\(k\)\}\)with at most one parent slot, and an ordered rule listgg\(the last rule is*else*\) that maps decisive\-slot values to a unit\. A*profile*θ\\thetaassigns every slot; its gold unit isy⋆=g\(θ\)y^\{\\star\}=g\(\\theta\)\. What the model has seen about slotkkis an evidence setEk⊆VkE\_\{k\}\\subseteq V\_\{k\}:\{θk\}\\\{\\theta\_\{k\}\\\}for a stated slot or a clean answer,\{θk,a\}\\\{\\theta\_\{k\},a\\\}for a*partial*answer \(“it might beaaor maybebb”\),VkV\_\{k\}for “I don’t know”; an*over\-informative*answer also reveals one unasked slot\. With a uniform likelihood inside each evidence set,
P\(y∣E\)∝∑θP\(θ\)1\[g\(θ\)=y\]∏k𝟙\[θk∈Ek\],P\(y\\mid E\)\\propto\\textstyle\\sum\_\{\\theta\}P\(\\theta\)\\,\\mathbb\{1\}\[g\(\\theta\)=y\]\\prod\_\{k\}\\mathbb\{1\}\[\\theta\_\{k\}\\in E\_\{k\}\],\(1\)computed exactly by enumeration\. The*Bayes ceiling*of a test set is the mean ofmaxyP\(y∣E\)\\max\_\{y\}P\(y\\mid E\); no text\-only model can exceed it in expectation\.
#### Value of information\.
WithAk\(⋅∣θ\)A\_\{k\}\(\\cdot\\mid\\theta\)the simulator’s answer distribution for slotkk, the*oracle VOI*is the expected increase in the probability of the correct unit:
VOI⋆\(k∣E\)=𝔼θ∼P\(⋅∣E\)𝔼a∼Ak\(⋅∣θ\)\[P\(g\(θ\)∣E⊕a\)\]−∑yP\(y∣E\)2\.\\begin\{split\}\\mathrm\{VOI\}^\{\\star\}\(k\\mid E\)=\{\}&\\mathbb\{E\}\_\{\\theta\\sim P\(\\cdot\\mid E\)\}\\,\\mathbb\{E\}\_\{a\\sim A\_\{k\}\(\\cdot\\mid\\theta\)\}\\big\[P\(g\(\\theta\)\\mid E\\oplus a\)\\big\]\\\\ &\-\\textstyle\\sum\_\{y\}P\(y\\mid E\)^\{2\}\.\\end\{split\}\(2\)
###### Proposition 1\(Gini bound\)\.
VOI⋆\(k∣E\)≤1−∑yP\(y∣E\)2=G\(P\(⋅∣E\)\)\\mathrm\{VOI\}^\{\\star\}\(k\\mid E\)\\leq 1\-\\sum\_\{y\}P\(y\\mid E\)^\{2\}=G\(P\(\\cdot\\mid E\)\), the Gini impurity\([Breiman et al\., 1984](https://arxiv.org/html/2609.30706#bib.bib14)\)of the posterior\.
###### Proof\.
The first term of Eq\. \([2](https://arxiv.org/html/2609.30706#S3.E2)\) is an expected probability, hence at most 1\. ∎
For a calibrated modelp≈P\(⋅∣E\)p\\approx P\(\\cdot\\mid E\), the value of any question therefore vanishes as the model becomes certain; §[7\.2](https://arxiv.org/html/2609.30706#S7.SS2)shows that a learned head does not discover this on its own out of distribution\.
#### Question policy\.
Givenppand predicted valuesv^k\\hat\{v\}\_\{k\}for the slots not yet asked, LAVOIR asks aboutk⋆=argmaxkv^kk^\{\\star\}=\\arg\\max\_\{k\}\\hat\{v\}\_\{k\}ifv^k⋆\>c\\hat\{v\}\_\{k^\{\\star\}\}\>cand fewer thanQQquestions have been asked, appends the question and answer to the state and re\-encodes; otherwise it hands off if1−maxypy\>ch1\-\\max\_\{y\}p\_\{y\}\>c\_\{h\}and decidesargmaxypy\\arg\\max\_\{y\}p\_\{y\}if not \(Figure[2](https://arxiv.org/html/2609.30706#S3.F2)\)\. A slot is never asked twice\.
Messagefirst message / answer addedModelppandv^k\\hat\{v\}\_\{k\}in one passmaxkv^k\>c\\max\_\{k\}\\hat\{v\}\_\{k\}\>c?e\.g\.request\.366→\\toyesAsk slotk⋆k^\{\\star\}fixed questionyessend question→\\towait for answer→\\tore\-encode1−maxypy\>ch1\-\\max\_\{y\}p\_\{y\}\>c\_\{h\}?still unsure?Decideargmaxypy\\arg\\max\_\{y\}p\_\{y\}Hand offto a humanno, orQQquestions askednoyesLOOP⋅\\cdotRUNS INSIDE THE MODELEXITFigure 2:The asking loop of §[3](https://arxiv.org/html/2609.30706#S3.SS0.SSS0.Px3)\. Every forward pass returns the decision distributionppand a valuev^k\\hat\{v\}\_\{k\}for each slot not yet asked\. While the best slot’s value exceeds the costccand the budgetQQis not used up, its fixed question is sent and the answer is appended to the conversation, which is re\-encoded\. Otherwise the model decides, or hands the case to a human if it is still unsure\. In the example of Figure[1](https://arxiv.org/html/2609.30706#S1.F1)the loop runs twice before exiting\.
## 4Model
#### Input\.
Laya’s sequence has\[CLS\], the question type and instruction, one\[MASK\]per option, and the state\. We append an optional slot block introduced by the plain\-text header “missing information:” \(Figure[3](https://arxiv.org/html/2609.30706#S4.F3)\)\. Option and slot orders are re\-shuffled every time an example is read, and the state is always a list of turns, so the target snapshot and the deployed model see the same format\.
\[CLS\]choice: Which team should handle this?\[SEP\]\[MASK\]privacy\[MASK\]legal\[MASK\]support\[SEP\]missing information:\[MASK\]data type\[MASK\]request\[MASK\]contract\[SEP\]My information was shared…\[SEP\]option markers→\\toscorer→p\\to pslot markers→\\toVOI head→v^k\\to\\hat\{v\}\_\{k\}Figure 3:Input sequence\. The second marker block lists the slots that could be asked; option markers \(blue\) feed the option scorer and slot markers \(green\) feed the VOI head in the same forward pass\. The state \(sand\) is the conversation so far\.
#### Architecture\.
We keep Laya’s decision head on ModernBERT\-large \(a question\-type embedding, two transformer layers and an option scorer;[Convai Innovations, 2026](https://arxiv.org/html/2609.30706#bib.bib20)\) and add three components\. A*segment embedding*of the block id is initialized to zero, so with an empty slot block the logits are bit\-identical to Laya’s \(torch\.equal, real weights, fp32\)\. The*VOI head*computeshk=MLP\(\[zk;f\(p\)\]\)h\_\{k\}=\\mathrm\{MLP\}\(\[z\_\{k\};f\(p\)\]\)\(LayerNorm overd\+4d\+4inputs, 256 GELU units, scalar output; 0\.27M parameters\), wherezkz\_\{k\}is the head output at slot markerkkandf\(p\)f\(p\)holds the largest probability, the top\-two margin, the entropy divided bylogm\\log m, andm/255m/255\(the features of Laya’s act head\), computed on the detached distribution\. The*Gini cap*gives
v^k=G\(softmax\(z/T\)\)⋅σ\(hk\),G\(p\)=1−∑ypy2,\\hat\{v\}\_\{k\}=G\\big\(\\mathrm\{softmax\}\(z/T\)\\big\)\\cdot\\sigma\(h\_\{k\}\),\\qquad G\(p\)=1\-\\textstyle\\sum\_\{y\}p\_\{y\}^\{2\},\(3\)withGGdetached, so a confident model predicts a structurally small value\. Uncapped variants outputhkh\_\{k\}directly\.
#### Losses\.
Laya’s decision loss is a soft cross\-entropy plus a policy\-gradient term that scores noisy logits with a proper reward\([Gneiting and Raftery, 2007](https://arxiv.org/html/2609.30706#bib.bib31);[Epstein, 1969](https://arxiv.org/html/2609.30706#bib.bib24)\)and applies REINFORCE with a group\-mean baseline; we expose its weightwRLw\_\{\\mathrm\{RL\}\}\(11is Laya’s recipe\)\. The VOI target of slotkkis
tk=p^k\(y⋆\)−p^0\(y⋆\),t\_\{k\}=\\hat\{p\}\_\{k\}\(y^\{\\star\}\)\-\\hat\{p\}\_\{0\}\(y^\{\\star\}\),\(4\)wherep^0\\hat\{p\}\_\{0\}is a frozen, temperature\-scaled snapshot on the example andp^k\\hat\{p\}\_\{k\}the same after a sampled answer to slotkk\(a*probe*\) is appended\. We minimize the MSE betweenv^k\\hat\{v\}\_\{k\}andtkt\_\{k\}, both divided by the target standard deviation, with weightλ=1\\lambda=1\. The target uses the gold unit, not the model’s own risk: a self\-referential target such as the drop in1−maxypy1\-\\max\_\{y\}p\_\{y\}rewards any answer that makes the model confident, including a misleading one\. It is also bounded in\[−1,1\]\[\-1,1\], unlike a log\-loss difference\. Because each message has several profiles, the regression estimates the snapshot’s version of Eq\. \([2](https://arxiv.org/html/2609.30706#S3.E2)\)\.
#### Training\.
A*warm*phase trains the decision model on all examples \(500 frozen\-encoder steps, then four epochs; learning rates2\.5×10−52\.5\\times 10^\{\-5\}encoder and10−410^\{\-4\}head, cosine, bf16, global batch 64\)\. Temperatures are fitted with L\-BFGS\([Liu and Nocedal, 1989](https://arxiv.org/html/2609.30706#bib.bib50)\)on a 5% held\-out slice\. The warm snapshot computes targets, two*joint*epochs trainℒdec\+λℒVOI\\mathcal\{L\}\_\{\\mathrm\{dec\}\}\+\\lambda\\mathcal\{L\}\_\{\\mathrm\{VOI\}\}with encoder learning rate5×10−65\\times 10^\{\-6\}, and targets are recomputed from that snapshot for one more joint epoch, so the head regresses on values of a decision model that has stopped moving\.
## 5Data
#### Schemas\.
We wrote twelve customer\-service schemas, eight for training and four held out for zero\-shot evaluation, each with five or six units, three or four decisive slots and one non\-decisive slot \(Appendix[A](https://arxiv.org/html/2609.30706#A1)\)\. Before generating any text, automatic gates checked reachability, unit masses, the share of units that need two or more slots, and the mean maximum posterior and VOI share of a dry run\. Three pilots and one discarded full run showed that data quality is governed by the*semantic independence*of slots, not by rule complexity: overlapping slots \(“an unrecognized transaction” and “did you authorize it”\) leak each other; a*not applicable*value is indistinguishable from “no” in text; default\-like values \(“technical help”\) are leaked by any vague message; non\-decisive dates or amounts imply decisive facts, so the final non\-decisive slot is the customer’s name; and the generator must be forbidden to invent specifics when the topic is hidden\.
#### Cases and answers\.
A*case*is a partial profile \(the slots stated in the first message\) with two to four complete profiles consistent with it, which differ in their hidden values and, typically, in their gold unit; without this pairing the regression would learn a deterministic mapping rather than an expectation\. Each profile yields a message\-only example, with probability 0\.7 one with a question–answer pair, and with probability 0\.4 one with two, plus a probe for every slot\. Answers are clean \(0\.675\), partial \(0\.125\), “I don’t know” \(0\.10\) or over\-informative \(0\.10\); probes of already known slots are always clean\.
#### Generation and two\-way checking\.
Messages \(ten styles, at most 120 words\) and answers are written by Qwen3\.6\-35B\-A3B\([Yang et al\., 2025](https://arxiv.org/html/2609.30706#bib.bib78);[Qwen Team, 2026](https://arxiv.org/html/2609.30706#bib.bib61)\)with vLLM\([Kwon et al\., 2023](https://arxiv.org/html/2609.30706#bib.bib44)\)\. The prompt lists the facts to state plainly and the topics not to mention, without their values\. Gemma\-4\-26B\-A4B\([Gemma Team, 2026](https://arxiv.org/html/2609.30706#bib.bib28)\)returns, for every decisive slot, a constrained guess and an evidence level; a text is regenerated \(up to three times\) if a hidden slot is guessed with explicit evidence \(*leakage*\) or a stated slot is not recovered \(*fidelity*\)\. The final run \(two GH200 nodes, 12 min\) produced 21,601 training examples \(8×\\times500 cases\), 5,379 seen\-schema test examples and 5,432 zero\-shot test examples\. First\-attempt message rejection was 1–10% for seven training schemas and 22% for privacy, whose “type of data” and “request” slots overlap irreducibly; answer rejection was at most 2\.8%, and partial answers were guessed correctly 46–62% of the time \(expected 50%\)\.
#### General single\-turn data\.
To keep general ability we add single\-turn data in Laya’s format \(details in Appendix[B](https://arxiv.org/html/2609.30706#A2)\): typed\-decisions, CLINC150, MultiNLI, Yelp and SGD first turns\([Convai Innovations, 2026](https://arxiv.org/html/2609.30706#bib.bib20);[Larson et al\., 2019](https://arxiv.org/html/2609.30706#bib.bib46);[Williams et al\., 2018](https://arxiv.org/html/2609.30706#bib.bib75);[Zhang et al\., 2015](https://arxiv.org/html/2609.30706#bib.bib83);[Rastogi et al\., 2020](https://arxiv.org/html/2609.30706#bib.bib63)\)\(19,000\); tool choice from Nemotron post\-training data\([NVIDIA, 2025](https://arxiv.org/html/2609.30706#bib.bib58)\)\(6,000, of which 5,447 train\); the six sources Laya trained on with question\-form yes/no items\([Zhang et al\., 2015](https://arxiv.org/html/2609.30706#bib.bib83);[Clark et al\., 2019](https://arxiv.org/html/2609.30706#bib.bib17);[Metsis et al\., 2006](https://arxiv.org/html/2609.30706#bib.bib54);[Liu, 2024](https://arxiv.org/html/2609.30706#bib.bib51);[Bajaj et al\., 2016](https://arxiv.org/html/2609.30706#bib.bib10);[Tobi\-Bueck, 2025](https://arxiv.org/html/2609.30706#bib.bib71)\)\(19,437\), with all 4,447 benchmark texts removed; and 27 open sources \(38,200, train splits only\)\.*Controlled*models use the schema data, the first 19,000 and tool choice \(46,048 examples\);*broad*models use everything \(103,685\)\.
#### Real conversations\.
For out\-of\-distribution tests we use 918 ABCD test dialogues\([Chen et al\., 2021](https://arxiv.org/html/2609.30706#bib.bib16)\)with ten flows as units, the first customer block as state and the agent’s first reply plus the customer’s second block as the real exchange, and 3,812 SGD first user turns\([Rastogi et al\., 2020](https://arxiv.org/html/2609.30706#bib.bib63)\)from 20 services \(2,766 from services unseen in training\) with the service’s arguments as slots, which cannot change the intent\.
## 6Experimental Setup
Table 1:Controlled study on seen and zero\-shot \(zs\) schemas\.*Gap*: accuracy minus Bayes ceiling \(95% CI\)\.ρ\\rho, top\-1: predicted vs\. oracle VOI\. AUC: area under the accuracy–questions curve \(0–2 questions\)\. b0\.5: best accuracy with≤\\leq0\.5 questions per conversation\. Upper block: group temperatures; lower: one temperature per type \(Laya\)\.Single turnVOI vs\. oracleAUCb0\.5TArmSplitAccCeilingGap \[95% CI\]ECEKLρ\\rhotop\-1B2B3B5VOIORACLEB2B3VOIgroupRL1seen\.653\.659−\-\.006 \[−\-\.016, \.004\]\.015\.016\.821\.876\.612\.766\.791\.800\.803\.612\.669\.741groupRL0seen\.652\.659−\-\.006 \[−\-\.017, \.004\]\.013\.012\.831\.923\.611\.763\.791\.799\.800\.611\.660\.755groupRL1zs\.517\.667−\-\.150 \[−\-\.162,−\-\.137\]\.130\.993\.690\.558\.502\.590\.607\.629\.658\.502\.566\.612groupRL0zs\.512\.667−\-\.154 \[−\-\.166,−\-\.142\]\.099\.747\.730\.745\.492\.574\.595\.616\.633\.492\.542\.579singleRL1seen\.652\.659−\-\.007\.078\.063\.817\.873\.614\.753\.772\.803\.802\.614\.614\.733singleRL0seen\.648\.659−\-\.010\.064\.043\.773\.934\.608\.751\.772\.797\.798\.608\.608\.762singleRL1zs\.514\.667−\-\.152\.071\.492\.682\.554\.499\.581\.580\.622\.653\.499\.499\.609singleRL0zs\.522\.667−\-\.145\.070\.481\.702\.649\.498\.578\.576\.624\.647\.498\.498\.549
#### Models\.
All models use ModernBERT\-large and a single seed:*controlled*models withwRL∈\{0,1\}w\_\{\\mathrm\{RL\}\}\\in\\\{0,1\\\}\(RL0/RL1\) and either one temperature per question type \(Laya’s procedure\) or per source group; the*broad*model \(wRL=0w\_\{\\mathrm\{RL\}\}=0, per\-source temperatures, uncapped\); and the final LAVOIR, the broad warm model with the joint phases re\-run with the Gini cap\. Training took 28–48 min on 8–16 GH200 GPUs of the JUPITER booster\. Laya and Jev numbers are those reported by Laya \(Jev’s are third\-party measurements\); apart from Figure[1](https://arxiv.org/html/2609.30706#S1.F1)we did not run Laya’s checkpoints\.
#### Policies\.
All policies share the decision head and allow at most two questions:B2never asks;B3asks about a random unresolved slot whenmaxypy<τ\\max\_\{y\}p\_\{y\}<\\tau;B5asks about the highest\-VOI slot when the prediction set, the smallest set of units whose probabilities sum to at least1−α1\-\\alpha, contains more than one unit, withα∈\{0\.5,0\.3,0\.2,0\.1,0\.05,0\.02,0\.01\}\\alpha\\in\\\{0\.5,0\.3,0\.2,0\.1,0\.05,0\.02,0\.01\\\}swept like the other thresholds \(the set is not calibrated on held\-out data, so it carries no coverage guarantee\);VOIis our policy;ORACLEasks about the slot with the largest exactVOI⋆\\mathrm\{VOI\}^\{\\star\}\. ORACLE is greedy \(it looks one question ahead\) and, like every policy, decides with the model’s own decision head, so it is a strong reference rather than an upper bound, and its value differs between models\. For every message\-only test example we run one dialogue per profile, answering with the profile’s probe, and sweepcc,τ\\tauandchc\_\{h\}\.
#### Metrics\.
Accuracy, Bayes ceiling and their gap with a bootstrap 95% interval\([Efron and Tibshirani, 1993](https://arxiv.org/html/2609.30706#bib.bib23)\), ECE\([Naeini et al\., 2015](https://arxiv.org/html/2609.30706#bib.bib56)\)and KL to the posterior; Spearmanρ\\rhoand top\-1 agreement betweenv^\\hat\{v\}andVOI⋆\\mathrm\{VOI\}^\{\\star\};AUC, the normalized area under the upper envelope of accuracy vs\. mean questions \(0–2\);b0\.5, the best accuracy with at most 0\.5 questions per conversation; and the share of*unnecessary*questions, those about a non\-decisive slot or a slot whose value is already determined by the evidence, among all questions asked\. AUC differences are given on the 0–1 scale, accuracy differences in percentage points\.
## 7Results
### 7\.1Controlled study
#### Decisions on seen schemas are Bayes\-optimal\.
Accuracy is 0\.006 below the Bayes ceiling with an interval that contains zero \(Table[1](https://arxiv.org/html/2609.30706#S6.T1)\); a model exploiting leaked hints would exceed the ceiling, and one under\-using the text would fall clearly below it\. Two schemas looked suspicious in argmax accuracy \(insurance 0\.649 vs\. a ceiling of 0\.610\), but comparing the probability of the gold unit with the posterior under a case\-clustered bootstrap shows no excess on any schema \(Appendix[C](https://arxiv.org/html/2609.30706#A3)\); the argmax excess comes from nearly tied posteriors\. Per\-schema leakage tests should therefore use probabilities, not argmax accuracy\.
#### The question policy matches the oracle\.
With group temperatures the VOI policy reaches an AUC of 0\.799–0\.800 against 0\.800–0\.803 for ORACLE, withρ=0\.82\\rho=0\.82–0\.830\.83and 88–92% top\-1 agreement\. With at most 0\.5 questions per conversation it is 12\.9–14\.4 points more accurate than never asking \(B2\) and 7\.2–9\.5 points more accurate than asking at random when uncertain \(B3\)\. B5 uses the same VOI ranking but its AUC is 0\.008–0\.009 lower: deciding*whether*to ask from the VOI itself beats deciding it from the size of a prediction set\. Atc=0\.02c=0\.02RL0 asks 1\.09 questions per conversation for an accuracy of 0\.889 with 1% unnecessary questions \(oracle: 1\.07 for 0\.890\); addingch=0\.3c\_\{h\}=0\.3hands off 18% of conversations and is 0\.973 accurate on the rest\.
#### Zero\-shot schemas\.
On the four held\-out schemas the VOI policy still beats every baseline \(AUC 0\.02–0\.05 above B5; with group temperatures, b0\.5 3\.7–4\.6 points above B3\), but decisions are weak: 0\.51–0\.52 accuracy against a ceiling of 0\.667\. Trained on eight schemas, the model does not learn to read unseen rules from option descriptions\.
### 7\.2Real conversations and the Gini cap
Table 2:Broad model without and with the Gini cap \(final LAVOIR\)\. Question rates atc=\.02/\.05/\.1c=\.02/\.05/\.1\. ABCD gains: accuracy change in points after the real first exchange in the cases where the policy \(atc=\.02c=\.02\) would and would not ask, with 95% paired bootstrap intervals over cases \(10,000 resamples; 529/389 cases for LAVOIR, 844/74 uncapped\)\. ORACLE is greedy and uses each model’s decision head, so it can lie below VOI\.UncappedLAVOIRSGD question rate\.93 / \.87 / \.69\.086 / \.063 / \.048SGD AUROC\(VOI→\\toerror\)\.70\.84SGD accuracy / ECE\.948 / \.041\.942 / \.044ABCD question rate\.92 / \.83 / \.71\.58 / \.49 / \.41ABCD AUROC\(VOI→\\toerror\)\.65\.71ABCD accuracy / ECE\.653 / \.222\.657 / \.186ABCD gain, asks\+2\.7 \[0\.5, 5\.0\]\+8\.3 \[4\.9, 11\.7\]ABCD gain, does not ask\+2\.7 \[0\.0, 6\.8\]−\-0\.3 \[−\-1\.5, 1\.0\]Seen: AUC VOI / ORACLE\.795 / \.797\.799/ \.797Seen: b0\.5\.757\.751Seen:ρ\\rho/ top\-1\.763 / \.960\.851/ \.937Zero\-shot:ρ\\rho/ b0\.5\.662 / \.535\.689 / \.581
#### Without the cap the model asks too much\.
The broad model is 0\.948 accurate on SGD and well calibrated \(ECE 0\.04\), yet atc=0\.02c=0\.02it would ask in 93% of SGD and 92% of ABCD conversations\. Among the 3,526 SGD cases with confidence≥0\.99\\geq 0\.99the median maximum VOI is 0\.133, although by Proposition[1](https://arxiv.org/html/2609.30706#Thmproposition1)the expected gain there is at most about 0\.02\. The head receives the decision distribution but has not learned the bound: out of distribution the slot representation dominates, and arguments that cannot change the intent get high value \(amount 0\.25, ride fare 0\.18\)\.
#### The cap fixes it structurally\.
Re\-running only the joint phases with Eq\. \([3](https://arxiv.org/html/2609.30706#S4.E3)\) leaves single\-turn decisions essentially unchanged \(calibration\-slice accuracy 0\.849→\\to0\.847, ABCD first\-message accuracy 0\.653→\\to0\.657\) and drops the question rate on confident cases to zero \(Table[2](https://arxiv.org/html/2609.30706#S7.T2)\)\. On SGD atc=0\.05c=0\.05the model asks in 40% of its errors and 4% of its correct decisions\. On ABCD the uncapped model’s questions were untargeted: the real exchange helped equally where it would and would not ask \(\+2\.7 each\)\. With the cap the whole gain falls where it asks \(\+8\.3, 95% CI \[4\.9, 11\.7\]\) and none where it does not \(−\-0\.3 \[−\-1\.5, 1\.0\]\); the difference between the groups is 8\.6 \[4\.9, 12\.2\] points, against 0\.0 \[−\-4\.7, 3\.9\] for the uncapped model, so the decision to ask now tracks the value of the answer\. Over all conversations the capped model gains about 4\.7 points from the exchange against 2\.7 for the uncapped one: re\-running the joint phases changed how the decision head reads a second turn, which our single\-turn checks do not measure, so we compare the two groups within each model rather than the totals across models\. On seen schemas the cap raises VOI quality \(ρ\\rho0\.76→\\to0\.85\) and keeps the AUC at ORACLE level; ORACLE is slightly below VOI here \(0\.797 vs\. 0\.799\) because it is greedy and shares the decision head\. With at most 0\.5 questions the final model reaches 0\.751, against 0\.610 for never asking and 0\.663 for random questions\. Zero\-shot b0\.5 improves by 4\.6 points\. What remains is calibration: on ABCD the model is wrong in 19% of the cases where its confidence is at least 0\.99, where the cap correctly stays silent but the confidence is wrong\.
### 7\.3External benchmarks
Table 3:Accuracy on Laya’s benchmarks \(seed 13, 400 cases per task\)\.*Ctrl*: controlled RL0\. Bold: LAVOIR above Laya\. Laya: best released checkpoint; Laya and Jev as reported by Laya \(Jev’s Banking77: 72 labels,n=100n=100\)\.∗Training split in our broad data and in Laya’s;†related train splits only\. Latencies come from different hardware \(ours GH200, Laya T4, Jev through its API\) and are not directly comparable\.SetCtrlLAVOIRLayaJevtyped\-decisions∗\.785\.774\.766\.727Banking77 \(77\)\.575\.533\.492\.870MASSIVE\-en \(20 opt\.\)\.790\.805\.783–jailbreak \(ToxicChat\)\.888\.825\.762–model routing†\.283\.754\.659–toxicity \(ToxicChat\)\.525\.605\.530–RAG relevance∗\.503\.665\.657–spam∗\.480\.993\.993–phishing∗\.625\.978\.993–AG News∗\.765\.905\.953\.910DAIR Emotion\.528\.575\.600\.480support triage∗\.285\.383\.522–latency, p5027 ms31 ms33–40 ms236 ms
We ran Laya’s benchmark builder and metrics unchanged on AG News\([Zhang et al\., 2015](https://arxiv.org/html/2609.30706#bib.bib83)\), DAIR Emotion\([Saravia et al\., 2018](https://arxiv.org/html/2609.30706#bib.bib65)\), Banking77\([Casanueva et al\., 2020](https://arxiv.org/html/2609.30706#bib.bib15)\), support tickets\([Tobi\-Bueck, 2025](https://arxiv.org/html/2609.30706#bib.bib71)\), Enron spam\([Metsis et al\., 2006](https://arxiv.org/html/2609.30706#bib.bib54)\), phishing\([Liu, 2024](https://arxiv.org/html/2609.30706#bib.bib51)\), ToxicChat\([Lin et al\., 2023](https://arxiv.org/html/2609.30706#bib.bib48)\), MS MARCO\([Bajaj et al\., 2016](https://arxiv.org/html/2609.30706#bib.bib10)\), a model\-routing task over GSM8K and MBPP items\([Cobbe et al\., 2021](https://arxiv.org/html/2609.30706#bib.bib18);[Austin et al\., 2021](https://arxiv.org/html/2609.30706#bib.bib8)\), MASSIVE\-en\([FitzGerald et al\., 2023](https://arxiv.org/html/2609.30706#bib.bib26)\)and the typed\-decisions test split\. LAVOIR is above Laya on seven of twelve sets, tied on spam and below on four \(Table[3](https://arxiv.org/html/2609.30706#S7.T3)\); Banking77, jailbreak and MASSIVE were never in our training data, and on typed\-decisions it is also above Jev\. Latencies were measured on different GPUs \(ours GH200, Laya T4\)\.
#### Yes/no collapse\.
The controlled model beat Laya on typed\-decisions, Banking77, MASSIVE and jailbreak but collapsed on every yes/no benchmark \(“false” on 398/400 spam cases, “true” on 399/400 RAG cases\)\. All its yes/no training items were*assertions*\(“This trace requires human review\.”\), while benchmark instructions are*questions*over dictionaries with named fields\. Question\-form yes/no items and the open sources removed the collapse \(spam 0\.480→\\to0\.993\) and lifted model routing from 0\.283 to 0\.754, at a small cost in VOI rank correlation on seen schemas\.
### 7\.4Lessons about Laya’s recipe
#### The policy\-gradient term does not help\.
In a miniature model, after the decision has converged the policy\-gradient term keeps sending gradients to the shared parameters that are one to two orders of magnitude larger than the VOI gradient; with a fixed exploration scale its variance does not vanish at the optimum\. WithwRL=0w\_\{\\mathrm\{RL\}\}=0the VOI loss reached10−410^\{\-4\}in 200 steps, withwRL=1w\_\{\\mathrm\{RL\}\}=1it was still 0\.06 after 1,000\. At full scale the RL0 arm is ahead in every warm epoch, reaches the same decision accuracy \(0\.918 vs\. 0\.919\), and learns the VOI head better \(top\-1 0\.92 vs\. 0\.88 on seen and 0\.75 vs\. 0\.56 on zero\-shot schemas; Appendix[C](https://arxiv.org/html/2609.30706#A3)\)\. Because Laya’s reward is a differentiable function of the reported distribution, its gradient can be taken directly; policy gradients would be needed only if the reward depended on which questions were asked\.
#### One temperature per type is not enough\.
Laya fits one temperature per question type\. On our mixture this gaveT=2\.25T=2\.25–2\.752\.75for choice because the general data has one\-hot targets, while our schemas, whose targets are soft posteriors, were already calibrated atT=1T=1; the shared temperature made their KL ten times worse \(0\.004→\\to0\.04–0\.06\)\. Per\-group temperatures \(T≈1\.02T\\approx 1\.02for our schemas\) improve ECE on seen schemas five\-fold, though not on zero\-shot schemas \(0\.070–0\.071→\\to0\.099–0\.130; Table[1](https://arxiv.org/html/2609.30706#S6.T1)\), yet the dialogue AUC barely moves: the VOI*ranking*is robust to temperature, and calibration matters mainly for hand\-off and, through the cap, for staying silent\.
## 8Discussion and Conclusion
LAVOIR turns a single\-pass decision encoder into one that knows when a question is worth asking and which one to ask, with the value of every candidate question predicted in the same forward pass as the decision\. Rule\-defined gold, several profiles per message, two\-way cross\-family checking and the Gini cap made this value learnable and trustworthy: on seen schemas the decisions are Bayes\-optimal and the questions oracle\-level, and on real conversations the questions go where answers help\. The cap matters beyond our setting: an uncapped head can learn to output zero when the model is certain, and in distribution it does, but out of distribution it values fields that merely look important\. Next steps are a Turkish model on MoganBERT\-TR\([Yılmaz et al\., 2026a](https://arxiv.org/html/2609.30706#bib.bib79)\), more and more varied schemas to close the zero\-shot gap, the ABCD training split for calibration on real dialogues, and a study with human participants\.
## Limitations
Synthetic dialogues:answers in the controlled study come from a simulator whose kinds and rates we chose, and LLM\-written text may be more cooperative than real users; the real\-data experiments test single exchanges, not the full loop with people\.Zero\-shot decisionsstay 15 points below the ceiling, andcalibration out of distributionis weak \(ABCD ECE 0\.19\)\.Residual overlapin the privacy schema caused 22% message rejection, a mild selection effect\.Missing baselines:a slow VOI that re\-encodes every possible answer, an LLM agent, and an ablation of the self\-referential target\. Questions are slottemplates; all runs useone seedandEnglishonly; Laya’s numbers are taken from its repository, and the two models are trained on different mixtures\.
## Ethics Statement
LAVOIR is a routing component: it chooses a unit, asks at most a few questions, or hands the case to a human\. In sensitive domains \(privacy, clinic routing, insurance\) the hand\-off threshold should be conservative and questions should be reviewed so that they do not request unnecessary personal information\. All synthetic profiles are fictitious; the generator and checker \(Qwen3\.6\-35B\-A3B, Gemma\-4\-26B\-A4B\) are Apache\-2\.0 models that place no restriction on generated text\. Most general sources are under permissive or share\-alike licenses \(Apache\-2\.0, MIT, CC0, CC BY, CC BY\-SA\), but several restrict use to non\-commercial research: ANLI and the support tickets \(CC BY\-NC 4\.0\), MS MARCO, the Yelp reviews, MARC and QQP \(their providers’ terms\); the phishing corpus is LGPL\-3\.0, and sources without a license on the Hugging Face Hub were used under their original distribution terms\. We therefore release the weights under CC BY\-NC 4\.0 and do not redistribute any third\-party data\.
## Code and Data Availability
The code is available at[https://github\.com/moganai/lavoir](https://github.com/moganai/lavoir)under Apache\-2\.0: the model, sequence builder, losses, VOI target computation, question policy, evaluation, training and fine\-tuning scripts, tests, and example workflow definitions with the data format\. The final model with its fitted temperatures is available at[https://huggingface\.co/moganai/lavoir](https://huggingface.co/moganai/lavoir)under CC BY\-NC 4\.0\. The general single\-turn mixture is not redistributed; it is built from the public sources listed in Appendix[B](https://arxiv.org/html/2609.30706#A2)\.
## Acknowledgments
We acknowledge the EuroHPC Joint Undertaking for awarding this project access to the EuroHPC supercomputer JUPITER, hosted by the Jülich Supercomputing Centre \(JSC\), through the EuroHPC AI Factories Playground access call \(project EHPC\-AIF\-2026PG01\-1296\)\. We thank Convai Innovations for releasing Laya’s code, data and benchmark harness under an open license\.
## References
- AbdelStark \(2026\)AbdelStark\. 2026\.jev\-benchmarks: Independent benchmarks of TypeSafe Jev\.GitHub repository,[https://github\.com/AbdelStark/jev\-benchmarks](https://github.com/AbdelStark/jev-benchmarks)\. Accessed September 2026\.
- Almeida \(2026\)Diogo Almeida\. 2026\.Introducing System One models and Jev\.TypeSafe AI blog, 15 September 2026,[https://typesafe\.ai/blog/introducing\-system\-one\-models\-and\-jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev); documentation:[https://docs\.typesafe\.ai/concepts/system\-one](https://docs.typesafe.ai/concepts/system-one)\.
- Aliannejadi et al\. \(2021\)Mohammad Aliannejadi, Julia Kiseleva, Aleksandr Chuklin, Jeff Dalton, and Mikhail Burtsev\. 2021\.Building and evaluating open\-domain dialogue corpora with clarifying questions\.In*Proceedings of EMNLP 2021*\.
- Aliannejadi et al\. \(2019\)Mohammad Aliannejadi, Hamed Zamani, Fabio Crestani, and W\. Bruce Croft\. 2019\.Asking clarifying questions in open\-domain information\-seeking conversations\.In*Proceedings of SIGIR 2019*, pages 475–484\.
- Andukuri et al\. \(2024\)Chinmaya Andukuri, Jan\-Philipp Fränken, Tobias Gerstenberg, and Noah D\. Goodman\. 2024\.STaR\-GATE: Teaching language models to ask clarifying questions\.In*Conference on Language Modeling \(COLM\)*\. arXiv:2403\.19154\.
- Angelopoulos and Bates \(2021\)Anastasios N\. Angelopoulos and Stephen Bates\. 2021\.A gentle introduction to conformal prediction and distribution\-free uncertainty quantification\.arXiv:2107\.07511\.
- Antypas et al\. \(2022\)Dimosthenis Antypas, Asahi Ushio, Jose Camacho\-Collados, Vitor Silva, Leonardo Neves, and Francesco Barbieri\. 2022\.Twitter topic classification\.In*Proceedings of COLING 2022*\.
- Austin et al\. \(2021\)Jacob Austin, Augustus Odena, Maxwell Nye, Maarten Bosma, Henryk Michalewski, David Dohan, Ellen Jiang, Carrie Cai, Michael Terry, Quoc Le, and Charles Sutton\. 2021\.Program synthesis with large language models\.arXiv:2108\.07732\.
- b\-mc2 \(2023\)b\-mc2\. 2023\.sql\-create\-context\.Hugging Face datasetb\-mc2/sql\-create\-context\(CC BY 4\.0\), built from WikiSQL and Spider\.
- Bajaj et al\. \(2016\)Payal Bajaj, Daniel Campos, Nick Craswell, Li Deng, Jianfeng Gao, Xiaodong Liu, Rangan Majumder, Andrew McNamara, Bhaskar Mitra, Tri Nguyen, et al\. 2016\.MS MARCO: A human generated machine reading comprehension dataset\.arXiv:1611\.09268\.
- Barbieri et al\. \(2020\)Francesco Barbieri, Jose Camacho\-Collados, Luis Espinosa Anke, and Leonardo Neves\. 2020\.TweetEval: Unified benchmark and comparative evaluation for tweet classification\.In*Findings of EMNLP 2020*\.
- Borkan et al\. \(2019\)Daniel Borkan, Lucas Dixon, Jeffrey Sorensen, Nithum Thain, and Lucy Vasserman\. 2019\.Nuanced metrics for measuring unintended bias with real data for text classification\.In*Companion Proceedings of The Web Conference \(WWW\) 2019*, pages 491–500\.
- Bowman et al\. \(2015\)Samuel R\. Bowman, Gabor Angeli, Christopher Potts, and Christopher D\. Manning\. 2015\.A large annotated corpus for learning natural language inference\.In*Proceedings of EMNLP 2015*, pages 632–642\.
- Breiman et al\. \(1984\)Leo Breiman, Jerome H\. Friedman, Richard A\. Olshen, and Charles J\. Stone\. 1984\.*Classification and Regression Trees*\.Wadsworth\.
- Casanueva et al\. \(2020\)Iñigo Casanueva, Tadas Temčinas, Daniela Gerz, Matthew Henderson, and Ivan Vulić\. 2020\.Efficient intent detection with dual sentence encoders\.In*Proceedings of the 2nd Workshop on NLP for Conversational AI*, pages 38–45\.
- Chen et al\. \(2021\)Derek Chen, Howard Chen, Yi Yang, Alexander Lin, and Zhou Yu\. 2021\.Action\-based conversations dataset: A corpus for building more in\-depth task\-oriented dialogue systems\.In*Proceedings of NAACL\-HLT 2021*\.
- Clark et al\. \(2019\)Christopher Clark, Kenton Lee, Ming\-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova\. 2019\.BoolQ: Exploring the surprising difficulty of natural yes/no questions\.In*Proceedings of NAACL\-HLT 2019*, pages 2924–2936\.
- Cobbe et al\. \(2021\)Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, Christopher Hesse, and John Schulman\. 2021\.Training verifiers to solve math word problems\.arXiv:2110\.14168\.
- Conover et al\. \(2023\)Mike Conover, Matt Hayes, Ankit Mathur, Jianwei Xie, Jun Wan, Sam Shah, Ali Ghodsi, Patrick Wendell, Matei Zaharia, and Reynold Xin\. 2023\.Free Dolly: Introducing the world’s first truly open instruction\-tuned LLM\.Databricks blog,[https://www\.databricks\.com/blog/2023/04/12/dolly\-first\-open\-commercially\-viable\-instruction\-tuned\-llm](https://www.databricks.com/blog/2023/04/12/dolly-first-open-commercially-viable-instruction-tuned-llm)\.
- Convai Innovations \(2026\)Convai Innovations\. 2026\.Laya: Multilingual non\-autoregressive System 1 decision model\.Software and model card, version 0\.3\.11 \(commit1e28ac2\),[https://huggingface\.co/convaiinnovations/laya](https://huggingface.co/convaiinnovations/laya); code:[https://github\.com/NandhaKishorM/laya](https://github.com/NandhaKishorM/laya)\. Accessed September 2026\.
- Demszky et al\. \(2020\)Dorottya Demszky, Dana Movshovitz\-Attias, Jeongwoo Ko, Alan Cowen, Gaurav Nemade, and Sujith Ravi\. 2020\.GoEmotions: A dataset of fine\-grained emotions\.In*Proceedings of ACL 2020*, pages 4040–4054\.
- Dolan and Brockett \(2005\)William B\. Dolan and Chris Brockett\. 2005\.Automatically constructing a corpus of sentential paraphrases\.In*Proceedings of the Third International Workshop on Paraphrasing \(IWP 2005\)*\.
- Efron and Tibshirani \(1993\)Bradley Efron and Robert J\. Tibshirani\. 1993\.*An Introduction to the Bootstrap*\.Chapman & Hall\.
- Epstein \(1969\)Edward S\. Epstein\. 1969\.A scoring system for probability forecasts of ranked categories\.*Journal of Applied Meteorology*, 8\(6\):985–987\.
- Fan et al\. \(2018\)Angela Fan, Mike Lewis, and Yann Dauphin\. 2018\.Hierarchical neural story generation\.In*Proceedings of ACL 2018*, pages 889–898\.
- FitzGerald et al\. \(2023\)Jack FitzGerald, Christopher Hench, Charith Peris, Scott Mackie, Kay Rottmann, Ana Sanchez, Aaron Nash, Liam Urbach, Vishesh Kakarala, Richa Singh, et al\. 2023\.MASSIVE: A 1M\-example multilingual natural language understanding dataset with 51 typologically\-diverse languages\.In*Proceedings of ACL 2023*, pages 4277–4302\.
- Foster et al\. \(2021\)Adam Foster, Desi R\. Ivanova, Ilyas Malik, and Tom Rainforth\. 2021\.Deep adaptive design: Amortizing sequential Bayesian experimental design\.In*Proceedings of ICML 2021*, pages 3384–3395\.
- Gemma Team \(2026\)Gemma Team, Google DeepMind\. 2026\.Gemma 4 technical report\.arXiv:2607\.02770\. Model used:google/gemma\-4\-26B\-A4B\-it\.
- Geifman and El\-Yaniv \(2017\)Yonatan Geifman and Ran El\-Yaniv\. 2017\.Selective classification for deep neural networks\.In*Advances in Neural Information Processing Systems 30*, pages 4878–4887\.
- Geirhos et al\. \(2020\)Robert Geirhos, Jörn\-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A\. Wichmann\. 2020\.Shortcut learning in deep neural networks\.*Nature Machine Intelligence*, 2:665–673\.
- Gneiting and Raftery \(2007\)Tilmann Gneiting and Adrian E\. Raftery\. 2007\.Strictly proper scoring rules, prediction, and estimation\.*Journal of the American Statistical Association*, 102\(477\):359–378\.
- Guo et al\. \(2017\)Chuan Guo, Geoff Pleiss, Yu Sun, and Kilian Q\. Weinberger\. 2017\.On calibration of modern neural networks\.In*Proceedings of ICML 2017*, pages 1321–1330\.
- Gururangan et al\. \(2018\)Suchin Gururangan, Swabha Swayamdipta, Omer Levy, Roy Schwartz, Samuel Bowman, and Noah A\. Smith\. 2018\.Annotation artifacts in natural language inference data\.In*Proceedings of NAACL\-HLT 2018*, pages 107–112\.
- Hendrycks et al\. \(2021\)Dan Hendrycks, Collin Burns, Saurav Kadavath, Akul Arora, Steven Basart, Eric Tang, Dawn Song, and Jacob Steinhardt\. 2021\.Measuring mathematical problem solving with the MATH dataset\.In*NeurIPS 2021 Datasets and Benchmarks Track*\.
- Howard \(1966\)Ronald A\. Howard\. 1966\.Information value theory\.*IEEE Transactions on Systems Science and Cybernetics*, 2\(1\):22–26\.
- Hu et al\. \(2024\)Zhiyuan Hu, Chumin Liu, Xidong Feng, Yilun Zhao, See\-Kiong Ng, Anh Tuan Luu, Junxian He, Pang Wei Koh, and Bryan Hooi\. 2024\.Uncertainty of thoughts: Uncertainty\-aware planning enhances information seeking in large language models\.In*Advances in Neural Information Processing Systems 37*\. arXiv:2402\.03271\.
- Iyer et al\. \(2017\)Shankar Iyer, Nikhil Dandekar, and Kornél Csernai\. 2017\.First Quora dataset release: Question pairs\.Quora blog\.
- Kahneman \(2011\)Daniel Kahneman\. 2011\.*Thinking, Fast and Slow*\.Farrar, Straus and Giroux\.
- Kennedy et al\. \(2020\)Chris J\. Kennedy, Geoff Bacon, Alexander Sahn, and Claudia von Vacano\. 2020\.Constructing interval variables via faceted Rasch measurement and multitask deep learning: A hate speech application\.arXiv:2009\.10277\. Dataset:ucberkeley\-dlab/measuring\-hate\-speech\(CC BY 4\.0\)\.
- Keung et al\. \(2020\)Phillip Keung, Yichao Lu, György Szarvas, and Noah A\. Smith\. 2020\.The multilingual Amazon reviews corpus\.In*Proceedings of EMNLP 2020*, pages 4563–4568\. English part viaSetFit/amazon\_reviews\_multi\_en; the original MARC license limits use to research\.
- Khot et al\. \(2018\)Tushar Khot, Ashish Sabharwal, and Peter Clark\. 2018\.SciTail: A textual entailment dataset from science question answering\.In*Proceedings of AAAI 2018*\.
- Kuhn et al\. \(2022\)Lorenz Kuhn, Yarin Gal, and Sebastian Farquhar\. 2022\.CLAM: Selective clarification for ambiguous questions with generative language models\.arXiv:2212\.07769\.
- Kwiatkowski et al\. \(2019\)Tom Kwiatkowski, Jennimaria Palomaki, Olivia Redfield, Michael Collins, Ankur Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, et al\. 2019\.Natural questions: A benchmark for question answering research\.*Transactions of the Association for Computational Linguistics*, 7:452–466\.
- Kwon et al\. \(2023\)Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph E\. Gonzalez, Hao Zhang, and Ion Stoica\. 2023\.Efficient memory management for large language model serving with PagedAttention\.In*Proceedings of SOSP 2023*, pages 611–626\.
- Lang \(1995\)Ken Lang\. 1995\.NewsWeeder: Learning to filter netnews\.In*Proceedings of ICML 1995*, pages 331–339\.
- Larson et al\. \(2019\)Stefan Larson, Anish Mahendran, Joseph J\. Peper, Christopher Clarke, Andrew Lee, Parker Hill, Jonathan K\. Kummerfeld, Kevin Leach, Michael A\. Laurenzano, Lingjia Tang, and Jason Mars\. 2019\.An evaluation dataset for intent classification and out\-of\-scope prediction\.In*Proceedings of EMNLP\-IJCNLP 2019*, pages 1311–1316\.
- Li et al\. \(2023\)Zhuoyan Li, Hangxiao Zhu, Zhuoran Lu, and Ming Yin\. 2023\.Synthetic data generation with large language models for text classification: Potential and limitations\.In*Proceedings of EMNLP 2023*\.
- Lin et al\. \(2023\)Zi Lin, Zihan Wang, Yongqi Tong, Yangkun Wang, Yuxin Guo, Yujia Wang, and Jingbo Shang\. 2023\.ToxicChat: Unveiling hidden challenges of toxicity detection in real\-world user\-AI conversation\.In*Findings of EMNLP 2023*\.
- Lindley \(1956\)Dennis V\. Lindley\. 1956\.On a measure of the information provided by an experiment\.*The Annals of Mathematical Statistics*, 27\(4\):986–1005\.
- Liu and Nocedal \(1989\)Dong C\. Liu and Jorge Nocedal\. 1989\.On the limited memory BFGS method for large scale optimization\.*Mathematical Programming*, 45:503–528\.
- Liu \(2024\)Zefang Liu\. 2024\.Phishing email dataset\.Hugging Face datasetzefang\-liu/phishing\-email\-dataset\(LGPL\-3\.0\)\.
- Maas et al\. \(2011\)Andrew L\. Maas, Raymond E\. Daly, Peter T\. Pham, Dan Huang, Andrew Y\. Ng, and Christopher Potts\. 2011\.Learning word vectors for sentiment analysis\.In*Proceedings of ACL\-HLT 2011*, pages 142–150\.
- Marone et al\. \(2025\)Marc Marone, Orion Weller, William Fleshman, Eugene Yang, Dawn Lawrie, and Benjamin Van Durme\. 2025\.mmBERT: A modern multilingual encoder with annealed language learning\.arXiv:2509\.06888\.
- Metsis et al\. \(2006\)Vangelis Metsis, Ion Androutsopoulos, and Georgios Paliouras\. 2006\.Spam filtering with naive Bayes – which naive Bayes?In*Third Conference on Email and Anti\-Spam \(CEAS 2006\)*\.
- Mozannar and Sontag \(2020\)Hussein Mozannar and David Sontag\. 2020\.Consistent estimators for learning to defer to an expert\.In*Proceedings of ICML 2020*, pages 7076–7087\.
- Naeini et al\. \(2015\)Mahdi Pakdaman Naeini, Gregory F\. Cooper, and Milos Hauskrecht\. 2015\.Obtaining well calibrated probabilities using Bayesian binning\.In*Proceedings of AAAI 2015*, pages 2901–2907\.
- Nie et al\. \(2020\)Yixin Nie, Adina Williams, Emily Dinan, Mohit Bansal, Jason Weston, and Douwe Kiela\. 2020\.Adversarial NLI: A new benchmark for natural language understanding\.In*Proceedings of ACL 2020*, pages 4885–4901\.
- NVIDIA \(2025\)NVIDIA\. 2025\.Nemotron\-Post\-Training\-Dataset\-v1\.Hugging Face datasetnvidia/Nemotron\-Post\-Training\-Dataset\-v1\.
- nibzard \(2026\)nibzard\. 2026\.decision\-model\-benchmark\.GitHub repository,[https://github\.com/nibzard/decision\-model\-benchmark](https://github.com/nibzard/decision-model-benchmark)\. Accessed September 2026\.
- Ong et al\. \(2024\)Isaac Ong, Amjad Almahairi, Vincent Wu, Wei\-Lin Chiang, Tianhao Wu, Joseph E\. Gonzalez, M\. Waleed Kadous, and Ion Stoica\. 2024\.RouteLLM: Learning to route LLMs with preference data\.arXiv:2406\.18665\.
- Qwen Team \(2026\)Qwen Team\. 2026\.Qwen3\.6\-35B\-A3B\.Model card,[https://huggingface\.co/Qwen/Qwen3\.6\-35B\-A3B](https://huggingface.co/Qwen/Qwen3.6-35B-A3B)\. FP8 variant used\.
- Rao and Daumé III \(2018\)Sudha Rao and Hal Daumé III\. 2018\.Learning to ask good questions: Ranking clarification questions using neural expected value of perfect information\.In*Proceedings of ACL 2018*, pages 2737–2746\.
- Rastogi et al\. \(2020\)Abhinav Rastogi, Xiaoxue Zang, Srinivas Sunkara, Raghav Gupta, and Pranav Khaitan\. 2020\.Towards scalable multi\-domain conversational agents: The schema\-guided dialogue dataset\.In*Proceedings of AAAI 2020*, pages 8689–8696\.
- Ren et al\. \(2023\)Allen Z\. Ren, Anushri Dixit, Alexandra Bodrova, Sumeet Singh, Stephen Tu, Noah Brown, Peng Xu, Leila Takayama, Fei Xia, Jake Varley, Zhenjia Xu, Dorsa Sadigh, Andy Zeng, and Anirudha Majumdar\. 2023\.Robots that ask for help: Uncertainty alignment for large language model planners\.In*Conference on Robot Learning \(CoRL\)*\. arXiv:2307\.01928\.
- Saravia et al\. \(2018\)Elvis Saravia, Hsien\-Chi Toby Liu, Yen\-Hao Huang, Junlin Wu, and Yi\-Shin Chen\. 2018\.CARER: Contextualized affect representations for emotion recognition\.In*Proceedings of EMNLP 2018*, pages 3687–3697\.
- Schatzmann et al\. \(2007\)Jost Schatzmann, Blaise Thomson, Karl Weilhammer, Hui Ye, and Steve Young\. 2007\.Agenda\-based user simulation for bootstrapping a POMDP dialogue system\.In*Proceedings of NAACL\-HLT 2007, Companion Volume \(Short Papers\)*, pages 149–152\.
- Shao et al\. \(2024\)Zhihong Shao, Peiyi Wang, Qihao Zhu, Runxin Xu, Junxiao Song, Xiao Bi, Haowei Zhang, Mingchuan Zhang, Y\. K\. Li, Y\. Wu, and Daya Guo\. 2024\.DeepSeekMath: Pushing the limits of mathematical reasoning in open language models\.arXiv:2402\.03300\.
- Socher et al\. \(2013\)Richard Socher, Alex Perelygin, Jean Wu, Jason Chuang, Christopher D\. Manning, Andrew Ng, and Christopher Potts\. 2013\.Recursive deep models for semantic compositionality over a sentiment treebank\.In*Proceedings of EMNLP 2013*, pages 1631–1642\.
- Stepanov et al\. \(2026\)Ihor Stepanov, Aleksandr Smechov, Mykhailo Shtopko, Dmytro Vodianytskyi, and Oleksandr Lukashov\. 2026\.SCX Router: Streaming zero\-shot model selection with a decoder\-KV classifier and a real\-world task ontology\.arXiv:2609\.02292\.
- Stepanov et al\. \(2025\)Ihor Stepanov, Mykhailo Shtopko, Dmytro Vodianytskyi, Oleksandr Lukashov, Alexander Yavorskyi, and Mykyta Yaroshenko\. 2025\.GLiClass: Generalist lightweight model for sequence classification tasks\.arXiv:2508\.07662\.
- Tobi\-Bueck \(2025\)Tobi\-Bueck\. 2025\.Customer support tickets\.Hugging Face datasetTobi\-Bueck/customer\-support\-tickets\(CC BY\-NC 4\.0\)\.
- Vovk et al\. \(2005\)Vladimir Vovk, Alex Gammerman, and Glenn Shafer\. 2005\.*Algorithmic Learning in a Random World*\.Springer\.
- Warner et al\. \(2024\)Benjamin Warner, Antoine Chaffin, Benjamin Clavié, Orion Weller, Oskar Hallström, Said Taghadouini, Alexis Gallagher, Raja Biswas, Faisal Ladhak, Tom Aarsen, Nathan Cooper, Griffin Adams, Jeremy Howard, and Iacopo Poli\. 2024\.Smarter, better, faster, longer: A modern bidirectional encoder for fast, memory efficient, and long context finetuning and inference\.arXiv:2412\.13663\.
- Warstadt et al\. \(2019\)Alex Warstadt, Amanpreet Singh, and Samuel R\. Bowman\. 2019\.Neural network acceptability judgments\.*Transactions of the Association for Computational Linguistics*, 7:625–641\.
- Williams et al\. \(2018\)Adina Williams, Nikita Nangia, and Samuel Bowman\. 2018\.A broad\-coverage challenge corpus for sentence understanding through inference\.In*Proceedings of NAACL\-HLT 2018*, pages 1112–1122\.
- Williams and Young \(2007\)Jason D\. Williams and Steve Young\. 2007\.Partially observable Markov decision processes for spoken dialog systems\.*Computer Speech & Language*, 21\(2\):393–422\.
- Williams \(1992\)Ronald J\. Williams\. 1992\.Simple statistical gradient\-following algorithms for connectionist reinforcement learning\.*Machine Learning*, 8:229–256\.
- Yang et al\. \(2025\)An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al\. 2025\.Qwen3 technical report\.arXiv:2505\.09388\.
- Yılmaz et al\. \(2026a\)Furkan Yılmaz, Habibe Aleyna Taşdemir, and Muhammed Faruk Gözay\. 2026a\.MoganBert\-TR: A Turkish encoder foundation model trained from scratch with a CLM\-to\-MLM curriculum\.arXiv:2608\.25768\.
- Yu et al\. \(2020\)Lili Yu, Howard Chen, Sida I\. Wang, Tao Lei, and Yoav Artzi\. 2020\.Interactive classification by asking informative questions\.In*Proceedings of ACL 2020*, pages 2664–2680\.
- Zaratiana et al\. \(2024\)Urchade Zaratiana, Nadi Tomeh, Pierre Holat, and Thierry Charnois\. 2024\.GLiNER: Generalist model for named entity recognition using bidirectional transformer\.In*Proceedings of NAACL 2024*, pages 5364–5376\.
- Zhang and Choi \(2023\)Michael J\. Q\. Zhang and Eunsol Choi\. 2023\.Clarify when necessary: Resolving ambiguity through interaction with LMs\.arXiv:2311\.09469\.
- Zhang et al\. \(2015\)Xiang Zhang, Junbo Zhao, and Yann LeCun\. 2015\.Character\-level convolutional networks for text classification\.In*Advances in Neural Information Processing Systems 28*, pages 649–657\.
- Zhang et al\. \(2019\)Yuan Zhang, Jason Baldridge, and Luheng He\. 2019\.PAWS: Paraphrase adversaries from word scrambling\.In*Proceedings of NAACL\-HLT 2019*, pages 1298–1308\.
## Appendix ASchemas
Table 4:The twelve schemas\.*Profiles*: decisive\-slot combinations with non\-zero prior\.*Multi\-slot*: prior mass of units that need at least two slots\.*Max post\.*: mean maximum posterior of message\-level evidence in the dry run\.*VOI\>\>\.05*: share of unknown slots with oracle VOI above 0\.05\.SchemaSplitUnitsDecisive/allProfilesUnit massMulti\-slotMax post\. / VOI\>\>\.05banking\_supporttrain54/544\.100–\.3501\.00\.474 / \.498ecommerce\_returnstrain54/532\.061–\.324\.70\.455 / \.474hr\_requeststrain63/424\.080–\.240\.80\.405 / \.532insurance\_claimstrain54/532\.109–\.2561\.00\.456 / \.518it\_helpdesktrain54/580\.130–\.292\.70\.400 / \.449privacy\_requeststrain54/576\.063–\.4401\.00\.549 / \.492telecom\_supporttrain64/532\.090–\.270\.80\.379 / \.410travel\_changestrain54/536\.053–\.350\.85\.516 / \.504saas\_supportzero\-shot54/544\.059–\.452\.88\.562 / \.468parcel\_deliveryzero\-shot54/532\.105–\.3411\.00\.485 / \.492student\_affairszero\-shot64/532\.067–\.282\.80\.441 / \.439clinic\_routingzero\-shot64/532\.050–\.242\.80\.400 / \.431#### Gates\.
Every unit is reachable, unit masses lie in\[0\.03,0\.5\]\[0\.03,0\.5\], at least 5% of the mass needs two or more slots, and a dry run of the exact posterior gives a mean maximum posterior in\[0\.35,0\.75\]\[0\.35,0\.75\]and a share of unknown slots withVOI⋆\>0\.05\\mathrm\{VOI\}^\{\\star\}\>0\.05in\[0\.2,0\.6\]\[0\.2,0\.6\]\. A validator enforces that the last rule is*else*and that non\-decisive slots occur in no rule; it caught a telecom slot used in no rule, fixed by adding a field\-technician unit\.
#### Pilot history\.
In pilot 1 the partial\-answer hint accuracy was 0\.70–0\.79 instead of 0\.5; the bias was confined to probes of slots already stated in the message, which the checker read from context, so such probes are now always clean\. Message leakage clustered on overlapping slot pairs \(banking*unrecognized transaction*/*authorized = no*\)\. Pilot 2 added*not applicable*values, which the generator rendered as “no” and produced 1,293–1,404 answer contradictions per schema; they were removed\. Pilot 3 cut message rejection to 0\.32–0\.35 in banking and privacy and answer rejection to 1\.6–3\.6%\. The first full run lost up to 36% of cases in telecom and clinic\_routing because non\-decisive dates implied decisive facts and the generator invented concrete problems when the topic was hidden; the final run replaced the non\-decisive slot with the customer’s name and forbade invented specifics\.
## Appendix BTraining Data
Table 5:Per\-schema leakage test on seen schemas: model minus posterior probability of the gold unit, and accuracy minus ceiling, with case\-clustered bootstrap 95% intervals where the interval excludes or nearly excludes zero\.RL0RL1Schemap\(gold\)p\(\\text\{gold\}\)model−\-post\.Acc−\-ceilingp\(gold\)p\(\\text\{gold\}\)model−\-post\.Acc−\-ceilingbanking−\-\.020 \[−\-\.037,−\-\.007\]−\-\.054 \[−\-\.100,−\-\.013\]−\-\.014 \[−\-\.027,−\-\.004\]−\-\.038 \[−\-\.074,−\-\.002\]ecommerce\+\.001 \[−\-\.002, \.003\]−\-\.016−\-\.000−\-\.010hr−\-\.002 \[−\-\.005, \.000\]−\-\.000\+\.003 \[\.000, \.005\]\+\.003insurance−\-\.005 \[−\-\.008,−\-\.002\]\+\.041 \[\.002, \.080\]−\-\.000 \[−\-\.005, \.004\]\+\.038 \[−\-\.001, \.076\]it\_helpdesk−\-\.003−\-\.006\+\.002−\-\.008privacy−\-\.009 \[−\-\.015,−\-\.005\]\+\.021 \[−\-\.015, \.056\]−\-\.004\+\.018telecom−\-\.004−\-\.025−\-\.001−\-\.034 \[−\-\.067,−\-\.001\]travel−\-\.002−\-\.016−\-\.003−\-\.017#### General v1 \(19,000\)\.
Typed\-decisions train split \(6,000\); CLINC150\-plus with random subsets of 5–15 intents \(5,000\); MultiNLI entailment/contradiction as yes/no with the hypothesis as instruction \(3,000\); Yelp reviews as class\-balanced score questions \(2,000\); SGD first user turns with the active intent among the intents of the service and two to four other services \(3,000\)\. Each source has 10–12 instruction wordings\.
#### Tool choice \(6,000\)\.
One of 13 shards of Nemotron\-Post\-Training\-Dataset\-v1: the state is the user message, the options are the row’s tool set with distractors, and the gold is the single tool called in the first assistant turn\. We removed rows whose gold tool name appears in the user message \(2,930, a dataset artifact\), rows without a call or with several distinct calls, and set sizes outside 2–12; train and test \(5,447 / 553\) are split by tool set\. LLM\-written tool schemas were explored as VOI schemas but dropped: only 2 of 200 passed the gates, and an assistant does not ask the user which tool to call\.
#### Laya sources \(19,437\)\.
AG News train \(2,500\), BoolQ \(3,000\), Enron spam train \(1,937\), phishing e\-mails \(2,000\), MS MARCO v1\.1 train \(2,500\) and support tickets \(2,500\), plus question\-form yes/no items from CLINC, Yelp, MultiNLI and SGD \(5,000\)\. All 4,447 state texts of Laya’s benchmark builder and 50 further Enron subject\-line overlaps were removed\.
#### Open diversity \(38,200, 27 sources, train splits only\)\.
Choice \(17,000\): DBpedia and Yahoo Answers\([Zhang et al\., 2015](https://arxiv.org/html/2609.30706#bib.bib83)\), 20 Newsgroups\([Lang, 1995](https://arxiv.org/html/2609.30706#bib.bib45)\), Tweet Topic\([Antypas et al\., 2022](https://arxiv.org/html/2609.30706#bib.bib7)\), GoEmotions\([Demszky et al\., 2020](https://arxiv.org/html/2609.30706#bib.bib21)\), TweetEval emotion and sentiment\([Barbieri et al\., 2020](https://arxiv.org/html/2609.30706#bib.bib11)\), MATH topic\([Hendrycks et al\., 2021](https://arxiv.org/html/2609.30706#bib.bib34)\), Dolly task type\([Conover et al\., 2023](https://arxiv.org/html/2609.30706#bib.bib19)\), and domain routing among GSM8K, MATH, MBPP, WritingPrompts, Natural Questions, SQL questions and Dolly\([Cobbe et al\., 2021](https://arxiv.org/html/2609.30706#bib.bib18);[Austin et al\., 2021](https://arxiv.org/html/2609.30706#bib.bib8);[Fan et al\., 2018](https://arxiv.org/html/2609.30706#bib.bib25);[Kwiatkowski et al\., 2019](https://arxiv.org/html/2609.30706#bib.bib43);[b\-mc2, 2023](https://arxiv.org/html/2609.30706#bib.bib9)\)\. Yes/no \(17,400\): SST\-2\([Socher et al\., 2013](https://arxiv.org/html/2609.30706#bib.bib68)\), IMDB\([Maas et al\., 2011](https://arxiv.org/html/2609.30706#bib.bib52)\), Civil Comments\([Borkan et al\., 2019](https://arxiv.org/html/2609.30706#bib.bib12)\), Measuring Hate Speech\([Kennedy et al\., 2020](https://arxiv.org/html/2609.30706#bib.bib39)\), TweetEval hate, offensive and irony, QQP\([Iyer et al\., 2017](https://arxiv.org/html/2609.30706#bib.bib37)\), PAWS\([Zhang et al\., 2019](https://arxiv.org/html/2609.30706#bib.bib84)\), MRPC\([Dolan and Brockett, 2005](https://arxiv.org/html/2609.30706#bib.bib22)\), SNLI\([Bowman et al\., 2015](https://arxiv.org/html/2609.30706#bib.bib13)\), ANLI\([Nie et al\., 2020](https://arxiv.org/html/2609.30706#bib.bib57)\), SciTail\([Khot et al\., 2018](https://arxiv.org/html/2609.30706#bib.bib41)\), CoLA\([Warstadt et al\., 2019](https://arxiv.org/html/2609.30706#bib.bib74)\)\. Score \(3,800\): Amazon star ratings from the English part of MARC\([Keung et al\., 2020](https://arxiv.org/html/2609.30706#bib.bib40)\)\(SetFit copy\), toxicity level from the Civil Comments toxicity score, and hate level from the Measuring Hate Speech score\. Multi\-class sources use random option subsets, 30% of negatively phrased questions have inverted targets, and states are often dictionaries with named fields\. None of Laya’s evaluation sets was used\.
## Appendix CAdditional Results
#### Leakage\.
Table[5](https://arxiv.org/html/2609.30706#A2.T5): for insurance the model assigns the gold unit slightly*less*probability than the posterior; the argmax excess comes from nearly tied posteriors \(e\.g\. 0\.44/0\.56\), and banking shows the same noise in the opposite direction\. The examples on which the model is most confident relative to the posterior contain no hidden clue, only partial answers ambiguous between two values\.
#### Policy gradient\.
Table[6](https://arxiv.org/html/2609.30706#A3.T6)shows the miniature model and Table[7](https://arxiv.org/html/2609.30706#A3.T7)the full\-scale controlled chains\. At full scale the ratio between the policy\-gradient and VOI gradient norms on the shared parameters was 1\.5–7×\\times\.
Table 6:Miniature model \(d=128d=128, 100 examples,wRL=1w\_\{\\mathrm\{RL\}\}=1\): decision KL, VOI loss and gradient norms sent to the shared parameters\.StepKLVOI MSE∥gRL∥\\lVert g\_\{\\mathrm\{RL\}\}\\rVert∥gCE∥\\lVert g\_\{\\mathrm\{CE\}\}\\rVert∥gVOI∥\\lVert g\_\{\\mathrm\{VOI\}\}\\rVert0\.5111\.7300\.0170\.00111\.50100\.0330\.98448\.490\.4910\.225300\.0110\.67328\.710\.1588\.384600\.0070\.12726\.950\.0982\.445Table 7:Full\-scale controlled training on the calibration slice: warm phase atT=1T=1, and joint phase \(VOI Spearman and MSE against sampled targets; KL on our schemas atT=1T=1and after fitting one temperature per type\)\.wRL=1w\_\{\\mathrm\{RL\}\}=1wRL=0w\_\{\\mathrm\{RL\}\}=0warm ep\. 1: acc / KL\.807 / \.230\.838 / \.102warm ep\. 2: acc / KL\.857 / \.073\.896 / \.018warm ep\. 3: acc / KL\.896 / \.019\.927 / \.006joint: Spearman / MSE\.489 / \.0108\.527 / \.0104joint: acc all / ours\.919 / \.965\.918 / \.956joint: KL,T=1→T\{=\}1\\tofitted\.0061→\\to\.0612\.0040→\\to\.0381
#### Temperatures\.
Per\-group temperatures giveT=1\.015T=1\.015–1\.0171\.017for our schemas andT=3\.78T=3\.78–4\.344\.34\(choice\),5\.525\.52–7\.517\.51\(score\) and2\.042\.04–4\.694\.69\(yes/no\) for the general data, whose KL drops from 0\.70–0\.94 to 0\.28–0\.31\.
#### Effect of the broad data\.
Broadening the data kept the dialogue behaviour on seen schemas \(AUC 0\.799→\\to0\.795, oracle 0\.797; accuracy 0\.652→\\to0\.648 against a ceiling of 0\.659\) but lowered VOI rank correlation \(0\.831→\\to0\.763\) and zero\-shot b0\.5 \(0\.579→\\to0\.535\), which the Gini cap restored to 0\.581\. Shuffling the options changes 1% of AG News, 5% of Emotion, 12% of MASSIVE and 29% of Banking77 decisions of the final model\.
## Appendix DImplementation
The experiments used an internal module on top of Laya’s code \(version 0\.3\.11, commit1e28ac2\) that leaves Laya’s files unchanged; the releasedlavoirpackage re\-implements it as a standalone package with its own tests\. The internal module’s 149 tests cover bit\-identical logits with no slots, marker positions under shuffling, the posterior against brute\-force enumeration, zero VOI for known and non\-decisive slots,0≤v^k≤G\(p\)0\\leq\\hat\{v\}\_\{k\}\\leq G\(p\)under the cap, no VOI gradient in the option scorer, and the policy\. Laya’s act/escalate head is trained with a zero\-weighted loss in its notebook, so we do not use it\.Similar Articles
Inference-Time Budget Control for LLM Search Agents
This paper introduces a two-stage inference-time budget control method for LLM search agents, using Value-of-Information scores to optimize tool-call and token allocation during multi-hop question answering.
CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning
This arXiv paper introduces CVPO, a reinforcement learning method for LLMs that adapts value-variance for advantage estimation and uses dynamic curriculum learning to match question difficulty, achieving better reasoning performance than VAPO on math tasks.
From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs
This paper studies how audio and visual information flow inside Audio-Visual Large Language Models (AVLLMs), revealing that AVLLMs follow sequential or parallel routing depending on input configuration, and that some tokens can be discarded after information transfer for efficiency.
Prompting is not enough: supervised baselines and leakage control for measuring shared decision-making with LLMs in pediatric encounters
This study compares zero-shot prompting of LLMs against supervised baselines for detecting shared decision-making in pediatric clinical encounters, finding that supervised models outperform zero-shot approaches and highlighting critical data leakage issues in evaluation pipelines.
VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization
This paper introduces a paradigm where Vision-Language Models (VLMs) act as test-time teachers to guide Video Generation Models (VGMs) via differentiable rewards and LoRA optimization, achieving a 16.7-point average improvement on video reasoning benchmarks.