TRACE: Business Rule-Grounded Reasoning Curriculum for Knowledge-Preserving Parametric Tool Retrieval in Enterprise LLMs
Summary
TRACE introduces a two-stage curriculum to preserve parametric tool knowledge in enterprise LLMs while enabling fast single-beam greedy decoding, achieving improved accuracy and recall over baselines.
View Cached Full Text
Cached at: 07/28/26, 02:25 PM
Paper page - TRACE: Business Rule-Grounded Reasoning Curriculum for Knowledge-Preserving Parametric Tool Retrieval in Enterprise LLMs
Source: https://huggingface.co/papers/2607.22639
Abstract
ParametricretrievalenablesLLMstoretrievetoolsimplicitlybyassigningeachAPIauniquevirtualtokenandtrainingthemodeltogenerateitviaconstrainedbeamsearch.Toolsenseshowsthatthisregimehastwocriticaldrawbacks:itdestroysparametrictoolknowledgeduringtraining,anditsbeam-searchdecodingistooslowforreal-timedeployment.WeintroduceTRACE(ToolRetrievalviaAugmentedChain-of-thoughtandEnterpriserules),atwo-stagecurriculumthatresolvesthisdissociation.Stage1reusesthemulti-formatmemorizationSFTfromToolSensetoseedtoolknowledgewithLoRA.Stage2isourcorecontribution:themodelistrainedtoemitathinkingtracebeforeproducingaJSONlistoftooltokens,usingtwodatasources--RRBpairsfromToolSenseandqueriessynthesizedtotargetbusinessrulescuratedbydomainexperts--bothaugmentedwithreasoningtraces.ThistrainingobjectivepreservesStage1MCQandQAprobingaccuracywhileenablingsingle-beamgreedydecodingatproductionlatency.Evaluatedonacombinedenterprisecatalogof8,300+toolsacrosstwoenterpriseproductlines,TRACEtrainingforStage2notonlypreservesbutimprovestoolunderstanding:MCQaccuracygains+3.2ppandQAprobinggains+9ppoverStage1.Onretrieval,TRACEachieves~86%recallonDomainAand~60%onDomainB--comparedtoembeddingbaselineperformanceof~27%&~52%--bothwithsingle-beamgreedydecoding,makingitdirectlydeployableatproductionlatency.
View arXiv pageView PDFAdd to collection
Get this paper in your agent:
hf papers read 2607\.22639
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2607.22639 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2607.22639 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2607.22639 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Beyond Reasoning: Reinforcement Learning Unlocks Parametric Knowledge in LLMs
This paper investigates whether reinforcement learning can improve the direct recall of parametric knowledge in LLMs beyond reasoning tasks. It demonstrates that RL with binary rewards yields significant gains in factual QA benchmarks by redistributing probability mass to unlock latent knowledge rather than acquiring new facts.
ToolSense: A Diagnostic Framework for Auditing Parametric Tool Knowledge in LLMs
ToolSense is an open-source diagnostic framework that generates three benchmarks (realistic retrieval, MCQ probing, QA probing) to audit LLMs' parametric tool knowledge, revealing a knowledge-retrieval dissociation where strong retrieval performance can coexist with poor factual understanding.
RIMRULE: Improving Tool-Using Language Agents via MDL-Guided Rule Learning
RimRule proposes a neuro-symbolic method that distills compact, interpretable rules from failure traces using the Minimum Description Length principle, improving LLM tool-use performance without modifying weights, and demonstrating rule portability across models.
SGR: A Stepwise Reasoning Framework for LLMs with External Subgraph Generation
Introduces SGR, a stepwise reasoning framework that enhances LLM reasoning by generating query-specific subgraphs from external knowledge bases, improving accuracy and factual reliability.
Stepwise Reasoning Enhancement for LLMs via External Subgraph Generation
This paper proposes SGR, a framework that enhances LLM stepwise reasoning by integrating external knowledge graphs through query-relevant subgraph generation, combining Cypher-based reasoning with collaborative reasoning integration. Experiments on CWQ, WebQSP, GrailQA, and KQA Pro show improved reasoning accuracy over standard prompting and knowledge-enhanced baselines.