Tag
The paper introduces Task Specialization Fine-Tuning (TSFT) for contextual reinforcement learning, a framework that uses a simple parametric model and integer linear programming to allocate fine-tuning budgets efficiently, significantly outperforming baselines in task coverage.
This paper introduces a framework that decomposes compound logical answer options into atomic judgments and uses an operator-constrained integer linear program to improve large language model reasoning over AND, OR, and NEITHER/NOR operators. It achieves significant F1 gains on LOGICAL-COMMONSENSEQA and a new benchmark LOGICAL-SATA.
This blog post presents an algorithm using integer linear programming to compute optimal tokenizers for language models, drawing parallels to solving the Traveling Salesman Problem. It notes that while the result is theoretically interesting, practical tokenizers are already near-optimal and the method may not generalize well.