inference-time-harness

Tag

Cards List
#inference-time-harness

Training with Harnesses: On-Policy Harness Self-Distillation for Complex Reasoning

arXiv cs.CL · 2026-05-12 Cached

This paper introduces On-Policy Harness Self-Distillation (OPHSD), a method that internalizes the capabilities of inference-time reasoning harnesses into the base model through self-distillation. The approach improves standalone performance on complex reasoning tasks, allowing the model to retain reasoning scaffolds without permanent external dependencies.

0 favorites 0 likes
← Back to home

Submit Feedback