agentic-benchmarks

Tag

Cards List
#agentic-benchmarks

Reason Wide, Not Deep: Amortizing the Reasoning Premium into Distilled Skills

arXiv cs.AI · 5h ago Cached

This paper proposes amortizing the high token cost of reasoning-mode LLMs by distilling domain-specific skills from existing trajectories into system prompts, recovering most of the reasoning gap on agentic benchmarks while emitting far fewer tokens.

0 favorites 0 likes
#agentic-benchmarks

CurveShift: Is Agent Progress Scalar? Separating Level from Shape

arXiv cs.CL · 2026-08-04 Cached

This paper introduces CurveShift, an analysis method that separates overall ability gains from difficulty-specific improvements in LLM agents. Using METR time-horizon data and LiveCodeBench, it finds that most apparent shifts toward harder tasks are ceiling effects, though a genuine hard-task effect exists for reasoning models in competitive programming.

0 favorites 0 likes
← Back to home

Submit Feedback