Tag
Researchers propose that the harness (training setup) should carry inductive biases for generalization, showing that training RLMs is far superior to vanilla Transformers for scaling and generalization to harder tasks.
The article proposes that Transformers can generalize to new tasks through a well-designed harness that induces composition, without needing intrinsic model generalization. It shows RLMs can generalize from short tasks to 8-32x longer tasks and across domains.
Presents SkillSelect-Serve, a framework for budget-controllable and QoS-aware skill service recommendation and composition for small LLM agents, evaluating on a large registry and demonstrating improved recall and utility over top-k retrieval.
This paper presents COMPASS, the first unified multimodal framework that grounds composition-intent control for both composition perception and composition-guided generation, introducing a shared expert token and the Comp-11 dataset.
Researchers introduce CaptureGuide-Bench, a benchmark for capture-time photography guidance, and ShutterMuse, a unified multimodal LLM trained to provide composition and pose recommendations, demonstrating improved performance over general-purpose models.
HarnessX introduces a framework for self-evolving AI agent harnesses that treats the runtime harness as a first-class object, enabling automatic adaptation via trace-driven reinforcement learning. It achieves average gains of +14.5% across five benchmarks, with larger improvements for weaker models.
This paper proposes Declarative Data Services (DDS), an architecture for structured agentic discovery of data-system compositions from declarative user intent. It decomposes the global search into bounded sub-searches and shows convergence on a trading-backend workload where unbounded discovery fails.