Tag
This paper introduces BehaviorBench, a comprehensive benchmark for evaluating foundation models on behavioral science tasks including behavior prediction, strategic decision-making, subject-trait inference, and behavioral knowledge application. It also presents Be.FM-1.5, a fine-tuned model that achieves strong distributional alignment, highlighting the gap between general-purpose and behaviorally adapted models.
This paper proposes using distributional alignment between task vector-based and in-context learning inference as a criterion for designing task vectors, and introduces Linear Task Vector (LTV) that minimizes next-token probability discrepancy via closed-form linear mapping. LTV achieves 9.2% average accuracy improvement over baselines across eight benchmarks and five LLMs.