behavioral-alignment

标签

Cards List
#behavioral-alignment

从提示到行为对齐:用于推荐评估的个性化LLM评判器

arXiv cs.AI · 19小时前 缓存

本文介绍了用于推荐评估的个性化LLM评判器的行为对齐框架,解决了双向合理化问题——即现成的LLM对同一项目既能论证用户参与的正向结果,也能论证负向结果。通过微调和偏好优化,该方法相比零样本基线在Macro-F1上提升了32.19%,并达到了生产环境中基于特征工程的基线水平。

0 人收藏 0 人点赞
#behavioral-alignment

On the use of foundation models in cognitive science

arXiv cs.CL · 2天前 缓存

This perspective paper from arXiv articulates a four-stage inferential framework for evaluating foundation models as cognitive and developmental models, emphasizing that behavioral alignment alone is insufficient and must be embedded within theoretical commitments and contrastive evaluation.

0 人收藏 0 人点赞
#behavioral-alignment

通过在精选数据集上进行训练来改进语言模型行为

OpenAI Blog · 2021-06-10 缓存

OpenAI 研究表明,通过在针对特定行为价值观的小型精选数据集(<100 个示例)上进行微调,可以显著改进语言模型的行为,且效果随着模型规模增大而提高。该方法为用户提供了工具,以便根据特定应用调整模型以符合《宪章》的价值观。

0 人收藏 0 人点赞
← 返回首页

提交意见反馈