taste-benchmark

Tag

Cards List
#taste-benchmark

The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

Hugging Face Daily Papers · yesterday Cached

This paper introduces Taste-Bench, a benchmark for measuring taste in LLM agents' long-horizon decisions, finding that frontier models have low accuracy and that taste can be improved through distillation training.

0 favorites 0 likes
← Back to home

Submit Feedback