decision-forking

Tag

Cards List
#decision-forking

How Good Are LLMs at Decision Forking? (GitHub Repo)

TLDR AI ↗ · 23h ago Cached

Taste-Bench is a benchmark that evaluates LLMs' ability to choose optimal paths at decision forks in long-horizon tasks, using trajectories from software engineering and machine learning research, with a leaderboard indicating current top models like GPT-5.6 Sol achieving 59.7% accuracy.

0 favorites 0 likes
← Back to home

Submit Feedback