gui-cli

Tag

Cards List
#gui-cli

@rohanpaul_ai: Long-horizon agent reliability has not arrived yet with better models. On WeaveBench's 114 hybrid GUI-CLI tasks, the be…

X AI KOLs Timeline · 23h ago Cached

This paper argues that long-horizon AI agent reliability is lacking despite better models, as seen on WeaveBench with only 41.2% pass rate, and proposes the LongHorizon-Harness to manage task state for improved performance.

0 favorites 0 likes
← Back to home

Submit Feedback