Tag
The paper introduces a knowledge-gated task-construction protocol to explicitly test LLM agents' dependence on hidden knowledge, validated through calibration tasks showing performance drops without access to private conventions.
The paper presents LongWoF-Bench, a benchmark for evaluating long-workflow tasks, and demonstrates that EvoMap Gene enhances task completion efficiency and reduces token costs by reusing verified execution experience.