Tag
The paper introduces a knowledge-gated task-construction protocol to explicitly test LLM agents' dependence on hidden knowledge, validated through calibration tasks showing performance drops without access to private conventions.