verifiable-tasks

Tag

Cards List
#verifiable-tasks

Ignorance or Incompetence? Constructing Knowledge-Gated, Verifiable Tasks for LLM Agents

Hugging Face Daily Papers · 2026-08-31 Cached

The paper introduces a knowledge-gated task-construction protocol to explicitly test LLM agents' dependence on hidden knowledge, validated through calibration tasks showing performance drops without access to private conventions.

0 favorites 0 likes
#verifiable-tasks

LongWoF-Bench: Evaluating EvoMap Genes for Verifiable Long-Workflow Tasks

Hugging Face Daily Papers · 2026-08-24 Cached

The paper presents LongWoF-Bench, a benchmark for evaluating long-workflow tasks, and demonstrates that EvoMap Gene enhances task completion efficiency and reduces token costs by reusing verified execution experience.

0 favorites 0 likes
← Back to home

Submit Feedback