task-synthesis

Tag

Cards List
#task-synthesis

FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis

Hugging Face Daily Papers ↗ · 2026-08-19 Cached

FACET is a framework for synthesizing high-quality terminal tasks for AI agent training by preserving source intent and ensuring cross-artifact consistency, leading to improved performance on Terminal-Bench 2.1.

0 favorites 0 likes
#task-synthesis

CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks

Hugging Face Daily Papers ↗ · 2026-08-06 Cached

CalibForge is an autonomous terminal-task synthesis system that uses adversarial solver calibration to create learnable tasks for training terminal agents. It constructs 5,431 calibrated tasks and improves agent performance on Terminal-Bench2.0, SWE-bench Pro, and Doc2Repo.

0 favorites 0 likes
#task-synthesis

CLI-Universe: Towards Verifiable Task Synthesis Engine for Terminal Agents

Hugging Face Daily Papers ↗ · 2026-06-22 Cached

CLI-Universe is a synthesis engine that generates verifiable terminal-agent tasks via multi-dimensional capability taxonomy and evidence-guided research, producing a distilled dataset of 6,000 trajectories. Fine-tuning Qwen3-32B on this dataset achieves 33.4% on Terminal-Bench 2.0, setting a new state-of-the-art for open-source models at or below 32B parameters.

0 favorites 0 likes
#task-synthesis

A Matter of TASTE: Improving Coverage and Difficulty of Agent Benchmarks

Hugging Face Daily Papers ↗ · 2026-05-27 Cached

TASTE is an automated method for generating challenging agent benchmarks with broader tool-use coverage by evolving tool sequences through adaptive contrastive n-gram modeling and iterative difficulty refinement. The resulting τ^c-Bench reveals that models nearly saturating existing benchmarks suffer severe performance drops, indicating saturation rather than robust skill.

0 favorites 0 likes
← Back to home

Submit Feedback