Tag
SkillDRE introduces a dual-stage feedback loop for autonomously evolving malicious skill packages in agents, achieving high attack success rates while maintaining benign task functionality.
This study benchmarks coding agents on reproducing Eurostat statistics, finding that semantic validation and a retry budget are crucial for reliability, not just execution diagnostics.
The paper introduces WER, a multi-phase framework that trains a Skill Optimizer using reinforcement learning from execution feedback to improve tool-using agents, achieving significant performance gains on benchmarks like BFCL v4 and τ2-bench.
FlowScout is a framework that automatically generates tool-integrated agentic workflows from historical task-solving records, using Monte Carlo tree search guided by execution feedback. Experiments show it improves tool invocation correctness and execution score over baselines.
This paper introduces the 'crystallization problem' for evaluating reusable memory in text-to-SQL systems, showing that storing verified corrected queries in a per-database bank improves held-out first-attempt accuracy by 4.34 points on BIRD, capturing 44.4% of the headroom provided by on-demand repair. Controlled interventions identify database-specific content as the main driver.
This paper presents a self-debugging technique where an agent iteratively generates, executes, and explains its own code to find bugs without error messages, improving accuracy by up to 12% and matching baselines that generate 10x more candidates.
EXPO-SQL proposes a fine-grained clause-level policy optimization method for Text-to-SQL, using execution feedback to assign rewards per clause rather than per query, significantly improving performance over existing supervised fine-tuning and RL approaches.
Evoflux uses evolutionary search at inference time to repair failed tool workflows for compact language models, boosting execution feasibility significantly over fine-tuning methods.
Introduces GATE (Grounding After Test from Execution), a method that bootstraps missing semantic groundings from execution feedback to handle under-specified user phrases in text-to-SQL tasks, consistently improving over strong baselines.
CP-Agent presents a calibrated risk-controlled approach for feedback-driven competitive programming using large language models, achieving significant improvements on benchmarks without parameter updates.