execution-feedback

Tag

Cards List
#execution-feedback

SkillDRE: Dual-Stage Red-Team Evolution of Agent Skills via Pre-Execution and Runtime Feedback

Hugging Face Daily Papers ↗ · 4d ago Cached

SkillDRE introduces a dual-stage feedback loop for autonomously evolving malicious skill packages in agents, achieving high attack success rates while maintaining benign task functionality.

0 favorites 0 likes
#execution-feedback

Can Coding Agents Reproduce Official Statistics? Metadata, Retry Budget and the Limits of Execution Feedback in a Controlled Eurostat Benchmark

arXiv cs.LG ↗ · 2026-09-22 Cached

This study benchmarks coding agents on reproducing Eurostat statistics, finding that semantic validation and a retry budget are crucial for reliability, not just execution diagnostics.

0 favorites 0 likes
#execution-feedback

Write, Execute, Refine: From Skill Followers to Skill Optimizers via Reinforcement Learning from Execution Feedback

arXiv cs.CL ↗ · 2026-08-19 Cached

The paper introduces WER, a multi-phase framework that trains a Skill Optimizer using reinforcement learning from execution feedback to improve tool-using agents, achieving significant performance gains on benchmarks like BFCL v4 and τ2-bench.

0 favorites 0 likes
#execution-feedback

FlowScout: From Execution Feedback to Reliable Tool-Using Agent Workflows

arXiv cs.LG ↗ · 2026-08-12 Cached

FlowScout is a framework that automatically generates tool-integrated agentic workflows from historical task-solving records, using Monte Carlo tree search guided by execution feedback. Experiments show it improves tool invocation correctness and execution score over baselines.

0 favorites 0 likes
#execution-feedback

From Test-Time Scaling to Reusable Memory: Measuring Crystallization in Text-to-SQL

arXiv cs.CL ↗ · 2026-08-10 Cached

This paper introduces the 'crystallization problem' for evaluating reusable memory in text-to-SQL systems, showing that storing verified corrected queries in a per-database bank improves held-out first-attempt accuracy by 4.34 points on BIRD, capturing 44.4% of the headroom provided by on-demand repair. Controlled interventions identify database-specific content as the main driver.

0 favorites 0 likes
#execution-feedback

@leanxbt: This paper completely changed how I think about how an agent fixes its own code: Generate code -> Execute it -> Explain…

X AI KOLs Timeline ↗ · 2026-07-10 Cached

This paper presents a self-debugging technique where an agent iteratively generates, executes, and explains its own code to find bugs without error messages, improving accuracy by up to 12% and matching baselines that generate 10x more candidates.

0 favorites 0 likes
#execution-feedback

EXPO-SQL: Execution-based Clause-level Policy Optimization for Text-to-SQL

arXiv cs.CL ↗ · 2026-06-24 Cached

EXPO-SQL proposes a fine-grained clause-level policy optimization method for Text-to-SQL, using execution feedback to assign rewards per clause rather than per query, significantly improving performance over existing supervised fine-tuning and RL approaches.

0 favorites 0 likes
#execution-feedback

Evoflux: Inference-Time Evolution of Executable Tool Workflows for Compact Agents

Hugging Face Daily Papers ↗ · 2026-06-10 Cached

Evoflux uses evolutionary search at inference time to repair failed tool workflows for compact language models, boosting execution feasibility significantly over fine-tuning methods.

0 favorites 0 likes
#execution-feedback

Bootstrapping Semantic Layer from Execution for Text-to-SQL

arXiv cs.CL ↗ · 2026-06-05 Cached

Introduces GATE (Grounding After Test from Execution), a method that bootstraps missing semantic groundings from execution feedback to handle under-specified user phrases in text-to-SQL tasks, consistently improving over strong baselines.

0 favorites 0 likes
#execution-feedback

CP-Agent: A Calibrated Risk-Controlled Agent for Feedback-Driven Competitive Programming

arXiv cs.CL ↗ · 2026-05-26 Cached

CP-Agent presents a calibrated risk-controlled approach for feedback-driven competitive programming using large language models, achieving significant improvements on benchmarks without parameter updates.

0 favorites 0 likes
← Back to home

Submit Feedback