agent-trajectories

Tag

Cards List
#agent-trajectories

trajectory-judge: What Outcome-Only LLM Judges Miss on Agent Trajectories

arXiv cs.CL · yesterday Cached

This paper introduces a deterministic testbed for evaluating LLM judges on agent trajectories, showing that outcome-only judges miss silent faults while step-based judges achieve higher recall with better calibration.

0 favorites 0 likes
#agent-trajectories

@omarsar0: NEW paper from Microsoft and colleagues. Debugging agent trajectories at scale is challenging. This is a clever approac…

X AI KOLs Following · 2026-07-15 Cached

This paper introduces OAT, a lightweight failure attribution tool for LLM-based agentic systems that trains only on successful trajectories and uses neural controlled differential equations to detect error steps, outperforming expensive baselines by orders of magnitude in speed and accuracy.

0 favorites 0 likes
#agent-trajectories

IdeaTrail: Full-Process Agent Trajectories for Scientific Ideation

arXiv cs.AI · 2026-07-14 Cached

IdeaTrail is a dataset of multi-turn process trajectories for scientific ideation, synthesizing research processes from evidence gathering to proposal construction using a Generator–Advisor loop to ensure grounding.

0 favorites 0 likes
#agent-trajectories

Dissecting model behavior through agent trajectories

arXiv cs.AI · 2026-06-17 Cached

This paper introduces the Simple Strands Agent (SSA), a minimal harness designed to reduce the intent-execution gap between AI models and their agentic behavior, and analyzes 138k trajectories across various model families to reveal fine-grained behavioral differences.

0 favorites 0 likes
#agent-trajectories

BraveGuard: From Open-World Threats to Safer Computer-Use Agents

Hugging Face Daily Papers · 2026-06-02 Cached

BraveGuard is a self-evolving defense framework that trains guard models using open-world threat signals and realistic agent trajectories to improve safety detection in computer-use agents, achieving significant accuracy gains on the AgentHazard benchmark.

0 favorites 0 likes
#agent-trajectories

TraceGraph: Shared Decision Landscapes for Diagnosing and Improving Agent Trajectories

arXiv cs.AI · 2026-06-01 Cached

TraceGraph is a graph-based framework that constructs shared decision landscapes from multi-model agent trajectories, enabling diagnosis of failure regions and improvement via trap-aware recovery pipelines.

0 favorites 0 likes
#agent-trajectories

Where Do Deep-Research Agents Go Wrong? Span-Level Error Localization in Agent Trajectories

Hugging Face Daily Papers · 2026-06-01 Cached

This paper introduces a claim-centric auditing framework for identifying error spans in deep-research agent trajectories, along with a new benchmark TELBench, improving process-level reliability assessment.

0 favorites 0 likes
#agent-trajectories

ACC: Compiling Agent Trajectories for Long-Context Training

Hugging Face Daily Papers · 2026-05-21 Cached

Agent Context Compilation (ACC) enhances long-context reasoning in LLMs by converting multi-turn agent trajectories into structured QA pairs, enabling direct supervision of distant context integration without additional annotation.

0 favorites 0 likes
← Back to home

Submit Feedback