trajectory

Tag

Cards List
#trajectory

AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation

arXiv cs.AI · 2026-07-09 Cached

AgentLens is a new open-source benchmark for evaluating coding agents that assesses the full trajectory of interactions, including instruction following, tool use, error recovery, and more, using formal verification and LLM-written reviews.

0 favorites 0 likes
#trajectory

Who Should Lead Decoding Now? Tracking Reliable Trajectories for Ensembling Masked Diffusion Language Models

Hugging Face Daily Papers · 2026-06-15 Cached

This paper proposes TIE, a knowledge fusion framework for masked diffusion language models that tracks confidence dynamics to identify reliable decoding trajectories and iteratively transfers partially denoised sequences between models, improving generation quality on reasoning tasks.

0 favorites 0 likes
#trajectory

@9hills: After trying many Agent Memory implementations, I found only two that are somewhat useful: 1. Hermes-style strictly length-limited entry-level memory and session recall, used to address personal assistant memory needs. But this has nothing to do with coding. 2. Skills precipitated from trajectories and skill evolution...

X AI KOLs Timeline · 2026-05-25 Cached

The author shares insights after trying various Agent Memory implementations, concluding that only strictly length-limited entry-level memory (like Hermes) and skill evolution based on trajectory precipitation are somewhat useful, while other graph-based or card-based methods are ineffective.

0 favorites 0 likes
#trajectory

@itsPaulAi: Woow Nvidia has just released a 2.6B open-source world model You can turn a single image, text prompt and trajectory in…

X AI KOLs Timeline · 2026-05-15 Cached

Nvidia released a 2.6B open-source world model that can generate controllable worlds from a single image, text prompt, and trajectory, running on a single GPU.

0 favorites 0 likes
#trajectory

From Spans to Trajectories: Observability for Long-Running Agents

YouTube AI Channels · 2026-06-27 Cached

The article discusses how traditional observability methods fail as AI agents run for days and hundreds of steps, proposes trajectory-based Observability-Driven Development (ODD) to monitor and debug long-running agents, and introduces related practices on the Honeycomb platform.

1 favorites 1 likes
← Back to home

Submit Feedback