trajectory

Tag

Cards List
#trajectory

@Saccc_c: Deepseek harness's trace mode is so interesting, now you can completely know the progress of Agent's work, and your Agent employees will never dare to slack off again. Its specific thinking steps, tool calls, and tool return results are all clearly visible. The most interesting thing is that it uses different colors to distinguish when AI executes tasks…

X AI KOLs Following · 2026-08-13 Cached

DeepSeek Harness has been released, featuring a trace mode that visualizes Agent's thinking steps, tool calls, and return results, and uses colors to distinguish different types of content. It can be installed via npm and is suitable for developers to observe Agent work progress.

0 favorites 0 likes
#trajectory

Beyond Single-Turn Confidence: Trajectory-Adapted Uncertainty Quantification for LLM Agents

arXiv cs.CL · 2026-08-13 Cached

This paper studies whether single-turn uncertainty quantification methods transfer to interactive LLM agent trajectories, evaluating white-box, black-box, and reflexive scorers across five LLMs and four tool-use datasets. Results show that transfer is uneven, with black-box self-consistency often strongest, and recommend revalidating UQ methods at the trajectory level.

0 favorites 0 likes
#trajectory

Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs

Hugging Face Daily Papers · 2026-08-03 Cached

Introduces AD-MCQ and DEFT-RLVR, a method for verifiable reasoning in autonomous driving VLMs that defers future trajectory exposure to post-decision verification, improving reasoning faithfulness while reducing hallucinations.

0 favorites 0 likes
#trajectory

AgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation

arXiv cs.AI · 2026-07-09 Cached

AgentLens is a new open-source benchmark for evaluating coding agents that assesses the full trajectory of interactions, including instruction following, tool use, error recovery, and more, using formal verification and LLM-written reviews.

0 favorites 0 likes
#trajectory

Who Should Lead Decoding Now? Tracking Reliable Trajectories for Ensembling Masked Diffusion Language Models

Hugging Face Daily Papers · 2026-06-15 Cached

This paper proposes TIE, a knowledge fusion framework for masked diffusion language models that tracks confidence dynamics to identify reliable decoding trajectories and iteratively transfers partially denoised sequences between models, improving generation quality on reasoning tasks.

0 favorites 0 likes
#trajectory

@9hills: After trying many Agent Memory implementations, I found only two that are somewhat useful: 1. Hermes-style strictly length-limited entry-level memory and session recall, used to address personal assistant memory needs. But this has nothing to do with coding. 2. Skills precipitated from trajectories and skill evolution...

X AI KOLs Timeline · 2026-05-25 Cached

The author shares insights after trying various Agent Memory implementations, concluding that only strictly length-limited entry-level memory (like Hermes) and skill evolution based on trajectory precipitation are somewhat useful, while other graph-based or card-based methods are ineffective.

0 favorites 0 likes
#trajectory

@itsPaulAi: Woow Nvidia has just released a 2.6B open-source world model You can turn a single image, text prompt and trajectory in…

X AI KOLs Timeline · 2026-05-15 Cached

Nvidia released a 2.6B open-source world model that can generate controllable worlds from a single image, text prompt, and trajectory, running on a single GPU.

0 favorites 0 likes
#trajectory

From Spans to Trajectories: Observability for Long-Running Agents

YouTube AI Channels · 2026-06-27 Cached

The article discusses how traditional observability methods fail as AI agents run for days and hundreds of steps, proposes trajectory-based Observability-Driven Development (ODD) to monitor and debug long-running agents, and introduces related practices on the Honeycomb platform.

1 favorites 1 likes
← Back to home

Submit Feedback