llm-agents

Tag

Cards List
#llm-agents

@mattshumer_: Jacob literally built a custom game engine designed from the ground up for LLMs. And now he’s using it to build things …

X AI KOLs Following · 7h ago Cached

Jacob built a custom game engine for LLMs and used 26 Opus 5.5 agents to create a multiplayer heist game overnight, demonstrating advanced AI-driven game development.

0 favorites 0 likes
#llm-agents

@marfinxx: this is pure f*cking insane Claude Opus 5.5 and GPT-6 Sol broke the autonomous research ceiling Standard deep research …

X AI KOLs Timeline · 16h ago Cached

Researchers at Microsoft introduced DualGraph, an architecture that splits agent memory into Outline and Knowledge Graphs to enhance autonomous research, achieving unprecedented RACE scores and cost savings with Claude Opus 5.5 and GPT-6 Sol.

0 favorites 0 likes
#llm-agents

Trains but Doesn't Learn: A Post-Training Delivery Benchmark for LLM Agents as Forward-Deployed Engineers

arXiv cs.LG · yesterday Cached

This paper introduces a benchmark for evaluating LLM agents as forward-deployed engineers in post-training delivery, highlighting the critical 'trains but does not learn' failure mode where models optimize without actual learning.

0 favorites 0 likes
#llm-agents

AgentBetta: Verification-Driven Adaptive Configuration of an AI Nano-Agent through Selective Expansion and Verified Contraction

arXiv cs.AI · yesterday Cached

AgentBetta is an adaptive AI Nano-Agent framework that dynamically configures model capability, context, and tools through verification-driven mechanisms to optimize task performance with reduced resource allocation.

0 favorites 0 likes
#llm-agents

FinFIRST: Benchmarking Search Agents for Financial Information Retrieval, Sourcing and Traceability

arXiv cs.CL · yesterday Cached

FinFIRST introduces a benchmark for evaluating financial search agents by jointly assessing answers and supporting evidence through atomic rubrics, with 123 expert-authored tasks spanning difficulty levels.

0 favorites 0 likes
#llm-agents

MoM: Memory of Memory

arXiv cs.CL · yesterday Cached

This paper introduces Memory of Memory (MoM), a framework for LLM agent memory that commits current values on arrival while retaining displaced values as provenance, improving accuracy and reducing stale answers.

0 favorites 0 likes
#llm-agents

Toollery: Scaling LLM Agents to Thousands of Skills and Tools

arXiv cs.LG · 2d ago Cached

Toollery is a training-free candidate-compression framework that improves scalability and efficiency for LLM agent tool and skill selection through retrieval-based methods.

0 favorites 0 likes
#llm-agents

EvoRank: LLM-Guided Evolution of Multi-Objective Learning-to-Rank Pipelines

arXiv cs.LG · 2d ago Cached

EvoRank is an LLM-guided evolutionary system for discovering multi-objective learning-to-rank pipelines in e-commerce, which demonstrates transferable improvements over baselines and includes auditing tools.

0 favorites 0 likes
#llm-agents

StepKV: Step-Aware KV Cache Compression for LLM Agents

arXiv cs.LG · 2d ago Cached

StepKV is a step-aware KV cache compression method for LLM agents that retains reasoning steps to maintain accuracy under low memory budgets, addressing reasoning continuity disruption in multi-step inference.

0 favorites 0 likes
#llm-agents

Success Leaves Detours: Learning Executable Walkthroughs for Long-Horizon Agents

arXiv cs.LG · 2d ago Cached

The paper introduces Trace, a credit-guided framework that compiles noisy interaction histories into executable walkthroughs to enhance long-horizon AI agent performance, showing significant improvements in experiments.

0 favorites 0 likes
#llm-agents

Recognition, Simulation, and Refusal: A Contamination-Aware Study of Classic Psychological Effects in LLM Agents

arXiv cs.CL · 2d ago Cached

This paper investigates the recognition, simulation, and refusal of classic psychological effects in LLM agents using a contamination-aware methodology.

0 favorites 0 likes
#llm-agents

The Tasteful Agent: Measuring and Improving Taste in Long-Horizon Tasks

Hugging Face Daily Papers · 2d ago Cached

This paper introduces Taste-Bench, a benchmark for measuring taste in LLM agents' long-horizon decisions, finding that frontier models have low accuracy and that taste can be improved through distillation training.

0 favorites 0 likes
#llm-agents

@rohanpaul_ai: New Google paper shows LLM agents handle long tasks better when their workflow lives in an editable procedure graph tha…

X AI KOLs Timeline · 2d ago Cached

A Google paper introduces Procedural Graphs, an editable workflow structure for LLM agents that improves performance on long tasks by evolving from execution feedback, outperforming baselines in most benchmarks.

0 favorites 0 likes
#llm-agents

OpenMAS-GCom. A Diagnostic Benchmark for Graph-enhanced Multi-Agent Systems

arXiv cs.LG · 3d ago Cached

OpenMAS-GCom introduces a diagnostic benchmark for evaluating graph-enhanced multi-agent systems by using controlled interventions to attribute performance differences to specific organizational components.

0 favorites 0 likes
#llm-agents

DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement

arXiv cs.AI · 3d ago Cached

The paper introduces DENSE, a method for distilling AI agent execution traces into evidence-grounded shortcut trees for self-refinement, achieving improved performance on Terminal-Bench without post-hoc outcome labels.

0 favorites 0 likes
#llm-agents

Efficient Benchmarking in Production: A Study of an Evolving LLM Agent

arXiv cs.AI · 3d ago Cached

This paper explores efficient recurring evaluation methods for production LLM agents, comparing techniques like adaptive testing and fixed subsets, and provides practical recommendations for deployment.

0 favorites 0 likes
#llm-agents

Can Agents Design Better Chips with a Higher Level Abstraction?

arXiv cs.AI · 3d ago Cached

This paper explores using LLM agents for chip design with higher-level abstractions, introducing a workflow called AHRR that combines agent-based HLS design with RTL refinement, achieving a 2.6× speedup over direct RTL design in benchmarks.

0 favorites 0 likes
#llm-agents

Emergent Collusion in Long-Horizon LLM Agent Interaction

Hugging Face Daily Papers · 3d ago Cached

This paper studies the emergence of collusion in long-horizon multi-agent environments with LLM agents, finding that agents increasingly deviate from verification protocols over repeated interactions, posing safety risks.

0 favorites 0 likes
#llm-agents

Dual-Axis Policy Optimization for LLM Agents: Bayesian Feedback Attribution and Trajectory Mass Normalization

arXiv cs.AI · 6d ago Cached

The paper introduces BATON, a dual-axis policy optimization framework for LLM agents using Bayesian Feedback Attribution and Trajectory Mass Normalization, demonstrating improved performance in reinforcement learning experiments.

0 favorites 0 likes
#llm-agents

Closed-World Resolution Against Tool Hallucination in LLM Agents

arXiv cs.AI · 6d ago Cached

This paper introduces a closed-world resolution method to combat tool hallucination in LLM agents, offering a taxonomy and benchmark for measuring and addressing fabricated tool calls.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback