llm-coding-agents

Tag

Cards List
#llm-coding-agents

When Agents Implement Systems: A Case Study in Defects, Detection, and Evaluation Rigor

arXiv cs.AI · 2026-09-03 Cached

A case study of an LLM coding agent implementing a multi-component data system, analyzing defects and evaluating retrieval strategies on HotpotQA, highlighting gaps in automated versus empirical testing.

0 favorites 0 likes
#llm-coding-agents

QuoteBench: How Matched Scores Can Hide Command-Path Failures

Hugging Face Daily Papers · 2026-08-13 Cached

QuoteBench reveals that execution-boundary parsing errors significantly reduce LLM coding agent success, and disclosing the boundary helps recover performance, showing that evaluation must account for deployment configuration rather than treating matched scores as intrinsic model properties.

0 favorites 0 likes
#llm-coding-agents

CORVUS: Context Optimization and Reduction Via Underlying Synchronization for LLM Coding Agents

arXiv cs.LG · 2026-07-28 Cached

CORVUS proposes a new trajectory architecture for LLM coding agents that decouples file-read actions from observations by maintaining a synchronized registry of relevant files, reducing input tokens by 9-50% and reasoning cycles by up to 37% while maintaining comparable pass rates on SWE-bench benchmarks.

0 favorites 0 likes
← Back to home

Submit Feedback