code-augmented

Tag

Cards List
#code-augmented

CombEval: A Framework for Evaluating Combinatorial Counting in Large Language Models

arXiv cs.AI · 2026-06-20 Cached

CombEval is a dynamic benchmark for evaluating combinatorial counting in large language models, using typed specifications to generate problems with solver-verified answers. It tests 11 LLMs under direct and code-augmented settings and finds brittleness on ordered objects, indistinguishable elements, relative constraints, and nested dependencies.

0 favorites 0 likes
#code-augmented

SPEAR: Code-Augmented Agentic Prompt Optimization

arXiv cs.CL · 2026-05-27 Cached

SPEAR is a code-augmented agentic prompt optimizer that uses a Python sandbox for structural error analysis, achieving state-of-the-art performance on multiple LLM evaluation suites including industrial judge tasks, BBH, and GSM8K.

0 favorites 0 likes
← Back to home

Submit Feedback