Tag
This paper evaluates the viability of 'vibe coding'—using natural language prompts with AI agents to generate code without human review—for greenfield software engineering tasks, and analyzes existing benchmarks for measuring LLM programming proficiency. The authors develop an evaluation suite for simple Python programming tasks to provide scoped insights.