programming-benchmarks

Tag

Cards List
#programming-benchmarks

Vibe Coding Ate My Homework: An evaluation of AI approaches to greenfield software engineering and programming

arXiv cs.AI ↗ · 2026-06-18 Cached

This paper evaluates the viability of 'vibe coding'—using natural language prompts with AI agents to generate code without human review—for greenfield software engineering tasks, and analyzes existing benchmarks for measuring LLM programming proficiency. The authors develop an evaluation suite for simple Python programming tasks to provide scoped insights.

0 favorites 0 likes
← Back to home

Submit Feedback