Tag
This paper evaluates a specification frame that improves LLM-generated code by reducing defects in critical areas like money arithmetic and access control, based on a pre-registered five-model study across finance, healthcare, and insurance tasks.
This blog post argues that GitHub's collaboration paradigm (branches, pull requests, code reviews) is ill-suited for the modern AI-driven software development era where LLMs and agents generate code at high velocity, calling for rethinking of tools and workflows.
This paper formalizes the 'patchwork problem' where LLM-generated code is locally correct but structurally incoherent across a codebase, proposes a taxonomy of eight failure categories and a hybrid verification framework, and demonstrates that many failures evade current tools.
This paper proposes a source-guided protocol where an LLM generates candidate modifications for a weak target model using a stronger same-family source model, showing substantial accuracy improvements on CIFAR-10 and SVHN benchmarks while disentangling transfer from adaptation effects.
Metal-Sci introduces a 10-task benchmark for optimizing scientific computing kernels on Apple Silicon, paired with an evolutionary search framework driven by large language models. The study evaluates models like Claude Opus 4.7, Gemini 3.1 Pro, and GPT 5.5, demonstrating significant speedups while using out-of-distribution testing to catch silent performance regressions.
KernelBench-X is a new benchmark for evaluating LLM-generated GPU kernels, revealing that task structure impacts correctness more than method design and that correctness does not guarantee hardware efficiency.