Tag
The article critiques the trend of giving AI agents more autonomy with less human oversight, arguing that it ignores hard-won software engineering practices like code review and staged rollouts, leading to silent failures and unexpected costs.
This research proposal investigates how coding LLM perplexity scales with codebase size across different programming languages, using Lean as a test case for formal languages. It suggests that Lean may have better scaling exponents, leading to safer and more secure software at scale.
Vitalik Buterin argues that AI can make formal verification more practical, helping generate specs and proofs to ensure software behaves correctly, potentially transforming critical software development beyond Ethereum.