Tag
The author created Any-Bench, a tool to benchmark AI models on personal codebases, addressing shortcomings in existing benchmarks like SWE-Bench.