GPT-5.6 Sol hits the ZeroBench human baseline at pass@5 without tools

Reddit r/singularity Models

Summary

GPT-5.6 Sol reportedly hits the ZeroBench human baseline at pass@5 without tools, meaning at least one of five attempts succeeds on the benchmark.

pass@5: Scores if at least one of the 5 attempts is correct pass^5: Scores only if all 5 attempts are correct "pass@1": Not a true single-attempt pass@1, it's the average score across 5 attempts. https://zerobench.github.io/
Original Article

Similar Articles

GPT 5.6 Sol benchmarks

Reddit r/singularity

GPT 5.6 Sol achieves new benchmark results, showcasing performance improvements in AI language modeling.