Tag
The Vals AI RSI Index indicates that the Opus 5.5 AI model has beaten the published human baseline on one of five tasks, shifting the predicted date for achieving full RSI from August 2027 to July 2027.
GPT-5.6 Sol reportedly hits the ZeroBench human baseline at pass@5 without tools, meaning at least one of five attempts succeeds on the benchmark.
This paper audits six large language models for gender stereotyping across English, Korean, Chinese, and Japanese, anchoring against human baselines. It finds that LLM stereotyping often exceeds human cross-country variation and can compound across languages, introducing a four-pattern framework to characterize such behaviors.