Tag
A humorous tweet commenting on the expected progression of AI model intelligence, comparing it from PhD-level to the level of someone who decided not to pursue a PhD.
OpenAI discusses the importance of evals (evaluations) for measuring and forecasting model progress, especially as benchmarks become saturated or gamed, featuring insights from Tejal Patwardhan and Andrew Mayne.