Tag
This paper introduces Behavioral Lift to measure how reasoning behaviors in AI thinking models correlate with correct answers, revealing an amplification-lift gap where training amplifies behaviors not most predictive of success.
讨论latent reasoning作为研究方向的价值,并质疑OpenAI o1模型宣称的thinking过程是否真实或只是对外宣传,认为模型可能尚未达到上限。
Google announces stable general availability of Gemini 2.5 Pro and Flash models, introduces new Gemini 2.5 Flash-Lite in preview with lower latency and cost, and updates pricing for the Flash family with adjusted input/output token rates.