@interjc: Gemini 3.5 Pro 快着点儿吧,我都要把你开除出御三家了
Summary
据泄露的基准测试结果,Gemini 3.5 Pro 在内部评估中性能超过 Claude Fable 5 和 GPT-5.6,零样本性能相比 3.1 Pro 有显著提升,目前处于私下验证测试阶段,即将公开推出。
View Cached Full Text
Cached at: 07/14/26, 06:29 PM
Gemini 3.5 Pro 快着点儿吧,我都要把你开除出御三家了
0xMarioNawfal (@RoundtableSpace): Gemini 3.5 Pro benchmark leak just dropped and the numbers are turning heads.
> Reportedly outperforming Claude Fable 5 and GPT-5.6 in internal evals > Significant zero-shot performance improvements over 3.1 Pro > Currently in private validation and testing
- Public rollout
Similar Articles
Gemini 3.5 Flash Looks Good For How Fast It Is (8 minute read)
Google released Gemini 3.5 Flash, a hybrid speed model that rivals Opus 4.7 and GPT-5.5 in speed and cost while performing well on agentic and coding benchmarks.
Gemini 3.5 flash scores, hasn’t even beat GPT 5.4 xhigh
Gemini 3.5 flash has achieved certain benchmark scores but has not yet surpassed GPT 5.4 xhigh in performance.
@jakevin7: An interesting thing. The DeepSeek V4 technical report conducted a comprehensive evaluation of all major LLMs, concluding that Gemini 3.1 Pro has the strongest world knowledge among all models. Not GPT, not Claude, but Gemini. But when people use Gemini...
According to the DeepSeek V4 technical report's evaluation of mainstream LLMs, Gemini 3.1 Pro is considered to have the strongest world knowledge, but users generally find it hard to use because the model does not proactively use search tools.
Gemini 3.5 Flash Benchmarks
Benchmark results for the Gemini 3.5 Flash model are discussed, likely showcasing its performance across various AI tasks.
What do you all think? Can we say qwen 3.6 27b beats gemini 2.5 pro? Or sonnet 3.7? Because when I tested, I found the 27b do better.
A user asks whether the 27B-parameter Qwen 3.6 model can outperform Gemini 2.5 Pro and Sonnet 3.7 on deep web search, coding, and agentic tasks, and seeks suggestions for the lowest-parameter model that can beat Gemini 2.5 Pro.