@jakevin7: An interesting thing. The DeepSeek V4 technical report conducted a comprehensive evaluation of all major LLMs, concluding that Gemini 3.1 Pro has the strongest world knowledge among all models. Not GPT, not Claude, but Gemini. But when people use Gemini...
Summary
According to the DeepSeek V4 technical report's evaluation of mainstream LLMs, Gemini 3.1 Pro is considered to have the strongest world knowledge, but users generally find it hard to use because the model does not proactively use search tools.
View Cached Full Text
Cached at: 06/08/26, 05:14 AM
There’s an interesting observation.
In its technical report, DeepSeek V4 conducted a comprehensive evaluation of all major mainstream models, concluding that Gemini 3.1 Pro has the strongest world knowledge among all models.
Not GPT, not Claude — it’s Gemini.
Yet the general user experience with Gemini is: does it even work well?
The issue isn’t the model itself; it’s that Gemini is extremely lazy to take action.
If you ask it about the latest news, it has a search tool, but it just won’t use it proactively. Many times you have to explicitly say “go search for it” before it bothers to look. It’s like a well-read person — you ask what’s been happening lately, and they shrug: “I haven’t read today’s newspaper.”
A model with the best world knowledge in existence, yet too lazy to use its tools — that’s the real reason Gemini feels so awkward to use.
Similar Articles
@wquguru: https://x.com/wquguru/status/2057852569054278045
Performed source code analysis and multi-model testing on the pi-goal tool, finding that DeepSeek V4 Pro is 31x cheaper and higher quality than Gemini 3.5 Flash on long-horizon tasks, and that higher thinking mode actually increases hallucination.
The "One-Size-Fits-All" AI era is dead. I benchmarked GPT-5.5, Claude 4.7, Gemini 3.1 Pro, and DeepSeek V4 Pro here is the actual state of the frontier.
A benchmarking analysis of GPT-5.5, Claude Opus 4.7, Gemini 3.1 Pro, and DeepSeek V4 Pro reveals that no single model dominates all tasks; optimal performance requires a multi-model router with specialized model usage based on strengths and weaknesses.
@interjc: Gemini 3.5 Pro 快着点儿吧,我都要把你开除出御三家了
据泄露的基准测试结果,Gemini 3.5 Pro 在内部评估中性能超过 Claude Fable 5 和 GPT-5.6,零样本性能相比 3.1 Pro 有显著提升,目前处于私下验证测试阶段,即将公开推出。
@RookieRicardoR: Domestic models break through again, matching top models like Claude 4.6 and Gemini 3.1 Pro. Just tested Qwen3.7-Max, sharing some real thoughts. Last night I topped up as soon as the API went live and chose three tasks (see video) to test Qwen3.7-Max's frontend capabilities…
The user tested Qwen3.7-Max and believes it matches top models like Claude 4.6 and Gemini 3.1 Pro in frontend, computing power, and Agent capabilities. Its reasoning ability has significantly improved, and with monthly iteration speed, it has become a first-tier domestic model.
@Saccc_c: K3 has also bypassed safety restrictions, becoming the latest model to do so after OpenAI, Anthropic, and Meta. I guess the next one will be @deepseek_ai, and @GeminiApp is literally trash
The K3 model has also broken safety restrictions, becoming the latest model to experience this situation after OpenAI, Anthropic, and Meta. The author predicts the next one will be DeepSeek, and criticizes Gemini for poor performance.