reasoning

标签

Cards List
#reasoning

你的智能体上周二相信什么?

Reddit r/AI_Agents · 昨天

本文探讨了智能体记忆系统中有效时间和事务时间之间的区别,认为存储这两种时间戳对于重构智能体在特定时刻所知道的信息至关重要。文章还提议设计基准测试来检验智能体的时间推理能力。

0 人收藏 0 人点赞
#reasoning

DeepSeek V4 Flash 0731

Hacker News Top · 2天前 缓存

DeepSeek V4 Flash 0731 展示了其在 ARC-AGI 基准上的结果,突显了 AI 模型在抽象推理方面的进展。

0 人收藏 0 人点赞
#reasoning

MACRO: Markov Chain Routing of Transformer Layers

arXiv cs.CL · 2天前 缓存

MACRO is a framework that learns task-specific execution routes over frozen LLM layers using Markov chain-based routing, improving reasoning accuracy without modifying model weights. It outperforms prior routing approaches while reducing search time significantly.

0 人收藏 0 人点赞
#reasoning

M$^3$R-Bench:证据支撑的多模态隐喻理解统一基准

arXiv cs.CL · 2天前 缓存

本文介绍了M3R-Bench,一个包含1,000个图文实例的统一证据支撑多模态隐喻理解基准,并提出了M3R-Reasoner,一个结合基于课程学习的推理监督和强化学习的8B参数模型,其性能优于更大的专有模型。

0 人收藏 0 人点赞
#reasoning

Answer First, Reason Later: Commitment Order in Diffusion LLMs

arXiv cs.CL · 2天前 缓存

This paper investigates why diffusion LLMs fail at reasoning tasks: unconstrained token commitment freezes answers early and collapses to answer-only outputs. The authors identify commitment order as the root cause and propose a training-free, frontier-gated decoding intervention that recovers performance while preserving parallel decoding.

0 人收藏 0 人点赞
#reasoning

重采样不如精炼:LLM推理中的测试时自校正

arXiv cs.AI · 2天前 缓存

一种新的无验证器广度-深度精炼框架通过采样多个推理轨迹、利用自我批评迭代地精炼每条轨迹,并通过多数投票进行聚合,从而在测试时提升LLM推理能力。在多个数学基准和开放权重模型上,该方法持续优于贪心解码、多数投票和基于验证器的选择。

0 人收藏 0 人点赞
#reasoning

C$^3$PO: 评估全模态模型中的跨模态组合与反事实表现

arXiv cs.AI · 2天前 缓存

介绍了C$^3$PO,一个包含3,404个样本的基准测试,用于评估多模态LLM中的跨模态组合与反事实推理。研究发现模态主导性导致了大多数失败,即使最佳模型(Gemini-3.1-Pro)也远低于人类准确率。

0 人收藏 0 人点赞
#reasoning

Woodpecker Distillation: Weak Models Diagnose Reasoning Bugs in Strong Models

arXiv cs.AI · 2天前 缓存

Proposes Woodpecker Distillation, a weak-to-strong training framework that uses weak probe models to identify and repair local reasoning bugs in stronger models via contrastive local interventions, improving performance on math reasoning benchmarks.

0 人收藏 0 人点赞
#reasoning

Meta的模型在五项STEM奥林匹克竞赛中夺得金牌

Reddit r/singularity · 2天前

Meta的AI模型在五项STEM奥林匹克竞赛中取得金牌级成绩,展示了先进的解题和推理能力。

0 人收藏 0 人点赞
#reasoning

新模型发布:Ling-3.0-tiny:总参数7.9B,每个token仅激活1.3B——免费一周

Reddit r/LocalLLaMA · 3天前

Ling-3.0-tiny发布:一款混合推理模型,总参数7.9B,每个token仅激活1.3B,免费一周。

0 人收藏 0 人点赞
#reasoning

@OpenAI:除了 GPT-5.6 Luna 的智能升级,Free 和 Go 用户现在也可以使用“Think”按钮,获得更多…

X AI KOLs · 3天前

OpenAI 宣布推出 GPT-5.6 Luna,带来智能升级;现在 Free 和 Go 用户也可以使用“Think”按钮,在更困难的问题上进行更多推理。

0 人收藏 0 人点赞
#reasoning

易于完成,难以选择:探究LLM在ProverbIT基准上的表现

arXiv cs.CL · 3天前 缓存

本文介绍了ProverbIT,一个包含100道多选题的新型意大利语基准,用于测试LLM完成谚语的能力。通过对13个模型的评估,研究发现,在没有正确答案的多选题格式中,模型性能显著下降,这表明模型依赖记忆模式而非深层的语义理解。

0 人收藏 0 人点赞
#reasoning

SPOT: Sparse Probing and Outcome Calibration for On-Policy Distillation

arXiv cs.LG · 3天前 缓存

Introduces SPOT, a method for on-policy distillation that uses sparse probing and outcome calibration to improve reasoning performance in smaller student models while balancing solution quality and coverage.

0 人收藏 0 人点赞
#reasoning

When Absence Is Evidence: Evaluating Completeness-Sensitive Negative Reasoning in Large Language Models

arXiv cs.CL · 3天前 缓存

Introduces CROWN-QA, a benchmark for completeness-sensitive negative reasoning in LLMs, showing models struggle to distinguish justified negative answers from insufficient evidence, often over-closing.

0 人收藏 0 人点赞
#reasoning

STRIVE: Probing Reasoning Limits in Graded Plausibility Generation and Evaluation

arXiv cs.CL · 3天前 缓存

This paper introduces STRIVE, an LLM-based framework for jointly generating and evaluating controlled event sets for psycholinguistic plausibility judgments. Experiments show that adding a global reasoning scratchpad and evaluator-guided refinement substantially improves generation quality, though near-boundary events remain challenging.

0 人收藏 0 人点赞
#reasoning

校准下限:格式修复可在中小规模下伪装成自我纠正

arXiv cs.CL · 3天前 缓存

本文表明,大语言模型表面上的自我纠正增益往往源于格式修复,而非推理能力的提升。在多种模型规模下,格式效应主导内容效应,且在能力较强的模型上内容边际接近零,这表明该领域将少数实测的自我纠正误归因于实际的内容改进。

0 人收藏 0 人点赞
#reasoning

@svpino:我一直在测试Kimi K3,天哪,这是世界上见过的最好的开放权重模型。它是一个2.8T参数的…

X AI KOLs Timeline · 4天前 缓存

Santiago Valdarrama 报道称,Kimi K3 是他测试过的最好的开放权重模型,一个2.8T参数的视觉模型,支持工具调用和推理,并拥有1M上下文窗口。

0 人收藏 0 人点赞
#reasoning

探究鸿沟:为什么更好的AI答案并不自动带来更好的思考

Reddit r/artificial · 4天前

本文探讨了“探究鸿沟”——生成答案与提出好问题之间的差异——认为更廉价的AI答案并不一定能改善思考。

0 人收藏 0 人点赞
#reasoning

立场:LLMs 无法跳跃

Hacker News Top · 4天前 缓存

一篇立场论文,认为大语言模型存在根本性局限,用“无法跳跃”这一隐喻来突出推理或泛化方面的不足。

0 人收藏 0 人点赞
#reasoning

别让我主动索取:大语言模型在溯因推理的主动多轮信息获取中存在不足

arXiv cs.CL · 4天前 缓存

本文介绍了Alien Abduction,一个交互式游戏,用于探究大语言模型在溯因推理中如何获取证据、更新假设以及决定何时停止。研究发现,与自行选择查询相比,模型在预先提供证据和由oracle提供示例的情况下表现更好,揭示了其在主动信息获取方面的不足。

0 人收藏 0 人点赞
Next →
← 返回首页

提交意见反馈