Tag
GPT 5.6 Sol can one-shot convert an arXiv paper into an interactive Marimo notebook, useful for hands-on understanding of papers in interpretability, inference engineering, and more.
FirstResearch introduces a structured framework for LLM scientific discovery agents that generates a Research Question Certificate containing primitive definitions, assumptions, mechanism, falsifiable hypothesis, and failure update rules, making the proposed research question inspectable before execution. Preliminary evaluations using LLM judges show that the certificate-centered approach outperforms baseline systems in audibility and score.
研究语言智能体在长期任务中世界模型塌缩的相变现象,发现状态负载和依赖密度等参数在临界点附近导致模型突然崩溃,而非逐渐退化。
This paper catalogues five recurring MCP server architectural patterns observed across fifteen independently developed servers, providing a taxonomy with context, problem, solution, and consequences. It also documents anti-patterns, cross-cutting concerns, and quantitative evaluations including inter-rater reliability and transport overhead.
S-GAI is a spectral geometry-aware initialization framework for one-hidden-layer sigmoidal MLPs that uses class-wise spectral geometry from image data to initialize weights, outperforming random initialization in terms of starting hidden state quality and achieving comparable final accuracy on benchmarks like MNIST and CIFAR-10.
This update to the RLM arXiv paper adds depth>1 experiments with recursive RLM calls, showing significant performance gains on OOLONG-Pairs and other benchmarks, along with new comparisons to OpenCode and Claude Code, additional training results on MRCRv2, and an expanded error analysis.
An OpenAI research preview explores learning from how people interact with their computers beyond chat, accompanied by a new arxiv paper on the topic.