All articles, most recently crawled first.
This article discusses the proof of Sendov's conjecture and its variants, focusing on polynomials with zeros in the unit disk and their critical points.
California has approved new tire efficiency standards to save drivers money and reduce emissions, with phased implementation starting in 2029 and stricter standards in 2033.
Israel created a fake think tank to produce reports designed to influence AI chatbots like Claude and Gemini, a practice referred to as LLM poisoning.
NaviDC-OCR is a unified vision-language framework that improves document parsing accuracy through deformation-aware learning and adaptive sampling, achieving state-of-the-art results on multiple benchmarks.
MegaParts introduces a scalable framework for part-aware 3D object generation using token-efficient vector-quantized tokens and autoregressive modeling, enabling generation of objects with up to 300 parts.
The paper systematically analyzes risks in agentic AI systems induced by expanding cognitive capabilities across physical, social, and self-referential levels, and proposes mitigation strategies to ensure their safe development.
This paper proposes VibeWorlding, a framework for benchmarking and training multimodal agents to construct 3D open worlds from user queries, showing that reinforcement learning improves open-source models to compete with closed-source frontiers.
GenRouter is a unified routing framework for agentic image generation that adaptively directs prompts to optimal workflows, significantly reducing costs and latency while improving visual alignment through demand profiling and self-evolution.
R^3-Bench introduces a benchmark showing that LLMs struggle to maintain performance when allocated shared computation budgets across multiple reasoning tasks like math, coding, and abstract reasoning.
This paper introduces a unified black-box reinforcement learning framework for stable and scalable optimization of agents through complex harnesses, using sandbox execution and trajectory reconstruction with improvements on benchmarks like ClawGym-Bench.
ENTLORE is a graph-grounded benchmark framework for enterprise question answering that evaluates latent organizational reasoning, showing that even with gold document sources, many implicit relation questions remain unanswered.
Ventor-QTest proposes a black-box audit method for vendor-hosted LLM APIs, using repeated and long-sequence probes to measure fidelity loss and detect degradation in long-horizon agentic tasks.
ConceptFormer learns continuous latent concept representations for visual document retrieval, bridging visual evidence and semantic relevance without text intermediates, achieving significant improvements over baselines.
This paper proposes D-SCAN, a detection framework for RAG poisoning that monitors attention collapse dynamics in language model generations, showing effectiveness against adversarial attacks even when outputs appear benign.
This position paper argues that AI agents in scientific teams should be studied as human-agent systems to enhance collaboration and mitigate risks such as reduced diversity in scientific inquiry.
UI-Mate is a foundation GUI agent that uses environment-grounded training and in-context demonstrations to improve reliability on long-horizon office tasks, achieving state-of-the-art results on computer-use benchmarks.
VideoGAIA introduces a benchmark for assessing agentic video understanding in multimodal models through complex, multi-turn tasks, revealing that even frontier models like GPT-5.5 achieve less than 60% accuracy.
PACE-Bench introduces a simulator-grounded benchmark for evaluating self-evolving agents on physics adaptation tasks involving iterative code redesign after environmental mutations, revealing that mechanism redesign is a major bottleneck compared to parameter inference.
The Qwen 3.8 27B AI model achieved a score of 52 on the Artificial Analysis Intelligence Index, as highlighted in a blog post by Simon Willison.
The paper proposes improvements to the optimization problem in combination loss analysis using modern techniques and AlphaEvolve, yielding an improved upper bound on the matrix multiplication exponent.