Tag
arXiv has received a 17.2 million investment from Simons Foundation International, Siegel Family Endowment, and XTX Markets to support its operations.
A study found that disclosed AI use in mathematics-related arXiv papers increased significantly from 1.39% to 14.09% in a six-month period.
This paper introduces the first method for continuous gradient descent optimization in machine learning models with p-adic parameters, using the Berkovich affine line to enable effective learning on tasks like modular arithmetic.
This paper distinguishes between aligning AI with human preferences versus human behavior, showing that preference alignment can reduce human-likeness and establishing a Turing-test gap in current alignment methods.
The paper introduces Self-Organizing Agent Teams (SAT), a method where AI agents learn reusable strategies to collaborate and reason together, achieving higher accuracy on mathematics and physics benchmarks compared to individual agents and baseline methods.
This paper introduces Memory of Memory (MoM), a framework for LLM agent memory that commits current values on arrival while retaining displaced values as provenance, improving accuracy and reducing stale answers.
The paper introduces FrontierMath Erdős, a benchmark of 68 open Erdős problems for evaluating AI models in mathematics using Lean proof assistant, aiming to address shortcomings in current AI research demonstrations.
This paper examines how error correlation among LLM judges reduces the statistical independence of consensus, showing that shared mistakes can make agreement appear stronger than it is, leading to incorrect conclusions in up to 28% of cases, and proposes using trusted examples to estimate and account for these dependencies.
This paper adapts the Music Lab experiment to study how social influence affects AI agents' selection of scientific papers, showing that social information reduces attention volume and breadth while increasing between-community variation.
The Austrian Academy of Science, partnering with Mistral and Sail Reply, releases Apollo, the first advanced large language model for Ancient Greek to help scholars restore tattered papyrus fragments by predicting missing words.
The paper introduces PAGE, a partition-aware gated KV-cache eviction method that uses a scalar metric to predict input classes and apply eviction only when safe, reducing accuracy degradation in large language models.
ZoAQ presents an adaptive zeroth-order optimization method that reuses past queries to reduce evaluation counts, saving 43–48% of queries in synthetic tasks and achieving high success in black-box attacks.
This 1996 paper introduces Lifestreams, a storage model for organizing personal digital data into a continuous, searchable stream.
OpenMAS-GCom introduces a diagnostic benchmark for evaluating graph-enhanced multi-agent systems by using controlled interventions to attribute performance differences to specific organizational components.
AutoRecLab is an autonomous Python-based system that automates recommender systems experiments from natural-language prompts, using retrieval-augmented generation (RAG), static verification, and tree search to generate and validate executable code.
This paper explores using a multimodal LLM to generate domain-specific feature pools for time series anomaly detection, showing that a simple statistical pipeline can match the performance of advanced pretrained models on benchmark datasets.
This article surveys the connection between lossless compression algorithms and machine learning, introducing a design framework that makes compression-based methods competitive with conventional baselines and particularly effective for malware classification.
The paper introduces DENSE, a method for distilling AI agent execution traces into evidence-grounded shortcut trees for self-refinement, achieving improved performance on Terminal-Bench without post-hoc outcome labels.
SAGE is a schema-guided system that uses large language models to automate grant review by structuring rubrics and linking evidence, with human-in-the-loop validation to improve accuracy.
The paper introduces ScriptMoE, a script-aware mixture-of-experts architecture for all-in-one multilingual scene text recognition, along with the TextMuSS-10M synthetic dataset, achieving state-of-the-art accuracy on benchmarks.