Tag
This paper introduces EFCA, a multi-timescale credit assignment method for agentic reinforcement learning that uses short-term feedback and medium-term state-history signals from environment interaction to improve task success and quality on ALFWorld and WebShop.
This paper introduces a deployable per-instance, multi-layer activation steering technique for large language models, showing that optimal layer selection varies per input and can be predicted from the prompt embedding without gold labels at inference.
This paper investigates whether pretraining LLMs on artificial languages (pre-pretraining) consistently improves token efficiency across multiple natural languages, finding that gains are highly dependent on experimental setup and random seed, though stable gains appear for small models with the Llama tokenizer.
The paper introduces STEMMA, a multi-agent framework that adversarially probes self-identity consistency in LLMs, motivated by concerns that knowledge distillation may transfer behavioral traits like identity representation from teacher to student models.
NeuPAT is a lightweight, architecture-agnostic framework that allocates neuron-wise update constraints during multimodal instruction tuning to preserve language capabilities in MLLMs, recovering 94.5% of language degradation from vanilla tuning.
This paper introduces Prompt Embedding Probes (PEP), a parameter-efficient extension of linear probes that uses learnable prompt embeddings on hidden states to detect hallucinations in frozen LLMs. Evaluations on TriviaQA, GSM8K, and MedQA with Qwen3 models show improvements over standard linear probes, including in pre-generation and cross-model settings.
This paper introduces an exam-style evaluation to study how reasoning models allocate a shared test-time compute budget across multiple questions. It finds that models fail to strategically ration compute, instead prioritizing questions by presentation order and ignoring value or difficulty.
This paper introduces a constant-aware comparison protocol for average-reward reinforcement learning regret bounds, deriving an explicit finite lower certificate for communicating MDPs and improving published coefficients.
Introduces CODS, an iterative critic-guided data selection method for offline reinforcement learning that retains task performance at low data budgets by selecting high-residual transitions over multiple rounds.
The paper introduces SkillConsist, a method using bidirectional graph alignment to detect inconsistencies between declared and implemented behavior in LLM agent skills, achieving strong F1 improvements over baselines.
This paper formalizes when benchmark contamination is detectable, deriving information-theoretic limits and proposing power-calibrated audits that distinguish a clean benchmark from a powerless detector. It reports two-sided empirical findings on calibration efficacy and validity gates.
This paper introduces GRACE, a framework that uses LLM-generated semantic descriptions at the attribute-value level to create unified metric spaces for clustering mixed tabular data, achieving scalability comparable to statistical baselines while improving clustering accuracy.
This paper shows that LLM judges embedded in reasoning pipelines often make poor decisions, and proposes Evidence-Locked Derive–Gate–Repair (EL-DGR) to constrain judge overrides with evidence certificates, improving accuracy over majority vote and first-candidate baselines.
Introduces AndroidReality, a perturbation-based framework for evaluating and improving the robustness of mobile agents, with a taxonomy of real-world interface perturbations and a training-free Test-Time Introspective Recovery (TTIR) mechanism.
QuantumMind presents an auditable agentic workflow that automatically generates and screens quantum speedup hypotheses using typed role-specialized actions and a deterministic validator.
This paper introduces the Mendel Gödel Machine, a recursive self-improving framework that applies comparative evolution to iteratively improve coding agents.
Agent-MD is a framework that selectively applies LLM reasoning to long-running molecular simulation campaigns, using a deterministic rule-based agent for routine tasks and event-triggered LLM review for exceptional conditions. Demonstrated in GCMC–MD water-vapor desorption simulations, it shows that auditable, reproducible scientific workflows can avoid placing every operation inside an LLM reasoning loop.
This paper studies how LLM agents negotiate in a dynamic supply chain bargaining problem, benchmarking nine models from OpenAI, Google, and Alibaba against a Bayesian equilibrium and finding that capability, provider identity, and prompt design shape surplus creation and division.
This paper presents a formal framework for constructing canonical interpretations from plural structure theories, motivated by structural failures in LLM-assisted reasoning. It distinguishes types of non-determinism and provides conditions for licensed canonicalization, without establishing full determinization for all cases.
This arXiv paper proposes a 'flow-by-flow' content-judgment bypass approach to govern AI outputs in high-loss domains, aiming to reduce harm from incorrect or unsafe AI-generated content.