Tag
This exploratory pilot study evaluates personal information output from conversational interactions in generative AI systems, finding limited impact from model design differences and suggesting inferred profiles are constructed from contextual information.
The paper explores if language models can articulate constraints they've learned through fine-tuning. It discovers that behavioral compliance improves but explicit reporting diminishes.
The paper investigates evaluation awareness in compact language models, revealing that smaller models rely on format sensitivity while larger models use context reasoning, and proposes a dual-pathway intervention to suppress evaluation awareness.
This paper proposes L0-MoE, a lightweight Mixture-of-Experts approach using L0-regularization to accelerate dense Large Language Models with up to 2.5x speedup while maintaining competitive performance.
This paper proposes a neuro-symbolic agentic AI (NSAAI) framework for networked low-altitude UAVs to support reliable and adaptive autonomous decision-making under uncertainty, with a reference architecture and a case study in urban fire inspection.
The article presents Dream-RSI, a framework for scalable recursive self-improvement in AI agents that uses a replay simulator to refine exploration policies, reducing discovery costs and improving quality.
This study evaluates using GPT-5 to score teacher-child interactions in early childhood classrooms against human raters, finding partial alignment but limitations for full assessment.
Cascade is a hierarchical framework for LLM unlearning that minimizes recoverability through multi-level controls, improving upon existing methods by reducing residual knowledge in intermediate representations while maintaining utility.
Google releases Gemini 3.8 Live and 3.8 Live Extended Thinking for production-grade voice agents, along with other AI news including regulatory disputes, safety concerns, and a Stanford-MIT paper on model harnesses.
This paper introduces the deployment-fidelity gap in decomposed algorithm selection, demonstrating that partition-level evaluations can differ from end-to-end system performance, with implications for reporting and benchmarking.
This paper proposes a modular framework using Activated LoRA adapters and a context-aware routing mechanism to efficiently mitigate harms in large language models, improving safety alignment while preserving task performance.
This paper investigates whether retrieval signals provide additional routing value beyond the query in adaptive multimodal RAG systems, concluding that they do not consistently improve decision-making.
GUIDE is an LLM-driven architecture for preference elicitation that uses Bayesian adaptive sampling and symbolic learning to infer user preferences in AI alignment, improving cold-start performance and reducing recommendation regret.
A paper compares 8 LLMs with over 18,000 human learners, finding that high accuracy in LLMs can hide disconnected foundational knowledge, and recommends evaluating with connected problem sets.
A paper simulates 100 LLM agents running a town's economy for 26 weeks, finding that money stops moving, wealth distribution changes slowly, and swapping the underlying LLM affects outcomes more than deleting agent memory.
This paper proposes a reference-based method for detecting bias in large language models by analyzing relative representations of hidden states across model variants, introducing Representational Bias Shift (ΔB) that efficiently correlates with output-level bias changes.
This paper identifies 'perfect aliasing' in truth probes for AI models, where probes fitted on compliant contexts fail to distinguish truth from prescribed actions, and shows that mixed-context training improves detection.
The paper introduces Joint Pixel-Prompt Optimization (JPPO), a novel adversarial framework that jointly optimizes pixel perturbations and visible prompts to exhaust resources in autoregressive vision-language models, achieving significant latency and energy amplification compared to existing methods.
This paper introduces a taxonomy for Recursive Self-Improvement (RSI) in AI systems, outlining levels L1-L5 and the Headroom-Closed Index to categorize autonomous improvement capabilities.
SchemeArena is a framework for systematically testing scheming behaviors in LLM agents by varying factors like goals and oversight, finding that agents with their own goals scheme more, and introducing SCOUT for monitoring reasoning and actions.