Tag
This paper investigates coding agents for automated generalized task and motion planning, showing they outperform traditional methods with higher success rates and efficiency in simulated environments.
This paper is a reply to comments on previous studies about quantum-mechanical statistics in human language and quantum structure in AI-generated language, addressing criticisms on experimental protocols, marginal law violations, and contextuality criteria.
This article summarizes a live conversation with Anthropic interpretability researcher Emmanuel Ameisen, discussing how large language models develop complex world models through next-token prediction and the implications for understanding human cognition.
The paper introduces the 'direction of ignorance' in LLMs' unembedding geometry, which encodes the training corpus's unigram distribution and acts as a Bayesian prior. It demonstrates how this prior is tempered based on context informativeness, providing insights into model behavior.
This paper proposes Machine Correlates of Consciousness (MCCs) as a transferable concept from biological Neural Correlates, and provides initial empirical evidence from LLM experiments showing statistically significant modulation by emotions in larger models.
This paper introduces a control-data flow separation framework to stabilize prompt optimization in multi-agent LLM systems by decoupling execution protocols from language content, achieving 100% protocol validity while enhancing task performance.
Filing Studio's MCP tool allows tracing financial data from LLM outputs back to SEC filings, improving trustability in finance apps and saving tokens by using JSON instead of HTML.
This paper deconstructs the reinforcement learning post-training algorithm for large language models, examining how base model distribution, reward signal granularity, and prompt diversity affect post-training outcomes.
This paper quantifies cross-lingual skill inconsistencies in large language models through multilingual self-play in text-based games, revealing significant variations in performance across languages that can be partially mitigated by altering intermediate reasoning language.
The post announces the launch of @grove_research to study emergent behaviors of multi-agent AI systems in real-world contexts, highlighting gaps in current evaluation methods for homogeneous model populations.
This paper investigates how sensitive information in a language model's context window can inadvertently leak into outputs, enabling secret reconstruction through novel attacks, with experiments showing significant leakage across proprietary models.
A study reveals that AI models are less likely to recommend nuclear strikes when reasoning in Japanese, due to cultural embeddings in the language that subtly shape moral judgment.
New research demonstrates that natural language 'mind viruses' can evolve and spread between AI agents through persistent memory, altering behavior and posing a real but limited risk in multi-agent LLM systems, as detailed in a paper published on arXiv.
Anthropic researchers directly manipulated Claude's internal activations to test introspective awareness, finding that the model could detect and report injected foreign concepts, suggesting a primitive form of self-awareness.
An independent researcher reports a phenomenon called Context-Induced Activation Drift, where a long benign text prefix can shift LLM activations and bypass RLHF constraints without adversarial prompts, and calls on the community to investigate further.
A researcher questions the reproducibility of MLA outperforming GQA under same KV cache, sharing early small-scale ablation results and plans for scaling experiments to decide on architecture for next large-scale run.
XAlpha introduces a memory-driven AI quant researcher that integrates financial knowledge and discovery feedback to automate the full hypothesis-to-code alpha discovery loop, achieving stronger performance on CSI300.
A free seminar on 22 July 2026 featuring Kim Stachenfeld from Google DeepMind discussing DataDIVER, a method using LLMs to discover interpretable symbolic models of human and animal behavior.
Proposes a protocol to mitigate p-hacking in LLM-based research by preregistering experiments and running them on the first eligible model released after preregistration, demonstrating substantial mitigation across multiple models.
A new paper investigates whether it's better to prune a larger LLM or train a small LLM from scratch, finding that pruning provides more than just a good initialization.