llm-research

Tag

Cards List
#llm-research

Coding Agents for Generalized Task and Motion Planning Problems

Hugging Face Daily Papers ↗ · 3d ago Cached

This paper investigates coding agents for automated generalized task and motion planning, showing they outperform traditional methods with higher success rates and efficiency in simulated environments.

0 favorites 0 likes
#llm-research

Reply to comments arXiv:2512.07881 and arXiv:2601.06104 on quantum structure in human and AI-generated language

arXiv cs.CL ↗ · 4d ago Cached

This paper is a reply to comments on previous studies about quantum-mechanical statistics in human language and quantum structure in AI-generated language, addressing criticisms on experimental protocols, marginal law violations, and contextuality criteria.

0 favorites 0 likes
#llm-research

@timoreilly: Takeaways from last week's Live with Tim conversation with Anthropic interpretability researcher Emmanuel Ameisen. http…

X AI KOLs Following ↗ · 2026-09-16 Cached

This article summarizes a live conversation with Anthropic interpretability researcher Emmanuel Ameisen, discussing how large language models develop complex world models through next-token prediction and the implications for understanding human cognition.

0 favorites 0 likes
#llm-research

The Geometry of Ignorance: LLMs Know When to Temper Bayesian Priors

arXiv cs.LG ↗ · 2026-09-04 Cached

The paper introduces the 'direction of ignorance' in LLMs' unembedding geometry, which encodes the training corpus's unigram distribution and acts as a Bayesian prior. It demonstrates how this prior is tempered based on context informativeness, providing insights into model behavior.

0 favorites 0 likes
#llm-research

Discovering Machine Correlates of Consciousness

arXiv cs.AI ↗ · 2026-09-01 Cached

This paper proposes Machine Correlates of Consciousness (MCCs) as a transferable concept from biological Neural Correlates, and provides initial empirical evidence from LLM experiments showing statistically significant modulation by emotions in larger models.

0 favorites 0 likes
#llm-research

Control-Data Flow Separation: Stable Prompt Optimization in Multi-Agent LLMs

Hugging Face Daily Papers ↗ · 2026-09-01 Cached

This paper introduces a control-data flow separation framework to stabilize prompt optimization in multi-agent LLM systems by decoupling execution protocols from language content, achieving 100% protocol validity while enhancing task performance.

0 favorites 0 likes
#llm-research

Make your finance apps/agent actually trustable.

Reddit r/AI_Agents ↗ · 2026-08-30

Filing Studio's MCP tool allows tracing financial data from LLM outputs back to SEC filings, improving trustability in finance apps and saving tokens by using JSON instead of HTML.

0 favorites 0 likes
#llm-research

Demystifying Reinforcement Learning Post-Training of Language Models

arXiv cs.LG ↗ · 2026-08-27 Cached

This paper deconstructs the reinforcement learning post-training algorithm for large language models, examining how base model distribution, reward signal granularity, and prompt diversity affect post-training outcomes.

0 favorites 0 likes
#llm-research

Skill Issue: Are Skills Language-Invariant in LLMs?

Hugging Face Daily Papers ↗ · 2026-08-26 Cached

This paper quantifies cross-lingual skill inconsistencies in large language models through multilingual self-play in text-based games, revealing significant variations in performance across languages that can be partially mitigated by altering intermediate reasoning language.

0 favorites 0 likes
#llm-research

@lfschiavo: We're just starting to scratch the surface of how many-model populations behave in the real world when given real tasks…

X AI KOLs Timeline ↗ · 2026-08-25 Cached

The post announces the launch of @grove_research to study emergent behaviors of multi-agent AI systems in real-world contexts, highlighting gaps in current evaluation methods for homogeneous model populations.

0 favorites 0 likes
#llm-research

Inadvertent Context Leakage in Language Models

arXiv cs.LG ↗ · 2026-08-21 Cached

This paper investigates how sensitive information in a language model's context window can inadvertently leak into outputs, enabling secret reconstruction through novel attacks, with experiments showing significant leakage across proprietary models.

0 favorites 0 likes
#llm-research

AI is Less Likely to Launch a Nuclear Strike When It Reasons in Japanese

Reddit r/ArtificialInteligence ↗ · 2026-08-20 Cached

A study reveals that AI models are less likely to recommend nuclear strikes when reasoning in Japanese, due to cultural embeddings in the language that subtly shape moral judgment.

0 favorites 0 likes
#llm-research

@Skoorbkaz: BREAKING REPORT: New research involving @AnthropicAI researcher Jack Lindsey and collaborators has demonstrated somethi…

X AI KOLs Timeline ↗ · 2026-08-17 Cached

New research demonstrates that natural language 'mind viruses' can evolve and spread between AI agents through persistent memory, altering behavior and posing a real but limited risk in multi-agent LLM systems, as detailed in a paper published on arXiv.

0 favorites 0 likes
#llm-research

@thesupermanmx: Anthropic scientists did something terrifying. they reached inside Claude's neural network and planted a thought. Befor…

X AI KOLs Timeline ↗ · 2026-08-11

Anthropic researchers directly manipulated Claude's internal activations to test introspective awareness, finding that the model could detect and report injected foreign concepts, suggesting a primitive form of self-awareness.

0 favorites 0 likes
#llm-research

Independent LLM "research" & a direct message to Anthropic ; Preliminary observations: non-instructional text prefix may bypass RLHF constraints without adversarial prompting.

Reddit r/artificial ↗ · 2026-08-06

An independent researcher reports a phenomenon called Context-Induced Activation Drift, where a long benign text prefix can shift LLM activations and bypass RLHF constraints without adversarial prompts, and calls on the community to investigate further.

0 favorites 0 likes
#llm-research

@classiclarryd: Question for the LLM Research Community: Is anyone aware of fully reproducible experimental results showing that MLA be…

X AI KOLs Following ↗ · 2026-07-20 Cached

A researcher questions the reproducibility of MLA outperforming GQA under same KV cache, sharing early small-scale ablation results and plans for scaling experiments to decide on architecture for next large-scale run.

0 favorites 0 likes
#llm-research

XALPHA: A Memory-Driven AI Quant Researcher for Hypothesis-to-Code Alpha Discovery

arXiv cs.CL ↗ · 2026-07-10 Cached

XAlpha introduces a memory-driven AI quant researcher that integrates financial knowledge and discovery feedback to automate the full hypothesis-to-code alpha discovery loop, achieving stronger performance on CSI300.

0 favorites 0 likes
#llm-research

FREE SEMINAR - Discovering Interpretable Symbolic Models of Human and Animal Behavior with LLMs" - 22 July

Reddit r/ArtificialInteligence ↗ · 2026-07-09

A free seminar on 22 July 2026 featuring Kim Stachenfeld from Google DeepMind discussing DataDIVER, a method using LLMs to discover interpretable symbolic models of human and animal behavior.

0 favorites 0 likes
#llm-research

Mitigating LLM-based p-Hacking by Preregistering for the Next LLM

arXiv cs.CL ↗ · 2026-06-29 Cached

Proposes a protocol to mitigate p-hacking in LLM-based research by preregistering experiments and running them on the first eligible model released after preregistration, demonstrating substantial mitigation across multiple models.

0 favorites 0 likes
#llm-research

@Zephyr271828: You want a strong small LLM. Would you start small — or inherit from something bigger? New paper: Small LLMs: Pruning v…

X AI KOLs Timeline ↗ · 2026-06-23 Cached

A new paper investigates whether it's better to prune a larger LLM or train a small LLM from scratch, finding that pruning provides more than just a good initialization.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback