llm-research

Tag

Cards List
#llm-research

@DanKornas: Keeping up with LLM systems research is messy when papers, reports, frameworks, and course links are scattered everywhe…

X AI KOLs Timeline ↗ · 2026-06-09 Cached

LLMSys-PaperList is a curated reading list on GitHub that organizes LLM systems research papers and resources into practical categories such as training systems, serving systems, and multi-modal coverage, helping AI/ML engineers and researchers stay updated.

0 favorites 0 likes
#llm-research

@MOSS_workshop: CfP for the 2nd version of MOSS at @COLM_conf! https://sites.google.com/view/moss-colm-2026/… (Deadline: 6/30) We welco…

X AI KOLs Following ↗ · 2026-06-03 Cached

Call for papers for the 2nd MOSS workshop at COLM 2026, focusing on small-scale research for algorithmic innovation and scientific understanding in LLMs, with deadline June 30.

0 favorites 0 likes
#llm-research

LLMs are not the black box you were promised

Hacker News Top ↗ · 2026-06-02 Cached

An article summarizing Anthropic's 2025 paper on mechanistic interpretability, showing that LLMs are not black boxes and that circuit tracing can reveal multi-step reasoning and human-identifiable concepts.

0 favorites 0 likes
#llm-research

@wlzh: Microsoft + UPenn Open-Source Multiplex Thinking: Let LLMs 'Clone' at Forks Then Merge. In a nutshell: When reasoning reaches a critical decision point, the model 'clones' into K pathfinders, each taking a different path. After one step, they merge back into a composite token and continue. With K=3, one token carries the information of three…

X AI KOLs Timeline ↗ · 2026-06-02 Cached

Microsoft and the University of Pennsylvania open-source Multiplex Thinking, which allows LLMs to split into K parallel paths during inference, explore, then merge, improving efficiency. On a 7B model, it achieves over 50% accuracy on AMC2023 (first 7B model to do so) and over 55% on AIME2025.

0 favorites 0 likes
#llm-research

Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents

arXiv cs.AI ↗ · 2026-06-01 Cached

This paper analyzes two capabilities in self-evolving LLM agents: harness-updating and harness-benefit. It finds that harness-updating is flat across base capability levels, while harness-benefit is non-monotonic, with mid-tier models benefiting most.

0 favorites 0 likes
#llm-research

Help interpreting metrics: a strong target text appears to induce a measurable latent-state shift in Gemma 3 12B IT

Reddit r/AI_Agents ↗ · 2026-05-29

A researcher presents evidence that strong target text can induce a measurable latent-state shift in Gemma 3 12B IT before final output, distinct from lexical or content overlaps, and discusses implications for AI safety beyond output-only evaluation.

0 favorites 0 likes
#llm-research

@cjziems: We're going live in 30 minutes, and we'd love to have you join Joined by @dorazhao9 and @Diyi_Yang, I'll be talking abo…

X AI KOLs Timeline ↗ · 2026-05-29 Cached

The article introduces the live discussion of the Augmented Mind podcast about the paper 'Reflections and New Directions for Human-Centered Large Language Models', emphasizing that AI development should shift from capability benchmarks to human flourishing and long-term well-being.

0 favorites 0 likes
#llm-research

@vintcessun: Actually, large language models' context windows are getting larger and larger, but costs are also skyrocketing. This paper simply treats context management as a deployment optimization problem and develops a unified framework called Efficiency Frontier. Simply put, they no longer look at performance or cost separately, but jointly model task performance, token overhead, and preprocessing reuse...

X AI KOLs Timeline ↗ · 2026-05-26 Cached

This paper proposes a unified framework called Efficiency Frontier, which treats large model context management as a deployment optimization problem, jointly modeling task performance, token overhead, and preprocessing reuse. On 5,000 HotpotQA instances, deployment optimization saves 25% of token usage, while memory compression is more than half the cost of full context in high-precision scenarios.

0 favorites 0 likes
#llm-research

@ihtesham2005: If you still think AI agents can't do real research, this paper will end that argument. Researchers from Google and Met…

X AI KOLs Following ↗ · 2026-05-13 Cached

Researchers from Google and Meta propose AutoTTS, a framework using AI agents to automatically discover and refine test-time scaling strategies for LLMs without human intervention. The agent successfully identified complex, coordinated reasoning mechanisms that outperformed manual baselines at a low computational cost.

0 favorites 0 likes
#llm-research

Three Regimes of Context-Parametric Conflict: A Predictive Framework and Empirical Validation

arXiv cs.CL ↗ · 2026-05-13 Cached

This paper proposes a three-regime framework to resolve empirical contradictions in how LLMs handle conflict between training knowledge and new documents, validated across five major models. It distinguishes between parametric strength and uniqueness and demonstrates how task framing and evidence coherence significantly impact model behavior.

0 favorites 0 likes
#llm-research

AutoLLMResearch: Training Research Agents for Automating LLM Experiment Configuration -- Learning from Cheap, Optimizing Expensive

Hugging Face Daily Papers ↗ · 2026-05-12 Cached

This paper introduces AutoLLMResearch, an agentic framework that automates the configuration of expensive LLM experiments by learning from low-fidelity environments and extrapolating to high-cost settings. It aims to reduce computational waste and reliance on expert intuition in scalable LLM research.

0 favorites 0 likes
#llm-research

@WilliamBarrHeld: To train better open models, we need predictable scaling. Delphi is Marin’s first step: we pretrained many small models…

X AI KOLs Following ↗ · 2026-05-11 Cached

Marin AI researchers, led by William Barr Held, introduce Delphi, a methodology that pretrains small models to accurately predict the training outcomes of larger 25B-parameter runs. This research aims to establish predictable scaling for more efficient open-source AI model development.

0 favorites 0 likes
#llm-research

Google's SkillOS for Self-Evolving AI Agents (22 minute read)

TLDR AI ↗ · 2026-05-11 Cached

Google Cloud AI Research introduces SkillOS, a reinforcement learning framework enabling LLM-based agents to self-evolve by curating reusable skills from past experiences.

0 favorites 0 likes
#llm-research

Shaping Schema via Language Representation as the Next Frontier for LLM Intelligence Expanding

Hugging Face Daily Papers ↗ · 2026-05-10 Cached

This paper argues that designing advanced language representations to shape cognitive schemas is a key frontier for expanding LLM intelligence without scaling parameters. It provides formalizations and empirical evidence showing that different linguistic structures significantly impact model performance and internal feature activations.

0 favorites 0 likes
#llm-research

Three key areas Anthropic is working on for their next models

Reddit r/singularity ↗ · 2026-05-06

Dianne Penn outlines three key focus areas for future Claude models: enhanced judgment and code quality, effectively infinite context windows with memory, and multi-agent coordination capabilities.

0 favorites 0 likes
← Previous
← Back to home

Submit Feedback