LLMs are General Asynchronous Agents
Summary
The paper proposes a general asynchronous LLM framework that lets users define inference coroutines with overlapping memory states, demonstrating that Qwen3.x models can handle streaming video, video games, and monitoring tasks asynchronously without task-specific training.
View Cached Full Text
Cached at: 09/30/26, 12:16 PM
Paper page - LLMs are General Asynchronous Agents
Source: https://huggingface.co/papers/2609.35427
Abstract
ModernLLMsareincreasinglycapableasautonomousagents,buttheyfollowsequentialinteractioncycles:read,think,replyorcalltools,repeat.Manyreal-worldusecasesarenotsequential:voiceassistants,embodiedagents,andmonitoringsystemsreceivenewinputswhiletheythinkorperformanothertask.ModernLLMsaddressthiswithspecializedarchitecturesforvoiceinteractionandvideostreams,VLAsforrobotcontrol,asynchronoustoolcallingforAPIusage,andothers.Inthiswork,wegeneralizefromdifferentasynchronoustaskstogeneralasynchronousagentsthatcanadapttodifferenttypesofconcurrency.Toachievethis,wedevelopanasynchronousLLMframeworkthatletsusers(ortheagentsthemselves)defineinferencecoroutineswithoverlappingmemorystates.WeshowcasethatQwen3.xmodelsarecapableofasynchronousoperationforstreamingvideounderstanding,videogames,andmonitoring,withouttask-specifictraining.
View arXiv pageView PDFProject pageGitHub4Add to collection
Get this paper in your agent:
hf papers read 2609\.35427
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2609.35427 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2609.35427 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2609.35427 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Your LLM shouldn’t be your coding-agent workflow
Argues that LLMs should be used for reasoning within coding-agent workflows, while deterministic infrastructure handles queues, state, retries, and recovery, so the process doesn't break when usage limits hit.
Multi-Stream LLMs: Unblocking Language Models with Parallel Streams of Thoughts, Inputs and Outputs
This paper proposes Multi-Stream LLMs, which transition from sequential message-based instruction tuning to parallel stream processing. This approach allows language models to simultaneously read, think, and generate across multiple concurrent data flows, addressing bottlenecks in autonomous agent applications.
Which API for general real-time LLM agents [D]
A Hacker News discussion exploring what API/interface should be used to build general real-time asynchronous LLM agents that can react to new stimuli (interruptions, live event streams) mid-inference, referencing the AsyncLLM preprint and existing vendor real-time APIs as partial solutions.
Choose what LLMs can and can’t do well
The article highlights that LLMs excel at ambiguous judgment tasks but are mediocre for consistent computation, advocating for task specialization in multi-agent systems.
Multi-Stream LLMs: new paper on parallelizing/separating prompts, thinking, I/O
This paper proposes Multi-Stream LLMs, which use multiple parallel input/output streams to allow models to read and generate simultaneously, unblocking limitations of sequential chat formats.