LLMs are General Asynchronous Agents

George Yakushev*, Denis Mazur*, Vladimir Bartenev, Vyacheslav Zhdanovskiy, Timofey Byzov, Vladimir Kaurkin, Vadim Pastushenko *Indicates Equal Contribution
Asynchronous agentsShared-memory inferenceTraining-free

Abstract

Modern LLMs are increasingly capable as autonomous agents, but they follow sequential interaction cycles: read, think, reply or call tools, repeat. Many real-world use cases are not sequential: voice assistants, embodied agents, and monitoring systems receive new inputs while they think or perform another task. Modern LLMs address this with specialized architectures for voice interaction and video streams, VLAs for robot control, asynchronous tool calling for API usage, and others. In this work, we generalize from different asynchronous tasks to general asynchronous agents that can adapt to different types of concurrency. To achieve this, we develop an asynchronous LLM framework that lets users (or the agents themselves) define inference coroutines with overlapping memory states. We showcase that Qwen 3.x models are capable of asynchronous operation for streaming video understanding, videogames, and monitoring, without task-specific training.

A training-free Doom agent. Observation and action coroutines share one cache and progress concurrently, so the agent keeps acting while it watches the game.

How AsyncLLM Works

AsyncLLM lets you define asynchronous agents in terms of coroutines communicating through shared cache blocks. Each coroutine writes to its own cache block and can access other coroutines’ outputs using cache views. Below, an agent reasons and simultaneously provides the user with a running summary of its progress. Colors mark cache blocks; the outlined block is where a coroutine writes.

Agent definition
In [1]:
api = AsyncLLM("Qwen/Qwen3.5-9B")
prompt_block, thinker_block, writer_block = [
    await api.create_block() for _ in range(3)]
paragraph_finished = asyncio.Event()  # thinker -> writer signal
thinker_done = asyncio.Event()
await api.forward(encode(PROMPT), write_to=prompt_block)
In [2]:
async def thinker_coro():
    async for token in api.generate(
        "<think>\n", [prompt_block, thinker_block],
        stop="</think>",
    ):
        if "\n\n" in token:
            paragraph_finished.set()  # notify the writer
    thinker_done.set()
    paragraph_finished.set()  # one last summary
In [3]:
async def writer_coro():
    while (not thinker_done.is_set()
           or paragraph_finished.is_set()):
        await paragraph_finished.wait()  # new paragraph
        paragraph_finished.clear()
        async for token in api.generate(
            "...</think> Summary:",
            [prompt_block, thinker_block, writer_block],
            stop="\n",
        ):
            send_to_user(token)  # e.g. speak via TTS
In [*]:
await asyncio.gather(thinker_coro(), writer_coro())
What happens
Inference engine
T decode W decode W prefill

One column per forward pass. When both coroutines are active they share a batch. Click a column to jump there.

Cache blocks
Sprompt_block Tthinker_block Wwriter_block
Thinker view In [2]
S
T

Reads the prompt and its own reasoning. Never sees the writer.

Writer view In [3]
S
T
W

Reads the thinker's block as it grows and writes the summary.

Demos

All of the demos below are built on top of AsyncLLM. Each agent is a set of coroutines communicating through shared cache blocks.

Streaming video. An event probe scores every incoming frame (bars along the bottom), and the agent describes new events as they happen while the video keeps playing.

Qwen3.5-35B-A3B finds how much RAM the listing with blue LED lights has.

The same model finds the email of the seller of the guitar in the red case.

Web agent on VisualWebArena Classifieds.

Self-Defining Agents

Because AsyncLLM inference is defined directly in Python, we can hand an agent its own runtime and let it define its own inference structure: which coroutines to run, which cache blocks they share, and how they react to the environment.

  1. 1WriteThe agent generates tool calls that rewrite its own inference code.
  2. 2PatchThe new code is applied to the already-running instance. Cache blocks and background tasks survive every rewrite.
  3. 3ActIts act() method runs on every environment step.
  4. 4EvaluateA harness scores the agent and reports the result in the next round's prompt.
next round

We run this loop with Qwen3.6-35B-A3B on two ViZDoom environments. The agent gets its runtime, the previous evaluation result, and the ability to modify its inference code, but no task-specific guidance on what inference structure to build. Self-defined agents show good initial results on our evaluations. The capability is not yet reliable, but it suggests that future LLMs may build self-adapting agents from an environment description.

More in our blog post, The KV cache as an agent runtime.

Results

SoccerNet (streaming)ProactiveVideoQA
ModelAUROCTrigger AccTimValPAUC (ω=0.5)
AsyncLLM Qwen3.5-9B0.67762.8239.620.541
AsyncLLM Qwen3.8-27B0.60854.8833.270.504
AsyncLLM Qwen3.6-35B-A3B0.65162.5537.460.539
Mage-VL0.55552.7927.870.428

Streaming video understanding. Training-free AsyncLLM agents outperform Mage-VL, a model trained specifically for video streaming (higher is better).

Qwen3.6-35B-A3B agentAccuracy % ↑# forward passes ↓Mean % events ↓
Sequential61.768452100
+ early answer47.06664158.71
+ skip rows26.47368665.86
AsyncLLM55.88383755.49

System monitoring on DevOps-Gym. AsyncLLM answers after 55% of the log events on average, using 45% of the sequential agent's forward passes.

BibTeX

@misc{yakushev2026llmsgeneralasynchronousagents,
      title={LLMs are General Asynchronous Agents},
      author={George Yakushev and Denis Mazur and Vladimir Bartenev and Vyacheslav Zhdanovskiy and Timofey Byzov and Vladimir Kaurkin and Vadim Pastushenko},
      year={2026},
      eprint={2609.35427},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2609.35427},
}