400个LLM智能体在一个MMO服务器中共存:关于感知延迟、即发即忘操作和负载削减的教训
摘要
一个在私人MMO服务器上运行400多个LLM智能体的项目,探索诸如陈旧感知、未确认操作和负载削减等挑战,使用本地部署的Qwen模型。
相似文章
在Claude Code中使用本地LLM作为代理
本文介绍了一种自定义的MCP设置,允许在同一会话中使用如llama.cpp等工具,将编码任务从Anthropic的Claude模型卸载到本地的Qwen3.8-27B模型。
帮助优化 llama.cpp + Qwen 27B 在 RTX PRO 6000 Blackwell 上用于编码代理的配置
用户详细介绍了他们在 RTX PRO 6000 Blackwell 上使用 llama.cpp 运行 Qwen 27B 进行本地编码代理的设置,与 Claude 模型进行了性能对比,并请求帮助解决频繁崩溃和响应格式错误的问题。
Poor Man's Agentic Modeling: Simulating Large LLM-Agent Societies on a Laptop
This paper proposes a method for simulating large LLM-agent societies on a laptop by fitting low-parameter surrogate models from a few hundred queries, using a statistical-physics-based taxonomy to predict when this approximation holds. The approach is validated on EconAgent and several other simulations using DeepSeek-elicited agent behaviors.
两块RTX 4090实际能同时运行多少个代理?三周的llama.cpp并发数据——软上限5 @ 64k,硬上限9,以及原因。
关于在两块RTX 4090 GPU上使用llama.cpp并发运行多个AI代理的基准测试报告,揭示了性能限制和Qwen模型的最佳配置。
8GB 显存跑 Qwen3.6 35B MoE 的 llama-server 配置 + 我踩的 max_tokens / thinking 陷阱
作者分享了一套在 8GB RTX 4060 上跑 35B-MoE Qwen3.6 的可用 llama-server 配置,重点提示因内部推理无限制而耗尽 max_tokens 的陷阱,并给出用 per-request thinking_budget_tokens 的解决方案。