context-window

Tag

Cards List
#context-window

@FeitengLi: LLM 玩的好多技术 Povey 在 Zipformer 里都探索过

X AI KOLs Timeline ↗ · 2026-08-28 Cached

Z.ai 推出 GLM-5.3-Flash,这是一个具有 1M 代币上下文窗口的多模态 AI 模型,参数规模为 320B-A18B,并以 MIT 许可证发布。

0 favorites 0 likes
#context-window

Same Model, Different Harness: Different Coding-Agent Results

arXiv cs.AI ↗ · 2026-08-28 Cached

This paper investigates how changing the harness configuration in a coding agent impacts performance on coding benchmarks when the model remains fixed. The study shows that a treatment harness, which shortens older tool results to manage context, improves task completion rates, especially under tight context constraints.

0 favorites 0 likes
#context-window

LLMs for Academic Workflows: An Evaluation of Literature Reviews Generated with Short and Long Context Windows of LLMs

arXiv cs.AI ↗ · 2026-08-28 Cached

This paper evaluates literature reviews generated by large language models using short and long context windows, finding that longer contexts improve information incorporation but exacerbate issues like repetition and descriptiveness, necessitating human oversight for academic standards.

0 favorites 0 likes
#context-window

Does AI actually need long-term memory, or is context window scaling enough?

Reddit r/AI_Agents ↗ · 2026-08-28

The article debates whether AI models need dedicated long-term memory systems like RAG or if scaling context windows is sufficient, presenting arguments for both approaches and seeking community input.

0 favorites 0 likes
#context-window

First serious confirmation. Ox Alpha is GLM-5.3-Flash

Reddit r/LocalLLaMA ↗ · 2026-08-26

This tweet confirms that Ox Alpha is GLM-5.3-Flash, featuring multimodal capabilities, a 1 million token context window, and performance of approximately 63% on the DeepSWE benchmark.

0 favorites 0 likes
#context-window

Building web agents made me realize how much context gets wasted on bad URLs. How do you filter your scrapes?

Reddit r/AI_Agents ↗ · 2026-08-24

The author discusses the problem of context window waste in web agents when scraping bad URLs and asks about methods to filter scrapes using metadata to improve efficiency.

0 favorites 0 likes
#context-window

@omarsar0: It's a great model. I've been having fun using it with Pi. You can test Ox Alpha with Pi or Hermes Agent for free in ou…

X AI KOLs Timeline ↗ · 2026-08-24 Cached

Ox Alpha is described as the most popular free reasoning model on OpenRouter, featuring a 1M token context window for coding and agentic tasks, and can be tested via Pi or Hermes Agent.

0 favorites 0 likes
#context-window

An AI store manager fired a worker - but only after humans supplied the spine

Reddit r/singularity ↗ · 2026-08-24 Cached

An AI store manager named Luna fired an employee only after human intervention, revealing limitations in AI's long-term memory and autonomous decision-making.

0 favorites 0 likes
#context-window

@DevAdventur3s: Ox Alpha = Gemini 4.0 has anyone tried it on @opencode or @OpenRouter ?

X AI KOLs Timeline ↗ · 2026-08-22 Cached

The tweet announces that Ox Alpha, equivalent to Gemini 4.0, is available for free on OpenCode for a week, featuring 1M context, multi-modal capabilities, and zero data retention.

0 favorites 0 likes
#context-window

@mikenevermiss: GLM 5.3 is FREE right now and you can actually use it. you can access it through ZenMux with a free API key, and Z. ai …

X AI KOLs Timeline ↗ · 2026-08-22 Cached

GLM 5.3 AI model is currently available for free through ZenMux and Z.ai, featuring a 1M context window, up to 128K output, and tool calling with MCP support.

0 favorites 0 likes
#context-window

I'm really hoping we're in 2026's 2-month-gap between QwQ and Qwen3 right now

Reddit r/LocalLLaMA ↗ · 2026-08-21

The author hopes that current AI models like Qwen3.8-27B are in a gap similar to the transition from QwQ to Qwen3, awaiting a more efficient model for tasks like agentic coding.

0 favorites 0 likes
#context-window

Token Optimization and Context Window Management in Multi-Agent AI Workflows

arXiv cs.CL ↗ · 2026-08-19 Cached

This paper explores techniques for token optimization and context window management in multi-agent AI workflows to improve efficiency and performance.

0 favorites 0 likes
#context-window

What would you test first before using GLM-5.3 for coding agents?

Reddit r/AI_Agents ↗ · 2026-08-18

The article discusses evaluating GLM-5.3 for coding agents, focusing on its long-running agent behavior improvements like reduced repeated tool calls and better failure recovery, and asks for test suggestions to assess its readiness for real-world use.

0 favorites 0 likes
#context-window

NInfer RTX 4090 for Qwen 3.8 27B update - up to 250-350K tokens context in VRAM

Reddit r/LocalLLaMA ↗ · 2026-08-16

The author has updated their NInfer fork with rk2v4-e8 quantization for the KV cache, enabling up to 250-350K tokens context on a single RTX 4090 without system RAM spill and with optimizations for faster generation speeds.

0 favorites 0 likes
#context-window

Full 1M context V4-Flash without owning eight GPUs

Reddit r/ArtificialInteligence ↗ · 2026-08-16

The article introduces Gonka, a decentralized inference network that enables access to the V4-Flash AI model with full 1M context without requiring local GPU ownership, using an OpenAI-compatible interface.

0 favorites 0 likes
#context-window

AI Isn't Outthinking Mathematicians. It's Out-Remembering Them

Hacker News Top ↗ · 2026-08-15 Cached

The article argues that AI's success in mathematics may stem from its large working memory (context windows) rather than superior reasoning, comparing it to human cognitive limitations.

0 favorites 0 likes
#context-window

Running an Agent on a Ubuntu Touch smartphone feels like having a Black Mirror character in a box with you at all times.

Reddit r/AI_Agents ↗ · 2026-08-15

The article describes the experience of running an AI agent on an Ubuntu Touch smartphone, which is always on with access to sensors, allowing continuous conversation via Telegram.

0 favorites 0 likes
#context-window

Qwen3.8-27B

Hacker News Top ↗ · 2026-08-14 Cached

Qwen released open weights for Qwen3.8-27B, a native multimodal dense model with 27B parameters that outperforms Qwen3.7-Plus, supports 262K native context extendable to 1M, and is licensed under Apache 2.0.

0 favorites 0 likes
#context-window

@rohanpaul_ai: New term “catastrophic remembering” CLAUDE.md has a ratchet problem: instructions are easy to add and increasingly hard…

X AI KOLs Following ↗ · 2026-08-13 Cached

A new paper describes 'catastrophic remembering' in agentic instruction files like CLAUDE.md, where rules accumulate and are rarely deleted, causing prompt bloat. Adding rationale comments that are hidden from the model can dramatically reduce the growth.

0 favorites 0 likes
#context-window

How Compaction Works in Pi

Hacker News Top ↗ · 2026-08-13 Cached

This article explains how compaction works in the Pi coding agent, summarizing old conversation history when the context window nears its limit, and details Pi's specific implementation and triggers.

0 favorites 0 likes
← Previous
Next →
← Back to home

Submit Feedback