Tag
Opus 5.5, launched by Anthropic, achieves 66.4% on Terminal-Bench, reduces costs by ~40% compared to Opus 5, and features a 1M context window with cache reads down 60%.
This article discusses user-reported issues with the Ling 3.0 AI model from Nous, focusing on problems with text output and context length limitations, potentially linked to quantization or rate limits.
The article describes a method to run the Qwen3.8-Flash-Next model with quantized n-grams to achieve over 170K token context on a single 96GB GPU card, with performance up to 110 tokens per second using INT4 quantization and memory-mapped disk access.
RadixArk/Qwen3.8-Flash-Next-NVFP4 is now supported in SGLang-V100, enabling full context operation on 4xV100 GPUs with performance metrics showing high throughput and context handling up to 256k tokens.
Introduces a Query-derived Erase Direction (QED) for linear attention models to improve retrieval and about double the usable context length, as demonstrated on the S-NIAH-1 benchmark.
The article synthesizes recent data and studies on LLMs in software development, revealing significant limitations in autonomy, context handling, and real-world outcomes, suggesting AI coding amplifies existing strengths rather than fixing weaknesses.
Glimmer reportedly hits 233.4 tps on an RTX 5090 using Dflash, with 256k context fitting on 24GB VRAM, sparking excitement among users.
A critique arguing that ARC-AGI 3 unfairly disables an agent's ability to maintain context across actions, making it an dishonest measure of general intelligence. It notes that allowing compaction triples scores while using far fewer tokens, and that real-world agents work that way.
A user shares their experience pushing Qwen 3.6 27B to 262K context with coherent results, and discusses using Rope/Yarn scaling to go higher, along with kv-cache swapping strategies for RTX 3090 Ti.
A Prime Intellect engineer pointed out that large language models like GPT-5.5 see retrieval accuracy drop from 80% at 256k tokens to 36% at one million tokens, indicating the 'context rot' problem—the model can accommodate but cannot effectively reason over long contexts, posing a challenge to agent applications.
A developer built a tool using MiniMax AI's M3 model to analyze entire GitHub repositories in a single prompt, producing code health reports and bug detection. It successfully processed react's 780k-token codebase for $0.23.
A user shares their attempts and configurations to achieve up to 115K context on a Q8-quantized Qwen3.6-27B model using 32GB VRAM on an RTX 5090, with benchmark results and trade-offs between context length and kv-cache quantization.
Running the quantized Gemma-4-31B model on an RTX 5090 increases context length from 35k to 80k, showcasing significant performance improvement.
Notes on Microsoft's FastContext technique and a small experiment on software engineering QA with retrieval hints.
Flowcat addresses the high cost and limited context of realtime voice models, achieving 4x lower cost and 7x more context.
A guide on optimizing VRAM usage on an AMD 7900XTX to run a 27B Qwen model with Q6K quantization and 131k context by compiling llama.cpp with OpenBLAS and CUDA_FA_ALL_QUANTS, and using kvcache quantization at q5_0/q4_0.
The video examines whether a claimed 12 million context model from subquadratic research is credible, analyzing its technical underpinnings and potential limitations.
A user asks how llama.cpp can serve multiple users each with full context length, noting that it seems to only share the context pool rather than providing dedicated context per user.
The author expresses surprise at how effective key-value cache quantization (q4_0) remains even with large context windows, citing accurate retrieval from a 100k context.
Discusses strategies and techniques for maximizing token usage in large language models to improve efficiency and output quality.