context-length

Tag

Cards List
#context-length

Opus 5.5 summary: 66.4% Terminal-Bench, ~40% cheaper to run than Opus 5, cache reads down 60%

Reddit r/artificial ↗ · 5d ago

Opus 5.5, launched by Anthropic, achieves 66.4% on Terminal-Bench, reduces costs by ~40% compared to Opus 5, and features a 1M context window with cache reads down 60%.

0 favorites 0 likes
#context-length

Ling 3.0 (on Nous)

Reddit r/AI_Agents ↗ · 2026-09-07

This article discusses user-reported issues with the Ling 3.0 AI model from Nous, focusing on problems with text output and context length limitations, potentially linked to quantization or rate limits.

0 favorites 0 likes
#context-length

Qwen3.8-Flash-Next at 170K context on a single 96 GB card. ~110 tok/s.

Reddit r/LocalLLaMA ↗ · 2026-08-30

The article describes a method to run the Qwen3.8-Flash-Next model with quantized n-grams to achieve over 170K token context on a single 96GB GPU card, with performance up to 110 tokens per second using INT4 quantization and memory-mapped disk access.

0 favorites 0 likes
#context-length

Qwen3.8-Flash-Next NVFP4 Day-3 support for 4xV100

Reddit r/LocalLLaMA ↗ · 2026-08-30

RadixArk/Qwen3.8-Flash-Next-NVFP4 is now supported in SGLang-V100, enabling full context operation on 4xV100 GPUs with performance metrics showing high throughput and context handling up to 256k tokens.

0 favorites 0 likes
#context-length

The Query Knows What to Forget: A Second Erase Direction for Linear Attention

arXiv cs.LG ↗ · 2026-08-17 Cached

Introduces a Query-derived Erase Direction (QED) for linear attention models to improve retrieval and about double the usable context length, as demonstrated on the S-NIAH-1 benchmark.

0 favorites 0 likes
#context-length

AI Software Development – What Does The Data Say?

Lobsters Hottest ↗ · 2026-08-17 Cached

The article synthesizes recent data and studies on LLMs in software development, revealing significant limitations in autonomy, context handling, and real-world outcomes, suggesting AI coding amplifies existing strengths rather than fixing weaknesses.

0 favorites 0 likes
#context-length

Glimmer: 233.4 tps on 5090 with Dflash

Reddit r/LocalLLaMA ↗ · 2026-08-10

Glimmer reportedly hits 233.4 tps on an RTX 5090 using Dflash, with 256k context fitting on 24GB VRAM, sparking excitement among users.

0 favorites 0 likes
#context-length

ARC-AGI 3 is not an honest measure of AGI

Reddit r/singularity ↗ · 2026-07-30

A critique arguing that ARC-AGI 3 unfairly disables an agent's ability to maintain context across actions, making it an dishonest measure of general intelligence. It notes that allowing compaction triples scores while using far fewer tokens, and that real-world agents work that way.

0 favorites 0 likes
#context-length

Qwen 3.6 27B is solid up to 262K context. How high have you guys gone above that using Rope/Yarn scaling?

Reddit r/LocalLLaMA ↗ · 2026-07-16

A user shares their experience pushing Qwen 3.6 27B to 262K context with coherent results, and discusses using Rope/Yarn scaling to go higher, along with kv-cache swapping strategies for RTX 3090 Ti.

0 favorites 0 likes
#context-length

@FinanceYF5: Everyone is hyping up the million-token context, but a Prime Intellect engineer spoke the truth: GPT-5.5 has 80% retrieval accuracy at 256k, but when extended to one million, it drops to 36%. The model doesn't fail to hold it—it fails to reason over it—the so-called context rot. Why more...

X AI KOLs Timeline ↗ · 2026-07-14 Cached

A Prime Intellect engineer pointed out that large language models like GPT-5.5 see retrieval accuracy drop from 80% at 256k tokens to 36% at one million tokens, indicating the 'context rot' problem—the model can accommodate but cannot effectively reason over long contexts, posing a challenge to agent applications.

0 favorites 0 likes
#context-length

@FareaNFts: i built a tool to test if @MiniMax_AI M3 model 1M context claim was actually real spoiler: it is here's what it does: >…

X AI KOLs Timeline ↗ · 2026-07-13 Cached

A developer built a tool using MiniMax AI's M3 model to analyze entire GitHub repositories in a single prompt, producing code health reports and bug detection. It successfully processed react's 780k-token codebase for $0.23.

0 favorites 0 likes
#context-length

Getting close to 100K context on 32GB VRAM with Qwen3.6-27 at Q8

Reddit r/LocalLLaMA ↗ · 2026-07-05

A user shares their attempts and configurations to achieve up to 115K context on a Q8-quantized Qwen3.6-27B model using 32GB VRAM on an RTX 5090, with benchmark results and trade-offs between context length and kv-cache quantization.

0 favorites 0 likes
#context-length

RTX5090, gemma-4-31B-it-Q6_K.gguf. Context: before - 35k, after - 80k!

Reddit r/LocalLLaMA ↗ · 2026-07-04

Running the quantized Gemma-4-31B model on an RTX 5090 increases context length from 35k to 80k, showcasing significant performance improvement.

0 favorites 0 likes
#context-length

Notes on Microsoft's FastContext, and a small SWE-QA experiment with retrieval hints

Reddit r/LocalLLaMA ↗ · 2026-06-30

Notes on Microsoft's FastContext technique and a small experiment on software engineering QA with retrieval hints.

0 favorites 0 likes
#context-length

Realtime voice models compounds on cost (and forgets)- "Flowcat" fixed both (4x cheaper, 7x more context)

Reddit r/AI_Agents ↗ · 2026-06-24

Flowcat addresses the high cost and limited context of realtime voice models, achieving 4x lower cost and 7x more context.

0 favorites 0 likes
#context-length

7900XTX 24GB vram, can finally fit Q6K+MTP with Qwen 3.6 27B at 131k context

Reddit r/LocalLLaMA ↗ · 2026-06-20

A guide on optimizing VRAM usage on an AMD 7900XTX to run a 27B Qwen model with Q6K quantization and 131k context by compiling llama.cpp with OpenBLAS and CUDA_FA_ALL_QUANTS, and using kvcache quantization at q5_0/q4_0.

0 favorites 0 likes
#context-length

Does subquadratic's 12 million context model claim hold any water?

Reddit r/singularity ↗ · 2026-06-17

The video examines whether a claimed 12 million context model from subquadratic research is credible, analyzing its technical underpinnings and potential limitations.

0 favorites 0 likes
#context-length

Maybe dumb question, but how do you serve multiple users with the full context length?

Reddit r/LocalLLaMA ↗ · 2026-06-15

A user asks how llama.cpp can serve multiple users each with full context length, noting that it seems to only share the context pool rather than providing dedicated context per user.

0 favorites 0 likes
#context-length

I'm still surprised on how good the kv quantization has become

Reddit r/LocalLLaMA ↗ · 2026-06-15

The author expresses surprise at how effective key-value cache quantization (q4_0) remains even with large context windows, citing accurate retrieval from a 100k context.

0 favorites 0 likes
#context-length

Token maxxing

Reddit r/singularity ↗ · 2026-06-06

Discusses strategies and techniques for maximizing token usage in large language models to improve efficiency and output quality.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback