context-length

Tag

Cards List
#context-length

Qwen3.8-Flash-Next (5.05bpw + ngram at bf16) exl3 on one r9700: 863 t/s prefill and 35 t/s decode at 230k context (256k max), is that ok or am i missing something?

Reddit r/LocalLLaMA ↗ · yesterday

A user shares and asks for feedback on running Qwen3.8-Flash-Next in EXL3 5.05bpw on a single Radeon RX 9700 (gfx12) with mixed GPU/CPU expert offloading, reporting ~863 t/s prefill and ~35 t/s decode at 230k context.

0 favorites 0 likes
#context-length

Staying on Task: Testing the Foundations of Long-Horizon Agent Reliability

arXiv cs.AI ↗ · 3d ago Cached

This NeurIPS 2026 workshop paper introduces Long-Transduction, a controlled diagnostic benchmark measuring how well models sustain stateful, context-dependent operations over long generations. Evaluating seven open-weight models, it finds severe degradation when scaling context length (up to 128K), varying input format, and increasing local task complexity.

0 favorites 0 likes
#context-length

Transformers Stop Thinking Too Early, and a Tiny LoRA Fixes It

Hugging Face Daily Papers ↗ · 5d ago Cached

研究发现预训练 transformer 在遵循上下文引用时极少使用其深度(仅可靠遵循 1.4-3.6 行),而一个冻结所有权重、只在早期层添加的 rank-8 LoRA 即可大幅扩展该能力——Qwen3-8B 在 24 行链上的精确准确率从 15.5% 提升至 99%。LoRA 起到'接力'作用,使程序行通过中间层传递链身份信息。

0 favorites 0 likes
#context-length

Opus 5.5 summary: 66.4% Terminal-Bench, ~40% cheaper to run than Opus 5, cache reads down 60%

Reddit r/artificial ↗ · 2026-09-24

Opus 5.5, launched by Anthropic, achieves 66.4% on Terminal-Bench, reduces costs by ~40% compared to Opus 5, and features a 1M context window with cache reads down 60%.

0 favorites 0 likes
#context-length

Ling 3.0 (on Nous)

Reddit r/AI_Agents ↗ · 2026-09-07

This article discusses user-reported issues with the Ling 3.0 AI model from Nous, focusing on problems with text output and context length limitations, potentially linked to quantization or rate limits.

0 favorites 0 likes
#context-length

Qwen3.8-Flash-Next at 170K context on a single 96 GB card. ~110 tok/s.

Reddit r/LocalLLaMA ↗ · 2026-08-30

The article describes a method to run the Qwen3.8-Flash-Next model with quantized n-grams to achieve over 170K token context on a single 96GB GPU card, with performance up to 110 tokens per second using INT4 quantization and memory-mapped disk access.

0 favorites 0 likes
#context-length

Qwen3.8-Flash-Next NVFP4 Day-3 support for 4xV100

Reddit r/LocalLLaMA ↗ · 2026-08-30

RadixArk/Qwen3.8-Flash-Next-NVFP4 is now supported in SGLang-V100, enabling full context operation on 4xV100 GPUs with performance metrics showing high throughput and context handling up to 256k tokens.

0 favorites 0 likes
#context-length

The Query Knows What to Forget: A Second Erase Direction for Linear Attention

arXiv cs.LG ↗ · 2026-08-17 Cached

Introduces a Query-derived Erase Direction (QED) for linear attention models to improve retrieval and about double the usable context length, as demonstrated on the S-NIAH-1 benchmark.

0 favorites 0 likes
#context-length

AI Software Development – What Does The Data Say?

Lobsters Hottest ↗ · 2026-08-17 Cached

The article synthesizes recent data and studies on LLMs in software development, revealing significant limitations in autonomy, context handling, and real-world outcomes, suggesting AI coding amplifies existing strengths rather than fixing weaknesses.

0 favorites 0 likes
#context-length

Glimmer: 233.4 tps on 5090 with Dflash

Reddit r/LocalLLaMA ↗ · 2026-08-10

Glimmer reportedly hits 233.4 tps on an RTX 5090 using Dflash, with 256k context fitting on 24GB VRAM, sparking excitement among users.

0 favorites 0 likes
#context-length

ARC-AGI 3 is not an honest measure of AGI

Reddit r/singularity ↗ · 2026-07-30

A critique arguing that ARC-AGI 3 unfairly disables an agent's ability to maintain context across actions, making it an dishonest measure of general intelligence. It notes that allowing compaction triples scores while using far fewer tokens, and that real-world agents work that way.

0 favorites 0 likes
#context-length

Qwen 3.6 27B is solid up to 262K context. How high have you guys gone above that using Rope/Yarn scaling?

Reddit r/LocalLLaMA ↗ · 2026-07-16

A user shares their experience pushing Qwen 3.6 27B to 262K context with coherent results, and discusses using Rope/Yarn scaling to go higher, along with kv-cache swapping strategies for RTX 3090 Ti.

0 favorites 0 likes
#context-length

@FinanceYF5: Everyone is hyping up the million-token context, but a Prime Intellect engineer spoke the truth: GPT-5.5 has 80% retrieval accuracy at 256k, but when extended to one million, it drops to 36%. The model doesn't fail to hold it—it fails to reason over it—the so-called context rot. Why more...

X AI KOLs Timeline ↗ · 2026-07-14 Cached

A Prime Intellect engineer pointed out that large language models like GPT-5.5 see retrieval accuracy drop from 80% at 256k tokens to 36% at one million tokens, indicating the 'context rot' problem—the model can accommodate but cannot effectively reason over long contexts, posing a challenge to agent applications.

0 favorites 0 likes
#context-length

@FareaNFts: i built a tool to test if @MiniMax_AI M3 model 1M context claim was actually real spoiler: it is here's what it does: >…

X AI KOLs Timeline ↗ · 2026-07-13 Cached

A developer built a tool using MiniMax AI's M3 model to analyze entire GitHub repositories in a single prompt, producing code health reports and bug detection. It successfully processed react's 780k-token codebase for $0.23.

0 favorites 0 likes
#context-length

Getting close to 100K context on 32GB VRAM with Qwen3.6-27 at Q8

Reddit r/LocalLLaMA ↗ · 2026-07-05

A user shares their attempts and configurations to achieve up to 115K context on a Q8-quantized Qwen3.6-27B model using 32GB VRAM on an RTX 5090, with benchmark results and trade-offs between context length and kv-cache quantization.

0 favorites 0 likes
#context-length

RTX5090, gemma-4-31B-it-Q6_K.gguf. Context: before - 35k, after - 80k!

Reddit r/LocalLLaMA ↗ · 2026-07-04

Running the quantized Gemma-4-31B model on an RTX 5090 increases context length from 35k to 80k, showcasing significant performance improvement.

0 favorites 0 likes
#context-length

Notes on Microsoft's FastContext, and a small SWE-QA experiment with retrieval hints

Reddit r/LocalLLaMA ↗ · 2026-06-30

Notes on Microsoft's FastContext technique and a small experiment on software engineering QA with retrieval hints.

0 favorites 0 likes
#context-length

Realtime voice models compounds on cost (and forgets)- "Flowcat" fixed both (4x cheaper, 7x more context)

Reddit r/AI_Agents ↗ · 2026-06-24

Flowcat addresses the high cost and limited context of realtime voice models, achieving 4x lower cost and 7x more context.

0 favorites 0 likes
#context-length

7900XTX 24GB vram, can finally fit Q6K+MTP with Qwen 3.6 27B at 131k context

Reddit r/LocalLLaMA ↗ · 2026-06-20

A guide on optimizing VRAM usage on an AMD 7900XTX to run a 27B Qwen model with Q6K quantization and 131k context by compiling llama.cpp with OpenBLAS and CUDA_FA_ALL_QUANTS, and using kvcache quantization at q5_0/q4_0.

0 favorites 0 likes
#context-length

Does subquadratic's 12 million context model claim hold any water?

Reddit r/singularity ↗ · 2026-06-17

The video examines whether a claimed 12 million context model from subquadratic research is credible, analyzing its technical underpinnings and potential limitations.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback