long-context

Tag

Cards List
#long-context

“What are we still missing in AI that is preventing us from fully automating digital work?”

Reddit r/singularity · 17h ago

The content questions what key limitations in AI, such as long memory and context windows, are preventing the full automation of digital jobs despite advancements in frontier models.

0 favorites 0 likes
#long-context

Context Poisoning as Extreme-Value Attention Interference in Long-Context Language Models

arXiv cs.CL · yesterday Cached

This paper introduces 'context poisoning' as a phenomenon where attention interference degrades performance in long-context language models, deriving bounds showing evidence margin must scale with log of distractor count and proposing mitigations like evidence bottlenecks.

0 favorites 0 likes
#long-context

XiaomiMiMo/MiMo-V2.6-Flash-RL · Hugging Face

Reddit r/LocalLLaMA · yesterday Cached

MiMo-V2.6-Flash-RL is a multimodal AI model that scales reinforcement learning for self-improvement, featuring a sparse mixture-of-experts architecture with 309B total parameters and 1M token context length.

0 favorites 0 likes
#long-context

@omarsar0: StepFun’s new Step 5 Preview model is impressive! Had a chance to test it early. I've been testing it as a coding agent…

X AI KOLs Following · yesterday Cached

StepFun's new Step 5 Preview model is tested as a coding agent, demonstrating competitive performance with models like GLM 5.3 and excelling in long-horizon tasks due to its effective stopping behavior.

0 favorites 0 likes
#long-context

TierKV: Long-Context On-Device LLMs via Predictive Multi-Tier KV Caching

arXiv cs.LG · 2d ago Cached

TierKV proposes a predictive multi-tier KV caching framework to optimize memory usage and throughput for long-context LLMs on mobile devices, achieving significant performance improvements with minimal accuracy degradation.

0 favorites 0 likes
#long-context

Elastic Threshold Attention: Learned Contextual Sparsity for Long-Context Decoding

arXiv cs.LG · 2d ago Cached

Elastic Threshold Attention (ETA) is a trainable sparse attention architecture that improves long-context decoding speed without quality degradation by using dynamic thresholds predicted from query representations.

0 favorites 0 likes
#long-context

RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models

arXiv cs.AI · 2d ago Cached

RBS-Attention introduces a training-free sparse-prefill method with dual-branch selection to mitigate mean dilution in long-context LLM inference, achieving up to 20.65× speedup on H100 GPUs while maintaining near-dense quality on benchmarks.

0 favorites 0 likes
#long-context

Boosting Deepresearch and LongContext Ability with Self-Generated Deepresearch Rollouts Traces

arXiv cs.CL · 2d ago Cached

The paper proposes a method called DLD-RL to enhance deep research and long-context abilities in AI agents by repurposing reinforcement learning trajectories into long-context QA data, showing significant performance improvements on benchmarks.

0 favorites 0 likes
#long-context

The bear can dance: Qwen 3.8 27B on one 3090 for 3 weeks

Reddit r/LocalLLaMA · 2d ago

An experiment demonstrated that a quantized Qwen 3.8 27B model, running locally on a single RTX 3090 GPU, autonomously pursued optimizing CUDA inference for over three weeks, producing functional kernels and benchmarks while maintaining coherent long-term goal-following.

0 favorites 0 likes
#long-context

Qwen3.8-Flash-Next at 1M context on Strix Halo: 38 tok/s decode, 18 min prefill (halogen 0.12.0)

Reddit r/LocalLLaMA · 3d ago

Halogen version 0.12.0 fixes performance degradation at high context depths, showing improved decode and prefill speeds for Qwen3.8-Flash-Next at 1 million tokens of context on AMD Ryzen AI Max+ hardware.

0 favorites 0 likes
#long-context

@omarsar0: Storytelling is finally solved. Neither ChatGPT nor Claude could solve it because it required long-horizon consistency,…

X AI KOLs Timeline · 4d ago Cached

Sherpa is an advanced fiction writing AI trained on 5.5 billion minutes of data, designed to solve long-horizon consistency in storytelling, outperforming existing models like ChatGPT and Claude.

0 favorites 0 likes
#long-context

@levie: Agents already make up the majority of inference. This will quickly trend toward nearly all inference over the next yea…

X AI KOLs Following · 5d ago Cached

The tweet asserts that AI agents now dominate inference traffic, with this trend expected to accelerate, leading to agents performing numerous tasks 24/7 and consuming vast amounts of tokens.

0 favorites 0 likes
#long-context

Block Parallelism For Efficient Distributed Long-Context Diffusion Language Model Training

arXiv cs.LG · 5d ago Cached

This paper introduces block parallelism and context-sharded block parallelism (CSBP) to efficiently train long-context diffusion language models, achieving significant throughput improvements and better performance on benchmarks like SWE-bench Verified.

0 favorites 0 likes
#long-context

XingChen-AGI/Xing4.0-29B-A4B MoE

Reddit r/LocalLLaMA · 6d ago

Xing4.0-29B-A4B is a next-generation MoE large language model developed by China Telecom, featuring 29B total parameters with 4B active per token, native support for 256K context length, and optimization for Ascend NPU with agent-oriented architecture for complex engineering tasks.

0 favorites 0 likes
#long-context

ASPIRE: Asynchronous Batched Self-Speculative Decoding for Long-Context LLM Inference

arXiv cs.LG · 6d ago Cached

ASPIRE is an asynchronous batched self-speculative decoding framework that enhances long-context LLM inference by enabling independent request scheduling and reducing attention staleness, achieving 1.70-4.58× speedup over baselines.

0 favorites 0 likes
#long-context

Long-Context Demonstration Selection Using State Space Models

arXiv cs.LG · 6d ago Cached

This paper proposes using state space models to efficiently select demonstrations for long-context language model prompts, reducing computational cost and improving performance.

0 favorites 0 likes
#long-context

DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression

Hacker News Top · 6d ago Cached

This article analyzes the DeepSeek-V4.1 Flash model, detailing its technical report on KV cache compression and architectural optimizations that enable efficient long-context processing and high-speed inference.

0 favorites 0 likes
#long-context

DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression

Hugging Face Daily Papers · 6d ago Cached

Introduces DeepSeek-V4.1-Flash, a multimodal Mixture-of-Experts model with 552B parameters, featuring advanced KV cache compression techniques to reduce deployment costs and improve efficiency for long-context agent workloads.

0 favorites 0 likes
#long-context

You can offload most of Qwen3.8-Flash-Next's KV cache to RAM with little decode slowdown

Reddit r/LocalLLaMA · 6d ago

A technique to offload the KV cache of Qwen3.8-Flash-Next to system RAM is demonstrated, allowing long-context inference with minimal decode slowdown by leveraging the model's efficient architecture.

0 favorites 0 likes
#long-context

Flattening Every Memory Peak in Long-Context Mixture-of-Experts Training

Hugging Face Daily Papers · 2026-09-13 Cached

This paper introduces techniques to manage memory peaks in training large Mixture-of-Experts models with long context lengths, including Pipelined LLEP, Ring-DTP, SCO, and OffloadStreamAdamW, which enable fixed GPU working sets and improve throughput up to 10.4x.

0 favorites 0 likes
Next →
← Back to home

Submit Feedback