latency-reduction

Tag

Cards List
#latency-reduction

Relativity Networks raises $22 million to bring a faster kind of fiber to data centers

TechCrunch AI · 4d ago Cached

Relativity Networks has raised $22 million to develop hollow-core fiber technology, which reduces latency by 30% and could transform data center geography for AI compute.

0 favorites 0 likes
#latency-reduction

We benchmarked MCP vs filesystem access across 20 production-agent scenarios. The filesystem setup cut LLM costs by 27% and latency by 32%

Reddit r/AI_Agents · 5d ago

A benchmark study comparing MCP and filesystem access for AI agents in 20 production scenarios found that filesystem access reduces LLM costs by 27% and latency by 32% while improving answer quality.

0 favorites 0 likes
#latency-reduction

Elbow-Based MoE Routing: A Training-Free Inference Time Plugin for Expert Selection

arXiv cs.LG · 2026-08-06 Cached

This paper introduces elbow-based routing, a training-free inference-time method for MoE models that dynamically adjusts the number of active experts per token by detecting the elbow point in router probability distributions, achieving a 5.3% average latency reduction while maintaining accuracy.

0 favorites 0 likes
#latency-reduction

@alibaba_cloud: One PolarDB, one MemOS—persistent, never-blackout memory for AI. PolarDB & MemTensor launch a one-stop AI memory soluti…

X AI KOLs Timeline · 2026-07-30 Cached

Alibaba Cloud launches a one-stop AI memory solution combining PolarDB, MemOS, and MemTensor, offering relational, vector, and graph retrieval in a single PolarDB-PG instance, reducing P99 latency by up to 89.2%.

0 favorites 0 likes
#latency-reduction

Workload-Aware Caching for Multi-Agent Systems

arXiv cs.AI · 2026-07-24 Cached

This paper presents a workload-aware cache eviction policy for multi-agent systems that uses recomputation cost, DAG dependency count, and agent invocation frequency to retain valuable cached entries, reducing latency by up to 64.7% over uncached baselines and 31.1% over the next best finite-capacity method.

0 favorites 0 likes
#latency-reduction

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications

Hugging Face Daily Papers · 2026-07-20 Cached

FlashRT is an agent harness that guides coding agents to automatically optimize and deploy real-time multimodal applications, achieving up to 70x latency reduction on NVIDIA B200 GPUs and 3.6x throughput improvement on AMD MI355X.

0 favorites 0 likes
#latency-reduction

Director: Accelerating Distributed MoE Serving via Online Proactive Expert Placement

arXiv cs.LG · 2026-07-13 Cached

This paper introduces Director, a distributed MoE serving system that minimizes end-to-end latency using prediction-driven, online proactive expert placement. It employs a lightweight predictor and a relaxation-based optimizer to achieve up to 55% latency reduction for models like Mistral, DeepSeek, and Qwen.

0 favorites 0 likes
#latency-reduction

@Alacritic_Super: If you are building production LLM applications, learn LLM Caching. Caching can reduce latency, GPU utilization, and AP…

X AI KOLs Timeline · 2026-07-12 Cached

This article emphasizes the importance of LLM caching in production systems to reduce latency, GPU utilization, and costs, and introduces LMCache, an open-source KV cache management layer for scalable LLM inference.

0 favorites 0 likes
#latency-reduction

@PyTorch: Normalization layers often introduce memory-bound bottlenecks in large language models and recommendation systems due t…

X AI KOLs Following · 2026-07-10 Cached

Meta introduces techniques like Lazy Pre-Norm, Multi-CTA Norm Fusion, and FlashNormAttention to fuse normalization operations with GEMM and Attention kernels, hiding up to 90% of normalization latency on NVIDIA B200 hardware and achieving up to 35% latency reduction in attention blocks.

0 favorites 0 likes
#latency-reduction

ProactiveLLM: Learning Active Interaction for Streaming Large Language Models

arXiv cs.CL · 2026-06-02 Cached

ProactiveLLM introduces a method for streaming LLMs to actively decide when to generate output based on endogenous cues, using mask-based streaming modeling and synchronized privileged self-distillation, reducing latency without external annotations.

0 favorites 0 likes
#latency-reduction

Skim: Speculative Execution for Fast and Efficient Web Agents

arXiv cs.AI · 2026-05-19 Cached

Accio is a speculative execution framework that reduces cost and latency for web agents by leveraging offline site-structure profiling and online selection of fast paths, achieving a 1.9x reduction in per-task cost and 33.4% latency reduction while maintaining accuracy.

0 favorites 0 likes
#latency-reduction

LiteFrame: Efficient Vision Encoders Unlock Frame Scaling in Video LLMs

Hugging Face Daily Papers · 2026-05-17 Cached

LiteFrame proposes a lightweight video encoder with Compressed Token Distillation training that reduces latency and enables processing 8x more frames for long-form video understanding in Video LLMs, improving accuracy while reducing compute.

0 favorites 0 likes
#latency-reduction

LatentRAG: Latent Reasoning and Retrieval for Efficient Agentic RAG

arXiv cs.CL · 2026-05-08 Cached

LatentRAG is a novel framework that shifts reasoning and retrieval for agentic RAG into continuous latent space, reducing inference latency by approximately 90% while maintaining performance comparable to explicit methods.

0 favorites 0 likes
#latency-reduction

Speeding up agentic workflows with WebSockets in the Responses API

OpenAI Blog · 2026-04-22 Cached

OpenAI details how WebSockets and API optimizations reduced latency by 40% for agentic workflows, enabling GPT-5.3-Codex-Spark to reach near 1,000 tokens per second.

0 favorites 0 likes
#latency-reduction

Prompt Caching in the API

OpenAI Blog · 2024-10-01 Cached

OpenAI introduces Prompt Caching, an automatic feature that reduces API costs by 50% and improves latency by reusing recently cached input tokens on GPT-4o, GPT-4o mini, o1-preview, and o1-mini models. The feature automatically applies to prompts longer than 1,024 tokens without requiring developer integration changes.

0 favorites 0 likes
← Back to home

Submit Feedback