agentic-workloads

Tag

Cards List
#agentic-workloads

@ArizePhoenix: Pricing is $10/$50 per million input/output tokens, with cache reads cut 75% to $0.25/M. Anthropic claims ~25% lower ty…

X AI KOLs Following · 6d ago

Anthropic announces pricing changes for its AI model API, with rates of $10 and $50 per million input and output tokens respectively, and a 75% reduction in cache read costs. The company claims these changes lead to about 25% lower typical costs and up to 45% savings for agentic workloads.

0 favorites 0 likes
#agentic-workloads

@SemiAnalysis_: AMD team is grinding hard Although, for most SLOs/open models, AMD is behind NVIDIA, we think that AMD's software will …

X AI KOLs Timeline · 2026-08-30 Cached

AMD is working hard to improve its software for AI workloads, especially agentic tasks, with potential for better performance per dollar than NVIDIA, but software and leadership issues currently hinder progress.

0 favorites 0 likes
#agentic-workloads

Watch out for cache read costs

Lobsters Hottest · 2026-08-10 Cached

A blog post explaining that cache read costs dominate LLM inference spending for agentic workloads, with cumulative costs growing quadratically as context is re-read each turn, and advice to reduce tool call count to cut costs.

0 favorites 0 likes
#agentic-workloads

@omarsar0: TokTier makes tokenization stateful for agentic serving. Across 153,951 real agent calls with a 94.1% prompt-cache hit …

X AI KOLs Following · 2026-08-03 Cached

TokTier is a stateful tokenization service that cuts tokenization overhead in agentic LLM serving by reusing stable token boundaries and running exact CPU/GPU tokenization, reducing time-to-first-token by 16–34% under vLLM.

0 favorites 0 likes
#agentic-workloads

Consistently unifying work from thousands of agents

Reddit r/AI_Agents · 2026-07-28

Describes a method for unifying outputs from thousands of agents in parallel forecasting tasks, achieving consistency and cost efficiency through post-processing and homogeneous task design.

0 favorites 0 likes
#agentic-workloads

@KVCache_AI: Mooncake now supports SSD Offloading for KV Cache. As agentic workloads become the norm, KV cache lifetimes are getting…

X AI KOLs Timeline · 2026-07-15 Cached

Mooncake announces support for SSD offloading of KV cache, enabling cost-effective scaling of KV cache capacity beyond DRAM for long-lived agentic workloads, with analysis showing bimodal reuse patterns that make tiered storage efficient.

0 favorites 0 likes
#agentic-workloads

I benchmarked 13 models at 65K-128K context to find out what actually matters for agentic workloads

Reddit r/LocalLLaMA · 2026-07-05

An extensive benchmark of 13 local LLMs at 65K-128K context shows that prefill speed dominates agentic workload performance (94-99% of wall-clock time), rendering tg128 misleading, and that KV head count is the key architectural factor over parameter count or MoE/dense design.

0 favorites 0 likes
#agentic-workloads

DualPath: Breaking the Storage Bandwidth Bottleneck in Agentic LLM Inference

Reddit r/singularity · 2026-06-24 Cached

DualPath is a system that breaks the storage bandwidth bottleneck in agentic LLM inference by introducing a dual-path KV-cache loading mechanism, improving throughput by up to 1.87x offline and 1.96x online.

0 favorites 0 likes
#agentic-workloads

@m_sirovatka: KV Cache re-use is the most important thing for agentic rollouts. We've integrated Mooncake Store into prime-rl with vL…

X AI KOLs Following · 2026-06-02 Cached

vLLM integrates Mooncake Store for distributed KV cache reuse, enabling cross-node prefix caching to efficiently serve agentic workloads with high token reuse.

0 favorites 0 likes
#agentic-workloads

@zhyncs42: Qwen inference team is super great — they achieved 540 TPS on TokenSpeed for agentic workloads Looking forward to them …

X AI KOLs Timeline · 2026-05-24 Cached

Qwen inference team announced TokenSpeed, a high-performance LLM inference engine for agentic workloads, achieving 540 TPS, with open-source preview available.

0 favorites 0 likes
#agentic-workloads

TokenSpeed: A Speed-of-Light LLM Inference Engine for Agentic Workloads (5 minute read)

TLDR AI · 2026-05-07 Cached

Lightseek releases TokenSpeed, a high-performance LLM inference engine optimized for agentic workloads, featuring compiler-backed parallelism and advanced kernel optimizations that have been adopted by vLLM.

0 favorites 0 likes
← Back to home

Submit Feedback