cpu-gpu

Tag

Cards List
#cpu-gpu

Architectural Implications of Agentic AI Workflows

arXiv cs.AI · 2026-08-06 Cached

This paper presents the first architectural characterization of agentic AI workflows, revealing fragmented, heterogeneous execution patterns that mismatch conventional server designs, and introduces a prototype server called Agora to improve CPU/GPU utilization and throughput.

0 favorites 0 likes
#cpu-gpu

@VikParuchuri: Marker 2 is out now - up to 5x faster and more accurate than mineru, docling, and liteparse with similar configs. Conve…

X AI KOLs Following · 2026-07-21 Cached

Marker 2 is released, offering up to 5x faster and more accurate conversion of PDFs, images, and DOCX to markdown compared to mineru, docling, and liteparse.

0 favorites 0 likes
#cpu-gpu

@techNmak: The smartest way to run a giant MoE model is not to add more GPUs. It is to stop treating every expert as GPU-worthy. L…

X AI KOLs Timeline · 2026-07-12 Cached

KTransformers is a framework that optimizes inference and fine-tuning of large Mixture-of-Experts models by dynamically placing only active experts on the GPU while keeping the rest in CPU memory, enabling large models like DeepSeek-V3 to run on limited consumer GPU memory.

0 favorites 0 likes
#cpu-gpu

@bastani_behnam: We just published how we unlocked +50% inference capacity on a 27B model — no new GPUs, no new nodes, at a fraction of …

X AI KOLs Following · 2026-04-21 Cached

OpenInfer demonstrates "vertical disaggregation" that boosts Qwen 3.5 27B throughput by ~50% by co-executing quantized layers across a single node’s AMD EPYC CPU and Nvidia L40S GPU with a custom SLA-aware scheduler.

0 favorites 0 likes
← Back to home

Submit Feedback