machine-learning-systems

Tag

Cards List
#machine-learning-systems

Learning Agent Execution for KV-Cache Management in Agentic Serving

arXiv cs.AI · 2026-08-18 Cached

CacheScout is an agent-aware KV-cache runtime layer for multi-agent LLM serving that learns agent execution semantics online to guide cache eviction and prefetching, improving cache hit rate and reducing latency.

0 favorites 0 likes
#machine-learning-systems

A Contract-Grade Verifier for LLM-Generated GPU Kernels, and a Native Blackwell Backward for the Gated-Linear-Recurrence Family

arXiv cs.LG · 2026-08-14 Cached

This paper introduces a contract-grade verifier of twelve adversarial gates for checking LLM-generated GPU kernels, finding that 39.5% of kernels accepted by standard loose tests are broken. It also presents the first native Blackwell training backward kernel for the GDN (gated-linear-recurrence) family.

0 favorites 0 likes
#machine-learning-systems

Hawk: Harnessing Hardware-Aware Knowledge for High-Performance NPU Kernel Generation

arXiv cs.AI · 2026-07-03 Cached

Hawk is a training-free framework that uses hardware-aware knowledge to improve NPU kernel generation via LLMs, raising generation accuracy from 49.4% to 80.0% and achieving up to 2.2× execution speedup over state-of-the-art baselines.

0 favorites 0 likes
#machine-learning-systems

@JunchenJiang: I was delighted to give three keynotes recently, one at the 3rd HotInfra workshop at ISCA'26 (hosted by Prof. Jian Huan…

X AI KOLs Timeline · 2026-06-30 Cached

Junchen Jiang delivered three keynotes arguing that KV cache is an underappreciated asset for LLM inference, enabling cost savings, latency reduction, and quality improvements, and should be treated as a core data layer in future inference infrastructure.

0 favorites 0 likes
#machine-learning-systems

ExecuTorch -- A Unified PyTorch Solution to Run AI Models On-Device

arXiv cs.LG · 2026-05-12 Cached

This article introduces ExecuTorch, a unified PyTorch-native deployment framework designed to run AI models on diverse edge devices without requiring model conversion or reimplementation.

0 favorites 0 likes
← Back to home

Submit Feedback