Tag
A practical guide covering common inference bottlenecks — including runtime selection (Ollama, llama.cpp, vLLM, SGLang), hardware-specific compilation flags, quantization choices (IQ vs Q_K vs NL), and model selection by task — with best practices for optimizing LLM inference within hardware constraints.
AgBench introduces a benchmark suite and open artifacts for evaluating agentic AI on personal devices, comparing local, hybrid, and cloud execution across task success, latency, API cost, and data exposure using over 162 million data points. The study finds no single architecture dominates, with local execution reducing cloud cost and privacy risk but generally lowering task success.
The article speculates on consumer interest in a hypothetical AI chip from Taalas that could run large language models like Qwen3.8-27B at high speeds, comparing it to past media formats.
This paper explores how prompt properties like cognitive load and phrasing pattern influence energy usage in on-device LLM inference, showing that cognitive load affects energy per token while phrasing impacts token usage, highlighting the need for model-aware prompt design for energy efficiency.
The author demonstrates how to run a 182B Qwen model on a 4080 GPU with 64GB DDR5 by offloading ngrams to SSD, achieving faster and more intelligent performance than a 27B model, making large model inference accessible on consumer hardware.
This article is a detailed guide on large language model deployment, covering key metrics such as latency, throughput, and memory usage, and illustrates how to optimize performance and choose hardware through practical cases.
A tweet sharing information or resources about the deployment of local large language models with GPUs.
This article discusses the trade-offs and considerations for developers deciding between using cloud-based AI services versus running large language models locally.
Fixed three bugs in a qMLX fork for running Qwen3.5-122B on Mac Studio, reducing prefill time from minutes to sub-seconds for long-context inference; open-sourced the fork and benchmark script.
Explored NVIDIA Dynamo, a tool for deploying LLMs across multiple GPU cluster nodes with features like model caching, autoscaling, multinode deployments, and Kubernetes integration.
Lilian Weng's blog post explores the concept of harness engineering as a key component for recursive self-improvement in AI systems, discussing design patterns, workflow automation, and the analogy to operating systems.
A comprehensive guide to running LLMs locally across various hardware and software setups is now available online for free, covering tools like llama.cpp, vLLM, and more.
This article provides a comprehensive overview of the complete technology stack for cloud deployment of Transformer inference, covering application scenarios, workload definition, models, inference engines, hardware, observability, and performance optimization, along with future trends.
This paper presents a two-stage methodology for end-to-end LLM deployment on spatial NPUs, progressing from human-guided development to an autonomous agent skill system. The system achieves speedups of 2.2x on prefill and 4.0x on decode for a reference model, and autonomously deploys eight additional LLMs on AMD XDNA 2 NPU with minimal human guidance.
A practitioner seeks advice on running AI agents 24/7 without high API costs, asking about local models, cloud GPUs, or hosted APIs, and wants cost-efficient setups balancing reliability and reasoning quality.
A lecture on LLM deployment techniques covering AWQ, vLLM, FlashAttention, quantization, and activation smoothing for efficient serving.
A discussion on the challenges consultants face when clients want to deploy LLMs despite having poor data governance, weighing the risks of fixing data first versus deploying quickly on messy data.
Skymizer announces the HTX301, a PCIe inference card capable of running 700B-parameter LLMs on-premises with high memory and low power consumption.
Anyscale published a technical guide on deploying production-ready AI agents using Ray Serve, MCP, and A2A protocols. The article addresses common infrastructure bottlenecks by proposing a decoupled microservices architecture that enables independent scaling of LLMs, tools, and agents.
This paper introduces geometric stability measures—based on pairwise distance consistency in representations—to predict language model steerability and detect structural drift. Supervised variants achieve near-perfect correlation (ρ=0.89-0.97) with linear steerability across 35-69 embedding models, while unsupervised variants outperform CKA and Procrustes for post-deployment drift detection.