Tag
The article shares a dedicated chapter on inference from the book 'How to scale your model', covering key optimization techniques for transformer models such as KV cache and latency analysis.
This article explains in detail MoE (Mixture of Experts) inference engineering, corrects misconceptions about activated parameters and deployment costs, and delves into technical details such as router selection, runtime grouping, GPU execution, memory management, and expert parallelism.
Google Startups released a new technical guide for building generative media applications using Google DeepMind's models, covering models, tools, architectures, quality assurance, and economic considerations.
A practical guide explaining how to calculate VRAM requirements for LLMs based on parameter count and quantization level, plus additional overhead from KV cache, activations, and batching.
The Perplexity team has published guidelines for the design, iteration, and maintenance of Agent Skills, emphasizing that writing Skills is not traditional coding but rather constructing context for the model. The article proposes a counter-intuitive methodology focused on evaluation-first approaches, progressive loading, and optimizing Agent behavior by handling edge cases (Gotchas).
Comprehensive educational guide explaining Wi-Fi standards from 802.11n (Wi-Fi 4) through 802.11bn (Wi-Fi 8), covering technical details like MIMO, DFS channels, throughput expectations, and practical router recommendations for consumers.