Tag
这篇论文对 NVIDIA Hopper GPU 架构进行了多层级微基准测试分析,评估了 L2 分区缓存、第四代张量核心(FP8)、DPX 指令、分布式共享内存(DSM)和张量内存加速器(TMA)等新特性的性能表现,结果显示 TMA 异步编程可实现 1.5 倍矩阵乘法加速,FP8 性能接近 FP16 的两倍,DPX 指令可加速生物信息学算法至少 4.75 倍。
NAND-16 is a computer built using 277,248 NAND gates, demonstrating the construction of a functional system from fundamental digital logic components.
This paper introduces AutoTuring, an AI agent for design-space exploration of GPU accelerators, and evaluates whether architectural knowledge improves performance by comparing hardware-aware and opaque conditions. The findings indicate that meaning aids optimization, but architectural knowledge and structured critique act as substitutes rather than complements.
PC Anatomy is an open source interactive 3D explorer for computer hardware, allowing users to dissect and inspect PC components from case to GPU core with detailed explanations.
This article benchmarks the performance of subnormal floating-point numbers on Intel, AMD, and ARM processors, revealing that Intel processors experience significant slowdowns with subnormals while AMD and ARM are unaffected.
This article explains the PDP-8 minicomputer's architecture and its use in teaching fundamental computer concepts, featuring a giveaway for a replica model.
This article explains how CHERIoT provides robust isolation without an MMU, focusing on its design and benefits for secure computing systems.
The article explains the history and reasoning behind the naming of the x86 undefined instruction 'ud2', tracing its evolution from unofficial opcodes to Intel's official implementation due to Hyrum's Law.
The tweet recommends Algorithmica's HPC series section on CPU cache, detailing concepts like latency, memory access patterns, and theoretical latency to help write faster software.
A small experiment comparing the RISC-V RVA23 and ARMv9 computer architectures.
A detailed technical explanation of how CPU caches work, covering the principle of locality, cache organization, indexing, and handling writes.
This paper presents Gauntlet, a multi-agent pipeline that uses LLMs to perform deep technical comprehension of computer architecture papers, and shows that its analyses are preferred over human analyses in 15 out of 20 comparisons.
An interview with Turing Award winner David Patterson covering RISC vs CISC history, GPU/TPU comparisons, Moore's Law, and career advice.
A cheatsheet covering memory segmentation concepts, likely useful for students and developers.
A retrospective analysis argues that the 8086 segmented memory architecture was a clever design that could have scaled gracefully, but software developers' insistence on treating memory as a flat space led to its perceived flaws.
This blog post argues for a return to rigorous full-system timing simulation in computer architecture to overcome the 'timing simulation wall' and accurately capture modern system behaviors, advocating for measuring the right execution intervals with statistically sound methods rather than simulating everything in detail.
This paper proposes a model-native computing architecture, envisioning future system design through the lens of computer architecture.
The article introduces core concepts from the book "Systems Performance" regarding latency, throughput, cache hierarchies, etc., and references latency numbers from experts like Jeff Dean, emphasizing the importance of hands-on practice for performance engineering.
This lecture introduces the flexible evolution of GPU architecture as a SIMD (vector/array) processor, discusses data parallelism, memory bank grouping, bank conflicts, serial bottlenecks, and the history of SIMD instructions (such as MMX), emphasizing how GPUs leverage data parallelism and deal with serial bottlenecks.