Tag
Animashree Anandkumar of Caltech and Accelerated Understanding will deliver the SC26 keynote on 'The Next Scaling Law after LLMs: AI that Understands the Physical World,' highlighting advances from FourCastNet to AI-driven scientific discovery in weather, fusion, and biology.
This paper explores the impact of different candidate-generation schedules on the energy consumption and performance of large language models during test-time scaling, demonstrating that larger batch sizes reduce energy use and latency.
The first release of the Guix-Science channel provides a dedicated, community-driven scientific software catalog for the Guix package manager, enhancing reproducibility and collaboration in scientific computing.
The tweet recommends Algorithmica's HPC series section on CPU cache, detailing concepts like latency, memory access patterns, and theoretical latency to help write faster software.
AMD releases ROCm 10.0, a major update to its open-source GPU compute platform, marking a decade and introducing native agentic AI developer experience with ROCm.AI.
This paper presents a validation-centric AI-assisted GPU porting workflow applied to a legacy 250,000-line weather simulation code, achieving a 5.1× application-level speedup and addressing numerical discrepancies through careful validation.
A look at the HP Z8 Fury workstation's insane 2TB DDR5 configuration costing over $200k just for RAM, and how current Nvidia RTX Pro 6000 price hikes make buying the base system with 4 GPUs a better deal than standalone cards.
Quantinuum and NVIDIA, with a pharmaceutical partner, validated a proof-of-principle Generative Quantum AI (GenQAI) framework that combines HPC, AI, and quantum computing to generate and execute quantum circuits for pharmaceutical R&D.
Google Cloud named primary HPC provider for NOAA's Weather and Climate Operational Supercomputing System, enabling operational numerical weather prediction in the public cloud for faster simulations and life-saving warnings.
The author rants about the clunkiness of using SSH and VPN to access university HPC clusters for computational chemistry, arguing that the friction discourages independent researchers from doing work and advocates for local computations.
A systematic exploration of FP32 matrix multiplication optimization on AMD Zen 3, achieving 85.30 GFLOPS (63.5% of theoretical peak) using AVX2/FMA intrinsics, surpassing naive implementation by 56.5x and matching optimized libraries.
This paper presents a nonparametric Bayesian inverse reinforcement learning approach using a Dirichlet process prior to infer multiple latent reward types from expert demonstrations, implementing a collapsed Gibbs sampler with parallelization via Ray for scalability.
A detailed blog post from the Guix HPC team describing how they identified and debugged a performance regression in the MPI stack on Slingshot interconnects, showcasing Guix's transparency and control.
l is a new runtime for k4, q, and qSQL that provides transparent SIMD, compressed vectors, and automatic parallelism while maintaining full compatibility with existing code. It targets high-performance computing on Wall Street.
Anthropic's Claude Science desktop application focuses on solving the digital grunt work in scientists' daily research, such as data integration and supercomputer scheduling, rather than pursuing a grand-narrative AI scientist. It lowers the barrier to scientific computing through a natural language interface.
A half-day tutorial at ISC High Performance 2026 on using compiler-assisted tools (FPChecker/LLVM) for floating-point error analysis and profiling in C/C++ scientific codes.
Chinese supercomputer LineShine has claimed the top spot in the global supercomputer rankings, marking a significant achievement in high-performance computing.
The LineShine supercomputer in Shenzhen, China, claims the number 1 spot on the TOP500 list with 2.198 Exaflops of sustained FP64 performance, powered by a custom Armv9 CPU with 13 million cores. It also leads the HPCG benchmark, surpassing El Capitan.
China has built the world's fastest supercomputer, LineShine, overtaking the US system El Capitan in the TOP500 ranking. The system uses only CPUs and entirely Chinese hardware/software, demonstrating technological self-sufficiency despite US export restrictions.
China's LineShine supercomputer becomes the world's fastest, displacing the US's El Capitan for the first time since 2017, marking a significant shift in high-performance computing rankings.