Tag
Apple's Mac Mini faces continued supply constraints due to high demand for local AI workloads, with shipping times extending up to three months for high-RAM configurations and higher prices at third-party retailers.
SkyPilot is an open-source tool that automatically finds and provisions the cheapest GPU across 18 cloud providers for AI workloads, reducing costs by up to 10x via spot instances and automatic management.
Modal is a serverless cloud platform designed for AI workloads, supporting inference, training, and sandboxes. It enables instant scaling from zero to thousands of GPUs with pure Python code, significantly reducing latency and accelerating time to market.
Tests NVLink on dual RTX 3090s for AI inference and training, finding significant speedups for tensor parallel prompt processing (30%) and FSDP training (3x), but minimal effect on token generation or layer split inference.
AMD has tagged the ROCm 7.14 'TheRock' tech preview, bringing AI training enhancements, performance improvements up to 16% for select AI workloads like Comfy UI, and ongoing Windows support for the open-source GPU compute stack.
A discussion about upgrading from dual RTX 3090s to alternatives like dual A6000s, RTX 5090, or 48GB RTX 4090, likely for AI/ML workloads.
Serverless GPUs may have higher hourly rates but can be more cost-effective overall depending on workload peak-to-average demand. The article on Modal's blog illustrates this with a widget.
Hugging Face Storage is now a first-class backend for SkyPilot, allowing users to mount Hugging Face repos and buckets into jobs on any cloud with zero egress fees, enabling flexible GPU compute across providers.
Reports indicate that Western companies are migrating AI workloads to Chinese models, and this is evolving into a trend at the procurement decision level.
China has surpassed the US with the world's fastest supercomputer, though the machine is not optimized for AI workloads.
LiteParse is a fast, open-source document parser written in Rust that provides high-quality spatial text extraction with bounding boxes, supporting multiple languages and platforms for AI document workloads.
Intel's discontinued Optane persistent memory technology is finding a second life in AI workloads, enabling a user to run a 1 trillion parameter model locally at ~4 tokens/second using cheap second-hand Optane modules. The article highlights Optane's lower latency compared to SSDs, making it suitable for large model inference despite being slower than DRAM.
Intel launched the Crescent Island GPU at Computex 2026, featuring up to 480GB VRAM and based on the Arc Xe 3P architecture, targeting next-generation AI workloads.
Alibaba unveiled the Zhenwu M890 AI chip, designed to handle the memory and communication demands of AI agent workloads, as part of its push for domestic alternatives.
AMD argues that agentic AI requires rethinking infrastructure planning, with a need for dedicated CPU racks for orchestration and control workloads, shifting the CPU:GPU ratio from 1:8 or 1:4 to 1:1 or higher, rather than simply adding more CPUs to GPU-dense servers.