Tag
Unconventional AI's first analog AI chip, using coupled resonators, generates an image with less than 900 nano joules, claiming a 1000x improvement in power efficiency for AI hardware.
NVIDIA reports up to 30× higher agentic throughput per megawatt with Vera Rubin compared to GB300, but emphasizes that production benchmarks should include additional metrics beyond tokens per watt.
Parasma Founder & CEO @vytalow argues that using human brain cells for compute could be far more energy-efficient than silicon chips, citing power efficiency, sample efficiency, and continual learning advantages.
This article compares the energy cost of running AI models on Strix Halo (max $0.48/day) versus Nvidia A6000, highlighting power efficiency and versatility.
Etched, an AI inference hardware startup, exited stealth after raising $800M and securing over $1B in customer contracts. Their first racks ship this summer, claiming state-of-the-art throughput, latency, and power efficiency.
Unconventional AI, led by former Databricks AI chief Naveen Rao, claims their oscillator-based computer architecture can reduce AI inference power consumption by up to 1,000x, demonstrated with their first image-generation model Un0.
Running Gemma 12B model on a Google Pixel 10 Pro using llama.cpp achieves 6.5 tokens per second prompt processing and 1.3 tokens per second generation with under 10 watts power consumption, demonstrating efficient on-device AI inference.
Nvidia unveiled its photonics co-packaged optics switch with Lambda, aiming to reduce power consumption and failure points in large GPU clusters for AI workloads.
A deep benchmark of 8 tiny LLMs (135M to 1B parameters) on a $250 Jetson Orin Nano Super across four power modes finds 25W to be Pareto-optimal, with SmolLM2-135M achieving 165.1 tok/s and best efficiency.
A user benchmarks RTX 5090 and RTX 6000 PRO GPUs for AI diffusion tasks, comparing performance at different power limits and showing tradeoffs between speed and power consumption.
A user shares power limit testing on a 4x RTX 3090 setup running Qwen3.6-27B with vLLM, finding 220W as the sweet spot for peak efficiency with minimal throughput loss.
A user benchmarks the Nvidia 5090 RTX GPU for LLM inference using llama.cpp, measuring prompt processing and token generation at various power levels, finding that prompt processing is more sensitive to power limits than token generation, and noting differences from the 4090 RTX.
User benchmarks dual Asus GX10 (DGX Spark) running MiniMax-M2.7-AWQ-4bit, achieving 30–40 tokens/s while drawing only ~100 W each, replacing noisy multi-GPU rigs.