5.6 Sol is coming to Cerebras at 750 tokens per second in July

Reddit r/singularity Models

Summary

The 5.6 Sol model is coming to Cerebras hardware in July, offering inference at 750 tokens per second.

https://preview.redd.it/8nbr61qjzn9h1.png?width=1853&format=png&auto=webp&s=a223073294a2498e7557061f8b3fc822eb677f96 Absolutely insane
Original Article

Similar Articles

AMD and Cerebras Launch AI Inference Solution (10 minute read)

TLDR AI

AMD and Cerebras announced a joint AI inference solution combining AMD Helios rackscale solutions with Cerebras Wafer-Scale Engine, aiming for ultra-low latency and high throughput. The disaggregated inference workflow is expected to deliver up to 5x higher tokens per second per watt.