Tag
OpenAI previews Ultrafast, a new service tier for GPT-5.6 Sol that runs up to 14× faster via Cerebras, generating up to 750 tokens per second in the OpenAI API.
AMD has signed a deal with AI chip startup Cerebras, signaling a strategic partnership in the AI hardware space.
Google Gemma announces that developers can now use the Gemma 4 31B model as the brain for voice AI, enabled by Hugging Face and Cerebras for ultra-fast inference, as part of an open-source cascaded speech-to-speech stack.
The author shares findings from Hermes Mixture-of-Agents experiments, including voter upgrades, GPU topology, and caching economics, showing that local prefix caching can make long agent sessions nearly free and that two independent GPU instances outperform a single partitioned one.
Thom Wolf and Cerebras released a fully open-source realtime voice demo with models and code, showcasing state-of-the-art speech-to-speech capabilities.
A claim that the Gemma-4-31B model running on Cerebras hardware outperforms ChatGPT's voice mode, demonstrated via a Hugging Face Space for real-time voice interaction.
Hugging Face and Cerebras demonstrate a real-time speech-to-speech pipeline combining open-source models (Nvidia's Parakeet, Gemma 4, Qwen3TTS) with Cerebras' fast inference, enabling natural conversational AI and powering robots like Reachy Mini.
Cerebras' deal with OpenAI to supply $20 billion worth of chips has pre-allocated near-term inference capacity, causing their API waitlist to become effectively infinite for other startups seeking fast inference.
Cerebras and Google DeepMind are hosting a 24-hour hackathon on June 28-29, featuring Gemma 4 models with ultra-fast Cerebras inference, with $5000 in prizes.
The 5.6 Sol model is coming to Cerebras hardware in July, offering inference at 750 tokens per second.
Cerebras Systems stock dropped nearly 20% after forecasting narrower gross margins despite better-than-expected Q1 earnings; CEO Andrew Feldman said the margin outlook was misunderstood due to equipment rental costs.
A summary of the LatePost interview, reviewing Baidu US R&D's early AI布局, including investing in Cerebras, nearly investing in OpenAI and Anthropic, and the flow of talent from Baidu to these companies.
A model labeled 'GPT 5.5' has appeared on Cerebras via OpenRouter statistics, suggesting a potential secret release or testing phase of a new GPT iteration.
In a tweet, Sarah Hooker argues that GPUs are ill-suited for the long-tail distribution of real-world data, suggesting a need for alternative AI hardware.
The article recounts Baidu Research US's investment in Cerebras, a wafer-scale chip company, a decade ago. It analyzes the shift in the AI chip market from training to inference and the importance of non-consensus investments.
Cerebras stock dropped over 31% within 17 days of its IPO at $311, with criticism about chip limitations and misleading claims.
The article argues that Cerebras chips are optimized for LLM inference and training, not general AI workloads, and cautions against overhyping their ability to challenge NVIDIA across all AI domains.
Cerebras co-founder explains the fundamental difference between WSE (Wafer Scale Engine) and NVIDIA GPU: GPU is designed for graphics, runs AI by stacking cores and NVLink interconnect, while WSE makes the entire wafer into a single chip, with on-chip interconnect bandwidth and memory bandwidth far exceeding GPU clusters, greatly leading in inference speed.
AI infrastructure startups Modal, Cerebras, Exa, and TurboPuffer have shown outstanding performance in the past week.
The co-founder of Cerebras explains how their Wafer-Scale Engine (WSE) simplifies design compared to traditional NVIDIA GPUs.