OpenAI says its Jalapeño chip can power faster AI responses than the competition

The Verge Products

Summary

OpenAI's Jalapeño AI chip delivers 1.5 to 1.9 times more work per watt and lower latency than Nvidia's chips, with deployment starting in small volumes by end of year.

<figure> <img alt="An image showing OpenAI’s Jalapeno chip" data-caption="" data-portal-copyright="Image: OpenAI" data-has-syndication-rights="1" src="https://platform.theverge.com/wp-content/uploads/sites/2/2026/08/Jalapeno-chip-final.jpeg?quality=90&#038;strip=all&#038;crop=0,0,100,100" /> <figcaption> </figcaption> </figure> <p class="wp-block-paragraph">OpenAI says its new AI chip, Jalape&ntilde;o, completes tasks more efficiently and returns responses faster than other AI systems, according to <a href="https://openai.com/index/jalapeno-first-results/">a blog post published on Tuesday</a>. During a briefing with reporters, OpenAI hardware vice president Richard Ho said Jalape&ntilde;o offers the "best of both worlds" with lower latency and higher throughput, as AI systems typically "have to make a trade-off between the two."</p> <p class="wp-block-paragraph"><a href="https://www.theverge.com/ai-artificial-intelligence/955939/openai-reveals-its-first-ai-processor-jalapeno">First introduced in June</a>, Jalape&ntilde;o is an Application-Specific Integrated Circuit (ASIC) made in partnership with Broadcom. It's designed for AI inference - the process of running a trained AI model to complete a task or deploy an agent.</p> <img src="https://platform.theverge.com/wp-content/uploads/sites/2/2026/08/tbt-jalapeno.png?quality=90&amp;strip=all&amp;crop=0,0,100,100" alt="" title="" data-has-syndication-rights="1" data-caption="&lt;em&gt;This chart measures Jalapeño&apos;s time between tokens (TBT) - or the&lt;/em&gt; &lt;em&gt;time it takes to deliver a response.&lt;/em&gt; | Image: OpenAI" data-portal-copyright="Image: OpenAI"> <p class="wp-block-paragraph">To mea …</p> <p><a href="https://www.theverge.com/ai-artificial-intelligence/984290/openai-jalapeno-ai-chip-benchmarks">Read the full story at The Verge.</a></p>
Original Article
View Cached Full Text

Cached at: 08/25/26, 04:59 PM

# OpenAI says its Jalapeño chip can power faster AI responses than the competition Source: [https://www.theverge.com/ai-artificial-intelligence/984290/openai-jalapeno-ai-chip-benchmarks](https://www.theverge.com/ai-artificial-intelligence/984290/openai-jalapeno-ai-chip-benchmarks) [![Emma Roth](https://platform.theverge.com/wp-content/uploads/sites/2/chorus/author_profile_images/195810/EMMA_ROTH.0.jpg?quality=90&strip=all&crop=0%2C0%2C100%2C100&w=96)](https://www.theverge.com/authors/emma-roth) Emma Roth is a news writer who covers the streaming wars, consumer tech, crypto, social media, and much more\. Previously, she was a writer and editor at MUO\. OpenAI says its new AI chip, Jalapeño, completes tasks more efficiently and returns responses faster than other AI systems, according to[a blog post published on Tuesday](https://openai.com/index/jalapeno-first-results/)\. During a briefing with reporters, OpenAI hardware vice president Richard Ho said Jalapeño offers the “best of both worlds” with lower latency and higher throughput, as AI systems typically “have to make a trade\-off between the two\.” [First introduced in June](https://www.theverge.com/ai-artificial-intelligence/955939/openai-reveals-its-first-ai-processor-jalapeno), Jalapeño is an Application\-Specific Integrated Circuit \(ASIC\) made in partnership with Broadcom\. It’s designed for AI inference — the process of running a trained AI model to complete a task or deploy an agent\. [![This chart measures Jalapeño’s time between tokens (TBT) — or the time it takes to deliver a response.](https://platform.theverge.com/wp-content/uploads/sites/2/2026/08/tbt-jalapeno.png?quality=90&strip=all&crop=0%2C0%2C100%2C100&w=2400)](https://platform.theverge.com/wp-content/uploads/sites/2/2026/08/tbt-jalapeno.png?quality=90&strip=all&crop=0,0,100,100) To measure Jalapeño’s performance, OpenAI[used InferenceX](https://semianalysis.com/about/), a benchmarking platform that shows how well AI systems handle inference\. The test compared Jalapeño’s performance against the best results recorded at the time, which were with[Nvidia’s GB200](https://www.theverge.com/2024/3/18/24105157/nvidia-blackwell-gpu-b200-ai)or[GB300](https://www.theverge.com/news/631835/nvidia-blackwell-ultra-ai-chip-gb300)superchips\. OpenAI says Jalapeño delivered 1\.5 to 1\.9 times more AI work per watt across GPT\-OSS 120B, DeepSeek R1, and Kimi K2\.5 1T than the comparison systems, while offering 1\.7 to 3\.6 times lower end\-to\-end latency across the three models\. That means the chip can provide users with “faster responses, more responsive agents, and more reliable access as the demand grows,” according to Ho\. [![Jalapeño can perform more work while using less energy, according to OpenAI.](https://platform.theverge.com/wp-content/uploads/sites/2/2026/08/jalapeno-throughput.png?quality=90&strip=all&crop=0%2C0%2C100%2C100&w=2400)](https://platform.theverge.com/wp-content/uploads/sites/2/2026/08/jalapeno-throughput.png?quality=90&strip=all&crop=0,0,100,100) OpenAI plans to deploy Jalapeño in “small volumes” by the end of this year, but will begin to “ramp the volume up” into 2027, Ho added\. The company doesn’t say how many chips it plans to deploy next year, however\. Even with these performance improvements, Ho said OpenAI doesn’t expect to replace its entire chip lineup with Jalapeño, saying its overall compute strategy includes “very good partners,” like Nvidia\. OpenAI will continue developing the second and third generations of the new chip\. **Follow topics and authors**from this story to see more like this in your personalized homepage feed and to receive email updates\. - Emma Roth

Similar Articles

OpenAI reveals its first AI processor: Jalapeño

The Verge

OpenAI has announced its first custom AI inference chip, Jalapeño, developed in partnership with Broadcom to reduce reliance on Nvidia GPUs, with deployment expected by the end of 2026.

OpenAI and Broadcom unveil LLM-optimized inference chip

OpenAI Blog

OpenAI and Broadcom unveiled Jalapeño, a custom LLM-optimized inference chip that promises substantially better performance per watt than current state-of-the-art, designed from the ground up for current and future AI models.