OpenAI's Jalapeño Chip Beats Nvidia On Inference Speed And Power

OpenAI's Jalapeño Chip Beats Nvidia On Inference Speed And Power

0:00 / 0:51
News

OpenAI's Jalapeño Chip Beats Nvidia On Inference Speed And Power

calendar_today Date:
schedule Duration: 0:51
database
Summary Report

OpenAI's Jalapeño chip delivers up to 3.6x lower latency and 4.1x higher performance on interactive workloads than comparison systems, tested across GPT, DeepSeek and Kimi models.

  • 01. Achieves both higher throughput and lower latency in one architecture, rather than trading one for the other
  • 02. OpenAI plans to deploy it in its own infrastructure by year end, with a second and third generation already in development
OpenAI has shared the first real test results for Jalapeño, its custom inference chip, delivering up to 1.9 times more AI work per watt and up to 3.6 times lower latency than comparison systems. It was tested across GPT, DeepSeek and Kimi models, and OpenAI plans to deploy it in its own infrastructure by the end of this year.