OpenAI published the first independent benchmark results for Jalapeño, its custom inference chip built with Broadcom. The chip outperformed current state-of-the-art systems, including Nvidia's Blackwell, on SemiAnalysis' InferenceX benchmark, according to OpenAI's blog and TechCrunch.

Across three models tested, Jalapeño delivered 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the comparison systems, TechCrunch reported. OpenAI said the chip's full-stack design specifically targets the prefill and communication phases of inference, the steps where data movement between chips typically creates bottlenecks.

"The bottom line is that the results show a very, very significant performance advance over state of the art," Richard Ho, OpenAI's head of hardware, told reporters, according to TechCrunch. "Jalapeño can serve more AI work per unit of power, while also returning responses more quickly. It's very efficient to serve a lot of customers, but it can also be very low latency."

Ho cautioned that competing hardware will keep improving before Jalapeño ships at scale. OpenAI plans small-volume deployment by the end of 2026, with broader rollout in 2027, according to the company's blog.

Inference cost and speed set the ceiling on what founders can build and how cheaply they can serve it. If OpenAI's own hardware materially undercuts Nvidia pricing for the workloads it runs, that pressure reaches every startup buying inference by the token, not just OpenAI's balance sheet.