OpenAI published the first performance results for Jalapeño, the inference processor it built with Broadcom, reporting 1.5 to 1.9 times more AI work per watt at peak throughput and 1.7 to 3.6 times lower end-to-end latency than the systems it was tested against. On workloads demanding high interactivity the gap widened to between 2.1 and 4.1 times. The comparison was against Nvidia’s Blackwell GPUs.
The tests ran on three models OpenAI did not build: GPT-OSS 120B, DeepSeek R1 at 670 billion parameters, and Moonshot’s Kimi K2.5 at a trillion. On GPT-OSS 120B, Jalapeño delivered roughly 1.9 times the peak mixed tokens per second per kilowatt and cut latency from 1.80 seconds to 1.03. DeepSeek R1 ran at about 1.7 times the efficiency, with latency dropping from 5.99 seconds to 1.65. Kimi K2.5 came in at roughly 1.5 times, from 5.31 seconds to 1.56.
The choice of models addresses the obvious objection. Custom silicon built by a model developer invites the assumption that it is tuned to that developer’s architecture, in the way a console is built around its own games. SemiAnalysis, the semiconductor research firm whose InferenceX benchmark suite OpenAI used, sent engineers to run the workloads in person and pushed back on that reading directly, describing Jalapeño as a generalized inference chip. OpenAI’s engineers demonstrated the point by porting Doom to it using Codex prompts.
SemiAnalysis found Jalapeño ahead of Blackwell on performance per watt in nearly every scenario it tested, and doing it without multi-token prediction, a technique the comparison systems were using. At low concurrency on DeepSeek R1 the chip passed 700 tokens per second per user on single-token prediction alone, with no speculative decoding.
The power figures are the part that matters for a company buying gigawatts. Jalapeño carries a 700-watt rating but drew at or below 550 watts sustained on the workloads tested. OpenAI has committed to $750 billion in compute spending through 2030, and inference is the recurring cost inside that number rather than the one-time training bill.
Broadcom built the chip on a nine-month development cycle, unusually fast for a reticle-sized ASIC. Deployment is planned for the end of 2026, which means these are pre-deployment numbers from the company that commissioned the silicon, with a third party present rather than a third party auditing independently.
The two announcements this week point in opposite directions. Nvidia spent $7 billion on Poolside on Monday to move into open-weight models, competing for the developers its own customers serve. OpenAI has spent nine months moving into the silicon Nvidia sells it. Neither has left the other’s market, and both now compete inside it.
Sources: Techerati, OfficeChai
–
By the Control Plane Editorial Team