Gadgets & Hardware
·By Seedwire Editorial·

OpenAI reports Jalapeño inference gains, with rollout to follow

OpenAI reports Jalapeño inference gains, with rollout to follow

Illustration, not documentary evidence of the event.

OpenAI presented benchmark results for its Jalapeño inference system at Hot Chips, claiming higher token output per user and higher throughput per kilowatt than an Nvidia Blackwell system. According to techcrunch.com, the results came from testing with SemiAnalysis’ InferenceX benchmark. These were OpenAI-reported results, not independently verified findings established by the supplied report.

The deployment schedule put broad availability beyond the presentation. OpenAI hardware chief Richard Ho estimated a small initial rollout at the end of 2026, followed by a larger deployment in 2027. Developed with Broadcom, Jalapeño was designed to reduce delays in prefill and communication. OpenAI said the system could keep model state, including the KV cache used during response generation, local to reduce data movement.

For teams choosing an inference provider, the useful distinction is between hardware efficiency and the service they can actually buy. Higher throughput per kilowatt could give OpenAI more capacity within a power budget. It does not establish lower customer prices or faster responses for a particular application. The supplied report gives no numerical performance margins or workload configurations, so it cannot support a calculation of potential savings or an application-specific speed comparison.

The more consequential design choice may be OpenAI’s plan to develop models, memory, chips and products together across multiple Jalapeño generations. That suggests the relevant test is how well the integrated system handles the workloads OpenAI intends to serve. A concrete follow-up would be results that identify the model, request lengths and serving load, alongside measurements from deployed services. Those details would help buyers judge whether the reported efficiency advantage translates into the response times and capacity their applications need.

Correction, September 19, 2026: This article was revised to attribute reported claims, remove unsupported statements and clarify the limits of the available evidence.

OpenAI
Jalapeño
inference chips
Broadcom
InferenceX
AI hardware
Seedwire Newsletter

Follow Seedwire by email

Request Seedwire news emails. There is no guaranteed delivery schedule. You can withdraw your request through the privacy contact.

By selecting Subscribe, you request Seedwire news emails. Privacy and withdrawal.