You are currently viewing Jalapeño’s First Results Show Industry-Leading Speed and Efficiency in AI Inference
Jalapeño AI inference chip

Jalapeño’s First Results Show Industry-Leading Speed and Efficiency in AI Inference

openAI says its first custom inference chip, Jalapeño, has demonstrated significant gains in AI performance, delivering higher throughput, lower latency and improved power efficiency across several leading language models.

The company’s testing found that Jalapeño can perform 1.5 to 1.9 times more AI work per watt at peak throughput while achieving 1.7 to 3.6 times lower end-to-end latency than comparison systems. For highly interactive workloads, performance was reported to be 2.1 to 4.1 times higher.

OpenAI evaluated Jalapeño using the public InferenceX benchmark from SemiAnalysis, comparing it with commercially available AI accelerator systems across different operating conditions. Tests covered GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T, demonstrating performance across models developed both inside and outside OpenAI.

On Kimi K2.5 1T, described by OpenAI as the largest public model tested, Jalapeño delivered approximately 1.5 times higher peak performance per watt and 3.4 times lower end-to-end latency than the comparison system. The chip is rated at 700 watts, while measured sustained power remained at or below 550 watts during the tested workloads.

OpenAI attributes the results to a full-stack approach that combines chip design, memory, networking, software and rack-scale systems around modern AI inference workloads. Jalapeño was specifically designed to reduce data movement and communication delays, while keeping model state such as the KV cache closer to the compute resources.

Artificial intelligence also played a role in the chip’s development. OpenAI says AI-assisted engineering helped the team move from initial design to tapeout in nine months. Using Codex with GPT-Astra, engineers also optimized three open-weight models that were not originally part of the production plan.

The company says selected AI-generated implementations for GPT-OSS attention and mixture-of-experts components performed 1.5 to 1.8 times faster than existing implementations written by human experts.

OpenAI plans to begin deploying Jalapeño within its compute infrastructure by the end of 2026. The company describes it as the first generation of a multigenerational hardware roadmap, with Gen 2 already in development and Gen 3 taking shape.

OpenAI says it will continue using accelerators from NVIDIA and other partners for training and inference while expanding its custom silicon infrastructure to improve AI speed, efficiency and capacity.

Full News