OpenAI Unveils Jalapeño Custom Silicon Benchmarks for AI Inference

OpenAI Unveils Jalapeño Custom Silicon Benchmarks for AI Inference

OpenAI has revealed performance data for Jalapeño, its first custom silicon chip designed exclusively for running AI models. Tested on open industry benchmarks against commercial Nvidia hardware, the accelerator delivered up to 1.9 times more output per kilowatt while cutting response delays by more than half. The company plans to begin rolling the silicon into its production data centers before the end of the year.

Testing was conducted on the InferenceX benchmark created by SemiAnalysis to measure full request cycles across 3 massive models: GPT OSS 120B, DeepSeek R1 670B, and Kimi K2.5 1T. On the 1T parameter Kimi model, Jalapeño achieved 1.5 times higher throughput per unit of power and 3.4 times lower latency than Nvidia GB300 systems. For DeepSeek R1, the chip produced 104.3 times more tokens per kilowatt when matched to existing decoding speeds. The hardware package carries a 700 watt rating, but internal measurements showed power draw staying at or below 550 watts during active tests.

Traditional processors force systems to balance raw speed against electrical efficiency. Jalapeño eliminates that compromise. Serving large language models requires 2 distinct stages: heavy computation during the initial prompt reading, followed by memory intensive token generation. By keeping the key value cache local and preventing data from stalling between chips, the architecture avoids the communication delays that slow down multi step AI agents.

Internal artificial intelligence models helped build and program the physical silicon from day 1. That automation allowed the engineering team to reach tapeout in just 9 months. The company also used Codex and GPT Astra to write custom software kernels, bringing 3 unreleased open weight models up to peak performance in 2 months. In specific testing on mixture of experts blocks, code generated by AI ran up to 1.8 times faster than manual code written by human software engineers.

Deploying proprietary silicon gives the company direct control over its operating expenses as user demand grows. Running more requests on less electricity lets the firm expand agentic features without letting server costs spiral out of control. Jalapeño represents the starting point of a broader hardware roadmap, with Gen 2 already deep in engineering and Gen 3 taking shape. Even with its own silicon coming online, OpenAI confirmed it will continue to purchase and deploy chips from Nvidia to meet expanding compute demands.

About the author

Majid T.
Majid T.
Owner of Technetbook | 10+ Years of Expertise in Technology | Seasoned Writer, Designer, and Programmer | Specialist in In-Depth Tech Reviews and Industry Insights | Passionate about Driving Innovation and Educating the Tech Community Technetbook

Join the conversation

Newsletter Subscription