1 week ago
OpenAI Reports Jalapeño Chip Gains in Speed and Efficiency
OpenAI made a special computer chip called Jalapeño for running trained AI models.
Running a trained model to produce an answer is called inference.
OpenAI says Jalapeño can do more AI work while using less electricity.
It also says the chip can return answers faster.
The company tested it with three large AI models using a public benchmark.
OpenAI reported better results than the systems used for comparison.
These results came from OpenAI's own testing, so the testing methods and comparison systems are important context.
The chip was built together with memory, networking and software so less information has to move around.
OpenAI says it will keep using NVIDIA and other companies' chips while beginning internal Jalapeño deployment by the end of 2026.
OpenAI says its Jalapeño custom chip delivers more AI work per watt and lower response times.
The inference chip was tested on GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T using the InferenceX benchmark.
OpenAI reported 1.5 to 1.9 times more work per watt, 1.7 to 3.6 times lower latency and 2.1 to 4.1 times higher performance for interactive workloads.
The company designed the chip, memory, networking, software and rack-scale system together to reduce data movement and communication delays.
OpenAI plans internal deployment by the end of 2026, while Gen 2 and Gen 3 are in development and broader deployment is referenced for 2027.
- Who
- OpenAI developed Jalapeño, with CEO Sam Altman announcing the chip; one article says it was developed in collaboration with Broadcom.
- What
- OpenAI reported benchmark results for its first custom AI inference chip.
- Where
- The announcement was reported from California, and the chip was presented at the Hot Chips conference.
- When
- The results were announced on Tuesday local time; OpenAI says internal deployment is planned by the end of 2026, while one headline refers to wider deployment in 2027.
- Why
- OpenAI says Jalapeño is intended to deliver faster responses, greater power efficiency and more computing capacity for AI models and agents.
OpenAI's Performance Claims
Evaluation Context
Speed and efficiency
OpenAI's Performance Claims
OpenAI says Jalapeño combines higher throughput and lower latency, delivering up to 1.9 times more work per watt and up to 3.6 times lower end-to-end latency across the tested models.
Evaluation Context
The articles report no independent verification; the figures come from OpenAI's testing and depend on its methodology, operating conditions and comparison systems.
Benchmark evidence
OpenAI's Performance Claims
OpenAI says it used the public InferenceX benchmark and tested three models, including models developed outside OpenAI, across both high-throughput and interactive workloads.
Evaluation Context
The articles do not provide the underlying benchmark data or an external assessment of whether the comparison systems were equivalent.
Role alongside commercial chips
OpenAI's Performance Claims
OpenAI presents Jalapeño as a way to tailor the hardware and broader infrastructure stack to its models and AI-agent workloads.
Evaluation Context
OpenAI says it will continue deploying NVIDIA and other partners' accelerators for both training and inference, so Jalapeño is not described as an immediate replacement.
Key facts
- Purpose
- Serving trained AI models during inference rather than training them
- Tested models
- GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T
- Benchmark
- InferenceX, a public benchmark from SemiAnalysis measuring the process of serving an AI request
- Reported results
- 1.5–1.9 times more work per watt, 1.7–3.6 times lower end-to-end latency and 2.1–4.1 times higher interactive performance
- Power
- Rated at 700 watts; measured sustained power was at or below 550 watts on tested workloads
- Development
- OpenAI says AI helped take the chip from initial design to tapeout in nine months
- Roadmap
- Jalapeño is Gen 1; Gen 2 is deep in development and Gen 3 is taking shape
Quotes
Sam Altman
OpenAI chief executive officer
“Jalapeño delivers both higher throughput and lower latency with one architecture. For customers, that can mean faster responses, more responsive agents, and more reliable access as demand grows,”
businesstoday.in
“Jalapeno delivers both higher throughput and lower latency with one architecture, where existing hardware systems often have to make a tradeoff between the two.”
freepressjournal.in
republicworld.com
Sources
OpenAI’s Jalapeño chip claims major AI speed boost: What to know
'We Made A Chip And It Is Fast': OpenAI Unveils Jalapeno AI Inference Chip With Major Gains In Speed, Latency And Power Efficiency
'We Made a Chip, and It Is Fast': OpenAI Unveils Jalapeno Custom AI Inference Chip





