1 week ago

OpenAI Reports Jalapeño Chip Gains in Speed and Efficiency

OpenAI Reports Jalapeño Chip Gains in Speed and Efficiency
'We Made A Chip And It Is Fast': OpenAI Unveils Jalapeno AI Inference Chip With Major Gains In Speed, Latency And Power Efficiency · freepressjournal.in

OpenAI made a special computer chip called Jalapeño for running trained AI models.

Running a trained model to produce an answer is called inference.

OpenAI says Jalapeño can do more AI work while using less electricity.

It also says the chip can return answers faster.

The company tested it with three large AI models using a public benchmark.

OpenAI reported better results than the systems used for comparison.

These results came from OpenAI's own testing, so the testing methods and comparison systems are important context.

The chip was built together with memory, networking and software so less information has to move around.

OpenAI says it will keep using NVIDIA and other companies' chips while beginning internal Jalapeño deployment by the end of 2026.

Key facts

Purpose
Serving trained AI models during inference rather than training them
Tested models
GPT-OSS 120B, DeepSeek R1 670B and Kimi K2.5 1T
Benchmark
InferenceX, a public benchmark from SemiAnalysis measuring the process of serving an AI request
Reported results
1.5–1.9 times more work per watt, 1.7–3.6 times lower end-to-end latency and 2.1–4.1 times higher interactive performance
Power
Rated at 700 watts; measured sustained power was at or below 550 watts on tested workloads
Development
OpenAI says AI helped take the chip from initial design to tapeout in nine months
Roadmap
Jalapeño is Gen 1; Gen 2 is deep in development and Gen 3 is taking shape

Quotes

Sam Altman

OpenAI chief executive officer

“Jalapeño delivers both higher throughput and lower latency with one architecture. For customers, that can mean faster responses, more responsive agents, and more reliable access as demand grows,”
businesstoday.in
“Jalapeno delivers both higher throughput and lower latency with one architecture, where existing hardware systems often have to make a tradeoff between the two.”
freepressjournal.in republicworld.com

Sources

Related news